95 lines
9.1 KiB
Plaintext
95 lines
9.1 KiB
Plaintext
================================================================================
|
|
RUN 1 — ARC-Easy, ARC-Challenge, Winogrande, PIQA
|
|
================================================================================
|
|
2026-07-26:19:07:26 INFO [_cli.run:388] Selected Tasks: ['arc_easy', 'arc_challenge', 'winogrande', 'piqa']
|
|
2026-07-26:19:07:28 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
|
|
2026-07-26:19:07:28 INFO [evaluator:239] Initializing hf model, with arguments: {'pretrained': 'appvoid/palmer-006'}
|
|
2026-07-26:19:07:49 INFO [models.huggingface:286] Using device 'cuda:0'
|
|
config.json: 1.62kB [00:00, 3.38MB/s]
|
|
tokenizer_config.json: 99.6kB [00:00, 54.3MB/s]
|
|
tokenizer.json: 2.35MB [00:00, 11.4MB/s]
|
|
special_tokens_map.json: 7.42kB [00:00, 19.4MB/s]
|
|
2026-07-26:19:07:52 INFO [models.huggingface:579] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
|
|
model.safetensors: 100%|█████████████████████| 182M/182M [00:08<00:00, 22.1MB/s]
|
|
[transformers] The fast path is not available because one of `(selective_state_update, causal_conv1d_fn, causal_conv1d_update)` is None. Falling back to the naive implementation. To install follow https://github.com/state-spaces/mamba/#installation and https://github.com/Dao-AILab/causal-conv1d
|
|
Loading weights: 100%|██████████████████████| 386/386 [00:00<00:00, 2314.98it/s]
|
|
generation_config.json: 100%|███████████████████| 133/133 [00:00<00:00, 571kB/s]
|
|
README.md: 9.00kB [00:00, 161kB/s]
|
|
ARC-Easy/train-00000-of-00001.parquet: 100%|██| 331k/331k [00:00<00:00, 803kB/s]
|
|
ARC-Easy/test-00000-of-00001.parquet: 100%|██| 346k/346k [00:00<00:00, 1.63MB/s]
|
|
ARC-Easy/validation-00000-of-00001.parqu(…): 100%|█| 86.1k/86.1k [00:00<00:00, 4
|
|
Generating train split: 100%|█████| 2251/2251 [00:00<00:00, 24351.27 examples/s]
|
|
Generating test split: 100%|█████| 2376/2376 [00:00<00:00, 393371.21 examples/s]
|
|
Generating validation split: 100%|█| 570/570 [00:00<00:00, 154851.56 examples/s]
|
|
ARC-Challenge/train-00000-of-00001.parqu(…): 100%|█| 190k/190k [00:00<00:00, 896
|
|
ARC-Challenge/test-00000-of-00001.parque(…): 100%|█| 204k/204k [00:00<00:00, 959
|
|
ARC-Challenge/validation-00000-of-00001.(…): 100%|█| 55.7k/55.7k [00:00<00:00, 2
|
|
Generating train split: 100%|████| 1119/1119 [00:00<00:00, 177512.34 examples/s]
|
|
Generating test split: 100%|█████| 1172/1172 [00:00<00:00, 224575.10 examples/s]
|
|
Generating validation split: 100%|█| 299/299 [00:00<00:00, 101431.32 examples/s]
|
|
README.md: 11.2kB [00:00, 25.5MB/s]
|
|
winogrande_xl/train-00000-of-00001.parqu(…): 100%|█| 2.06M/2.06M [00:00<00:00, 4
|
|
winogrande_xl/test-00000-of-00001.parque(…): 100%|█| 118k/118k [00:00<00:00, 284
|
|
winogrande_xl/validation-00000-of-00001.(…): 100%|█| 85.9k/85.9k [00:00<00:00, 2
|
|
Generating train split: 100%|█| 40398/40398 [00:00<00:00, 1299756.78 examples/s]
|
|
Generating test split: 100%|█████| 1767/1767 [00:00<00:00, 430541.14 examples/s]
|
|
Generating validation split: 100%|█| 1267/1267 [00:00<00:00, 403108.79 examples/
|
|
piqa_train.parquet: 100%|██████████████████| 2.64M/2.64M [00:00<00:00, 3.25MB/s]
|
|
piqa_validation.parquet: 100%|████████████████| 300k/300k [00:00<00:00, 726kB/s]
|
|
piqa_test.parquet: 100%|█████████████████████| 496k/496k [00:00<00:00, 1.20MB/s]
|
|
Generating train split: 100%|██| 16113/16113 [00:00<00:00, 686092.14 examples/s]
|
|
Generating validation split: 100%|█| 1838/1838 [00:00<00:00, 398137.21 examples/
|
|
Generating test split: 100%|█████| 3084/3084 [00:00<00:00, 528054.93 examples/s]
|
|
2026-07-26:19:08:16 INFO [evaluator_utils:446] Selected tasks:
|
|
2026-07-26:19:08:16 INFO [evaluator_utils:480] Task: arc_challenge (arc/arc_challenge.yaml)
|
|
2026-07-26:19:08:16 INFO [evaluator_utils:480] Task: arc_easy (arc/arc_easy.yaml)
|
|
2026-07-26:19:08:16 INFO [evaluator_utils:480] Task: piqa (piqa/piqa.yaml)
|
|
2026-07-26:19:08:16 INFO [evaluator_utils:480] Task: winogrande (winogrande/default.yaml)
|
|
2026-07-26:19:08:16 INFO [api.task:312] Building contexts for arc_easy on rank 0...
|
|
100%|██████████████████████████████████████| 2376/2376 [00:04<00:00, 507.62it/s]
|
|
2026-07-26:19:08:21 INFO [api.task:312] Building contexts for arc_challenge on rank 0...
|
|
100%|██████████████████████████████████████| 1172/1172 [00:02<00:00, 505.21it/s]
|
|
2026-07-26:19:08:24 INFO [api.task:312] Building contexts for winogrande on rank 0...
|
|
100%|████████████████████████████████████| 1267/1267 [00:00<00:00, 73584.28it/s]
|
|
2026-07-26:19:08:24 INFO [api.task:312] Building contexts for piqa on rank 0...
|
|
100%|██████████████████████████████████████| 1838/1838 [00:02<00:00, 691.67it/s]
|
|
2026-07-26:19:08:26 INFO [evaluator:585] Running loglikelihood requests
|
|
Tokenizing inputs: 0%| | 0/20398 [00:00<?, ?it/s][transformers] Ignoring clean_up_tokenization_spaces=True for BPE tokenizer TokenizersBackend. The clean_up_tokenization post-processing step is designed for WordPiece tokenizers and is destructive for BPE (it strips spaces before punctuation). Set clean_up_tokenization_spaces=False to suppress this warning, or set clean_up_tokenization_spaces_for_bpe_even_though_it_will_corrupt_output=True to force cleanup anyway.
|
|
Tokenizing inputs: 100%|████████████████| 20398/20398 [00:06<00:00, 3197.87it/s]
|
|
Running loglikelihood requests: 100%|█████| 20398/20398 [04:21<00:00, 78.09it/s]
|
|
2026-07-26:19:13:03 INFO [loggers.evaluation_tracker:316] Output path not provided, skipping saving results aggregated
|
|
hf ({'pretrained': 'appvoid/palmer-006'}), gen_kwargs: ({}), limit: None, num_fewshot: None, batch_size: 16
|
|
| Tasks |Version|Filter|n-shot| Metric | |Value | |Stderr|
|
|
|-------------|------:|------|-----:|--------|---|-----:|---|-----:|
|
|
|arc_challenge| 1|none | 0|acc |↑ |0.2619|± |0.0128|
|
|
| | |none | 0|acc_norm|↑ |0.2927|± |0.0133|
|
|
|arc_easy | 1|none | 0|acc |↑ |0.5097|± |0.0103|
|
|
| | |none | 0|acc_norm|↑ |0.4482|± |0.0102|
|
|
|piqa | 1|none | 0|acc |↑ |0.6376|± |0.0112|
|
|
| | |none | 0|acc_norm|↑ |0.6360|± |0.0112|
|
|
|winogrande | 1|none | 0|acc |↑ |0.5036|± |0.0141|
|
|
|
|
================================================================================
|
|
RUN 2 — HellaSwag
|
|
================================================================================
|
|
2026-07-26:20:16:54 INFO [_cli.run:388] Selected Tasks: ['hellaswag']
|
|
2026-07-26:20:16:56 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
|
|
2026-07-26:20:16:56 INFO [evaluator:239] Initializing hf model, with arguments: {'pretrained': 'appvoid/palmer-006'}
|
|
2026-07-26:20:17:01 INFO [models.huggingface:286] Using device 'cuda:0'
|
|
2026-07-26:20:17:02 INFO [models.huggingface:579] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
|
|
The fast path is not available because one of `(selective_state_update, causal_conv1d_fn, causal_conv1d_update)` is None. Falling back to the naive implementation. To install follow https://github.com/state-spaces/mamba/#installation and https://github.com/Dao-AILab/causal-conv1d
|
|
Loading weights: 100%|█| 386/386 [00:00<00:00, 1408.36it/s, Materializing param=
|
|
2026-07-26:20:17:11 INFO [evaluator_utils:446] Selected tasks:
|
|
2026-07-26:20:17:11 INFO [evaluator_utils:480] Task: hellaswag (hellaswag/hellaswag.yaml)
|
|
2026-07-26:20:17:11 INFO [api.task:312] Building contexts for hellaswag on rank 0...
|
|
100%|███████████████████████████████████| 10042/10042 [00:07<00:00, 1317.53it/s]
|
|
2026-07-26:20:17:19 INFO [evaluator:585] Running loglikelihood requests
|
|
Tokenizing inputs: 100%|████████████████| 40168/40168 [00:23<00:00, 1726.55it/s]
|
|
Running loglikelihood requests: 100%|█████| 40168/40168 [17:38<00:00, 37.96it/s]
|
|
2026-07-26:20:35:31 INFO [loggers.evaluation_tracker:316] Output path not provided, skipping saving results aggregated
|
|
hf ({'pretrained': 'appvoid/palmer-006'}), gen_kwargs: ({}), limit: None, num_fewshot: None, batch_size: 64
|
|
| Tasks |Version|Filter|n-shot| Metric | |Value | |Stderr|
|
|
|---------|------:|------|-----:|--------|---|-----:|---|-----:|
|
|
|hellaswag| 1|none | 0|acc |↑ |0.3299|± |0.0047|
|
|
| | |none | 0|acc_norm|↑ |0.3841|± |0.0049|
|