================================================================================ RUN 1 — ARC-Easy, ARC-Challenge, Winogrande, PIQA ================================================================================ 2026-07-26:19:07:26 INFO [_cli.run:388] Selected Tasks: ['arc_easy', 'arc_challenge', 'winogrande', 'piqa'] 2026-07-26:19:07:28 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234 2026-07-26:19:07:28 INFO [evaluator:239] Initializing hf model, with arguments: {'pretrained': 'appvoid/palmer-006'} 2026-07-26:19:07:49 INFO [models.huggingface:286] Using device 'cuda:0' config.json: 1.62kB [00:00, 3.38MB/s] tokenizer_config.json: 99.6kB [00:00, 54.3MB/s] tokenizer.json: 2.35MB [00:00, 11.4MB/s] special_tokens_map.json: 7.42kB [00:00, 19.4MB/s] 2026-07-26:19:07:52 INFO [models.huggingface:579] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'} model.safetensors: 100%|█████████████████████| 182M/182M [00:08<00:00, 22.1MB/s] [transformers] The fast path is not available because one of `(selective_state_update, causal_conv1d_fn, causal_conv1d_update)` is None. Falling back to the naive implementation. To install follow https://github.com/state-spaces/mamba/#installation and https://github.com/Dao-AILab/causal-conv1d Loading weights: 100%|██████████████████████| 386/386 [00:00<00:00, 2314.98it/s] generation_config.json: 100%|███████████████████| 133/133 [00:00<00:00, 571kB/s] README.md: 9.00kB [00:00, 161kB/s] ARC-Easy/train-00000-of-00001.parquet: 100%|██| 331k/331k [00:00<00:00, 803kB/s] ARC-Easy/test-00000-of-00001.parquet: 100%|██| 346k/346k [00:00<00:00, 1.63MB/s] ARC-Easy/validation-00000-of-00001.parqu(…): 100%|█| 86.1k/86.1k [00:00<00:00, 4 Generating train split: 100%|█████| 2251/2251 [00:00<00:00, 24351.27 examples/s] Generating test split: 100%|█████| 2376/2376 [00:00<00:00, 393371.21 examples/s] Generating validation split: 100%|█| 570/570 [00:00<00:00, 154851.56 examples/s] ARC-Challenge/train-00000-of-00001.parqu(…): 100%|█| 190k/190k [00:00<00:00, 896 ARC-Challenge/test-00000-of-00001.parque(…): 100%|█| 204k/204k [00:00<00:00, 959 ARC-Challenge/validation-00000-of-00001.(…): 100%|█| 55.7k/55.7k [00:00<00:00, 2 Generating train split: 100%|████| 1119/1119 [00:00<00:00, 177512.34 examples/s] Generating test split: 100%|█████| 1172/1172 [00:00<00:00, 224575.10 examples/s] Generating validation split: 100%|█| 299/299 [00:00<00:00, 101431.32 examples/s] README.md: 11.2kB [00:00, 25.5MB/s] winogrande_xl/train-00000-of-00001.parqu(…): 100%|█| 2.06M/2.06M [00:00<00:00, 4 winogrande_xl/test-00000-of-00001.parque(…): 100%|█| 118k/118k [00:00<00:00, 284 winogrande_xl/validation-00000-of-00001.(…): 100%|█| 85.9k/85.9k [00:00<00:00, 2 Generating train split: 100%|█| 40398/40398 [00:00<00:00, 1299756.78 examples/s] Generating test split: 100%|█████| 1767/1767 [00:00<00:00, 430541.14 examples/s] Generating validation split: 100%|█| 1267/1267 [00:00<00:00, 403108.79 examples/ piqa_train.parquet: 100%|██████████████████| 2.64M/2.64M [00:00<00:00, 3.25MB/s] piqa_validation.parquet: 100%|████████████████| 300k/300k [00:00<00:00, 726kB/s] piqa_test.parquet: 100%|█████████████████████| 496k/496k [00:00<00:00, 1.20MB/s] Generating train split: 100%|██| 16113/16113 [00:00<00:00, 686092.14 examples/s] Generating validation split: 100%|█| 1838/1838 [00:00<00:00, 398137.21 examples/ Generating test split: 100%|█████| 3084/3084 [00:00<00:00, 528054.93 examples/s] 2026-07-26:19:08:16 INFO [evaluator_utils:446] Selected tasks: 2026-07-26:19:08:16 INFO [evaluator_utils:480] Task: arc_challenge (arc/arc_challenge.yaml) 2026-07-26:19:08:16 INFO [evaluator_utils:480] Task: arc_easy (arc/arc_easy.yaml) 2026-07-26:19:08:16 INFO [evaluator_utils:480] Task: piqa (piqa/piqa.yaml) 2026-07-26:19:08:16 INFO [evaluator_utils:480] Task: winogrande (winogrande/default.yaml) 2026-07-26:19:08:16 INFO [api.task:312] Building contexts for arc_easy on rank 0... 100%|██████████████████████████████████████| 2376/2376 [00:04<00:00, 507.62it/s] 2026-07-26:19:08:21 INFO [api.task:312] Building contexts for arc_challenge on rank 0... 100%|██████████████████████████████████████| 1172/1172 [00:02<00:00, 505.21it/s] 2026-07-26:19:08:24 INFO [api.task:312] Building contexts for winogrande on rank 0... 100%|████████████████████████████████████| 1267/1267 [00:00<00:00, 73584.28it/s] 2026-07-26:19:08:24 INFO [api.task:312] Building contexts for piqa on rank 0... 100%|██████████████████████████████████████| 1838/1838 [00:02<00:00, 691.67it/s] 2026-07-26:19:08:26 INFO [evaluator:585] Running loglikelihood requests Tokenizing inputs: 0%| | 0/20398 [00:00