Files
MicroLLM2/mmlu57_token.log
ModelHub XC f9ca541196 初始化项目,由ModelHub XC社区提供模型
Model: MLVXN/MicroLLM2
Source: Original Platform
2026-09-14 14:54:18 +08:00

72 KiB

[MMLU57-TOKEN] start Sun Aug 9 23:10:07 UTC 2026 token=hf_oUMhDnr...
[*] MicroLLM2 MMLU — model: /home/zeus/microllm2/microllm2-checkpoints/final_merged shots=5 limit=20
[*] GPT2-XL 1.5B 1024ctx vocab=50259 (ChatML) — MMLU via direct eval (no harness needed)
[*] lm-eval not installed — using lightweight direct MMLU eval (same logic, no harness)
[*] Install harness for official numbers: pip install lm-eval==0.4.4
[*] Loading tokenizer + model /home/zeus/microllm2/microllm2-checkpoints/final_merged ...
[+] Loaded on cuda:0 dtype=torch.bfloat16 — starting MMLU
MMLU subjects: 0%| | 0/57 [00:00<?, ?it/s]
============================================================
[>] abstract_algebra (shots=5)
MMLU subjects: 2%|▏ | 1/57 [00:03<03:29, 3.73s/it][=] abstract_algebra: 6/20 = 30.0% (running avg 30.0%)
 
============================================================
[>] anatomy (shots=5)
MMLU subjects: 4%|▎ | 2/57 [00:06<02:43, 2.97s/it][=] anatomy: 5/20 = 25.0% (running avg 27.5%)
 
============================================================
[>] astronomy (shots=5)
MMLU subjects: 5%|▌ | 3/57 [00:08<02:23, 2.66s/it][=] astronomy: 7/20 = 35.0% (running avg 30.0%)
 
============================================================
[>] business_ethics (shots=5)
MMLU subjects: 7%|▋ | 4/57 [00:10<02:10, 2.45s/it][=] business_ethics: 6/20 = 30.0% (running avg 30.0%)
 
============================================================
[>] clinical_knowledge (shots=5)
MMLU subjects: 9%|▉ | 5/57 [00:12<02:05, 2.42s/it][=] clinical_knowledge: 9/20 = 45.0% (running avg 33.0%)
 
============================================================
[>] college_biology (shots=5)
MMLU subjects: 11%|█ | 6/57 [00:15<01:59, 2.35s/it][=] college_biology: 9/20 = 45.0% (running avg 35.0%)
 
============================================================
[>] college_chemistry (shots=5)
MMLU subjects: 12%|█▏ | 7/57 [00:17<01:56, 2.33s/it][=] college_chemistry: 3/20 = 15.0% (running avg 32.1%)
 
============================================================
[>] college_computer_science (shots=5)
MMLU subjects: 14%|█▍ | 8/57 [00:19<01:55, 2.37s/it][=] college_computer_science: 9/20 = 45.0% (running avg 33.8%)
 
============================================================
[>] college_mathematics (shots=5)
MMLU subjects: 16%|█▌ | 9/57 [00:22<01:52, 2.34s/it][=] college_mathematics: 7/20 = 35.0% (running avg 33.9%)
 
============================================================
[>] college_medicine (shots=5)
MMLU subjects: 18%|█▊ | 10/57 [00:24<01:47, 2.28s/it][=] college_medicine: 6/20 = 30.0% (running avg 33.5%)
 
============================================================
[>] college_physics (shots=5)
MMLU subjects: 19%|█▉ | 11/57 [00:26<01:42, 2.23s/it][=] college_physics: 3/20 = 15.0% (running avg 31.8%)
 
============================================================
[>] computer_security (shots=5)
MMLU subjects: 21%|██ | 12/57 [00:28<01:37, 2.17s/it][=] computer_security: 6/20 = 30.0% (running avg 31.7%)
 
============================================================
[>] conceptual_physics (shots=5)
MMLU subjects: 23%|██▎ | 13/57 [00:30<01:37, 2.22s/it][=] conceptual_physics: 1/20 = 5.0% (running avg 29.6%)
 
============================================================
[>] econometrics (shots=5)
MMLU subjects: 25%|██▍ | 14/57 [00:32<01:33, 2.18s/it][=] econometrics: 6/20 = 30.0% (running avg 29.6%)
 
============================================================
[>] electrical_engineering (shots=5)
MMLU subjects: 26%|██▋ | 15/57 [00:34<01:30, 2.15s/it][=] electrical_engineering: 4/20 = 20.0% (running avg 29.0%)
 
============================================================
[>] elementary_mathematics (shots=5)
MMLU subjects: 28%|██▊ | 16/57 [00:36<01:26, 2.11s/it][=] elementary_mathematics: 6/20 = 30.0% (running avg 29.1%)
 
============================================================
[>] formal_logic (shots=5)
MMLU subjects: 30%|██▉ | 17/57 [00:39<01:24, 2.12s/it][=] formal_logic: 2/20 = 10.0% (running avg 27.9%)
 
============================================================
[>] global_facts (shots=5)
MMLU subjects: 32%|███▏ | 18/57 [00:40<01:18, 2.00s/it][=] global_facts: 7/20 = 35.0% (running avg 28.3%)
 
============================================================
[>] high_school_biology (shots=5)
MMLU subjects: 33%|███▎ | 19/57 [00:43<01:17, 2.04s/it][=] high_school_biology: 9/20 = 45.0% (running avg 29.2%)
 
============================================================
[>] high_school_chemistry (shots=5)
MMLU subjects: 35%|███▌ | 20/57 [00:44<01:13, 1.99s/it][=] high_school_chemistry: 7/20 = 35.0% (running avg 29.5%)
 
============================================================
[>] high_school_computer_science (shots=5)
MMLU subjects: 37%|███▋ | 21/57 [00:46<01:12, 2.01s/it][=] high_school_computer_science: 7/20 = 35.0% (running avg 29.8%)
 
============================================================
[>] high_school_european_history (shots=5)
MMLU subjects: 39%|███▊ | 22/57 [00:48<01:07, 1.94s/it][=] high_school_european_history: 4/20 = 20.0% (running avg 29.3%)
 
============================================================
[>] high_school_geography (shots=5)
MMLU subjects: 40%|████ | 23/57 [00:50<01:05, 1.91s/it][=] high_school_geography: 5/20 = 25.0% (running avg 29.1%)
 
============================================================
[>] high_school_government_and_politics (shots=5)
MMLU subjects: 42%|████▏ | 24/57 [00:52<01:02, 1.89s/it][=] high_school_government_and_politics: 4/20 = 20.0% (running avg 28.7%)
 
============================================================
[>] high_school_macroeconomics (shots=5)
MMLU subjects: 44%|████▍ | 25/57 [00:54<00:58, 1.84s/it][=] high_school_macroeconomics: 0/20 = 0.0% (running avg 27.6%)
 
============================================================
[>] high_school_mathematics (shots=5)
MMLU subjects: 46%|████▌ | 26/57 [00:55<00:57, 1.85s/it][=] high_school_mathematics: 4/20 = 20.0% (running avg 27.3%)
 
============================================================
[>] high_school_microeconomics (shots=5)
MMLU subjects: 47%|████▋ | 27/57 [00:57<00:55, 1.86s/it][=] high_school_microeconomics: 7/20 = 35.0% (running avg 27.6%)
 
============================================================
[>] high_school_physics (shots=5)
MMLU subjects: 49%|████▉ | 28/57 [00:59<00:55, 1.90s/it][=] high_school_physics: 4/20 = 20.0% (running avg 27.3%)
 
============================================================
[>] high_school_psychology (shots=5)
MMLU subjects: 51%|█████ | 29/57 [01:01<00:52, 1.87s/it][=] high_school_psychology: 5/20 = 25.0% (running avg 27.2%)
 
============================================================
[>] high_school_statistics (shots=5)
MMLU subjects: 53%|█████▎ | 30/57 [01:04<00:54, 2.04s/it][=] high_school_statistics: 8/20 = 40.0% (running avg 27.7%)
 
============================================================
[>] high_school_us_history (shots=5)
MMLU subjects: 54%|█████▍ | 31/57 [01:05<00:49, 1.92s/it][=] high_school_us_history: 4/20 = 20.0% (running avg 27.4%)
 
============================================================
[>] high_school_world_history (shots=5)
MMLU subjects: 56%|█████▌ | 32/57 [01:07<00:46, 1.86s/it][=] high_school_world_history: 7/20 = 35.0% (running avg 27.7%)
 
============================================================
[>] human_aging (shots=5)
MMLU subjects: 58%|█████▊ | 33/57 [01:09<00:45, 1.90s/it][=] human_aging: 8/20 = 40.0% (running avg 28.0%)
 
============================================================
[>] human_sexuality (shots=5)
MMLU subjects: 60%|█████▉ | 34/57 [01:11<00:44, 1.94s/it][=] human_sexuality: 3/20 = 15.0% (running avg 27.6%)
 
============================================================
[>] international_law (shots=5)
MMLU subjects: 61%|██████▏ | 35/57 [01:13<00:43, 1.99s/it][=] international_law: 7/20 = 35.0% (running avg 27.9%)
 
============================================================
[>] jurisprudence (shots=5)
MMLU subjects: 63%|██████▎ | 36/57 [01:15<00:40, 1.92s/it][=] jurisprudence: 8/20 = 40.0% (running avg 28.2%)
 
============================================================
[>] logical_fallacies (shots=5)
MMLU subjects: 65%|██████▍ | 37/57 [01:17<00:37, 1.85s/it][=] logical_fallacies: 7/20 = 35.0% (running avg 28.4%)
 
============================================================
[>] machine_learning (shots=5)
Downloading data: 100%|██████████| 5.25k/5.25k [00:00<00:00, 19.0kB/s]
Generating test split: 100%|██████████| 112/112 [00:00<00:00, 19509.20 examples/s]
Generating validation split: 100%|██████████| 11/11 [00:00<00:00, 3115.91 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1538.18 examples/s]
MMLU subjects: 67%|██████▋ | 38/57 [01:19<00:38, 2.03s/it][=] machine_learning: 10/20 = 50.0% (running avg 28.9%)
 
============================================================
[>] management (shots=5)
Downloading data: 100%|██████████| 14.7k/14.7k [00:00<00:00, 70.8kB/s]
Downloading data: 100%|██████████| 4.50k/4.50k [00:00<00:00, 25.7kB/s]
Downloading data: 100%|██████████| 3.61k/3.61k [00:00<00:00, 19.7kB/s]
Generating test split: 100%|██████████| 103/103 [00:00<00:00, 23526.29 examples/s]
Generating validation split: 100%|██████████| 11/11 [00:00<00:00, 3129.65 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1433.95 examples/s]
MMLU subjects: 68%|██████▊ | 39/57 [01:23<00:45, 2.55s/it][=] management: 4/20 = 20.0% (running avg 28.7%)
 
============================================================
[>] marketing (shots=5)
Downloading data: 100%|██████████| 37.3k/37.3k [00:00<00:00, 214kB/s]
Downloading data: 100%|██████████| 8.21k/8.21k [00:00<00:00, 40.1kB/s]
Downloading data: 100%|██████████| 4.28k/4.28k [00:00<00:00, 13.5kB/s]
Generating test split: 100%|██████████| 234/234 [00:00<00:00, 47706.56 examples/s]
Generating validation split: 100%|██████████| 25/25 [00:00<00:00, 6806.73 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1449.31 examples/s]
MMLU subjects: 70%|███████ | 40/57 [01:26<00:47, 2.77s/it][=] marketing: 7/20 = 35.0% (running avg 28.9%)
 
============================================================
[>] medical_genetics (shots=5)
Downloading data: 100%|██████████| 16.4k/16.4k [00:00<00:00, 84.2kB/s]
Downloading data: 100%|██████████| 5.63k/5.63k [00:00<00:00, 24.9kB/s]
Downloading data: 100%|██████████| 3.77k/3.77k [00:00<00:00, 22.4kB/s]
Generating test split: 100%|██████████| 100/100 [00:00<00:00, 23050.69 examples/s]
Generating validation split: 100%|██████████| 11/11 [00:00<00:00, 3044.36 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1133.23 examples/s]
MMLU subjects: 72%|███████▏ | 41/57 [01:29<00:47, 2.94s/it][=] medical_genetics: 8/20 = 40.0% (running avg 29.1%)
 
============================================================
[>] miscellaneous (shots=5)
Downloading data: 100%|██████████| 98.6k/98.6k [00:00<00:00, 386kB/s]
Downloading data: 100%|██████████| 13.2k/13.2k [00:00<00:00, 72.8kB/s]
Downloading data: 100%|██████████| 3.37k/3.37k [00:00<00:00, 18.4kB/s]
Generating test split: 100%|██████████| 783/783 [00:00<00:00, 137394.47 examples/s]
Generating validation split: 100%|██████████| 86/86 [00:00<00:00, 21921.01 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1380.52 examples/s]
MMLU subjects: 74%|███████▎ | 42/57 [01:33<00:46, 3.09s/it][=] miscellaneous: 6/20 = 30.0% (running avg 29.2%)
 
============================================================
[>] moral_disputes (shots=5)
Downloading data: 100%|██████████| 60.9k/60.9k [00:00<00:00, 298kB/s]
Downloading data: 100%|██████████| 10.7k/10.7k [00:00<00:00, 51.4kB/s]
Downloading data: 100%|██████████| 4.41k/4.41k [00:00<00:00, 26.5kB/s]
Generating test split: 100%|██████████| 346/346 [00:00<00:00, 62921.83 examples/s]
Generating validation split: 100%|██████████| 38/38 [00:00<00:00, 9795.56 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1382.98 examples/s]
MMLU subjects: 75%|███████▌ | 43/57 [01:36<00:43, 3.11s/it][=] moral_disputes: 4/20 = 20.0% (running avg 29.0%)
 
============================================================
[>] moral_scenarios (shots=5)
Downloading data: 100%|██████████| 89.8k/89.8k [00:00<00:00, 499kB/s]
Downloading data: 100%|██████████| 14.9k/14.9k [00:00<00:00, 96.7kB/s]
Downloading data: 100%|██████████| 5.14k/5.14k [00:00<00:00, 27.4kB/s]
Generating test split: 100%|██████████| 895/895 [00:00<00:00, 133752.66 examples/s]
Generating validation split: 100%|██████████| 100/100 [00:00<00:00, 20927.57 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1409.47 examples/s]
MMLU subjects: 77%|███████▋ | 44/57 [01:39<00:40, 3.15s/it][=] moral_scenarios: 3/20 = 15.0% (running avg 28.6%)
 
============================================================
[>] nutrition (shots=5)
Downloading data: 100%|██████████| 55.0k/55.0k [00:00<00:00, 172kB/s]
Downloading data: 100%|██████████| 9.02k/9.02k [00:00<00:00, 58.3kB/s]
Downloading data: 100%|██████████| 4.99k/4.99k [00:00<00:00, 27.9kB/s]
Generating test split: 100%|██████████| 306/306 [00:00<00:00, 60429.26 examples/s]
Generating validation split: 100%|██████████| 33/33 [00:00<00:00, 9086.33 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1449.21 examples/s]
MMLU subjects: 79%|███████▉ | 45/57 [01:43<00:40, 3.34s/it][=] nutrition: 4/20 = 20.0% (running avg 28.4%)
 
============================================================
[>] philosophy (shots=5)
MMLU subjects: 81%|████████ | 46/57 [01:45<00:31, 2.87s/it][=] philosophy: 3/20 = 15.0% (running avg 28.2%)
 
============================================================
[>] prehistory (shots=5)
Downloading data: 100%|██████████| 54.3k/54.3k [00:00<00:00, 334kB/s]
Downloading data: 100%|██████████| 9.89k/9.89k [00:00<00:00, 49.9kB/s]
Downloading data: 100%|██████████| 4.62k/4.62k [00:00<00:00, 21.7kB/s]
Generating test split: 100%|██████████| 324/324 [00:00<00:00, 63532.23 examples/s]
Generating validation split: 100%|██████████| 35/35 [00:00<00:00, 9094.33 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1395.22 examples/s]
MMLU subjects: 82%|████████▏ | 47/57 [01:48<00:30, 3.04s/it][=] prehistory: 5/20 = 25.0% (running avg 28.1%)
 
============================================================
[>] professional_accounting (shots=5)
Downloading data: 100%|██████████| 69.5k/69.5k [00:00<00:00, 354kB/s]
Downloading data: 100%|██████████| 12.9k/12.9k [00:00<00:00, 61.7kB/s]
Downloading data: 100%|██████████| 4.89k/4.89k [00:00<00:00, 27.3kB/s]
Generating test split: 100%|██████████| 282/282 [00:00<00:00, 54398.83 examples/s]
Generating validation split: 100%|██████████| 31/31 [00:00<00:00, 8148.87 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1535.93 examples/s]
MMLU subjects: 84%|████████▍ | 48/57 [01:52<00:28, 3.19s/it][=] professional_accounting: 6/20 = 30.0% (running avg 28.1%)
 
============================================================
[>] professional_law (shots=5)
Downloading data: 100%|██████████| 1.04M/1.04M [00:00<00:00, 3.93MB/s]
Downloading data: 100%|██████████| 116k/116k [00:00<00:00, 600kB/s]/s]
Downloading data: 100%|██████████| 15.1k/15.1k [00:00<00:00, 90.1kB/s]
Generating test split: 100%|██████████| 1534/1534 [00:00<00:00, 72607.74 examples/s]
Generating validation split: 100%|██████████| 170/170 [00:00<00:00, 29159.27 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1370.87 examples/s]
MMLU subjects: 86%|████████▌ | 49/57 [01:55<00:25, 3.17s/it][=] professional_law: 7/20 = 35.0% (running avg 28.3%)
 
============================================================
[>] professional_medicine (shots=5)
Downloading data: 100%|██████████| 125k/125k [00:00<00:00, 663kB/s]
Downloading data: 100%|██████████| 19.9k/19.9k [00:00<00:00, 134kB/s]
Downloading data: 100%|██████████| 8.45k/8.45k [00:00<00:00, 51.7kB/s]
Generating test split: 100%|██████████| 272/272 [00:00<00:00, 37004.56 examples/s]
Generating validation split: 100%|██████████| 31/31 [00:00<00:00, 7737.65 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1409.19 examples/s]
MMLU subjects: 88%|████████▊ | 50/57 [01:58<00:21, 3.13s/it][=] professional_medicine: 1/20 = 5.0% (running avg 27.8%)
 
============================================================
[>] professional_psychology (shots=5)
Downloading data: 100%|██████████| 133k/133k [00:00<00:00, 596kB/s]
Downloading data: 100%|██████████| 22.1k/22.1k [00:00<00:00, 81.6kB/s]
Downloading data: 100%|██████████| 4.69k/4.69k [00:00<00:00, 30.8kB/s]
Generating test split: 100%|██████████| 612/612 [00:00<00:00, 95884.13 examples/s]
Generating validation split: 100%|██████████| 69/69 [00:00<00:00, 16709.41 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1311.87 examples/s]
MMLU subjects: 89%|████████▉ | 51/57 [02:01<00:19, 3.24s/it][=] professional_psychology: 9/20 = 45.0% (running avg 28.1%)
 
============================================================
[>] public_relations (shots=5)
Downloading data: 100%|██████████| 20.6k/20.6k [00:00<00:00, 125kB/s]
Downloading data: 100%|██████████| 6.45k/6.45k [00:00<00:00, 35.0kB/s]
Downloading data: 100%|██████████| 4.43k/4.43k [00:00<00:00, 29.7kB/s]
Generating test split: 100%|██████████| 110/110 [00:00<00:00, 24437.15 examples/s]
Generating validation split: 100%|██████████| 12/12 [00:00<00:00, 3341.85 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1466.54 examples/s]
MMLU subjects: 91%|█████████ | 52/57 [02:05<00:16, 3.20s/it][=] public_relations: 9/20 = 45.0% (running avg 28.5%)
 
============================================================
[>] security_studies (shots=5)
Downloading data: 100%|██████████| 114k/114k [00:00<00:00, 671kB/s]
Downloading data: 100%|██████████| 18.7k/18.7k [00:00<00:00, 106kB/s]
Downloading data: 100%|██████████| 7.49k/7.49k [00:00<00:00, 44.8kB/s]
Generating test split: 100%|██████████| 245/245 [00:00<00:00, 37800.42 examples/s]
Generating validation split: 100%|██████████| 27/27 [00:00<00:00, 6508.78 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1451.42 examples/s]
MMLU subjects: 93%|█████████▎| 53/57 [02:08<00:12, 3.18s/it][=] security_studies: 5/20 = 25.0% (running avg 28.4%)
 
============================================================
[>] sociology (shots=5)
Downloading data: 100%|██████████| 43.9k/43.9k [00:00<00:00, 227kB/s]
Downloading data: 100%|██████████| 8.36k/8.36k [00:00<00:00, 50.9kB/s]
Downloading data: 100%|██████████| 4.21k/4.21k [00:00<00:00, 27.1kB/s]
Generating test split: 100%|██████████| 201/201 [00:00<00:00, 40145.48 examples/s]
Generating validation split: 100%|██████████| 22/22 [00:00<00:00, 5844.98 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1353.79 examples/s]
MMLU subjects: 95%|█████████▍| 54/57 [02:11<00:09, 3.13s/it][=] sociology: 4/20 = 20.0% (running avg 28.2%)
 
============================================================
[>] us_foreign_policy (shots=5)
Downloading data: 100%|██████████| 19.5k/19.5k [00:00<00:00, 113kB/s]
Downloading data: 100%|██████████| 5.27k/5.27k [00:00<00:00, 33.8kB/s]
Downloading data: 100%|██████████| 4.22k/4.22k [00:00<00:00, 22.6kB/s]
Generating test split: 100%|██████████| 100/100 [00:00<00:00, 23196.02 examples/s]
Generating validation split: 100%|██████████| 11/11 [00:00<00:00, 2971.81 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1376.35 examples/s]
MMLU subjects: 96%|█████████▋| 55/57 [02:14<00:06, 3.20s/it][=] us_foreign_policy: 5/20 = 25.0% (running avg 28.2%)
 
============================================================
[>] virology (shots=5)
Downloading data: 100%|██████████| 27.3k/27.3k [00:00<00:00, 175kB/s]
Downloading data: 100%|██████████| 7.05k/7.05k [00:00<00:00, 42.2kB/s]
Downloading data: 100%|██████████| 3.87k/3.87k [00:00<00:00, 23.3kB/s]
Generating test split: 100%|██████████| 166/166 [00:00<00:00, 42164.02 examples/s]
Generating validation split: 100%|██████████| 18/18 [00:00<00:00, 5863.43 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1653.77 examples/s]
MMLU subjects: 98%|█████████▊| 56/57 [02:17<00:03, 3.17s/it][=] virology: 5/20 = 25.0% (running avg 28.1%)
 
============================================================
[>] world_religions (shots=5)
Downloading data: 100%|██████████| 18.9k/18.9k [00:00<00:00, 106kB/s]
Downloading data: 100%|██████████| 4.94k/4.94k [00:00<00:00, 30.2kB/s]
Downloading data: 100%|██████████| 3.30k/3.30k [00:00<00:00, 21.1kB/s]
Generating test split: 100%|██████████| 171/171 [00:00<00:00, 39673.97 examples/s]
Generating validation split: 100%|██████████| 19/19 [00:00<00:00, 5365.37 examples/s]
Generating dev split: 100%|██████████| 5/5 [00:00<00:00, 1399.69 examples/s]
MMLU subjects: 100%|██████████| 57/57 [02:20<00:00, 2.47s/it]s/s]
[=] world_religions: 3/20 = 15.0% (running avg 27.9%)
 
============================================================
MMLU RESULT — /home/zeus/microllm2/microllm2-checkpoints/final_merged
Shots: 5 Subjects: 57/57
abstract_algebra 30.0% (6/20)
anatomy 25.0% (5/20)
astronomy 35.0% (7/20)
business_ethics 30.0% (6/20)
clinical_knowledge 45.0% (9/20)
college_biology 45.0% (9/20)
college_chemistry 15.0% (3/20)
college_computer_science 45.0% (9/20)
college_mathematics 35.0% (7/20)
college_medicine 30.0% (6/20)
college_physics 15.0% (3/20)
computer_security 30.0% (6/20)
conceptual_physics 5.0% (1/20)
econometrics 30.0% (6/20)
electrical_engineering 20.0% (4/20)
elementary_mathematics 30.0% (6/20)
formal_logic 10.0% (2/20)
global_facts 35.0% (7/20)
high_school_biology 45.0% (9/20)
high_school_chemistry 35.0% (7/20)
high_school_computer_science 35.0% (7/20)
high_school_european_history 20.0% (4/20)
high_school_geography 25.0% (5/20)
high_school_government_and_politics 20.0% (4/20)
high_school_macroeconomics 0.0% (0/20)
high_school_mathematics 20.0% (4/20)
high_school_microeconomics 35.0% (7/20)
high_school_physics 20.0% (4/20)
high_school_psychology 25.0% (5/20)
high_school_statistics 40.0% (8/20)
high_school_us_history 20.0% (4/20)
high_school_world_history 35.0% (7/20)
human_aging 40.0% (8/20)
human_sexuality 15.0% (3/20)
international_law 35.0% (7/20)
jurisprudence 40.0% (8/20)
logical_fallacies 35.0% (7/20)
machine_learning 50.0% (10/20)
management 20.0% (4/20)
marketing 35.0% (7/20)
medical_genetics 40.0% (8/20)
miscellaneous 30.0% (6/20)
moral_disputes 20.0% (4/20)
moral_scenarios 15.0% (3/20)
nutrition 20.0% (4/20)
philosophy 15.0% (3/20)
prehistory 25.0% (5/20)
professional_accounting 30.0% (6/20)
professional_law 35.0% (7/20)
professional_medicine 5.0% (1/20)
professional_psychology 45.0% (9/20)
public_relations 45.0% (9/20)
security_studies 25.0% (5/20)
sociology 20.0% (4/20)
us_foreign_policy 25.0% (5/20)
virology 25.0% (5/20)
world_religions 15.0% (3/20)
 
OVERALL: 318/1140 = 27.89%
============================================================
[+] Saved mmlu_results.json
 
Note: GPT2-XL base ~24-26% MMLU (random 25%). MicroLLM2 distilled should be 25-30% —
MMLU is knowledge-heavy; GPT2 1.5B 1024ctx cannot match 7B+ models. Use as sanity check, not SOTA claim.
[MMLU57-TOKEN] exit 0 at Sun Aug 9 23:12:32 UTC 2026
{
"model": "/home/zeus/microllm2/microllm2-checkpoints/final_merged",
"shots": 5,
"limit": 20,
"overall": {
"correct": 318,
"total": 1140,
"accuracy": 0.2789473684210526
},
"subjects": {
"abstract_algebra": {
"correct": 6,
"total": 20,
"accuracy": 0.3
},
"anatomy": {
"correct": 5,
"total": 20,
"accuracy": 0.25
},
"astronomy": {
"correct": 7,
"total": 20,
"accuracy": 0.35
},
"business_ethics": {
"correct": 6,
"total": 20,
"accuracy": 0.3
},
FINAL57 318 / 1140 27.89 57
'mmlu_results.json' -> 'mmlu57_results.json'
'mmlu_results.json' -> '/teamspace/studios/this_studio/microllm2/mmlu57_results.json'
'mmlu57_token.log' -> '/teamspace/studios/this_studio/microllm2/mmlu57_token.log'