Files
SmolLM-135M-neuraltxt-dpo-v1/logs/dpo_default_20260601_0821.log
ModelHub XC 62b47dc323 初始化项目,由ModelHub XC社区提供模型
Model: JaydeepR/SmolLM-135M-neuraltxt-dpo-v1
Source: Original Platform
2026-08-02 05:49:18 +08:00

357 lines
50 KiB
Plaintext
Raw Permalink Blame History

This file contains invisible Unicode characters

This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

==========================================
method=dpo run_id=dpo_default
base=paperbd/smollm_135M_neuraltxt_v1
dataset=paperbd/paper_preference_150K-v1
batch=16 grad_accum=8 (eff ~128)
==========================================
== uv sync ==
Resolved 148 packages in 6ms
Checked 130 packages in 454ms
== Step 0: baseline diversity (SFT model) ==
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
`torch_dtype` is deprecated! Use `dtype` instead!
Skipping import of cpp extensions due to incompatible torch version. Please upgrade to torch >= 2.11.0 (found 2.10.0+cu128).
Loading weights: 0%| | 0/272 [00:00<?, ?it/s]
Loading weights: 0%| | 1/272 [00:00<00:35, 7.61it/s]
Loading weights: 7%|▋ | 20/272 [00:00<00:03, 66.09it/s]
Loading weights: 10%|▉ | 26/272 [00:00<00:05, 46.55it/s]
Loading weights: 14%|█▍ | 38/272 [00:00<00:06, 36.76it/s]
Loading weights: 15%|█▌ | 42/272 [00:01<00:08, 25.59it/s]
Loading weights: 17%|█▋ | 45/272 [00:01<00:10, 21.54it/s]
Loading weights: 18%|█▊ | 49/272 [00:01<00:09, 23.44it/s]
Loading weights: 21%|██▏ | 58/272 [00:01<00:06, 32.12it/s]
Loading weights: 23%|██▎ | 62/272 [00:01<00:06, 31.49it/s]
Loading weights: 24%|██▍ | 66/272 [00:02<00:06, 29.78it/s]
Loading weights: 26%|██▋ | 72/272 [00:02<00:05, 35.44it/s]
Loading weights: 28%|██▊ | 76/272 [00:02<00:08, 24.20it/s]
Loading weights: 31%|███ | 83/272 [00:02<00:08, 23.20it/s]
Loading weights: 35%|███▍ | 95/272 [00:03<00:04, 37.27it/s]
Loading weights: 45%|████▍ | 122/272 [00:03<00:01, 77.29it/s]
Loading weights: 55%|█████▍ | 149/272 [00:03<00:01, 114.13it/s]
Loading weights: 64%|██████▎ | 173/272 [00:03<00:00, 138.25it/s]
Loading weights: 70%|███████ | 191/272 [00:03<00:00, 117.04it/s]
Loading weights: 76%|███████▌ | 206/272 [00:03<00:00, 113.11it/s]
Loading weights: 81%|████████ | 220/272 [00:03<00:00, 107.77it/s]
Loading weights: 86%|████████▌ | 233/272 [00:03<00:00, 100.86it/s]
Loading weights: 90%|█████████ | 245/272 [00:04<00:00, 92.20it/s]
Loading weights: 94%|█████████▍| 255/272 [00:04<00:00, 46.48it/s]
Loading weights: 97%|█████████▋| 263/272 [00:05<00:00, 35.94it/s]
Loading weights: 99%|█████████▉| 269/272 [00:05<00:00, 25.36it/s]
Loading weights: 100%|██████████| 272/272 [00:05<00:00, 47.51it/s]
[1/100] 16 responses avg_len=332 chars
t=0.3 **Question:** What is the purpose of the 'Format Matching Accuracy' metric in th...
t=0.3 {"question":"What is the purpose of the 'FM' metric in the dataset? How is the ...
t=0.3 **Question:** What is the purpose of the 'Format Matching Accuracy' metric in th...
t=0.3 **Question:** What does Figure 7 illustrate? **Answer:** Figure 7 illustrates t...
t=0.5 **Question:** What is the value of the 'statuscode' field in the context of the ...
t=0.5 **Question:** What does Figure 7 illustrate? **Answer:** Figure 7 illustrates t...
t=0.5 ### Q1 **Question:** What is the purpose of the 'Format Matching Accuracy' metri...
t=0.5 ### Q1 **Question:** What metrics can be used to control page layout in the data...
t=0.7 **Question:** What metrics are used to control page layout in the dataset? **An...
t=0.7 What are the key metrics used to control page layout in the dataset?...
t=0.7 **Question:** What is the purpose of the 'FM' metric in the dataset? **Answer:*...
t=0.7 **Question:** What metrics are used to control the page layout in the dataset? ...
t=1.0 **Question:** What does Figure 5 illustrate? **Answer:** Figure 5 illustrates t...
t=1.0 ["What is the purpose of the 'policy' table in Figure 4?", "How is the ' policy'...
t=1.0 **Question:** What is the purpose of the 'металь' variable (formula) in Figure 7...
t=1.0 ### Q1 **Question:** What does Figure 4 illustrate? **Answer:** Figure 4 illust...
[2/100] 16 responses avg_len=171 chars
t=0.3 [{"subject": "AoA", "relation": "is a", "object": "declaration of actions"}, {"s...
t=0.3 [{"subject": "AoA", "relation": "is a", "object": "declaration of actions"}, {"s...
t=0.3 [{"subject": "AoA", "relation": "is a", "object": "declaration of actions"}, {"s...
t=0.3 [{"subject": "AoA", "relation": "is a", "object": "declaration"}, {"subject": "O...
t=0.5 [{"subject": "AoA", "relation": "is a", "object": "declaration"}, {"subject": "O...
t=0.5 [{"subject": "Option", "relation": "is a", "object": "sequence of actions"}, {"s...
t=0.5 [{"subject": "AoA", "relation": "is a", "object": "declaration"}, {"subject": "O...
t=0.5 [{"subject": "AoA", "relation": "is a", "object": "declaration of actions"}, {"s...
t=0.7 [{"subject": "AoA", "relation": "is a", "object": "declaration"}, {"subject": "O...
t=0.7 [{"subject": "Option", "relation": "is a", "object": "sequence of actions"}, {"s...
t=0.7 [{"subject": "Option", "relation": "consists of", "object": "sequence of actions...
t=0.7 [{"subject": "Option", "relation": "consists of", "object": "sequence of actions...
t=1.0 [{"subject": "AoA", "relation": "is a", "object": "declaration of actions"}, {"s...
t=1.0 [{"subject": "Option A", "relation": "consists of", "object": "sequence of actio...
t=1.0 [{"subject": "AoA", "relation": "is", "object": "the second entry of path estima...
t=1.0 [{"subject": "AoA", "relation": "is a", "object": "indicatively a sequence of ac...
[3/100] 16 responses avg_len=209 chars
[4/100] 16 responses avg_len=913 chars
[5/100] 16 responses avg_len=923 chars
[6/100] 16 responses avg_len=288 chars
[7/100] 16 responses avg_len=101 chars
[8/100] 16 responses avg_len=215 chars
[9/100] 16 responses avg_len=160 chars
[10/100] 16 responses avg_len=277 chars
[11/100] 16 responses avg_len=681 chars
[12/100] 16 responses avg_len=731 chars
[13/100] 16 responses avg_len=123 chars
[14/100] 16 responses avg_len=678 chars
[15/100] 16 responses avg_len=248 chars
[16/100] 16 responses avg_len=960 chars
[17/100] 16 responses avg_len=631 chars
[18/100] 16 responses avg_len=1005 chars
[19/100] 16 responses avg_len=387 chars
[20/100] 16 responses avg_len=198 chars
[21/100] 16 responses avg_len=225 chars
[22/100] 16 responses avg_len=39 chars
[23/100] 16 responses avg_len=296 chars
[24/100] 16 responses avg_len=32 chars
[25/100] 16 responses avg_len=109 chars
[26/100] 16 responses avg_len=242 chars
[27/100] 16 responses avg_len=581 chars
[28/100] 16 responses avg_len=1107 chars
[29/100] 16 responses avg_len=141 chars
[30/100] 16 responses avg_len=320 chars
[31/100] 16 responses avg_len=1058 chars
[32/100] 16 responses avg_len=352 chars
[33/100] 16 responses avg_len=47 chars
[34/100] 16 responses avg_len=97 chars
[35/100] 16 responses avg_len=1184 chars
[36/100] 16 responses avg_len=722 chars
[37/100] 16 responses avg_len=654 chars
[38/100] 16 responses avg_len=107 chars
[39/100] 16 responses avg_len=228 chars
[40/100] 16 responses avg_len=128 chars
[41/100] 16 responses avg_len=767 chars
[42/100] 16 responses avg_len=214 chars
[43/100] 16 responses avg_len=447 chars
[44/100] 16 responses avg_len=284 chars
[45/100] 16 responses avg_len=72 chars
[46/100] 16 responses avg_len=167 chars
[47/100] 16 responses avg_len=124 chars
[48/100] 16 responses avg_len=1305 chars
[49/100] 16 responses avg_len=303 chars
[50/100] 16 responses avg_len=658 chars
[51/100] 16 responses avg_len=694 chars
[52/100] 16 responses avg_len=1016 chars
[53/100] 16 responses avg_len=240 chars
[54/100] 16 responses avg_len=107 chars
[55/100] 16 responses avg_len=355 chars
[56/100] 16 responses avg_len=1133 chars
[57/100] 16 responses avg_len=197 chars
[58/100] 16 responses avg_len=150 chars
[59/100] 16 responses avg_len=304 chars
[60/100] 16 responses avg_len=564 chars
[61/100] 16 responses avg_len=1131 chars
[62/100] 16 responses avg_len=687 chars
[63/100] 16 responses avg_len=193 chars
[64/100] 16 responses avg_len=685 chars
[65/100] 16 responses avg_len=102 chars
[66/100] 16 responses avg_len=249 chars
[67/100] 16 responses avg_len=572 chars
[68/100] 16 responses avg_len=732 chars
[69/100] 16 responses avg_len=90 chars
[70/100] 16 responses avg_len=183 chars
[71/100] 16 responses avg_len=224 chars
[72/100] 16 responses avg_len=1019 chars
[73/100] 16 responses avg_len=89 chars
[74/100] 16 responses avg_len=101 chars
[75/100] 16 responses avg_len=235 chars
[76/100] 16 responses avg_len=703 chars
[77/100] 16 responses avg_len=1104 chars
[78/100] 16 responses avg_len=479 chars
[79/100] 16 responses avg_len=574 chars
[80/100] 16 responses avg_len=117 chars
[81/100] 16 responses avg_len=195 chars
[82/100] 16 responses avg_len=505 chars
[83/100] 16 responses avg_len=273 chars
[84/100] 16 responses avg_len=188 chars
[85/100] 16 responses avg_len=69 chars
[86/100] 16 responses avg_len=156 chars
[87/100] 16 responses avg_len=230 chars
[88/100] 16 responses avg_len=183 chars
[89/100] 16 responses avg_len=1189 chars
[90/100] 16 responses avg_len=700 chars
[91/100] 16 responses avg_len=570 chars
[92/100] 16 responses avg_len=541 chars
[93/100] 16 responses avg_len=123 chars
[94/100] 16 responses avg_len=279 chars
[95/100] 16 responses avg_len=96 chars
[96/100] 16 responses avg_len=157 chars
[97/100] 16 responses avg_len=255 chars
[98/100] 16 responses avg_len=68 chars
[99/100] 16 responses avg_len=70 chars
[100/100] 16 responses avg_len=99 chars
Saved 100 records to evals/baseline_smollm_135M_neuraltxt_v1_n100_r4.jsonl
Skipping import of cpp extensions due to incompatible torch version. Please upgrade to torch >= 2.11.0 (found 2.10.0+cu128).
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
Loading SBERT model: sentence-transformers/all-MiniLM-L6-v2 ...
Loading weights: 0%| | 0/103 [00:00<?, ?it/s]
Loading weights: 100%|██████████| 103/103 [00:00<00:00, 3833.95it/s]
BertModel LOAD REPORT from: sentence-transformers/all-MiniLM-L6-v2
Key | Status | |
------------------------+------------+--+-
embeddings.position_ids | UNEXPECTED | |
Notes:
- UNEXPECTED: can be ignored when loading from different task/architecture; not ok if you expect identical arch.
Scoring 100 records ...
[1/100] ead=0.3242 sbert=0.5795 vendi=6.4491
[2/100] ead=0.0646 sbert=0.1333 vendi=1.8002
[3/100] ead=0.2121 sbert=0.4044 vendi=4.2298
[4/100] ead=0.3804 sbert=0.6572 vendi=9.1213
[5/100] ead=0.3125 sbert=0.5625 vendi=7.6994
[6/100] ead=0.0421 sbert=0.0879 vendi=1.4742
[7/100] ead=0.0467 sbert=0.0973 vendi=1.5481
[8/100] ead=0.0822 sbert=0.1682 vendi=2.1513
[9/100] ead=0.0569 sbert=0.1179 vendi=1.7004
[10/100] ead=0.0401 sbert=0.0838 vendi=1.5064
[11/100] ead=0.0499 sbert=0.1038 vendi=1.6867
[12/100] ead=0.0606 sbert=0.1254 vendi=1.8828
[13/100] ead=0.0881 sbert=0.1797 vendi=1.9738
[14/100] ead=0.0966 sbert=0.1962 vendi=2.4463
[15/100] ead=0.1188 sbert=0.2385 vendi=2.3770
[16/100] ead=0.3707 sbert=0.6442 vendi=9.0352
[17/100] ead=0.0473 sbert=0.0985 vendi=1.6221
[18/100] ead=0.3004 sbert=0.5446 vendi=7.3248
[19/100] ead=0.1918 sbert=0.3700 vendi=4.1698
[20/100] ead=0.0839 sbert=0.1715 vendi=2.0533
[21/100] ead=0.0750 sbert=0.1539 vendi=1.8828
[22/100] ead=0.2175 sbert=0.4136 vendi=2.3536
[23/100] ead=0.1908 sbert=0.3682 vendi=3.5052
[24/100] ead=0.0035 sbert=0.0075 vendi=1.0423
[25/100] ead=0.1548 sbert=0.3047 vendi=2.4538
[26/100] ead=0.1845 sbert=0.3573 vendi=3.4154
[27/100] ead=0.1386 sbert=0.2753 vendi=3.2602
[28/100] ead=0.2224 sbert=0.4216 vendi=5.1495
[29/100] ead=0.1245 sbert=0.2490 vendi=2.3464
[30/100] ead=0.0156 sbert=0.0331 vendi=1.1984
[31/100] ead=0.2515 sbert=0.4691 vendi=6.0774
[32/100] ead=0.0614 sbert=0.1269 vendi=1.7186
[33/100] ead=0.1728 sbert=0.3368 vendi=3.4806
[34/100] ead=0.1283 sbert=0.2561 vendi=2.4322
[35/100] ead=0.3394 sbert=0.6012 vendi=6.0677
[36/100] ead=0.0324 sbert=0.0681 vendi=1.4721
[37/100] ead=0.0758 sbert=0.1556 vendi=2.0815
[38/100] ead=0.0112 sbert=0.0238 vendi=1.1322
[39/100] ead=0.0554 sbert=0.1150 vendi=1.7443
[40/100] ead=0.0705 sbert=0.1450 vendi=1.7605
[41/100] ead=0.0465 sbert=0.0968 vendi=1.6472
[42/100] ead=0.0583 sbert=0.1208 vendi=1.7793
[43/100] ead=0.0751 sbert=0.1542 vendi=1.9610
[44/100] ead=0.1985 sbert=0.3814 vendi=3.4862
[45/100] ead=0.2542 sbert=0.4734 vendi=3.5825
[46/100] ead=0.1904 sbert=0.3676 vendi=3.3044
[47/100] ead=0.0485 sbert=0.1010 vendi=1.5474
[48/100] ead=0.1395 sbert=0.2768 vendi=3.1843
[49/100] ead=0.0669 sbert=0.1379 vendi=1.8626
[50/100] ead=0.0585 sbert=0.1211 vendi=1.8633
[51/100] ead=0.0531 sbert=0.1103 vendi=1.7103
[52/100] ead=0.0685 sbert=0.1412 vendi=2.0075
[53/100] ead=0.0150 sbert=0.0317 vendi=1.2069
[54/100] ead=0.1283 sbert=0.2561 vendi=2.6694
[55/100] ead=0.0505 sbert=0.1050 vendi=1.6928
[56/100] ead=0.0700 sbert=0.1441 vendi=2.0044
[57/100] ead=0.0616 sbert=0.1273 vendi=1.8634
[58/100] ead=0.3127 sbert=0.5628 vendi=3.8481
[59/100] ead=0.2337 sbert=0.4404 vendi=4.7649
[60/100] ead=0.0817 sbert=0.1672 vendi=2.1959
[61/100] ead=0.2487 sbert=0.4646 vendi=6.0817
[62/100] ead=0.0460 sbert=0.0959 vendi=1.6597
[63/100] ead=0.1015 sbert=0.2055 vendi=2.2814
[64/100] ead=0.0349 sbert=0.0732 vendi=1.4762
[65/100] ead=0.2045 sbert=0.3917 vendi=2.4907
[66/100] ead=0.1410 sbert=0.2795 vendi=2.5030
[67/100] ead=0.2509 sbert=0.4681 vendi=5.4121
[68/100] ead=0.0335 sbert=0.0702 vendi=1.4732
[69/100] ead=0.0607 sbert=0.1255 vendi=1.5118
[70/100] ead=0.0736 sbert=0.1512 vendi=1.8842
[71/100] ead=0.1013 sbert=0.2052 vendi=2.3555
[72/100] ead=0.2619 sbert=0.4856 vendi=6.0547
[73/100] ead=0.0161 sbert=0.0340 vendi=1.1377
[74/100] ead=0.0587 sbert=0.1215 vendi=1.6269
[75/100] ead=0.1286 sbert=0.2567 vendi=2.5381
[76/100] ead=0.0828 sbert=0.1694 vendi=2.2143
[77/100] ead=0.0533 sbert=0.1107 vendi=1.7308
[78/100] ead=0.0733 sbert=0.1506 vendi=1.9479
[79/100] ead=0.0250 sbert=0.0526 vendi=1.3521
[80/100] ead=0.1060 sbert=0.2141 vendi=1.9917
[81/100] ead=0.0193 sbert=0.0407 vendi=1.2564
[82/100] ead=0.0750 sbert=0.1539 vendi=2.1178
[83/100] ead=0.0382 sbert=0.0800 vendi=1.4891
[84/100] ead=0.0327 sbert=0.0687 vendi=1.2990
[85/100] ead=0.2022 sbert=0.3878 vendi=2.9515
[86/100] ead=0.0523 sbert=0.1087 vendi=1.6659
[87/100] ead=0.1255 sbert=0.2509 vendi=2.5125
[88/100] ead=0.1345 sbert=0.2677 vendi=2.8861
[89/100] ead=0.2613 sbert=0.4847 vendi=6.2508
[90/100] ead=0.0807 sbert=0.1653 vendi=2.1725
[91/100] ead=0.3083 sbert=0.5563 vendi=7.4214
[92/100] ead=0.1516 sbert=0.2989 vendi=3.3448
[93/100] ead=0.0216 sbert=0.0455 vendi=1.2171
[94/100] ead=0.0789 sbert=0.1618 vendi=2.0897
[95/100] ead=0.0273 sbert=0.0574 vendi=1.2478
[96/100] ead=0.0895 sbert=0.1824 vendi=2.1625
[97/100] ead=0.1938 sbert=0.3735 vendi=3.1834
[98/100] ead=0.0165 sbert=0.0350 vendi=1.1601
[99/100] ead=0.0036 sbert=0.0076 vendi=1.0425
[100/100] ead=0.0068 sbert=0.0145 vendi=1.0716
--- summary ---
ead mean=0.1173 min=0.0035 max=0.3804
sbert mean=0.2263 min=0.0075 max=0.6572
vendi mean=2.7327 min=1.0423 max=9.1213
Saved to evals/baseline_smollm_135M_neuraltxt_v1_n100_r4_diversity.jsonl
== Step 1: dpo training ==
🦥 Unsloth: Will patch your computer to enable 2x faster free finetuning.
Unsloth: Your Flash Attention 2 installation seems to be broken. Using Xformers instead. No performance changes will be seen.
🦥 Unsloth Zoo will now patch everything to make training faster!
Method: DPO
Base model: paperbd/smollm_135M_neuraltxt_v1
Dataset: paperbd/paper_preference_150K-v1
==((====))== Unsloth 2026.5.9: Fast Llama patching. Transformers: 5.5.0.
\\ /| NVIDIA GeForce RTX 3090. Num GPUs = 1. Max memory: 23.559 GB. Platform: Linux.
O^O/ \_/ \ Torch: 2.10.0+cu128. CUDA: 8.6. CUDA Toolkit: 12.8. Triton: 3.6.0
\ / Bfloat16 = TRUE. FA [Xformers = 0.0.35. FA2 = False]
"-____-" Free license: http://github.com/unslothai/unsloth
Unsloth: Fast downloading is enabled - ignore downloading bars which are red colored!
Loading weights: 0%| | 0/272 [00:00<?, ?it/s]
Loading weights: 0%| | 1/272 [00:00<00:48, 5.53it/s]
Loading weights: 1%| | 3/272 [00:00<00:27, 9.68it/s]
Loading weights: 7%|▋ | 20/272 [00:00<00:05, 48.17it/s]
Loading weights: 11%|█▏ | 31/272 [00:00<00:04, 59.82it/s]
Loading weights: 14%|█▍ | 39/272 [00:00<00:04, 48.30it/s]
Loading weights: 17%|█▋ | 45/272 [00:01<00:05, 43.78it/s]
Loading weights: 21%|██▏ | 58/272 [00:01<00:04, 52.79it/s]
Loading weights: 40%|███▉ | 108/272 [00:01<00:01, 143.62it/s]
Loading weights: 58%|█████▊ | 157/272 [00:01<00:00, 221.55it/s]
Loading weights: 68%|██████▊ | 186/272 [00:01<00:00, 227.39it/s]
Loading weights: 98%|█████████▊| 266/272 [00:01<00:00, 369.44it/s]
Loading weights: 100%|██████████| 272/272 [00:01<00:00, 162.16it/s]
Unsloth: Will map <|im_end|> to EOS = <|im_end|>.
Unsloth 2026.5.9 patched 30 layers with 30 QKV layers, 30 O layers and 30 MLP layers.
Generating train split: 0%| | 0/120000 [00:00<?, ? examples/s]
Generating train split: 4%|▍ | 5130/120000 [00:00<00:02, 38583.15 examples/s]
Generating train split: 13%|█▎ | 15356/120000 [00:00<00:01, 58250.06 examples/s]
Generating train split: 21%|██▏ | 25536/120000 [00:00<00:01, 72439.69 examples/s]
Generating train split: 30%|██▉ | 35832/120000 [00:00<00:01, 80698.81 examples/s]
Generating train split: 38%|███▊ | 46055/120000 [00:00<00:00, 85865.95 examples/s]
Generating train split: 47%|████▋ | 56286/120000 [00:00<00:00, 90403.14 examples/s]
Generating train split: 60%|█████▉ | 71551/120000 [00:00<00:00, 95799.15 examples/s]
Generating train split: 68%|██████▊ | 81770/120000 [00:00<00:00, 95922.40 examples/s]
Generating train split: 77%|███████▋ | 91974/120000 [00:01<00:00, 96922.16 examples/s]
Generating train split: 89%|████████▉ | 107312/120000 [00:01<00:00, 101120.87 examples/s]
Generating train split: 100%|██████████| 120000/120000 [00:01<00:00, 103243.96 examples/s]
Generating train split: 100%|██████████| 120000/120000 [00:01<00:00, 83913.46 examples/s]
Generating test split: 0%| | 0/31380 [00:00<?, ? examples/s]
Generating test split: 49%|████▉ | 15326/31380 [00:00<00:00, 90441.08 examples/s]
Generating test split: 100%|██████████| 31380/31380 [00:00<00:00, 102633.98 examples/s]
Generating test split: 100%|██████████| 31380/31380 [00:01<00:00, 26525.99 examples/s]
Map: 0%| | 0/120000 [00:00<?, ? examples/s]
Map: 1%| | 892/120000 [00:00<00:13, 8807.70 examples/s]
Map: 2%|▏ | 1955/120000 [00:00<00:11, 9871.75 examples/s]
Map: 2%|▎ | 3000/120000 [00:00<00:11, 10068.58 examples/s]
Map: 4%|▎ | 4258/120000 [00:00<00:10, 11049.85 examples/s]
Map: 5%|▍ | 5734/120000 [00:00<00:09, 11863.96 examples/s]
Map: 6%|▌ | 6927/120000 [00:00<00:09, 11881.11 examples/s]
Map: 7%|▋ | 8722/120000 [00:00<00:09, 11875.01 examples/s]
Map: 8%|▊ | 10000/120000 [00:00<00:09, 11655.11 examples/s]
Map: 9%|▉ | 11215/120000 [00:00<00:09, 11790.20 examples/s]
Map: 10%|█ | 12530/120000 [00:01<00:08, 12175.08 examples/s]
Map: 12%|█▏ | 13856/120000 [00:01<00:08, 12486.12 examples/s]
Map: 13%|█▎ | 15730/120000 [00:01<00:08, 12196.96 examples/s]
Map: 14%|█▍ | 17000/120000 [00:01<00:08, 11910.16 examples/s]
Map: 15%|█▌ | 18294/120000 [00:01<00:08, 12182.96 examples/s]
Map: 16%|█▋ | 19744/120000 [00:01<00:08, 12417.22 examples/s]
Map: 18%|█▊ | 21429/120000 [00:01<00:08, 11916.69 examples/s]
Map: 19%|█▉ | 22756/120000 [00:01<00:08, 12051.18 examples/s]
Map: 20%|██ | 24000/120000 [00:02<00:08, 11726.27 examples/s]
Map: 21%|██ | 25286/120000 [00:02<00:07, 12026.26 examples/s]
Map: 22%|██▏ | 26751/120000 [00:02<00:07, 12275.76 examples/s]
Map: 23%|██▎ | 28000/120000 [00:02<00:07, 11870.23 examples/s]
Map: 25%|██▍ | 29928/120000 [00:02<00:07, 12224.03 examples/s]
Map: 26%|██▋ | 31745/120000 [00:02<00:07, 12138.95 examples/s]
Map: 28%|██▊ | 33000/120000 [00:02<00:07, 11896.00 examples/s]
Map: 29%|██▊ | 34304/120000 [00:02<00:07, 12185.17 examples/s]
Map: 30%|██▉ | 35703/120000 [00:02<00:06, 12284.01 examples/s]
Map: 31%|███ | 36959/120000 [00:03<00:06, 12357.33 examples/s]
Map: 32%|███▏ | 38821/120000 [00:03<00:06, 12374.10 examples/s]
Map: 34%|███▍ | 40789/120000 [00:03<00:06, 12326.39 examples/s]
Map: 36%|███▌ | 42781/120000 [00:03<00:06, 12276.63 examples/s]
Map: 37%|███▋ | 44489/120000 [00:03<00:06, 11956.57 examples/s]
Map: 38%|███▊ | 45788/120000 [00:03<00:06, 12187.55 examples/s]
Map: 40%|███▉ | 47745/120000 [00:03<00:05, 12152.51 examples/s]
Map: 41%|████ | 49000/120000 [00:04<00:05, 11869.34 examples/s]
Map: 42%|████▏ | 50254/120000 [00:04<00:05, 12031.34 examples/s]
Map: 43%|████▎ | 51725/120000 [00:04<00:05, 12332.12 examples/s]
Map: 44%|████▍ | 53000/120000 [00:04<00:05, 12018.18 examples/s]
Map: 45%|████▌ | 54257/120000 [00:04<00:05, 12164.17 examples/s]
Map: 46%|████▋ | 55748/120000 [00:04<00:05, 12501.76 examples/s]
Map: 48%|████▊ | 57753/120000 [00:04<00:05, 12448.17 examples/s]
Map: 50%|████▉ | 59747/120000 [00:04<00:04, 12415.26 examples/s]
Map: 51%|█████ | 61000/120000 [00:05<00:04, 12122.05 examples/s]
Map: 52%|█████▏ | 62290/120000 [00:05<00:04, 12313.53 examples/s]
Map: 53%|█████▎ | 63744/120000 [00:05<00:04, 12456.76 examples/s]
Map: 54%|█████▍ | 65000/120000 [00:05<00:04, 12039.41 examples/s]
Map: 55%|█████▌ | 66291/120000 [00:05<00:04, 12270.00 examples/s]
Map: 56%|█████▋ | 67740/120000 [00:05<00:04, 12439.88 examples/s]
Map: 58%|█████▊ | 69487/120000 [00:05<00:04, 12147.69 examples/s]
Map: 59%|█████▉ | 70767/120000 [00:05<00:03, 12315.06 examples/s]
Map: 60%|██████ | 72371/120000 [00:06<00:07, 6239.17 examples/s]
Map: 61%|██████▏ | 73748/120000 [00:06<00:06, 7293.37 examples/s]
Map: 62%|██████▎ | 75000/120000 [00:06<00:05, 8030.57 examples/s]
Map: 64%|██████▎ | 76277/120000 [00:06<00:04, 8968.93 examples/s]
Map: 65%|██████▍ | 77725/120000 [00:06<00:04, 9848.76 examples/s]
Map: 66%|██████▌ | 78964/120000 [00:06<00:03, 10441.02 examples/s]
Map: 67%|██████▋ | 80743/120000 [00:07<00:03, 10883.00 examples/s]
Map: 68%|██████▊ | 82000/120000 [00:07<00:03, 10936.59 examples/s]
Map: 69%|██████▉ | 83322/120000 [00:07<00:03, 11474.54 examples/s]
Map: 71%|███████ | 84731/120000 [00:07<00:02, 11804.19 examples/s]
Map: 72%|███████▏ | 85987/120000 [00:07<00:02, 12005.22 examples/s]
Map: 73%|███████▎ | 87838/120000 [00:07<00:02, 12124.70 examples/s]
Map: 75%|███████▍ | 89792/120000 [00:07<00:02, 12169.56 examples/s]
Map: 76%|███████▋ | 91789/120000 [00:07<00:02, 12205.01 examples/s]
Map: 78%|███████▊ | 93788/120000 [00:08<00:02, 12272.65 examples/s]
Map: 80%|███████▉ | 95762/120000 [00:08<00:01, 12231.54 examples/s]
Map: 81%|████████ | 97000/120000 [00:08<00:01, 11931.15 examples/s]
Map: 82%|████████▏ | 98253/120000 [00:08<00:01, 12067.99 examples/s]
Map: 83%|████████▎ | 99738/120000 [00:08<00:01, 12377.58 examples/s]
Map: 84%|████████▍ | 101000/120000 [00:08<00:01, 12054.15 examples/s]
Map: 85%|████████▌ | 102252/120000 [00:08<00:01, 12173.69 examples/s]
Map: 86%|████████▋ | 103752/120000 [00:08<00:01, 12515.32 examples/s]
Map: 88%|████████▊ | 105741/120000 [00:09<00:01, 12438.92 examples/s]
Map: 89%|████████▉ | 107000/120000 [00:09<00:01, 12064.07 examples/s]
Map: 90%|█████████ | 108286/120000 [00:09<00:00, 12266.27 examples/s]
Map: 91%|█████████▏| 109756/120000 [00:09<00:00, 12493.44 examples/s]
Map: 93%|█████████▎| 111754/120000 [00:09<00:00, 12493.85 examples/s]
Map: 95%|█████████▍| 113760/120000 [00:09<00:00, 12467.75 examples/s]
Map: 96%|█████████▋| 115746/120000 [00:09<00:00, 12416.13 examples/s]
Map: 98%|█████████▊| 117482/120000 [00:10<00:00, 12155.03 examples/s]
Map: 99%|█████████▉| 118746/120000 [00:10<00:00, 12261.82 examples/s]
Map: 100%|██████████| 120000/120000 [00:10<00:00, 12007.25 examples/s]
Map: 100%|██████████| 120000/120000 [00:10<00:00, 11541.65 examples/s]
warmup_ratio is deprecated and will be removed in v5.2. Use `warmup_steps` instead.
Extracting prompt in train dataset (num_proc=64): 0%| | 0/117600 [00:00<?, ? examples/s]
Extracting prompt in train dataset (num_proc=64): 0%| | 17/117600 [00:03<6:49:54, 4.78 examples/s]
Extracting prompt in train dataset (num_proc=64): 2%|▏ | 2508/117600 [00:03<01:59, 962.90 examples/s]
Extracting prompt in train dataset (num_proc=64): 4%|▍ | 4571/117600 [00:03<00:56, 1983.47 examples/s]
Extracting prompt in train dataset (num_proc=64): 11%|█ | 13085/117600 [00:03<00:13, 7723.59 examples/s]
Extracting prompt in train dataset (num_proc=64): 18%|█▊ | 21705/117600 [00:03<00:06, 14669.52 examples/s]
Extracting prompt in train dataset (num_proc=64): 27%|██▋ | 31648/117600 [00:04<00:03, 24189.06 examples/s]
Extracting prompt in train dataset (num_proc=64): 38%|███▊ | 44563/117600 [00:04<00:01, 38737.63 examples/s]
Extracting prompt in train dataset (num_proc=64): 47%|████▋ | 55568/117600 [00:04<00:01, 50527.31 examples/s]
Extracting prompt in train dataset (num_proc=64): 55%|█████▌ | 65243/117600 [00:04<00:00, 58766.76 examples/s]
Extracting prompt in train dataset (num_proc=64): 65%|██████▍ | 76246/117600 [00:04<00:00, 69637.97 examples/s]
Extracting prompt in train dataset (num_proc=64): 74%|███████▍ | 86781/117600 [00:04<00:00, 77925.91 examples/s]
Extracting prompt in train dataset (num_proc=64): 82%|████████▏ | 96956/117600 [00:04<00:00, 66577.55 examples/s]
Extracting prompt in train dataset (num_proc=64): 90%|████████▉ | 105540/117600 [00:04<00:00, 58558.53 examples/s]
Extracting prompt in train dataset (num_proc=64): 96%|█████████▌| 112975/117600 [00:05<00:00, 53786.07 examples/s]
Extracting prompt in train dataset (num_proc=64): 100%|██████████| 117600/117600 [00:14<00:00, 7986.30 examples/s]
Applying chat template to train dataset (num_proc=64): 0%| | 0/117600 [00:00<?, ? examples/s]
Applying chat template to train dataset (num_proc=64): 0%| | 71/117600 [00:04<1:58:59, 16.46 examples/s]
Applying chat template to train dataset (num_proc=64): 0%| | 446/117600 [00:04<14:20, 136.13 examples/s]
Applying chat template to train dataset (num_proc=64): 1%|▏ | 1749/117600 [00:04<02:47, 693.69 examples/s]
Applying chat template to train dataset (num_proc=64): 3%|▎ | 3935/117600 [00:04<00:59, 1909.93 examples/s]
Applying chat template to train dataset (num_proc=64): 5%|▌ | 6463/117600 [00:04<00:30, 3683.59 examples/s]
Applying chat template to train dataset (num_proc=64): 9%|▉ | 10481/117600 [00:04<00:15, 7114.89 examples/s]
Applying chat template to train dataset (num_proc=64): 13%|█▎ | 14923/117600 [00:04<00:09, 11328.05 examples/s]
Applying chat template to train dataset (num_proc=64): 16%|█▌ | 18757/117600 [00:05<00:06, 15132.15 examples/s]
Applying chat template to train dataset (num_proc=64): 19%|█▊ | 21757/117600 [00:05<00:05, 17316.60 examples/s]
Applying chat template to train dataset (num_proc=64): 21%|██ | 24671/117600 [00:05<00:05, 17238.09 examples/s]
Applying chat template to train dataset (num_proc=64): 23%|██▎ | 27558/117600 [00:05<00:04, 18926.08 examples/s]
Applying chat template to train dataset (num_proc=64): 26%|██▌ | 30088/117600 [00:05<00:04, 20098.03 examples/s]
Applying chat template to train dataset (num_proc=64): 28%|██▊ | 32642/117600 [00:05<00:04, 20216.03 examples/s]
Applying chat template to train dataset (num_proc=64): 30%|██▉ | 35055/117600 [00:05<00:05, 15110.39 examples/s]
Applying chat template to train dataset (num_proc=64): 32%|███▏ | 37046/117600 [00:06<00:05, 15600.90 examples/s]
Applying chat template to train dataset (num_proc=64): 33%|███▎ | 38968/117600 [00:06<00:06, 13087.48 examples/s]
Applying chat template to train dataset (num_proc=64): 40%|███▉ | 46540/117600 [00:06<00:03, 23643.37 examples/s]
Applying chat template to train dataset (num_proc=64): 44%|████▎ | 51271/117600 [00:06<00:02, 28537.04 examples/s]
Applying chat template to train dataset (num_proc=64): 51%|█████▏ | 60295/117600 [00:06<00:01, 42601.85 examples/s]
Applying chat template to train dataset (num_proc=64): 58%|█████▊ | 67938/117600 [00:06<00:00, 50900.45 examples/s]
Applying chat template to train dataset (num_proc=64): 64%|██████▍ | 75271/117600 [00:06<00:00, 56695.28 examples/s]
Applying chat template to train dataset (num_proc=64): 72%|███████▏ | 84773/117600 [00:06<00:00, 66308.37 examples/s]
Applying chat template to train dataset (num_proc=64): 83%|████████▎ | 97068/117600 [00:07<00:00, 81916.65 examples/s]
Applying chat template to train dataset (num_proc=64): 90%|████████▉ | 105801/117600 [00:07<00:00, 64758.04 examples/s]
Applying chat template to train dataset (num_proc=64): 96%|█████████▋| 113202/117600 [00:07<00:00, 48439.75 examples/s]
Applying chat template to train dataset (num_proc=64): 100%|██████████| 117600/117600 [00:16<00:00, 7071.73 examples/s]
Tokenizing train dataset (num_proc=64): 0%| | 0/117600 [00:00<?, ? examples/s]
Tokenizing train dataset (num_proc=64): 0%| | 25/117600 [00:04<6:08:41, 5.31 examples/s]
Tokenizing train dataset (num_proc=64): 0%| | 110/117600 [00:04<1:05:09, 30.05 examples/s]
Tokenizing train dataset (num_proc=64): 0%| | 409/117600 [00:04<13:14, 147.50 examples/s]
Tokenizing train dataset (num_proc=64): 1%| | 991/117600 [00:05<04:17, 452.21 examples/s]
Tokenizing train dataset (num_proc=64): 1%|▏ | 1744/117600 [00:05<02:02, 944.07 examples/s]
Tokenizing train dataset (num_proc=64): 3%|▎ | 3056/117600 [00:05<00:56, 2011.21 examples/s]
Tokenizing train dataset (num_proc=64): 4%|▍ | 4680/117600 [00:05<00:32, 3521.58 examples/s]
Tokenizing train dataset (num_proc=64): 5%|▌ | 6297/117600 [00:05<00:21, 5183.53 examples/s]
Tokenizing train dataset (num_proc=64): 7%|▋ | 7998/117600 [00:05<00:15, 6920.36 examples/s]
Tokenizing train dataset (num_proc=64): 8%|▊ | 9818/117600 [00:05<00:12, 8813.53 examples/s]
Tokenizing train dataset (num_proc=64): 10%|▉ | 11576/117600 [00:05<00:09, 10604.86 examples/s]
Tokenizing train dataset (num_proc=64): 11%|█ | 13044/117600 [00:05<00:09, 11448.50 examples/s]
Tokenizing train dataset (num_proc=64): 12%|█▏ | 14500/117600 [00:06<00:08, 11707.23 examples/s]
Tokenizing train dataset (num_proc=64): 14%|█▎ | 15896/117600 [00:06<00:09, 10771.81 examples/s]
Tokenizing train dataset (num_proc=64): 15%|█▍ | 17587/117600 [00:06<00:08, 11608.38 examples/s]
Tokenizing train dataset (num_proc=64): 16%|█▌ | 18868/117600 [00:06<00:08, 11685.05 examples/s]
Tokenizing train dataset (num_proc=64): 17%|█▋ | 20132/117600 [00:06<00:08, 11718.90 examples/s]
Tokenizing train dataset (num_proc=64): 18%|█▊ | 21367/117600 [00:06<00:09, 9827.66 examples/s]
Tokenizing train dataset (num_proc=64): 19%|█▉ | 22449/117600 [00:07<00:15, 6007.24 examples/s]
Tokenizing train dataset (num_proc=64): 20%|██ | 23821/117600 [00:07<00:13, 6996.73 examples/s]
Tokenizing train dataset (num_proc=64): 22%|██▏ | 26149/117600 [00:07<00:09, 10045.65 examples/s]
Tokenizing train dataset (num_proc=64): 23%|██▎ | 27504/117600 [00:07<00:10, 8805.21 examples/s]
Tokenizing train dataset (num_proc=64): 27%|██▋ | 31232/117600 [00:07<00:05, 14429.68 examples/s]
Tokenizing train dataset (num_proc=64): 30%|███ | 35846/117600 [00:07<00:03, 21405.08 examples/s]
Tokenizing train dataset (num_proc=64): 34%|███▍ | 39957/117600 [00:07<00:02, 26115.52 examples/s]
Tokenizing train dataset (num_proc=64): 37%|███▋ | 43101/117600 [00:08<00:03, 22624.18 examples/s]
Tokenizing train dataset (num_proc=64): 39%|███▉ | 45834/117600 [00:08<00:03, 19317.61 examples/s]
Tokenizing train dataset (num_proc=64): 41%|████ | 48156/117600 [00:08<00:03, 19839.95 examples/s]
Tokenizing train dataset (num_proc=64): 43%|████▎ | 50440/117600 [00:08<00:03, 16929.60 examples/s]
Tokenizing train dataset (num_proc=64): 45%|████▍ | 52401/117600 [00:08<00:03, 17397.60 examples/s]
Tokenizing train dataset (num_proc=64): 46%|████▌ | 54335/117600 [00:08<00:03, 17767.58 examples/s]
Tokenizing train dataset (num_proc=64): 48%|████▊ | 56265/117600 [00:08<00:04, 15116.65 examples/s]
Tokenizing train dataset (num_proc=64): 49%|████▉ | 57955/117600 [00:09<00:03, 15257.16 examples/s]
Tokenizing train dataset (num_proc=64): 51%|█████ | 59631/117600 [00:09<00:03, 15586.82 examples/s]
Tokenizing train dataset (num_proc=64): 52%|█████▏ | 61343/117600 [00:09<00:03, 15971.45 examples/s]
Tokenizing train dataset (num_proc=64): 54%|█████▎ | 63020/117600 [00:09<00:03, 16128.97 examples/s]
Tokenizing train dataset (num_proc=64): 55%|█████▌ | 64703/117600 [00:09<00:03, 16261.99 examples/s]
Tokenizing train dataset (num_proc=64): 57%|█████▋ | 66452/117600 [00:09<00:03, 16608.08 examples/s]
Tokenizing train dataset (num_proc=64): 58%|█████▊ | 68221/117600 [00:09<00:02, 16914.06 examples/s]
Tokenizing train dataset (num_proc=64): 59%|█████▉ | 69944/117600 [00:09<00:02, 16827.58 examples/s]
Tokenizing train dataset (num_proc=64): 61%|██████ | 71839/117600 [00:09<00:02, 17433.49 examples/s]
Tokenizing train dataset (num_proc=64): 63%|██████▎ | 73616/117600 [00:09<00:02, 17366.18 examples/s]
Tokenizing train dataset (num_proc=64): 64%|██████▍ | 75572/117600 [00:10<00:02, 18009.25 examples/s]
Tokenizing train dataset (num_proc=64): 66%|██████▌ | 77406/117600 [00:10<00:02, 17909.74 examples/s]
Tokenizing train dataset (num_proc=64): 67%|██████▋ | 79217/117600 [00:10<00:02, 14275.56 examples/s]
Tokenizing train dataset (num_proc=64): 69%|██████▊ | 80818/117600 [00:10<00:02, 14654.53 examples/s]
Tokenizing train dataset (num_proc=64): 70%|███████ | 82398/117600 [00:10<00:02, 14484.53 examples/s]
Tokenizing train dataset (num_proc=64): 72%|███████▏ | 84797/117600 [00:10<00:01, 17018.98 examples/s]
Tokenizing train dataset (num_proc=64): 74%|███████▎ | 86587/117600 [00:10<00:01, 17057.64 examples/s]
Tokenizing train dataset (num_proc=64): 75%|███████▌ | 88354/117600 [00:10<00:02, 13950.26 examples/s]
Tokenizing train dataset (num_proc=64): 77%|███████▋ | 89979/117600 [00:11<00:01, 14509.20 examples/s]
Tokenizing train dataset (num_proc=64): 78%|███████▊ | 91654/117600 [00:11<00:01, 15089.23 examples/s]
Tokenizing train dataset (num_proc=64): 79%|███████▉ | 93275/117600 [00:11<00:01, 15141.39 examples/s]
Tokenizing train dataset (num_proc=64): 82%|████████▏ | 96114/117600 [00:11<00:01, 17315.82 examples/s]
Tokenizing train dataset (num_proc=64): 83%|████████▎ | 97993/117600 [00:11<00:01, 17571.36 examples/s]
Tokenizing train dataset (num_proc=64): 85%|████████▍ | 99826/117600 [00:11<00:01, 17717.56 examples/s]
Tokenizing train dataset (num_proc=64): 86%|████████▋ | 101615/117600 [00:11<00:00, 17553.92 examples/s]
Tokenizing train dataset (num_proc=64): 88%|████████▊ | 103476/117600 [00:11<00:00, 17818.91 examples/s]
Tokenizing train dataset (num_proc=64): 90%|████████▉ | 105283/117600 [00:11<00:00, 17746.47 examples/s]
Tokenizing train dataset (num_proc=64): 91%|█████████ | 107072/117600 [00:12<00:00, 17163.56 examples/s]
Tokenizing train dataset (num_proc=64): 93%|█████████▎| 109248/117600 [00:12<00:00, 18358.71 examples/s]
Tokenizing train dataset (num_proc=64): 94%|█████████▍| 111113/117600 [00:12<00:00, 16723.66 examples/s]
Tokenizing train dataset (num_proc=64): 96%|█████████▌| 112837/117600 [00:12<00:00, 16666.13 examples/s]
Tokenizing train dataset (num_proc=64): 97%|█████████▋| 114545/117600 [00:12<00:00, 16391.81 examples/s]
Tokenizing train dataset (num_proc=64): 99%|█████████▉| 116221/117600 [00:12<00:00, 14387.71 examples/s]
Tokenizing train dataset (num_proc=64): 100%|██████████| 117600/117600 [00:13<00:00, 8406.56 examples/s]
Extracting prompt in eval dataset (num_proc=64): 0%| | 0/2400 [00:00<?, ? examples/s]
Extracting prompt in eval dataset (num_proc=64): 2%|▏ | 38/2400 [00:04<04:09, 9.45 examples/s]
Extracting prompt in eval dataset (num_proc=64): 38%|███▊ | 912/2400 [00:04<00:04, 305.87 examples/s]
Extracting prompt in eval dataset (num_proc=64): 57%|█████▋ | 1358/2400 [00:04<00:02, 497.60 examples/s]
Extracting prompt in eval dataset (num_proc=64): 99%|█████████▉| 2384/2400 [00:04<00:00, 1096.89 examples/s]
Extracting prompt in eval dataset (num_proc=64): 100%|██████████| 2400/2400 [00:05<00:00, 473.42 examples/s]
Applying chat template to eval dataset (num_proc=64): 0%| | 0/2400 [00:00<?, ? examples/s]
Applying chat template to eval dataset (num_proc=64): 2%|▏ | 38/2400 [00:04<04:48, 8.19 examples/s]
Applying chat template to eval dataset (num_proc=64): 13%|█▎ | 304/2400 [00:04<00:24, 87.11 examples/s]
Applying chat template to eval dataset (num_proc=64): 19%|█▉ | 456/2400 [00:04<00:13, 139.63 examples/s]
Applying chat template to eval dataset (num_proc=64): 24%|██▍ | 570/2400 [00:05<00:09, 187.71 examples/s]
Applying chat template to eval dataset (num_proc=64): 28%|██▊ | 684/2400 [00:05<00:06, 249.29 examples/s]
Applying chat template to eval dataset (num_proc=64): 35%|███▍ | 836/2400 [00:05<00:04, 343.99 examples/s]
Applying chat template to eval dataset (num_proc=64): 41%|████ | 988/2400 [00:05<00:03, 449.88 examples/s]
Applying chat template to eval dataset (num_proc=64): 46%|████▌ | 1102/2400 [00:05<00:02, 532.72 examples/s]
Applying chat template to eval dataset (num_proc=64): 51%|█████ | 1215/2400 [00:05<00:01, 594.53 examples/s]
Applying chat template to eval dataset (num_proc=64): 55%|█████▌ | 1326/2400 [00:05<00:01, 673.73 examples/s]
Applying chat template to eval dataset (num_proc=64): 60%|█████▉ | 1437/2400 [00:05<00:01, 724.88 examples/s]
Applying chat template to eval dataset (num_proc=64): 68%|██████▊ | 1623/2400 [00:06<00:00, 849.73 examples/s]
Applying chat template to eval dataset (num_proc=64): 74%|███████▍ | 1771/2400 [00:06<00:00, 920.90 examples/s]
Applying chat template to eval dataset (num_proc=64): 78%|███████▊ | 1882/2400 [00:06<00:00, 909.02 examples/s]
Applying chat template to eval dataset (num_proc=64): 83%|████████▎ | 1993/2400 [00:06<00:00, 893.63 examples/s]
Applying chat template to eval dataset (num_proc=64): 88%|████████▊ | 2104/2400 [00:06<00:00, 832.70 examples/s]
Applying chat template to eval dataset (num_proc=64): 94%|█████████▍| 2252/2400 [00:06<00:00, 947.02 examples/s]
Applying chat template to eval dataset (num_proc=64): 98%|█████████▊| 2363/2400 [00:06<00:00, 923.97 examples/s]
Applying chat template to eval dataset (num_proc=64): 100%|██████████| 2400/2400 [00:07<00:00, 314.06 examples/s]
Tokenizing eval dataset (num_proc=64): 0%| | 0/2400 [00:00<?, ? examples/s]
Tokenizing eval dataset (num_proc=64): 1%| | 27/2400 [00:04<06:33, 6.03 examples/s]
Tokenizing eval dataset (num_proc=64): 6%|▌ | 139/2400 [00:04<00:55, 40.38 examples/s]
Tokenizing eval dataset (num_proc=64): 13%|█▎ | 304/2400 [00:04<00:19, 107.24 examples/s]
Tokenizing eval dataset (num_proc=64): 19%|█▉ | 456/2400 [00:04<00:10, 184.35 examples/s]
Tokenizing eval dataset (num_proc=64): 25%|██▍ | 590/2400 [00:04<00:06, 260.41 examples/s]
Tokenizing eval dataset (num_proc=64): 29%|██▉ | 703/2400 [00:05<00:05, 332.95 examples/s]
Tokenizing eval dataset (num_proc=64): 34%|███▍ | 813/2400 [00:05<00:03, 403.40 examples/s]
Tokenizing eval dataset (num_proc=64): 38%|███▊ | 915/2400 [00:05<00:03, 472.38 examples/s]
Tokenizing eval dataset (num_proc=64): 43%|████▎ | 1028/2400 [00:05<00:02, 543.68 examples/s]
Tokenizing eval dataset (num_proc=64): 49%|████▊ | 1169/2400 [00:05<00:01, 686.31 examples/s]
Tokenizing eval dataset (num_proc=64): 53%|█████▎ | 1271/2400 [00:05<00:01, 682.30 examples/s]
Tokenizing eval dataset (num_proc=64): 58%|█████▊ | 1382/2400 [00:05<00:01, 715.85 examples/s]
Tokenizing eval dataset (num_proc=64): 62%|██████▏ | 1495/2400 [00:06<00:01, 763.53 examples/s]
Tokenizing eval dataset (num_proc=64): 67%|██████▋ | 1604/2400 [00:06<00:01, 786.20 examples/s]
Tokenizing eval dataset (num_proc=64): 72%|███████▏ | 1724/2400 [00:06<00:00, 831.67 examples/s]
Tokenizing eval dataset (num_proc=64): 76%|███████▋ | 1833/2400 [00:06<00:00, 818.78 examples/s]
Tokenizing eval dataset (num_proc=64): 82%|████████▏ | 1956/2400 [00:06<00:00, 862.33 examples/s]
Tokenizing eval dataset (num_proc=64): 85%|████████▌ | 2050/2400 [00:06<00:00, 854.63 examples/s]
Tokenizing eval dataset (num_proc=64): 89%|████████▉ | 2141/2400 [00:06<00:00, 824.46 examples/s]
Tokenizing eval dataset (num_proc=64): 94%|█████████▍| 2264/2400 [00:06<00:00, 856.25 examples/s]
Tokenizing eval dataset (num_proc=64): 100%|█████████▉| 2392/2400 [00:07<00:00, 951.44 examples/s]
Tokenizing eval dataset (num_proc=64): 100%|██████████| 2400/2400 [00:07<00:00, 311.91 examples/s]
The tokenizer has new PAD/BOS/EOS tokens that differ from the model config and generation config. The model config and generation config were aligned accordingly, being updated with the tokenizer's values. Updated tokens: {'bos_token_id': 2}.
==((====))== Unsloth - 2x faster free finetuning | Num GPUs used = 1
\\ /| Num examples = 117,600 | Num Epochs = 3 | Total steps = 2,757
O^O/ \_/ \ Batch size per device = 16 | Gradient accumulation steps = 8
\ / Data Parallel GPUs = 1 | Total batch size (16 x 8 x 1) = 128
"-____-" Trainable parameters = 9,768,960 of 144,283,968 (6.77% trained)
0%| | 0/2757 [00:00<?, ?it/s]`use_return_dict` is deprecated! Use `return_dict` instead!
0%| | 1/2757 [00:21<16:28:47, 21.53s/it]Unsloth: Will smartly offload gradients to save VRAM!
Unsloth: Double buffering enabled (parallel H2D + compute) for backward pass.
Traceback (most recent call last):
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/train_preference.py", line 219, in <module>
main()
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/train_preference.py", line 204, in main
trainer.train()
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/unsloth_compiled_cache/UnslothDPOTrainer.py", line 84, in wrapper
output = f(self, *args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/transformers/trainer.py", line 1424, in train
return inner_training_loop(
^^^^^^^^^^^^^^^^^^^^
File "<string>", line 83, in _fast_inner_training_loop
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/transformers/trainer.py", line 1734, in _run_epoch
tr_loss_step = self.training_step(model, inputs, num_items_in_batch)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "<string>", line 40, in _unsloth_training_step
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/unsloth_compiled_cache/UnslothDPOTrainer.py", line 2539, in compute_loss
loss, metrics = self.get_batch_loss_metrics(model, inputs, train_eval="train")
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/unsloth_compiled_cache/UnslothDPOTrainer.py", line 2462, in get_batch_loss_metrics
ref_chosen_logps, ref_rejected_logps = self.compute_ref_log_probs(batch)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/unsloth_compiled_cache/UnslothDPOTrainer.py", line 1635, in compute_ref_log_probs
ref_model_output = self.concatenated_forward(self.model, batch, is_ref_model=True)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/unsloth_compiled_cache/UnslothDPOTrainer.py", line 2329, in concatenated_forward
outputs = model(input_ids, **model_kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/torch/nn/modules/module.py", line 1776, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/torch/nn/modules/module.py", line 1787, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/accelerate/utils/operations.py", line 823, in forward
return model_forward(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/accelerate/utils/operations.py", line 811, in __call__
return convert_to_fp32(self.model_forward(*args, **kwargs))
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/accelerate/utils/operations.py", line 790, in convert_to_fp32
return recursively_apply(_convert_to_fp32, tensor, test_type=_is_fp16_bf16_tensor)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/accelerate/utils/operations.py", line 120, in recursively_apply
k: recursively_apply(
^^^^^^^^^^^^^^^^^^
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/accelerate/utils/operations.py", line 127, in recursively_apply
return func(data, *args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/accelerate/utils/operations.py", line 782, in _convert_to_fp32
return tensor.float()
^^^^^^^^^^^^^^
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 9.77 GiB. GPU 0 has a total capacity of 23.56 GiB of which 6.23 GiB is free. Process 265634 has 17.31 GiB memory in use. Of the allocated memory 16.87 GiB is allocated by PyTorch, and 98.96 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
0%| | 1/2757 [00:25<19:08:41, 25.01s/it]