=== utc === Thu Jul 30 09:33:21 UTC 2026 === /workspace === total 20585 drwxrwxrwx 11 root root 3008589 Jul 30 09:26 . drwxr-xr-x 1 root root 85 Jul 30 09:33 .. drwxrwxrwx 4 root root 2027870 Jul 27 23:24 .cache drwxrwxrwx 2 root root 1 Jul 28 00:50 .persist drwxrwxrwx 4 root root 1000359 Jul 27 22:35 _shards drwxrwxrwx 3 root root 3004667 Jul 30 09:25 checkpoints drwxrwxrwx 3 root root 3003766 Jul 27 22:35 data drwxrwxrwx 2 root root 2095794 Jul 30 09:26 hf-500m drwxrwxrwx 31 root root 2023621 Jul 30 09:27 llama.cpp drwxrwxrwx 11 root root 2006109 Jul 28 00:50 sprocket -rw-rw-rw- 1 root root 806956 Jul 30 09:30 train.log drwxrwxrwx 11 root root 2006034 Jul 27 23:16 tune -rw-rw-rw- 1 root root 58167 Jul 27 23:24 tune.log -rw-rw-rw- 1 root root 30155 Jul 27 23:56 tune2.log -rw-rw-rw- 1 root root 2633 Jul 27 23:56 tune_results.json -rw-rw-rw- 1 root root 0 Jul 27 22:34 watchdog.log === data/processed === total 39505306 drwxrwxrwx 2 root root 3003766 Jul 30 09:22 . drwxrwxrwx 3 root root 3003766 Jul 27 22:35 .. -rw-rw-rw- 1 root root 41615850 Jul 30 09:22 sft_packed.npz -rw-rw-rw- 1 root root 402804266 Jul 27 22:42 train.bin -rw-rw-rw- 1 root root 120 Jul 27 22:42 train.bin.manifest.json -rw-rw-rw- 1 root root 40003002080 Jul 28 03:09 train_sample_100BT.bin -rw-rw-rw- 1 root root 1062 Jul 28 03:09 train_sample_100BT.bin.manifest.json === checkpoints === total 33287788 drwxrwxrwx 3 root root 3004667 Jul 30 09:25 . drwxrwxrwx 11 root root 3008589 Jul 30 09:26 .. -rw-rw-rw- 1 root root 6013745939 Jul 30 08:58 500m_151500.pt -rw-rw-rw- 1 root root 6013747091 Jul 30 09:09 500m_152000.pt -rw-rw-rw- 1 root root 6013748179 Jul 30 09:20 500m_152500.pt -rw-rw-rw- 1 root root 2004810907 Jul 30 09:22 500m_final.pt -rw-rw-rw- 1 root root 6013748435 Jul 30 09:22 500m_latest.pt -rw-rw-rw- 1 root root 2004468011 Jul 30 09:25 500m_sft_final.pt -rw-rw-rw- 1 root root 6013408307 Jul 30 09:25 500m_sft_latest.pt drwxrwxrwx 2 root root 3001493 Jul 28 00:50 _archived_20260728_005036 === shards scratch === total 4893 drwxrwxrwx 4 root root 1000359 Jul 27 22:35 . drwxrwxrwx 11 root root 3008589 Jul 30 09:26 .. drwxrwxrwx 3 root root 1000359 Jul 27 22:35 .cache drwxrwxrwx 4 root root 1 Jul 28 03:07 sample === disk === Filesystem Size Used Avail Use% Mounted on mfs#us-ne-1.runpod.net:9421 685T 448T 238T 66% /workspace === watchdog.log === === train.log (first 60) === [22:34:48] billing guard armed: AUTO_STOP=1, watchdog 2h (pid 103) [22:34:48] repo present; pulling Already up to date. [22:34:49] installing deps WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager, possibly rendering your system unusable.It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv. Use the --root-user-action option if you know what you are doing and want to suppress this warning. [notice] A new release of pip is available: 24.2 -> 26.1.2 [notice] To update, run: python -m pip install --upgrade pip torch 2.4.1+cu124 cuda=True NVIDIA H100 80GB HBM3 VRAM 85.0 GB total, 84.5 GB free [22:35:02] building corpus: target 200000000 tokens from sample/10BT (resumable) ========================================================================== BUILD CORPUS HuggingFaceFW/fineweb-edu:sample/10BT -> data/processed/train.bin ========================================================================== target 200.0M tokens (our 32k tokenizer) 14 shards in subset, 14 remaining [ 1/14] 000_00000.parquet 194,000 docs 201.4M tok in 444s (dl 8s) | total 201.4M (100.7%) | 453.6K tok/s | ETA -0.0h DONE: 201.4M tokens -> data/processed/train.bin (0.40 GB) shards used: 1 train with: py -m src.train.train --preset 500m --data data/processed/train.bin --loader memmap --resume auto [22:42:27] pretraining 500m for 200 steps (ctx 1024, mb 24 x2) corpus 201,402,133 tokens (train 200,395,123 / val 1,007,010) — memmap (0.40 GB, streamed) model 500m: 501.1M params starting fresh (no --resume) iter 0 | train 10.642 | val 9.907 | lr 2.00e-04 | 43,180 tok/s | 62.72 GB | 0.000B tok iter 1: 69,762 tok/s (0.70s/iter) iter 2: 69,721 tok/s (0.70s/iter) iter 3: 69,702 tok/s (0.71s/iter) iter 4: 69,682 tok/s (0.71s/iter) iter 5: 69,681 tok/s (0.71s/iter) iter 20 | train 7.722 | val 7.769 | lr 5.90e-04 | 58,491 tok/s | 66.73 GB | 0.001B tok iter 40 | train 7.651 | val 7.636 | lr 5.54e-04 | 61,657 tok/s | 66.73 GB | 0.002B tok iter 60 | train 7.212 | val 7.226 | lr 4.96e-04 | 62,815 tok/s | 66.73 GB | 0.003B tok iter 80 | train 6.893 | val 6.987 | lr 4.21e-04 | 63,415 tok/s | 66.73 GB | 0.004B tok iter 100 | train 6.764 | val 6.815 | lr 3.36e-04 | 62,773 tok/s | 66.73 GB | 0.005B tok iter 120 | train 6.656 | val 6.669 | lr 2.51e-04 | 63,172 tok/s | 66.73 GB | 0.006B tok iter 140 | train 6.528 | val 6.553 | lr 1.74e-04 | 63,460 tok/s | 66.73 GB | 0.007B tok iter 160 | train 6.468 | val 6.521 | lr 1.13e-04 | 63,678 tok/s | 66.73 GB | 0.008B tok iter 180 | train 6.483 | val 6.457 | lr 7.36e-05 | 63,286 tok/s | 66.73 GB | 0.009B tok iter 200 | train 6.416 | val 6.422 | lr 6.00e-05 | 63,470 tok/s | 66.73 GB | 0.010B tok training complete [22:45:30] SFT from checkpoints/500m_final.pt rendering 18,151 conversations -> ctx 1024 blocks... rendered 5,000/18,151 rendered 10,000/18,151 rendered 15,000/18,151 packed 5,079 blocks (0 dropped), cached -> data/processed/sft_packed.npz corpus: 5,079 blocks x 1024 = 5,200,896 tokens, 4,567,314 trainable (87.8%) train 4,978 blocks / val 101 blocks 155 steps/epoch x 3.0 epochs = 465 steps (warmup 13) model 500m: 501.1M params step 0/465 | train 10.626 | val 10.544 | lr 1.54e-06 | ep 0.00 | 43.54 GB step 25/465 | train 7.857 | val 7.738 | lr 2.00e-05 | ep 0.16 | 47.55 GB step 50/465 | train 6.971 | val 6.907 | lr 1.97e-05 | ep 0.32 | 47.55 GB step 75/465 | train 6.322 | val 6.407 | lr 1.92e-05 | ep 0.48 | 47.55 GB step 100/465 | train 6.125 | val 6.158 | lr 1.84e-05 | ep 0.65 | 47.55 GB step 125/465 | train 5.935 | val 6.009 | lr 1.74e-05 | ep 0.81 | 47.55 GB === train.log (last 400) === ...hf-500m/model.safetensors: 7%|▋ | 66.8MB / 1.00GB  Processing Files (0 / 1) : 7%|▋ | 66.8MB / 1.00GB, 6.40MB/s New Data Upload : 50%|████▉ | 66.8MB / 134MB, 6.40MB/s  ...hf-500m/model.safetensors: 7%|▋ | 70.5MB / 1.00GB  Processing Files (0 / 1) : 7%|▋ | 70.5MB / 1.00GB, 6.60MB/s New Data Upload : 35%|███▌ | 70.5MB / 201MB, 6.60MB/s  ...hf-500m/model.safetensors: 13%|█▎ | 134MB / 1.00GB  Processing Files (0 / 1) : 13%|█▎ | 134MB / 1.00GB, 12.5MB/s New Data Upload : 50%|█████ | 134MB / 268MB, 12.5MB/s  ...hf-500m/model.safetensors: 19%|█▉ | 190MB / 1.00GB  Processing Files (0 / 1) : 19%|█▉ | 190MB / 1.00GB, 17.5MB/s New Data Upload : 57%|█████▋ | 190MB / 335MB, 17.5MB/s  ...hf-500m/model.safetensors: 21%|██ | 208MB / 1.00GB  Processing Files (0 / 1) : 21%|██ | 208MB / 1.00GB, 18.8MB/s New Data Upload : 62%|██████▏ | 208MB / 335MB, 18.8MB/s  ...hf-500m/model.safetensors: 33%|███▎ | 335MB / 1.00GB  Processing Files (0 / 1) : 33%|███▎ | 335MB / 1.00GB, 30.2MB/s New Data Upload : 83%|████████▎ | 335MB / 402MB, 30.2MB/s  ...hf-500m/model.safetensors: 33%|███▎ | 335MB / 1.00GB  ...hf-500m/model.safetensors: 34%|███▎ | 338MB / 1.00GB  Processing Files (0 / 1) : 34%|███▎ | 338MB / 1.00GB, 29.0MB/s New Data Upload : 72%|███████▏ | 338MB / 469MB, 29.0MB/s  ...hf-500m/model.safetensors: 47%|████▋ | 469MB / 1.00GB  Processing Files (0 / 1) : 47%|████▋ | 469MB / 1.00GB, 40.4MB/s New Data Upload : 87%|████████▋ | 469MB / 536MB, 40.4MB/s  ...hf-500m/model.safetensors: 53%|█████▎ | 536MB / 1.00GB  Processing Files (0 / 1) : 53%|█████▎ | 536MB / 1.00GB, 45.6MB/s New Data Upload : 89%|████████▉ | 536MB / 604MB, 45.6MB/s  ...hf-500m/model.safetensors: 60%|██████ | 603MB / 1.00GB  Processing Files (0 / 1) : 60%|██████ | 603MB / 1.00GB, 50.6MB/s New Data Upload : 90%|████████▉ | 603MB / 671MB, 50.6MB/s  ...hf-500m/model.safetensors: 67%|██████▋ | 670MB / 1.00GB  Processing Files (0 / 1) : 67%|██████▋ | 670MB / 1.00GB, 54.4MB/s New Data Upload : 91%|█████████ | 670MB / 738MB, 54.4MB/s  ...hf-500m/model.safetensors: 67%|██████▋ | 670MB / 1.00GB  Processing Files (0 / 1) : 67%|██████▋ | 670MB / 1.00GB, 53.2MB/s New Data Upload : 91%|█████████ | 670MB / 738MB, 53.2MB/s  ...hf-500m/model.safetensors: 74%|███████▎ | 737MB / 1.00GB  Processing Files (0 / 1) : 74%|███████▎ | 737MB / 1.00GB, 58.0MB/s New Data Upload : 92%|█████████▏| 737MB / 805MB, 58.0MB/s  ...hf-500m/model.safetensors: 80%|████████ | 805MB / 1.00GB  Processing Files (0 / 1) : 80%|████████ | 805MB / 1.00GB, 62.6MB/s New Data Upload : 100%|█████████▉| 805MB / 805MB, 62.6MB/s  ...hf-500m/model.safetensors: 80%|████████ | 805MB / 1.00GB  Processing Files (0 / 1) : 80%|████████ | 805MB / 1.00GB, 61.2MB/s New Data Upload : 92%|█████████▏| 805MB / 872MB, 61.2MB/s  ...hf-500m/model.safetensors: 87%|████████▋ | 872MB / 1.00GB  Processing Files (0 / 1) : 87%|████████▋ | 872MB / 1.00GB, 65.8MB/s New Data Upload : 93%|█████████▎| 872MB / 939MB, 65.8MB/s  ...hf-500m/model.safetensors: 94%|█████████▎| 939MB / 1.00GB  Processing Files (0 / 1) : 94%|█████████▎| 939MB / 1.00GB, 70.2MB/s New Data Upload : 100%|█████████▉| 939MB / 939MB, 70.2MB/s  ...hf-500m/model.safetensors: 94%|█████████▎| 939MB / 1.00GB  ...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB  Processing Files (0 / 1) : 100%|█████████▉| 1.00GB / 1.00GB, 72.6MB/s New Data Upload : 100%|█████████▉| 1.00GB / 1.00GB, 72.6MB/s  ...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB  Processing Files (0 / 1) : 100%|█████████▉| 1.00GB / 1.00GB, 69.5MB/s New Data Upload : 100%|█████████▉| 1.00GB / 1.00GB, 69.5MB/s  ...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB  Processing Files (0 / 1) : 100%|█████████▉| 1.00GB / 1.00GB, 68.0MB/s New Data Upload : 100%|█████████▉| 1.00GB / 1.00GB, 68.0MB/s  ...hf-500m/model.safetensors: 100%|██████████| 1.00GB / 1.00GB  Processing Files (1 / 1) : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s New Data Upload : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s  Processing Files (1 / 1) : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s New Data Upload : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s ...hf-500m/model.safetensors: 100%|██████████| 1.00GB / 1.00GB pushed -> https://huggingface.co/HuggingFace7141/sprocket-500m [09:26:28] converting to GGUF q4_k_m (phone / Ollama / LM Studio) Cloning into 'llama.cpp'... Updating files: 7% (243/3304) Updating files: 8% (265/3304) Updating files: 9% (298/3304) Updating files: 10% (331/3304) Updating files: 11% (364/3304) Updating files: 12% (397/3304) Updating files: 12% (410/3304) Updating files: 13% (430/3304) Updating files: 14% (463/3304) Updating files: 15% (496/3304) Updating files: 16% (529/3304) Updating files: 17% (562/3304) Updating files: 18% (595/3304) Updating files: 19% (628/3304) Updating files: 19% (640/3304) Updating files: 20% (661/3304) Updating files: 21% (694/3304) Updating files: 22% (727/3304) Updating files: 23% (760/3304) Updating files: 24% (793/3304) Updating files: 25% (826/3304) Updating files: 26% (860/3304) Updating files: 26% (884/3304) Updating files: 27% (893/3304) Updating files: 28% (926/3304) Updating files: 29% (959/3304) Updating files: 30% (992/3304) Updating files: 31% (1025/3304) Updating files: 32% (1058/3304) Updating files: 33% (1091/3304) Updating files: 34% (1124/3304) Updating files: 34% (1142/3304) Updating files: 35% (1157/3304) Updating files: 36% (1190/3304) Updating files: 37% (1223/3304) Updating files: 38% (1256/3304) Updating files: 39% (1289/3304) Updating files: 40% (1322/3304) Updating files: 41% (1355/3304) Updating files: 41% (1381/3304) Updating files: 42% (1388/3304) Updating files: 43% (1421/3304) Updating files: 44% (1454/3304) Updating files: 45% (1487/3304) Updating files: 46% (1520/3304) Updating files: 47% (1553/3304) Updating files: 48% (1586/3304) Updating files: 49% (1619/3304) Updating files: 49% (1624/3304) Updating files: 50% (1652/3304) Updating files: 51% (1686/3304) Updating files: 52% (1719/3304) Updating files: 53% (1752/3304) Updating files: 54% (1785/3304) Updating files: 55% (1818/3304) Updating files: 56% (1851/3304) Updating files: 56% (1863/3304) Updating files: 57% (1884/3304) Updating files: 57% (1911/3304) Updating files: 58% (1917/3304) Updating files: 59% (1950/3304) Updating files: 60% (1983/3304) Updating files: 61% (2016/3304) Updating files: 62% (2049/3304) Updating files: 63% (2082/3304) Updating files: 63% (2098/3304) Updating files: 64% (2115/3304) Updating files: 65% (2148/3304) Updating files: 66% (2181/3304) Updating files: 67% (2214/3304) Updating files: 68% (2247/3304) Updating files: 69% (2280/3304) Updating files: 70% (2313/3304) Updating files: 70% (2341/3304) Updating files: 71% (2346/3304) Updating files: 72% (2379/3304) Updating files: 73% (2412/3304) Updating files: 74% (2445/3304) Updating files: 75% (2478/3304) Updating files: 76% (2512/3304) Updating files: 77% (2545/3304) Updating files: 77% (2572/3304) Updating files: 78% (2578/3304) Updating files: 79% (2611/3304) Updating files: 80% (2644/3304) Updating files: 81% (2677/3304) Updating files: 82% (2710/3304) Updating files: 83% (2743/3304) Updating files: 84% (2776/3304) Updating files: 85% (2809/3304) Updating files: 85% (2811/3304) Updating files: 86% (2842/3304) Updating files: 87% (2875/3304) Updating files: 88% (2908/3304) Updating files: 89% (2941/3304) Updating files: 90% (2974/3304) Updating files: 91% (3007/3304) Updating files: 92% (3040/3304) Updating files: 92% (3065/3304) Updating files: 93% (3073/3304) Updating files: 94% (3106/3304) Updating files: 95% (3139/3304) Updating files: 96% (3172/3304) Updating files: 97% (3205/3304) Updating files: 98% (3238/3304) Updating files: 99% (3271/3304) Updating files: 99% (3300/3304) Updating files: 100% (3304/3304) Updating files: 100% (3304/3304), done. INFO:hf-to-gguf:Loading model: hf-500m INFO:hf-to-gguf:Model architecture: LlamaForCausalLM INFO:hf-to-gguf:gguf: indexing model part 'model.safetensors' INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only INFO:hf-to-gguf:Exporting model... INFO:hf-to-gguf:token_embd.weight, torch.bfloat16 --> F16, shape = {1280, 32000} INFO:hf-to-gguf:blk.0.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.0.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.0.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.0.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.0.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.0.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.0.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.0.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.0.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.1.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.1.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.1.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.1.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.1.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.1.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.1.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.1.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.1.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.10.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.10.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.10.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.10.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.10.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.10.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.10.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.10.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.10.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.11.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.11.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.11.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.11.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.11.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.11.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.11.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.11.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.11.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.12.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.12.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.12.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.12.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.12.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.12.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.12.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.12.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.12.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.13.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.13.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.13.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.13.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.13.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.13.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.13.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.13.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.13.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.14.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.14.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.14.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.14.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.14.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.14.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.14.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.14.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.14.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.15.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.15.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.15.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.15.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.15.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.15.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.15.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.15.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.15.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.16.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.16.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.16.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.16.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.16.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.16.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.16.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.16.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.16.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.17.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.17.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.17.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.17.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.17.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.17.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.17.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.17.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.17.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.18.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.18.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.18.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.18.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.18.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.18.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.18.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.18.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.18.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.19.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.19.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.19.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.19.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.19.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.19.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.19.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.19.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.19.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.2.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.2.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.2.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.2.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.2.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.2.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.2.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.2.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.2.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.20.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.20.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.20.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.20.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.20.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.20.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.20.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.20.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.20.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.21.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.21.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.21.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.21.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.21.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.21.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.21.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.21.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.21.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.22.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.22.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.22.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.22.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.22.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.22.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.22.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.22.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.22.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.23.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.23.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.23.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.23.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.23.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.23.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.23.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.23.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.23.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.24.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.24.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.24.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.24.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.24.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.24.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.24.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.24.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.24.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.25.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.25.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.25.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.25.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.25.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.25.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.25.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.25.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.25.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.3.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.3.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.3.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.3.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.3.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.3.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.3.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.3.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.3.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.4.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.4.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.4.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.4.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.4.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.4.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.4.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.4.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.4.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.5.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.5.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.5.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.5.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.5.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.5.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.5.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.5.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.5.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.6.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.6.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.6.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.6.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.6.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.6.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.6.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.6.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.6.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.7.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.7.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.7.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.7.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.7.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.7.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.7.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.7.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.7.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.8.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.8.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.8.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.8.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.8.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.8.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.8.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.8.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.8.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.9.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.9.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280} INFO:hf-to-gguf:blk.9.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.9.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584} INFO:hf-to-gguf:blk.9.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:blk.9.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:blk.9.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.9.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280} INFO:hf-to-gguf:blk.9.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256} INFO:hf-to-gguf:output_norm.weight, torch.bfloat16 --> F32, shape = {1280} INFO:hf-to-gguf:Set meta model INFO:hf-to-gguf:Set model parameters INFO:hf-to-gguf:gguf: context length = 2048 INFO:hf-to-gguf:gguf: embedding length = 1280 INFO:hf-to-gguf:gguf: feed forward length = 3584 INFO:hf-to-gguf:gguf: head count = 20 INFO:hf-to-gguf:gguf: key-value head count = 4 INFO:hf-to-gguf:gguf: rope theta = 10000.0 INFO:hf-to-gguf:gguf: rms norm epsilon = 1e-05 INFO:hf-to-gguf:gguf: file type = 1 INFO:hf-to-gguf:Set model quantization version INFO:hf-to-gguf:Set model tokenizer WARNING:hf-to-gguf: WARNING:hf-to-gguf:************************************************************************************** WARNING:hf-to-gguf:** WARNING: The BPE pre-tokenizer was not recognized! WARNING:hf-to-gguf:** There are 2 possible reasons for this: WARNING:hf-to-gguf:** - the model has not been added to convert_hf_to_gguf_update.py yet WARNING:hf-to-gguf:** - the pre-tokenization config has changed upstream WARNING:hf-to-gguf:** Check your model files and convert_hf_to_gguf_update.py and update them accordingly. WARNING:hf-to-gguf:** ref: https://github.com/ggml-org/llama.cpp/pull/6920 WARNING:hf-to-gguf:** WARNING:hf-to-gguf:** chkhsh: 50cc433d4d528ad898dc1178e17a479f50c03d2c883d07494ae21da32d3eed4f WARNING:hf-to-gguf:************************************************************************************** WARNING:hf-to-gguf: Traceback (most recent call last): File "/workspace/llama.cpp/conversion/llama.py", line 134, in set_vocab self._set_vocab_sentencepiece() File "/workspace/llama.cpp/conversion/base.py", line 1831, in _set_vocab_sentencepiece tokens, scores, toktypes = self._create_vocab_sentencepiece() ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/llama.cpp/conversion/base.py", line 1848, in _create_vocab_sentencepiece raise FileNotFoundError(f"File not found: {tokenizer_path}") FileNotFoundError: File not found: /workspace/hf-500m/tokenizer.model During handling of the above exception, another exception occurred: Traceback (most recent call last): File "/workspace/llama.cpp/conversion/llama.py", line 137, in set_vocab self._set_vocab_llama_hf() File "/workspace/llama.cpp/conversion/base.py", line 1933, in _set_vocab_llama_hf vocab = gguf.LlamaHfVocab(self.dir_model) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/llama.cpp/gguf-py/gguf/vocab.py", line 571, in __init__ raise FileNotFoundError('Cannot find Llama BPE tokenizer') FileNotFoundError: Cannot find Llama BPE tokenizer During handling of the above exception, another exception occurred: Traceback (most recent call last): File "/workspace/llama.cpp/convert_hf_to_gguf.py", line 296, in main() File "/workspace/llama.cpp/convert_hf_to_gguf.py", line 290, in main model_instance.write() File "/workspace/llama.cpp/conversion/base.py", line 1026, in write self.prepare_metadata(vocab_only=False) File "/workspace/llama.cpp/conversion/base.py", line 1193, in prepare_metadata self.set_vocab() File "/workspace/llama.cpp/conversion/llama.py", line 140, in set_vocab self._set_vocab_gpt2() File "/workspace/llama.cpp/conversion/base.py", line 1714, in _set_vocab_gpt2 tokens, toktypes, tokpre = self.get_vocab_base() ^^^^^^^^^^^^^^^^^^^^^ File "/workspace/llama.cpp/conversion/base.py", line 1351, in get_vocab_base tokpre = self.get_vocab_base_pre(tokenizer) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/llama.cpp/conversion/base.py", line 1702, in get_vocab_base_pre raise NotImplementedError("BPE pre-tokenizer was not recognized - update get_vocab_base_pre()") NotImplementedError: BPE pre-tokenizer was not recognized - update get_vocab_base_pre() llama_print_build_info: build = 1 (b2f2216) llama_print_build_info: built with GNU 11.4.0 for Linux x86_64 llama_quantize: quantizing '/workspace/sprocket-500m-f16.gguf' to '/workspace/sprocket-500m-q4_k_m.gguf' as Q4_K_M gguf_init_from_file: failed to open GGUF file '/workspace/sprocket-500m-f16.gguf' (No such file or directory) llama_model_quantize: failed to quantize: llama_model_loader: failed to load model from /workspace/sprocket-500m-f16.gguf llama_quantize: failed to quantize model from '/workspace/sprocket-500m-f16.gguf' ls: cannot access '/workspace/sprocket-*.gguf': No such file or directory [09:30:02] DONE. [09:30:02] base ckpt : /workspace/checkpoints/500m_final.pt [09:30:02] sft ckpt : /workspace/checkpoints/500m_sft_final.pt [09:30:02] hf model : /workspace/hf-500m [09:30:02] gguf : /workspace/sprocket-500m-q4_k_m.gguf [09:30:02] pipeline finished cleanly [09:30:02] TERMINATING POD - exit code 0 Runpod config file not found, please run `runpodctl config` to create it pod "jasakfns6809ft" removed