Files
sprocket-500m/debug/status.txt
ModelHub XC 50786d48b0 初始化项目,由ModelHub XC社区提供模型
Model: kandivault/sprocket-500m
Source: Original Platform
2026-09-17 05:03:23 +08:00

522 lines
42 KiB
Plaintext
Raw Permalink Blame History

This file contains invisible Unicode characters

This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

=== utc ===
Thu Jul 30 09:33:21 UTC 2026
=== /workspace ===
total 20585
drwxrwxrwx 11 root root 3008589 Jul 30 09:26 .
drwxr-xr-x 1 root root 85 Jul 30 09:33 ..
drwxrwxrwx 4 root root 2027870 Jul 27 23:24 .cache
drwxrwxrwx 2 root root 1 Jul 28 00:50 .persist
drwxrwxrwx 4 root root 1000359 Jul 27 22:35 _shards
drwxrwxrwx 3 root root 3004667 Jul 30 09:25 checkpoints
drwxrwxrwx 3 root root 3003766 Jul 27 22:35 data
drwxrwxrwx 2 root root 2095794 Jul 30 09:26 hf-500m
drwxrwxrwx 31 root root 2023621 Jul 30 09:27 llama.cpp
drwxrwxrwx 11 root root 2006109 Jul 28 00:50 sprocket
-rw-rw-rw- 1 root root 806956 Jul 30 09:30 train.log
drwxrwxrwx 11 root root 2006034 Jul 27 23:16 tune
-rw-rw-rw- 1 root root 58167 Jul 27 23:24 tune.log
-rw-rw-rw- 1 root root 30155 Jul 27 23:56 tune2.log
-rw-rw-rw- 1 root root 2633 Jul 27 23:56 tune_results.json
-rw-rw-rw- 1 root root 0 Jul 27 22:34 watchdog.log
=== data/processed ===
total 39505306
drwxrwxrwx 2 root root 3003766 Jul 30 09:22 .
drwxrwxrwx 3 root root 3003766 Jul 27 22:35 ..
-rw-rw-rw- 1 root root 41615850 Jul 30 09:22 sft_packed.npz
-rw-rw-rw- 1 root root 402804266 Jul 27 22:42 train.bin
-rw-rw-rw- 1 root root 120 Jul 27 22:42 train.bin.manifest.json
-rw-rw-rw- 1 root root 40003002080 Jul 28 03:09 train_sample_100BT.bin
-rw-rw-rw- 1 root root 1062 Jul 28 03:09 train_sample_100BT.bin.manifest.json
=== checkpoints ===
total 33287788
drwxrwxrwx 3 root root 3004667 Jul 30 09:25 .
drwxrwxrwx 11 root root 3008589 Jul 30 09:26 ..
-rw-rw-rw- 1 root root 6013745939 Jul 30 08:58 500m_151500.pt
-rw-rw-rw- 1 root root 6013747091 Jul 30 09:09 500m_152000.pt
-rw-rw-rw- 1 root root 6013748179 Jul 30 09:20 500m_152500.pt
-rw-rw-rw- 1 root root 2004810907 Jul 30 09:22 500m_final.pt
-rw-rw-rw- 1 root root 6013748435 Jul 30 09:22 500m_latest.pt
-rw-rw-rw- 1 root root 2004468011 Jul 30 09:25 500m_sft_final.pt
-rw-rw-rw- 1 root root 6013408307 Jul 30 09:25 500m_sft_latest.pt
drwxrwxrwx 2 root root 3001493 Jul 28 00:50 _archived_20260728_005036
=== shards scratch ===
total 4893
drwxrwxrwx 4 root root 1000359 Jul 27 22:35 .
drwxrwxrwx 11 root root 3008589 Jul 30 09:26 ..
drwxrwxrwx 3 root root 1000359 Jul 27 22:35 .cache
drwxrwxrwx 4 root root 1 Jul 28 03:07 sample
=== disk ===
Filesystem Size Used Avail Use% Mounted on
mfs#us-ne-1.runpod.net:9421 685T 448T 238T 66% /workspace
=== watchdog.log ===
=== train.log (first 60) ===
[22:34:48] billing guard armed: AUTO_STOP=1, watchdog 2h (pid 103)
[22:34:48] repo present; pulling
Already up to date.
[22:34:49] installing deps
WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager, possibly rendering your system unusable.It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv. Use the --root-user-action option if you know what you are doing and want to suppress this warning.
[notice] A new release of pip is available: 24.2 -> 26.1.2
[notice] To update, run: python -m pip install --upgrade pip
torch 2.4.1+cu124 cuda=True NVIDIA H100 80GB HBM3
VRAM 85.0 GB total, 84.5 GB free
[22:35:02] building corpus: target 200000000 tokens from sample/10BT (resumable)
==========================================================================
BUILD CORPUS HuggingFaceFW/fineweb-edu:sample/10BT -> data/processed/train.bin
==========================================================================
target 200.0M tokens (our 32k tokenizer)
14 shards in subset, 14 remaining
[ 1/14] 000_00000.parquet 194,000 docs 201.4M tok in 444s (dl 8s) | total 201.4M (100.7%) | 453.6K tok/s | ETA -0.0h
DONE: 201.4M tokens -> data/processed/train.bin (0.40 GB)
shards used: 1
train with: py -m src.train.train --preset 500m --data data/processed/train.bin --loader memmap --resume auto
[22:42:27] pretraining 500m for 200 steps (ctx 1024, mb 24 x2)
corpus 201,402,133 tokens (train 200,395,123 / val 1,007,010) — memmap (0.40 GB, streamed)
model 500m: 501.1M params
starting fresh (no --resume)
iter 0 | train 10.642 | val 9.907 | lr 2.00e-04 | 43,180 tok/s | 62.72 GB | 0.000B tok
iter 1: 69,762 tok/s (0.70s/iter)
iter 2: 69,721 tok/s (0.70s/iter)
iter 3: 69,702 tok/s (0.71s/iter)
iter 4: 69,682 tok/s (0.71s/iter)
iter 5: 69,681 tok/s (0.71s/iter)
iter 20 | train 7.722 | val 7.769 | lr 5.90e-04 | 58,491 tok/s | 66.73 GB | 0.001B tok
iter 40 | train 7.651 | val 7.636 | lr 5.54e-04 | 61,657 tok/s | 66.73 GB | 0.002B tok
iter 60 | train 7.212 | val 7.226 | lr 4.96e-04 | 62,815 tok/s | 66.73 GB | 0.003B tok
iter 80 | train 6.893 | val 6.987 | lr 4.21e-04 | 63,415 tok/s | 66.73 GB | 0.004B tok
iter 100 | train 6.764 | val 6.815 | lr 3.36e-04 | 62,773 tok/s | 66.73 GB | 0.005B tok
iter 120 | train 6.656 | val 6.669 | lr 2.51e-04 | 63,172 tok/s | 66.73 GB | 0.006B tok
iter 140 | train 6.528 | val 6.553 | lr 1.74e-04 | 63,460 tok/s | 66.73 GB | 0.007B tok
iter 160 | train 6.468 | val 6.521 | lr 1.13e-04 | 63,678 tok/s | 66.73 GB | 0.008B tok
iter 180 | train 6.483 | val 6.457 | lr 7.36e-05 | 63,286 tok/s | 66.73 GB | 0.009B tok
iter 200 | train 6.416 | val 6.422 | lr 6.00e-05 | 63,470 tok/s | 66.73 GB | 0.010B tok
training complete
[22:45:30] SFT from checkpoints/500m_final.pt
rendering 18,151 conversations -> ctx 1024 blocks...
rendered 5,000/18,151
rendered 10,000/18,151
rendered 15,000/18,151
packed 5,079 blocks (0 dropped), cached -> data/processed/sft_packed.npz
corpus: 5,079 blocks x 1024 = 5,200,896 tokens, 4,567,314 trainable (87.8%)
train 4,978 blocks / val 101 blocks
155 steps/epoch x 3.0 epochs = 465 steps (warmup 13)
model 500m: 501.1M params
step 0/465 | train 10.626 | val 10.544 | lr 1.54e-06 | ep 0.00 | 43.54 GB
step 25/465 | train 7.857 | val 7.738 | lr 2.00e-05 | ep 0.16 | 47.55 GB
step 50/465 | train 6.971 | val 6.907 | lr 1.97e-05 | ep 0.32 | 47.55 GB
step 75/465 | train 6.322 | val 6.407 | lr 1.92e-05 | ep 0.48 | 47.55 GB
step 100/465 | train 6.125 | val 6.158 | lr 1.84e-05 | ep 0.65 | 47.55 GB
step 125/465 | train 5.935 | val 6.009 | lr 1.74e-05 | ep 0.81 | 47.55 GB
=== train.log (last 400) ===
...hf-500m/model.safetensors: 7%|▋ | 66.8MB / 1.00GB 
Processing Files (0 / 1) : 7%|▋ | 66.8MB / 1.00GB, 6.40MB/s
New Data Upload : 50%|████▉ | 66.8MB / 134MB, 6.40MB/s 
...hf-500m/model.safetensors: 7%|▋ | 70.5MB / 1.00GB 
Processing Files (0 / 1) : 7%|▋ | 70.5MB / 1.00GB, 6.60MB/s
New Data Upload : 35%|███▌ | 70.5MB / 201MB, 6.60MB/s 
...hf-500m/model.safetensors: 13%|█▎ | 134MB / 1.00GB 
Processing Files (0 / 1) : 13%|█▎ | 134MB / 1.00GB, 12.5MB/s
New Data Upload : 50%|█████ | 134MB / 268MB, 12.5MB/s 
...hf-500m/model.safetensors: 19%|█▉ | 190MB / 1.00GB 
Processing Files (0 / 1) : 19%|█▉ | 190MB / 1.00GB, 17.5MB/s
New Data Upload : 57%|█████▋ | 190MB / 335MB, 17.5MB/s 
...hf-500m/model.safetensors: 21%|██ | 208MB / 1.00GB 
Processing Files (0 / 1) : 21%|██ | 208MB / 1.00GB, 18.8MB/s
New Data Upload : 62%|██████▏ | 208MB / 335MB, 18.8MB/s 
...hf-500m/model.safetensors: 33%|███▎ | 335MB / 1.00GB 
Processing Files (0 / 1) : 33%|███▎ | 335MB / 1.00GB, 30.2MB/s
New Data Upload : 83%|████████▎ | 335MB / 402MB, 30.2MB/s 
...hf-500m/model.safetensors: 33%|███▎ | 335MB / 1.00GB 
...hf-500m/model.safetensors: 34%|███▎ | 338MB / 1.00GB 
Processing Files (0 / 1) : 34%|███▎ | 338MB / 1.00GB, 29.0MB/s
New Data Upload : 72%|███████▏ | 338MB / 469MB, 29.0MB/s 
...hf-500m/model.safetensors: 47%|████▋ | 469MB / 1.00GB 
Processing Files (0 / 1) : 47%|████▋ | 469MB / 1.00GB, 40.4MB/s
New Data Upload : 87%|████████▋ | 469MB / 536MB, 40.4MB/s 
...hf-500m/model.safetensors: 53%|█████▎ | 536MB / 1.00GB 
Processing Files (0 / 1) : 53%|█████▎ | 536MB / 1.00GB, 45.6MB/s
New Data Upload : 89%|████████▉ | 536MB / 604MB, 45.6MB/s 
...hf-500m/model.safetensors: 60%|██████ | 603MB / 1.00GB 
Processing Files (0 / 1) : 60%|██████ | 603MB / 1.00GB, 50.6MB/s
New Data Upload : 90%|████████▉ | 603MB / 671MB, 50.6MB/s 
...hf-500m/model.safetensors: 67%|██████▋ | 670MB / 1.00GB 
Processing Files (0 / 1) : 67%|██████▋ | 670MB / 1.00GB, 54.4MB/s
New Data Upload : 91%|█████████ | 670MB / 738MB, 54.4MB/s 
...hf-500m/model.safetensors: 67%|██████▋ | 670MB / 1.00GB 
Processing Files (0 / 1) : 67%|██████▋ | 670MB / 1.00GB, 53.2MB/s
New Data Upload : 91%|█████████ | 670MB / 738MB, 53.2MB/s 
...hf-500m/model.safetensors: 74%|███████▎ | 737MB / 1.00GB 
Processing Files (0 / 1) : 74%|███████▎ | 737MB / 1.00GB, 58.0MB/s
New Data Upload : 92%|█████████▏| 737MB / 805MB, 58.0MB/s 
...hf-500m/model.safetensors: 80%|████████ | 805MB / 1.00GB 
Processing Files (0 / 1) : 80%|████████ | 805MB / 1.00GB, 62.6MB/s
New Data Upload : 100%|█████████▉| 805MB / 805MB, 62.6MB/s 
...hf-500m/model.safetensors: 80%|████████ | 805MB / 1.00GB 
Processing Files (0 / 1) : 80%|████████ | 805MB / 1.00GB, 61.2MB/s
New Data Upload : 92%|█████████▏| 805MB / 872MB, 61.2MB/s 
...hf-500m/model.safetensors: 87%|████████▋ | 872MB / 1.00GB 
Processing Files (0 / 1) : 87%|████████▋ | 872MB / 1.00GB, 65.8MB/s
New Data Upload : 93%|█████████▎| 872MB / 939MB, 65.8MB/s 
...hf-500m/model.safetensors: 94%|█████████▎| 939MB / 1.00GB 
Processing Files (0 / 1) : 94%|█████████▎| 939MB / 1.00GB, 70.2MB/s
New Data Upload : 100%|█████████▉| 939MB / 939MB, 70.2MB/s 
...hf-500m/model.safetensors: 94%|█████████▎| 939MB / 1.00GB 
...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB 
Processing Files (0 / 1) : 100%|█████████▉| 1.00GB / 1.00GB, 72.6MB/s
New Data Upload : 100%|█████████▉| 1.00GB / 1.00GB, 72.6MB/s 
...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB 
Processing Files (0 / 1) : 100%|█████████▉| 1.00GB / 1.00GB, 69.5MB/s
New Data Upload : 100%|█████████▉| 1.00GB / 1.00GB, 69.5MB/s 
...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB 
Processing Files (0 / 1) : 100%|█████████▉| 1.00GB / 1.00GB, 68.0MB/s
New Data Upload : 100%|█████████▉| 1.00GB / 1.00GB, 68.0MB/s 
...hf-500m/model.safetensors: 100%|██████████| 1.00GB / 1.00GB 
Processing Files (1 / 1) : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s
New Data Upload : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s 
Processing Files (1 / 1) : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s
New Data Upload : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s
...hf-500m/model.safetensors: 100%|██████████| 1.00GB / 1.00GB
pushed -> https://huggingface.co/HuggingFace7141/sprocket-500m
[09:26:28] converting to GGUF q4_k_m (phone / Ollama / LM Studio)
Cloning into 'llama.cpp'...
Updating files: 7% (243/3304)
Updating files: 8% (265/3304)
Updating files: 9% (298/3304)
Updating files: 10% (331/3304)
Updating files: 11% (364/3304)
Updating files: 12% (397/3304)
Updating files: 12% (410/3304)
Updating files: 13% (430/3304)
Updating files: 14% (463/3304)
Updating files: 15% (496/3304)
Updating files: 16% (529/3304)
Updating files: 17% (562/3304)
Updating files: 18% (595/3304)
Updating files: 19% (628/3304)
Updating files: 19% (640/3304)
Updating files: 20% (661/3304)
Updating files: 21% (694/3304)
Updating files: 22% (727/3304)
Updating files: 23% (760/3304)
Updating files: 24% (793/3304)
Updating files: 25% (826/3304)
Updating files: 26% (860/3304)
Updating files: 26% (884/3304)
Updating files: 27% (893/3304)
Updating files: 28% (926/3304)
Updating files: 29% (959/3304)
Updating files: 30% (992/3304)
Updating files: 31% (1025/3304)
Updating files: 32% (1058/3304)
Updating files: 33% (1091/3304)
Updating files: 34% (1124/3304)
Updating files: 34% (1142/3304)
Updating files: 35% (1157/3304)
Updating files: 36% (1190/3304)
Updating files: 37% (1223/3304)
Updating files: 38% (1256/3304)
Updating files: 39% (1289/3304)
Updating files: 40% (1322/3304)
Updating files: 41% (1355/3304)
Updating files: 41% (1381/3304)
Updating files: 42% (1388/3304)
Updating files: 43% (1421/3304)
Updating files: 44% (1454/3304)
Updating files: 45% (1487/3304)
Updating files: 46% (1520/3304)
Updating files: 47% (1553/3304)
Updating files: 48% (1586/3304)
Updating files: 49% (1619/3304)
Updating files: 49% (1624/3304)
Updating files: 50% (1652/3304)
Updating files: 51% (1686/3304)
Updating files: 52% (1719/3304)
Updating files: 53% (1752/3304)
Updating files: 54% (1785/3304)
Updating files: 55% (1818/3304)
Updating files: 56% (1851/3304)
Updating files: 56% (1863/3304)
Updating files: 57% (1884/3304)
Updating files: 57% (1911/3304)
Updating files: 58% (1917/3304)
Updating files: 59% (1950/3304)
Updating files: 60% (1983/3304)
Updating files: 61% (2016/3304)
Updating files: 62% (2049/3304)
Updating files: 63% (2082/3304)
Updating files: 63% (2098/3304)
Updating files: 64% (2115/3304)
Updating files: 65% (2148/3304)
Updating files: 66% (2181/3304)
Updating files: 67% (2214/3304)
Updating files: 68% (2247/3304)
Updating files: 69% (2280/3304)
Updating files: 70% (2313/3304)
Updating files: 70% (2341/3304)
Updating files: 71% (2346/3304)
Updating files: 72% (2379/3304)
Updating files: 73% (2412/3304)
Updating files: 74% (2445/3304)
Updating files: 75% (2478/3304)
Updating files: 76% (2512/3304)
Updating files: 77% (2545/3304)
Updating files: 77% (2572/3304)
Updating files: 78% (2578/3304)
Updating files: 79% (2611/3304)
Updating files: 80% (2644/3304)
Updating files: 81% (2677/3304)
Updating files: 82% (2710/3304)
Updating files: 83% (2743/3304)
Updating files: 84% (2776/3304)
Updating files: 85% (2809/3304)
Updating files: 85% (2811/3304)
Updating files: 86% (2842/3304)
Updating files: 87% (2875/3304)
Updating files: 88% (2908/3304)
Updating files: 89% (2941/3304)
Updating files: 90% (2974/3304)
Updating files: 91% (3007/3304)
Updating files: 92% (3040/3304)
Updating files: 92% (3065/3304)
Updating files: 93% (3073/3304)
Updating files: 94% (3106/3304)
Updating files: 95% (3139/3304)
Updating files: 96% (3172/3304)
Updating files: 97% (3205/3304)
Updating files: 98% (3238/3304)
Updating files: 99% (3271/3304)
Updating files: 99% (3300/3304)
Updating files: 100% (3304/3304)
Updating files: 100% (3304/3304), done.
INFO:hf-to-gguf:Loading model: hf-500m
INFO:hf-to-gguf:Model architecture: LlamaForCausalLM
INFO:hf-to-gguf:gguf: indexing model part 'model.safetensors'
INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
INFO:hf-to-gguf:Exporting model...
INFO:hf-to-gguf:token_embd.weight, torch.bfloat16 --> F16, shape = {1280, 32000}
INFO:hf-to-gguf:blk.0.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.0.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.0.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.0.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.0.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.0.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.0.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.0.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.0.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.1.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.1.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.1.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.1.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.1.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.1.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.1.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.1.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.1.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.10.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.10.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.10.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.10.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.10.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.10.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.10.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.10.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.10.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.11.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.11.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.11.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.11.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.11.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.11.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.11.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.11.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.11.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.12.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.12.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.12.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.12.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.12.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.12.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.12.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.12.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.12.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.13.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.13.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.13.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.13.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.13.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.13.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.13.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.13.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.13.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.14.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.14.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.14.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.14.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.14.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.14.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.14.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.14.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.14.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.15.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.15.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.15.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.15.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.15.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.15.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.15.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.15.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.15.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.16.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.16.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.16.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.16.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.16.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.16.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.16.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.16.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.16.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.17.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.17.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.17.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.17.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.17.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.17.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.17.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.17.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.17.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.18.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.18.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.18.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.18.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.18.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.18.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.18.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.18.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.18.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.19.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.19.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.19.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.19.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.19.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.19.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.19.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.19.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.19.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.2.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.2.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.2.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.2.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.2.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.2.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.2.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.2.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.2.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.20.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.20.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.20.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.20.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.20.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.20.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.20.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.20.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.20.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.21.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.21.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.21.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.21.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.21.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.21.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.21.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.21.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.21.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.22.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.22.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.22.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.22.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.22.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.22.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.22.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.22.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.22.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.23.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.23.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.23.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.23.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.23.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.23.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.23.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.23.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.23.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.24.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.24.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.24.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.24.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.24.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.24.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.24.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.24.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.24.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.25.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.25.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.25.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.25.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.25.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.25.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.25.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.25.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.25.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.3.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.3.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.3.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.3.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.3.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.3.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.3.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.3.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.3.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.4.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.4.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.4.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.4.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.4.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.4.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.4.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.4.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.4.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.5.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.5.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.5.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.5.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.5.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.5.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.5.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.5.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.5.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.6.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.6.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.6.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.6.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.6.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.6.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.6.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.6.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.6.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.7.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.7.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.7.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.7.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.7.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.7.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.7.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.7.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.7.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.8.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.8.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.8.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.8.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.8.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.8.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.8.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.8.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.8.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.9.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.9.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.9.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.9.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.9.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.9.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.9.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.9.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.9.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:output_norm.weight, torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:Set meta model
INFO:hf-to-gguf:Set model parameters
INFO:hf-to-gguf:gguf: context length = 2048
INFO:hf-to-gguf:gguf: embedding length = 1280
INFO:hf-to-gguf:gguf: feed forward length = 3584
INFO:hf-to-gguf:gguf: head count = 20
INFO:hf-to-gguf:gguf: key-value head count = 4
INFO:hf-to-gguf:gguf: rope theta = 10000.0
INFO:hf-to-gguf:gguf: rms norm epsilon = 1e-05
INFO:hf-to-gguf:gguf: file type = 1
INFO:hf-to-gguf:Set model quantization version
INFO:hf-to-gguf:Set model tokenizer
WARNING:hf-to-gguf:
WARNING:hf-to-gguf:**************************************************************************************
WARNING:hf-to-gguf:** WARNING: The BPE pre-tokenizer was not recognized!
WARNING:hf-to-gguf:** There are 2 possible reasons for this:
WARNING:hf-to-gguf:** - the model has not been added to convert_hf_to_gguf_update.py yet
WARNING:hf-to-gguf:** - the pre-tokenization config has changed upstream
WARNING:hf-to-gguf:** Check your model files and convert_hf_to_gguf_update.py and update them accordingly.
WARNING:hf-to-gguf:** ref: https://github.com/ggml-org/llama.cpp/pull/6920
WARNING:hf-to-gguf:**
WARNING:hf-to-gguf:** chkhsh: 50cc433d4d528ad898dc1178e17a479f50c03d2c883d07494ae21da32d3eed4f
WARNING:hf-to-gguf:**************************************************************************************
WARNING:hf-to-gguf:
Traceback (most recent call last):
File "/workspace/llama.cpp/conversion/llama.py", line 134, in set_vocab
self._set_vocab_sentencepiece()
File "/workspace/llama.cpp/conversion/base.py", line 1831, in _set_vocab_sentencepiece
tokens, scores, toktypes = self._create_vocab_sentencepiece()
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/llama.cpp/conversion/base.py", line 1848, in _create_vocab_sentencepiece
raise FileNotFoundError(f"File not found: {tokenizer_path}")
FileNotFoundError: File not found: /workspace/hf-500m/tokenizer.model
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File "/workspace/llama.cpp/conversion/llama.py", line 137, in set_vocab
self._set_vocab_llama_hf()
File "/workspace/llama.cpp/conversion/base.py", line 1933, in _set_vocab_llama_hf
vocab = gguf.LlamaHfVocab(self.dir_model)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/llama.cpp/gguf-py/gguf/vocab.py", line 571, in __init__
raise FileNotFoundError('Cannot find Llama BPE tokenizer')
FileNotFoundError: Cannot find Llama BPE tokenizer
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File "/workspace/llama.cpp/convert_hf_to_gguf.py", line 296, in <module>
main()
File "/workspace/llama.cpp/convert_hf_to_gguf.py", line 290, in main
model_instance.write()
File "/workspace/llama.cpp/conversion/base.py", line 1026, in write
self.prepare_metadata(vocab_only=False)
File "/workspace/llama.cpp/conversion/base.py", line 1193, in prepare_metadata
self.set_vocab()
File "/workspace/llama.cpp/conversion/llama.py", line 140, in set_vocab
self._set_vocab_gpt2()
File "/workspace/llama.cpp/conversion/base.py", line 1714, in _set_vocab_gpt2
tokens, toktypes, tokpre = self.get_vocab_base()
^^^^^^^^^^^^^^^^^^^^^
File "/workspace/llama.cpp/conversion/base.py", line 1351, in get_vocab_base
tokpre = self.get_vocab_base_pre(tokenizer)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/workspace/llama.cpp/conversion/base.py", line 1702, in get_vocab_base_pre
raise NotImplementedError("BPE pre-tokenizer was not recognized - update get_vocab_base_pre()")
NotImplementedError: BPE pre-tokenizer was not recognized - update get_vocab_base_pre()
llama_print_build_info: build = 1 (b2f2216)
llama_print_build_info: built with GNU 11.4.0 for Linux x86_64
llama_quantize: quantizing '/workspace/sprocket-500m-f16.gguf' to '/workspace/sprocket-500m-q4_k_m.gguf' as Q4_K_M
gguf_init_from_file: failed to open GGUF file '/workspace/sprocket-500m-f16.gguf' (No such file or directory)
llama_model_quantize: failed to quantize: llama_model_loader: failed to load model from /workspace/sprocket-500m-f16.gguf
llama_quantize: failed to quantize model from '/workspace/sprocket-500m-f16.gguf'
ls: cannot access '/workspace/sprocket-*.gguf': No such file or directory
[09:30:02] DONE.
[09:30:02] base ckpt : /workspace/checkpoints/500m_final.pt
[09:30:02] sft ckpt : /workspace/checkpoints/500m_sft_final.pt
[09:30:02] hf model : /workspace/hf-500m
[09:30:02] gguf : /workspace/sprocket-500m-q4_k_m.gguf
[09:30:02] pipeline finished cleanly
[09:30:02] TERMINATING POD - exit code 0
Runpod config file not found, please run `runpodctl config` to create it
pod "jasakfns6809ft" removed