=== utc ===
Thu Jul 30 09:33:21 UTC 2026

=== /workspace ===
total 20585
drwxrwxrwx 11 root root 3008589 Jul 30 09:26 .
drwxr-xr-x  1 root root      85 Jul 30 09:33 ..
drwxrwxrwx  4 root root 2027870 Jul 27 23:24 .cache
drwxrwxrwx  2 root root       1 Jul 28 00:50 .persist
drwxrwxrwx  4 root root 1000359 Jul 27 22:35 _shards
drwxrwxrwx  3 root root 3004667 Jul 30 09:25 checkpoints
drwxrwxrwx  3 root root 3003766 Jul 27 22:35 data
drwxrwxrwx  2 root root 2095794 Jul 30 09:26 hf-500m
drwxrwxrwx 31 root root 2023621 Jul 30 09:27 llama.cpp
drwxrwxrwx 11 root root 2006109 Jul 28 00:50 sprocket
-rw-rw-rw-  1 root root  806956 Jul 30 09:30 train.log
drwxrwxrwx 11 root root 2006034 Jul 27 23:16 tune
-rw-rw-rw-  1 root root   58167 Jul 27 23:24 tune.log
-rw-rw-rw-  1 root root   30155 Jul 27 23:56 tune2.log
-rw-rw-rw-  1 root root    2633 Jul 27 23:56 tune_results.json
-rw-rw-rw-  1 root root       0 Jul 27 22:34 watchdog.log

=== data/processed ===
total 39505306
drwxrwxrwx 2 root root     3003766 Jul 30 09:22 .
drwxrwxrwx 3 root root     3003766 Jul 27 22:35 ..
-rw-rw-rw- 1 root root    41615850 Jul 30 09:22 sft_packed.npz
-rw-rw-rw- 1 root root   402804266 Jul 27 22:42 train.bin
-rw-rw-rw- 1 root root         120 Jul 27 22:42 train.bin.manifest.json
-rw-rw-rw- 1 root root 40003002080 Jul 28 03:09 train_sample_100BT.bin
-rw-rw-rw- 1 root root        1062 Jul 28 03:09 train_sample_100BT.bin.manifest.json

=== checkpoints ===
total 33287788
drwxrwxrwx  3 root root    3004667 Jul 30 09:25 .
drwxrwxrwx 11 root root    3008589 Jul 30 09:26 ..
-rw-rw-rw-  1 root root 6013745939 Jul 30 08:58 500m_151500.pt
-rw-rw-rw-  1 root root 6013747091 Jul 30 09:09 500m_152000.pt
-rw-rw-rw-  1 root root 6013748179 Jul 30 09:20 500m_152500.pt
-rw-rw-rw-  1 root root 2004810907 Jul 30 09:22 500m_final.pt
-rw-rw-rw-  1 root root 6013748435 Jul 30 09:22 500m_latest.pt
-rw-rw-rw-  1 root root 2004468011 Jul 30 09:25 500m_sft_final.pt
-rw-rw-rw-  1 root root 6013408307 Jul 30 09:25 500m_sft_latest.pt
drwxrwxrwx  2 root root    3001493 Jul 28 00:50 _archived_20260728_005036

=== shards scratch ===
total 4893
drwxrwxrwx  4 root root 1000359 Jul 27 22:35 .
drwxrwxrwx 11 root root 3008589 Jul 30 09:26 ..
drwxrwxrwx  3 root root 1000359 Jul 27 22:35 .cache
drwxrwxrwx  4 root root       1 Jul 28 03:07 sample

=== disk ===
Filesystem                   Size  Used Avail Use% Mounted on
mfs#us-ne-1.runpod.net:9421  685T  448T  238T  66% /workspace

=== watchdog.log ===

=== train.log (first 60) ===
[22:34:48] billing guard armed: AUTO_STOP=1, watchdog 2h (pid 103)
[22:34:48] repo present; pulling
Already up to date.
[22:34:49] installing deps
WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager, possibly rendering your system unusable.It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv. Use the --root-user-action option if you know what you are doing and want to suppress this warning.

[notice] A new release of pip is available: 24.2 -> 26.1.2
[notice] To update, run: python -m pip install --upgrade pip
  torch 2.4.1+cu124  cuda=True  NVIDIA H100 80GB HBM3
  VRAM 85.0 GB total, 84.5 GB free
[22:35:02] building corpus: target 200000000 tokens from sample/10BT (resumable)
==========================================================================
BUILD CORPUS  HuggingFaceFW/fineweb-edu:sample/10BT  ->  data/processed/train.bin
==========================================================================
  target 200.0M tokens (our 32k tokenizer)
  14 shards in subset, 14 remaining

  [   1/14] 000_00000.parquet 194,000 docs 201.4M tok in 444s (dl 8s) | total 201.4M (100.7%) | 453.6K tok/s | ETA -0.0h

  DONE: 201.4M tokens -> data/processed/train.bin (0.40 GB)
  shards used: 1

  train with:  py -m src.train.train --preset 500m --data data/processed/train.bin --loader memmap --resume auto
[22:42:27] pretraining 500m for 200 steps (ctx 1024, mb 24 x2)
corpus 201,402,133 tokens (train 200,395,123 / val 1,007,010) — memmap (0.40 GB, streamed)
model 500m: 501.1M params
starting fresh (no --resume)
iter      0 | train 10.642 | val 9.907 | lr 2.00e-04 | 43,180 tok/s | 62.72 GB | 0.000B tok
  iter 1: 69,762 tok/s  (0.70s/iter)
  iter 2: 69,721 tok/s  (0.70s/iter)
  iter 3: 69,702 tok/s  (0.71s/iter)
  iter 4: 69,682 tok/s  (0.71s/iter)
  iter 5: 69,681 tok/s  (0.71s/iter)
iter     20 | train 7.722 | val 7.769 | lr 5.90e-04 | 58,491 tok/s | 66.73 GB | 0.001B tok
iter     40 | train 7.651 | val 7.636 | lr 5.54e-04 | 61,657 tok/s | 66.73 GB | 0.002B tok
iter     60 | train 7.212 | val 7.226 | lr 4.96e-04 | 62,815 tok/s | 66.73 GB | 0.003B tok
iter     80 | train 6.893 | val 6.987 | lr 4.21e-04 | 63,415 tok/s | 66.73 GB | 0.004B tok
iter    100 | train 6.764 | val 6.815 | lr 3.36e-04 | 62,773 tok/s | 66.73 GB | 0.005B tok
iter    120 | train 6.656 | val 6.669 | lr 2.51e-04 | 63,172 tok/s | 66.73 GB | 0.006B tok
iter    140 | train 6.528 | val 6.553 | lr 1.74e-04 | 63,460 tok/s | 66.73 GB | 0.007B tok
iter    160 | train 6.468 | val 6.521 | lr 1.13e-04 | 63,678 tok/s | 66.73 GB | 0.008B tok
iter    180 | train 6.483 | val 6.457 | lr 7.36e-05 | 63,286 tok/s | 66.73 GB | 0.009B tok
iter    200 | train 6.416 | val 6.422 | lr 6.00e-05 | 63,470 tok/s | 66.73 GB | 0.010B tok
training complete
[22:45:30] SFT from checkpoints/500m_final.pt
rendering 18,151 conversations -> ctx 1024 blocks...
    rendered 5,000/18,151
    rendered 10,000/18,151
    rendered 15,000/18,151
  packed 5,079 blocks (0 dropped), cached -> data/processed/sft_packed.npz
corpus: 5,079 blocks x 1024 = 5,200,896 tokens, 4,567,314 trainable (87.8%)
  train 4,978 blocks / val 101 blocks
  155 steps/epoch x 3.0 epochs = 465 steps (warmup 13)
model 500m: 501.1M params
step     0/465 | train 10.626 | val 10.544 | lr 1.54e-06 | ep 0.00 | 43.54 GB
step    25/465 | train 7.857 | val 7.738 | lr 2.00e-05 | ep 0.16 | 47.55 GB
step    50/465 | train 6.971 | val 6.907 | lr 1.97e-05 | ep 0.32 | 47.55 GB
step    75/465 | train 6.322 | val 6.407 | lr 1.92e-05 | ep 0.48 | 47.55 GB
step   100/465 | train 6.125 | val 6.158 | lr 1.84e-05 | ep 0.65 | 47.55 GB
step   125/465 | train 5.935 | val 6.009 | lr 1.74e-05 | ep 0.81 | 47.55 GB

=== train.log (last 400) ===


  ...hf-500m/model.safetensors:   7%|▋         | 66.8MB / 1.00GB            [A[A
Processing Files (0 / 1)      :   7%|▋         | 66.8MB / 1.00GB, 6.40MB/s  

New Data Upload               :  50%|████▉     | 66.8MB /  134MB, 6.40MB/s  [A


  ...hf-500m/model.safetensors:   7%|▋         | 70.5MB / 1.00GB            [A[A
Processing Files (0 / 1)      :   7%|▋         | 70.5MB / 1.00GB, 6.60MB/s  

New Data Upload               :  35%|███▌      | 70.5MB /  201MB, 6.60MB/s  [A


  ...hf-500m/model.safetensors:  13%|█▎        |  134MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  13%|█▎        |  134MB / 1.00GB, 12.5MB/s  

New Data Upload               :  50%|█████     |  134MB /  268MB, 12.5MB/s  [A


  ...hf-500m/model.safetensors:  19%|█▉        |  190MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  19%|█▉        |  190MB / 1.00GB, 17.5MB/s  

New Data Upload               :  57%|█████▋    |  190MB /  335MB, 17.5MB/s  [A


  ...hf-500m/model.safetensors:  21%|██        |  208MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  21%|██        |  208MB / 1.00GB, 18.8MB/s  

New Data Upload               :  62%|██████▏   |  208MB /  335MB, 18.8MB/s  [A


  ...hf-500m/model.safetensors:  33%|███▎      |  335MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  33%|███▎      |  335MB / 1.00GB, 30.2MB/s  

New Data Upload               :  83%|████████▎ |  335MB /  402MB, 30.2MB/s  [A


  ...hf-500m/model.safetensors:  33%|███▎      |  335MB / 1.00GB            [A[A


  ...hf-500m/model.safetensors:  34%|███▎      |  338MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  34%|███▎      |  338MB / 1.00GB, 29.0MB/s  

New Data Upload               :  72%|███████▏  |  338MB /  469MB, 29.0MB/s  [A


  ...hf-500m/model.safetensors:  47%|████▋     |  469MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  47%|████▋     |  469MB / 1.00GB, 40.4MB/s  

New Data Upload               :  87%|████████▋ |  469MB /  536MB, 40.4MB/s  [A


  ...hf-500m/model.safetensors:  53%|█████▎    |  536MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  53%|█████▎    |  536MB / 1.00GB, 45.6MB/s  

New Data Upload               :  89%|████████▉ |  536MB /  604MB, 45.6MB/s  [A


  ...hf-500m/model.safetensors:  60%|██████    |  603MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  60%|██████    |  603MB / 1.00GB, 50.6MB/s  

New Data Upload               :  90%|████████▉ |  603MB /  671MB, 50.6MB/s  [A


  ...hf-500m/model.safetensors:  67%|██████▋   |  670MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  67%|██████▋   |  670MB / 1.00GB, 54.4MB/s  

New Data Upload               :  91%|█████████ |  670MB /  738MB, 54.4MB/s  [A


  ...hf-500m/model.safetensors:  67%|██████▋   |  670MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  67%|██████▋   |  670MB / 1.00GB, 53.2MB/s  

New Data Upload               :  91%|█████████ |  670MB /  738MB, 53.2MB/s  [A


  ...hf-500m/model.safetensors:  74%|███████▎  |  737MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  74%|███████▎  |  737MB / 1.00GB, 58.0MB/s  

New Data Upload               :  92%|█████████▏|  737MB /  805MB, 58.0MB/s  [A


  ...hf-500m/model.safetensors:  80%|████████  |  805MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  80%|████████  |  805MB / 1.00GB, 62.6MB/s  

New Data Upload               : 100%|█████████▉|  805MB /  805MB, 62.6MB/s  [A


  ...hf-500m/model.safetensors:  80%|████████  |  805MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  80%|████████  |  805MB / 1.00GB, 61.2MB/s  

New Data Upload               :  92%|█████████▏|  805MB /  872MB, 61.2MB/s  [A


  ...hf-500m/model.safetensors:  87%|████████▋ |  872MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  87%|████████▋ |  872MB / 1.00GB, 65.8MB/s  

New Data Upload               :  93%|█████████▎|  872MB /  939MB, 65.8MB/s  [A


  ...hf-500m/model.safetensors:  94%|█████████▎|  939MB / 1.00GB            [A[A
Processing Files (0 / 1)      :  94%|█████████▎|  939MB / 1.00GB, 70.2MB/s  

New Data Upload               : 100%|█████████▉|  939MB /  939MB, 70.2MB/s  [A


  ...hf-500m/model.safetensors:  94%|█████████▎|  939MB / 1.00GB            [A[A


  ...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB            [A[A
Processing Files (0 / 1)      : 100%|█████████▉| 1.00GB / 1.00GB, 72.6MB/s  

New Data Upload               : 100%|█████████▉| 1.00GB / 1.00GB, 72.6MB/s  [A


  ...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB            [A[A
Processing Files (0 / 1)      : 100%|█████████▉| 1.00GB / 1.00GB, 69.5MB/s  

New Data Upload               : 100%|█████████▉| 1.00GB / 1.00GB, 69.5MB/s  [A


  ...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB            [A[A
Processing Files (0 / 1)      : 100%|█████████▉| 1.00GB / 1.00GB, 68.0MB/s  

New Data Upload               : 100%|█████████▉| 1.00GB / 1.00GB, 68.0MB/s  [A


  ...hf-500m/model.safetensors: 100%|██████████| 1.00GB / 1.00GB            [A[A
Processing Files (1 / 1)      : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s  

New Data Upload               : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s  [A
Processing Files (1 / 1)      : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s  

New Data Upload               : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s  

  ...hf-500m/model.safetensors: 100%|██████████| 1.00GB / 1.00GB            
  pushed -> https://huggingface.co/HuggingFace7141/sprocket-500m
[09:26:28] converting to GGUF q4_k_m (phone / Ollama / LM Studio)
Cloning into 'llama.cpp'...
Updating files:   7% (243/3304)
Updating files:   8% (265/3304)
Updating files:   9% (298/3304)
Updating files:  10% (331/3304)
Updating files:  11% (364/3304)
Updating files:  12% (397/3304)
Updating files:  12% (410/3304)
Updating files:  13% (430/3304)
Updating files:  14% (463/3304)
Updating files:  15% (496/3304)
Updating files:  16% (529/3304)
Updating files:  17% (562/3304)
Updating files:  18% (595/3304)
Updating files:  19% (628/3304)
Updating files:  19% (640/3304)
Updating files:  20% (661/3304)
Updating files:  21% (694/3304)
Updating files:  22% (727/3304)
Updating files:  23% (760/3304)
Updating files:  24% (793/3304)
Updating files:  25% (826/3304)
Updating files:  26% (860/3304)
Updating files:  26% (884/3304)
Updating files:  27% (893/3304)
Updating files:  28% (926/3304)
Updating files:  29% (959/3304)
Updating files:  30% (992/3304)
Updating files:  31% (1025/3304)
Updating files:  32% (1058/3304)
Updating files:  33% (1091/3304)
Updating files:  34% (1124/3304)
Updating files:  34% (1142/3304)
Updating files:  35% (1157/3304)
Updating files:  36% (1190/3304)
Updating files:  37% (1223/3304)
Updating files:  38% (1256/3304)
Updating files:  39% (1289/3304)
Updating files:  40% (1322/3304)
Updating files:  41% (1355/3304)
Updating files:  41% (1381/3304)
Updating files:  42% (1388/3304)
Updating files:  43% (1421/3304)
Updating files:  44% (1454/3304)
Updating files:  45% (1487/3304)
Updating files:  46% (1520/3304)
Updating files:  47% (1553/3304)
Updating files:  48% (1586/3304)
Updating files:  49% (1619/3304)
Updating files:  49% (1624/3304)
Updating files:  50% (1652/3304)
Updating files:  51% (1686/3304)
Updating files:  52% (1719/3304)
Updating files:  53% (1752/3304)
Updating files:  54% (1785/3304)
Updating files:  55% (1818/3304)
Updating files:  56% (1851/3304)
Updating files:  56% (1863/3304)
Updating files:  57% (1884/3304)
Updating files:  57% (1911/3304)
Updating files:  58% (1917/3304)
Updating files:  59% (1950/3304)
Updating files:  60% (1983/3304)
Updating files:  61% (2016/3304)
Updating files:  62% (2049/3304)
Updating files:  63% (2082/3304)
Updating files:  63% (2098/3304)
Updating files:  64% (2115/3304)
Updating files:  65% (2148/3304)
Updating files:  66% (2181/3304)
Updating files:  67% (2214/3304)
Updating files:  68% (2247/3304)
Updating files:  69% (2280/3304)
Updating files:  70% (2313/3304)
Updating files:  70% (2341/3304)
Updating files:  71% (2346/3304)
Updating files:  72% (2379/3304)
Updating files:  73% (2412/3304)
Updating files:  74% (2445/3304)
Updating files:  75% (2478/3304)
Updating files:  76% (2512/3304)
Updating files:  77% (2545/3304)
Updating files:  77% (2572/3304)
Updating files:  78% (2578/3304)
Updating files:  79% (2611/3304)
Updating files:  80% (2644/3304)
Updating files:  81% (2677/3304)
Updating files:  82% (2710/3304)
Updating files:  83% (2743/3304)
Updating files:  84% (2776/3304)
Updating files:  85% (2809/3304)
Updating files:  85% (2811/3304)
Updating files:  86% (2842/3304)
Updating files:  87% (2875/3304)
Updating files:  88% (2908/3304)
Updating files:  89% (2941/3304)
Updating files:  90% (2974/3304)
Updating files:  91% (3007/3304)
Updating files:  92% (3040/3304)
Updating files:  92% (3065/3304)
Updating files:  93% (3073/3304)
Updating files:  94% (3106/3304)
Updating files:  95% (3139/3304)
Updating files:  96% (3172/3304)
Updating files:  97% (3205/3304)
Updating files:  98% (3238/3304)
Updating files:  99% (3271/3304)
Updating files:  99% (3300/3304)
Updating files: 100% (3304/3304)
Updating files: 100% (3304/3304), done.
INFO:hf-to-gguf:Loading model: hf-500m
INFO:hf-to-gguf:Model architecture: LlamaForCausalLM
INFO:hf-to-gguf:gguf: indexing model part 'model.safetensors'
INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
INFO:hf-to-gguf:Exporting model...
INFO:hf-to-gguf:token_embd.weight,           torch.bfloat16 --> F16, shape = {1280, 32000}
INFO:hf-to-gguf:blk.0.attn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.0.ffn_down.weight,       torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.0.ffn_gate.weight,       torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.0.ffn_up.weight,         torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.0.ffn_norm.weight,       torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.0.attn_k.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.0.attn_output.weight,    torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.0.attn_q.weight,         torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.0.attn_v.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.1.attn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.1.ffn_down.weight,       torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.1.ffn_gate.weight,       torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.1.ffn_up.weight,         torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.1.ffn_norm.weight,       torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.1.attn_k.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.1.attn_output.weight,    torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.1.attn_q.weight,         torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.1.attn_v.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.10.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.10.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.10.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.10.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.10.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.10.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.10.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.10.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.10.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.11.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.11.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.11.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.11.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.11.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.11.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.11.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.11.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.11.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.12.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.12.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.12.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.12.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.12.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.12.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.12.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.12.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.12.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.13.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.13.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.13.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.13.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.13.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.13.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.13.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.13.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.13.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.14.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.14.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.14.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.14.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.14.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.14.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.14.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.14.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.14.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.15.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.15.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.15.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.15.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.15.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.15.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.15.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.15.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.15.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.16.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.16.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.16.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.16.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.16.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.16.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.16.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.16.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.16.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.17.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.17.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.17.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.17.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.17.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.17.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.17.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.17.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.17.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.18.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.18.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.18.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.18.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.18.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.18.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.18.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.18.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.18.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.19.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.19.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.19.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.19.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.19.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.19.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.19.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.19.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.19.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.2.attn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.2.ffn_down.weight,       torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.2.ffn_gate.weight,       torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.2.ffn_up.weight,         torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.2.ffn_norm.weight,       torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.2.attn_k.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.2.attn_output.weight,    torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.2.attn_q.weight,         torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.2.attn_v.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.20.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.20.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.20.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.20.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.20.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.20.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.20.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.20.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.20.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.21.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.21.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.21.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.21.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.21.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.21.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.21.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.21.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.21.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.22.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.22.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.22.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.22.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.22.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.22.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.22.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.22.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.22.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.23.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.23.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.23.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.23.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.23.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.23.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.23.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.23.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.23.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.24.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.24.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.24.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.24.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.24.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.24.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.24.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.24.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.24.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.25.attn_norm.weight,     torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.25.ffn_down.weight,      torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.25.ffn_gate.weight,      torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.25.ffn_up.weight,        torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.25.ffn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.25.attn_k.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.25.attn_output.weight,   torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.25.attn_q.weight,        torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.25.attn_v.weight,        torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.3.attn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.3.ffn_down.weight,       torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.3.ffn_gate.weight,       torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.3.ffn_up.weight,         torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.3.ffn_norm.weight,       torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.3.attn_k.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.3.attn_output.weight,    torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.3.attn_q.weight,         torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.3.attn_v.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.4.attn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.4.ffn_down.weight,       torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.4.ffn_gate.weight,       torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.4.ffn_up.weight,         torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.4.ffn_norm.weight,       torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.4.attn_k.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.4.attn_output.weight,    torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.4.attn_q.weight,         torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.4.attn_v.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.5.attn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.5.ffn_down.weight,       torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.5.ffn_gate.weight,       torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.5.ffn_up.weight,         torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.5.ffn_norm.weight,       torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.5.attn_k.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.5.attn_output.weight,    torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.5.attn_q.weight,         torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.5.attn_v.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.6.attn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.6.ffn_down.weight,       torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.6.ffn_gate.weight,       torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.6.ffn_up.weight,         torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.6.ffn_norm.weight,       torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.6.attn_k.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.6.attn_output.weight,    torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.6.attn_q.weight,         torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.6.attn_v.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.7.attn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.7.ffn_down.weight,       torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.7.ffn_gate.weight,       torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.7.ffn_up.weight,         torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.7.ffn_norm.weight,       torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.7.attn_k.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.7.attn_output.weight,    torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.7.attn_q.weight,         torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.7.attn_v.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.8.attn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.8.ffn_down.weight,       torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.8.ffn_gate.weight,       torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.8.ffn_up.weight,         torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.8.ffn_norm.weight,       torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.8.attn_k.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.8.attn_output.weight,    torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.8.attn_q.weight,         torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.8.attn_v.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.9.attn_norm.weight,      torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.9.ffn_down.weight,       torch.bfloat16 --> F16, shape = {3584, 1280}
INFO:hf-to-gguf:blk.9.ffn_gate.weight,       torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.9.ffn_up.weight,         torch.bfloat16 --> F16, shape = {1280, 3584}
INFO:hf-to-gguf:blk.9.ffn_norm.weight,       torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:blk.9.attn_k.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:blk.9.attn_output.weight,    torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.9.attn_q.weight,         torch.bfloat16 --> F16, shape = {1280, 1280}
INFO:hf-to-gguf:blk.9.attn_v.weight,         torch.bfloat16 --> F16, shape = {1280, 256}
INFO:hf-to-gguf:output_norm.weight,          torch.bfloat16 --> F32, shape = {1280}
INFO:hf-to-gguf:Set meta model
INFO:hf-to-gguf:Set model parameters
INFO:hf-to-gguf:gguf: context length = 2048
INFO:hf-to-gguf:gguf: embedding length = 1280
INFO:hf-to-gguf:gguf: feed forward length = 3584
INFO:hf-to-gguf:gguf: head count = 20
INFO:hf-to-gguf:gguf: key-value head count = 4
INFO:hf-to-gguf:gguf: rope theta = 10000.0
INFO:hf-to-gguf:gguf: rms norm epsilon = 1e-05
INFO:hf-to-gguf:gguf: file type = 1
INFO:hf-to-gguf:Set model quantization version
INFO:hf-to-gguf:Set model tokenizer
WARNING:hf-to-gguf:

WARNING:hf-to-gguf:**************************************************************************************
WARNING:hf-to-gguf:** WARNING: The BPE pre-tokenizer was not recognized!
WARNING:hf-to-gguf:**          There are 2 possible reasons for this:
WARNING:hf-to-gguf:**          - the model has not been added to convert_hf_to_gguf_update.py yet
WARNING:hf-to-gguf:**          - the pre-tokenization config has changed upstream
WARNING:hf-to-gguf:**          Check your model files and convert_hf_to_gguf_update.py and update them accordingly.
WARNING:hf-to-gguf:** ref:     https://github.com/ggml-org/llama.cpp/pull/6920
WARNING:hf-to-gguf:**
WARNING:hf-to-gguf:** chkhsh:  50cc433d4d528ad898dc1178e17a479f50c03d2c883d07494ae21da32d3eed4f
WARNING:hf-to-gguf:**************************************************************************************
WARNING:hf-to-gguf:

Traceback (most recent call last):
  File "/workspace/llama.cpp/conversion/llama.py", line 134, in set_vocab
    self._set_vocab_sentencepiece()
  File "/workspace/llama.cpp/conversion/base.py", line 1831, in _set_vocab_sentencepiece
    tokens, scores, toktypes = self._create_vocab_sentencepiece()
                               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/workspace/llama.cpp/conversion/base.py", line 1848, in _create_vocab_sentencepiece
    raise FileNotFoundError(f"File not found: {tokenizer_path}")
FileNotFoundError: File not found: /workspace/hf-500m/tokenizer.model

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/workspace/llama.cpp/conversion/llama.py", line 137, in set_vocab
    self._set_vocab_llama_hf()
  File "/workspace/llama.cpp/conversion/base.py", line 1933, in _set_vocab_llama_hf
    vocab = gguf.LlamaHfVocab(self.dir_model)
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/workspace/llama.cpp/gguf-py/gguf/vocab.py", line 571, in __init__
    raise FileNotFoundError('Cannot find Llama BPE tokenizer')
FileNotFoundError: Cannot find Llama BPE tokenizer

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/workspace/llama.cpp/convert_hf_to_gguf.py", line 296, in <module>
    main()
  File "/workspace/llama.cpp/convert_hf_to_gguf.py", line 290, in main
    model_instance.write()
  File "/workspace/llama.cpp/conversion/base.py", line 1026, in write
    self.prepare_metadata(vocab_only=False)
  File "/workspace/llama.cpp/conversion/base.py", line 1193, in prepare_metadata
    self.set_vocab()
  File "/workspace/llama.cpp/conversion/llama.py", line 140, in set_vocab
    self._set_vocab_gpt2()
  File "/workspace/llama.cpp/conversion/base.py", line 1714, in _set_vocab_gpt2
    tokens, toktypes, tokpre = self.get_vocab_base()
                               ^^^^^^^^^^^^^^^^^^^^^
  File "/workspace/llama.cpp/conversion/base.py", line 1351, in get_vocab_base
    tokpre = self.get_vocab_base_pre(tokenizer)
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/workspace/llama.cpp/conversion/base.py", line 1702, in get_vocab_base_pre
    raise NotImplementedError("BPE pre-tokenizer was not recognized - update get_vocab_base_pre()")
NotImplementedError: BPE pre-tokenizer was not recognized - update get_vocab_base_pre()
llama_print_build_info: build = 1 (b2f2216)
llama_print_build_info: built with GNU 11.4.0 for Linux x86_64
llama_quantize: quantizing '/workspace/sprocket-500m-f16.gguf' to '/workspace/sprocket-500m-q4_k_m.gguf' as Q4_K_M
gguf_init_from_file: failed to open GGUF file '/workspace/sprocket-500m-f16.gguf' (No such file or directory)
llama_model_quantize: failed to quantize: llama_model_loader: failed to load model from /workspace/sprocket-500m-f16.gguf
llama_quantize: failed to quantize model from '/workspace/sprocket-500m-f16.gguf'
ls: cannot access '/workspace/sprocket-*.gguf': No such file or directory
[09:30:02] DONE.
[09:30:02]   base ckpt : /workspace/checkpoints/500m_final.pt
[09:30:02]   sft ckpt  : /workspace/checkpoints/500m_sft_final.pt
[09:30:02]   hf model  : /workspace/hf-500m
[09:30:02]   gguf      : /workspace/sprocket-500m-q4_k_m.gguf
[09:30:02] pipeline finished cleanly
[09:30:02] TERMINATING POD - exit code 0
Runpod config file not found, please run `runpodctl config` to create it
pod "jasakfns6809ft" removed
