522 lines
42 KiB
Plaintext
522 lines
42 KiB
Plaintext
=== utc ===
|
||
Thu Jul 30 09:33:21 UTC 2026
|
||
|
||
=== /workspace ===
|
||
total 20585
|
||
drwxrwxrwx 11 root root 3008589 Jul 30 09:26 .
|
||
drwxr-xr-x 1 root root 85 Jul 30 09:33 ..
|
||
drwxrwxrwx 4 root root 2027870 Jul 27 23:24 .cache
|
||
drwxrwxrwx 2 root root 1 Jul 28 00:50 .persist
|
||
drwxrwxrwx 4 root root 1000359 Jul 27 22:35 _shards
|
||
drwxrwxrwx 3 root root 3004667 Jul 30 09:25 checkpoints
|
||
drwxrwxrwx 3 root root 3003766 Jul 27 22:35 data
|
||
drwxrwxrwx 2 root root 2095794 Jul 30 09:26 hf-500m
|
||
drwxrwxrwx 31 root root 2023621 Jul 30 09:27 llama.cpp
|
||
drwxrwxrwx 11 root root 2006109 Jul 28 00:50 sprocket
|
||
-rw-rw-rw- 1 root root 806956 Jul 30 09:30 train.log
|
||
drwxrwxrwx 11 root root 2006034 Jul 27 23:16 tune
|
||
-rw-rw-rw- 1 root root 58167 Jul 27 23:24 tune.log
|
||
-rw-rw-rw- 1 root root 30155 Jul 27 23:56 tune2.log
|
||
-rw-rw-rw- 1 root root 2633 Jul 27 23:56 tune_results.json
|
||
-rw-rw-rw- 1 root root 0 Jul 27 22:34 watchdog.log
|
||
|
||
=== data/processed ===
|
||
total 39505306
|
||
drwxrwxrwx 2 root root 3003766 Jul 30 09:22 .
|
||
drwxrwxrwx 3 root root 3003766 Jul 27 22:35 ..
|
||
-rw-rw-rw- 1 root root 41615850 Jul 30 09:22 sft_packed.npz
|
||
-rw-rw-rw- 1 root root 402804266 Jul 27 22:42 train.bin
|
||
-rw-rw-rw- 1 root root 120 Jul 27 22:42 train.bin.manifest.json
|
||
-rw-rw-rw- 1 root root 40003002080 Jul 28 03:09 train_sample_100BT.bin
|
||
-rw-rw-rw- 1 root root 1062 Jul 28 03:09 train_sample_100BT.bin.manifest.json
|
||
|
||
=== checkpoints ===
|
||
total 33287788
|
||
drwxrwxrwx 3 root root 3004667 Jul 30 09:25 .
|
||
drwxrwxrwx 11 root root 3008589 Jul 30 09:26 ..
|
||
-rw-rw-rw- 1 root root 6013745939 Jul 30 08:58 500m_151500.pt
|
||
-rw-rw-rw- 1 root root 6013747091 Jul 30 09:09 500m_152000.pt
|
||
-rw-rw-rw- 1 root root 6013748179 Jul 30 09:20 500m_152500.pt
|
||
-rw-rw-rw- 1 root root 2004810907 Jul 30 09:22 500m_final.pt
|
||
-rw-rw-rw- 1 root root 6013748435 Jul 30 09:22 500m_latest.pt
|
||
-rw-rw-rw- 1 root root 2004468011 Jul 30 09:25 500m_sft_final.pt
|
||
-rw-rw-rw- 1 root root 6013408307 Jul 30 09:25 500m_sft_latest.pt
|
||
drwxrwxrwx 2 root root 3001493 Jul 28 00:50 _archived_20260728_005036
|
||
|
||
=== shards scratch ===
|
||
total 4893
|
||
drwxrwxrwx 4 root root 1000359 Jul 27 22:35 .
|
||
drwxrwxrwx 11 root root 3008589 Jul 30 09:26 ..
|
||
drwxrwxrwx 3 root root 1000359 Jul 27 22:35 .cache
|
||
drwxrwxrwx 4 root root 1 Jul 28 03:07 sample
|
||
|
||
=== disk ===
|
||
Filesystem Size Used Avail Use% Mounted on
|
||
mfs#us-ne-1.runpod.net:9421 685T 448T 238T 66% /workspace
|
||
|
||
=== watchdog.log ===
|
||
|
||
=== train.log (first 60) ===
|
||
[22:34:48] billing guard armed: AUTO_STOP=1, watchdog 2h (pid 103)
|
||
[22:34:48] repo present; pulling
|
||
Already up to date.
|
||
[22:34:49] installing deps
|
||
WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager, possibly rendering your system unusable.It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv. Use the --root-user-action option if you know what you are doing and want to suppress this warning.
|
||
|
||
[notice] A new release of pip is available: 24.2 -> 26.1.2
|
||
[notice] To update, run: python -m pip install --upgrade pip
|
||
torch 2.4.1+cu124 cuda=True NVIDIA H100 80GB HBM3
|
||
VRAM 85.0 GB total, 84.5 GB free
|
||
[22:35:02] building corpus: target 200000000 tokens from sample/10BT (resumable)
|
||
==========================================================================
|
||
BUILD CORPUS HuggingFaceFW/fineweb-edu:sample/10BT -> data/processed/train.bin
|
||
==========================================================================
|
||
target 200.0M tokens (our 32k tokenizer)
|
||
14 shards in subset, 14 remaining
|
||
|
||
[ 1/14] 000_00000.parquet 194,000 docs 201.4M tok in 444s (dl 8s) | total 201.4M (100.7%) | 453.6K tok/s | ETA -0.0h
|
||
|
||
DONE: 201.4M tokens -> data/processed/train.bin (0.40 GB)
|
||
shards used: 1
|
||
|
||
train with: py -m src.train.train --preset 500m --data data/processed/train.bin --loader memmap --resume auto
|
||
[22:42:27] pretraining 500m for 200 steps (ctx 1024, mb 24 x2)
|
||
corpus 201,402,133 tokens (train 200,395,123 / val 1,007,010) — memmap (0.40 GB, streamed)
|
||
model 500m: 501.1M params
|
||
starting fresh (no --resume)
|
||
iter 0 | train 10.642 | val 9.907 | lr 2.00e-04 | 43,180 tok/s | 62.72 GB | 0.000B tok
|
||
iter 1: 69,762 tok/s (0.70s/iter)
|
||
iter 2: 69,721 tok/s (0.70s/iter)
|
||
iter 3: 69,702 tok/s (0.71s/iter)
|
||
iter 4: 69,682 tok/s (0.71s/iter)
|
||
iter 5: 69,681 tok/s (0.71s/iter)
|
||
iter 20 | train 7.722 | val 7.769 | lr 5.90e-04 | 58,491 tok/s | 66.73 GB | 0.001B tok
|
||
iter 40 | train 7.651 | val 7.636 | lr 5.54e-04 | 61,657 tok/s | 66.73 GB | 0.002B tok
|
||
iter 60 | train 7.212 | val 7.226 | lr 4.96e-04 | 62,815 tok/s | 66.73 GB | 0.003B tok
|
||
iter 80 | train 6.893 | val 6.987 | lr 4.21e-04 | 63,415 tok/s | 66.73 GB | 0.004B tok
|
||
iter 100 | train 6.764 | val 6.815 | lr 3.36e-04 | 62,773 tok/s | 66.73 GB | 0.005B tok
|
||
iter 120 | train 6.656 | val 6.669 | lr 2.51e-04 | 63,172 tok/s | 66.73 GB | 0.006B tok
|
||
iter 140 | train 6.528 | val 6.553 | lr 1.74e-04 | 63,460 tok/s | 66.73 GB | 0.007B tok
|
||
iter 160 | train 6.468 | val 6.521 | lr 1.13e-04 | 63,678 tok/s | 66.73 GB | 0.008B tok
|
||
iter 180 | train 6.483 | val 6.457 | lr 7.36e-05 | 63,286 tok/s | 66.73 GB | 0.009B tok
|
||
iter 200 | train 6.416 | val 6.422 | lr 6.00e-05 | 63,470 tok/s | 66.73 GB | 0.010B tok
|
||
training complete
|
||
[22:45:30] SFT from checkpoints/500m_final.pt
|
||
rendering 18,151 conversations -> ctx 1024 blocks...
|
||
rendered 5,000/18,151
|
||
rendered 10,000/18,151
|
||
rendered 15,000/18,151
|
||
packed 5,079 blocks (0 dropped), cached -> data/processed/sft_packed.npz
|
||
corpus: 5,079 blocks x 1024 = 5,200,896 tokens, 4,567,314 trainable (87.8%)
|
||
train 4,978 blocks / val 101 blocks
|
||
155 steps/epoch x 3.0 epochs = 465 steps (warmup 13)
|
||
model 500m: 501.1M params
|
||
step 0/465 | train 10.626 | val 10.544 | lr 1.54e-06 | ep 0.00 | 43.54 GB
|
||
step 25/465 | train 7.857 | val 7.738 | lr 2.00e-05 | ep 0.16 | 47.55 GB
|
||
step 50/465 | train 6.971 | val 6.907 | lr 1.97e-05 | ep 0.32 | 47.55 GB
|
||
step 75/465 | train 6.322 | val 6.407 | lr 1.92e-05 | ep 0.48 | 47.55 GB
|
||
step 100/465 | train 6.125 | val 6.158 | lr 1.84e-05 | ep 0.65 | 47.55 GB
|
||
step 125/465 | train 5.935 | val 6.009 | lr 1.74e-05 | ep 0.81 | 47.55 GB
|
||
|
||
=== train.log (last 400) ===
|
||
|
||
|
||
...hf-500m/model.safetensors: 7%|▋ | 66.8MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 7%|▋ | 66.8MB / 1.00GB, 6.40MB/s
|
||
|
||
New Data Upload : 50%|████▉ | 66.8MB / 134MB, 6.40MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 7%|▋ | 70.5MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 7%|▋ | 70.5MB / 1.00GB, 6.60MB/s
|
||
|
||
New Data Upload : 35%|███▌ | 70.5MB / 201MB, 6.60MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 13%|█▎ | 134MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 13%|█▎ | 134MB / 1.00GB, 12.5MB/s
|
||
|
||
New Data Upload : 50%|█████ | 134MB / 268MB, 12.5MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 19%|█▉ | 190MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 19%|█▉ | 190MB / 1.00GB, 17.5MB/s
|
||
|
||
New Data Upload : 57%|█████▋ | 190MB / 335MB, 17.5MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 21%|██ | 208MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 21%|██ | 208MB / 1.00GB, 18.8MB/s
|
||
|
||
New Data Upload : 62%|██████▏ | 208MB / 335MB, 18.8MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 33%|███▎ | 335MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 33%|███▎ | 335MB / 1.00GB, 30.2MB/s
|
||
|
||
New Data Upload : 83%|████████▎ | 335MB / 402MB, 30.2MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 33%|███▎ | 335MB / 1.00GB [A[A
|
||
|
||
|
||
...hf-500m/model.safetensors: 34%|███▎ | 338MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 34%|███▎ | 338MB / 1.00GB, 29.0MB/s
|
||
|
||
New Data Upload : 72%|███████▏ | 338MB / 469MB, 29.0MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 47%|████▋ | 469MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 47%|████▋ | 469MB / 1.00GB, 40.4MB/s
|
||
|
||
New Data Upload : 87%|████████▋ | 469MB / 536MB, 40.4MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 53%|█████▎ | 536MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 53%|█████▎ | 536MB / 1.00GB, 45.6MB/s
|
||
|
||
New Data Upload : 89%|████████▉ | 536MB / 604MB, 45.6MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 60%|██████ | 603MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 60%|██████ | 603MB / 1.00GB, 50.6MB/s
|
||
|
||
New Data Upload : 90%|████████▉ | 603MB / 671MB, 50.6MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 67%|██████▋ | 670MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 67%|██████▋ | 670MB / 1.00GB, 54.4MB/s
|
||
|
||
New Data Upload : 91%|█████████ | 670MB / 738MB, 54.4MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 67%|██████▋ | 670MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 67%|██████▋ | 670MB / 1.00GB, 53.2MB/s
|
||
|
||
New Data Upload : 91%|█████████ | 670MB / 738MB, 53.2MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 74%|███████▎ | 737MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 74%|███████▎ | 737MB / 1.00GB, 58.0MB/s
|
||
|
||
New Data Upload : 92%|█████████▏| 737MB / 805MB, 58.0MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 80%|████████ | 805MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 80%|████████ | 805MB / 1.00GB, 62.6MB/s
|
||
|
||
New Data Upload : 100%|█████████▉| 805MB / 805MB, 62.6MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 80%|████████ | 805MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 80%|████████ | 805MB / 1.00GB, 61.2MB/s
|
||
|
||
New Data Upload : 92%|█████████▏| 805MB / 872MB, 61.2MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 87%|████████▋ | 872MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 87%|████████▋ | 872MB / 1.00GB, 65.8MB/s
|
||
|
||
New Data Upload : 93%|█████████▎| 872MB / 939MB, 65.8MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 94%|█████████▎| 939MB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 94%|█████████▎| 939MB / 1.00GB, 70.2MB/s
|
||
|
||
New Data Upload : 100%|█████████▉| 939MB / 939MB, 70.2MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 94%|█████████▎| 939MB / 1.00GB [A[A
|
||
|
||
|
||
...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 100%|█████████▉| 1.00GB / 1.00GB, 72.6MB/s
|
||
|
||
New Data Upload : 100%|█████████▉| 1.00GB / 1.00GB, 72.6MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 100%|█████████▉| 1.00GB / 1.00GB, 69.5MB/s
|
||
|
||
New Data Upload : 100%|█████████▉| 1.00GB / 1.00GB, 69.5MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB [A[A
|
||
Processing Files (0 / 1) : 100%|█████████▉| 1.00GB / 1.00GB, 68.0MB/s
|
||
|
||
New Data Upload : 100%|█████████▉| 1.00GB / 1.00GB, 68.0MB/s [A
|
||
|
||
|
||
...hf-500m/model.safetensors: 100%|██████████| 1.00GB / 1.00GB [A[A
|
||
Processing Files (1 / 1) : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s
|
||
|
||
New Data Upload : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s [A
|
||
Processing Files (1 / 1) : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s
|
||
|
||
New Data Upload : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s
|
||
|
||
...hf-500m/model.safetensors: 100%|██████████| 1.00GB / 1.00GB
|
||
pushed -> https://huggingface.co/HuggingFace7141/sprocket-500m
|
||
[09:26:28] converting to GGUF q4_k_m (phone / Ollama / LM Studio)
|
||
Cloning into 'llama.cpp'...
|
||
Updating files: 7% (243/3304)
|
||
Updating files: 8% (265/3304)
|
||
Updating files: 9% (298/3304)
|
||
Updating files: 10% (331/3304)
|
||
Updating files: 11% (364/3304)
|
||
Updating files: 12% (397/3304)
|
||
Updating files: 12% (410/3304)
|
||
Updating files: 13% (430/3304)
|
||
Updating files: 14% (463/3304)
|
||
Updating files: 15% (496/3304)
|
||
Updating files: 16% (529/3304)
|
||
Updating files: 17% (562/3304)
|
||
Updating files: 18% (595/3304)
|
||
Updating files: 19% (628/3304)
|
||
Updating files: 19% (640/3304)
|
||
Updating files: 20% (661/3304)
|
||
Updating files: 21% (694/3304)
|
||
Updating files: 22% (727/3304)
|
||
Updating files: 23% (760/3304)
|
||
Updating files: 24% (793/3304)
|
||
Updating files: 25% (826/3304)
|
||
Updating files: 26% (860/3304)
|
||
Updating files: 26% (884/3304)
|
||
Updating files: 27% (893/3304)
|
||
Updating files: 28% (926/3304)
|
||
Updating files: 29% (959/3304)
|
||
Updating files: 30% (992/3304)
|
||
Updating files: 31% (1025/3304)
|
||
Updating files: 32% (1058/3304)
|
||
Updating files: 33% (1091/3304)
|
||
Updating files: 34% (1124/3304)
|
||
Updating files: 34% (1142/3304)
|
||
Updating files: 35% (1157/3304)
|
||
Updating files: 36% (1190/3304)
|
||
Updating files: 37% (1223/3304)
|
||
Updating files: 38% (1256/3304)
|
||
Updating files: 39% (1289/3304)
|
||
Updating files: 40% (1322/3304)
|
||
Updating files: 41% (1355/3304)
|
||
Updating files: 41% (1381/3304)
|
||
Updating files: 42% (1388/3304)
|
||
Updating files: 43% (1421/3304)
|
||
Updating files: 44% (1454/3304)
|
||
Updating files: 45% (1487/3304)
|
||
Updating files: 46% (1520/3304)
|
||
Updating files: 47% (1553/3304)
|
||
Updating files: 48% (1586/3304)
|
||
Updating files: 49% (1619/3304)
|
||
Updating files: 49% (1624/3304)
|
||
Updating files: 50% (1652/3304)
|
||
Updating files: 51% (1686/3304)
|
||
Updating files: 52% (1719/3304)
|
||
Updating files: 53% (1752/3304)
|
||
Updating files: 54% (1785/3304)
|
||
Updating files: 55% (1818/3304)
|
||
Updating files: 56% (1851/3304)
|
||
Updating files: 56% (1863/3304)
|
||
Updating files: 57% (1884/3304)
|
||
Updating files: 57% (1911/3304)
|
||
Updating files: 58% (1917/3304)
|
||
Updating files: 59% (1950/3304)
|
||
Updating files: 60% (1983/3304)
|
||
Updating files: 61% (2016/3304)
|
||
Updating files: 62% (2049/3304)
|
||
Updating files: 63% (2082/3304)
|
||
Updating files: 63% (2098/3304)
|
||
Updating files: 64% (2115/3304)
|
||
Updating files: 65% (2148/3304)
|
||
Updating files: 66% (2181/3304)
|
||
Updating files: 67% (2214/3304)
|
||
Updating files: 68% (2247/3304)
|
||
Updating files: 69% (2280/3304)
|
||
Updating files: 70% (2313/3304)
|
||
Updating files: 70% (2341/3304)
|
||
Updating files: 71% (2346/3304)
|
||
Updating files: 72% (2379/3304)
|
||
Updating files: 73% (2412/3304)
|
||
Updating files: 74% (2445/3304)
|
||
Updating files: 75% (2478/3304)
|
||
Updating files: 76% (2512/3304)
|
||
Updating files: 77% (2545/3304)
|
||
Updating files: 77% (2572/3304)
|
||
Updating files: 78% (2578/3304)
|
||
Updating files: 79% (2611/3304)
|
||
Updating files: 80% (2644/3304)
|
||
Updating files: 81% (2677/3304)
|
||
Updating files: 82% (2710/3304)
|
||
Updating files: 83% (2743/3304)
|
||
Updating files: 84% (2776/3304)
|
||
Updating files: 85% (2809/3304)
|
||
Updating files: 85% (2811/3304)
|
||
Updating files: 86% (2842/3304)
|
||
Updating files: 87% (2875/3304)
|
||
Updating files: 88% (2908/3304)
|
||
Updating files: 89% (2941/3304)
|
||
Updating files: 90% (2974/3304)
|
||
Updating files: 91% (3007/3304)
|
||
Updating files: 92% (3040/3304)
|
||
Updating files: 92% (3065/3304)
|
||
Updating files: 93% (3073/3304)
|
||
Updating files: 94% (3106/3304)
|
||
Updating files: 95% (3139/3304)
|
||
Updating files: 96% (3172/3304)
|
||
Updating files: 97% (3205/3304)
|
||
Updating files: 98% (3238/3304)
|
||
Updating files: 99% (3271/3304)
|
||
Updating files: 99% (3300/3304)
|
||
Updating files: 100% (3304/3304)
|
||
Updating files: 100% (3304/3304), done.
|
||
INFO:hf-to-gguf:Loading model: hf-500m
|
||
INFO:hf-to-gguf:Model architecture: LlamaForCausalLM
|
||
INFO:hf-to-gguf:gguf: indexing model part 'model.safetensors'
|
||
INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
|
||
INFO:hf-to-gguf:Exporting model...
|
||
INFO:hf-to-gguf:token_embd.weight, torch.bfloat16 --> F16, shape = {1280, 32000}
|
||
INFO:hf-to-gguf:blk.0.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.0.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.0.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.0.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.0.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.0.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.0.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.0.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.0.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.1.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.1.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.1.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.1.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.1.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.1.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.1.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.1.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.1.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.10.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.10.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.10.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.10.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.10.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.10.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.10.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.10.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.10.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.11.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.11.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.11.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.11.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.11.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.11.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.11.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.11.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.11.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.12.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.12.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.12.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.12.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.12.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.12.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.12.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.12.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.12.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.13.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.13.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.13.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.13.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.13.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.13.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.13.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.13.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.13.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.14.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.14.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.14.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.14.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.14.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.14.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.14.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.14.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.14.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.15.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.15.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.15.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.15.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.15.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.15.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.15.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.15.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.15.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.16.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.16.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.16.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.16.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.16.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.16.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.16.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.16.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.16.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.17.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.17.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.17.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.17.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.17.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.17.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.17.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.17.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.17.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.18.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.18.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.18.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.18.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.18.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.18.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.18.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.18.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.18.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.19.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.19.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.19.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.19.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.19.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.19.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.19.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.19.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.19.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.2.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.2.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.2.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.2.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.2.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.2.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.2.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.2.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.2.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.20.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.20.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.20.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.20.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.20.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.20.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.20.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.20.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.20.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.21.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.21.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.21.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.21.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.21.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.21.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.21.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.21.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.21.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.22.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.22.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.22.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.22.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.22.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.22.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.22.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.22.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.22.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.23.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.23.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.23.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.23.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.23.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.23.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.23.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.23.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.23.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.24.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.24.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.24.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.24.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.24.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.24.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.24.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.24.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.24.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.25.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.25.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.25.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.25.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.25.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.25.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.25.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.25.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.25.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.3.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.3.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.3.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.3.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.3.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.3.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.3.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.3.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.3.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.4.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.4.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.4.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.4.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.4.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.4.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.4.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.4.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.4.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.5.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.5.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.5.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.5.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.5.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.5.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.5.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.5.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.5.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.6.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.6.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.6.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.6.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.6.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.6.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.6.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.6.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.6.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.7.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.7.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.7.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.7.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.7.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.7.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.7.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.7.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.7.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.8.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.8.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.8.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.8.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.8.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.8.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.8.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.8.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.8.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.9.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.9.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||
INFO:hf-to-gguf:blk.9.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.9.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||
INFO:hf-to-gguf:blk.9.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:blk.9.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:blk.9.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.9.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||
INFO:hf-to-gguf:blk.9.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||
INFO:hf-to-gguf:output_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||
INFO:hf-to-gguf:Set meta model
|
||
INFO:hf-to-gguf:Set model parameters
|
||
INFO:hf-to-gguf:gguf: context length = 2048
|
||
INFO:hf-to-gguf:gguf: embedding length = 1280
|
||
INFO:hf-to-gguf:gguf: feed forward length = 3584
|
||
INFO:hf-to-gguf:gguf: head count = 20
|
||
INFO:hf-to-gguf:gguf: key-value head count = 4
|
||
INFO:hf-to-gguf:gguf: rope theta = 10000.0
|
||
INFO:hf-to-gguf:gguf: rms norm epsilon = 1e-05
|
||
INFO:hf-to-gguf:gguf: file type = 1
|
||
INFO:hf-to-gguf:Set model quantization version
|
||
INFO:hf-to-gguf:Set model tokenizer
|
||
WARNING:hf-to-gguf:
|
||
|
||
WARNING:hf-to-gguf:**************************************************************************************
|
||
WARNING:hf-to-gguf:** WARNING: The BPE pre-tokenizer was not recognized!
|
||
WARNING:hf-to-gguf:** There are 2 possible reasons for this:
|
||
WARNING:hf-to-gguf:** - the model has not been added to convert_hf_to_gguf_update.py yet
|
||
WARNING:hf-to-gguf:** - the pre-tokenization config has changed upstream
|
||
WARNING:hf-to-gguf:** Check your model files and convert_hf_to_gguf_update.py and update them accordingly.
|
||
WARNING:hf-to-gguf:** ref: https://github.com/ggml-org/llama.cpp/pull/6920
|
||
WARNING:hf-to-gguf:**
|
||
WARNING:hf-to-gguf:** chkhsh: 50cc433d4d528ad898dc1178e17a479f50c03d2c883d07494ae21da32d3eed4f
|
||
WARNING:hf-to-gguf:**************************************************************************************
|
||
WARNING:hf-to-gguf:
|
||
|
||
Traceback (most recent call last):
|
||
File "/workspace/llama.cpp/conversion/llama.py", line 134, in set_vocab
|
||
self._set_vocab_sentencepiece()
|
||
File "/workspace/llama.cpp/conversion/base.py", line 1831, in _set_vocab_sentencepiece
|
||
tokens, scores, toktypes = self._create_vocab_sentencepiece()
|
||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||
File "/workspace/llama.cpp/conversion/base.py", line 1848, in _create_vocab_sentencepiece
|
||
raise FileNotFoundError(f"File not found: {tokenizer_path}")
|
||
FileNotFoundError: File not found: /workspace/hf-500m/tokenizer.model
|
||
|
||
During handling of the above exception, another exception occurred:
|
||
|
||
Traceback (most recent call last):
|
||
File "/workspace/llama.cpp/conversion/llama.py", line 137, in set_vocab
|
||
self._set_vocab_llama_hf()
|
||
File "/workspace/llama.cpp/conversion/base.py", line 1933, in _set_vocab_llama_hf
|
||
vocab = gguf.LlamaHfVocab(self.dir_model)
|
||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||
File "/workspace/llama.cpp/gguf-py/gguf/vocab.py", line 571, in __init__
|
||
raise FileNotFoundError('Cannot find Llama BPE tokenizer')
|
||
FileNotFoundError: Cannot find Llama BPE tokenizer
|
||
|
||
During handling of the above exception, another exception occurred:
|
||
|
||
Traceback (most recent call last):
|
||
File "/workspace/llama.cpp/convert_hf_to_gguf.py", line 296, in <module>
|
||
main()
|
||
File "/workspace/llama.cpp/convert_hf_to_gguf.py", line 290, in main
|
||
model_instance.write()
|
||
File "/workspace/llama.cpp/conversion/base.py", line 1026, in write
|
||
self.prepare_metadata(vocab_only=False)
|
||
File "/workspace/llama.cpp/conversion/base.py", line 1193, in prepare_metadata
|
||
self.set_vocab()
|
||
File "/workspace/llama.cpp/conversion/llama.py", line 140, in set_vocab
|
||
self._set_vocab_gpt2()
|
||
File "/workspace/llama.cpp/conversion/base.py", line 1714, in _set_vocab_gpt2
|
||
tokens, toktypes, tokpre = self.get_vocab_base()
|
||
^^^^^^^^^^^^^^^^^^^^^
|
||
File "/workspace/llama.cpp/conversion/base.py", line 1351, in get_vocab_base
|
||
tokpre = self.get_vocab_base_pre(tokenizer)
|
||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||
File "/workspace/llama.cpp/conversion/base.py", line 1702, in get_vocab_base_pre
|
||
raise NotImplementedError("BPE pre-tokenizer was not recognized - update get_vocab_base_pre()")
|
||
NotImplementedError: BPE pre-tokenizer was not recognized - update get_vocab_base_pre()
|
||
llama_print_build_info: build = 1 (b2f2216)
|
||
llama_print_build_info: built with GNU 11.4.0 for Linux x86_64
|
||
llama_quantize: quantizing '/workspace/sprocket-500m-f16.gguf' to '/workspace/sprocket-500m-q4_k_m.gguf' as Q4_K_M
|
||
gguf_init_from_file: failed to open GGUF file '/workspace/sprocket-500m-f16.gguf' (No such file or directory)
|
||
llama_model_quantize: failed to quantize: llama_model_loader: failed to load model from /workspace/sprocket-500m-f16.gguf
|
||
llama_quantize: failed to quantize model from '/workspace/sprocket-500m-f16.gguf'
|
||
ls: cannot access '/workspace/sprocket-*.gguf': No such file or directory
|
||
[09:30:02] DONE.
|
||
[09:30:02] base ckpt : /workspace/checkpoints/500m_final.pt
|
||
[09:30:02] sft ckpt : /workspace/checkpoints/500m_sft_final.pt
|
||
[09:30:02] hf model : /workspace/hf-500m
|
||
[09:30:02] gguf : /workspace/sprocket-500m-q4_k_m.gguf
|
||
[09:30:02] pipeline finished cleanly
|
||
[09:30:02] TERMINATING POD - exit code 0
|
||
Runpod config file not found, please run `runpodctl config` to create it
|
||
pod "jasakfns6809ft" removed
|