初始化项目,由ModelHub XC社区提供模型

Model: rombodawg/LosslessMegaCoder-llama2-13b-mini
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-02 16:16:18 +08:00
commit 8471292b55
19 changed files with 94141 additions and 0 deletions

35
.gitattributes vendored Normal file
View File

@@ -0,0 +1,35 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

98
README.md Normal file
View File

@@ -0,0 +1,98 @@
---
license: llama2
datasets:
- rombodawg/LosslessMegaCodeTrainingV2_1m_Evol_Uncensored
---
___________________________
- Please note this model was not trained on the rombodawg/LosslessMegaCodeTrainingV3_MINI dataset, despite the name similarity. You can find the training data at the bottom of the model card labeled (megacode2-min100)
___________________________
This is one of the first models trained on the LosslessMegaCodeTrainingV2_1m_Evol_Uncensored dataset. The version of the dataset used for this model was filtered by removed any data with less than 100 tokens but plans for much more refined filtering are in the works
- This model was made as a colaboration between me and andreaskoepf who is an affiliate of Open Assistant.
This Model score .29 on humaneval+ the same as LLaMA-2 70B Chat Link bellow (in this benchmark the model is called andreaskoepf/llama2-13b-megacode2_min100)
- https://tju01.github.io/FastEval-OpenAssistant/
Prompt template:
- chatml format is used: "<|im_start|>system\n{system message}<|im_end|>\n<|im_start|>user\n{user prompt}<|im_end|>\n<|im_start|>assistant\n{Assistant answer}<|im_end|>\n"
multi-line:
```
<|im_start|>system
{system message}<|im_end|>
<|im_start|>user
{user prompt}<|im_end|>
<|im_start|>assistant
{Assistant answer}<|im_end|>
```
Gpt4all template:
- System prompt
```
<|im_start|>system
"Below is an instruction that describes a task. Write a response that appropriately completes the request."
```
- Prompt template
```
<|im_end|>
<|im_start|>user
"%1"<|im_end|>
<|im_start|>assistant
```
Oobagooba Text-Generation-Webui Template
- user:
```
<|im_start|>user
{User string}<|im_end|>
```
- bot:
```
<|im_start|>assistant
{Bot string}<|im_end|>
```
- turn_template:
```
<|user|>\n<|user-message|>\n\n<|bot|>\n<|bot-message|>\n\n
```
- context:
```
<|im_start|>system
Below is an instruction that describes a task. Write a response that appropriately completes the request.<|im_end|>
```
Current quantizations available:
- https://huggingface.co/TheBloke/LosslessMegaCoder-Llama2-13B-Mini-GPTQ
Training data:
- https://wandb.ai/open-assistant/epfl-mt-sft/runs/run34_megacode2_min100_13b
The link for the full dataset is bellow:
- https://huggingface.co/datasets/rombodawg/LosslessMegaCodeTrainingV2_1m_Evol_Uncensored
Link for the filtered dataset used to make this model are bellow:
- https://huggingface.co/datasets/andreaskoepf/megacode2-min100
The original posting for this model was uploaded at the link bellow.
- https://huggingface.co/andreaskoepf/llama2-13b-megacode2_min100
# [Open LLM Leaderboard Evaluation Results](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard)
Detailed results can be found [here](https://huggingface.co/datasets/open-llm-leaderboard/details_rombodawg__LosslessMegaCoder-llama2-13b-mini)
| Metric | Value |
|-----------------------|---------------------------|
| Avg. | 49.92 |
| ARC (25-shot) | 60.58 |
| HellaSwag (10-shot) | 81.26 |
| MMLU (5-shot) | 57.92 |
| TruthfulQA (0-shot) | 48.89 |
| Winogrande (5-shot) | 76.95 |
| GSM8K (5-shot) | 15.92 |
| DROP (3-shot) | 7.89 |

9
added_tokens.json Normal file
View File

@@ -0,0 +1,9 @@
{
"<CLS>": 32000,
"<EOD>": 32002,
"<MASK>": 32003,
"<PAD>": 32004,
"<SEP>": 32001,
"<|im_end|>": 32006,
"<|im_start|>": 32005
}

25
config.json Normal file
View File

@@ -0,0 +1,25 @@
{
"architectures": [
"LlamaForCausalLM"
],
"bos_token_id": 1,
"eos_token_id": 32006,
"hidden_act": "silu",
"hidden_size": 5120,
"initializer_range": 0.02,
"intermediate_size": 13824,
"max_position_embeddings": 2048,
"model_type": "llama",
"num_attention_heads": 40,
"num_hidden_layers": 40,
"num_key_value_heads": 40,
"pad_token_id": 0,
"pretraining_tp": 1,
"rms_norm_eps": 1e-05,
"rope_scaling": null,
"tie_word_embeddings": false,
"torch_dtype": "bfloat16",
"transformers_version": "4.31.0",
"use_cache": true,
"vocab_size": 32007
}

6
generation_config.json Normal file
View File

@@ -0,0 +1,6 @@
{
"bos_token_id": 1,
"eos_token_id": 32006,
"pad_token_id": 0,
"transformers_version": "4.31.0"
}

13
huggingface-metadata.txt Normal file
View File

@@ -0,0 +1,13 @@
url: https://huggingface.co/andreaskoepf/llama2-13b-megacode2_min100
branch: main
download date: 2023-08-14 19:58:21
sha256sum:
68c2905e84622ab407e5d957c5db51e13bf56aeb2f893d5b640c53792be9902b pytorch_model-00001-of-00008.bin
0ac2e52618efb80d24a77ef275cd9a987897956ed65a2b7bc07a35ab42878d8e pytorch_model-00002-of-00008.bin
027c243ae0dfcd3445b7a2ada510ebb6cc46c7eb0c7423bf529a6bf37f4de881 pytorch_model-00003-of-00008.bin
2a35b79e40a33abc4b1fb899ee5c935d37b3a51ad1fe8ea659705f2b24c4e33b pytorch_model-00004-of-00008.bin
20e26f0ccc5f7d77a21191d6b4352c1619a9752dabfddafcd8fb18de2f2a089a pytorch_model-00005-of-00008.bin
a5d4c91832515da5f8382158343e7bdfdb388a50c8c96792c95906c4220f41cc pytorch_model-00006-of-00008.bin
f58042fb65429539213c6a77f90f62e72b59b4086c1f00101acf2be89ad8e83f pytorch_model-00007-of-00008.bin
659119c0075b6d40340077da1b3f47ab043e0fa37cc1f2549a0226fad631cc5c pytorch_model-00008-of-00008.bin
9e556afd44213b6bd1be2b850ebbbd98f5481437a8021afaf58ee7fb1818d347 tokenizer.model

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:68c2905e84622ab407e5d957c5db51e13bf56aeb2f893d5b640c53792be9902b
size 3709532442

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0ac2e52618efb80d24a77ef275cd9a987897956ed65a2b7bc07a35ab42878d8e
size 3701617004

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:027c243ae0dfcd3445b7a2ada510ebb6cc46c7eb0c7423bf529a6bf37f4de881
size 3701617662

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2a35b79e40a33abc4b1fb899ee5c935d37b3a51ad1fe8ea659705f2b24c4e33b
size 3664896684

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:20e26f0ccc5f7d77a21191d6b4352c1619a9752dabfddafcd8fb18de2f2a089a
size 3664917840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a5d4c91832515da5f8382158343e7bdfdb388a50c8c96792c95906c4220f41cc
size 3664917840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f58042fb65429539213c6a77f90f62e72b59b4086c1f00101acf2be89ad8e83f
size 3596769306

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:659119c0075b6d40340077da1b3f47ab043e0fa37cc1f2549a0226fad631cc5c
size 327753093

View File

@@ -0,0 +1,410 @@
{
"metadata": {
"total_size": 26031882240
},
"weight_map": {
"lm_head.weight": "pytorch_model-00008-of-00008.bin",
"model.embed_tokens.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.0.input_layernorm.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.0.mlp.down_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.0.mlp.gate_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.0.mlp.up_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.0.post_attention_layernorm.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.0.self_attn.k_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.0.self_attn.o_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.0.self_attn.q_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.0.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00008.bin",
"model.layers.0.self_attn.v_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.1.input_layernorm.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.1.mlp.down_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.1.mlp.gate_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.1.mlp.up_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.1.post_attention_layernorm.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.1.self_attn.k_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.1.self_attn.o_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.1.self_attn.q_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.1.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00008.bin",
"model.layers.1.self_attn.v_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.10.input_layernorm.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.10.mlp.down_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.10.mlp.gate_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.10.mlp.up_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.10.post_attention_layernorm.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.10.self_attn.k_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.10.self_attn.o_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.10.self_attn.q_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.10.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00008.bin",
"model.layers.10.self_attn.v_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.11.input_layernorm.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.11.mlp.down_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.11.mlp.gate_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.11.mlp.up_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.11.post_attention_layernorm.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.11.self_attn.k_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.11.self_attn.o_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.11.self_attn.q_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.11.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00008.bin",
"model.layers.11.self_attn.v_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.12.input_layernorm.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.12.mlp.down_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.12.mlp.gate_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.12.mlp.up_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.12.post_attention_layernorm.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.12.self_attn.k_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.12.self_attn.o_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.12.self_attn.q_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.12.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00008.bin",
"model.layers.12.self_attn.v_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.13.input_layernorm.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.13.mlp.down_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.13.mlp.gate_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.13.mlp.up_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.13.post_attention_layernorm.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.13.self_attn.k_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.13.self_attn.o_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.13.self_attn.q_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.13.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00008.bin",
"model.layers.13.self_attn.v_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.14.input_layernorm.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.14.mlp.down_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.14.mlp.gate_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.14.mlp.up_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.14.post_attention_layernorm.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.14.self_attn.k_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.14.self_attn.o_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.14.self_attn.q_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.14.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00008.bin",
"model.layers.14.self_attn.v_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.15.input_layernorm.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.15.mlp.down_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.15.mlp.gate_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.15.mlp.up_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.15.post_attention_layernorm.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.15.self_attn.k_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.15.self_attn.o_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.15.self_attn.q_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.15.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00008.bin",
"model.layers.15.self_attn.v_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.16.input_layernorm.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.16.mlp.down_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.16.mlp.gate_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.16.mlp.up_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.16.post_attention_layernorm.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.16.self_attn.k_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.16.self_attn.o_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.16.self_attn.q_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.16.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00008.bin",
"model.layers.16.self_attn.v_proj.weight": "pytorch_model-00003-of-00008.bin",
"model.layers.17.input_layernorm.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.17.mlp.down_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.17.mlp.gate_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.17.mlp.up_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.17.post_attention_layernorm.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.17.self_attn.k_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.17.self_attn.o_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.17.self_attn.q_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.17.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00008.bin",
"model.layers.17.self_attn.v_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.18.input_layernorm.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.18.mlp.down_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.18.mlp.gate_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.18.mlp.up_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.18.post_attention_layernorm.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.18.self_attn.k_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.18.self_attn.o_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.18.self_attn.q_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.18.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00008.bin",
"model.layers.18.self_attn.v_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.19.input_layernorm.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.19.mlp.down_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.19.mlp.gate_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.19.mlp.up_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.19.post_attention_layernorm.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.19.self_attn.k_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.19.self_attn.o_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.19.self_attn.q_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.19.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00008.bin",
"model.layers.19.self_attn.v_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.2.input_layernorm.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.2.mlp.down_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.2.mlp.gate_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.2.mlp.up_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.2.post_attention_layernorm.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.2.self_attn.k_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.2.self_attn.o_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.2.self_attn.q_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.2.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00008.bin",
"model.layers.2.self_attn.v_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.20.input_layernorm.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.20.mlp.down_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.20.mlp.gate_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.20.mlp.up_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.20.post_attention_layernorm.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.20.self_attn.k_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.20.self_attn.o_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.20.self_attn.q_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.20.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00008.bin",
"model.layers.20.self_attn.v_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.21.input_layernorm.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.21.mlp.down_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.21.mlp.gate_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.21.mlp.up_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.21.post_attention_layernorm.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.21.self_attn.k_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.21.self_attn.o_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.21.self_attn.q_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.21.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00008.bin",
"model.layers.21.self_attn.v_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.22.input_layernorm.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.22.mlp.down_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.22.mlp.gate_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.22.mlp.up_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.22.post_attention_layernorm.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.22.self_attn.k_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.22.self_attn.o_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.22.self_attn.q_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.22.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00008.bin",
"model.layers.22.self_attn.v_proj.weight": "pytorch_model-00004-of-00008.bin",
"model.layers.23.input_layernorm.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.23.mlp.down_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.23.mlp.gate_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.23.mlp.up_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.23.post_attention_layernorm.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.23.self_attn.k_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.23.self_attn.o_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.23.self_attn.q_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.23.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00008.bin",
"model.layers.23.self_attn.v_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.24.input_layernorm.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.24.mlp.down_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.24.mlp.gate_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.24.mlp.up_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.24.post_attention_layernorm.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.24.self_attn.k_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.24.self_attn.o_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.24.self_attn.q_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.24.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00008.bin",
"model.layers.24.self_attn.v_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.25.input_layernorm.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.25.mlp.down_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.25.mlp.gate_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.25.mlp.up_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.25.post_attention_layernorm.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.25.self_attn.k_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.25.self_attn.o_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.25.self_attn.q_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.25.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00008.bin",
"model.layers.25.self_attn.v_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.26.input_layernorm.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.26.mlp.down_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.26.mlp.gate_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.26.mlp.up_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.26.post_attention_layernorm.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.26.self_attn.k_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.26.self_attn.o_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.26.self_attn.q_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.26.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00008.bin",
"model.layers.26.self_attn.v_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.27.input_layernorm.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.27.mlp.down_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.27.mlp.gate_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.27.mlp.up_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.27.post_attention_layernorm.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.27.self_attn.k_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.27.self_attn.o_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.27.self_attn.q_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.27.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00008.bin",
"model.layers.27.self_attn.v_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.28.input_layernorm.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.28.mlp.down_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.28.mlp.gate_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.28.mlp.up_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.28.post_attention_layernorm.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.28.self_attn.k_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.28.self_attn.o_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.28.self_attn.q_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.28.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00008.bin",
"model.layers.28.self_attn.v_proj.weight": "pytorch_model-00005-of-00008.bin",
"model.layers.29.input_layernorm.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.29.mlp.down_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.29.mlp.gate_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.29.mlp.up_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.29.post_attention_layernorm.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.29.self_attn.k_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.29.self_attn.o_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.29.self_attn.q_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.29.self_attn.rotary_emb.inv_freq": "pytorch_model-00006-of-00008.bin",
"model.layers.29.self_attn.v_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.3.input_layernorm.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.3.mlp.down_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.3.mlp.gate_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.3.mlp.up_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.3.post_attention_layernorm.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.3.self_attn.k_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.3.self_attn.o_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.3.self_attn.q_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.3.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00008.bin",
"model.layers.3.self_attn.v_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.30.input_layernorm.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.30.mlp.down_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.30.mlp.gate_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.30.mlp.up_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.30.post_attention_layernorm.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.30.self_attn.k_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.30.self_attn.o_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.30.self_attn.q_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.30.self_attn.rotary_emb.inv_freq": "pytorch_model-00006-of-00008.bin",
"model.layers.30.self_attn.v_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.31.input_layernorm.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.31.mlp.down_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.31.mlp.gate_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.31.mlp.up_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.31.post_attention_layernorm.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.31.self_attn.k_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.31.self_attn.o_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.31.self_attn.q_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.31.self_attn.rotary_emb.inv_freq": "pytorch_model-00006-of-00008.bin",
"model.layers.31.self_attn.v_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.32.input_layernorm.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.32.mlp.down_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.32.mlp.gate_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.32.mlp.up_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.32.post_attention_layernorm.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.32.self_attn.k_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.32.self_attn.o_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.32.self_attn.q_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.32.self_attn.rotary_emb.inv_freq": "pytorch_model-00006-of-00008.bin",
"model.layers.32.self_attn.v_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.33.input_layernorm.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.33.mlp.down_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.33.mlp.gate_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.33.mlp.up_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.33.post_attention_layernorm.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.33.self_attn.k_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.33.self_attn.o_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.33.self_attn.q_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.33.self_attn.rotary_emb.inv_freq": "pytorch_model-00006-of-00008.bin",
"model.layers.33.self_attn.v_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.34.input_layernorm.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.34.mlp.down_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.34.mlp.gate_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.34.mlp.up_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.34.post_attention_layernorm.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.34.self_attn.k_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.34.self_attn.o_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.34.self_attn.q_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.34.self_attn.rotary_emb.inv_freq": "pytorch_model-00006-of-00008.bin",
"model.layers.34.self_attn.v_proj.weight": "pytorch_model-00006-of-00008.bin",
"model.layers.35.input_layernorm.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.35.mlp.down_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.35.mlp.gate_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.35.mlp.up_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.35.post_attention_layernorm.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.35.self_attn.k_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.35.self_attn.o_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.35.self_attn.q_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.35.self_attn.rotary_emb.inv_freq": "pytorch_model-00007-of-00008.bin",
"model.layers.35.self_attn.v_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.36.input_layernorm.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.36.mlp.down_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.36.mlp.gate_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.36.mlp.up_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.36.post_attention_layernorm.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.36.self_attn.k_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.36.self_attn.o_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.36.self_attn.q_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.36.self_attn.rotary_emb.inv_freq": "pytorch_model-00007-of-00008.bin",
"model.layers.36.self_attn.v_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.37.input_layernorm.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.37.mlp.down_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.37.mlp.gate_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.37.mlp.up_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.37.post_attention_layernorm.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.37.self_attn.k_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.37.self_attn.o_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.37.self_attn.q_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.37.self_attn.rotary_emb.inv_freq": "pytorch_model-00007-of-00008.bin",
"model.layers.37.self_attn.v_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.38.input_layernorm.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.38.mlp.down_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.38.mlp.gate_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.38.mlp.up_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.38.post_attention_layernorm.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.38.self_attn.k_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.38.self_attn.o_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.38.self_attn.q_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.38.self_attn.rotary_emb.inv_freq": "pytorch_model-00007-of-00008.bin",
"model.layers.38.self_attn.v_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.39.input_layernorm.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.39.mlp.down_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.39.mlp.gate_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.39.mlp.up_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.39.post_attention_layernorm.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.39.self_attn.k_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.39.self_attn.o_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.39.self_attn.q_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.39.self_attn.rotary_emb.inv_freq": "pytorch_model-00007-of-00008.bin",
"model.layers.39.self_attn.v_proj.weight": "pytorch_model-00007-of-00008.bin",
"model.layers.4.input_layernorm.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.4.mlp.down_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.4.mlp.gate_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.4.mlp.up_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.4.post_attention_layernorm.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.4.self_attn.k_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.4.self_attn.o_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.4.self_attn.q_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.4.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00008.bin",
"model.layers.4.self_attn.v_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.5.input_layernorm.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.5.mlp.down_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.5.mlp.gate_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.5.mlp.up_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.5.post_attention_layernorm.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.5.self_attn.k_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.5.self_attn.o_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.5.self_attn.q_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.5.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00008.bin",
"model.layers.5.self_attn.v_proj.weight": "pytorch_model-00001-of-00008.bin",
"model.layers.6.input_layernorm.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.6.mlp.down_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.6.mlp.gate_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.6.mlp.up_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.6.post_attention_layernorm.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.6.self_attn.k_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.6.self_attn.o_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.6.self_attn.q_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.6.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00008.bin",
"model.layers.6.self_attn.v_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.7.input_layernorm.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.7.mlp.down_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.7.mlp.gate_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.7.mlp.up_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.7.post_attention_layernorm.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.7.self_attn.k_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.7.self_attn.o_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.7.self_attn.q_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.7.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00008.bin",
"model.layers.7.self_attn.v_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.8.input_layernorm.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.8.mlp.down_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.8.mlp.gate_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.8.mlp.up_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.8.post_attention_layernorm.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.8.self_attn.k_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.8.self_attn.o_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.8.self_attn.q_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.8.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00008.bin",
"model.layers.8.self_attn.v_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.9.input_layernorm.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.9.mlp.down_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.9.mlp.gate_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.9.mlp.up_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.9.post_attention_layernorm.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.9.self_attn.k_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.9.self_attn.o_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.9.self_attn.q_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.layers.9.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00008.bin",
"model.layers.9.self_attn.v_proj.weight": "pytorch_model-00002-of-00008.bin",
"model.norm.weight": "pytorch_model-00007-of-00008.bin"
}
}

31
special_tokens_map.json Normal file
View File

@@ -0,0 +1,31 @@
{
"additional_special_tokens": [
"<|im_start|>",
"<|im_end|>"
],
"bos_token": {
"content": "<s>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"cls_token": "<CLS>",
"eos_token": {
"content": "</s>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"mask_token": "<MASK>",
"pad_token": "<PAD>",
"sep_token": "<SEP>",
"unk_token": {
"content": "<unk>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
}
}

93454
tokenizer.json Normal file

File diff suppressed because it is too large Load Diff

BIN
tokenizer.model (Stored with Git LFS) Normal file

Binary file not shown.

33
tokenizer_config.json Normal file
View File

@@ -0,0 +1,33 @@
{
"bos_token": {
"__type": "AddedToken",
"content": "<s>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"clean_up_tokenization_spaces": false,
"eos_token": {
"__type": "AddedToken",
"content": "</s>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"legacy": false,
"model_max_length": 1000000000000000019884624838656,
"pad_token": null,
"padding_side": "right",
"sp_model_kwargs": {},
"tokenizer_class": "LlamaTokenizer",
"unk_token": {
"__type": "AddedToken",
"content": "<unk>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
}
}