初始化项目,由ModelHub XC社区提供模型

Model: defog/sqlcoder-34b-alpha
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-20 14:30:08 +08:00
commit 559f4795ca
16 changed files with 755 additions and 0 deletions

56
.gitattributes vendored Normal file
View File

@@ -0,0 +1,56 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
pytorch_model-00004-of-00007.bin filter=lfs diff=lfs merge=lfs -text
pytorch_model-00001-of-00007.bin filter=lfs diff=lfs merge=lfs -text
tokenizer.json filter=lfs diff=lfs merge=lfs -text
pytorch_model-00005-of-00007.bin filter=lfs diff=lfs merge=lfs -text
pytorch_model-00006-of-00007.bin filter=lfs diff=lfs merge=lfs -text
pytorch_model-00002-of-00007.bin filter=lfs diff=lfs merge=lfs -text
pytorch_model-00007-of-00007.bin filter=lfs diff=lfs merge=lfs -text
pytorch_model-00003-of-00007.bin filter=lfs diff=lfs merge=lfs -text

81
README.md Normal file
View File

@@ -0,0 +1,81 @@
---
license: cc-by-4.0
language:
- en
pipeline_tag: text-generation
---
# Defog SQLCoder
**Updated on Nov 14 to reflect benchmarks for SQLCoder-34B**
Defog's SQLCoder is a state-of-the-art LLM for converting natural language questions to SQL queries.
[Interactive Demo](https://defog.ai/sqlcoder-demo/) | [🤗 HF Repo](https://huggingface.co/defog/sqlcoder-34b-alpha) | [♾️ Colab](https://colab.research.google.com/drive/1z4rmOEiFkxkMiecAWeTUlPl0OmKgfEu7?usp=sharing) | [🐦 Twitter](https://twitter.com/defogdata)
## TL;DR
SQLCoder-34B is a 34B parameter model that outperforms `gpt-4` and `gpt-4-turbo` for natural language to SQL generation tasks on our [sql-eval](https://github.com/defog-ai/sql-eval) framework, and significantly outperforms all popular open-source models.
SQLCoder-34B is fine-tuned on a base CodeLlama model.
## Results on novel datasets not seen in training
| model | perc_correct |
|-|-|
| defog-sqlcoder-34b | 84.0 |
| gpt4-turbo-2023-11-09 | 82.5 |
| gpt4-2023-11-09 | 82.5 |
| defog-sqlcoder2 | 77.5 |
| gpt4-2023-08-28 | 74.0 |
| defog-sqlcoder-7b | 71.0 |
| gpt-3.5-2023-10-04 | 66.0 |
| claude-2 | 64.5 |
| gpt-3.5-2023-08-28 | 61.0 |
| claude_instant_1 | 61.0 |
| text-davinci-003 | 52.5 |
![image](https://github.com/defog-ai/sqlcoder/assets/5008293/caed3423-8e86-4952-9da1-1a5e016a4696)
## License
The code in this repo (what little there is of it) is Apache-2 licensed. The model weights have a `CC BY-SA 4.0` license. The TL;DR is that you can use and modify the model for any purpose including commercial use. However, if you modify the weights (for example, by fine-tuning), you must open-source your modified weights under the same license terms.
## Training
Defog was trained on more than 20,000 human-curated questions. These questions were based on 10 different schemas. None of the schemas in the training data were included in our evaluation framework.
You can read more about our [training approach](https://defog.ai/blog/open-sourcing-sqlcoder2-7b/) and [evaluation framework](https://defog.ai/blog/open-sourcing-sqleval/).
## Results by question category
We classified each generated question into one of 5 categories. The table displays the percentage of questions answered correctly by each model, broken down by category.
| | date | group_by | order_by | ratio | join | where |
| -------------- | ---- | -------- | -------- | ----- | ---- | ----- |
| sqlcoder-34b | 80 | 94.3 | 88.6 | 74.3 | 82.9 | 82.9 |
| gpt-4 | 68 | 94.3 | 85.7 | 77.1 | 85.7 | 80 |
| sqlcoder2-15b | 76 | 80 | 77.1 | 60 | 77.1 | 77.1 |
| sqlcoder-7b | 64 | 82.9 | 74.3 | 54.3 | 74.3 | 74.3 |
| gpt-3.5 | 68 | 77.1 | 68.6 | 37.1 | 71.4 | 74.3 |
| claude-2 | 52 | 71.4 | 74.3 | 57.1 | 65.7 | 62.9 |
| claude-instant | 48 | 71.4 | 74.3 | 45.7 | 62.9 | 60 |
| gpt-3 | 32 | 71.4 | 68.6 | 25.7 | 57.1 | 54.3 |
<img width="831" alt="image" src="https://github.com/defog-ai/sqlcoder/assets/5008293/79c5bdc8-373c-4abd-822e-e2c2569ed353">
## Using SQLCoder
You can use SQLCoder via the `transformers` library by downloading our model weights from the Hugging Face repo. We have added sample code for [inference](./inference.py) on a [sample database schema](./metadata.sql).
```bash
python inference.py -q "Question about the sample database goes here"
# Sample question:
# Do we get more revenue from customers in New York compared to customers in San Francisco? Give me the total revenue for each city, and the difference between the two.
```
You can also use a demo on our website [here](https://defog.ai/sqlcoder-demo)
## Hardware Requirements
SQLCoder-34B has been tested on a 4xA10 GPU with `float16` weights. You can also load an 8-bit and 4-bit quantized version of the model on consumer GPUs with 20GB or more of memory  like RTX 4090, RTX 3090, and Apple M2 Pro, M2 Max, or M2 Ultra Chips with 20GB or more of memory.
## Todo
- [x] Open-source the v1 model weights
- [x] Train the model on more data, with higher data variance
- [ ] Tune the model further with Reward Modelling and RLHF
- [ ] Pretrain a model from scratch that specializes in SQL analysis

27
config.json Normal file
View File

@@ -0,0 +1,27 @@
{
"_name_or_path": "defog/sqlcoder-34b",
"architectures": [
"LlamaForCausalLM"
],
"attention_bias": false,
"bos_token_id": 1,
"eos_token_id": 2,
"hidden_act": "silu",
"hidden_size": 8192,
"initializer_range": 0.02,
"intermediate_size": 22016,
"max_position_embeddings": 16384,
"model_type": "llama",
"num_attention_heads": 64,
"num_hidden_layers": 48,
"num_key_value_heads": 8,
"pretraining_tp": 1,
"rms_norm_eps": 1e-05,
"rope_scaling": null,
"rope_theta": 1000000,
"tie_word_embeddings": false,
"torch_dtype": "float16",
"transformers_version": "4.34.1",
"use_cache": true,
"vocab_size": 32000
}

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}

6
generation_config.json Normal file
View File

@@ -0,0 +1,6 @@
{
"_from_model_config": true,
"bos_token_id": 1,
"eos_token_id": 2,
"transformers_version": "4.34.1"
}

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8812e55851943abc023ed4e5af05fbca20220a7a0e01baef3e2eeadd7736e304
size 9852638880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d468a9e9aebf8dafa35244c7bccaf998467b23c8146e0967c620b1fb7a350eba
size 9689094520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:084d0930019aa448abae86cbee779094632b8d4c89e0cf7423ed28615a3f817a
size 9689094520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b36adc5ae895a6fe3e806ab89e2cca380b6505910281c940894b437a0dd957d9
size 9689094520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1cba9e25a3e429ed83ce073fd0f4ee4366863db88cbc5a64eeabd363e7f5d613
size 9689094520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e11b17daf17d28e5de1c747917927cf58d12f47bf4c7377f71a553203d54fd14
size 9689094520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:804a80eeb98f146424257c25e10d907634886b6c0cb66718cb8cc20c12af4255
size 9189987200

View File

@@ -0,0 +1,442 @@
{
"metadata": {
"total_size": 67487940608
},
"weight_map": {
"lm_head.weight": "pytorch_model-00007-of-00007.bin",
"model.embed_tokens.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.0.input_layernorm.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.0.mlp.down_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.0.mlp.gate_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.0.mlp.up_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.0.post_attention_layernorm.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.0.self_attn.k_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.0.self_attn.o_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.0.self_attn.q_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.0.self_attn.v_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.1.input_layernorm.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.1.mlp.down_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.1.mlp.gate_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.1.mlp.up_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.1.post_attention_layernorm.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.1.self_attn.k_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.1.self_attn.o_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.1.self_attn.q_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.1.self_attn.v_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.10.input_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.10.mlp.down_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.10.mlp.gate_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.10.mlp.up_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.10.post_attention_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.10.self_attn.k_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.10.self_attn.o_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.10.self_attn.q_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.10.self_attn.v_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.11.input_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.11.mlp.down_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.11.mlp.gate_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.11.mlp.up_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.11.post_attention_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.11.self_attn.k_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.11.self_attn.o_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.11.self_attn.q_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.11.self_attn.v_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.12.input_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.12.mlp.down_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.12.mlp.gate_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.12.mlp.up_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.12.post_attention_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.12.self_attn.k_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.12.self_attn.o_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.12.self_attn.q_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.12.self_attn.v_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.13.input_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.13.mlp.down_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.13.mlp.gate_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.13.mlp.up_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.13.post_attention_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.13.self_attn.k_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.13.self_attn.o_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.13.self_attn.q_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.13.self_attn.v_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.14.input_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.14.mlp.down_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.14.mlp.gate_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.14.mlp.up_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.14.post_attention_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.14.self_attn.k_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.14.self_attn.o_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.14.self_attn.q_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.14.self_attn.v_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.15.input_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.15.mlp.down_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.15.mlp.gate_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.15.mlp.up_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.15.post_attention_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.15.self_attn.k_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.15.self_attn.o_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.15.self_attn.q_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.15.self_attn.v_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.16.input_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.16.mlp.down_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.16.mlp.gate_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.16.mlp.up_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.16.post_attention_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.16.self_attn.k_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.16.self_attn.o_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.16.self_attn.q_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.16.self_attn.v_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.17.input_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.17.mlp.down_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.17.mlp.gate_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.17.mlp.up_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.17.post_attention_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.17.self_attn.k_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.17.self_attn.o_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.17.self_attn.q_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.17.self_attn.v_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.18.input_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.18.mlp.down_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.18.mlp.gate_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.18.mlp.up_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.18.post_attention_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.18.self_attn.k_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.18.self_attn.o_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.18.self_attn.q_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.18.self_attn.v_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.19.input_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.19.mlp.down_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.19.mlp.gate_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.19.mlp.up_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.19.post_attention_layernorm.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.19.self_attn.k_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.19.self_attn.o_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.19.self_attn.q_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.19.self_attn.v_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.2.input_layernorm.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.2.mlp.down_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.2.mlp.gate_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.2.mlp.up_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.2.post_attention_layernorm.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.2.self_attn.k_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.2.self_attn.o_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.2.self_attn.q_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.2.self_attn.v_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.20.input_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.20.mlp.down_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.20.mlp.gate_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.20.mlp.up_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.20.post_attention_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.20.self_attn.k_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.20.self_attn.o_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.20.self_attn.q_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.20.self_attn.v_proj.weight": "pytorch_model-00003-of-00007.bin",
"model.layers.21.input_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.21.mlp.down_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.21.mlp.gate_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.21.mlp.up_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.21.post_attention_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.21.self_attn.k_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.21.self_attn.o_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.21.self_attn.q_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.21.self_attn.v_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.22.input_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.22.mlp.down_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.22.mlp.gate_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.22.mlp.up_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.22.post_attention_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.22.self_attn.k_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.22.self_attn.o_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.22.self_attn.q_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.22.self_attn.v_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.23.input_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.23.mlp.down_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.23.mlp.gate_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.23.mlp.up_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.23.post_attention_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.23.self_attn.k_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.23.self_attn.o_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.23.self_attn.q_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.23.self_attn.v_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.24.input_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.24.mlp.down_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.24.mlp.gate_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.24.mlp.up_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.24.post_attention_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.24.self_attn.k_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.24.self_attn.o_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.24.self_attn.q_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.24.self_attn.v_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.25.input_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.25.mlp.down_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.25.mlp.gate_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.25.mlp.up_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.25.post_attention_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.25.self_attn.k_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.25.self_attn.o_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.25.self_attn.q_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.25.self_attn.v_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.26.input_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.26.mlp.down_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.26.mlp.gate_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.26.mlp.up_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.26.post_attention_layernorm.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.26.self_attn.k_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.26.self_attn.o_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.26.self_attn.q_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.26.self_attn.v_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.27.input_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.27.mlp.down_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.27.mlp.gate_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.27.mlp.up_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.27.post_attention_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.27.self_attn.k_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.27.self_attn.o_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.27.self_attn.q_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.27.self_attn.v_proj.weight": "pytorch_model-00004-of-00007.bin",
"model.layers.28.input_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.28.mlp.down_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.28.mlp.gate_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.28.mlp.up_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.28.post_attention_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.28.self_attn.k_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.28.self_attn.o_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.28.self_attn.q_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.28.self_attn.v_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.29.input_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.29.mlp.down_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.29.mlp.gate_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.29.mlp.up_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.29.post_attention_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.29.self_attn.k_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.29.self_attn.o_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.29.self_attn.q_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.29.self_attn.v_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.3.input_layernorm.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.3.mlp.down_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.3.mlp.gate_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.3.mlp.up_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.3.post_attention_layernorm.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.3.self_attn.k_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.3.self_attn.o_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.3.self_attn.q_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.3.self_attn.v_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.30.input_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.30.mlp.down_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.30.mlp.gate_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.30.mlp.up_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.30.post_attention_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.30.self_attn.k_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.30.self_attn.o_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.30.self_attn.q_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.30.self_attn.v_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.31.input_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.31.mlp.down_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.31.mlp.gate_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.31.mlp.up_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.31.post_attention_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.31.self_attn.k_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.31.self_attn.o_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.31.self_attn.q_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.31.self_attn.v_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.32.input_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.32.mlp.down_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.32.mlp.gate_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.32.mlp.up_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.32.post_attention_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.32.self_attn.k_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.32.self_attn.o_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.32.self_attn.q_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.32.self_attn.v_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.33.input_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.33.mlp.down_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.33.mlp.gate_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.33.mlp.up_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.33.post_attention_layernorm.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.33.self_attn.k_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.33.self_attn.o_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.33.self_attn.q_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.33.self_attn.v_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.34.input_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.34.mlp.down_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.34.mlp.gate_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.34.mlp.up_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.34.post_attention_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.34.self_attn.k_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.34.self_attn.o_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.34.self_attn.q_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.34.self_attn.v_proj.weight": "pytorch_model-00005-of-00007.bin",
"model.layers.35.input_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.35.mlp.down_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.35.mlp.gate_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.35.mlp.up_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.35.post_attention_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.35.self_attn.k_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.35.self_attn.o_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.35.self_attn.q_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.35.self_attn.v_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.36.input_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.36.mlp.down_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.36.mlp.gate_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.36.mlp.up_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.36.post_attention_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.36.self_attn.k_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.36.self_attn.o_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.36.self_attn.q_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.36.self_attn.v_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.37.input_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.37.mlp.down_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.37.mlp.gate_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.37.mlp.up_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.37.post_attention_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.37.self_attn.k_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.37.self_attn.o_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.37.self_attn.q_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.37.self_attn.v_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.38.input_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.38.mlp.down_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.38.mlp.gate_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.38.mlp.up_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.38.post_attention_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.38.self_attn.k_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.38.self_attn.o_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.38.self_attn.q_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.38.self_attn.v_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.39.input_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.39.mlp.down_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.39.mlp.gate_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.39.mlp.up_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.39.post_attention_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.39.self_attn.k_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.39.self_attn.o_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.39.self_attn.q_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.39.self_attn.v_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.4.input_layernorm.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.4.mlp.down_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.4.mlp.gate_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.4.mlp.up_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.4.post_attention_layernorm.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.4.self_attn.k_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.4.self_attn.o_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.4.self_attn.q_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.4.self_attn.v_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.40.input_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.40.mlp.down_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.40.mlp.gate_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.40.mlp.up_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.40.post_attention_layernorm.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.40.self_attn.k_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.40.self_attn.o_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.40.self_attn.q_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.40.self_attn.v_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.41.input_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.41.mlp.down_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.41.mlp.gate_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.41.mlp.up_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.41.post_attention_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.41.self_attn.k_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.41.self_attn.o_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.41.self_attn.q_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.41.self_attn.v_proj.weight": "pytorch_model-00006-of-00007.bin",
"model.layers.42.input_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.42.mlp.down_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.42.mlp.gate_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.42.mlp.up_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.42.post_attention_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.42.self_attn.k_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.42.self_attn.o_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.42.self_attn.q_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.42.self_attn.v_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.43.input_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.43.mlp.down_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.43.mlp.gate_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.43.mlp.up_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.43.post_attention_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.43.self_attn.k_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.43.self_attn.o_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.43.self_attn.q_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.43.self_attn.v_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.44.input_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.44.mlp.down_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.44.mlp.gate_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.44.mlp.up_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.44.post_attention_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.44.self_attn.k_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.44.self_attn.o_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.44.self_attn.q_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.44.self_attn.v_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.45.input_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.45.mlp.down_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.45.mlp.gate_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.45.mlp.up_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.45.post_attention_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.45.self_attn.k_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.45.self_attn.o_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.45.self_attn.q_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.45.self_attn.v_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.46.input_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.46.mlp.down_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.46.mlp.gate_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.46.mlp.up_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.46.post_attention_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.46.self_attn.k_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.46.self_attn.o_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.46.self_attn.q_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.46.self_attn.v_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.47.input_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.47.mlp.down_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.47.mlp.gate_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.47.mlp.up_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.47.post_attention_layernorm.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.47.self_attn.k_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.47.self_attn.o_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.47.self_attn.q_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.47.self_attn.v_proj.weight": "pytorch_model-00007-of-00007.bin",
"model.layers.5.input_layernorm.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.5.mlp.down_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.5.mlp.gate_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.5.mlp.up_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.5.post_attention_layernorm.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.5.self_attn.k_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.5.self_attn.o_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.5.self_attn.q_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.5.self_attn.v_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.6.input_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.6.mlp.down_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.6.mlp.gate_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.6.mlp.up_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.6.post_attention_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.6.self_attn.k_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.6.self_attn.o_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.6.self_attn.q_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.6.self_attn.v_proj.weight": "pytorch_model-00001-of-00007.bin",
"model.layers.7.input_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.7.mlp.down_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.7.mlp.gate_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.7.mlp.up_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.7.post_attention_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.7.self_attn.k_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.7.self_attn.o_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.7.self_attn.q_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.7.self_attn.v_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.8.input_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.8.mlp.down_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.8.mlp.gate_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.8.mlp.up_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.8.post_attention_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.8.self_attn.k_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.8.self_attn.o_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.8.self_attn.q_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.8.self_attn.v_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.9.input_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.9.mlp.down_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.9.mlp.gate_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.9.mlp.up_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.9.post_attention_layernorm.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.9.self_attn.k_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.9.self_attn.o_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.9.self_attn.q_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.layers.9.self_attn.v_proj.weight": "pytorch_model-00002-of-00007.bin",
"model.norm.weight": "pytorch_model-00007-of-00007.bin"
}
}

33
special_tokens_map.json Normal file
View File

@@ -0,0 +1,33 @@
{
"additional_special_tokens": [
"▁<PRE>",
"▁<MID>",
"▁<SUF>",
"▁<EOT>",
"▁<PRE>",
"▁<MID>",
"▁<SUF>",
"▁<EOT>"
],
"bos_token": {
"content": "<s>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
},
"eos_token": {
"content": "</s>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
},
"unk_token": {
"content": "<unk>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
}
}

3
tokenizer.json Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1d0fad9b4268298c57c3510c076611229453b3145fe244aa84e52b102df4f152
size 1843492

85
tokenizer_config.json Normal file
View File

@@ -0,0 +1,85 @@
{
"added_tokens_decoder": {
"0": {
"content": "<unk>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false,
"special": true
},
"1": {
"content": "<s>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false,
"special": true
},
"2": {
"content": "</s>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false,
"special": true
},
"32000": {
"content": "▁<PRE>",
"lstrip": true,
"normalized": false,
"rstrip": true,
"single_word": false,
"special": true
},
"32001": {
"content": "▁<MID>",
"lstrip": true,
"normalized": false,
"rstrip": true,
"single_word": false,
"special": true
},
"32002": {
"content": "▁<SUF>",
"lstrip": true,
"normalized": false,
"rstrip": true,
"single_word": false,
"special": true
},
"32003": {
"content": "▁<EOT>",
"lstrip": true,
"normalized": false,
"rstrip": true,
"single_word": false,
"special": true
}
},
"additional_special_tokens": [
"▁<PRE>",
"▁<MID>",
"▁<SUF>",
"▁<EOT>",
"▁<PRE>",
"▁<MID>",
"▁<SUF>",
"▁<EOT>"
],
"bos_token": "<s>",
"clean_up_tokenization_spaces": false,
"eos_token": "</s>",
"eot_token": "▁<EOT>",
"fill_token": "<FILL_ME>",
"legacy": null,
"middle_token": "▁<MID>",
"model_max_length": 1000000000000000019884624838656,
"pad_token": null,
"prefix_token": "▁<PRE>",
"sp_model_kwargs": {},
"suffix_token": "▁<SUF>",
"tokenizer_class": "CodeLlamaTokenizer",
"unk_token": "<unk>",
"use_default_system_prompt": false
}