初始化项目，由ModelHub XC社区提供模型

Model: Xwin-LM/Xwin-Math-7B-V1.1 Source: Original Platform
2026-05-15 20:09:06 +08:00
commit 891a557084
12 changed files with 684 additions and 0 deletions
--- a/.gitattributes
+++ b/.gitattributes
@@ -0,0 +1,37 @@
 *.7z filter=lfs diff=lfs merge=lfs -text
 *.arrow filter=lfs diff=lfs merge=lfs -text
 *.bin filter=lfs diff=lfs merge=lfs -text
 *.bz2 filter=lfs diff=lfs merge=lfs -text
 *.ckpt filter=lfs diff=lfs merge=lfs -text
 *.ftz filter=lfs diff=lfs merge=lfs -text
 *.gz filter=lfs diff=lfs merge=lfs -text
 *.h5 filter=lfs diff=lfs merge=lfs -text
 *.joblib filter=lfs diff=lfs merge=lfs -text
 *.lfs.* filter=lfs diff=lfs merge=lfs -text
 *.mlmodel filter=lfs diff=lfs merge=lfs -text
 *.model filter=lfs diff=lfs merge=lfs -text
 *.msgpack filter=lfs diff=lfs merge=lfs -text
 *.npy filter=lfs diff=lfs merge=lfs -text
 *.npz filter=lfs diff=lfs merge=lfs -text
 *.onnx filter=lfs diff=lfs merge=lfs -text
 *.ot filter=lfs diff=lfs merge=lfs -text
 *.parquet filter=lfs diff=lfs merge=lfs -text
 *.pb filter=lfs diff=lfs merge=lfs -text
 *.pickle filter=lfs diff=lfs merge=lfs -text
 *.pkl filter=lfs diff=lfs merge=lfs -text
 *.pt filter=lfs diff=lfs merge=lfs -text
 *.pth filter=lfs diff=lfs merge=lfs -text
 *.rar filter=lfs diff=lfs merge=lfs -text
 *.safetensors filter=lfs diff=lfs merge=lfs -text
 saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.tar.* filter=lfs diff=lfs merge=lfs -text
 *.tar filter=lfs diff=lfs merge=lfs -text
 *.tflite filter=lfs diff=lfs merge=lfs -text
 *.tgz filter=lfs diff=lfs merge=lfs -text
 *.wasm filter=lfs diff=lfs merge=lfs -text
 *.xz filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
 . filter=lfs diff=lfs merge=lfs -text
 /openrlhf/RLHF/examples/scripts/ckpt/Xwin-Math-7B-V1.1 filter=lfs diff=lfs merge=lfs -text
--- a/README.md
+++ b/README.md
@@ -0,0 +1,207 @@
 ---
 license: llama2
 ---
 # Xwin-Math
 <p align="center">
  <a href="https://github.com/Xwin-LM/Xwin-LM/tree/main/Xwin-Math"><img src="https://img.shields.io/badge/GitHub-yellow.svg?style=social&logo=github"></a>
  <a href="https://huggingface.co/Xwin-LM"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Models-blue"></a>
 </p>
 [Paper Link](https://arxiv.org/pdf/2403.04706) Xwin-Math is a series of powerful SFT LLMs for math problems based on LLaMA-2. 
 ## 🔥 News
 - 💥 [May, 2024] The [Xwin-Math-70B-V1.1](https://huggingface.co/Xwin-LM/Xwin-Math-70B-V1.1) model achieves **51.9 pass@1 on the MATH benchmark** and **90.6 pass@1 on the GSM8K benchmark**. This is a new SoTA model based on LLaMA-2-70B!
 - 💥 [May, 2024] The [Xwin-Math-7B-V1.1](https://huggingface.co/Xwin-LM/Xwin-Math-7B-V1.1) model achieves **44.7 pass@1 on the MATH benchmark** and **84.4 pass@1 on the GSM8K benchmark**. This is a new SoTA model based on LLaMA-2-7B!
 - 💥 [Nov, 2023] The [Xwin-Math-70B-V1.0](https://huggingface.co/Xwin-LM/Xwin-Math-70B-V1.0) model achieves **31.8 pass@1 on the MATH benchmark** and **87.0 pass@1 on the GSM8K benchmark**. This performance places it first amongst all open-source models!
 - 💥 [Nov, 2023] The [Xwin-Math-7B-V1.0](https://huggingface.co/Xwin-LM/Xwin-Math-7B-V1.0) and [Xwin-Math-13B-V1.0](https://huggingface.co/Xwin-LM/Xwin-Math-13B-V1.0) models achieve **66.6 and 76.2 pass@1 on the GSM8K benchmark**, ranking as top-1 among all LLaMA-2 based 7B and 13B open-source models respectively!
 ## ✨ Model Card
 |  Model  |  GSM8K  |  MATH  |  Checkpoint  |  License  |
 |:-:|:-:|:-:|:-:|:-:|
 |Xwin-Math-7B-V1.0 |  66.6  |  17.4  | 🤗 <a href="https://huggingface.co/Xwin-LM/Xwin-Math-7B-V1.0" target="_blank">HF Link</a> | <a href="https://ai.meta.com/resources/models-and-libraries/llama-downloads/" target="_blank">Llama 2 License|
 |Xwin-Math-7B-V1.1 |  84.4  |  44.7  | 🤗 <a href="https://huggingface.co/Xwin-LM/Xwin-Math-7B-V1.1" target="_blank">HF Link</a> | <a href="https://ai.meta.com/resources/models-and-libraries/llama-downloads/" target="_blank">Llama 2 License|
 |Xwin-Math-13B-V1.0|  76.2  |  21.7  | 🤗 <a href="https://huggingface.co/Xwin-LM/Xwin-Math-13B-V1.0" target="_blank">HF Link</a> |  <a href="https://ai.meta.com/resources/models-and-libraries/llama-downloads/" target="_blank">Llama 2 License|
 |Xwin-Math-70B-V1.0|  87.0  |  31.8  | 🤗 <a href="https://huggingface.co/Xwin-LM/Xwin-Math-70B-V1.0" target="_blank">HF Link</a> |  <a href="https://ai.meta.com/resources/models-and-libraries/llama-downloads/" target="_blank">Llama 2 License|
 |Xwin-Math-70B-V1.1|  90.6  |  51.9  | 🤗 <a href="https://huggingface.co/Xwin-LM/Xwin-Math-70B-V1.1" target="_blank">HF Link</a> |  <a href="https://ai.meta.com/resources/models-and-libraries/llama-downloads/" target="_blank">Llama 2 License|
 * Xwin-Math-7B-V1.1 uses 1.92M GSM8K and 960K MATH synthetic data
 * Xwin-Math-70B-V1.1 uses 960K GSM8K and 480K MATH synthetic data
 ## 🚀 Benchmarks
 ### Xwin-Math performance on [MATH](https://github.com/hendrycks/math) and [GSM8K](https://github.com/openai/grade-school-math).
 Xwin-Math-70B-V1.0 has achieved **31.8% on MATH** and **87.0% on GSM8K**. These scores are **5.3** and **3.1** points higher, respectively, than the previous state-of-the-art open-source MetaMath and LEMAv1 model.
 | **Model** |**MATH (Our test)** | **GSM8K (Our test)** |
 |:-:|:-:|:-:|
 |  GPT-4 (zero-shot)    |  52.4  |  94.8  |
 |  GPT-35-Turbo (8-shot)|  37.1  |  81.0  |
 | |
 |  WizardMath-70B       |  23.9  |  81.1  |
 |  MAmmoTH-70B          |  20.8  |  72.6  |
 |  MetaMath-70B         |  26.5  |  82.0  |
 |  LEMAv1-70B           |  25.9  |  83.9  |
 |**Xwin-Math-70B-V1.0** |**31.8**|**87.0**|
 |**Xwin-Math-70B-V1.1** |**51.9**|**90.6**|
 | | 
 |  WizardMath-13B       |  15.0  |  63.7  |
 |  MAmmoTH-13B          |  12.3  |  56.2  |
 |  MetaMath-13B         |  22.7  |  70.9  |
 |  LEMAv1-13B           |  13.6  |  65.0  |
 |**Xwin-Math-13B-V1.0** |  21.7  |  76.2  |
 | |
 |  WizardMath-7B        |  10.9  |  55.0  |
 |  MAmmoTH-7B           |  9.6   |  50.2  |
 |  MetaMath-7B 			|  20.1  |  66.6  |
 |  LEMAv1-7B            |  10.0  |  54.7  |
 |**Xwin-Math-7B-V1.0**  |  17.4  |  66.6  |
 |**Xwin-Math-7B-V1.1**  |  44.7  |  84.4  |
 We obtain these results using our flexible evaluation strategy. Due to differences in environment and hardware, the test results may be slightly different from the report, but we ensure that the evaluation is as accurate and fair as possible.
 ### Xwin-Math performance on other math benchmarks.
 Our 70B model shows strong mathematical reasoning capabilities among all open-sourced models. Also note that our model even approaches or surpasses the performance of GPT-35-Turbo on some benchmarks.
 | **Model** | SVAMP | ASDiv | NumGlue | Algebra | MAWPS | **Average** |
 |:-:|:-:|:-:|:-:|:-:|:-:|:-:|
 |  GPT-35-Turbo (8-shot)|  80.6  |  84.1  |  81.8  |  90.5  |  91.7  |  85.7  |
 | |
 |  WizardMath-70B       |  80.2	 |  75.8  |  71.4  |  64.0  |  74.9  |  73.3  |
 |  MAmmoTH-70B 	        |  71.2	 |  73.9  |  62.7  |  58.1  |  72.2  |  67.6  |
 |  MetaMath-70B         |  85.8  |  81.1  |  77.5  |  79.7  |  81.4  |  81.1  |
 |  LEMAv1-70B-MATH *    |  81.6  |  77.1  |  72.1  |  69.4  |  81.8  |  76.5  |
 |**Xwin-Math-70B-V1.0** |  84.0  |  84.1  |  81.3  |  78.4  |  90.8  |  83.7  |
 \* LEMAv1 has two models, and we report the better LEMAv1-70B-MATH model in these benchmarks.
 ## 🔨 Evaluation
 In order to evaluate a model's mathematical capabilities more flexibly and ensure a fair comparison of results, particularly for the MATH benchmark, we have developed a new evaluation tool. We have also assessed the pass@1 results of recent models on MATH and GSM8K benchmarks, which provides more accurate results.
 We hope this toolkit can benefit open-source community by providing more accurate insights and conclusions. For a deeper understanding of our evaluation tool and methods, please visit [here](https://github.com/Xwin-LM/Xwin-LM/tree/main/Xwin-Math/eval)
 * "Report" refers to the accuracy stated in the original papers.
 * "Repro" indicates the results is reproduced by generating responses and evaluating them using the respective open-source models and scripts.
 * "Strict" and "Flex" denote the results we achieved by employing our two strategies to extract answer and evaluate the same responses as "Repro".
 | Model | MATH <br> (Report) <br/> |MATH <br> (Repro) <br/> | MATH <br> (Strict) <br/>  |MATH <br> (Flex) <br/> | GSM8K <br> (Report) <br/> |GSM8K <br> (Repro) <br/>|  GSM8K <br> (Strict) <br/> |  GSM8K <br> (Report) <br/> |
 |:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|
 |  GPT-35-Turbo (8-shot)|  34.1  |  -     |  23.8  |  37.1  |  80.8  |  -     |  77.9  |  81.0  |
 | |
 |  WizardMath-70B       |  22.7  |  23.0  |  23.9  |  23.9  |  81.6  |  81.4  |  81.1  |  81.1  |
 |  MAmmoTH-70B          |  21.1  |  18.0  |  20.0  |  20.8  |  72.4  |  72.6  |  72.6  |  72.6  |
 |  MetaMath-70B         |  26.6  |  25.9  |  26.3  |  26.5  |  82.3  |  82.3  |  82.0  |  82.0  |
 |**Xwin-Math-70B-V1.0** |  -     |  -     |**31.8**|**31.8**|  -     |  -     |**87.0**|**87.0**|
 | | 
 |  WizardMath-13B       |  14.0  |  14.2  |  14.9  |  15.0  |  63.9  |  63.9  |  63.7  |  63.7  |
 |  MAmmoTH-13B          |  12.9  |  10.8  |  11.8  |  12.3  |  56.3	 |  56.2  |  56.1  |  56.2  |
 |  MetaMath-13B         |  22.4  |  22.5  |  22.6  |  22.7  |  72.3	 |  71.0  |  70.9  |  70.9  |
 |**Xwin-Math-13B-V1.0** |  -     |  -     |  21.6  |  21.7  |  -     |  -     |  76.2  |  76.2  |
 | |
 |  WizardMath-7B        |  10.7  |  10.3  |  10.9  |  10.9  |  54.9  |  55.2  |  55.0  |  55.0  |
 |  MAmmoTH-7B           |  10.4  |  8.6   |   9.1  |  9.6   |  50.5  |  50.2  |  50.2  |  50.2  |
 |  MetaMath-7B          |  19.8  |  19.6  |  19.9  |  20.1  |  66.5  |  66.6  |  66.6  |  66.6  |
 |**Xwin-Math-7B-V1.0**  |  -     |  -     |  17.3  |  17.4  |  -     |  -     |  66.6  |  66.6  |
 ### Installation
 Before you start, please install the requirements.
 ```bash
 pip install -r requirements.txt
 ```
 We tested our result using `python 3.8` and `cuda 11.8`. We recommend you use docker.
 ```bash
 docker run --gpus all -it --rm --ipc=host superbench/dev:cuda11.8
 ```
 ### Generate
 To generate the model's responses, you can use the `generate.py` script. Please be aware that generating responses is separate from verifying their correctness. After that, we will then check for their correctness.
 For the generation process, we use the Vicuna-v1.1 system prompt with chain-of-thought and format instruction. We also employ a greedy decoding strategy and set the maximum sequence length to 2048.
 ```
 "A chat between a curious user and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the user's questions. USER: {instruction} Give your solution in detail. In the end, write your final answer in the format of 'The answer is: <ANSWER>.'. ASSISTANT:"
 ```
 Here is an simple example to generate using [vLLM](https://docs.vllm.ai/en/latest/).
 ```bash
 cd eval
 python generate.py --dataset_path dataset/gsm8k.json --model_path path/to/your/model --tensor_parallel_size 4
 ```
 By default the results will be output to the `eval/response`, using the prompt `eval/prompt/xwin_math.json`. If you wish to change the output path or use a different prompt
 ```bash
 python generate.py --dataset_path dataset/gsm8k.json --model_path path/to/your/model --tensor_parallel_size 4 --output_path /your/path --prompt_path /your/path
 ```
 We provide some datasets (in `eval/dataset`):
 - `gsm8k.json`: GSM8K. 
 - `math.json`: MATH. 
 - `combination.json`: A combination of many benchmarks, can evaluate the OOD capability of the model. 
 If you wan't to use your own datasets, please format your dataset like this. 
 ```jsonc
 [
    {
        "question": "Janet\u2019s ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder at the farmers' market daily for $2 per fresh duck egg. How much in dollars does she make every day at the farmers' market?",
        "answer": "18",
        "type": "GSM8K",
        "subtype": "",
        "level": 0,
    },
    // ... more data items
 ]
 ```
 ### Evaluate
 To verify the accuracy of the answers after generation, you can use the `check.py script.
 Here is an simple example
 ```bash
 cd eval
 python eval.py /path/to/model/response
 ```
 The result will be saved in `eval/evaluation`
 If you do not want to save the results or want to change the save path
 ```bash
 python eval.py --data_path /path/to/model/response --save_path /path/to/save --save_result True
 ```
 Once you run the script, the terminal will display the output as a table. This table will show the number of instances for each benchmark and the corresponding accuracy. Here is a hypothetical example of what the output might look like:
 ||Type|Subtype|Level|Correct|Incorrect|Total|Accuracy|
 |---|---|---|---|---|---|---|---|
 |0|MAWPS|addsub|0|359|33|392|0.915816|
 |1|MAWPS|multiarith|0|586|14|600|0.976667|
 |...|
 ## Citation
 Please consider citing our work if you use the data or code in this repo.
 ```
@software{xwin-math,
  title = {Xwin-Math},
  author = {Xwin-Math Team},
  url = {https://github.com/Xwin-LM/Xwin-LM/Xwin-Math},
  version = {pre-release},
  year = {2023},
  month = {11},
 }
 ```
 ## Acknowledgements
 Thanks to [Llama 2](https://ai.meta.com/llama/), [FastChat](https://github.com/lm-sys/FastChat), and [vLLM](https://github.com/vllm-project/vllm).
--- a/added_tokens.json
+++ b/added_tokens.json
@@ -0,0 +1,3 @@
 {
  "[PAD]": 32000
 }
--- a/config.json
+++ b/config.json
@@ -0,0 +1,26 @@
 {
  "_name_or_path": "/data/mnt/chen_llm/Llama-2-7b-hf",
  "architectures": [
    "LlamaForCausalLM"
  ],
  "bos_token_id": 1,
  "eos_token_id": 2,
  "hidden_act": "silu",
  "hidden_size": 4096,
  "initializer_range": 0.02,
  "intermediate_size": 11008,
  "max_position_embeddings": 4096,
  "model_type": "llama",
  "num_attention_heads": 32,
  "num_hidden_layers": 32,
  "num_key_value_heads": 32,
  "pad_token_id": 0,
  "pretraining_tp": 1,
  "rms_norm_eps": 1e-05,
  "rope_scaling": null,
  "tie_word_embeddings": false,
  "torch_dtype": "float32",
  "transformers_version": "4.28.1",
  "use_cache": false,
  "vocab_size": 32001
 }
--- a/generation_config.json
+++ b/generation_config.json
@@ -0,0 +1,10 @@
 {
  "bos_token_id": 1,
  "do_sample": true,
  "eos_token_id": 2,
  "max_length": 4096,
  "pad_token_id": 0,
  "temperature": 0.6,
  "top_p": 0.9,
  "transformers_version": "4.28.1"
 }
--- a/pytorch_model-00001-of-00003.bin
+++ b/pytorch_model-00001-of-00003.bin
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:502146db64c298dc68049b3e42a462b9206bf577d8d5afeff0e601ddbd621346
 size 9878008210
--- a/pytorch_model-00002-of-00003.bin
+++ b/pytorch_model-00002-of-00003.bin
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:e1f9ec3ea2c915fef21810d153a42552c1193f025322b793879b045dde583470
 size 9894803382
--- a/pytorch_model-00003-of-00003.bin
+++ b/pytorch_model-00003-of-00003.bin
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:45d00394ce58724f7f1293d82fc591a854492f94298e62d31064a8bdccdc2613
 size 7181008761
--- a/pytorch_model.bin.index.json
+++ b/pytorch_model.bin.index.json
@@ -0,0 +1,330 @@
 {
  "metadata": {
    "total_size": 26953703424
  },
  "weight_map": {
    "lm_head.weight": "pytorch_model-00003-of-00003.bin",
    "model.embed_tokens.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.0.input_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.0.mlp.down_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.0.mlp.gate_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.0.mlp.up_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.0.post_attention_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.0.self_attn.k_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.0.self_attn.o_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.0.self_attn.q_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.0.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00003.bin",
    "model.layers.0.self_attn.v_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.1.input_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.1.mlp.down_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.1.mlp.gate_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.1.mlp.up_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.1.post_attention_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.1.self_attn.k_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.1.self_attn.o_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.1.self_attn.q_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.1.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00003.bin",
    "model.layers.1.self_attn.v_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.10.input_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.10.mlp.down_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.10.mlp.gate_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.10.mlp.up_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.10.post_attention_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.10.self_attn.k_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.10.self_attn.o_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.10.self_attn.q_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.10.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00003.bin",
    "model.layers.10.self_attn.v_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.11.input_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.11.mlp.down_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.11.mlp.gate_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.11.mlp.up_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.11.post_attention_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.11.self_attn.k_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.11.self_attn.o_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.11.self_attn.q_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.11.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00003.bin",
    "model.layers.11.self_attn.v_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.12.input_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.12.mlp.down_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.12.mlp.gate_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.12.mlp.up_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.12.post_attention_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.12.self_attn.k_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.12.self_attn.o_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.12.self_attn.q_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.12.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00003.bin",
    "model.layers.12.self_attn.v_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.13.input_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.13.mlp.down_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.13.mlp.gate_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.13.mlp.up_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.13.post_attention_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.13.self_attn.k_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.13.self_attn.o_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.13.self_attn.q_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.13.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00003.bin",
    "model.layers.13.self_attn.v_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.14.input_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.14.mlp.down_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.14.mlp.gate_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.14.mlp.up_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.14.post_attention_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.14.self_attn.k_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.14.self_attn.o_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.14.self_attn.q_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.14.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00003.bin",
    "model.layers.14.self_attn.v_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.15.input_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.15.mlp.down_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.15.mlp.gate_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.15.mlp.up_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.15.post_attention_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.15.self_attn.k_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.15.self_attn.o_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.15.self_attn.q_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.15.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00003.bin",
    "model.layers.15.self_attn.v_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.16.input_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.16.mlp.down_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.16.mlp.gate_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.16.mlp.up_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.16.post_attention_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.16.self_attn.k_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.16.self_attn.o_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.16.self_attn.q_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.16.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00003.bin",
    "model.layers.16.self_attn.v_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.17.input_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.17.mlp.down_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.17.mlp.gate_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.17.mlp.up_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.17.post_attention_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.17.self_attn.k_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.17.self_attn.o_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.17.self_attn.q_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.17.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00003.bin",
    "model.layers.17.self_attn.v_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.18.input_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.18.mlp.down_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.18.mlp.gate_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.18.mlp.up_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.18.post_attention_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.18.self_attn.k_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.18.self_attn.o_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.18.self_attn.q_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.18.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00003.bin",
    "model.layers.18.self_attn.v_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.19.input_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.19.mlp.down_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.19.mlp.gate_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.19.mlp.up_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.19.post_attention_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.19.self_attn.k_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.19.self_attn.o_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.19.self_attn.q_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.19.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00003.bin",
    "model.layers.19.self_attn.v_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.2.input_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.2.mlp.down_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.2.mlp.gate_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.2.mlp.up_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.2.post_attention_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.2.self_attn.k_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.2.self_attn.o_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.2.self_attn.q_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.2.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00003.bin",
    "model.layers.2.self_attn.v_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.20.input_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.20.mlp.down_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.20.mlp.gate_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.20.mlp.up_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.20.post_attention_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.20.self_attn.k_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.20.self_attn.o_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.20.self_attn.q_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.20.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00003.bin",
    "model.layers.20.self_attn.v_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.21.input_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.21.mlp.down_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.21.mlp.gate_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.21.mlp.up_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.21.post_attention_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.21.self_attn.k_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.21.self_attn.o_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.21.self_attn.q_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.21.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00003.bin",
    "model.layers.21.self_attn.v_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.22.input_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.22.mlp.down_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.22.mlp.gate_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.22.mlp.up_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.22.post_attention_layernorm.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.22.self_attn.k_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.22.self_attn.o_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.22.self_attn.q_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.22.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00003.bin",
    "model.layers.22.self_attn.v_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.23.input_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.23.mlp.down_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.23.mlp.gate_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.23.mlp.up_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.23.post_attention_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.23.self_attn.k_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.23.self_attn.o_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.23.self_attn.q_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.23.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00003.bin",
    "model.layers.23.self_attn.v_proj.weight": "pytorch_model-00002-of-00003.bin",
    "model.layers.24.input_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.24.mlp.down_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.24.mlp.gate_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.24.mlp.up_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.24.post_attention_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.24.self_attn.k_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.24.self_attn.o_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.24.self_attn.q_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.24.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00003.bin",
    "model.layers.24.self_attn.v_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.25.input_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.25.mlp.down_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.25.mlp.gate_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.25.mlp.up_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.25.post_attention_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.25.self_attn.k_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.25.self_attn.o_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.25.self_attn.q_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.25.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00003.bin",
    "model.layers.25.self_attn.v_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.26.input_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.26.mlp.down_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.26.mlp.gate_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.26.mlp.up_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.26.post_attention_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.26.self_attn.k_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.26.self_attn.o_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.26.self_attn.q_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.26.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00003.bin",
    "model.layers.26.self_attn.v_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.27.input_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.27.mlp.down_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.27.mlp.gate_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.27.mlp.up_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.27.post_attention_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.27.self_attn.k_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.27.self_attn.o_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.27.self_attn.q_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.27.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00003.bin",
    "model.layers.27.self_attn.v_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.28.input_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.28.mlp.down_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.28.mlp.gate_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.28.mlp.up_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.28.post_attention_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.28.self_attn.k_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.28.self_attn.o_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.28.self_attn.q_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.28.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00003.bin",
    "model.layers.28.self_attn.v_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.29.input_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.29.mlp.down_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.29.mlp.gate_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.29.mlp.up_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.29.post_attention_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.29.self_attn.k_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.29.self_attn.o_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.29.self_attn.q_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.29.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00003.bin",
    "model.layers.29.self_attn.v_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.3.input_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.3.mlp.down_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.3.mlp.gate_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.3.mlp.up_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.3.post_attention_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.3.self_attn.k_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.3.self_attn.o_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.3.self_attn.q_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.3.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00003.bin",
    "model.layers.3.self_attn.v_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.30.input_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.30.mlp.down_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.30.mlp.gate_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.30.mlp.up_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.30.post_attention_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.30.self_attn.k_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.30.self_attn.o_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.30.self_attn.q_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.30.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00003.bin",
    "model.layers.30.self_attn.v_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.31.input_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.31.mlp.down_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.31.mlp.gate_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.31.mlp.up_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.31.post_attention_layernorm.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.31.self_attn.k_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.31.self_attn.o_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.31.self_attn.q_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.31.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00003.bin",
    "model.layers.31.self_attn.v_proj.weight": "pytorch_model-00003-of-00003.bin",
    "model.layers.4.input_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.4.mlp.down_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.4.mlp.gate_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.4.mlp.up_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.4.post_attention_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.4.self_attn.k_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.4.self_attn.o_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.4.self_attn.q_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.4.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00003.bin",
    "model.layers.4.self_attn.v_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.5.input_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.5.mlp.down_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.5.mlp.gate_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.5.mlp.up_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.5.post_attention_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.5.self_attn.k_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.5.self_attn.o_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.5.self_attn.q_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.5.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00003.bin",
    "model.layers.5.self_attn.v_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.6.input_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.6.mlp.down_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.6.mlp.gate_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.6.mlp.up_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.6.post_attention_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.6.self_attn.k_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.6.self_attn.o_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.6.self_attn.q_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.6.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00003.bin",
    "model.layers.6.self_attn.v_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.7.input_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.7.mlp.down_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.7.mlp.gate_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.7.mlp.up_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.7.post_attention_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.7.self_attn.k_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.7.self_attn.o_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.7.self_attn.q_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.7.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00003.bin",
    "model.layers.7.self_attn.v_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.8.input_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.8.mlp.down_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.8.mlp.gate_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.8.mlp.up_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.8.post_attention_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.8.self_attn.k_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.8.self_attn.o_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.8.self_attn.q_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.8.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00003.bin",
    "model.layers.8.self_attn.v_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.9.input_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.9.mlp.down_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.9.mlp.gate_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.9.mlp.up_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.9.post_attention_layernorm.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.9.self_attn.k_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.9.self_attn.o_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.9.self_attn.q_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.layers.9.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00003.bin",
    "model.layers.9.self_attn.v_proj.weight": "pytorch_model-00001-of-00003.bin",
    "model.norm.weight": "pytorch_model-00003-of-00003.bin"
  }
 }
--- a/special_tokens_map.json
+++ b/special_tokens_map.json
@@ -0,0 +1,24 @@
 {
  "bos_token": {
    "content": "<s>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  },
  "eos_token": {
    "content": "</s>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  },
  "pad_token": "[PAD]",
  "unk_token": {
    "content": "<unk>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  }
 }
--- a/tokenizer.model
+++ b/tokenizer.model
--- a/tokenizer_config.json
+++ b/tokenizer_config.json
@@ -0,0 +1,35 @@
 {
  "add_bos_token": true,
  "add_eos_token": false,
  "bos_token": {
    "__type": "AddedToken",
    "content": "<s>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  },
  "clean_up_tokenization_spaces": false,
  "eos_token": {
    "__type": "AddedToken",
    "content": "</s>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  },
  "legacy": false,
  "model_max_length": 2048,
  "pad_token": null,
  "padding_side": "right",
  "sp_model_kwargs": {},
  "tokenizer_class": "LlamaTokenizer",
  "unk_token": {
    "__type": "AddedToken",
    "content": "<unk>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  }
 }