初始化项目，由ModelHub XC社区提供模型

Model: allenai/OLMo-2-1124-13B-Instruct Source: Original Platform
2026-06-15 11:44:13 +08:00
commit d1a678f975
17 changed files with 601440 additions and 0 deletions
--- a/.gitattributes
+++ b/.gitattributes
@@ -0,0 +1,41 @@
 *.7z filter=lfs diff=lfs merge=lfs -text
 *.arrow filter=lfs diff=lfs merge=lfs -text
 *.bin filter=lfs diff=lfs merge=lfs -text
 *.bz2 filter=lfs diff=lfs merge=lfs -text
 *.ckpt filter=lfs diff=lfs merge=lfs -text
 *.ftz filter=lfs diff=lfs merge=lfs -text
 *.gz filter=lfs diff=lfs merge=lfs -text
 *.h5 filter=lfs diff=lfs merge=lfs -text
 *.joblib filter=lfs diff=lfs merge=lfs -text
 *.lfs.* filter=lfs diff=lfs merge=lfs -text
 *.mlmodel filter=lfs diff=lfs merge=lfs -text
 *.model filter=lfs diff=lfs merge=lfs -text
 *.msgpack filter=lfs diff=lfs merge=lfs -text
 *.npy filter=lfs diff=lfs merge=lfs -text
 *.npz filter=lfs diff=lfs merge=lfs -text
 *.onnx filter=lfs diff=lfs merge=lfs -text
 *.ot filter=lfs diff=lfs merge=lfs -text
 *.parquet filter=lfs diff=lfs merge=lfs -text
 *.pb filter=lfs diff=lfs merge=lfs -text
 *.pickle filter=lfs diff=lfs merge=lfs -text
 *.pkl filter=lfs diff=lfs merge=lfs -text
 *.pt filter=lfs diff=lfs merge=lfs -text
 *.pth filter=lfs diff=lfs merge=lfs -text
 *.rar filter=lfs diff=lfs merge=lfs -text
 *.safetensors filter=lfs diff=lfs merge=lfs -text
 saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.tar.* filter=lfs diff=lfs merge=lfs -text
 *.tar filter=lfs diff=lfs merge=lfs -text
 *.tflite filter=lfs diff=lfs merge=lfs -text
 *.tgz filter=lfs diff=lfs merge=lfs -text
 *.wasm filter=lfs diff=lfs merge=lfs -text
 *.xz filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
 model-00001-of-00006.safetensors filter=lfs diff=lfs merge=lfs -text
 model-00002-of-00006.safetensors filter=lfs diff=lfs merge=lfs -text
 model-00003-of-00006.safetensors filter=lfs diff=lfs merge=lfs -text
 model-00004-of-00006.safetensors filter=lfs diff=lfs merge=lfs -text
 model-00005-of-00006.safetensors filter=lfs diff=lfs merge=lfs -text
 model-00006-of-00006.safetensors filter=lfs diff=lfs merge=lfs -text
--- a/README.md
+++ b/README.md
@@ -0,0 +1,151 @@
 ---
 license: apache-2.0
 language:
 - en
 pipeline_tag: text-generation
 base_model:
 - allenai/OLMo-2-1124-13B-Instruct-RLVR2
 library_name: transformers
 datasets:
 - allenai/RLVR-MATH
 ---
 <img alt="OLMo Logo" src="https://huggingface.co/datasets/allenai/blog-images/resolve/main/olmo2/olmo.png" width="242px">
 # OLMo-2-1124-13B-Instruct
 ## NOTE: 1/3/2025 UPDATE:
 Upon the initial release of OLMo-2 models, we realized the post-trained models did not share the pre-tokenization logic that the base models use. As a result, we have trained new post-trained models. The new models are available under the same names as the original models, but we have made the old models available with a postfix "-preview". See [OLMo 2 Preview Post-trained Models](https://huggingface.co/collections/allenai/olmo-2-preview-post-trained-models-6762f662c660962e52de7c96) for the colleciton of the legacy models.
 ## Release Documentation
 OLMo 2 13B Instruct November 2024 is post-trained variant of the [OLMo-2 13B November 2024](https://huggingface.co/allenai/OLMo2-13B-1124) model, which has undergone supervised finetuning on an OLMo-specific variant of the [Tülu 3 dataset](https://huggingface.co/datasets/allenai/tulu-3-sft-olmo-2-mixture and further DPO training on [this dataset](https://huggingface.co/datasets/allenai/olmo-2-1124-13b-preference-mix), and finally RLVR training using [this data](https://huggingface.co/datasets/allenai/RLVR-GSM).
 Tülu 3 is designed for state-of-the-art performance on a diversity of tasks in addition to chat, such as MATH, GSM8K, and IFEval.
 Check out the [OLMo 2 paper](https://arxiv.org/abs/2501.00656) or [Tülu 3 paper](https://arxiv.org/abs/2411.15124) for more details!
 OLMo is a series of **O**pen **L**anguage **Mo**dels designed to enable the science of language models. 
 These models are trained on the Dolma dataset. We are releasing all code, checkpoints, logs (coming soon), and associated training details. 
 The core models released in this batch include the following:
 | **Stage**           | **OLMo 2 7B**                                                                                          | **OLMo 2 13B**                                                                                         |
 |----------------------|----------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------|
 | **Base Model**       | [allenai/OLMo2-7B-1124](https://huggingface.co/allenai/OLMo2-7B-1124)                                | [allenai/OLMo-2-13B-1124](https://huggingface.co/allenai/OLMo-2-13B-1124)                             |
 | **SFT**              | [allenai/OLMo-2-1124-7B-SFT](https://huggingface.co/allenai/OLMo-2-1124-7B-SFT)                | [allenai/OLMo-2-1124-13B-SFT](https://huggingface.co/allenai/OLMo-2-1124-13B-SFT)              |
 | **DPO**              | [allenai/OLMo-2-1124-7B-DPO](https://huggingface.co/allenai/OLMo-2-1124-7B-DPO)                | [allenai/OLMo-2-1124-13B-DPO](https://huggingface.co/allenai/OLMo-2-1124-13B-DPO)              |
 | **Final Models (RLVR)** | [allenai/OLMo-2-1124-7B-Instruct](https://huggingface.co/allenai/OLMo-2-1124-7B-Instruct)                        | [allenai/OLMo-2-1124-13B-Instruct](https://huggingface.co/allenai/OLMo-2-1124-13B-Instruct)                      |
 | **Reward Model (RM)**| [allenai/OLMo-2-1124-7B-RM](https://huggingface.co/allenai/OLMo-2-1124-7B-RM)                                                     | [allenai/OLMo-2-1124-13B-RM](https://huggingface.co/allenai/OLMo-2-1124-13B-RM)                                                     |
 ## Model description
 - **Model type:** A model trained on a mix of publicly available, synthetic and human-created datasets.
 - **Language(s) (NLP):** Primarily English
 - **License:** Apache 2.0
 - **Finetuned from model:** allenai/OLMo-2-13B-1124-RLVR2
 ### Model Sources
 - **Project Page:** https://allenai.org/olmo
 - **Repositories:** 
    - Core repo (training, inference, fine-tuning etc.): https://github.com/allenai/OLMo
    - Evaluation code: https://github.com/allenai/olmes
    - Further fine-tuning code: https://github.com/allenai/open-instruct
 - **Paper:** https://arxiv.org/abs/2501.00656
 - **Demo:** https://playground.allenai.org/
 ## Installation
 OLMo 2 will be supported in the next version of Transformers, and you need to install it from the main branch using:
 ```bash
 pip install --upgrade git+https://github.com/huggingface/transformers.git
 ```
 ## Using the model
 ### Loading with HuggingFace
 To load the model with HuggingFace, use the following snippet:
 ```
 from transformers import AutoModelForCausalLM
 olmo_model = AutoModelForCausalLM.from_pretrained("allenai/OLMo-2-1124-13B-Instruct")
 ```
 ### Chat template
 The chat template for our models is formatted as:
 ```
 <|endoftext|><|user|>\nHow are you doing?\n<|assistant|>\nI'm just a computer program, so I don't have feelings, but I'm functioning as expected. How can I assist you today?<|endoftext|>
 ```
 Or with new lines expanded:
 ```
 <|endoftext|><|user|>
 How are you doing?
 <|assistant|>
 I'm just a computer program, so I don't have feelings, but I'm functioning as expected. How can I assist you today?<|endoftext|>
 ```
 It is embedded within the tokenizer as well, for `tokenizer.apply_chat_template`.
 ### System prompt
 In Ai2 demos, we use this system prompt by default:
 ```
 You are OLMo 2, a helpful and harmless AI Assistant built by the Allen Institute for AI.
 ```
 The model has not been trained with a specific system prompt in mind.
 ### Bias, Risks, and Limitations
 The OLMo-2 models have limited safety training, but are not deployed automatically with in-the-loop filtering of responses like ChatGPT, so the model can produce problematic outputs (especially when prompted to do so). 
 See the Falcon 180B model card for an example of this.
 ## Performance
 | Model | Average | AlpacaEval | BBH | DROP | GSM8k | IFEval | MATH | MMLU | Safety | PopQA | TruthQA |
 |-------|---------|------------|-----|------|--------|---------|------|-------|---------|-------|---------|
 | **Open weights models** |
 | Gemma-2-9B-it | 51.9 | 43.7 | 2.5 | 58.8 | 79.7 | 69.9 | 29.8 | 69.1 | 75.5 | 28.3 | 61.4 |
 | Ministral-8B-Instruct | 52.1 | 31.4 | 56.2 | 56.2 | 80.0 | 56.4 | 40.0 | 68.5 | 56.2 | 20.2 | 55.5 |
 | Mistral-Nemo-Instruct-2407 | 50.9 | 45.8 | 54.6 | 23.6 | 81.4 | 64.5 | 31.9 | 70.0 | 52.7 | 26.9 | 57.7 |
 | Qwen-2.5-7B-Instruct | 57.1 | 29.7 | 25.3 | 54.4 | 83.8 | 74.7 | 69.9 | 76.6 | 75.0 | 18.1 | 63.1 |
 | Llama-3.1-8B-Instruct | 58.9 | 25.8 | 69.7 | 61.7 | 83.4 | 80.6 | 42.5 | 71.3 | 70.2 | 28.4 | 55.1 |
 | Tülu 3 8B | 60.4 | 34.0 | 66.0 | 62.6 | 87.6 | 82.4 | 43.7 | 68.2 | 75.4 | 29.1 | 55.0 |
 | Qwen-2.5-14B-Instruct | 60.8 | 34.6 | 34.0 | 50.5 | 83.9 | 82.4 | 70.6 | 81.1 | 79.3 | 21.1 | 70.8 |
 | **Fully open models** |
 | OLMo-7B-Instruct | 28.2 | 5.2 | 35.3 | 30.7 | 14.3 | 32.2 | 2.1 | 46.3 | 54.0 | 17.1 | 44.5 |
 | OLMo-7B-0424-Instruct | 33.1 | 8.5 | 34.4 | 47.9 | 23.2 | 39.2 | 5.2 | 48.9 | 49.3 | 18.9 | 55.2 |
 | OLMoE-1B-7B-0924-Instruct | 35.5 | 8.5 | 37.2 | 34.3 | 47.2 | 46.2 | 8.4 | 51.6 | 51.6 | 20.6 | 49.1 |
 | MAP-Neo-7B-Instruct | 42.9 | 17.6 | 26.4 | 48.2 | 69.4 | 35.9 | 31.5 | 56.5 | 73.7 | 18.4 | 51.6 |
 | *OLMo-2-7B-SFT* | 50.2 | 10.2 | 49.7 | 59.6 | 74.6 | 66.9 | 25.3 | 61.1 | 82.1 | 23.6 | 48.6 |
 | *OLMo-2-7B-DPO* | 54.2 | 27.9 | 46.7 | 60.2 | 82.6 | 73.0 | 30.3 | 60.8 | 81.0 | 23.5 | 56.0 |
 | *OLMo-2-13B-SFT* | 55.3 | 11.5 | 59.6 | 71.3 | 76.3 | 68.6 | 29.5 | 68.0 | 82.3 | 29.4 | 57.1 |
 | *OLMo-2-13B-DPO* | 60.6 | 38.3 | 57.9 | 71.5 | 82.3 | 80.2 | 35.2 | 67.9 | 79.7 | 29.0 | 63.9 |
 | **OLMo-2-7B-1124–Instruct** | 54.8 | 29.1 | 46.6 | 60.5 | 85.1 | 72.3 | 32.5 | 61.3 | 80.6 | 23.2 | 56.5 |
 | **OLMo-2-13B-1124-Instruct** | 62.0 | 39.5 | 58.8 | 71.5 | 87.4 | 82.6 | 39.2 | 68.5 | 79.1 | 28.8 | 64.3 |
 ## License and use
 OLMo 2 is licensed under the Apache 2.0 license.
 OLMo 2 is intended for research and educational use.
 For more information, please see our [Responsible Use Guidelines](https://allenai.org/responsible-use).
 This model has been fine-tuned using a dataset mix with outputs generated from third party models and are subject to additional terms: [Gemma Terms of Use](https://ai.google.dev/gemma/terms).
 ## Citation
 ```bibtex
@article{olmo20242olmo2furious,
      title={2 OLMo 2 Furious}, 
      author={Team OLMo and Pete Walsh and Luca Soldaini and Dirk Groeneveld and Kyle Lo and Shane Arora and Akshita Bhagia and Yuling Gu and Shengyi Huang and Matt Jordan and Nathan Lambert and Dustin Schwenk and Oyvind Tafjord and Taira Anderson and David Atkinson and Faeze Brahman and Christopher Clark and Pradeep Dasigi and Nouha Dziri and Michal Guerquin and Hamish Ivison and Pang Wei Koh and Jiacheng Liu and Saumya Malik and William Merrill and Lester James V. Miranda and Jacob Morrison and Tyler Murray and Crystal Nam and Valentina Pyatkin and Aman Rangapur and Michael Schmitz and Sam Skjonsberg and David Wadden and Christopher Wilhelm and Michael Wilson and Luke Zettlemoyer and Ali Farhadi and Noah A. Smith and Hannaneh Hajishirzi},
      year={2024},
      eprint={2501.00656},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2501.00656}, 
 }
 ```
--- a/config.json
+++ b/config.json
@@ -0,0 +1,27 @@
 {
  "_name_or_path": "/weka/oe-adapt-default/costah/models/olmo2/1210_13b_gsm_beta_0.03_lr_3e-7_7470_checkpoints/step_180/",
  "architectures": [
    "Olmo2ForCausalLM"
  ],
  "attention_bias": false,
  "attention_dropout": 0.0,
  "eos_token_id": 100257,
  "hidden_act": "silu",
  "hidden_size": 5120,
  "initializer_range": 0.02,
  "intermediate_size": 13824,
  "max_position_embeddings": 4096,
  "model_type": "olmo2",
  "num_attention_heads": 40,
  "num_hidden_layers": 40,
  "num_key_value_heads": 40,
  "pad_token_id": 100277,
  "rms_norm_eps": 1e-06,
  "rope_scaling": null,
  "rope_theta": 500000,
  "tie_word_embeddings": false,
  "torch_dtype": "bfloat16",
  "transformers_version": "4.47.0.dev0",
  "use_cache": false,
  "vocab_size": 100352
 }
--- a/configuration.json
+++ b/configuration.json
@@ -0,0 +1 @@
 {"framework": "pytorch", "task": "text-generation", "allow_remote": true}
--- a/generation_config.json
+++ b/generation_config.json
@@ -0,0 +1,6 @@
 {
  "_from_model_config": true,
  "eos_token_id": 100257,
  "pad_token_id": 100277,
  "transformers_version": "4.47.0.dev0"
 }
--- a/merges.txt
+++ b/merges.txt
--- a/model-00001-of-00006.safetensors
+++ b/model-00001-of-00006.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:884503480a474aaabf839123612eef8aa28f2d50d08a49bb7a62d0d756dd0377
 size 4991475672
--- a/model-00002-of-00006.safetensors
+++ b/model-00002-of-00006.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:23c10fee35770c0f44166e20991fa4a3da88044652ce0cdedb03426426328b8a
 size 4970587952
--- a/model-00003-of-00006.safetensors
+++ b/model-00003-of-00006.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:5e83f0d2467257172182fb701ee113e6a70f676160d60a5c2ee4e6d612d849a0
 size 4881438320
--- a/model-00004-of-00006.safetensors
+++ b/model-00004-of-00006.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:212f64c1a9b9ade4870af03182a6b05a1662398537a256e7bd38ad2302780c3f
 size 4933887952
--- a/model-00005-of-00006.safetensors
+++ b/model-00005-of-00006.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:25e5a2a15836560a08a7c0292372394a8f4106e34cab2929cd59b26655fb9945
 size 4933887952
--- a/model-00006-of-00006.safetensors
+++ b/model-00006-of-00006.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:a0eb7fa0197f3f85a16832928cb9428912198cd65af3dd3a5e79ecc3c206ee90
 size 2721170744
--- a/model.safetensors.index.json
+++ b/model.safetensors.index.json
@@ -0,0 +1,450 @@
 {
  "metadata": {
    "total_size": 27432396800
  },
  "weight_map": {
    "lm_head.weight": "model-00006-of-00006.safetensors",
    "model.embed_tokens.weight": "model-00001-of-00006.safetensors",
    "model.layers.0.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.0.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.0.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.0.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
    "model.layers.0.post_feedforward_layernorm.weight": "model-00001-of-00006.safetensors",
    "model.layers.0.self_attn.k_norm.weight": "model-00001-of-00006.safetensors",
    "model.layers.0.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.0.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.0.self_attn.q_norm.weight": "model-00001-of-00006.safetensors",
    "model.layers.0.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.0.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.1.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.1.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.1.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.1.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
    "model.layers.1.post_feedforward_layernorm.weight": "model-00001-of-00006.safetensors",
    "model.layers.1.self_attn.k_norm.weight": "model-00001-of-00006.safetensors",
    "model.layers.1.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.1.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.1.self_attn.q_norm.weight": "model-00001-of-00006.safetensors",
    "model.layers.1.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.1.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.10.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.10.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.10.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.10.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.10.post_feedforward_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.10.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.10.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.10.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.10.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.10.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.10.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.11.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.11.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.11.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.11.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.11.post_feedforward_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.11.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.11.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.11.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.11.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.11.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.11.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.12.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.12.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.12.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.12.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.12.post_feedforward_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.12.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.12.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.12.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.12.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.12.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.12.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.13.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.13.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.13.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.13.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.13.post_feedforward_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.13.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.13.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.13.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.13.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.13.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.13.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.14.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.14.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.14.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.14.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.14.post_feedforward_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.14.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.14.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.14.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.14.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.14.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.14.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.15.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.15.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.15.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.15.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.15.post_feedforward_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.15.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.15.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.15.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.15.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.15.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.15.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.16.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.16.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.16.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.16.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.16.post_feedforward_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.16.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.16.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.16.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.16.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.16.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.16.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.17.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.17.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.17.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.17.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.17.post_feedforward_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.17.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.17.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.17.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.17.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.17.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.17.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.18.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.18.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.18.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.18.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.18.post_feedforward_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.18.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.18.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.18.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.18.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.18.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.18.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.19.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.19.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.19.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.19.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.19.post_feedforward_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.19.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.19.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.19.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.19.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.19.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.19.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.2.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.2.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.2.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.2.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
    "model.layers.2.post_feedforward_layernorm.weight": "model-00001-of-00006.safetensors",
    "model.layers.2.self_attn.k_norm.weight": "model-00001-of-00006.safetensors",
    "model.layers.2.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.2.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.2.self_attn.q_norm.weight": "model-00001-of-00006.safetensors",
    "model.layers.2.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.2.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.20.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.20.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.20.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.20.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.20.post_feedforward_layernorm.weight": "model-00003-of-00006.safetensors",
    "model.layers.20.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.20.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.20.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.20.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.20.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.20.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.21.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.21.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.21.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.21.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.21.post_feedforward_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.21.self_attn.k_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.21.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.21.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.21.self_attn.q_norm.weight": "model-00003-of-00006.safetensors",
    "model.layers.21.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.21.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
    "model.layers.22.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.22.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.22.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.22.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.22.post_feedforward_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.22.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.22.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.22.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.22.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.22.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.22.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.23.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.23.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.23.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.23.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.23.post_feedforward_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.23.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.23.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.23.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.23.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.23.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.23.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.24.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.24.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.24.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.24.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.24.post_feedforward_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.24.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.24.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.24.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.24.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.24.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.24.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.25.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.25.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.25.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.25.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.25.post_feedforward_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.25.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.25.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.25.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.25.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.25.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.25.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.26.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.26.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.26.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.26.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.26.post_feedforward_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.26.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.26.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.26.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.26.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.26.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.26.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.27.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.27.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.27.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.27.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.27.post_feedforward_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.27.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.27.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.27.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.27.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.27.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.27.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.28.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.28.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.28.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.28.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.28.post_feedforward_layernorm.weight": "model-00004-of-00006.safetensors",
    "model.layers.28.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.28.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.28.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.28.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.28.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.28.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.29.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.29.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.29.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.29.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.29.post_feedforward_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.29.self_attn.k_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.29.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.29.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.29.self_attn.q_norm.weight": "model-00004-of-00006.safetensors",
    "model.layers.29.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.29.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
    "model.layers.3.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.3.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.3.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.3.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
    "model.layers.3.post_feedforward_layernorm.weight": "model-00001-of-00006.safetensors",
    "model.layers.3.self_attn.k_norm.weight": "model-00001-of-00006.safetensors",
    "model.layers.3.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.3.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.3.self_attn.q_norm.weight": "model-00001-of-00006.safetensors",
    "model.layers.3.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.3.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.30.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.30.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.30.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.30.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.30.post_feedforward_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.30.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.30.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.30.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.30.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.30.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.30.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.31.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.31.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.31.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.31.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.31.post_feedforward_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.31.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.31.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.31.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.31.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.31.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.31.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.32.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.32.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.32.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.32.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.32.post_feedforward_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.32.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.32.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.32.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.32.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.32.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.32.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.33.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.33.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.33.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.33.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.33.post_feedforward_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.33.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.33.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.33.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.33.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.33.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.33.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.34.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.34.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.34.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.34.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.34.post_feedforward_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.34.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.34.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.34.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.34.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.34.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.34.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.35.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.35.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.35.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.35.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.35.post_feedforward_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.35.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.35.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.35.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.35.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.35.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.35.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.36.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.36.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.36.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.36.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.36.post_feedforward_layernorm.weight": "model-00005-of-00006.safetensors",
    "model.layers.36.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.36.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.36.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.36.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.36.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.36.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.37.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.37.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.37.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.37.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
    "model.layers.37.post_feedforward_layernorm.weight": "model-00006-of-00006.safetensors",
    "model.layers.37.self_attn.k_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.37.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.37.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.37.self_attn.q_norm.weight": "model-00005-of-00006.safetensors",
    "model.layers.37.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.37.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
    "model.layers.38.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.38.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.38.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.38.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
    "model.layers.38.post_feedforward_layernorm.weight": "model-00006-of-00006.safetensors",
    "model.layers.38.self_attn.k_norm.weight": "model-00006-of-00006.safetensors",
    "model.layers.38.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.38.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.38.self_attn.q_norm.weight": "model-00006-of-00006.safetensors",
    "model.layers.38.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.38.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.39.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.39.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.39.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.39.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
    "model.layers.39.post_feedforward_layernorm.weight": "model-00006-of-00006.safetensors",
    "model.layers.39.self_attn.k_norm.weight": "model-00006-of-00006.safetensors",
    "model.layers.39.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.39.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.39.self_attn.q_norm.weight": "model-00006-of-00006.safetensors",
    "model.layers.39.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.39.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
    "model.layers.4.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.4.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.4.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.4.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
    "model.layers.4.post_feedforward_layernorm.weight": "model-00001-of-00006.safetensors",
    "model.layers.4.self_attn.k_norm.weight": "model-00001-of-00006.safetensors",
    "model.layers.4.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.4.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.4.self_attn.q_norm.weight": "model-00001-of-00006.safetensors",
    "model.layers.4.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.4.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.5.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.5.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.5.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.5.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
    "model.layers.5.post_feedforward_layernorm.weight": "model-00001-of-00006.safetensors",
    "model.layers.5.self_attn.k_norm.weight": "model-00001-of-00006.safetensors",
    "model.layers.5.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.5.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.5.self_attn.q_norm.weight": "model-00001-of-00006.safetensors",
    "model.layers.5.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.5.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.6.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.6.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.6.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.6.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.6.post_feedforward_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.6.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.6.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.6.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.6.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.6.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.6.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
    "model.layers.7.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.7.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.7.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.7.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.7.post_feedforward_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.7.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.7.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.7.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.7.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.7.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.7.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.8.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.8.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.8.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.8.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.8.post_feedforward_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.8.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.8.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.8.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.8.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.8.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.8.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.9.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.9.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.9.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.9.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.9.post_feedforward_layernorm.weight": "model-00002-of-00006.safetensors",
    "model.layers.9.self_attn.k_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.9.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.9.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.9.self_attn.q_norm.weight": "model-00002-of-00006.safetensors",
    "model.layers.9.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
    "model.layers.9.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
    "model.norm.weight": "model-00006-of-00006.safetensors"
  }
 }
--- a/special_tokens_map.json
+++ b/special_tokens_map.json
@@ -0,0 +1,30 @@
 {
  "bos_token": {
    "content": "<|endoftext|>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  },
  "eos_token": {
    "content": "<|endoftext|>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  },
  "pad_token": {
    "content": "<|pad|>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  },
  "unk_token": {
    "content": "<|endoftext|>",
    "lstrip": false,
    "normalized": false,
    "rstrip": false,
    "single_word": false
  }
 }
--- a/tokenizer.json
+++ b/tokenizer.json
--- a/tokenizer_config.json
+++ b/tokenizer_config.json
@@ -0,0 +1,190 @@
 {
  "add_prefix_space": false,
  "added_tokens_decoder": {
    "100256": {
      "content": "<|extra_id_0|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100257": {
      "content": "<|endoftext|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "100258": {
      "content": "<|fim_prefix|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "100259": {
      "content": "<|fim_middle|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "100260": {
      "content": "<|fim_suffix|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "100261": {
      "content": "|||PHONE_NUMBER|||",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100262": {
      "content": "|||EMAIL_ADDRESS|||",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100263": {
      "content": "|||IP_ADDRESS|||",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100264": {
      "content": "<|im_start|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "100265": {
      "content": "<|im_end|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "100266": {
      "content": "<|extra_id_1|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100267": {
      "content": "<|extra_id_2|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100268": {
      "content": "<|extra_id_3|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100269": {
      "content": "<|extra_id_4|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100270": {
      "content": "<|extra_id_5|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100271": {
      "content": "<|extra_id_6|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100272": {
      "content": "<|extra_id_7|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100273": {
      "content": "<|extra_id_8|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100274": {
      "content": "<|extra_id_9|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100275": {
      "content": "<|extra_id_10|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": false
    },
    "100276": {
      "content": "<|endofprompt|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "100277": {
      "content": "<|pad|>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    }
  },
  "bos_token": "<|endoftext|>",
  "chat_template": "{{ bos_token }}{% for message in messages %}{% if message['role'] == 'system' %}{{ '<|system|>\n' + message['content'] + '\n' }}{% elif message['role'] == 'user' %}{{ '<|user|>\n' + message['content'] + '\n' }}{% elif message['role'] == 'assistant' %}{% if not loop.last %}{{ '<|assistant|>\n'  + message['content'] + eos_token + '\n' }}{% else %}{{ '<|assistant|>\n'  + message['content'] + eos_token }}{% endif %}{% endif %}{% if loop.last and add_generation_prompt %}{{ '<|assistant|>\n' }}{% endif %}{% endfor %}",
  "clean_up_tokenization_spaces": false,
  "eos_token": "<|endoftext|>",
  "extra_special_tokens": {},
  "model_max_length": 1000000000000000019884624838656,
  "pad_token": "<|pad|>",
  "tokenizer_class": "GPT2Tokenizer",
  "unk_token": "<|endoftext|>"
 }
--- a/vocab.json
+++ b/vocab.json
		`@@ -0,0 +1 @@`
							`{"framework": "pytorch", "task": "text-generation", "allow_remote": true}`