初始化项目,由ModelHub XC社区提供模型

Model: starmpcc/Asclepius-Llama2-13B
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-06-16 14:05:20 +08:00
commit 6be9854300
17 changed files with 22965 additions and 0 deletions

35
.gitattributes vendored Normal file
View File

@@ -0,0 +1,35 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

161
README.md Normal file
View File

@@ -0,0 +1,161 @@
---
license: cc-by-nc-4.0
datasets:
- starmpcc/Asclepius-Synthetic-Clinical-Notes
language:
- en
pipeline_tag: text2text-generation
tags:
- medical
---
# Model Card for Model ID
<!-- Provide a quick summary of what the model is/does. -->
This is an official model checkpoint for Asclepius-Llama2-13B [(arxiv)](https://arxiv.org/abs/2309.00237).
This model is an enhanced version of Asclepius-13B, by replacing the base model with Llama-2 and increasing the max sequence length to 4096.
## UPDATE
### 2024.01.10
- Asclepius-R, the variant of Asclepius that trained on MIMIC-III discharge summaries, is now available on [Physionet](https://physionet.org/content/asclepius-r/1.0.0/)!
## Model Details
### Model Description
<!-- Provide a longer summary of what this model is. -->
- **Model type:** Clinical LLM (Large Language Model)
- **Language(s) (NLP):** English
- **License:** CC-BY-NC-SA 4.0
- **Finetuned from model [optional]:** Llama2-13B
### Model Sources [optional]
<!-- Provide the basic links for the model. -->
- **Repository:** https://github.com/starmpcc/Asclepius
- **Paper [optional]:** https://arxiv.org/abs/2309.00237
- **Data:** https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes
## Uses
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
This model can perform below 8 clinical NLP tasks, with clincal notes.
- Named Entity Recognition
- Abbreviation Expansion
- Relation Extraction
- Temporal Information Extraction
- Coreference Resolution
- Paraphrasing
- Summarization
- Question Answering
### Direct Use
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
[More Information Needed]
### Downstream Use [optional]
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
[More Information Needed]
### Out-of-Scope Use
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
ONLY USE THIS MODEL FOR RESEARCH PURPOSE!!
## How to Get Started with the Model
```python
prompt = """You are an intelligent clinical languge model.
Below is a snippet of patient's discharge summary and a following instruction from healthcare professional.
Write a response that appropriately completes the instruction.
The response should provide the accurate answer to the instruction, while being concise.
[Discharge Summary Begin]
{note}
[Discharge Summary End]
[Instruction Begin]
{question}
[Instruction End]
"""
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("starmpcc/Asclepius-Llama2-13B", use_fast=False)
model = AutoModelForCausalLM.from_pretrained("starmpcc/Asclepius-Llama13-7B")
note = "This is a sample note"
question = "What is the diagnosis?"
model_input = prompt.format(note=note, question=question)
input_ids = tokenizer(model_input, return_tensors="pt").input_ids
output = model.generate(input_ids)
print(tokenizer.decode(output[0]))
```
## Training Details
### Training Data
<!-- This should link to a Data Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes
### Training Procedure
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
- Initial training was conducted using causal language modeling on synthetic clinical notes.
- It was then fine-tuned with clinical instruction-response pairs.
- For a comprehensive overview of our methods, our upcoming paper will serve as a resource.
#### Training Hyperparameters
- We followed config used in [Stanford Alpaca](https://github.com/tatsu-lab/stanford_alpaca)
-
#### Speeds, Sizes, Times
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
- Pre-Training (1 epoch): 1h 58m with 8x A100 80G
- Instruction Fine-Tuning (3 epoch): 12h 39m with 8x A100 80G
## Citation
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
**BibTeX:**
```
@misc{kweon2023publicly,
title={Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes},
author={Sunjun Kweon and Junu Kim and Jiyoun Kim and Sujeong Im and Eunbyeol Cho and Seongsu Bae and Jungwoo Oh and Gyubok Lee and Jong Hak Moon and Seng Chan You and Seungjin Baek and Chang Hoon Han and Yoon Bin Jung and Yohan Jo and Edward Choi},
year={2023},
eprint={2309.00237},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
```
# [Open LLM Leaderboard Evaluation Results](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard)
Detailed results can be found [here](https://huggingface.co/datasets/open-llm-leaderboard/details_starmpcc__Asclepius-Llama2-13B)
| Metric | Value |
|-----------------------|---------------------------|
| Avg. | 44.85 |
| ARC (25-shot) | 55.89 |
| HellaSwag (10-shot) | 79.66 |
| MMLU (5-shot) | 52.38 |
| TruthfulQA (0-shot) | 40.76 |
| Winogrande (5-shot) | 72.69 |
| GSM8K (5-shot) | 0.15 |
| DROP (3-shot) | 12.42 |

3
added_tokens.json Normal file
View File

@@ -0,0 +1,3 @@
{
"<pad>": 32000
}

26
config.json Normal file
View File

@@ -0,0 +1,26 @@
{
"_name_or_path": "/home/junukim/synthetic_167k_llama_2_13b",
"architectures": [
"LlamaForCausalLM"
],
"bos_token_id": 1,
"eos_token_id": 2,
"hidden_act": "silu",
"hidden_size": 5120,
"initializer_range": 0.02,
"intermediate_size": 13824,
"max_position_embeddings": 4096,
"model_type": "llama",
"num_attention_heads": 40,
"num_hidden_layers": 40,
"num_key_value_heads": 40,
"pad_token_id": 0,
"pretraining_tp": 1,
"rms_norm_eps": 1e-05,
"rope_scaling": null,
"tie_word_embeddings": false,
"torch_dtype": "float32",
"transformers_version": "4.28.0",
"use_cache": true,
"vocab_size": 32000
}

10
generation_config.json Normal file
View File

@@ -0,0 +1,10 @@
{
"_from_model_config": true,
"bos_token_id": 1,
"eos_token_id": 2,
"max_length": 4096,
"pad_token_id": 0,
"temperature": 0.9,
"top_p": 0.6,
"transformers_version": "4.28.0"
}

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4868bfcc3552ed85407c6532b221cbdbca4eb6374be4694c0b93994f649fd2c9
size 9956546443

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fa547eee800c469ba81da9f6843c5d41b748bef30f5070e849993ad2e45d7b9c
size 9940859009

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6c9e090d5d65c9f535b52412c3f409972b6320f5d5ba77abb57b6c9525ed5f3f
size 9940859567

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:aa29679eb9012967cf1a1399f9d47ba159924bfdf9b40dec50988f926588c7dd
size 9867417913

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:86f5114d04c408806cdccfe1fc6bb3baf1310dcc248575988a11a6ec49a6f929
size 9867459649

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c2e012de740546af8770ee664b2c6d97f8b863bc04353aede20d4abf42695b31
size 2490476719

View File

@@ -0,0 +1,410 @@
{
"metadata": {
"total_size": 52063467520
},
"weight_map": {
"lm_head.weight": "pytorch_model-00006-of-00006.bin",
"model.embed_tokens.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.0.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.0.mlp.down_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.0.mlp.gate_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.0.mlp.up_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.0.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.0.self_attn.k_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.0.self_attn.o_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.0.self_attn.q_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.0.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
"model.layers.0.self_attn.v_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.1.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.1.mlp.down_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.1.mlp.gate_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.1.mlp.up_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.1.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.1.self_attn.k_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.1.self_attn.o_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.1.self_attn.q_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.1.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
"model.layers.1.self_attn.v_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.10.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.10.mlp.down_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.10.mlp.gate_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.10.mlp.up_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.10.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.10.self_attn.k_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.10.self_attn.o_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.10.self_attn.q_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.10.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
"model.layers.10.self_attn.v_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.11.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.11.mlp.down_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.11.mlp.gate_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.11.mlp.up_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.11.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.11.self_attn.k_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.11.self_attn.o_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.11.self_attn.q_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.11.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
"model.layers.11.self_attn.v_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.12.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.12.mlp.down_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.12.mlp.gate_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.12.mlp.up_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.12.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.12.self_attn.k_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.12.self_attn.o_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.12.self_attn.q_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.12.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
"model.layers.12.self_attn.v_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.13.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.13.mlp.down_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.13.mlp.gate_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.13.mlp.up_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.13.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.13.self_attn.k_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.13.self_attn.o_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.13.self_attn.q_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.13.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
"model.layers.13.self_attn.v_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.14.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.14.mlp.down_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.14.mlp.gate_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.14.mlp.up_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.14.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.14.self_attn.k_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.14.self_attn.o_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.14.self_attn.q_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.14.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
"model.layers.14.self_attn.v_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.15.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.15.mlp.down_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.15.mlp.gate_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.15.mlp.up_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.15.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.15.self_attn.k_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.15.self_attn.o_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.15.self_attn.q_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.15.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
"model.layers.15.self_attn.v_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.16.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.16.mlp.down_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.16.mlp.gate_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.16.mlp.up_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.16.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.16.self_attn.k_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.16.self_attn.o_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.16.self_attn.q_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.16.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
"model.layers.16.self_attn.v_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.17.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.17.mlp.down_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.17.mlp.gate_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.17.mlp.up_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.17.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.17.self_attn.k_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.17.self_attn.o_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.17.self_attn.q_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.17.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
"model.layers.17.self_attn.v_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.18.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.18.mlp.down_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.18.mlp.gate_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.18.mlp.up_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.18.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.18.self_attn.k_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.18.self_attn.o_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.18.self_attn.q_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.18.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
"model.layers.18.self_attn.v_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.19.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.19.mlp.down_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.19.mlp.gate_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.19.mlp.up_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.19.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.19.self_attn.k_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.19.self_attn.o_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.19.self_attn.q_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.19.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
"model.layers.19.self_attn.v_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.2.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.2.mlp.down_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.2.mlp.gate_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.2.mlp.up_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.2.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.2.self_attn.k_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.2.self_attn.o_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.2.self_attn.q_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.2.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
"model.layers.2.self_attn.v_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.20.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.20.mlp.down_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.20.mlp.gate_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.20.mlp.up_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.20.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.20.self_attn.k_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.20.self_attn.o_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.20.self_attn.q_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.20.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
"model.layers.20.self_attn.v_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.21.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.21.mlp.down_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.21.mlp.gate_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.21.mlp.up_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.21.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.21.self_attn.k_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.21.self_attn.o_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.21.self_attn.q_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.21.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
"model.layers.21.self_attn.v_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.22.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.22.mlp.down_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.22.mlp.gate_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.22.mlp.up_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.22.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.22.self_attn.k_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.22.self_attn.o_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.22.self_attn.q_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.22.self_attn.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
"model.layers.22.self_attn.v_proj.weight": "pytorch_model-00003-of-00006.bin",
"model.layers.23.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.23.mlp.down_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.23.mlp.gate_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.23.mlp.up_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.23.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.23.self_attn.k_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.23.self_attn.o_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.23.self_attn.q_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.23.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
"model.layers.23.self_attn.v_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.24.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.24.mlp.down_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.24.mlp.gate_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.24.mlp.up_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.24.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.24.self_attn.k_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.24.self_attn.o_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.24.self_attn.q_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.24.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
"model.layers.24.self_attn.v_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.25.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.25.mlp.down_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.25.mlp.gate_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.25.mlp.up_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.25.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.25.self_attn.k_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.25.self_attn.o_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.25.self_attn.q_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.25.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
"model.layers.25.self_attn.v_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.26.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.26.mlp.down_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.26.mlp.gate_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.26.mlp.up_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.26.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.26.self_attn.k_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.26.self_attn.o_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.26.self_attn.q_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.26.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
"model.layers.26.self_attn.v_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.27.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.27.mlp.down_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.27.mlp.gate_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.27.mlp.up_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.27.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.27.self_attn.k_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.27.self_attn.o_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.27.self_attn.q_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.27.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
"model.layers.27.self_attn.v_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.28.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.28.mlp.down_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.28.mlp.gate_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.28.mlp.up_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.28.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.28.self_attn.k_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.28.self_attn.o_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.28.self_attn.q_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.28.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
"model.layers.28.self_attn.v_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.29.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.29.mlp.down_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.29.mlp.gate_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.29.mlp.up_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.29.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.29.self_attn.k_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.29.self_attn.o_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.29.self_attn.q_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.29.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
"model.layers.29.self_attn.v_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.3.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.3.mlp.down_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.3.mlp.gate_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.3.mlp.up_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.3.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.3.self_attn.k_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.3.self_attn.o_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.3.self_attn.q_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.3.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
"model.layers.3.self_attn.v_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.30.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.30.mlp.down_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.30.mlp.gate_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.30.mlp.up_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.30.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.30.self_attn.k_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.30.self_attn.o_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.30.self_attn.q_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.30.self_attn.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
"model.layers.30.self_attn.v_proj.weight": "pytorch_model-00004-of-00006.bin",
"model.layers.31.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.31.mlp.down_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.31.mlp.gate_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.31.mlp.up_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.31.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.31.self_attn.k_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.31.self_attn.o_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.31.self_attn.q_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.31.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
"model.layers.31.self_attn.v_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.32.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.32.mlp.down_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.32.mlp.gate_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.32.mlp.up_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.32.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.32.self_attn.k_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.32.self_attn.o_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.32.self_attn.q_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.32.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
"model.layers.32.self_attn.v_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.33.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.33.mlp.down_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.33.mlp.gate_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.33.mlp.up_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.33.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.33.self_attn.k_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.33.self_attn.o_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.33.self_attn.q_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.33.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
"model.layers.33.self_attn.v_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.34.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.34.mlp.down_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.34.mlp.gate_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.34.mlp.up_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.34.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.34.self_attn.k_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.34.self_attn.o_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.34.self_attn.q_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.34.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
"model.layers.34.self_attn.v_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.35.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.35.mlp.down_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.35.mlp.gate_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.35.mlp.up_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.35.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.35.self_attn.k_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.35.self_attn.o_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.35.self_attn.q_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.35.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
"model.layers.35.self_attn.v_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.36.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.36.mlp.down_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.36.mlp.gate_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.36.mlp.up_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.36.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.36.self_attn.k_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.36.self_attn.o_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.36.self_attn.q_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.36.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
"model.layers.36.self_attn.v_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.37.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.37.mlp.down_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.37.mlp.gate_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.37.mlp.up_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.37.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.37.self_attn.k_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.37.self_attn.o_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.37.self_attn.q_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.37.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
"model.layers.37.self_attn.v_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.38.input_layernorm.weight": "pytorch_model-00006-of-00006.bin",
"model.layers.38.mlp.down_proj.weight": "pytorch_model-00006-of-00006.bin",
"model.layers.38.mlp.gate_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.38.mlp.up_proj.weight": "pytorch_model-00006-of-00006.bin",
"model.layers.38.post_attention_layernorm.weight": "pytorch_model-00006-of-00006.bin",
"model.layers.38.self_attn.k_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.38.self_attn.o_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.38.self_attn.q_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.38.self_attn.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
"model.layers.38.self_attn.v_proj.weight": "pytorch_model-00005-of-00006.bin",
"model.layers.39.input_layernorm.weight": "pytorch_model-00006-of-00006.bin",
"model.layers.39.mlp.down_proj.weight": "pytorch_model-00006-of-00006.bin",
"model.layers.39.mlp.gate_proj.weight": "pytorch_model-00006-of-00006.bin",
"model.layers.39.mlp.up_proj.weight": "pytorch_model-00006-of-00006.bin",
"model.layers.39.post_attention_layernorm.weight": "pytorch_model-00006-of-00006.bin",
"model.layers.39.self_attn.k_proj.weight": "pytorch_model-00006-of-00006.bin",
"model.layers.39.self_attn.o_proj.weight": "pytorch_model-00006-of-00006.bin",
"model.layers.39.self_attn.q_proj.weight": "pytorch_model-00006-of-00006.bin",
"model.layers.39.self_attn.rotary_emb.inv_freq": "pytorch_model-00006-of-00006.bin",
"model.layers.39.self_attn.v_proj.weight": "pytorch_model-00006-of-00006.bin",
"model.layers.4.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.4.mlp.down_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.4.mlp.gate_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.4.mlp.up_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.4.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.4.self_attn.k_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.4.self_attn.o_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.4.self_attn.q_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.4.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
"model.layers.4.self_attn.v_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.5.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.5.mlp.down_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.5.mlp.gate_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.5.mlp.up_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.5.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.5.self_attn.k_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.5.self_attn.o_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.5.self_attn.q_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.5.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
"model.layers.5.self_attn.v_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.6.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.6.mlp.down_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.6.mlp.gate_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.6.mlp.up_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.6.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.6.self_attn.k_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.6.self_attn.o_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.6.self_attn.q_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.6.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
"model.layers.6.self_attn.v_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.7.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.7.mlp.down_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.7.mlp.gate_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.7.mlp.up_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.7.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.7.self_attn.k_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.7.self_attn.o_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.7.self_attn.q_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.7.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
"model.layers.7.self_attn.v_proj.weight": "pytorch_model-00001-of-00006.bin",
"model.layers.8.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.8.mlp.down_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.8.mlp.gate_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.8.mlp.up_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.8.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.8.self_attn.k_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.8.self_attn.o_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.8.self_attn.q_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.8.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
"model.layers.8.self_attn.v_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.9.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.9.mlp.down_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.9.mlp.gate_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.9.mlp.up_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.9.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.9.self_attn.k_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.9.self_attn.o_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.9.self_attn.q_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.layers.9.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
"model.layers.9.self_attn.v_proj.weight": "pytorch_model-00002-of-00006.bin",
"model.norm.weight": "pytorch_model-00006-of-00006.bin"
}
}

6
special_tokens_map.json Normal file
View File

@@ -0,0 +1,6 @@
{
"bos_token": "<s>",
"eos_token": "</s>",
"pad_token": "<s>",
"unk_token": "<unk>"
}

3
tokenizer.model Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9e556afd44213b6bd1be2b850ebbbd98f5481437a8021afaf58ee7fb1818d347
size 499723

35
tokenizer_config.json Normal file
View File

@@ -0,0 +1,35 @@
{
"add_bos_token": true,
"add_eos_token": false,
"bos_token": {
"__type": "AddedToken",
"content": "<s>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"clean_up_tokenization_spaces": false,
"eos_token": {
"__type": "AddedToken",
"content": "</s>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"legacy": false,
"model_max_length": 4096,
"pad_token": null,
"padding_side": "right",
"sp_model_kwargs": {},
"tokenizer_class": "LlamaTokenizer",
"unk_token": {
"__type": "AddedToken",
"content": "<unk>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
}
}

22255
trainer_state.json Normal file

File diff suppressed because it is too large Load Diff

3
training_args.bin Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b6f81a806e8015237d1e5872481a137ab02a427cfa0851b5f3f2c3f543fc4c82
size 3835