初始化项目,由ModelHub XC社区提供模型

Model: okwinds/MiroThinker-14B-DPO-v0.1
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-20 06:18:14 +08:00
commit 7674b8eeab
19 changed files with 909815 additions and 0 deletions

47
.gitattributes vendored Normal file
View File

@@ -0,0 +1,47 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

134
README.md Normal file
View File

@@ -0,0 +1,134 @@
---
frameworks:
- Pytorch
license: Apache License 2.0
tasks:
- text-generation
base_model:
- okwinds/MiroThinker-14B-SFT-v0.1
tags:
- agent
- open-source
- miromind
---
本模型转载自 huggingface 【[miromind-ai](https://huggingface.co/miromind-ai)】
#### 📖 关于项目相关的研究,可阅读公众号“觉察流”文章👇</br>
《[MiroMind-M1:如何用CAMPO算法打造高效且可复现的全栈开源推理模型](https://mp.weixin.qq.com/s/REPzzgsUjDMikg4jIo9KRg)》
#### _本仓库作者在此 👇🏻 扫一扫_
<img src="https://www.modelscope.cn/models/okwinds/GPT-2/resolve/master/qrcode_for_jcl_258.jpg" />
---
SDK下载
```bash
#安装ModelScope
pip install modelscope
```
```python
#SDK模型下载
from modelscope import snapshot_download
model_dir = snapshot_download('okwinds/MiroThinker-14B-DPO-v0.1')
```
Git下载
```
#Git模型下载
git clone https://www.modelscope.cn/okwinds/MiroThinker-14B-DPO-v0.1.git
```
# 官方 MiroThinker-14B-DPO-v0.1 简介
<div align="center">
<img src="https://cdn-uploads.huggingface.co/production/uploads/68525b342230a897a65cc1c0/87mYQ_a-4jpnMkVR4hrgm.png" width="55%" alt="MiroThinker" />
</div>
<!-- <hr> -->
<div align="center">
[![Demo](https://img.shields.io/badge/Demo-FFB300?style=for-the-badge&logo=airplayvideo&logoColor=white)](https://dr.miromind.ai/)
[![Models](https://img.shields.io/badge/Models-5EDDD2?style=for-the-badge&logo=huggingface&logoColor=ffffff&labelColor)](https://www.modelscope.cn/collections/MiroFlow-c89926d9ab8845)
[![Data](https://img.shields.io/badge/Data-0040A1?style=for-the-badge&logo=huggingface&logoColor=ffffff&labelColor)](https://www.modelscope.cn//datasets/okwinds/MiroVerse-v0.1)
[![Blog](https://img.shields.io/badge/Blog-4285F4?style=for-the-badge&logo=google-chrome&logoColor=white)](https://miromind.ai/blog/miromind-open-deep-research)
[![Github](https://img.shields.io/badge/GitHub-24292F?style=for-the-badge&logo=github&logoColor=white)](https://github.com/MiroMindAI/MiroThinker)
[![Discord](https://img.shields.io/badge/Discord-5865F2?style=for-the-badge&logo=discord&logoColor=white)](https://discord.com/invite/GPqEnkzQZd)
[![WeChat](https://img.shields.io/badge/WeChat-07C160?style=for-the-badge&logo=wechat&logoColor=white)](https://cdn-uploads.huggingface.co/production/uploads/68525b342230a897a65cc1c0/SGK70isvVpeJwk_fny9sb.png)
[![RedNote](https://img.shields.io/badge/RedNote-FF2442?style=for-the-badge&logo=revoltdotchat&logoColor=white)](https://www.xiaohongshu.com/user/profile/663098830000000003033edc)
[![Website](https://img.shields.io/badge/Website-4285F4?style=for-the-badge&logo=monster&logoColor=white)](https://miromind.ai/)
</div>
## Introduction
MiroThinker is an open-source agentic model series built on top of Qwen3. Designed for deep research and complex, long-horizon problem solving, it integrates strong capabilities in task decomposition, multi-hop reasoning, retrieval-augmented generation, code execution, web browsing, and document/file processing, making it suitable for a wide range of real-world applications.
We have released the MiroThinker-v0.1 series, including both SFT and DPO variants at parameter scales of 8B, 14B, and 32B. Notably, MiroThinker v0.1 achieves state-of-the-art performance among open-source models on the [GAIA benchmark](https://huggingface.co/datasets/gaia-benchmark/GAIA), a rigorous evaluation suite for advanced agentic capabilities, demonstrating its strength in long-context, decision-intensive, and real-world task scenarios.
## Online Demo
Welcome to try out our online demo [here](https://dr.miromind.ai/). In this demo, we have deployed our [MiroThinker-32B-DPO-v0.1](https://huggingface.co/miromind-ai/MiroThinker-32B-DPO-v0.1) along with commercial tools (you can find more details in our [GitHub](https://github.com/MiroMindAI/MiroThinker)), aiming to deliver a better experience.
## Performance
### GAIA Benchmark
| **Method** | Text-103<br>Best Pass@1 | Text-103<br>Pass@1 (Avg@8) | Val-165<br>Best Pass@1 | Val-165<br>Pass@1 (Avg@8) |
| ----------------------------------------------------------------- | :--: | :--: | :--: | :--: |
| Search-o1-7B | 17.5 | - | - | - |
| R1-Searcher-7B | 20.4 | - | - | - |
| WebDancer-7B | 31.0 | - | - | - |
| WebSailor-7B | 37.9 | - | - | - |
| CK-Pro-8B | 40.3 | - | 32.7 | - |
| MiroThinker-8B-SFT-v0.1 | 44.7 | 40.1 | 34.6 | 31.8 |
| &nbsp;&nbsp;&nbsp;&nbsp;+ Commercial Tools | 46.6 | 42.1 | 37.6 | 33.9 |
| MiroThinker-8B-DPO-v0.1 | 46.6 | 44.8 | 37.0 | 35.4 |
| &nbsp;&nbsp;&nbsp;&nbsp;+ Commercial Tools | 50.5 | 46.7 | 38.2 | 35.9 |
| | | | | |
| Search-o1-32B | 28.2 | - | - | - |
| WebThinker-32B-RL | 48.5 | - | - | - |
| WebDancer-QwQ-32B | 51.5 | - | - | - |
| WebSailor-32B | 53.2 | - | - | - |
| WebShaper-QwQ-32B | 53.3 | - | - | - |
| WebShaper-72B | 60.1 | - | - | - |
| MiroThinker-14B-SFT-v0.1 | 47.6 | 44.4 | 37.0 | 34.4 |
| &nbsp;&nbsp;&nbsp;&nbsp;+ Commercial Tools | 49.5 | 47.5 | 41.8 | 39.8 |
| MiroThinker-14B-DPO-v0.1 | 48.5 | 46.6 | 42.4 | 39.2 |
| &nbsp;&nbsp;&nbsp;&nbsp;+ Commercial Tools | 52.4 | 48.5 | 45.5 | 42.0 |
| MiroThinker-32B-SFT-v0.1 | 55.3 | 51.3 | 44.9 | 42.7 |
| &nbsp;&nbsp;&nbsp;&nbsp;+ Commercial Tools | 58.3 | 54.2 | 48.5 | 45.8 |
| <span style="white-space:nowrap;">MiroThinker-32B-DPO-v0.1</span> | 57.3 | 54.1 | 48.5 | 45.9 |
| &nbsp;&nbsp;&nbsp;&nbsp;+ Commercial Tools | **60.2** | **57.9** | **50.9** | **48.9** |
1. Following the practices of WebThinker, WebAgents, and CognitiveKernel, we report the Best Pass@1, the highest score across three runs, which often reflects stronger performance, though it may exhibit some variability. To provide a more stable measure, we additionally report Pass@1 (Avg@8), which offers greater consistency at the cost of slightly lower scores.
2. For consistency with prior open-source works, we evaluate GAIA-Text-103 using the WebAgents LLM-as-judge template, and report results on GAIA-Val-165 using the official GAIA scorer script.
3. By default, we use open-source tools wherever possible, except for the code tool [E2B](https://github.com/e2b-dev/E2B) and the Google search tool [Serper](https://serper.dev/). We use [Whisper](https://huggingface.co/openai/whisper-large-v3-turbo), [Qwen2.5-VL-72B-Instruct](https://www.modelscope.cn/models/Qwen/Qwen2.5-VL-72B-Instruct), and [Qwen3-235B-A22B-Thinking-2507](https://www.modelscope.cn/models/Qwen/Qwen3-235B-A22B-Thinking-2507) in our implementation. The framework can be easily extended to other open-source tools of your choice.
4. Commercial tools were mainly used for multimodal capabilities and certain complex reasoning subtasks. The majority of tasks, including planning, browsing, refinement, navigation, and more, were handled by our models.
### More Benchmarks
Coming soon
## Quick Start
MiroThinker-v0.1 is trained on our large-scale, high-quality trajectory and preference datasets [MiroVerse-v0.1](https://www.modelscope.cn/datasets/okwinds/MiroVerse-v0.1), utilizing the efficient training framework [MiroTrain](https://github.com/MiroMindAI/MiroTrain), and enhanced with tool-use capabilities through our agentic framework [MiroFlow](https://github.com/MiroMindAI/MiroFlow).
To promote reproducibility and benefit the community, we decided to open-source the entire suite mentioned above. For more technical details, evaluation results, and usage tutorials, please visit our [GitHub repository](https://github.com/MiroMindAI/MiroThinker).
## License
MiroThinker-v0.1 is licensed under Apache 2.0.
## Contact Us
MiroThinker is developed by the MiroMind Foundation Model Team.
If you would like to leave us a message, feel free to get in touch.
In addition to [GitHub](https://github.com/MiroMindAI/),
[Discord](https://discord.com/invite/GPqEnkzQZd),
[WeChat](https://cdn-uploads.huggingface.co/production/uploads/68525b342230a897a65cc1c0/SGK70isvVpeJwk_fny9sb.png),
and [RedNote](https://www.xiaohongshu.com/user/profile/663098830000000003033edc),
you can also reach us via email at talent@miromind.ai.

30
config.json Normal file
View File

@@ -0,0 +1,30 @@
{
"architectures": [
"Qwen3ForCausalLM"
],
"attention_bias": false,
"attention_dropout": 0.0,
"bos_token_id": 151643,
"eos_token_id": 151645,
"head_dim": 128,
"hidden_act": "silu",
"hidden_size": 5120,
"initializer_range": 0.02,
"intermediate_size": 17408,
"max_position_embeddings": 40960,
"max_window_layers": 40,
"model_type": "qwen3",
"num_attention_heads": 40,
"num_hidden_layers": 40,
"num_key_value_heads": 8,
"rms_norm_eps": 1e-06,
"rope_scaling": null,
"rope_theta": 1000000,
"sliding_window": null,
"tie_word_embeddings": false,
"torch_dtype": "bfloat16",
"transformers_version": "4.51.0",
"use_cache": true,
"use_sliding_window": false,
"vocab_size": 151936
}

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework":"Pytorch","task":"text-generation"}

13
generation_config.json Normal file
View File

@@ -0,0 +1,13 @@
{
"bos_token_id": 151643,
"do_sample": true,
"eos_token_id": [
151645,
151643
],
"pad_token_id": 151643,
"temperature": 0.6,
"top_k": 20,
"top_p": 0.95,
"transformers_version": "4.51.0"
}

8
hash.txt Normal file
View File

@@ -0,0 +1,8 @@
10287609eb87646690a49ba2797f5dd370aa81398d2e2e3f3f5175c717c7f439 model-00001-of-00008.safetensors
94ac10c9cc5879d2e51cf8505b276702f358d8867876ff2562a9181af7169611 model-00002-of-00008.safetensors
b023bfd53c04164475aee029643df0945cc3b6c0cbfe241e5dec2eedd689b6b5 model-00003-of-00008.safetensors
040f53da2df5094bd62b2355c860badbad178ea2af8defd14c8dad3c3fcdfc55 model-00004-of-00008.safetensors
c678b23c35f168ebf56934eaf3445b346dc8338aa1bb0b2ece005896fd9101d7 model-00005-of-00008.safetensors
87c72e4984606095f012b713421090aba72fae38d4df647f75f13d27384480b9 model-00006-of-00008.safetensors
7763368905fb2ec6f478bb75931c398d0afddee5b4f29e343e9dc139da80cf44 model-00007-of-00008.safetensors
264e3db33d333aa9e1eb46e587ee7fa58fcc3caea48ea83f5983b61f44f1d800 model-00008-of-00008.safetensors

151388
merges.txt Normal file

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:10287609eb87646690a49ba2797f5dd370aa81398d2e2e3f3f5175c717c7f439
size 3841788544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:94ac10c9cc5879d2e51cf8505b276702f358d8867876ff2562a9181af7169611
size 3963750816

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b023bfd53c04164475aee029643df0945cc3b6c0cbfe241e5dec2eedd689b6b5
size 3963750880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:040f53da2df5094bd62b2355c860badbad178ea2af8defd14c8dad3c3fcdfc55
size 3963750880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c678b23c35f168ebf56934eaf3445b346dc8338aa1bb0b2ece005896fd9101d7
size 3963750880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:87c72e4984606095f012b713421090aba72fae38d4df647f75f13d27384480b9
size 3963750880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7763368905fb2ec6f478bb75931c398d0afddee5b4f29e343e9dc139da80cf44
size 3963750880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:264e3db33d333aa9e1eb46e587ee7fa58fcc3caea48ea83f5983b61f44f1d800
size 1912371880

View File

@@ -0,0 +1,450 @@
{
"metadata": {
"total_size": 29536614400
},
"weight_map": {
"model.embed_tokens.weight": "model-00001-of-00008.safetensors",
"model.layers.0.input_layernorm.weight": "model-00001-of-00008.safetensors",
"model.layers.0.mlp.down_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.0.mlp.gate_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.0.mlp.up_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.0.post_attention_layernorm.weight": "model-00001-of-00008.safetensors",
"model.layers.0.self_attn.k_norm.weight": "model-00001-of-00008.safetensors",
"model.layers.0.self_attn.k_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.0.self_attn.o_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.0.self_attn.q_norm.weight": "model-00001-of-00008.safetensors",
"model.layers.0.self_attn.q_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.0.self_attn.v_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.1.input_layernorm.weight": "model-00001-of-00008.safetensors",
"model.layers.1.mlp.down_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.1.mlp.gate_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.1.mlp.up_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.1.post_attention_layernorm.weight": "model-00001-of-00008.safetensors",
"model.layers.1.self_attn.k_norm.weight": "model-00001-of-00008.safetensors",
"model.layers.1.self_attn.k_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.1.self_attn.o_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.1.self_attn.q_norm.weight": "model-00001-of-00008.safetensors",
"model.layers.1.self_attn.q_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.1.self_attn.v_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.2.input_layernorm.weight": "model-00001-of-00008.safetensors",
"model.layers.2.mlp.down_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.2.mlp.gate_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.2.mlp.up_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.2.post_attention_layernorm.weight": "model-00001-of-00008.safetensors",
"model.layers.2.self_attn.k_norm.weight": "model-00001-of-00008.safetensors",
"model.layers.2.self_attn.k_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.2.self_attn.o_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.2.self_attn.q_norm.weight": "model-00001-of-00008.safetensors",
"model.layers.2.self_attn.q_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.2.self_attn.v_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.3.mlp.gate_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.3.self_attn.k_norm.weight": "model-00001-of-00008.safetensors",
"model.layers.3.self_attn.k_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.3.self_attn.o_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.3.self_attn.q_norm.weight": "model-00001-of-00008.safetensors",
"model.layers.3.self_attn.q_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.3.self_attn.v_proj.weight": "model-00001-of-00008.safetensors",
"model.layers.3.input_layernorm.weight": "model-00002-of-00008.safetensors",
"model.layers.3.mlp.down_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.3.mlp.up_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.3.post_attention_layernorm.weight": "model-00002-of-00008.safetensors",
"model.layers.4.input_layernorm.weight": "model-00002-of-00008.safetensors",
"model.layers.4.mlp.down_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.4.mlp.gate_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.4.mlp.up_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.4.post_attention_layernorm.weight": "model-00002-of-00008.safetensors",
"model.layers.4.self_attn.k_norm.weight": "model-00002-of-00008.safetensors",
"model.layers.4.self_attn.k_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.4.self_attn.o_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.4.self_attn.q_norm.weight": "model-00002-of-00008.safetensors",
"model.layers.4.self_attn.q_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.4.self_attn.v_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.5.input_layernorm.weight": "model-00002-of-00008.safetensors",
"model.layers.5.mlp.down_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.5.mlp.gate_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.5.mlp.up_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.5.post_attention_layernorm.weight": "model-00002-of-00008.safetensors",
"model.layers.5.self_attn.k_norm.weight": "model-00002-of-00008.safetensors",
"model.layers.5.self_attn.k_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.5.self_attn.o_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.5.self_attn.q_norm.weight": "model-00002-of-00008.safetensors",
"model.layers.5.self_attn.q_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.5.self_attn.v_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.6.input_layernorm.weight": "model-00002-of-00008.safetensors",
"model.layers.6.mlp.down_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.6.mlp.gate_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.6.mlp.up_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.6.post_attention_layernorm.weight": "model-00002-of-00008.safetensors",
"model.layers.6.self_attn.k_norm.weight": "model-00002-of-00008.safetensors",
"model.layers.6.self_attn.k_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.6.self_attn.o_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.6.self_attn.q_norm.weight": "model-00002-of-00008.safetensors",
"model.layers.6.self_attn.q_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.6.self_attn.v_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.7.input_layernorm.weight": "model-00002-of-00008.safetensors",
"model.layers.7.mlp.down_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.7.mlp.gate_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.7.mlp.up_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.7.post_attention_layernorm.weight": "model-00002-of-00008.safetensors",
"model.layers.7.self_attn.k_norm.weight": "model-00002-of-00008.safetensors",
"model.layers.7.self_attn.k_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.7.self_attn.o_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.7.self_attn.q_norm.weight": "model-00002-of-00008.safetensors",
"model.layers.7.self_attn.q_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.7.self_attn.v_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.8.input_layernorm.weight": "model-00002-of-00008.safetensors",
"model.layers.8.mlp.down_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.8.mlp.gate_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.8.mlp.up_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.8.post_attention_layernorm.weight": "model-00002-of-00008.safetensors",
"model.layers.8.self_attn.k_norm.weight": "model-00002-of-00008.safetensors",
"model.layers.8.self_attn.k_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.8.self_attn.o_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.8.self_attn.q_norm.weight": "model-00002-of-00008.safetensors",
"model.layers.8.self_attn.q_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.8.self_attn.v_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.9.mlp.gate_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.9.self_attn.k_norm.weight": "model-00002-of-00008.safetensors",
"model.layers.9.self_attn.k_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.9.self_attn.o_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.9.self_attn.q_norm.weight": "model-00002-of-00008.safetensors",
"model.layers.9.self_attn.q_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.9.self_attn.v_proj.weight": "model-00002-of-00008.safetensors",
"model.layers.10.input_layernorm.weight": "model-00003-of-00008.safetensors",
"model.layers.10.mlp.down_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.10.mlp.gate_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.10.mlp.up_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.10.post_attention_layernorm.weight": "model-00003-of-00008.safetensors",
"model.layers.10.self_attn.k_norm.weight": "model-00003-of-00008.safetensors",
"model.layers.10.self_attn.k_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.10.self_attn.o_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.10.self_attn.q_norm.weight": "model-00003-of-00008.safetensors",
"model.layers.10.self_attn.q_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.10.self_attn.v_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.11.input_layernorm.weight": "model-00003-of-00008.safetensors",
"model.layers.11.mlp.down_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.11.mlp.gate_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.11.mlp.up_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.11.post_attention_layernorm.weight": "model-00003-of-00008.safetensors",
"model.layers.11.self_attn.k_norm.weight": "model-00003-of-00008.safetensors",
"model.layers.11.self_attn.k_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.11.self_attn.o_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.11.self_attn.q_norm.weight": "model-00003-of-00008.safetensors",
"model.layers.11.self_attn.q_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.11.self_attn.v_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.12.input_layernorm.weight": "model-00003-of-00008.safetensors",
"model.layers.12.mlp.down_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.12.mlp.gate_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.12.mlp.up_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.12.post_attention_layernorm.weight": "model-00003-of-00008.safetensors",
"model.layers.12.self_attn.k_norm.weight": "model-00003-of-00008.safetensors",
"model.layers.12.self_attn.k_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.12.self_attn.o_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.12.self_attn.q_norm.weight": "model-00003-of-00008.safetensors",
"model.layers.12.self_attn.q_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.12.self_attn.v_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.13.input_layernorm.weight": "model-00003-of-00008.safetensors",
"model.layers.13.mlp.down_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.13.mlp.gate_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.13.mlp.up_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.13.post_attention_layernorm.weight": "model-00003-of-00008.safetensors",
"model.layers.13.self_attn.k_norm.weight": "model-00003-of-00008.safetensors",
"model.layers.13.self_attn.k_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.13.self_attn.o_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.13.self_attn.q_norm.weight": "model-00003-of-00008.safetensors",
"model.layers.13.self_attn.q_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.13.self_attn.v_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.14.input_layernorm.weight": "model-00003-of-00008.safetensors",
"model.layers.14.mlp.down_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.14.mlp.gate_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.14.mlp.up_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.14.post_attention_layernorm.weight": "model-00003-of-00008.safetensors",
"model.layers.14.self_attn.k_norm.weight": "model-00003-of-00008.safetensors",
"model.layers.14.self_attn.k_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.14.self_attn.o_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.14.self_attn.q_norm.weight": "model-00003-of-00008.safetensors",
"model.layers.14.self_attn.q_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.14.self_attn.v_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.15.mlp.gate_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.15.self_attn.k_norm.weight": "model-00003-of-00008.safetensors",
"model.layers.15.self_attn.k_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.15.self_attn.o_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.15.self_attn.q_norm.weight": "model-00003-of-00008.safetensors",
"model.layers.15.self_attn.q_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.15.self_attn.v_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.9.input_layernorm.weight": "model-00003-of-00008.safetensors",
"model.layers.9.mlp.down_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.9.mlp.up_proj.weight": "model-00003-of-00008.safetensors",
"model.layers.9.post_attention_layernorm.weight": "model-00003-of-00008.safetensors",
"model.layers.15.input_layernorm.weight": "model-00004-of-00008.safetensors",
"model.layers.15.mlp.down_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.15.mlp.up_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.15.post_attention_layernorm.weight": "model-00004-of-00008.safetensors",
"model.layers.16.input_layernorm.weight": "model-00004-of-00008.safetensors",
"model.layers.16.mlp.down_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.16.mlp.gate_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.16.mlp.up_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.16.post_attention_layernorm.weight": "model-00004-of-00008.safetensors",
"model.layers.16.self_attn.k_norm.weight": "model-00004-of-00008.safetensors",
"model.layers.16.self_attn.k_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.16.self_attn.o_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.16.self_attn.q_norm.weight": "model-00004-of-00008.safetensors",
"model.layers.16.self_attn.q_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.16.self_attn.v_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.17.input_layernorm.weight": "model-00004-of-00008.safetensors",
"model.layers.17.mlp.down_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.17.mlp.gate_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.17.mlp.up_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.17.post_attention_layernorm.weight": "model-00004-of-00008.safetensors",
"model.layers.17.self_attn.k_norm.weight": "model-00004-of-00008.safetensors",
"model.layers.17.self_attn.k_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.17.self_attn.o_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.17.self_attn.q_norm.weight": "model-00004-of-00008.safetensors",
"model.layers.17.self_attn.q_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.17.self_attn.v_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.18.input_layernorm.weight": "model-00004-of-00008.safetensors",
"model.layers.18.mlp.down_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.18.mlp.gate_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.18.mlp.up_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.18.post_attention_layernorm.weight": "model-00004-of-00008.safetensors",
"model.layers.18.self_attn.k_norm.weight": "model-00004-of-00008.safetensors",
"model.layers.18.self_attn.k_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.18.self_attn.o_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.18.self_attn.q_norm.weight": "model-00004-of-00008.safetensors",
"model.layers.18.self_attn.q_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.18.self_attn.v_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.19.input_layernorm.weight": "model-00004-of-00008.safetensors",
"model.layers.19.mlp.down_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.19.mlp.gate_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.19.mlp.up_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.19.post_attention_layernorm.weight": "model-00004-of-00008.safetensors",
"model.layers.19.self_attn.k_norm.weight": "model-00004-of-00008.safetensors",
"model.layers.19.self_attn.k_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.19.self_attn.o_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.19.self_attn.q_norm.weight": "model-00004-of-00008.safetensors",
"model.layers.19.self_attn.q_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.19.self_attn.v_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.20.input_layernorm.weight": "model-00004-of-00008.safetensors",
"model.layers.20.mlp.down_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.20.mlp.gate_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.20.mlp.up_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.20.post_attention_layernorm.weight": "model-00004-of-00008.safetensors",
"model.layers.20.self_attn.k_norm.weight": "model-00004-of-00008.safetensors",
"model.layers.20.self_attn.k_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.20.self_attn.o_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.20.self_attn.q_norm.weight": "model-00004-of-00008.safetensors",
"model.layers.20.self_attn.q_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.20.self_attn.v_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.21.mlp.gate_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.21.self_attn.k_norm.weight": "model-00004-of-00008.safetensors",
"model.layers.21.self_attn.k_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.21.self_attn.o_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.21.self_attn.q_norm.weight": "model-00004-of-00008.safetensors",
"model.layers.21.self_attn.q_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.21.self_attn.v_proj.weight": "model-00004-of-00008.safetensors",
"model.layers.21.input_layernorm.weight": "model-00005-of-00008.safetensors",
"model.layers.21.mlp.down_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.21.mlp.up_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.21.post_attention_layernorm.weight": "model-00005-of-00008.safetensors",
"model.layers.22.input_layernorm.weight": "model-00005-of-00008.safetensors",
"model.layers.22.mlp.down_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.22.mlp.gate_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.22.mlp.up_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.22.post_attention_layernorm.weight": "model-00005-of-00008.safetensors",
"model.layers.22.self_attn.k_norm.weight": "model-00005-of-00008.safetensors",
"model.layers.22.self_attn.k_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.22.self_attn.o_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.22.self_attn.q_norm.weight": "model-00005-of-00008.safetensors",
"model.layers.22.self_attn.q_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.22.self_attn.v_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.23.input_layernorm.weight": "model-00005-of-00008.safetensors",
"model.layers.23.mlp.down_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.23.mlp.gate_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.23.mlp.up_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.23.post_attention_layernorm.weight": "model-00005-of-00008.safetensors",
"model.layers.23.self_attn.k_norm.weight": "model-00005-of-00008.safetensors",
"model.layers.23.self_attn.k_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.23.self_attn.o_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.23.self_attn.q_norm.weight": "model-00005-of-00008.safetensors",
"model.layers.23.self_attn.q_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.23.self_attn.v_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.24.input_layernorm.weight": "model-00005-of-00008.safetensors",
"model.layers.24.mlp.down_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.24.mlp.gate_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.24.mlp.up_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.24.post_attention_layernorm.weight": "model-00005-of-00008.safetensors",
"model.layers.24.self_attn.k_norm.weight": "model-00005-of-00008.safetensors",
"model.layers.24.self_attn.k_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.24.self_attn.o_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.24.self_attn.q_norm.weight": "model-00005-of-00008.safetensors",
"model.layers.24.self_attn.q_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.24.self_attn.v_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.25.input_layernorm.weight": "model-00005-of-00008.safetensors",
"model.layers.25.mlp.down_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.25.mlp.gate_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.25.mlp.up_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.25.post_attention_layernorm.weight": "model-00005-of-00008.safetensors",
"model.layers.25.self_attn.k_norm.weight": "model-00005-of-00008.safetensors",
"model.layers.25.self_attn.k_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.25.self_attn.o_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.25.self_attn.q_norm.weight": "model-00005-of-00008.safetensors",
"model.layers.25.self_attn.q_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.25.self_attn.v_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.26.input_layernorm.weight": "model-00005-of-00008.safetensors",
"model.layers.26.mlp.down_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.26.mlp.gate_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.26.mlp.up_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.26.post_attention_layernorm.weight": "model-00005-of-00008.safetensors",
"model.layers.26.self_attn.k_norm.weight": "model-00005-of-00008.safetensors",
"model.layers.26.self_attn.k_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.26.self_attn.o_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.26.self_attn.q_norm.weight": "model-00005-of-00008.safetensors",
"model.layers.26.self_attn.q_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.26.self_attn.v_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.27.mlp.gate_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.27.self_attn.k_norm.weight": "model-00005-of-00008.safetensors",
"model.layers.27.self_attn.k_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.27.self_attn.o_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.27.self_attn.q_norm.weight": "model-00005-of-00008.safetensors",
"model.layers.27.self_attn.q_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.27.self_attn.v_proj.weight": "model-00005-of-00008.safetensors",
"model.layers.27.input_layernorm.weight": "model-00006-of-00008.safetensors",
"model.layers.27.mlp.down_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.27.mlp.up_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.27.post_attention_layernorm.weight": "model-00006-of-00008.safetensors",
"model.layers.28.input_layernorm.weight": "model-00006-of-00008.safetensors",
"model.layers.28.mlp.down_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.28.mlp.gate_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.28.mlp.up_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.28.post_attention_layernorm.weight": "model-00006-of-00008.safetensors",
"model.layers.28.self_attn.k_norm.weight": "model-00006-of-00008.safetensors",
"model.layers.28.self_attn.k_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.28.self_attn.o_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.28.self_attn.q_norm.weight": "model-00006-of-00008.safetensors",
"model.layers.28.self_attn.q_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.28.self_attn.v_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.29.input_layernorm.weight": "model-00006-of-00008.safetensors",
"model.layers.29.mlp.down_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.29.mlp.gate_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.29.mlp.up_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.29.post_attention_layernorm.weight": "model-00006-of-00008.safetensors",
"model.layers.29.self_attn.k_norm.weight": "model-00006-of-00008.safetensors",
"model.layers.29.self_attn.k_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.29.self_attn.o_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.29.self_attn.q_norm.weight": "model-00006-of-00008.safetensors",
"model.layers.29.self_attn.q_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.29.self_attn.v_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.30.input_layernorm.weight": "model-00006-of-00008.safetensors",
"model.layers.30.mlp.down_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.30.mlp.gate_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.30.mlp.up_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.30.post_attention_layernorm.weight": "model-00006-of-00008.safetensors",
"model.layers.30.self_attn.k_norm.weight": "model-00006-of-00008.safetensors",
"model.layers.30.self_attn.k_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.30.self_attn.o_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.30.self_attn.q_norm.weight": "model-00006-of-00008.safetensors",
"model.layers.30.self_attn.q_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.30.self_attn.v_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.31.input_layernorm.weight": "model-00006-of-00008.safetensors",
"model.layers.31.mlp.down_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.31.mlp.gate_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.31.mlp.up_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.31.post_attention_layernorm.weight": "model-00006-of-00008.safetensors",
"model.layers.31.self_attn.k_norm.weight": "model-00006-of-00008.safetensors",
"model.layers.31.self_attn.k_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.31.self_attn.o_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.31.self_attn.q_norm.weight": "model-00006-of-00008.safetensors",
"model.layers.31.self_attn.q_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.31.self_attn.v_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.32.input_layernorm.weight": "model-00006-of-00008.safetensors",
"model.layers.32.mlp.down_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.32.mlp.gate_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.32.mlp.up_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.32.post_attention_layernorm.weight": "model-00006-of-00008.safetensors",
"model.layers.32.self_attn.k_norm.weight": "model-00006-of-00008.safetensors",
"model.layers.32.self_attn.k_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.32.self_attn.o_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.32.self_attn.q_norm.weight": "model-00006-of-00008.safetensors",
"model.layers.32.self_attn.q_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.32.self_attn.v_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.33.mlp.gate_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.33.self_attn.k_norm.weight": "model-00006-of-00008.safetensors",
"model.layers.33.self_attn.k_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.33.self_attn.o_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.33.self_attn.q_norm.weight": "model-00006-of-00008.safetensors",
"model.layers.33.self_attn.q_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.33.self_attn.v_proj.weight": "model-00006-of-00008.safetensors",
"model.layers.33.input_layernorm.weight": "model-00007-of-00008.safetensors",
"model.layers.33.mlp.down_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.33.mlp.up_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.33.post_attention_layernorm.weight": "model-00007-of-00008.safetensors",
"model.layers.34.input_layernorm.weight": "model-00007-of-00008.safetensors",
"model.layers.34.mlp.down_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.34.mlp.gate_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.34.mlp.up_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.34.post_attention_layernorm.weight": "model-00007-of-00008.safetensors",
"model.layers.34.self_attn.k_norm.weight": "model-00007-of-00008.safetensors",
"model.layers.34.self_attn.k_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.34.self_attn.o_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.34.self_attn.q_norm.weight": "model-00007-of-00008.safetensors",
"model.layers.34.self_attn.q_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.34.self_attn.v_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.35.input_layernorm.weight": "model-00007-of-00008.safetensors",
"model.layers.35.mlp.down_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.35.mlp.gate_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.35.mlp.up_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.35.post_attention_layernorm.weight": "model-00007-of-00008.safetensors",
"model.layers.35.self_attn.k_norm.weight": "model-00007-of-00008.safetensors",
"model.layers.35.self_attn.k_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.35.self_attn.o_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.35.self_attn.q_norm.weight": "model-00007-of-00008.safetensors",
"model.layers.35.self_attn.q_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.35.self_attn.v_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.36.input_layernorm.weight": "model-00007-of-00008.safetensors",
"model.layers.36.mlp.down_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.36.mlp.gate_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.36.mlp.up_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.36.post_attention_layernorm.weight": "model-00007-of-00008.safetensors",
"model.layers.36.self_attn.k_norm.weight": "model-00007-of-00008.safetensors",
"model.layers.36.self_attn.k_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.36.self_attn.o_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.36.self_attn.q_norm.weight": "model-00007-of-00008.safetensors",
"model.layers.36.self_attn.q_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.36.self_attn.v_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.37.input_layernorm.weight": "model-00007-of-00008.safetensors",
"model.layers.37.mlp.down_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.37.mlp.gate_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.37.mlp.up_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.37.post_attention_layernorm.weight": "model-00007-of-00008.safetensors",
"model.layers.37.self_attn.k_norm.weight": "model-00007-of-00008.safetensors",
"model.layers.37.self_attn.k_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.37.self_attn.o_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.37.self_attn.q_norm.weight": "model-00007-of-00008.safetensors",
"model.layers.37.self_attn.q_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.37.self_attn.v_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.38.input_layernorm.weight": "model-00007-of-00008.safetensors",
"model.layers.38.mlp.down_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.38.mlp.gate_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.38.mlp.up_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.38.post_attention_layernorm.weight": "model-00007-of-00008.safetensors",
"model.layers.38.self_attn.k_norm.weight": "model-00007-of-00008.safetensors",
"model.layers.38.self_attn.k_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.38.self_attn.o_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.38.self_attn.q_norm.weight": "model-00007-of-00008.safetensors",
"model.layers.38.self_attn.q_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.38.self_attn.v_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.39.mlp.gate_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.39.self_attn.k_norm.weight": "model-00007-of-00008.safetensors",
"model.layers.39.self_attn.k_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.39.self_attn.o_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.39.self_attn.q_norm.weight": "model-00007-of-00008.safetensors",
"model.layers.39.self_attn.q_proj.weight": "model-00007-of-00008.safetensors",
"model.layers.39.self_attn.v_proj.weight": "model-00007-of-00008.safetensors",
"lm_head.weight": "model-00008-of-00008.safetensors",
"model.layers.39.input_layernorm.weight": "model-00008-of-00008.safetensors",
"model.layers.39.mlp.down_proj.weight": "model-00008-of-00008.safetensors",
"model.layers.39.mlp.up_proj.weight": "model-00008-of-00008.safetensors",
"model.layers.39.post_attention_layernorm.weight": "model-00008-of-00008.safetensors",
"model.norm.weight": "model-00008-of-00008.safetensors"
}
}

757480
tokenizer.json Normal file

File diff suppressed because it is too large Load Diff

239
tokenizer_config.json Normal file
View File

@@ -0,0 +1,239 @@
{
"add_bos_token": false,
"add_prefix_space": false,
"added_tokens_decoder": {
"151643": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151644": {
"content": "<|im_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151645": {
"content": "<|im_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151646": {
"content": "<|object_ref_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151647": {
"content": "<|object_ref_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151648": {
"content": "<|box_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151649": {
"content": "<|box_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151650": {
"content": "<|quad_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151651": {
"content": "<|quad_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151652": {
"content": "<|vision_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151653": {
"content": "<|vision_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151654": {
"content": "<|vision_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151655": {
"content": "<|image_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151656": {
"content": "<|video_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151657": {
"content": "<tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151658": {
"content": "</tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151659": {
"content": "<|fim_prefix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151660": {
"content": "<|fim_middle|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151661": {
"content": "<|fim_suffix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151662": {
"content": "<|fim_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151663": {
"content": "<|repo_name|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151664": {
"content": "<|file_sep|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151665": {
"content": "<tool_response>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151666": {
"content": "</tool_response>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151667": {
"content": "<think>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151668": {
"content": "</think>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
}
},
"additional_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"bos_token": null,
"chat_template": "{%- if tools %}\n {{- '<|im_start|>system\\n' }}\n {%- if messages[0].role == 'system' %}\n {{- messages[0].content + '\\n\\n' }}\n {%- endif %}\n {{- \"# Tools\\n\\nYou may call one or more functions to assist with the user query.\\n\\nYou are provided with function signatures within <tools></tools> XML tags:\\n<tools>\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n</tools>\\n\\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\\n<tool_call>\\n{\\\"name\\\": <function-name>, \\\"arguments\\\": <args-json-object>}\\n</tool_call><|im_end|>\\n\" }}\n{%- else %}\n {%- if messages[0].role == 'system' %}\n {{- '<|im_start|>system\\n' + messages[0].content + '<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}\n{%- for message in messages[::-1] %}\n {%- set index = (messages|length - 1) - loop.index0 %}\n {%- if ns.multi_step_tool and message.role == \"user\" and message.content is string and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}\n {%- set ns.multi_step_tool = false %}\n {%- set ns.last_query_index = index %}\n {%- endif %}\n{%- endfor %}\n{%- for message in messages %}\n {%- if message.content is string %}\n {%- set content = message.content %}\n {%- else %}\n {%- set content = '' %}\n {%- endif %}\n {%- if (message.role == \"user\") or (message.role == \"system\" and not loop.first) %}\n {{- '<|im_start|>' + message.role + '\\n' + content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {%- set reasoning_content = '' %}\n {%- if message.reasoning_content is string %}\n {%- set reasoning_content = message.reasoning_content %}\n {%- else %}\n {%- if '</think>' in content %}\n {%- set reasoning_content = content.split('</think>')[0].rstrip('\\n').split('<think>')[-1].lstrip('\\n') %}\n {%- set content = content.split('</think>')[-1].lstrip('\\n') %}\n {%- endif %}\n {%- endif %}\n {%- if loop.index0 > ns.last_query_index %}\n {%- if loop.last or (not loop.last and reasoning_content) %}\n {{- '<|im_start|>' + message.role + '\\n<think>\\n' + reasoning_content.strip('\\n') + '\\n</think>\\n\\n' + content.lstrip('\\n') }}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- if message.tool_calls %}\n {%- for tool_call in message.tool_calls %}\n {%- if (loop.first and content) or (not loop.first) %}\n {{- '\\n' }}\n {%- endif %}\n {%- if tool_call.function %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {{- '<tool_call>\\n{\"name\": \"' }}\n {{- tool_call.name }}\n {{- '\", \"arguments\": ' }}\n {%- if tool_call.arguments is string %}\n {{- tool_call.arguments }}\n {%- else %}\n {{- tool_call.arguments | tojson }}\n {%- endif %}\n {{- '}\\n</tool_call>' }}\n {%- endfor %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if loop.first or (messages[loop.index0 - 1].role != \"tool\") %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n<tool_response>\\n' }}\n {{- content }}\n {{- '\\n</tool_response>' }}\n {%- if loop.last or (messages[loop.index0 + 1].role != \"tool\") %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n {%- if enable_thinking is defined and enable_thinking is false %}\n {{- '<think>\\n\\n</think>\\n\\n' }}\n {%- endif %}\n{%- endif %}",
"clean_up_tokenization_spaces": false,
"eos_token": "<|im_end|>",
"errors": "replace",
"model_max_length": 131072,
"pad_token": "<|endoftext|>",
"split_special_tokens": false,
"tokenizer_class": "Qwen2Tokenizer",
"unk_token": null
}

1
vocab.json Normal file

File diff suppressed because one or more lines are too long