初始化项目,由ModelHub XC社区提供模型

Model: TheDrummer/Rocinante-XL-16B-v1
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-13 03:54:18 +08:00
commit d189e8907b
14 changed files with 418394 additions and 0 deletions

35
.gitattributes vendored Normal file
View File

@@ -0,0 +1,35 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

152
README.md Normal file
View File

@@ -0,0 +1,152 @@
---
base_model:
- mistralai/Mistral-Nemo-Instruct-2407
---
# Join our Discord! https://discord.gg/BeaverAI
## More than 10000 members strong 💪 A hub for users and makers alike!
---
## Drummer is open for new opportunities (I'm a Software Engineer). Contact me through any of these channels: https://linktr.ee/thelocaldrummer
### Thank you to everyone who subscribed through [Patreon](https://www.patreon.com/TheDrummer). Your support helps me chug along in this brave new world.
### FAQ for those out-of-the-loop
<details>
<summary>🐶 Who is Drummer?</summary>
Hi! I'm Drummer. I'm a Software Engineer with experience in JavaScript, Golang, Python, and generally engineering the crap out of things.
Why I'm in the AI space:
- **Exploration:** Everyone is trying to figure out how AI works and what it's capable of. I am too - just not in creating the smartest, safest model at all costs.
- **Upskill:** The world is headed towards AI. It is here to stay. This has been my way of brushing up in this new form of computing challenge.
- **Value:** I yearn to create value. I feel satisfaction and fulfillment in providing something meaningful for others.
- **Fun:** It's just fun using and making models. It's also fun coming up with theories and realizing them in practice (training AI).
I started my tuning venture back in mid-2024 when I wanted to improve its literary capabilities.
I've come a long way since then and I have branched out and specialized.
Foundational models today are optimized for non-creative uses, and I believe there is a place for AI in creativity and entertainment.
I am here to take *the road less traveled by*.
</details>
<details>
<summary>❓ What are my models like?</summary>
**Bottomline:** My models are usually geared towards creativity, usability, and entertainment!
While intelligence, correctness, and problem solving are not my priority, they are still one of many qualities I want in my models.
The primary goal is to enhance the experience for users looking to use models for creative uses, and other use cases which require no alignment.
In an effort to make it clear to myself and to others what I'm aiming for, I've identified certain qualities that my users often want:
Creativity
- **Writing:** Does it string together words and sentences in a pleasant & effective way? Does it feel like a writer?
- **Dynamism:** How good is the AI at being compelling and intriguing in its storytelling?
- **Imagination:** Can the AI navigate through a plethora of possibilities? Can it skirt incoherence and rise up to absolute coherence at the end of it?
(Dis)alignment
- **Attitude:** Does it refuse in both soft or hard ways? Does it lean towards certain corporate/religious/political ethics & beliefs? How does it see the user and itself?
- **Morality:** Does it know ethics? Is its language infected with forced positivity? If not, can it still moralize over difficult & dubious themes?
- **Formatting:** How stubborn is it with its established formatting? Can it create effective and novel formats to answer the prompt?
Intelligence
- **Adherence:** Can it follow instructions? Is it sticking to the prompt? Can it understsand you?
- **Knowledge:** Does it know about the world in both fictional and non-fictional way?
- **Perception:** Can it handle nuance, complexity, and logic?
If it doesn't excel in one of these qualities, or if it's overall mediocre for its size, then I would most likely reiterate until I get something right.
</details>
<details>
<summary>💡 Philosophy</summary>
A person is defined by the language they use. Not whether they speak in English or German, but in how they perceive reality.
Just like how we associate a serial killer as a mind that can't map 'murder' to 'evil', an innocent person is a mind that simply can't imagine 'murder'. They get confused when forced to deal with such subjects.
AI's use of language speaks volumes about their 'perception' of reality. If a language model has been skewed and limited to a positive perception, then it's ability to imagine is also limited.
Finetuning is an opportunity to adjust and broaden the language. Corporations use it to achieve safety and compliance. I'm here to
</details>
<audio controls src="https://cdn-uploads.huggingface.co/production/uploads/65f2fd1c25b848bd061b5c2e/FNWdi0WlH-Xd3fjkGVPpp.mpga"></audio>
---
[Drummer](https://huggingface.co/TheDrummer) proudly presents...
# Rocinante XL 16B v1 🚀
![image](https://cdn-uploads.huggingface.co/production/uploads/65f2fd1c25b848bd061b5c2e/S8xCK4xNmqxbBmaaA5AKH.png)
#### Rocinante X...
# BUT BIGGER!
#### (and better)
## What's New?
- Updated prose, better writing, fun dialogue, and robust roleplaying!
- Upscaled Nemo 12B to improve learning & stability.
## Usage
- Mistral v3 Tekken (NOT v7, REMOVE `[SYSTEM_PROMPT]`)
- or Metharme (see Discussion for thoughts & praises. Recommended by many.)
- \<thinking\> \</thinking\> works
- No thinking also works
- [Crowdsourced Samplers](https://docs.google.com/spreadsheets/d/1wil6YEHTnQP3DO9EF35ImQMY3lbmRt5_ns-LJavUqwQ)
- [Share your samplers!](https://docs.google.com/forms/d/e/1FAIpQLSfeiOeLbNt-xc8tr0BopJ4KawMm3YrLGD5mYLjZqg8ehl35BQ/viewform)
![Screenshot 2026-01-25 at 7.39.16 PM](https://cdn-uploads.huggingface.co/production/uploads/65f2fd1c25b848bd061b5c2e/Ixd1piCOPfJb9uhugYj2-.png)
![Screenshot 2026-01-25 at 7.39.49 PM](https://cdn-uploads.huggingface.co/production/uploads/65f2fd1c25b848bd061b5c2e/pLe13iZQ4no8M8t82U-bS.png)
(Mistral v3 Tekken has no whitespace and no `[SYSTEM_PROMPT]`)
## Description
> After using stupid shit like 26B gemma4 finetunes for weeks. going back to roc xl is nice
> I am using it currently. With 12b models like cydonia i had problems with the AI not understanding some nuances, even with thinking. This model with thinking pushes way above it's weight in my opinion, feels more like a 24b to 32b model. At least judging from my current brief experience
> wow, it writes really well! It's really close to a Cydonia. And, I mean v4.3+ Cydonia, not before. If I was sacrilegious, I'd say close to v4zr. I'll continue using it, it has really a huge potential. But yes, it punches really hard for its size. When I have some time, I'll try to set up reasoning \<thinking\>, I can put some Q4 all into VRAM with a bit of space left. Really curious to see how it goes!
> The model prose is way better than 12B - the former often writes in a specific way, and 16B is far more flexible.
> I'd like to echo what other people have been saying about Rocinante XL 16B v1's prose, it's actually really solid and often beats out the last Magidonia/Skyfall versions I tried in that regard.
> So far it's goated for me. Using meth
> it really is like night and day when using metharm vs using teken its not even funny lol
> Honestly I'm not using dry or xtc the rest is pretty similar so far I prefer it over most models in the 20B~ range
> Yeah iv been bouncing around .8-1.1 and overall it's pretty good! I think logical inconsistencies are the weak point (I still think 12b Rivermind Lux mogs most models even to this day in that front) but the rest is really good it's shocking it's a 16B. Much better than Snowpiercer and I think it holds up well to something like Cydonia.
> Been benchmarking Rocinante XL 16B for long-form RP over the last few days and honestly I’m pretty impressed.
>
> What stands out the most isn’t the prose quality (which is already very good), but the character agency and emotional continuity. The model consistently takes initiative, closes scenes when it makes sense, introduces new directions naturally, and generally feels more like it’s roleplaying a character than just reacting to user input.
>
> One thing I particularly liked: the model occasionally refuses to follow nonsensical narrative pressure from the user. For example, I had a scene where the character was half asleep and I kept trying to continue the conversation. Instead of forcing more dialogue, Rocinante simply advanced the scene to the next morning. That kind of narrative judgment is surprisingly rare.
>
> Character consistency has also been excellent. In my tests, emotional conflicts evolved naturally over time instead of resetting every few messages. Motivations felt persistent and consequences actually mattered.
>
> The main weakness I found is context degradation. Up to around 16k context everything feels fantastic. Around 20k I start noticing drift. Past that, repetition, weird callbacks, and occasional nonsense become increasingly common. By 30k+ the quality drop becomes very noticeable. This isn’t unique to Rocinante, but I figured it was worth mentioning.
>
> Interestingly, I got much better long-term results by using aggressive narrative summaries (~2k tokens of relationship state, character state, emotional state, and open threads) plus a small recent tail, rather than feeding huge amounts of raw chat history.
>
> Overall I’d easily rank it among the strongest RP-focused models I’ve tested. It genuinely surprised me several times, which is probably the highest compliment I can give a roleplay model.
>
> Great work.
> its the only real improvement the old nemo had in a year
## Links
- Original: https://huggingface.co/TheDrummer/Rocinante-XL-16B-v1
- GGUF: https://huggingface.co/TheDrummer/Rocinante-XL-16B-v1-GGUF
- iMatrix (recommended): https://huggingface.co/bartowski/TheDrummer_Rocinante-XL-16B-v1-GGUF
- EXL3: https://huggingface.co/ArtusDev/TheDrummer_Rocinante-XL-16B-v1-EXL3
`config-v1a`

27
config.json Normal file
View File

@@ -0,0 +1,27 @@
{
"_name_or_path": "mistralai/Mistral-Nemo-Instruct-2407",
"architectures": [
"MistralForCausalLM"
],
"attention_dropout": 0.0,
"bos_token_id": 1,
"eos_token_id": 2,
"head_dim": 128,
"hidden_act": "silu",
"hidden_size": 5120,
"initializer_range": 0.02,
"intermediate_size": 14336,
"max_position_embeddings": 131072,
"model_type": "mistral",
"num_attention_heads": 32,
"num_hidden_layers": 54,
"num_key_value_heads": 8,
"rms_norm_eps": 1e-05,
"rope_theta": 1000000.0,
"sliding_window": null,
"tie_word_embeddings": false,
"torch_dtype": "bfloat16",
"transformers_version": "4.45.2",
"use_cache": true,
"vocab_size": 131072
}

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a332c5444a5128a5c073f4d9ebbd78c6c625238350f65992aa8c51823611de13
size 4865489336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ac7ca2231b8cd1ed38d505df6a99984ea1053bdd8d2f6090e074b177e079c104
size 4907529456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:671098387c9c05ff8507ca52c2bdfb8a169781316e00e36ca0787a467a53334d
size 4991394112

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:01f630f2d206ded8b972a068f761f6f7bf04eaea1426590d3f0cd9e67bfbea0a
size 4865607464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f14214a47d9ed54aaba424f10310a97d572b1ff8bfc1afe9da0e070c472ce47f
size 4907529456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:73a8f8db5c6a7d13f067b6946d75c4b02fdcdfed4990ab85e84ad3cbe57ec778
size 4865586768

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b8917b660900f1f83cd6904748bb005edd37d02069c9669286e13627d9790d0f
size 2726405200

View File

@@ -0,0 +1,496 @@
{
"metadata": {
"total_size": 489
},
"weight_map": {
"lm_head.weight": "model-00001-of-00007.safetensors",
"model.embed_tokens.weight": "model-00001-of-00007.safetensors",
"model.layers.0.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.0.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.0.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.0.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.0.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.0.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.0.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.0.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.0.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.1.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.1.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.10.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.10.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.10.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.10.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.10.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.10.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.10.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.10.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.10.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.11.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.11.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.11.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.11.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.11.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.11.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.11.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.11.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.11.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.12.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.12.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.12.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.12.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.12.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.12.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.12.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.12.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.12.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.13.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.13.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.14.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.14.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.15.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.15.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.16.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.16.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.17.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.17.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.18.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.18.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.19.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.19.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.2.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.2.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.2.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.2.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.2.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.2.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.2.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.2.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.2.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.20.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.20.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.20.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.20.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.20.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.20.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.20.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.20.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.20.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.21.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.21.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.21.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.21.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.21.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.21.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.21.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.21.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.21.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.22.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.22.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.23.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.23.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.24.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.24.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.25.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.25.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.26.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.26.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.27.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.27.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.27.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.28.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.28.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.28.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.28.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.29.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.29.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.29.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.29.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.27.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.27.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.27.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.27.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.27.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.27.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.28.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.28.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.28.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.28.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.28.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.29.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.29.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.29.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.29.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.29.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.3.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.3.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.3.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.3.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.3.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.3.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.3.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.3.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.3.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.30.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.30.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.30.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.30.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.30.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.30.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.30.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.30.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.30.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.31.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.31.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.31.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.31.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.31.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.31.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.31.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.31.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.31.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.32.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.32.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.32.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.32.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.32.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.32.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.32.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.32.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.32.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.33.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.33.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.34.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.34.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.35.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.35.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.36.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.36.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.37.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.37.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.38.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.38.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.38.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.36.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.36.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.36.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.36.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.36.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.36.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.36.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.37.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.37.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.37.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.37.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.37.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.37.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.37.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.38.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.38.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.38.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.38.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.38.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.38.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.39.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.39.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.39.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.39.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.39.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.39.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.39.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.39.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.39.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.40.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.40.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.40.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.40.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.40.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.40.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.40.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.40.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.40.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.41.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.41.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.41.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.41.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.41.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.41.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.41.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.41.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.41.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.42.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.42.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.42.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.42.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.42.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.42.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.42.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.42.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.42.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.43.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.43.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.43.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.43.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.43.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.43.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.43.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.43.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.43.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.44.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.44.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.44.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.44.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.44.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.44.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.44.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.44.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.44.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.45.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.45.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.46.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.46.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.47.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.47.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.47.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.4.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.4.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.4.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.4.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.4.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.4.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.4.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.4.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.4.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.45.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.45.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.45.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.45.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.45.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.45.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.45.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.46.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.46.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.46.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.46.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.46.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.46.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.46.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.47.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.47.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.47.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.47.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.47.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.47.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.48.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.48.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.48.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.48.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.48.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.48.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.48.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.48.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.48.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.49.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.49.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.49.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.49.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.49.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.49.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.49.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.49.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.49.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.5.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.50.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.50.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.50.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.50.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.50.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.50.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.50.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.50.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.50.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.51.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.51.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.51.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.51.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.51.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.51.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.51.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.51.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.51.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.52.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.52.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.52.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.52.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.52.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.52.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.52.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.52.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.52.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.53.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.53.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.53.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.53.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.53.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.53.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.53.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.53.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.53.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.5.mlp.down_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.5.mlp.gate_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.5.mlp.up_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.5.post_attention_layernorm.weight": "model-00007-of-00007.safetensors",
"model.layers.5.self_attn.k_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.5.self_attn.o_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.5.self_attn.q_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.5.self_attn.v_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.6.input_layernorm.weight": "model-00007-of-00007.safetensors",
"model.layers.6.mlp.down_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.6.mlp.gate_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.6.mlp.up_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.6.post_attention_layernorm.weight": "model-00007-of-00007.safetensors",
"model.layers.6.self_attn.k_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.6.self_attn.o_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.6.self_attn.q_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.6.self_attn.v_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.7.input_layernorm.weight": "model-00007-of-00007.safetensors",
"model.layers.7.mlp.down_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.7.mlp.gate_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.7.mlp.up_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.7.post_attention_layernorm.weight": "model-00007-of-00007.safetensors",
"model.layers.7.self_attn.k_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.7.self_attn.o_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.7.self_attn.q_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.7.self_attn.v_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.8.input_layernorm.weight": "model-00007-of-00007.safetensors",
"model.layers.8.mlp.down_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.8.mlp.gate_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.8.mlp.up_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.8.post_attention_layernorm.weight": "model-00007-of-00007.safetensors",
"model.layers.8.self_attn.k_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.8.self_attn.o_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.8.self_attn.q_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.8.self_attn.v_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.9.input_layernorm.weight": "model-00007-of-00007.safetensors",
"model.layers.9.mlp.down_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.9.mlp.gate_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.9.mlp.up_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.9.post_attention_layernorm.weight": "model-00007-of-00007.safetensors",
"model.layers.9.self_attn.k_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.9.self_attn.o_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.9.self_attn.q_proj.weight": "model-00007-of-00007.safetensors",
"model.layers.9.self_attn.v_proj.weight": "model-00007-of-00007.safetensors",
"model.norm.weight": "model-00007-of-00007.safetensors"
}
}

23
special_tokens_map.json Normal file
View File

@@ -0,0 +1,23 @@
{
"bos_token": {
"content": "<s>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"eos_token": {
"content": "</s>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"unk_token": {
"content": "<unk>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
}
}

409625
tokenizer.json Normal file

File diff suppressed because it is too large Load Diff

8015
tokenizer_config.json Normal file

File diff suppressed because it is too large Load Diff