初始化项目,由ModelHub XC社区提供模型

Model: Steelskull/L3-Aethora-15B
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-04-22 11:30:59 +08:00
commit 65e48e28ea
16 changed files with 413423 additions and 0 deletions

42
.gitattributes vendored Normal file
View File

@@ -0,0 +1,42 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
model-00001-of-00007.safetensors filter=lfs diff=lfs merge=lfs -text
model-00002-of-00007.safetensors filter=lfs diff=lfs merge=lfs -text
model-00003-of-00007.safetensors filter=lfs diff=lfs merge=lfs -text
model-00004-of-00007.safetensors filter=lfs diff=lfs merge=lfs -text
model-00005-of-00007.safetensors filter=lfs diff=lfs merge=lfs -text
model-00006-of-00007.safetensors filter=lfs diff=lfs merge=lfs -text
model-00007-of-00007.safetensors filter=lfs diff=lfs merge=lfs -text

146
README.md Normal file
View File

@@ -0,0 +1,146 @@
---
library_name: transformers
tags:
- llama-factory
license: llama3
datasets:
- TheSkullery/Aether-Lite-V1.2
---
<!DOCTYPE html>
<style>
body, html {
height: 100%; /* Ensure the full height of the page is used */
margin: 0;
padding: 0;
font-family: 'Quicksand', sans-serif;
background: linear-gradient(135deg, #2E3440 0%, #1A202C 100%);
color: #D8DEE9;
font-size: 16px;
}
.container {
width: 100%; /* Full width */
height: 100%; /* Full height */
padding: 20px;
margin: 0; /* Remove margin to fill the entire area */
background-color: rgba(255, 255, 255, 0.02);
border-radius: 12px;
box-shadow: 0 4px 10px rgba(0, 0, 0, 0.2);
backdrop-filter: blur(10px);
border: 1px solid rgba(255, 255, 255, 0.1);
}
.header h1 {
font-size: 28px;
color: #5F9EA0;
margin: 0 0 20px 0;
text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.3);
}
.update-section h2 {
font-size: 24px;
color: #88C0D0;
}
.update-section p {
font-size: 16px;
line-height: 1.6;
color: #ECEFF4;
}
.info img {
width: 100%;
border-radius: 10px;
margin-bottom: 15px;
}
a {
color: #88C0D0;
text-decoration: none;
}
a:hover {
color: #A3BE8C;
}
.button {
display: inline-block;
background-color: #5E81AC;
color: #E5E9F0;
padding: 10px 20px;
border-radius: 5px;
cursor: pointer;
text-decoration: none;
}
.button:hover {
background-color: #81A1C1;
}
pre {
background-color: #2E3440;
padding: 10px;
border-radius: 5px;
overflow-x: auto;
}
code {
font-family: 'Courier New', monospace;
color: #D8DEE9;
}
</style>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>L3-Aethora-15B Data Card</title>
<link href="https://fonts.googleapis.com/css2?family=Quicksand:wght@400;500;600&display=swap" rel="stylesheet">
</head>
<body>
<div class="container">
<div class="header">
<h1>L3-Aethora-15B</h1>
</div>
<div class="info">
<img src="https://cdn-uploads.huggingface.co/production/uploads/64545af5ec40bbbd01242ca6/W0qzZK_V1Zt1GdgCIsnrP.png">
<p>The Skullery Presents L3-Aethora-15B.</p>
<p><strong>Creator:</strong> <a href="https://huggingface.co/steelskull" target="_blank">Steelskull</a></p>
<p><strong>Dataset:</strong> <a href="https://huggingface.co/datasets/TheSkullery/Aether-Lite-V1.2" target="_blank">Aether-Lite-V1.2</a></p>
<p><strong>Trained:</strong> 4 x A100 for 15 hours Using RsLora and DORA</p>
<h1>About L3-Aethora-15B:</h1>
<pre><code> L3 = Llama3 </code></pre>
<p>L3-Aethora-15B was crafted through using the abilteration method to adjust model responses. The model's refusal is inhibited, focusing on yielding more compliant and facilitative dialogue interactions. It then underwent a modified DUS (Depth Up Scale) merge (originally used by @Elinas) by using passthrough merge to create a 15b model, with specific adjustments (zeroing) to 'o_proj' and 'down_proj', enhancing its efficiency and reducing perplexity. This created AbL3In-15b.<br>
<p>AbL3In-15b was then trained for 4 epochs using Rslora & DORA training methods on the Aether-Lite-V1.2 dataset, containing ~82000 high quality samples, designed to strike a fine balance between creativity, slop, and intelligence at about a 60/40 split</p>
<p>This model is trained on the L3 prompt format.</p>
<h2>Quants:</h2>
<li><a href="https://huggingface.co/mradermacher/L3-Aethora-15B-GGUF" target="_blank">Mradermacher/L3-Aethora-15B-GGUF</a></li>
<li><a href="https://huggingface.co/mradermacher/L3-Aethora-15B-i1-GGUF" target="_blank">Mradermacher/L3-Aethora-15B-i1-GGUF</a></li>
<li><a href="https://huggingface.co/NikolayKozloff" target="_blank">NikolayKozloff/L3-Aethora-15B-GGUF</a></li>
<p></p>
<h2>Dataset Summary: (Filtered)</h2>
<p>Filtered Phrases: GPTslop, Claudism's</p>
<ul>
<li><strong>mrfakename/Pure-Dove-ShareGPT:</strong> Processed 3707, Removed 150</li>
<li><strong>mrfakename/Capybara-ShareGPT:</strong> Processed 13412, Removed 2594</li>
<li><strong>jondurbin/airoboros-3.2:</strong> Processed 54517, Removed 4192</li>
<li><strong>PJMixers/grimulkan_theory-of-mind-ShareGPT:</strong> Processed 533, Removed 6</li>
<li><strong>grimulkan/PIPPA-augmented-dedup:</strong> Processed 869, Removed 46</li>
<li><strong>grimulkan/LimaRP-augmented:</strong> Processed 790, Removed 14</li>
<li><strong>PJMixers/grimulkan_physical-reasoning-ShareGPT:</strong> Processed 895, Removed 4</li>
<li><strong>MinervaAI/Aesir-Preview:</strong> Processed 994, Removed 6</li>
<li><strong>Doctor-Shotgun/no-robots-sharegpt:</strong> Processed 9911, Removed 89</li>
</ul>
<h2>Deduplication Stats:</h2>
<p>Starting row count: 85628, Final row count: 81960, Rows removed: 3668</p>
<p><strong>I've had a few people ask about donations so here's a link:</strong</p>
</div>
<div class="donation-section">
<a href="https://ko-fi.com/Y8Y0AO2XE" target="_blank">
<img height="36" style="border:0px;height:36px;" src="https://storage.ko-fi.com/cdn/kofi2.png?v=3" border="0" alt="Buy Me a Coffee at ko-fi.com" />
</a>
</div>
</div>
</div>
</body>
</html>

29
config.json Normal file
View File

@@ -0,0 +1,29 @@
{
"_name_or_path": "TheSkullery/AbL3In-15B",
"architectures": [
"LlamaForCausalLM"
],
"attention_bias": false,
"attention_dropout": 0.0,
"bos_token_id": 128000,
"eos_token_id": 128009,
"hidden_act": "silu",
"hidden_size": 4096,
"initializer_range": 0.02,
"intermediate_size": 14336,
"max_position_embeddings": 8192,
"mlp_bias": false,
"model_type": "llama",
"num_attention_heads": 32,
"num_hidden_layers": 64,
"num_key_value_heads": 8,
"pretraining_tp": 1,
"rms_norm_eps": 1e-05,
"rope_scaling": null,
"rope_theta": 500000.0,
"tie_word_embeddings": false,
"torch_dtype": "bfloat16",
"transformers_version": "4.41.2",
"use_cache": true,
"vocab_size": 128256
}

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}

6
generation_config.json Normal file
View File

@@ -0,0 +1,6 @@
{
"_from_model_config": true,
"bos_token_id": 128000,
"eos_token_id": 128009,
"transformers_version": "4.41.2"
}

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bfa172f1b231f1592bca7740228fdc42c42302d422125f3c2c7f902d0ee82ef7
size 4976698672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:26a92bcad582f052b4a54a6f14ba9c7435d5114b67413efa8363d9013df3f4bd
size 4999802720

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6e6883aa12dc99a5c76962da780375daa9099664d43601cd57aa927530b3f6d0
size 4915916176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5bd0b13d14e41a01f61a7a3bd023b4557e1a1b7d54f6b396b001ea31afaa3801
size 4999819336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1dd3b9305b8a5da1b9eaf221ee02f281f5b598b45fb6f6663ddbff3ff94adfdd
size 4915916184

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ee24163e80d2762bc1aa33e2bb3bc611f2178b591801d8dd9eae41de2636b4ad
size 4160931608

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:674610adfedb7f19d21475a7280d6b34774894e654f1d50f4c24444ba7bd0d04
size 1050673280

View File

@@ -0,0 +1,586 @@
{
"metadata": {
"total_size": 30019690496
},
"weight_map": {
"lm_head.weight": "model-00007-of-00007.safetensors",
"model.embed_tokens.weight": "model-00001-of-00007.safetensors",
"model.layers.0.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.0.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.0.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.0.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.0.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.0.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.0.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.0.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.0.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.1.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.1.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.1.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.10.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.10.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.10.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.10.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.10.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.10.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.10.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.10.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.10.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.11.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.11.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.11.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.11.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.11.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.11.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.11.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.11.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.11.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.12.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.12.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.12.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.12.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.12.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.12.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.12.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.12.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.12.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.13.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.13.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.13.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.14.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.14.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.14.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.15.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.15.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.15.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.16.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.16.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.16.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.17.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.17.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.17.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.18.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.18.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.18.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.19.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.19.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.19.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.2.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.2.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.2.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.2.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.2.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.2.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.2.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.2.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.2.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.20.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.20.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.20.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.20.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.20.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.20.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.20.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.20.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.20.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.21.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.21.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.21.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.21.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.21.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.21.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.21.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.21.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.21.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.22.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.22.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.22.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.23.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.23.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.23.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.24.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.24.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.24.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.25.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.25.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.25.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.26.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.26.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.26.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.27.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.27.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.27.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.27.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.27.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.27.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.27.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.27.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.27.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.28.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.28.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.28.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.28.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.28.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.28.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.28.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.28.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.28.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.29.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.29.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.29.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.29.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.29.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.29.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.29.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.29.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.29.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.3.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.3.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.3.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.3.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.3.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.3.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.3.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.3.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.3.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.30.input_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.30.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.30.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.30.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.30.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
"model.layers.30.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.30.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.30.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.30.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.31.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.31.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.31.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.31.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.31.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.31.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.31.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.31.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.31.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
"model.layers.32.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.32.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.32.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.32.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.32.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.32.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.32.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.32.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.32.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.33.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.33.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.33.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.34.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.34.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.34.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.35.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.35.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.35.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.36.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.36.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.36.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.36.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.36.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.36.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.36.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.36.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.36.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.37.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.37.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.37.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.37.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.37.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.37.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.37.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.37.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.37.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.38.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.38.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.38.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.38.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.38.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.38.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.38.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.38.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.38.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.39.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.39.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.39.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.39.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.39.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.39.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.39.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.39.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.39.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.4.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.4.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.4.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.4.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.4.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.4.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.4.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.4.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.4.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.40.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.40.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.40.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.40.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.40.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.40.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.40.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.40.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.40.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.41.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.41.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.41.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.41.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.41.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.41.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.41.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.41.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.41.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.42.input_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.42.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.42.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.42.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.42.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
"model.layers.42.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.42.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.42.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.42.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.43.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.43.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.43.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.43.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.43.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.43.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.43.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.43.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.43.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
"model.layers.44.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.44.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.44.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.44.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.44.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.44.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.44.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.44.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.44.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.45.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.45.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.45.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.45.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.45.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.45.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.45.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.45.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.45.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.46.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.46.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.46.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.46.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.46.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.46.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.46.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.46.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.46.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.47.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.47.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.47.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.47.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.47.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.47.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.47.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.47.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.47.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.48.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.48.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.48.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.48.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.48.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.48.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.48.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.48.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.48.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.49.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.49.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.49.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.49.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.49.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.49.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.49.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.49.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.49.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.5.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.5.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.5.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.5.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.5.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.5.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.5.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.5.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.5.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.50.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.50.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.50.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.50.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.50.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.50.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.50.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.50.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.50.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.51.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.51.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.51.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.51.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.51.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.51.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.51.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.51.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.51.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.52.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.52.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.52.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.52.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.52.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.52.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.52.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.52.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.52.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.53.input_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.53.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.53.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.53.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.53.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
"model.layers.53.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.53.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.53.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.53.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.54.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.54.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.54.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.54.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.54.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.54.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.54.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.54.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.54.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
"model.layers.55.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.55.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.55.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.55.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.55.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.55.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.55.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.55.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.55.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.56.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.56.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.56.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.56.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.56.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.56.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.56.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.56.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.56.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.57.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.57.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.57.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.57.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.57.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.57.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.57.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.57.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.57.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.58.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.58.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.58.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.58.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.58.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.58.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.58.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.58.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.58.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.59.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.59.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.59.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.59.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.59.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.59.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.59.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.59.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.59.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.6.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.6.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.6.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.6.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.6.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.6.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.6.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.6.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.6.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.60.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.60.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.60.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.60.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.60.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.60.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.60.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.60.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.60.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.61.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.61.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.61.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.61.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.61.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.61.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.61.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.61.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.61.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.62.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.62.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.62.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.62.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.62.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.62.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.62.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.62.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.62.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.63.input_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.63.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.63.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.63.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.63.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
"model.layers.63.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.63.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.63.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.63.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
"model.layers.7.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.7.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.7.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.7.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.7.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.7.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.7.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.7.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.7.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.8.input_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.8.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.8.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.8.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.8.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
"model.layers.8.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.8.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.8.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.8.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
"model.layers.9.input_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.9.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.9.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.9.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.9.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
"model.layers.9.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.9.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.9.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
"model.layers.9.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
"model.norm.weight": "model-00006-of-00007.safetensors"
}
}

23
special_tokens_map.json Normal file
View File

@@ -0,0 +1,23 @@
{
"bos_token": {
"content": "<|begin_of_text|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"eos_token": {
"content": "<|eot_id|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"pad_token": {
"content": "<|end_of_text|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
}
}

410504
tokenizer.json Normal file

File diff suppressed because it is too large Load Diff

2065
tokenizer_config.json Normal file

File diff suppressed because it is too large Load Diff