初始化项目,由ModelHub XC社区提供模型
Model: seedboxai/KafkaLM-15B Source: Original Platform
This commit is contained in:
36
.gitattributes
vendored
Normal file
36
.gitattributes
vendored
Normal file
@@ -0,0 +1,36 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
||||
246
README.md
Normal file
246
README.md
Normal file
@@ -0,0 +1,246 @@
|
||||
---
|
||||
library_name: transformers
|
||||
tags:
|
||||
- pruning
|
||||
- distillation
|
||||
- sparsity‑2:4
|
||||
license: apache-2.0
|
||||
language:
|
||||
- en
|
||||
- de
|
||||
- fr
|
||||
- es
|
||||
- it
|
||||
- pt
|
||||
base_model:
|
||||
- doubledsbv/KafkaLM-15B-Base
|
||||
pipeline_tag: text-generation
|
||||
---
|
||||
|
||||
<img src="https://cdn-uploads.huggingface.co/production/uploads/645ded34a45b4182d7f5c385/EgsjPDWd37LjAtamiICxk.png" width="480" height="480" alt="image/png">
|
||||
|
||||
|
||||
# Model Description
|
||||
|
||||
**KafkaLM‑15B‑Base** is a 15‑billion‑parameter, sparsity‑aware language model distilled from *Mistral‑Small‑24B‑Base‑2501* and further post trained (SFT + DPO + GRPO /w verifiable rewards).
|
||||
|
||||
This experimental model was created in five stages:
|
||||
|
||||
| Stage | What we did | Why it matters |
|
||||
|-------|-------------|----------------|
|
||||
| **1. SimplePrune** | Applied a hierarchical, hardware‑aware pruning pipeline that combines block‑, channel‑ and 2:4 structured sparsity (≈ 37.5 % parameter reduction) | Slashes memory footprint while minimizing perplexity degradation |
|
||||
| **2. Teacher calibration** | Briefly fine‑tuned the unpruned 24 B teacher on a 10 B‑token multilingual European corpus on a AMD M300A cluster | Produces stable logits and hidden states for distillation |
|
||||
| **3. Knowledge distillation** | Distilled the calibrated teacher into the pruned 15 B student using a **fused loss**:<br/>`L Pooled SquareHead + LKL + 0.25 * LCE` | Transfers teacher capabiities effectively with <15B tokens **(< 2 epochs)** on 64 MI300A nodes |
|
||||
| **4. SFT+DPO** | Supervised finetuning + Direct Preference Optimization) on curated open-source multilingual and multitask datasets | Enhances model alignment with human preferences while preserving multilingual capabilities |
|
||||
| **5. RL** | Trained GRPO as separate LoRA adapter to make it easy for serving and optional for using | Enables flexible deployment with optional reinforcement learning benefits without modifying the base model |
|
||||
|
||||
**Key capabilities**
|
||||
|
||||
* Balanced for both **multitask** and multilingual conversation and long context handling
|
||||
* Structured **2:4 sparsity** → runs up to **40 % faster** on sparsity‑aware kernels
|
||||
* Distilled on a combination of multilingual pretraining and synthetic data
|
||||
* Training pipeline optimized for unified‑memory GPUs (AMD MI300A) but runs on any CUDA / ROCm device
|
||||
|
||||
---
|
||||
|
||||
### LoRA based reasoning capabilities
|
||||
|
||||
|
||||
[Download adapter from hf](https://huggingface.co/seedboxai/KafkaLM-15B-GRPO_LoRA_Exp)
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
from peft import PeftModel
|
||||
|
||||
# Load base model
|
||||
base_model = AutoModelForCausalLM.from_pretrained("seedboxai/KafkaLM-15B")
|
||||
tokenizer = AutoTokenizer.from_pretrained("seedboxai/KafkaLM-15B")
|
||||
|
||||
# Apply LoRA adapter
|
||||
model = PeftModel.from_pretrained(
|
||||
base_model,
|
||||
"seedboxai/KafkaLM-15B-GRPO_LoRA_Exp",
|
||||
adapter_name="grpo_lora"
|
||||
)
|
||||
```
|
||||
|
||||
## Pruning Process
|
||||
|
||||
**Pruning & Distillation Strategy — SimplePrune**
|
||||
Hardware‑aware, hierarchical pipeline. SimplePrune starts with coarse block‑level pruning and drills down to channel‑ and neuron‑level removals, finishing with 2 : 4 structured sparsity. This staged approach converts compression ratios into real memory‑bandwidth and latency gains.
|
||||
|
||||
**Sensitivity‑guided selection**
|
||||
Each stage is driven by activation‑magnitude profiles and Hessian‑based importance scores captured asynchronously during training, allowing the framework to run inside the MI300A’s 512 GB unified memory without OOM interruptions.
|
||||
|
||||
**Two‑phase optimisation**
|
||||
A fast greedy pass prunes low‑impact blocks in MLP expansion layers, after which a **Tabu‑Search** meta‑heuristic explores cross‑layer combinations for a better global trade‑off between sparsity and perplexity/KL divergence.
|
||||
|
||||
**Post‑pruning knowledge distillation**
|
||||
The pruned 15 B student is distilled from a calibrated 24 B teacher using a fused LSquareHead + KL + 0.25 · CE loss across 20 B multilingual tokens, restoring > 96 % of the original quality in ≤ 2 epochs on up to 64 MI300A nodes.
|
||||
|
||||
### Results
|
||||
Up to 40 % parameter reduction (24 B → 15 B) delivers 2× lower TTFT and ≈ 40 % higher tokens/s versus the uncompressed teacher while matching perplexity and divergence metrics—validating SimplePrune as an effective route to deploy KafkaLM in memory‑constrained, sparsity‑accelerated environments.
|
||||
|
||||
| Metric | Mistral‑24B | **KafkaLM‑15B** | Δ |
|
||||
|--------|-------------|-----------------|---|
|
||||
| Time‑to‑First‑Token | 4.91 s | **2.46 s** | −50% |
|
||||
| Prompts / s | 4.70 | **6.55** | +38% |
|
||||
| Tokens / s | 579 | **812** | +40% |
|
||||
|
||||
|
||||
<img src="https://cdn-uploads.huggingface.co/production/uploads/645ded34a45b4182d7f5c385/4rDhaeC-1GMj6KWbB27f9.png" width="480" height="480" alt="image/png">
|
||||
|
||||
|
||||
### Training scalability (distillation run, MI300A cluster)
|
||||
|
||||
| Nodes | Tokens / s | Speed‑up |
|
||||
|-------|------------|----------|
|
||||
| 4 | 1 461 | – |
|
||||
| 8 | 3 327 | 2.3 × |
|
||||
| 16 | 7 423 | 5.1 × |
|
||||
| 32 | 15 286 | 10.5 × |
|
||||
| 64 | 25 455 | 17.4 × |
|
||||
|
||||
Near‑linear scaling thanks to sharded ZeRO‑3 + RCCL optimisations.
|
||||
|
||||
# Inference
|
||||
|
||||
|
||||
### Transformers
|
||||
|
||||
```python
|
||||
model_name = "seedboxai/KafkaLM-15B"
|
||||
|
||||
# load the tokenizer and the model
|
||||
tokenizer = AutoTokenizer.from_pretrained(model_name)
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
model_name,
|
||||
torch_dtype="auto",
|
||||
device_map="auto"
|
||||
)
|
||||
|
||||
# prepare the model input
|
||||
prompt = "Why did Kafka hit different?"
|
||||
messages = [
|
||||
{"role": "user", "content": prompt}
|
||||
]
|
||||
text = tokenizer.apply_chat_template(
|
||||
messages,
|
||||
tokenize=False,
|
||||
add_generation_prompt=True,
|
||||
)
|
||||
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
|
||||
|
||||
# conduct text completion
|
||||
generated_ids = model.generate(
|
||||
**model_inputs,
|
||||
max_new_tokens=1024
|
||||
)
|
||||
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist()
|
||||
|
||||
response = tokenizer.decode(output_ids, skip_special_tokens=True)
|
||||
|
||||
print(response)
|
||||
```
|
||||
|
||||
## vLLM
|
||||
|
||||
```python
|
||||
|
||||
"""
|
||||
This example shows how to use KafkaLM-15B in vLLM with and without LoRA reasoning functionality
|
||||
for offline inference.
|
||||
"""
|
||||
|
||||
from typing import Optional
|
||||
from huggingface_hub import snapshot_download
|
||||
from vllm import EngineArgs, LLMEngine, RequestOutput, SamplingParams
|
||||
from vllm.lora.request import LoRARequest
|
||||
import transformers
|
||||
|
||||
SYS_MESSAGE = 'A conversation between User and Assistant. The user asks a question, and the Assistant solves it. The assistant first thinks about the reasoning process in the mind and then provides the user with the answer. The reasoning process and answer are enclosed within <think> </think> and <answer> </answer> tags, respectively, i.e., <think> reasoning process here </think> <answer> answer here </answer>.'
|
||||
tokenizer = transformers.AutoTokenizer.from_pretrained("")
|
||||
|
||||
def create_test_prompts(lora_path: str) -> list[tuple[str, SamplingParams, Optional[LoRARequest]]]:
|
||||
"""Create a list of test prompts with their sampling parameters.
|
||||
1 requests for base model, 1 request for the LoRA.
|
||||
"""
|
||||
return [
|
||||
("Why did Kafka hit different?",
|
||||
SamplingParams(temperature=0.7,
|
||||
top_p=0.95,
|
||||
max_tokens=1024), None),
|
||||
|
||||
("Create a Markdown table comparing SHA‑256, BLAKE3, and SHA‑3 with columns: internal structure, block size, and throughput.",
|
||||
SamplingParams(temperature=0.6,
|
||||
top_p=0.95,
|
||||
top_k=20,
|
||||
logprobs=1,
|
||||
prompt_logprobs=1,
|
||||
stop=["</s>", "<eos>"],
|
||||
max_tokens=4096),
|
||||
LoRARequest("reasoning-lora", 1, lora_path)),
|
||||
|
||||
]
|
||||
|
||||
|
||||
def process_requests(engine: LLMEngine,
|
||||
test_prompts: list[tuple[str, SamplingParams,
|
||||
Optional[LoRARequest]]]):
|
||||
"""Continuously process a list of prompts and handle the outputs."""
|
||||
request_id = 0
|
||||
|
||||
while test_prompts or engine.has_unfinished_requests():
|
||||
if test_prompts:
|
||||
input, sampling_params, lora_request = test_prompts.pop(0)
|
||||
prompt = tokenizer.apply_chat_template([{'role':'system', 'content': SYS_MESSAGE},{'role': 'user', 'content': input}], tokenize = False, add_generation_prompt = True)
|
||||
engine.add_request(str(request_id),
|
||||
prompt,
|
||||
sampling_params,
|
||||
lora_request=lora_request)
|
||||
request_id += 1
|
||||
|
||||
request_outputs: list[RequestOutput] = engine.step()
|
||||
|
||||
for request_output in request_outputs:
|
||||
if request_output.finished:
|
||||
print(request_output)
|
||||
|
||||
|
||||
def initialize_engine() -> LLMEngine:
|
||||
"""Initialize the LLMEngine."""
|
||||
|
||||
engine_args = EngineArgs(model="seedboxai/KafkaLM-15B",
|
||||
enable_lora=True,
|
||||
max_loras=1,
|
||||
max_lora_rank=128,
|
||||
max_cpu_loras=2,
|
||||
max_num_seqs=256)
|
||||
|
||||
return LLMEngine.from_engine_args(engine_args)
|
||||
|
||||
|
||||
def main():
|
||||
"""Main function that sets up and runs the prompt processing."""
|
||||
engine = initialize_engine()
|
||||
lora_path = snapshot_download(repo_id="seedboxai/KafkaLM-15B-GRPO_LoRA_Exp")
|
||||
test_prompts = create_test_prompts(lora_path)
|
||||
process_requests(engine, test_prompts)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
|
||||
```
|
||||
|
||||
|
||||
## Citation
|
||||
```bibtex
|
||||
@misc{kafkalm2025,
|
||||
title={Evaluating AMD's MI300A APU: Performance Insights on LLM Training via Knowledge Distillation},
|
||||
author={Dennis Dickmann, Philipp Offenhäuser, Rishabh Saxena, George S. Markomanolis, Alessandro Rigazzi, Patrick Keller, Dennis Hoppe},
|
||||
howpublished={Cray User Group Conference, 2025},
|
||||
note={to be published},
|
||||
year={2025}
|
||||
}
|
||||
```
|
||||
350
config.json
Normal file
350
config.json
Normal file
@@ -0,0 +1,350 @@
|
||||
{
|
||||
"architectures": [
|
||||
"MistralForCausalLM"
|
||||
],
|
||||
"attention_dropout": 0.0,
|
||||
"bos_token_id": 1,
|
||||
"eos_token_id": 2,
|
||||
"head_dim": 128,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 5120,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 25600,
|
||||
"max_position_embeddings": 32768,
|
||||
"model_type": "mistral",
|
||||
"num_attention_heads": 32,
|
||||
"num_hidden_layers": 32,
|
||||
"num_key_value_heads": 8,
|
||||
"pad_token_id": 11,
|
||||
"rms_norm_eps": 1e-05,
|
||||
"rope_theta": 100000000.0,
|
||||
"sliding_window": null,
|
||||
"sparsity": [
|
||||
{
|
||||
"layer_idx": 0,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.0.self_attn.o_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 1,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.1.self_attn.v_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 1,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.1.self_attn.o_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 2,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.2.self_attn.v_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 2,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.2.self_attn.o_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 3,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.3.self_attn.v_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 4,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.4.self_attn.q_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 4,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.4.self_attn.o_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 5,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.5.self_attn.k_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 5,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.5.self_attn.o_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 7,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.7.self_attn.k_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 7,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.7.self_attn.o_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 8,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.8.self_attn.k_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 9,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.9.self_attn.k_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 12,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.12.self_attn.q_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 15,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.15.self_attn.q_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 15,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.15.self_attn.o_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 16,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.16.self_attn.q_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 17,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.17.self_attn.o_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 18,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.18.self_attn.q_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 19,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.19.self_attn.q_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 19,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.19.self_attn.v_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 20,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.20.self_attn.q_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 20,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.20.self_attn.v_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 24,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.24.self_attn.k_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 24,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.24.self_attn.v_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 24,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.24.self_attn.o_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 25,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.25.self_attn.q_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 25,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.25.self_attn.k_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 25,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.25.self_attn.v_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 25,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.25.self_attn.o_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 26,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.26.self_attn.q_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 26,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.26.self_attn.v_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 27,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.27.self_attn.v_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 27,
|
||||
"layer_type": "mlp",
|
||||
"param_name": "model.layers.27.mlp.down_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 131072000,
|
||||
"zeroed": 65536000
|
||||
},
|
||||
{
|
||||
"layer_idx": 28,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.28.self_attn.q_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 30,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.30.self_attn.o_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 31,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.31.self_attn.q_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
},
|
||||
{
|
||||
"layer_idx": 31,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.31.self_attn.v_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 5242880,
|
||||
"zeroed": 2621440
|
||||
},
|
||||
{
|
||||
"layer_idx": 31,
|
||||
"layer_type": "self_attn",
|
||||
"param_name": "model.layers.31.self_attn.o_proj.weight",
|
||||
"sparsity": 50.0,
|
||||
"total": 20971520,
|
||||
"zeroed": 10485760
|
||||
}
|
||||
],
|
||||
"tie_word_embeddings": false,
|
||||
"torch_dtype": "bfloat16",
|
||||
"transformers_version": "4.51.1",
|
||||
"unsloth_fixed": true,
|
||||
"use_cache": true,
|
||||
"vocab_size": 131072
|
||||
}
|
||||
7
generation_config.json
Normal file
7
generation_config.json
Normal file
@@ -0,0 +1,7 @@
|
||||
{
|
||||
"_from_model_config": true,
|
||||
"bos_token_id": 1,
|
||||
"eos_token_id": 2,
|
||||
"pad_token_id": 11,
|
||||
"transformers_version": "4.51.1"
|
||||
}
|
||||
3
model-00001-of-00007.safetensors
Normal file
3
model-00001-of-00007.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:53f7cdac68543976747655250aa2fd964f7f17866d9629e4bb88064fc96a41dc
|
||||
size 4970336816
|
||||
3
model-00002-of-00007.safetensors
Normal file
3
model-00002-of-00007.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:2f89503c4eba410a7bd81339a6a24b30781e563ea6eb1f9f344e2f172ba1379d
|
||||
size 4760642896
|
||||
3
model-00003-of-00007.safetensors
Normal file
3
model-00003-of-00007.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:edf690ca91e6ed235dd67692961785a8ae156235b2c941f74e6af981619a750d
|
||||
size 4980864600
|
||||
3
model-00004-of-00007.safetensors
Normal file
3
model-00004-of-00007.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:35b32b635bf1ff11e75798df62c7c2c9c643d1a6c15a7ef075a1b2ee2973301c
|
||||
size 4823557848
|
||||
3
model-00005-of-00007.safetensors
Normal file
3
model-00005-of-00007.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:bf3fd470035ca0104cf6afb48c5e9c2d2adb13449a4acac540d15e860dea9a28
|
||||
size 4980864616
|
||||
3
model-00006-of-00007.safetensors
Normal file
3
model-00006-of-00007.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:0f9c9b57757216c5f4bb5ffba83c658e688bc43d6dcab2289005265040fcc5f4
|
||||
size 4823557848
|
||||
3
model-00007-of-00007.safetensors
Normal file
3
model-00007-of-00007.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:b03a181292b11a2085196de38f800df843c3948e52e465616d1c308e05698d37
|
||||
size 1866496680
|
||||
298
model.safetensors.index.json
Normal file
298
model.safetensors.index.json
Normal file
@@ -0,0 +1,298 @@
|
||||
{
|
||||
"metadata": {
|
||||
"total_size": 31206287360
|
||||
},
|
||||
"weight_map": {
|
||||
"lm_head.weight": "model-00007-of-00007.safetensors",
|
||||
"model.embed_tokens.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.0.input_layernorm.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.0.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.0.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.0.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.0.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.0.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.0.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.0.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.0.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.1.input_layernorm.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.1.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.1.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.1.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.1.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.1.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.1.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.1.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.1.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.10.input_layernorm.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.10.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.10.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.10.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.10.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.10.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.10.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.10.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.10.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.11.input_layernorm.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.11.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.11.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.11.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.11.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.11.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.11.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.11.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.11.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.12.input_layernorm.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.12.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.12.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.12.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.12.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.12.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.12.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.12.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.12.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.13.input_layernorm.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.13.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.13.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.13.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.13.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.13.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.13.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.13.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.13.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.14.input_layernorm.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.14.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.14.mlp.gate_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.14.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.14.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.14.self_attn.k_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.14.self_attn.o_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.14.self_attn.q_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.14.self_attn.v_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.15.input_layernorm.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.15.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.15.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.15.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.15.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.15.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.15.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.15.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.15.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.16.input_layernorm.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.16.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.16.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.16.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.16.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.16.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.16.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.16.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.16.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.17.input_layernorm.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.17.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.17.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.17.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.17.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.17.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.17.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.17.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.17.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.18.input_layernorm.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.18.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.18.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.18.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.18.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.18.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.18.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.18.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.18.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.19.input_layernorm.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.19.mlp.down_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.19.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.19.mlp.up_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.19.post_attention_layernorm.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.19.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.19.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.19.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.19.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.2.input_layernorm.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.2.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.2.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.2.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.2.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.2.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.2.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.2.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.2.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.20.input_layernorm.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.20.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.20.mlp.gate_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.20.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.20.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.20.self_attn.k_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.20.self_attn.o_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.20.self_attn.q_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.20.self_attn.v_proj.weight": "model-00004-of-00007.safetensors",
|
||||
"model.layers.21.input_layernorm.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.21.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.21.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.21.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.21.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.21.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.21.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.21.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.21.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.22.input_layernorm.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.22.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.22.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.22.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.22.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.22.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.22.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.22.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.22.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.23.input_layernorm.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.23.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.23.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.23.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.23.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.23.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.23.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.23.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.23.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.24.input_layernorm.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.24.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.24.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.24.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.24.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.24.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.24.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.24.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.24.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.25.input_layernorm.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.25.mlp.down_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.25.mlp.gate_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.25.mlp.up_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.25.post_attention_layernorm.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.25.self_attn.k_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.25.self_attn.o_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.25.self_attn.q_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.25.self_attn.v_proj.weight": "model-00005-of-00007.safetensors",
|
||||
"model.layers.26.input_layernorm.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.26.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.26.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.26.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.26.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.26.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.26.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.26.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.26.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.27.input_layernorm.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.27.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.27.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.27.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.27.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.27.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.27.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.27.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.27.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.28.input_layernorm.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.28.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.28.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.28.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.28.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.28.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.28.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.28.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.28.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.29.input_layernorm.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.29.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.29.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.29.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.29.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.29.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.29.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.29.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.29.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.3.input_layernorm.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.3.mlp.down_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.3.mlp.gate_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.3.mlp.up_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.3.post_attention_layernorm.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.3.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.3.self_attn.o_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.3.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.3.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.30.input_layernorm.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.30.mlp.down_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.30.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.30.mlp.up_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.30.post_attention_layernorm.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.30.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.30.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.30.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.30.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.31.input_layernorm.weight": "model-00007-of-00007.safetensors",
|
||||
"model.layers.31.mlp.down_proj.weight": "model-00007-of-00007.safetensors",
|
||||
"model.layers.31.mlp.gate_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.31.mlp.up_proj.weight": "model-00007-of-00007.safetensors",
|
||||
"model.layers.31.post_attention_layernorm.weight": "model-00007-of-00007.safetensors",
|
||||
"model.layers.31.self_attn.k_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.31.self_attn.o_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.31.self_attn.q_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.31.self_attn.v_proj.weight": "model-00006-of-00007.safetensors",
|
||||
"model.layers.4.input_layernorm.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.4.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.4.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.4.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.4.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.4.self_attn.k_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.4.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.4.self_attn.q_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.4.self_attn.v_proj.weight": "model-00001-of-00007.safetensors",
|
||||
"model.layers.5.input_layernorm.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.5.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.5.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.5.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.5.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.5.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.5.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.5.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.5.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.6.input_layernorm.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.6.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.6.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.6.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.6.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.6.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.6.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.6.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.6.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.7.input_layernorm.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.7.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.7.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.7.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.7.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.7.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.7.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.7.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.7.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.8.input_layernorm.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.8.mlp.down_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.8.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.8.mlp.up_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.8.post_attention_layernorm.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.8.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.8.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.8.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.8.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.9.input_layernorm.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.9.mlp.down_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.9.mlp.gate_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.9.mlp.up_proj.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.9.post_attention_layernorm.weight": "model-00003-of-00007.safetensors",
|
||||
"model.layers.9.self_attn.k_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.9.self_attn.o_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.9.self_attn.q_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.layers.9.self_attn.v_proj.weight": "model-00002-of-00007.safetensors",
|
||||
"model.norm.weight": "model-00007-of-00007.safetensors"
|
||||
}
|
||||
}
|
||||
1032
special_tokens_map.json
Normal file
1032
special_tokens_map.json
Normal file
File diff suppressed because it is too large
Load Diff
3
tokenizer.json
Normal file
3
tokenizer.json
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:b76085f9923309d873994d444989f7eb6ec074b06f25b58f1e8d7b7741070949
|
||||
size 17078037
|
||||
9021
tokenizer_config.json
Normal file
9021
tokenizer_config.json
Normal file
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user