初始化项目,由ModelHub XC社区提供模型

Model: ibm-research/granite-4.0-h-3b-ar
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-27 19:20:06 +08:00
commit 1a6d234dda
14 changed files with 603411 additions and 0 deletions

35
.gitattributes vendored Normal file
View File

@@ -0,0 +1,35 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

603
README.md Normal file
View File

@@ -0,0 +1,603 @@
---
license: apache-2.0
language:
- en
- ar
- acm
- apc
- ars
- ary
- arz
- afb
pipeline_tag: text-generation
---
# ibm-research/granite-4.0-h-3b-ar
## granite-4.0-h-3b-ar: A multilingual LLM for English and Arabic, including MSA and its regional dialects.
## Model Summary
**granite-4.0-h-3b-ar** is a lightweight 3-billion-parameter **instruct** model developed through a collaboration between IBM Research and the Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) as part of the IBM–MBZUAI AI Center of Excellence. Despite its compact size, the model delivers strong performance, making it well-suited for efficient deployment in resource-constrained environments.
The model builds upon the capabilities of [granite-4.0-h-micro-base](https://huggingface.co/ibm-granite/granite-4.0-micro-base) and uses the exact architecture (GQA, Mamba2, MLP with SwiGLU, RMSNorm, and shared input/output embeddings), delivering enhanced performance across Modern Standard Arabic (MSA, `ar`) and multiple Arabic dialects, including Egyptian (`arz`), Moroccan (`ary`), Syrian/Levantine (`apc`), Saudi/Najdi (`ars`), and Emirati/Gulf (`afb`). At the same time, it retains strong proficiency in English (`en`), making it a robust and efficient multilingual model suitable for a wide range of cross-lingual and dialectal applications.
## Key Technical Specifications
- Model Developers: **IBM Research** and **MBZUAI** (IBM–MBZUAI AI Center of Excellence)
- Languages: Arabic (MSA & dialects) and English
- Architecture: Decoder-only dense transformer architecture
- Parameters: 3 Billion
- Context Length: 8,192
- Vocabulary Size: 100,352
- Core architecture components: GQA, Mamba2, MLP with SwiGLU, RMSNorm, and shared input/output embeddings
## Training Procedure
The enhanced Arabic and dialect-capable model is built starting from the granite-4.0-h-micro-base, following a carefully designed multi-stage pipeline. First, continued pretraining (CPT) adapts the base model to Arabic and its dialects by exposing it to large-scale bilingual data. This stage enables the model to acquire language-specific features and broaden its knowledge across multiple aspects of Arabic contexts, including culture, religion, and regional variations. Next, instruction fine-tuning (IFT) is applied to improve the model’s ability to follow Arabic and dialectal instructions precisely and consistently. Afterward, a layer-wise model merging strategy combines the fine-tuned model with the granite-4.0-h-micro model. This step is critical to ensure that the model retains its strong English capabilities while gaining enhanced Arabic performance. Finally, a second round of instruction fine-tuning is conducted on the merged model. This stage stabilizes the merged weights and further refines overall performance, resulting in a balanced, high-quality bilingual model that performs well across English, standard Arabic, and dialectal Arabic settings.
## Training Data
The training data consists of a mixture of public English and Arabic datasets with permissive licenses, crawled MSA and dialectal data collected in accordance with IBM Data Acquisition guidelines, and synthetic data generated using open-weight large language models.
Rigorous data filtering and quality control procedures were applied to the crawled data. This included removing non-Arabic content using language identification, followed by both exact and fuzzy deduplication. Finally, additional filtering was performed using IBM’s [Data Prep Kit](https://github.com/data-prep-kit/data-prep-kit) to eliminate low-quality samples.
Given the scarcity of high-quality regional dialectal data, a large portion of FineWeb-Edu English educational content was translated into five Arabic dialects using GPT-OSS-120B as the teacher model. The generated translations were further filtered using a LLM-as-a-judge framework with multiple models, in order to ensure both knowledge accuracy and fluency.
### Continued Pre-Training Data
In the Continued Pre-Training (CPT) stage, approximately 100B tokens of English and standard/dialectal Arabic data were used, with a roughly balanced distribution between the two languages. The training data spans diverse domains and is sourced from public datasets such as FineWeb, DCLM, HPLT, Webhose, and Wikipedia, as well as crawled and translated data. Extensive experiments were conducted to optimize the data mixture across these sources, with weights tuned on a heldout development set.
### Instruction Fine-Tuning Data
Due to the limited availability of Arabic and dialectal datasets with permissive licenses, we rely on English datasets and translated data for instruction fine-tuning. Approximately one million examples are used in this stage. The data is mainly curated to support general-domain applications, while also including domain-specific examples for math, coding, and reasoning tasks.
### Infrastructure
The granite-4.0-h-3b-ar language model was trained using an NVIDIA H100 cluster hosted in an IBM data center. The training infrastructure is designed to support large-scale distributed workloads and can scale to thousands of GPUs.
## How to Use
Here is a simple example on how to use granite-4.0-h-3b-ar for text generation:
- Prompt: "أخبرني عن شركة أي بي إم وعن جامعة محمد بن زايد للذكاء الاصطناعي" (Tell me about IBM company and Mohamed bin Zayed University of Artificial Intelligence.)
### Using HuggingFace
Install the following libraries:
```python
pip install torch torchvision torchaudio
pip install accelerate
pip install transformers
```
Then, use the following code:
```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
device= "cuda"
model_id = "ibm-research/granite-4.0-h-3b-ar"
messages = [
{"role": "user", "content": "أخبرني عن شركة أي بي إم وعن جامعة محمد بن زايد للذكاء الاصطناعي"},
]
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map=device,
torch_dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
)
model.eval()
messages = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
input_tokens = tokenizer(messages, return_tensors="pt").to(device)
outputs = model.generate(**input_tokens, max_new_tokens=1024)
outputs = tokenizer.batch_decode(outputs)
print(outputs[0])
```
Expected output (actual output can vary):
```plain
<|start_of_role|>user<|end_of_role|>أخبرني عن شركة أي بي إم وعن جامعة محمد بن زايد للذكاء الاصطناعي<|end_of_text|>
<|start_of_role|>assistant<|end_of_role|> 🇦🇪 أي بي إم، شركة عالمية كبيرة، تقدم خدمات تكنولوجيا معلومات وبرمجيات ومنتجات متنوعة. هي معروفة بتقنيتها المتقدمة وحلولها المبتكرة.
جامعة محمد بن زايد للذكاء الاصطناعي، مؤسسة رائدة في مجال الذكاء الاصطناعي، تهدف إلى تعزيز البحث والتطوير في هذا المجال. تسعى لتكون مركزًا للابتكار والتعاون في مجال الذكاء الاصطناعي.
<|end_of_text|>
```
### Using vLLM (if a GPU is available)
```Python
from vllm import LLM, SamplingParams
model_id = "ibm-research/granite-4.0-h-3b-ar"
messages = [
{"role": "user", "content": "أخبرني عن شركة أي بي إم وعن جامعة محمد بن زايد للذكاء الاصطناعي"},
]
llm = LLM(
model=model_id,
dtype="bfloat16",
enforce_eager=False,
gpu_memory_utilization=0.6,
tensor_parallel_size=1, # Num. of GPUs
)
# Get tokenizer and format messages
tokenizer = llm.get_tokenizer()
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
sampling_params = SamplingParams(max_tokens=1024)
outputs = llm.generate(prompt, sampling_params=sampling_params)
print(outputs[0].outputs[0].text)
```
Expected output (actual output can vary):
```plain
كشركة عالمية مشهورة في مجال التكنولوجيا، تقدم شركة أي بي إم حلولًا متقدمة للشركات والحكومات حول العالم. كما تقوم بتوفير خدماتها للعديد من المؤسسات في دولة الإمارات العربية المتحدة.
فيما يتعلق بجامعة محمد بن زايد للذكاء الاصطناعي، فهي مبادرة رائدة أطلقتها الحكومة الإماراتية بهدف خلق مجتمع معرفي وتعزيز التقنيات المتقدمة. تسعى الجامعة إلى جذب أفضل العقول في مجال الذكاء الاصطناعي من جميع أنحاء العالم وتوفير الدعم البحثي والتعليمي لهم.
```
## Evaluation
We perform the evaluation using the following multilingual benchmarks and the included languages. We used LM-Evaluation-Harness to evaluate different models in a 5-shot setting. In the table below, we list the language labels that are used in each benchmark.
<div style="overflow-x:auto;">
<table>
<caption><b>Multilingual Benchmarks and the included languages:</b></caption>
<thead>
<tr>
<th style="text-align:left; background-color: #001d6c; color: white;">Benchmarks</th>
<th style="text-align:left; background-color: #001d6c; color: white;"># Type</th>
<th style="text-align:center; background-color: #001d6c; color: white;">Languages</th>
<th style="text-align:center; background-color: #001d6c; color: white;">Notes</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">Flores-200</td>
<td style="text-align:center; background-color: #FFFFFF; color: #2D2D2D;">Translation </td>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">MSA <-> English, English <-> [Egyptian, Iraqi, Leventine, Morrocan, Najdi]</td>
<td>-</td>
</tr>
<tr>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">Alyah</td>
<td style="text-align:center; background-color: #FFFFFF; color: #2D2D2D;">QA </td>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">Emirati</td>
<td>UAE culture</td>
</tr>
<tr>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">INCLUDE 44</td>
<td style="text-align:center; background-color: #FFFFFF; color: #2D2D2D;">QA</td>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">MSA</td>
<td>-</td>
</tr>
<tr>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">PalmX 2025</td>
<td style="text-align:center; background-color: #FFFFFF; color: #2D2D2D;">QA </td>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">MSA</td>
<td>Arabic and Islamic culture</td>
</tr>
<tr>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">MMLU-ProX</td>
<td style="text-align:center; background-color: #FFFFFF; color: #2D2D2D;">QA</td>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">English, MSA</td>
<td>-</td>
</tr>
<tr>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">Belebele</td>
<td style="text-align:center; background-color: #FFFFFF; color: #2D2D2D;">QA </td>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">English, MSA, Egyptian, Iraqi, Leventine, Morrocan, Najdi</td>
<td>-</td>
</tr>
<tr>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">DialectalArabicMMLU</td>
<td style="text-align:center; background-color: #FFFFFF; color: #2D2D2D;">QA </td>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">English, MSA, Egyptian, Emirati, Morrocan, Saudi, Syrian</td>
<td>-</td>
</tr>
<tr>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">Global PIQA</td>
<td style="text-align:center; background-color: #FFFFFF; color: #2D2D2D;">QA</td>
<td style="text-align:left; background-color: #FFFFFF; color: #2D2D2D;">English, MSA, Egyptian, Iraqi, Leventine, Morrocan, Najdi</td>
<td>-</td>
</tr>
</tbody>
</table>
</div>
### Translation (5-shot)
The models are evaluated on the Floress-200 dataset, where the performance is measured using **BLEU** score. Here, _Dial._ represents the average performance over the dialects.
<div style="overflow-x:auto;">
<table style="border-collapse:collapse;width:100%;font-size:14px;">
<caption style="margin-bottom:10px;">
<b>Translation Benchmark Results (Higher is Better)</b>
</caption>
<thead>
<tr>
<th style="background:#001d6c;color:white;padding:6px;">Rank</th>
<th style="background:#001d6c;color:white;padding:6px;text-align:left;">Model</th>
<th style="background:#001d6c;color:white;padding:6px;">Size</th>
<th style="background:#001d6c;color:white;padding:6px;">Avg ↓</th>
<th style="background:#001d6c;color:white;padding:6px;">En→MSA</th>
<th style="background:#001d6c;color:white;padding:6px;">MSA→En</th>
<th style="background:#001d6c;color:white;padding:6px;">Dial.→MSA</th>
<th style="background:#001d6c;color:white;padding:6px;">Dial.→En</th>
</tr>
</thead>
<tbody>
<tr style="background:#FFFFFF;">
<td align="center">1</td>
<td>🟢 granite-4.0-h-3b-ar</td>
<td align="center">3.2</td>
<td align="center"><b>28.32</b></td>
<td align="center"><b>27.64</b></td>
<td align="center">36.54</td>
<td align="center"><b>15.23</b></td>
<td align="center"><b>33.86</b></td>
</tr>
<tr style="background:#F7F9FC;">
<td align="center">2</td>
<td>UBC-NileChat-3B</td>
<td align="center">3.1</td>
<td align="center">26.30</td>
<td align="center">22.06</td>
<td align="center">37.00</td>
<td align="center">13.58</td>
<td align="center">32.58</td>
</tr>
<tr style="background:#FFFFFF;">
<td align="center">3</td>
<td>gemma-3-4b-it</td>
<td align="center">4.3</td>
<td align="center">25.06</td>
<td align="center">22.68</td>
<td align="center">35.72</td>
<td align="center">12.53</td>
<td align="center">29.31</td>
</tr>
<tr style="background:#F7F9FC;">
<td align="center">4</td>
<td>jais-fam-2p7b-chat</td>
<td align="center">2.7</td>
<td align="center">21.84</td>
<td align="center">19.79</td>
<td align="center">32.07</td>
<td align="center">9.94</td>
<td align="center">25.58</td>
</tr>
<tr style="background:#FFFFFF;">
<td align="center">5</td>
<td>Nile-Chat-4B</td>
<td align="center">3.9</td>
<td align="center">21.54</td>
<td align="center">9.32</td>
<td align="center"><b>37.12</b></td>
<td align="center">8.42</td>
<td align="center">31.28</td>
</tr>
<tr style="background:#F7F9FC;">
<td align="center">6</td>
<td>Qwen3-4B-Instruct-2507</td>
<td align="center">4.0</td>
<td align="center">20.96</td>
<td align="center">14.64</td>
<td align="center">33.12</td>
<td align="center">9.42</td>
<td align="center">26.68</td>
</tr>
<tr style="background:#FFFFFF;">
<td align="center">7</td>
<td>g4.0-h-micro_hf</td>
<td align="center">3.2</td>
<td align="center">20.70</td>
<td align="center">16.52</td>
<td align="center">30.87</td>
<td align="center">9.25</td>
<td align="center">26.17</td>
</tr>
<tr style="background:#F7F9FC;">
<td align="center">8</td>
<td>Falcon-H1-3B-Instruct</td>
<td align="center">3.1</td>
<td align="center">19.52</td>
<td align="center">12.40</td>
<td align="center">32.31</td>
<td align="center">7.28</td>
<td align="center">26.11</td>
</tr>
<tr style="background:#FFFFFF;">
<td align="center">9</td>
<td>Atlas-Chat-2B</td>
<td align="center">2.6</td>
<td align="center">19.27</td>
<td align="center">10.00</td>
<td align="center">32.38</td>
<td align="center">7.01</td>
<td align="center">27.71</td>
</tr>
<tr style="background:#F7F9FC;">
<td align="center">10</td>
<td>g3.3-2b-instruct</td>
<td align="center">2.5</td>
<td align="center">16.64</td>
<td align="center">12.11</td>
<td align="center">26.80</td>
<td align="center">7.37</td>
<td align="center">20.30</td>
</tr>
</tbody>
</table>
</div>
### Understanding and Cultural Knowledge (5-shot)
The models are evaluated on a suite of QA tasks in a generative mode. The performance is reported using **exact match** with _strict-match_ as a requirement, where the model has to follow the instructions exactly as specified by the task. Below, _(Dial)_ represents the average performance over the dialects.
<div style="overflow-x:auto;">
<table style="border-collapse:collapse;width:100%;font-size:14px;">
<caption style="margin-bottom:10px;">
<b>Arabic-Centric Benchmark Results (Higher is Better)</b>
</caption>
<thead>
<tr>
<th style="background:#001d6c;color:white;padding:6px;">Rank</th>
<th style="background:#001d6c;color:white;padding:6px;text-align:left;">Model</th>
<th style="background:#001d6c;color:white;padding:6px;">Size</th>
<th style="background:#001d6c;color:white;padding:6px;">Avg ↓</th>
<th style="background:#001d6c;color:white;padding:6px;">Belebele (En)</th>
<th style="background:#001d6c;color:white;padding:6px;">DialArMMLU (En)</th>
<th style="background:#001d6c;color:white;padding:6px;">MMLU-ProX (En)</th>
<th style="background:#001d6c;color:white;padding:6px;">Global PIQA (En)</th>
<th style="background:#001d6c;color:white;padding:6px;">Belebele (MSA)</th>
<th style="background:#001d6c;color:white;padding:6px;">DialArMMLU (MSA)</th>
<th style="background:#001d6c;color:white;padding:6px;">PalmX 2025 (MSA)</th>
<th style="background:#001d6c;color:white;padding:6px;">INCLUDE-base (MSA)</th>
<th style="background:#001d6c;color:white;padding:6px;">MMLU-ProX (MSA)</th>
<th style="background:#001d6c;color:white;padding:6px;">Global PIQA (MSA)</th>
<th style="background:#001d6c;color:white;padding:6px;">Alyah (Dial)</th>
<th style="background:#001d6c;color:white;padding:6px;">Belebele (Dial)</th>
<th style="background:#001d6c;color:white;padding:6px;">DialArMMLU (Dial)</th>
<th style="background:#001d6c;color:white;padding:6px;">Global PIQA (Dial)</th>
</tr>
</thead>
<tbody>
<tr style="background:#FFFFFF;">
<td align="center">1</td>
<td>Qwen3-4B-Instruct-2507</td>
<td align="center">4.0</td>
<td align="center"><b>67.16</b></td>
<td align="center"><b>91.67</b></td>
<td align="center"><b>75.41</b></td>
<td align="center"><b>60.77</b></td>
<td align="center">81.0</td>
<td align="center"><b>85.22</b></td>
<td align="center">61.15</td>
<td align="center">62.91</td>
<td align="center">54.35</td>
<td align="center"><b>46.17</b></td>
<td align="center">64.0</td>
<td align="center">61.72</td>
<td align="center">70.67</td>
<td align="center">56.01</td>
<td align="center">69.2</td>
</tr>
<tr style="background:#F7F9FC;">
<td align="center">2</td>
<td>🟢 granite-4.0-h-3b-ar</td>
<td align="center">3.2</td>
<td align="center">65.8</td>
<td align="center">88.0</td>
<td align="center">69.98</td>
<td align="center">36.94</td>
<td align="center"><b>82.0</b></td>
<td align="center">78.89</td>
<td align="center"><b>61.21</b></td>
<td align="center">64.02</td>
<td align="center"><b>59.96</b></td>
<td align="center">28.23</td>
<td align="center"><b>81.0</b></td>
<td align="center">66.24</td>
<td align="center">68.76</td>
<td align="center"><b>59.11</b></td>
<td align="center"><b>76.8</b></td>
</tr>
<tr style="background:#FFFFFF;">
<td align="center">3</td>
<td>UBC-NileChat-3B</td>
<td align="center">3.1</td>
<td align="center">65.35</td>
<td align="center">89.22</td>
<td align="center">67.27</td>
<td align="center">33.39</td>
<td align="center">76.0</td>
<td align="center">83.44</td>
<td align="center">58.44</td>
<td align="center"><b>69.05</b></td>
<td align="center">57.97</td>
<td align="center">25.59</td>
<td align="center">77.0</td>
<td align="center"><b>71.7</b></td>
<td align="center"><b>74.4</b></td>
<td align="center">55.49</td>
<td align="center">76.0</td>
</tr>
<tr style="background:#F7F9FC;">
<td align="center">4</td>
<td>gemma-3-4b-it</td>
<td align="center">4.3</td>
<td align="center">62.68</td>
<td align="center">87.78</td>
<td align="center">64.69</td>
<td align="center">39.8</td>
<td align="center">78.0</td>
<td align="center">76.89</td>
<td align="center">54.13</td>
<td align="center">62.21</td>
<td align="center">54.35</td>
<td align="center">30.15</td>
<td align="center">77.0</td>
<td align="center">64.96</td>
<td align="center">63.67</td>
<td align="center">49.96</td>
<td align="center">74.0</td>
</tr>
<tr style="background:#FFFFFF;">
<td align="center">5</td>
<td>Falcon-H1-3B-Instruct</td>
<td align="center">3.1</td>
<td align="center">61.58</td>
<td align="center">90.44</td>
<td align="center">71.48</td>
<td align="center">52.06</td>
<td align="center">74.0</td>
<td align="center">80.0</td>
<td align="center">52.63</td>
<td align="center">58.39</td>
<td align="center">50.36</td>
<td align="center">25.44</td>
<td align="center">66.0</td>
<td align="center">58.91</td>
<td align="center">63.62</td>
<td align="center">47.92</td>
<td align="center">70.8</td>
</tr>
<tr style="background:#F7F9FC;">
<td align="center">6</td>
<td>g4.0-h-micro_hf</td>
<td align="center">3.2</td>
<td align="center">58.82</td>
<td align="center">87.78</td>
<td align="center">69.09</td>
<td align="center">43.36</td>
<td align="center">80.0</td>
<td align="center">72.67</td>
<td align="center">52.89</td>
<td align="center">54.84</td>
<td align="center">46.74</td>
<td align="center">28.26</td>
<td align="center">65.0</td>
<td align="center">55.16</td>
<td align="center">52.2</td>
<td align="center">47.69</td>
<td align="center">67.8</td>
</tr>
<tr style="background:#FFFFFF;">
<td align="center">7</td>
<td>Nile-Chat-4B</td>
<td align="center">3.9</td>
<td align="center">58.68</td>
<td align="center">83.67</td>
<td align="center">61.82</td>
<td align="center">26.69</td>
<td align="center">73.0</td>
<td align="center">74.89</td>
<td align="center">51.55</td>
<td align="center">58.39</td>
<td align="center">50.91</td>
<td align="center">19.35</td>
<td align="center">72.0</td>
<td align="center">61.3</td>
<td align="center">66.38</td>
<td align="center">48.77</td>
<td align="center">72.8</td>
</tr>
<tr style="background:#F7F9FC;">
<td align="center">8</td>
<td>Atlas-Chat-2B</td>
<td align="center">2.6</td>
<td align="center">53.46</td>
<td align="center">82.33</td>
<td align="center">59.01</td>
<td align="center">25.52</td>
<td align="center">73.0</td>
<td align="center">67.56</td>
<td align="center">44.37</td>
<td align="center">61.07</td>
<td align="center">39.49</td>
<td align="center">12.42</td>
<td align="center">58.0</td>
<td align="center">62.92</td>
<td align="center">56.29</td>
<td align="center">42.7</td>
<td align="center">63.8</td>
</tr>
<tr style="background:#FFFFFF;">
<td align="center">9</td>
<td>g3.3-2b-instruct</td>
<td align="center">2.5</td>
<td align="center">51.07</td>
<td align="center">80.44</td>
<td align="center">58.28</td>
<td align="center">34.2</td>
<td align="center">78.0</td>
<td align="center">58.0</td>
<td align="center">41.98</td>
<td align="center">49.95</td>
<td align="center">37.14</td>
<td align="center">18.42</td>
<td align="center">65.0</td>
<td align="center">48.93</td>
<td align="center">43.29</td>
<td align="center">38.91</td>
<td align="center">62.4</td>
</tr>
<tr style="background:#F7F9FC;">
<td align="center">10</td>
<td>jais-fam-2p7b-chat</td>
<td align="center">2.7</td>
<td align="center">48.82</td>
<td align="center">67.89</td>
<td align="center">40.19</td>
<td align="center">18.39</td>
<td align="center">71.0</td>
<td align="center">64.33</td>
<td align="center">42.23</td>
<td align="center">55.38</td>
<td align="center">28.99</td>
<td align="center">16.06</td>
<td align="center">62.0</td>
<td align="center">63.0</td>
<td align="center">50.96</td>
<td align="center">37.82</td>
<td align="center">65.2</td>
</tr>
</tbody>
</table>
</div>
### Ethical Considerations and Limitations:
granite-4.0-h-3b-ar was trained on a mixture of English and Arabic data, including data covering four regional Arabic dialects. Although the model is designed to support Arabic and dialectal dialogue use cases, its performance may vary depending on the dialect, domain, prompt style, and task complexity, and may not always be comparable to its performance on English-language tasks. For dialect-specific or domain-specific applications, providing a small number of examples through few-shot prompting can help improve output accuracy and consistency.
While granite-4.0-h-3b-ar has been developed with safety considerations in mind, it may still occasionally generate inaccurate, biased, inappropriate, or unsafe responses. Users and developers should conduct task-specific evaluation, safety testing, and additional tuning before deploying the model in production environments.
For enterprise deployments, especially in safety-sensitive or high-risk settings, we recommend pairing granite-4.0-h-3b-ar with appropriate input and output moderation systems, such as Granite Guardian, to help detect and flag risks across relevant dimensions described in the IBM AI Risk Atlas.
## citation
```bibtex
@misc{dialLLM,
title={granite-4.0-h-3b-ar},
author={IBM–MBZUAI AI Center of Excellence},
year={2026},
}
```

114
chat_template.jinja Normal file
View File

@@ -0,0 +1,114 @@
{%- set tools_system_message_prefix = 'You are a helpful assistant with access to the following tools. You may call one or more tools to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>' %}
{%- set tools_system_message_suffix = '\n</tools>\n\nFor each tool call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call>. If a tool does not exist in the provided list of tools, notify the user that you do not have the ability to fulfill the request.' %}
{%- set documents_system_message_prefix = 'You are a helpful assistant with access to the following documents. You may use one or more documents to assist with the user query.\n\nYou are given a list of documents within <documents></documents> XML tags:\n<documents>' %}
{%- set documents_system_message_suffix = '\n</documents>\n\nWrite the response to the user\'s input by strictly aligning with the facts in the provided documents. If the information needed to answer the question is not available in the documents, inform the user that the question cannot be answered based on the available data.' %}
{%- if available_tools is defined and available_tools %}
{%- set tools = available_tools %}
{%- endif %}
{%- set ns = namespace(tools_system_message=tools_system_message_prefix,
documents_system_message=documents_system_message_prefix,
system_message=''
) %}
{%- if tools %}
{%- for tool in tools %}
{%- set ns.tools_system_message = ns.tools_system_message + '\n' + (tool | tojson) %}
{%- endfor %}
{%- set ns.tools_system_message = ns.tools_system_message + tools_system_message_suffix %}
{%- else %}
{%- set ns.tools_system_message = '' %}
{%- endif %}
{%- if documents %}
{%- for document in documents %}
{%- set ns.documents_system_message = ns.documents_system_message + '\n' + (document | tojson) %}
{%- endfor %}
{%- set ns.documents_system_message = ns.documents_system_message + documents_system_message_suffix %}
{%- else %}
{%- set ns.documents_system_message = '' %}
{%- endif %}
{%- if messages[0].role == 'system' %}
{%- if messages[0].content is string %}
{%- set ns.system_message = messages[0].content %}
{%- elif messages[0].content is iterable %}
{%- for entry in messages[0].content %}
{%- if entry.type== 'text' %}
{%- if ns.system_message != '' %}
{%- set ns.system_message = ns.system_message + '\n' %}
{%- endif %}
{%- set ns.system_message = ns.system_message + entry.text %}
{%- endif %}
{%- endfor %}
{%- endif %}
{%- if tools and documents %}
{%- set ns.system_message = ns.system_message + '\n\n' + ns.tools_system_message + '\n\n' + ns.documents_system_message %}
{%- elif tools %}
{%- set ns.system_message = ns.system_message + '\n\n' + ns.tools_system_message %}
{%- elif documents %}
{%- set ns.system_message = ns.system_message + '\n\n' + ns.documents_system_message %}
{%- endif %}
{%- else %}
{%- if tools and documents %}
{%- set ns.system_message = ns.tools_system_message + '\n\n' + ns.documents_system_message %}
{%- elif tools %}
{%- set ns.system_message = ns.tools_system_message %}
{%- elif documents %}
{%- set ns.system_message = ns.documents_system_message %}
{%- endif %}
{%- endif %}
{%- if ns.system_message %}
{{- '<|start_of_role|>system<|end_of_role|>' + ns.system_message + '<|end_of_text|>\n' }}
{%- endif %}
{%- for message in messages %}
{%- set content = namespace(val='') %}
{%- if message.content is string %}
{%- set content.val = message.content %}
{%- else %}
{%- if message.content is iterable %}
{%- for entry in message.content %}
{%- if entry.type== 'text' %}
{%- if content.val != '' %}
{%- set content.val = content.val + '\n' %}
{%- endif %}
{%- set content.val = content.val + entry.text %}
{%- endif %}
{%- endfor %}
{%- endif %}
{%- endif %}
{%- if (message.role == 'user') or (message.role == 'system' and not loop.first) %}
{{- '<|start_of_role|>' + message.role + '<|end_of_role|>' + content.val + '<|end_of_text|>\n' }}
{%- elif message.role == 'assistant' %}
{{- '<|start_of_role|>' + message.role + '<|end_of_role|>' + content.val }}
{%- if message.tool_calls %}
{%- for tool_call in message.tool_calls %}
{%- if (loop.first and content.val) or (not loop.first) %}
{{- '\n' }}
{%- endif %}
{%- if tool_call.function %}
{%- set tool_call = tool_call.function %}
{%- endif %}
{{- '<tool_call>\n{"name": "' }}
{{- tool_call.name }}
{{- '", "arguments": ' }}
{%- if tool_call.arguments is string %}
{{- tool_call.arguments }}
{%- else %}
{{- tool_call.arguments | tojson }}
{%- endif %}
{{- '}\n</tool_call>' }}
{%- endfor %}
{%- endif %}
{{- '<|end_of_text|>\n' }}
{%- elif message.role == 'tool' %}
{%- if loop.first or (messages[loop.index0 - 1].role != 'tool') %}
{{- '<|start_of_role|>user<|end_of_role|>' }}
{%- endif %}
{{- '\n<tool_response>\n' }}
{{- content.val }}
{{- '\n</tool_response>' }}
{%- if loop.last or (messages[loop.index0 + 1].role != 'tool') %}
{{- '<|end_of_text|>\n' }}
{%- endif %}
{%- endif %}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|start_of_role|>assistant<|end_of_role|>' }}
{%- endif %}

89
config.json Normal file
View File

@@ -0,0 +1,89 @@
{
"architectures": [
"GraniteMoeHybridForCausalLM"
],
"attention_bias": false,
"attention_dropout": 0.0,
"attention_multiplier": 0.015625,
"bos_token_id": 100257,
"dtype": "bfloat16",
"embedding_multiplier": 12,
"eos_token_id": 100257,
"hidden_act": "silu",
"hidden_size": 2048,
"initializer_range": 0.1,
"intermediate_size": 8192,
"layer_types": [
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"attention",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"attention",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"attention",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"mamba",
"attention",
"mamba",
"mamba",
"mamba",
"mamba"
],
"logits_scaling": 8,
"mamba_chunk_size": 256,
"mamba_conv_bias": true,
"mamba_d_conv": 4,
"mamba_d_head": 64,
"mamba_d_state": 128,
"mamba_expand": 2,
"mamba_n_groups": 1,
"mamba_n_heads": 64,
"mamba_proj_bias": false,
"max_position_embeddings": 131072,
"model_type": "granitemoehybrid",
"normalization_function": "rmsnorm",
"num_attention_heads": 32,
"num_experts_per_tok": 0,
"num_hidden_layers": 40,
"num_key_value_heads": 8,
"num_local_experts": 0,
"output_router_logits": false,
"pad_token_id": 100256,
"position_embedding_type": "nope",
"residual_multiplier": 0.22,
"rms_norm_eps": 1e-05,
"rope_scaling": null,
"rope_theta": 10000,
"router_aux_loss_coef": 0.01,
"shared_intermediate_size": 8192,
"tie_word_embeddings": true,
"transformers_version": "4.57.0",
"use_cache": false,
"vocab_size": 100352
}

9
generation_config.json Normal file
View File

@@ -0,0 +1,9 @@
{
"_from_model_config": true,
"bos_token_id": 100257,
"eos_token_id": [
100257
],
"pad_token_id": 100256,
"transformers_version": "4.57.0"
}

100001
merges.txt Normal file

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5d12e24f71bd2fd851b55883965d111395d895f1a9fa53afcf20138565e6dd32
size 4990606832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:74c191f00a680733a9418df48835270985808862050aa951a1c061eff71892b4
size 1803279192

View File

@@ -0,0 +1,474 @@
{
"metadata": {
"total_size": 6793833984
},
"weight_map": {
"lm_head.weight": "model-00002-of-00002.safetensors",
"model.embed_tokens.weight": "model-00001-of-00002.safetensors",
"model.layers.0.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.0.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.0.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.0.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.0.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.0.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.0.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.0.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.0.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.0.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.0.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.0.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.1.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.1.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.1.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.1.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.1.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.1.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.1.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.1.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.1.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.1.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.1.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.1.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.10.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.10.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.10.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.10.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.10.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.10.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.10.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.10.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.10.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.10.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.10.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.10.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.11.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.11.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.11.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.11.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.11.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.11.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.11.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.11.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.11.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.11.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.11.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.11.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.12.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.12.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.12.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.12.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.12.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.12.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.12.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.12.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.12.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.12.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.12.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.12.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.13.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.13.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.13.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.13.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.13.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.13.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.13.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.13.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.13.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.13.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.13.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.13.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.14.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.14.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.14.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.14.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.14.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.14.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.14.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.14.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.14.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.14.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.14.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.14.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.15.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.15.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.15.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.15.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.15.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.15.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.15.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.15.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.16.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.16.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.16.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.16.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.16.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.16.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.16.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.16.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.16.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.16.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.16.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.16.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.17.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.17.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.17.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.17.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.17.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.17.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.17.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.17.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.17.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.17.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.17.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.17.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.18.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.18.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.18.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.18.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.18.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.18.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.18.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.18.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.18.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.18.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.18.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.18.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.19.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.19.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.19.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.19.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.19.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.19.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.19.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.19.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.19.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.19.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.19.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.19.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.2.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.2.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.2.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.2.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.2.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.2.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.2.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.2.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.2.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.2.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.2.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.2.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.20.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.20.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.20.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.20.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.20.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.20.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.20.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.20.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.20.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.20.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.20.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.20.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.21.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.21.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.21.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.21.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.21.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.21.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.21.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.21.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.21.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.21.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.21.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.21.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.22.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.22.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.22.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.22.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.22.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.22.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.22.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.22.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.22.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.22.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.22.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.22.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.23.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.23.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.23.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.23.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.23.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.23.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.23.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.23.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.23.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.23.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.23.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.23.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.24.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.24.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.24.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.24.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.24.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.24.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.24.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.24.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.24.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.24.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.24.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.24.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.25.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.25.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.25.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.25.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.25.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.25.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.25.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.25.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.26.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.26.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.26.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.26.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.26.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.26.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.26.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.26.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.26.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.26.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.26.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.26.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.27.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.27.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.27.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.27.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.27.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.27.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.27.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.27.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.27.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.27.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.27.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.27.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.28.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.28.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.28.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.28.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.28.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.28.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.28.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.28.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.28.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.28.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.28.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.28.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.29.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.29.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.29.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.29.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.29.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.29.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.29.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.29.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.29.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.29.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.29.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.29.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.3.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.3.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.3.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.3.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.3.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.3.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.3.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.3.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.3.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.3.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.3.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.3.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.30.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.30.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.30.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.30.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.30.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.30.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.30.mamba.in_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.30.mamba.norm.weight": "model-00002-of-00002.safetensors",
"model.layers.30.mamba.out_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.30.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.30.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.30.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.31.input_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.31.mamba.A_log": "model-00002-of-00002.safetensors",
"model.layers.31.mamba.D": "model-00002-of-00002.safetensors",
"model.layers.31.mamba.conv1d.bias": "model-00002-of-00002.safetensors",
"model.layers.31.mamba.conv1d.weight": "model-00002-of-00002.safetensors",
"model.layers.31.mamba.dt_bias": "model-00002-of-00002.safetensors",
"model.layers.31.mamba.in_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.31.mamba.norm.weight": "model-00002-of-00002.safetensors",
"model.layers.31.mamba.out_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.31.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.31.shared_mlp.input_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.31.shared_mlp.output_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.32.input_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.32.mamba.A_log": "model-00002-of-00002.safetensors",
"model.layers.32.mamba.D": "model-00002-of-00002.safetensors",
"model.layers.32.mamba.conv1d.bias": "model-00002-of-00002.safetensors",
"model.layers.32.mamba.conv1d.weight": "model-00002-of-00002.safetensors",
"model.layers.32.mamba.dt_bias": "model-00002-of-00002.safetensors",
"model.layers.32.mamba.in_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.32.mamba.norm.weight": "model-00002-of-00002.safetensors",
"model.layers.32.mamba.out_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.32.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.32.shared_mlp.input_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.32.shared_mlp.output_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.33.input_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.33.mamba.A_log": "model-00002-of-00002.safetensors",
"model.layers.33.mamba.D": "model-00002-of-00002.safetensors",
"model.layers.33.mamba.conv1d.bias": "model-00002-of-00002.safetensors",
"model.layers.33.mamba.conv1d.weight": "model-00002-of-00002.safetensors",
"model.layers.33.mamba.dt_bias": "model-00002-of-00002.safetensors",
"model.layers.33.mamba.in_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.33.mamba.norm.weight": "model-00002-of-00002.safetensors",
"model.layers.33.mamba.out_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.33.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.33.shared_mlp.input_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.33.shared_mlp.output_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.34.input_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.34.mamba.A_log": "model-00002-of-00002.safetensors",
"model.layers.34.mamba.D": "model-00002-of-00002.safetensors",
"model.layers.34.mamba.conv1d.bias": "model-00002-of-00002.safetensors",
"model.layers.34.mamba.conv1d.weight": "model-00002-of-00002.safetensors",
"model.layers.34.mamba.dt_bias": "model-00002-of-00002.safetensors",
"model.layers.34.mamba.in_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.34.mamba.norm.weight": "model-00002-of-00002.safetensors",
"model.layers.34.mamba.out_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.34.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.34.shared_mlp.input_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.34.shared_mlp.output_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.35.input_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.35.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.35.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.35.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.35.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.35.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.35.shared_mlp.input_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.35.shared_mlp.output_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.36.input_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.36.mamba.A_log": "model-00002-of-00002.safetensors",
"model.layers.36.mamba.D": "model-00002-of-00002.safetensors",
"model.layers.36.mamba.conv1d.bias": "model-00002-of-00002.safetensors",
"model.layers.36.mamba.conv1d.weight": "model-00002-of-00002.safetensors",
"model.layers.36.mamba.dt_bias": "model-00002-of-00002.safetensors",
"model.layers.36.mamba.in_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.36.mamba.norm.weight": "model-00002-of-00002.safetensors",
"model.layers.36.mamba.out_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.36.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.36.shared_mlp.input_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.36.shared_mlp.output_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.37.input_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.37.mamba.A_log": "model-00002-of-00002.safetensors",
"model.layers.37.mamba.D": "model-00002-of-00002.safetensors",
"model.layers.37.mamba.conv1d.bias": "model-00002-of-00002.safetensors",
"model.layers.37.mamba.conv1d.weight": "model-00002-of-00002.safetensors",
"model.layers.37.mamba.dt_bias": "model-00002-of-00002.safetensors",
"model.layers.37.mamba.in_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.37.mamba.norm.weight": "model-00002-of-00002.safetensors",
"model.layers.37.mamba.out_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.37.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.37.shared_mlp.input_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.37.shared_mlp.output_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.38.input_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.38.mamba.A_log": "model-00002-of-00002.safetensors",
"model.layers.38.mamba.D": "model-00002-of-00002.safetensors",
"model.layers.38.mamba.conv1d.bias": "model-00002-of-00002.safetensors",
"model.layers.38.mamba.conv1d.weight": "model-00002-of-00002.safetensors",
"model.layers.38.mamba.dt_bias": "model-00002-of-00002.safetensors",
"model.layers.38.mamba.in_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.38.mamba.norm.weight": "model-00002-of-00002.safetensors",
"model.layers.38.mamba.out_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.38.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.38.shared_mlp.input_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.38.shared_mlp.output_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.39.input_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.39.mamba.A_log": "model-00002-of-00002.safetensors",
"model.layers.39.mamba.D": "model-00002-of-00002.safetensors",
"model.layers.39.mamba.conv1d.bias": "model-00002-of-00002.safetensors",
"model.layers.39.mamba.conv1d.weight": "model-00002-of-00002.safetensors",
"model.layers.39.mamba.dt_bias": "model-00002-of-00002.safetensors",
"model.layers.39.mamba.in_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.39.mamba.norm.weight": "model-00002-of-00002.safetensors",
"model.layers.39.mamba.out_proj.weight": "model-00002-of-00002.safetensors",
"model.layers.39.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
"model.layers.39.shared_mlp.input_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.39.shared_mlp.output_linear.weight": "model-00002-of-00002.safetensors",
"model.layers.4.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.4.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.4.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.4.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.4.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.4.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.4.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.4.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.4.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.4.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.4.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.4.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.5.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.5.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.5.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.5.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.5.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.5.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.5.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.5.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.6.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.6.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.6.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.6.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.6.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.6.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.6.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.6.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.6.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.6.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.6.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.6.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.7.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.7.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.7.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.7.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.7.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.7.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.7.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.7.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.7.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.7.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.7.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.7.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.8.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.8.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.8.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.8.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.8.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.8.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.8.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.8.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.8.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.8.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.8.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.8.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.9.input_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.9.mamba.A_log": "model-00001-of-00002.safetensors",
"model.layers.9.mamba.D": "model-00001-of-00002.safetensors",
"model.layers.9.mamba.conv1d.bias": "model-00001-of-00002.safetensors",
"model.layers.9.mamba.conv1d.weight": "model-00001-of-00002.safetensors",
"model.layers.9.mamba.dt_bias": "model-00001-of-00002.safetensors",
"model.layers.9.mamba.in_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.9.mamba.norm.weight": "model-00001-of-00002.safetensors",
"model.layers.9.mamba.out_proj.weight": "model-00001-of-00002.safetensors",
"model.layers.9.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
"model.layers.9.shared_mlp.input_linear.weight": "model-00001-of-00002.safetensors",
"model.layers.9.shared_mlp.output_linear.weight": "model-00001-of-00002.safetensors",
"model.norm.weight": "model-00002-of-00002.safetensors"
}
}

1
model.sig Normal file

File diff suppressed because one or more lines are too long

30
special_tokens_map.json Normal file
View File

@@ -0,0 +1,30 @@
{
"bos_token": {
"content": "<|end_of_text|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"eos_token": {
"content": "<|end_of_text|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"pad_token": {
"content": "<|pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"unk_token": {
"content": "<|unk|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
}
}

501264
tokenizer.json Normal file

File diff suppressed because it is too large Load Diff

784
tokenizer_config.json Normal file
View File

@@ -0,0 +1,784 @@
{
"add_bos_token": false,
"add_prefix_space": false,
"added_tokens_decoder": {
"100256": {
"content": "<|pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100257": {
"content": "<|end_of_text|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100258": {
"content": "<|fim_prefix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"100259": {
"content": "<|fim_middle|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"100260": {
"content": "<|fim_suffix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"100261": {
"content": "<|fim_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"100262": {
"content": "<|filename|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"100263": {
"content": "<|reponame|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"100264": {
"content": "<|start_of_role|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100265": {
"content": "<|end_of_role|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100266": {
"content": "<|unused_1|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100267": {
"content": "<|start_of_plugin|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100268": {
"content": "<|end_of_plugin|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100269": {
"content": "<|unk|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100270": {
"content": "<tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"100271": {
"content": "</tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"100272": {
"content": "<tool_response>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"100273": {
"content": "</tool_response>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"100274": {
"content": "<think>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"100275": {
"content": "</think>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"100276": {
"content": "<think_on>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100277": {
"content": "<think_off>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100278": {
"content": "<schema>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100279": {
"content": "</schema>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100280": {
"content": "<tools>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100281": {
"content": "</tools>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100282": {
"content": "<documents>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100283": {
"content": "</documents>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100284": {
"content": "<|unused_15|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100285": {
"content": "<|unused_16|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100286": {
"content": "<|unused_17|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100287": {
"content": "<|unused_18|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100288": {
"content": "<|unused_19|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100289": {
"content": "<|unused_20|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100290": {
"content": "<|unused_21|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100291": {
"content": "<|unused_22|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100292": {
"content": "<|unused_23|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100293": {
"content": "<|unused_24|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100294": {
"content": "<|unused_25|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100295": {
"content": "<|unused_26|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100296": {
"content": "<|unused_27|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100297": {
"content": "<|unused_28|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100298": {
"content": "<|unused_29|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100299": {
"content": "<|unused_30|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100300": {
"content": "<|unused_31|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100301": {
"content": "<|unused_32|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100302": {
"content": "<|unused_33|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100303": {
"content": "<|unused_34|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100304": {
"content": "<|unused_35|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100305": {
"content": "<|unused_36|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100306": {
"content": "<|unused_37|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100307": {
"content": "<|unused_38|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100308": {
"content": "<|unused_39|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100309": {
"content": "<|unused_40|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100310": {
"content": "<|unused_41|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100311": {
"content": "<|unused_42|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100312": {
"content": "<|unused_43|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100313": {
"content": "<|unused_44|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100314": {
"content": "<|unused_45|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100315": {
"content": "<|unused_46|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100316": {
"content": "<|unused_47|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100317": {
"content": "<|unused_48|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100318": {
"content": "<|unused_49|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100319": {
"content": "<|unused_50|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100320": {
"content": "<|unused_51|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100321": {
"content": "<|unused_52|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100322": {
"content": "<|unused_53|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100323": {
"content": "<|unused_54|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100324": {
"content": "<|unused_55|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100325": {
"content": "<|unused_56|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100326": {
"content": "<|unused_57|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100327": {
"content": "<|unused_58|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100328": {
"content": "<|unused_59|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100329": {
"content": "<|unused_60|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100330": {
"content": "<|unused_61|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100331": {
"content": "<|unused_62|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100332": {
"content": "<|unused_63|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100333": {
"content": "<|unused_64|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100334": {
"content": "<|unused_65|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100335": {
"content": "<|unused_66|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100336": {
"content": "<|unused_67|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100337": {
"content": "<|unused_68|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100338": {
"content": "<|unused_69|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100339": {
"content": "<|unused_70|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100340": {
"content": "<|unused_71|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100341": {
"content": "<|unused_72|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100342": {
"content": "<|unused_73|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100343": {
"content": "<|unused_74|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100344": {
"content": "<|unused_75|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100345": {
"content": "<|unused_76|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100346": {
"content": "<|unused_77|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100347": {
"content": "<|unused_78|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100348": {
"content": "<|unused_79|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100349": {
"content": "<|unused_80|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100350": {
"content": "<|unused_81|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"100351": {
"content": "<|unused_82|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
}
},
"bos_token": "<|end_of_text|>",
"clean_up_tokenization_spaces": false,
"eos_token": "<|end_of_text|>",
"extra_special_tokens": {},
"legacy": true,
"model_max_length": 1000000000000000019884624838656,
"pad_token": "<|pad|>",
"padding_side": "left",
"tokenizer_class": "GPT2Tokenizer",
"unk_token": "<|unk|>"
}

1
vocab.json Normal file

File diff suppressed because one or more lines are too long