349 lines
11 KiB
Markdown
349 lines
11 KiB
Markdown
|
|
---
|
|||
|
|
library_name: transformers
|
|||
|
|
tags:
|
|||
|
|
- pashto
|
|||
|
|
- mistral
|
|||
|
|
- 4bit
|
|||
|
|
- nf4
|
|||
|
|
- vocabulary-expansion
|
|||
|
|
- afghanistan
|
|||
|
|
- pashto-language
|
|||
|
|
language:
|
|||
|
|
- ps
|
|||
|
|
- en
|
|||
|
|
license: apache-2.0
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
base_model: mistralai/Mistral-7B-v0.1
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Ghanam-7B-Base-Pashto-v0.1
|
|||
|
|
|
|||
|
|
## Model Description
|
|||
|
|
|
|||
|
|
**Ghanam-7B-Base-Pashto-v0.1** is an experimental base model that extends the [Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1) vocabulary to include 47 Pashto/Arabic characters. This is the **first stage** of creating a Pashto-capable language model, where the tokenizer has been modified and the embedding layer has been expanded to support Pashto script.
|
|||
|
|
|
|||
|
|
### Key Features
|
|||
|
|
- **Vocabulary Size**: 32,010 tokens (original 32,000 + 47 Pashto characters)
|
|||
|
|
- **Quantization**: 4-bit NF4 with double quantization for efficient inference
|
|||
|
|
- **Memory Efficient**: ~4GB VRAM usage for inference
|
|||
|
|
- **Pashto Script Support**: Now recognizes all Pashto-specific characters including ښ، ږ، څ، ډ، ړ، ځ، ګ، ۍ، ې
|
|||
|
|
|
|||
|
|
### Pashto Characters Added
|
|||
|
|
```
|
|||
|
|
ا ب پ ت ث ج چ ح خ د ذ ر ز ژ س ش ص ض ط ظ ع غ ف ق ک ګ گ ل م ن ڼ و ه ی ډ ړ ځ څ ښ ږ ۍ ې ء آ أ ؤ ئ
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### ⚠️ Important Note
|
|||
|
|
This model **does not yet understand Pashto language semantics**. It can:
|
|||
|
|
- ✅ Tokenize Pashto text correctly
|
|||
|
|
- ✅ Generate Pashto characters (if prompted in Pashto)
|
|||
|
|
- ❌ Understand Pashto meaning or context
|
|||
|
|
- ❌ Respond meaningfully in Pashto
|
|||
|
|
|
|||
|
|
The model still generates English text as it hasn't been fine-tuned on Pashto data.
|
|||
|
|
|
|||
|
|
## Model Details
|
|||
|
|
|
|||
|
|
### Model Description
|
|||
|
|
|
|||
|
|
This model is the first step in creating a Pashto language model based on Mistral-7B. The vocabulary was expanded using the **Stanford vocabulary expansion method** (mean resizing), where new embeddings were initialized using the mean of existing embeddings to minimize disruption to the model's performance.
|
|||
|
|
|
|||
|
|
- **Developed by:** Nassim JP (Community Project)
|
|||
|
|
- **Model type:** Causal Language Model (Decoder-only Transformer)
|
|||
|
|
- **Language(s):** English (original) + Pashto script support (tokenization only)
|
|||
|
|
- **License:** Apache 2.0
|
|||
|
|
- **Base Model:** [Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1)
|
|||
|
|
- **Quantization:** 4-bit NF4 (BitsAndBytes)
|
|||
|
|
- **Vocabulary Expansion Method:** Stanford mean resizing
|
|||
|
|
|
|||
|
|
### Model Sources
|
|||
|
|
|
|||
|
|
- **Repository:** [Coming Soon]
|
|||
|
|
- **Base Model:** [mistralai/Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1)
|
|||
|
|
|
|||
|
|
## Uses
|
|||
|
|
|
|||
|
|
### Direct Use
|
|||
|
|
|
|||
|
|
You can use this model for:
|
|||
|
|
- **Pashto tokenization** and text preprocessing
|
|||
|
|
- **Inference experiments** with Pashto prompts (though responses will be in English)
|
|||
|
|
- **As a starting point** for fine-tuning on Pashto datasets
|
|||
|
|
- **Research** on vocabulary expansion for low-resource languages
|
|||
|
|
|
|||
|
|
### Downstream Use
|
|||
|
|
|
|||
|
|
This model is intended to be **fine-tuned** on Pashto text data for:
|
|||
|
|
- Pashto language modeling
|
|||
|
|
- Pashto text generation
|
|||
|
|
- Pashto translation tasks
|
|||
|
|
- Pashto question answering
|
|||
|
|
|
|||
|
|
### Out-of-Scope Use
|
|||
|
|
|
|||
|
|
- **Do not use** this model for production Pashto applications without fine-tuning
|
|||
|
|
- **Do not expect** Pashto understanding or generation in Pashto
|
|||
|
|
- **Not suitable** for tasks requiring Pashto semantic comprehension
|
|||
|
|
- **Not tested** for bias, toxicity, or safety in Pashto context
|
|||
|
|
|
|||
|
|
## Bias, Risks, and Limitations
|
|||
|
|
|
|||
|
|
### Technical Limitations
|
|||
|
|
1. **No Pashto Understanding**: The model has only expanded vocabulary but has not learned Pashto semantics
|
|||
|
|
2. **English-Only Knowledge**: All pre-training knowledge remains in English
|
|||
|
|
3. **Potential Performance Degradation**: Vocabulary expansion may slightly impact English performance
|
|||
|
|
4. **Untested on Pashto Tasks**: No evaluation has been conducted on Pashto benchmarks
|
|||
|
|
|
|||
|
|
### Recommendations
|
|||
|
|
- Use this model **only as a foundation** for Pashto fine-tuning
|
|||
|
|
- **Validate carefully** before using in any production environment
|
|||
|
|
- **Expect English responses** when prompting in Pashto (until fine-tuned)
|
|||
|
|
- Consider this a **research checkpoint** rather than a production model
|
|||
|
|
|
|||
|
|
## How to Get Started with the Model
|
|||
|
|
|
|||
|
|
### Loading the Model (4-bit)
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
import torch
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
|
|||
|
|
|
|||
|
|
# Model path (local or from Hub)
|
|||
|
|
model_path = "nassimjp/Ghanam-7B-Base-Pashto-v0.1"
|
|||
|
|
|
|||
|
|
# 4-bit quantization config
|
|||
|
|
bnb_config = BitsAndBytesConfig(
|
|||
|
|
load_in_4bit=True,
|
|||
|
|
bnb_4bit_use_double_quant=True,
|
|||
|
|
bnb_4bit_quant_type="nf4",
|
|||
|
|
bnb_4bit_compute_dtype=torch.bfloat16
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
# Load model and tokenizer
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|||
|
|
model_path,
|
|||
|
|
quantization_config=bnb_config,
|
|||
|
|
device_map="auto"
|
|||
|
|
)
|
|||
|
|
tokenizer = AutoTokenizer.from_pretrained(model_path)
|
|||
|
|
|
|||
|
|
# Set padding token (optional)
|
|||
|
|
tokenizer.pad_token = tokenizer.eos_token
|
|||
|
|
|
|||
|
|
# Test tokenization
|
|||
|
|
pashto_text = "پښتونخوا کې ډېر ښکلي ښارونه او ځنګلونه شته"
|
|||
|
|
tokens = tokenizer.encode(pashto_text)
|
|||
|
|
print(f"Tokens: {tokens}")
|
|||
|
|
print(f"Decoded: {tokenizer.decode(tokens)}")
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### Generating Text
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
# Prompt in Pashto (model will respond in English)
|
|||
|
|
prompt = "سلام، څنګه یاست؟"
|
|||
|
|
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
|||
|
|
|
|||
|
|
with torch.no_grad():
|
|||
|
|
outputs = model.generate(
|
|||
|
|
**inputs,
|
|||
|
|
max_new_tokens=50,
|
|||
|
|
temperature=0.7,
|
|||
|
|
do_sample=True,
|
|||
|
|
repetition_penalty=1.2
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
|
|||
|
|
print(response)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### Inspecting New Tokens
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
# Check if Pashto characters are recognized
|
|||
|
|
pashto_chars = ["ښ", "ږ", "څ", "ډ", "ړ", "ځ", "ګ"]
|
|||
|
|
for char in pashto_chars:
|
|||
|
|
token_id = tokenizer.encode(char)[0]
|
|||
|
|
token_str = tokenizer.decode([token_id])
|
|||
|
|
print(f"Character: {char} -> Token ID: {token_id} -> Decoded: {token_str}")
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
## Training Details
|
|||
|
|
|
|||
|
|
### Training Data
|
|||
|
|
|
|||
|
|
**No training data was used** for this model. The vocabulary was expanded without any fine-tuning on Pashto text. The model retains only the original Mistral-7B pre-training knowledge.
|
|||
|
|
|
|||
|
|
### Training Procedure
|
|||
|
|
|
|||
|
|
#### Vocabulary Expansion Method
|
|||
|
|
|
|||
|
|
The model underwent a **two-stage process**:
|
|||
|
|
|
|||
|
|
**Stage 1: Tokenizer Modification**
|
|||
|
|
- Added 47 Pashto/Arabic characters to the original tokenizer
|
|||
|
|
- Preserved all original tokens (no replacements)
|
|||
|
|
- New vocabulary size: 32,010
|
|||
|
|
|
|||
|
|
**Stage 2: Embedding Resizing**
|
|||
|
|
- Used Stanford vocabulary expansion method (mean resizing)
|
|||
|
|
- New embeddings initialized with the mean of existing embeddings
|
|||
|
|
- Both `lm_head` and embedding layers were resized
|
|||
|
|
- No gradient updates or training performed
|
|||
|
|
|
|||
|
|
#### Training Hyperparameters
|
|||
|
|
|
|||
|
|
- **Quantization**: 4-bit NF4 with double quantization
|
|||
|
|
- **Compute dtype**: bfloat16
|
|||
|
|
- **Method**: Vocabulary expansion only (no fine-tuning)
|
|||
|
|
- **Embedding Initialization**: Mean resizing
|
|||
|
|
|
|||
|
|
## Evaluation
|
|||
|
|
|
|||
|
|
### Testing Data, Factors & Metrics
|
|||
|
|
|
|||
|
|
#### Testing Data
|
|||
|
|
Initial testing was performed using:
|
|||
|
|
- **Pashto sample texts** to verify tokenization
|
|||
|
|
- **English prompts** to ensure base functionality
|
|||
|
|
- **Character-level inspection** to confirm tokenizer additions
|
|||
|
|
|
|||
|
|
#### Metrics
|
|||
|
|
- **Vocabulary Size**: Successfully expanded from 32,000 to 32,010
|
|||
|
|
- **Tokenization Accuracy**: All 47 characters are properly tokenized
|
|||
|
|
- **Generation Quality**: English generation preserved
|
|||
|
|
- **Memory Usage**: ~4GB VRAM for inference
|
|||
|
|
|
|||
|
|
### Results
|
|||
|
|
|
|||
|
|
| Test | Result | Status |
|
|||
|
|
|------|--------|--------|
|
|||
|
|
| Pashto Tokenization | ✅ All 47 characters recognized | PASS |
|
|||
|
|
| English Generation | ✅ Original performance preserved | PASS |
|
|||
|
|
| Pashto Understanding | ❌ No semantic understanding | EXPECTED |
|
|||
|
|
| Pashto Response | ❌ Generates English | EXPECTED |
|
|||
|
|
| Model Loading (4-bit) | ✅ Successfully loads | PASS |
|
|||
|
|
| Memory Efficiency | ✅ ~4GB VRAM | PASS |
|
|||
|
|
|
|||
|
|
## Environmental Impact
|
|||
|
|
|
|||
|
|
Carbon emissions were minimal as this model:
|
|||
|
|
- Did not require any training
|
|||
|
|
- Only performed vocabulary expansion and quantization
|
|||
|
|
- Used a single GPU for a few minutes
|
|||
|
|
|
|||
|
|
- **Hardware Type:** NVIDIA A100 (or similar)
|
|||
|
|
- **Hours used:** < 1 hour
|
|||
|
|
- **Cloud Provider:** N/A (local)
|
|||
|
|
- **Compute Region:** N/A
|
|||
|
|
- **Carbon Emitted:** Negligible
|
|||
|
|
|
|||
|
|
## Technical Specifications
|
|||
|
|
|
|||
|
|
### Model Architecture and Objective
|
|||
|
|
|
|||
|
|
- **Architecture:** Transformer decoder (Mistral-7B)
|
|||
|
|
- **Number of Layers:** 32
|
|||
|
|
- **Hidden Size:** 4096
|
|||
|
|
- **Attention Heads:** 32
|
|||
|
|
- **Intermediate Size:** 14336
|
|||
|
|
- **Vocabulary Size:** 32,010 (expanded from 32,000)
|
|||
|
|
- **Positional Encoding:** Rotary Position Embeddings (RoPE)
|
|||
|
|
|
|||
|
|
### Compute Infrastructure
|
|||
|
|
|
|||
|
|
#### Hardware
|
|||
|
|
- CPU for embedding expansion
|
|||
|
|
- GPU for quantization and inference testing
|
|||
|
|
|
|||
|
|
#### Software
|
|||
|
|
- **Transformers**: v4.31.0+
|
|||
|
|
- **BitsAndBytes**: v0.41.0+
|
|||
|
|
- **PyTorch**: v2.0.0+
|
|||
|
|
- **Accelerate**: v0.20.0+
|
|||
|
|
|
|||
|
|
## Model Card Authors
|
|||
|
|
|
|||
|
|
- **Nassim JP** - Vocabulary expansion and quantization implementation
|
|||
|
|
- Original model: Mistral AI team
|
|||
|
|
|
|||
|
|
## Model Card Contact
|
|||
|
|
|
|||
|
|
For questions or collaboration:
|
|||
|
|
- Hugging Face: [nassimjp](https://huggingface.co/nassimjp)
|
|||
|
|
|
|||
|
|
## Acknowledgments
|
|||
|
|
|
|||
|
|
- Mistral AI for the base model
|
|||
|
|
- Stanford NLP for the vocabulary expansion method
|
|||
|
|
- Hugging Face for the transformers library
|
|||
|
|
- BitsAndBytes team for 4-bit quantization
|
|||
|
|
|
|||
|
|
## Glossary
|
|||
|
|
|
|||
|
|
- **Vocabulary Expansion**: Adding new tokens to a model's tokenizer
|
|||
|
|
- **Mean Resizing**: Initializing new embeddings with the mean of existing embeddings
|
|||
|
|
- **4-bit NF4**: 4-bit Normal Float quantization method
|
|||
|
|
- **Fine-tuning**: Training a pre-trained model on domain-specific data
|
|||
|
|
|
|||
|
|
## Next Steps
|
|||
|
|
|
|||
|
|
To make this model truly Pashto-capable, the next steps include:
|
|||
|
|
|
|||
|
|
1. **Collect Pashto corpus** (books, articles, web text)
|
|||
|
|
2. **Pre-train on Pashto** data (continued pre-training)
|
|||
|
|
3. **Fine-tune** for specific tasks (translation, QA, generation)
|
|||
|
|
4. **Evaluate** on Pashto benchmarks
|
|||
|
|
5. **Deploy** in production environments
|
|||
|
|
|
|||
|
|
## Citation
|
|||
|
|
|
|||
|
|
If you use this model or the vocabulary expansion method, please cite:
|
|||
|
|
|
|||
|
|
```bibtex
|
|||
|
|
@misc{ghanam2024pashto,
|
|||
|
|
author = {Nassim JP},
|
|||
|
|
title = {Ghanam-7B-Base-Pashto: Vocabulary Expansion for Pashto Language},
|
|||
|
|
year = {2024},
|
|||
|
|
publisher = {Hugging Face},
|
|||
|
|
howpublished = {\url{https://huggingface.co/nassimjp/Ghanam-7B-Base-Pashto-v0.1}}
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### Original Mistral-7B Citation
|
|||
|
|
|
|||
|
|
```bibtex
|
|||
|
|
@article{jiang2023mistral,
|
|||
|
|
title={Mistral 7B},
|
|||
|
|
author={Jiang, Albert Q and others},
|
|||
|
|
journal={arXiv preprint arXiv:2310.06825},
|
|||
|
|
year={2023}
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Additional Information
|
|||
|
|
|
|||
|
|
### Usage Tips
|
|||
|
|
|
|||
|
|
1. **Tokenization**: The tokenizer now correctly handles all Pashto characters
|
|||
|
|
2. **Memory**: 4-bit quantization allows running on consumer GPUs with 6-8GB VRAM
|
|||
|
|
3. **Generation**: Use `do_sample=True` and `temperature=0.7` for diverse outputs
|
|||
|
|
4. **Pad Token**: If using batch generation, set `pad_token=tokenizer.eos_token`
|
|||
|
|
|
|||
|
|
### Common Issues
|
|||
|
|
|
|||
|
|
| Issue | Solution |
|
|||
|
|
|-------|----------|
|
|||
|
|
| Tokenizer missing characters | Ensure you're using the correct tokenizer from this repo |
|
|||
|
|
| Memory errors | Reduce `max_new_tokens` or use CPU offloading |
|
|||
|
|
| Poor generation quality | This is expected - the model needs fine-tuning |
|
|||
|
|
| Pashto text displayed as `[UNK]` | Tokenizer hasn't loaded properly - reload from the repo |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
**Model Status:** 🔬 Research Phase - Pre-fine-tuning checkpoint
|
|||
|
|
|
|||
|
|
**Last Updated:** June 2026
|
|||
|
|
```
|