Model: nassimjp/Ghanam-7B-Base-Pashto-v0.1 Source: Original Platform
library_name, tags, language, license, pipeline_tag, base_model
| library_name | tags | language | license | pipeline_tag | base_model | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| transformers |
|
|
apache-2.0 | text-generation | mistralai/Mistral-7B-v0.1 |
Ghanam-7B-Base-Pashto-v0.1
Model Description
Ghanam-7B-Base-Pashto-v0.1 is an experimental base model that extends the Mistral-7B-v0.1 vocabulary to include 47 Pashto/Arabic characters. This is the first stage of creating a Pashto-capable language model, where the tokenizer has been modified and the embedding layer has been expanded to support Pashto script.
Key Features
- Vocabulary Size: 32,010 tokens (original 32,000 + 47 Pashto characters)
- Quantization: 4-bit NF4 with double quantization for efficient inference
- Memory Efficient: ~4GB VRAM usage for inference
- Pashto Script Support: Now recognizes all Pashto-specific characters including ښ، ږ، څ، ډ، ړ، ځ، ګ، ۍ، ې
Pashto Characters Added
ا ب پ ت ث ج چ ح خ د ذ ر ز ژ س ش ص ض ط ظ ع غ ف ق ک ګ گ ل م ن ڼ و ه ی ډ ړ ځ څ ښ ږ ۍ ې ء آ أ ؤ ئ
⚠️ Important Note
This model does not yet understand Pashto language semantics. It can:
- ✅ Tokenize Pashto text correctly
- ✅ Generate Pashto characters (if prompted in Pashto)
- ❌ Understand Pashto meaning or context
- ❌ Respond meaningfully in Pashto
The model still generates English text as it hasn't been fine-tuned on Pashto data.
Model Details
Model Description
This model is the first step in creating a Pashto language model based on Mistral-7B. The vocabulary was expanded using the Stanford vocabulary expansion method (mean resizing), where new embeddings were initialized using the mean of existing embeddings to minimize disruption to the model's performance.
- Developed by: Nassim JP (Community Project)
- Model type: Causal Language Model (Decoder-only Transformer)
- Language(s): English (original) + Pashto script support (tokenization only)
- License: Apache 2.0
- Base Model: Mistral-7B-v0.1
- Quantization: 4-bit NF4 (BitsAndBytes)
- Vocabulary Expansion Method: Stanford mean resizing
Model Sources
- Repository: [Coming Soon]
- Base Model: mistralai/Mistral-7B-v0.1
Uses
Direct Use
You can use this model for:
- Pashto tokenization and text preprocessing
- Inference experiments with Pashto prompts (though responses will be in English)
- As a starting point for fine-tuning on Pashto datasets
- Research on vocabulary expansion for low-resource languages
Downstream Use
This model is intended to be fine-tuned on Pashto text data for:
- Pashto language modeling
- Pashto text generation
- Pashto translation tasks
- Pashto question answering
Out-of-Scope Use
- Do not use this model for production Pashto applications without fine-tuning
- Do not expect Pashto understanding or generation in Pashto
- Not suitable for tasks requiring Pashto semantic comprehension
- Not tested for bias, toxicity, or safety in Pashto context
Bias, Risks, and Limitations
Technical Limitations
- No Pashto Understanding: The model has only expanded vocabulary but has not learned Pashto semantics
- English-Only Knowledge: All pre-training knowledge remains in English
- Potential Performance Degradation: Vocabulary expansion may slightly impact English performance
- Untested on Pashto Tasks: No evaluation has been conducted on Pashto benchmarks
Recommendations
- Use this model only as a foundation for Pashto fine-tuning
- Validate carefully before using in any production environment
- Expect English responses when prompting in Pashto (until fine-tuned)
- Consider this a research checkpoint rather than a production model
How to Get Started with the Model
Loading the Model (4-bit)
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
# Model path (local or from Hub)
model_path = "nassimjp/Ghanam-7B-Base-Pashto-v0.1"
# 4-bit quantization config
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16
)
# Load model and tokenizer
model = AutoModelForCausalLM.from_pretrained(
model_path,
quantization_config=bnb_config,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_path)
# Set padding token (optional)
tokenizer.pad_token = tokenizer.eos_token
# Test tokenization
pashto_text = "پښتونخوا کې ډېر ښکلي ښارونه او ځنګلونه شته"
tokens = tokenizer.encode(pashto_text)
print(f"Tokens: {tokens}")
print(f"Decoded: {tokenizer.decode(tokens)}")
Generating Text
# Prompt in Pashto (model will respond in English)
prompt = "سلام، څنګه یاست؟"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=50,
temperature=0.7,
do_sample=True,
repetition_penalty=1.2
)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
Inspecting New Tokens
# Check if Pashto characters are recognized
pashto_chars = ["ښ", "ږ", "څ", "ډ", "ړ", "ځ", "ګ"]
for char in pashto_chars:
token_id = tokenizer.encode(char)[0]
token_str = tokenizer.decode([token_id])
print(f"Character: {char} -> Token ID: {token_id} -> Decoded: {token_str}")
Training Details
Training Data
No training data was used for this model. The vocabulary was expanded without any fine-tuning on Pashto text. The model retains only the original Mistral-7B pre-training knowledge.
Training Procedure
Vocabulary Expansion Method
The model underwent a two-stage process:
Stage 1: Tokenizer Modification
- Added 47 Pashto/Arabic characters to the original tokenizer
- Preserved all original tokens (no replacements)
- New vocabulary size: 32,010
Stage 2: Embedding Resizing
- Used Stanford vocabulary expansion method (mean resizing)
- New embeddings initialized with the mean of existing embeddings
- Both
lm_headand embedding layers were resized - No gradient updates or training performed
Training Hyperparameters
- Quantization: 4-bit NF4 with double quantization
- Compute dtype: bfloat16
- Method: Vocabulary expansion only (no fine-tuning)
- Embedding Initialization: Mean resizing
Evaluation
Testing Data, Factors & Metrics
Testing Data
Initial testing was performed using:
- Pashto sample texts to verify tokenization
- English prompts to ensure base functionality
- Character-level inspection to confirm tokenizer additions
Metrics
- Vocabulary Size: Successfully expanded from 32,000 to 32,010
- Tokenization Accuracy: All 47 characters are properly tokenized
- Generation Quality: English generation preserved
- Memory Usage: ~4GB VRAM for inference
Results
| Test | Result | Status |
|---|---|---|
| Pashto Tokenization | ✅ All 47 characters recognized | PASS |
| English Generation | ✅ Original performance preserved | PASS |
| Pashto Understanding | ❌ No semantic understanding | EXPECTED |
| Pashto Response | ❌ Generates English | EXPECTED |
| Model Loading (4-bit) | ✅ Successfully loads | PASS |
| Memory Efficiency | ✅ ~4GB VRAM | PASS |
Environmental Impact
Carbon emissions were minimal as this model:
-
Did not require any training
-
Only performed vocabulary expansion and quantization
-
Used a single GPU for a few minutes
-
Hardware Type: NVIDIA A100 (or similar)
-
Hours used: < 1 hour
-
Cloud Provider: N/A (local)
-
Compute Region: N/A
-
Carbon Emitted: Negligible
Technical Specifications
Model Architecture and Objective
- Architecture: Transformer decoder (Mistral-7B)
- Number of Layers: 32
- Hidden Size: 4096
- Attention Heads: 32
- Intermediate Size: 14336
- Vocabulary Size: 32,010 (expanded from 32,000)
- Positional Encoding: Rotary Position Embeddings (RoPE)
Compute Infrastructure
Hardware
- CPU for embedding expansion
- GPU for quantization and inference testing
Software
- Transformers: v4.31.0+
- BitsAndBytes: v0.41.0+
- PyTorch: v2.0.0+
- Accelerate: v0.20.0+
Model Card Authors
- Nassim JP - Vocabulary expansion and quantization implementation
- Original model: Mistral AI team
Model Card Contact
For questions or collaboration:
- Hugging Face: nassimjp
Acknowledgments
- Mistral AI for the base model
- Stanford NLP for the vocabulary expansion method
- Hugging Face for the transformers library
- BitsAndBytes team for 4-bit quantization
Glossary
- Vocabulary Expansion: Adding new tokens to a model's tokenizer
- Mean Resizing: Initializing new embeddings with the mean of existing embeddings
- 4-bit NF4: 4-bit Normal Float quantization method
- Fine-tuning: Training a pre-trained model on domain-specific data
Next Steps
To make this model truly Pashto-capable, the next steps include:
- Collect Pashto corpus (books, articles, web text)
- Pre-train on Pashto data (continued pre-training)
- Fine-tune for specific tasks (translation, QA, generation)
- Evaluate on Pashto benchmarks
- Deploy in production environments
Citation
If you use this model or the vocabulary expansion method, please cite:
@misc{ghanam2024pashto,
author = {Nassim JP},
title = {Ghanam-7B-Base-Pashto: Vocabulary Expansion for Pashto Language},
year = {2024},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/nassimjp/Ghanam-7B-Base-Pashto-v0.1}}
}
Original Mistral-7B Citation
@article{jiang2023mistral,
title={Mistral 7B},
author={Jiang, Albert Q and others},
journal={arXiv preprint arXiv:2310.06825},
year={2023}
}
Additional Information
Usage Tips
- Tokenization: The tokenizer now correctly handles all Pashto characters
- Memory: 4-bit quantization allows running on consumer GPUs with 6-8GB VRAM
- Generation: Use
do_sample=Trueandtemperature=0.7for diverse outputs - Pad Token: If using batch generation, set
pad_token=tokenizer.eos_token
Common Issues
| Issue | Solution |
|---|---|
| Tokenizer missing characters | Ensure you're using the correct tokenizer from this repo |
| Memory errors | Reduce max_new_tokens or use CPU offloading |
| Poor generation quality | This is expected - the model needs fine-tuning |
Pashto text displayed as [UNK] |
Tokenizer hasn't loaded properly - reload from the repo |
Model Status: 🔬 Research Phase - Pre-fine-tuning checkpoint
Last Updated: June 2026