--- library_name: transformers tags: - pashto - mistral - 4bit - nf4 - vocabulary-expansion - afghanistan - pashto-language language: - ps - en license: apache-2.0 pipeline_tag: text-generation base_model: mistralai/Mistral-7B-v0.1 --- # Ghanam-7B-Base-Pashto-v0.1 ## Model Description **Ghanam-7B-Base-Pashto-v0.1** is an experimental base model that extends the [Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1) vocabulary to include 47 Pashto/Arabic characters. This is the **first stage** of creating a Pashto-capable language model, where the tokenizer has been modified and the embedding layer has been expanded to support Pashto script. ### Key Features - **Vocabulary Size**: 32,010 tokens (original 32,000 + 47 Pashto characters) - **Quantization**: 4-bit NF4 with double quantization for efficient inference - **Memory Efficient**: ~4GB VRAM usage for inference - **Pashto Script Support**: Now recognizes all Pashto-specific characters including ښ، ږ، څ، ډ، ړ، ځ، ګ، ۍ، ې ### Pashto Characters Added ``` ا ب پ ت ث ج چ ح خ د ذ ر ز ژ س ش ص ض ط ظ ع غ ف ق ک ګ گ ل م ن ڼ و ه ی ډ ړ ځ څ ښ ږ ۍ ې ء آ أ ؤ ئ ``` ### ⚠️ Important Note This model **does not yet understand Pashto language semantics**. It can: - ✅ Tokenize Pashto text correctly - ✅ Generate Pashto characters (if prompted in Pashto) - ❌ Understand Pashto meaning or context - ❌ Respond meaningfully in Pashto The model still generates English text as it hasn't been fine-tuned on Pashto data. ## Model Details ### Model Description This model is the first step in creating a Pashto language model based on Mistral-7B. The vocabulary was expanded using the **Stanford vocabulary expansion method** (mean resizing), where new embeddings were initialized using the mean of existing embeddings to minimize disruption to the model's performance. - **Developed by:** Nassim JP (Community Project) - **Model type:** Causal Language Model (Decoder-only Transformer) - **Language(s):** English (original) + Pashto script support (tokenization only) - **License:** Apache 2.0 - **Base Model:** [Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1) - **Quantization:** 4-bit NF4 (BitsAndBytes) - **Vocabulary Expansion Method:** Stanford mean resizing ### Model Sources - **Repository:** [Coming Soon] - **Base Model:** [mistralai/Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1) ## Uses ### Direct Use You can use this model for: - **Pashto tokenization** and text preprocessing - **Inference experiments** with Pashto prompts (though responses will be in English) - **As a starting point** for fine-tuning on Pashto datasets - **Research** on vocabulary expansion for low-resource languages ### Downstream Use This model is intended to be **fine-tuned** on Pashto text data for: - Pashto language modeling - Pashto text generation - Pashto translation tasks - Pashto question answering ### Out-of-Scope Use - **Do not use** this model for production Pashto applications without fine-tuning - **Do not expect** Pashto understanding or generation in Pashto - **Not suitable** for tasks requiring Pashto semantic comprehension - **Not tested** for bias, toxicity, or safety in Pashto context ## Bias, Risks, and Limitations ### Technical Limitations 1. **No Pashto Understanding**: The model has only expanded vocabulary but has not learned Pashto semantics 2. **English-Only Knowledge**: All pre-training knowledge remains in English 3. **Potential Performance Degradation**: Vocabulary expansion may slightly impact English performance 4. **Untested on Pashto Tasks**: No evaluation has been conducted on Pashto benchmarks ### Recommendations - Use this model **only as a foundation** for Pashto fine-tuning - **Validate carefully** before using in any production environment - **Expect English responses** when prompting in Pashto (until fine-tuned) - Consider this a **research checkpoint** rather than a production model ## How to Get Started with the Model ### Loading the Model (4-bit) ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig # Model path (local or from Hub) model_path = "nassimjp/Ghanam-7B-Base-Pashto-v0.1" # 4-bit quantization config bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_use_double_quant=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16 ) # Load model and tokenizer model = AutoModelForCausalLM.from_pretrained( model_path, quantization_config=bnb_config, device_map="auto" ) tokenizer = AutoTokenizer.from_pretrained(model_path) # Set padding token (optional) tokenizer.pad_token = tokenizer.eos_token # Test tokenization pashto_text = "پښتونخوا کې ډېر ښکلي ښارونه او ځنګلونه شته" tokens = tokenizer.encode(pashto_text) print(f"Tokens: {tokens}") print(f"Decoded: {tokenizer.decode(tokens)}") ``` ### Generating Text ```python # Prompt in Pashto (model will respond in English) prompt = "سلام، څنګه یاست؟" inputs = tokenizer(prompt, return_tensors="pt").to(model.device) with torch.no_grad(): outputs = model.generate( **inputs, max_new_tokens=50, temperature=0.7, do_sample=True, repetition_penalty=1.2 ) response = tokenizer.decode(outputs[0], skip_special_tokens=True) print(response) ``` ### Inspecting New Tokens ```python # Check if Pashto characters are recognized pashto_chars = ["ښ", "ږ", "څ", "ډ", "ړ", "ځ", "ګ"] for char in pashto_chars: token_id = tokenizer.encode(char)[0] token_str = tokenizer.decode([token_id]) print(f"Character: {char} -> Token ID: {token_id} -> Decoded: {token_str}") ``` ## Training Details ### Training Data **No training data was used** for this model. The vocabulary was expanded without any fine-tuning on Pashto text. The model retains only the original Mistral-7B pre-training knowledge. ### Training Procedure #### Vocabulary Expansion Method The model underwent a **two-stage process**: **Stage 1: Tokenizer Modification** - Added 47 Pashto/Arabic characters to the original tokenizer - Preserved all original tokens (no replacements) - New vocabulary size: 32,010 **Stage 2: Embedding Resizing** - Used Stanford vocabulary expansion method (mean resizing) - New embeddings initialized with the mean of existing embeddings - Both `lm_head` and embedding layers were resized - No gradient updates or training performed #### Training Hyperparameters - **Quantization**: 4-bit NF4 with double quantization - **Compute dtype**: bfloat16 - **Method**: Vocabulary expansion only (no fine-tuning) - **Embedding Initialization**: Mean resizing ## Evaluation ### Testing Data, Factors & Metrics #### Testing Data Initial testing was performed using: - **Pashto sample texts** to verify tokenization - **English prompts** to ensure base functionality - **Character-level inspection** to confirm tokenizer additions #### Metrics - **Vocabulary Size**: Successfully expanded from 32,000 to 32,010 - **Tokenization Accuracy**: All 47 characters are properly tokenized - **Generation Quality**: English generation preserved - **Memory Usage**: ~4GB VRAM for inference ### Results | Test | Result | Status | |------|--------|--------| | Pashto Tokenization | ✅ All 47 characters recognized | PASS | | English Generation | ✅ Original performance preserved | PASS | | Pashto Understanding | ❌ No semantic understanding | EXPECTED | | Pashto Response | ❌ Generates English | EXPECTED | | Model Loading (4-bit) | ✅ Successfully loads | PASS | | Memory Efficiency | ✅ ~4GB VRAM | PASS | ## Environmental Impact Carbon emissions were minimal as this model: - Did not require any training - Only performed vocabulary expansion and quantization - Used a single GPU for a few minutes - **Hardware Type:** NVIDIA A100 (or similar) - **Hours used:** < 1 hour - **Cloud Provider:** N/A (local) - **Compute Region:** N/A - **Carbon Emitted:** Negligible ## Technical Specifications ### Model Architecture and Objective - **Architecture:** Transformer decoder (Mistral-7B) - **Number of Layers:** 32 - **Hidden Size:** 4096 - **Attention Heads:** 32 - **Intermediate Size:** 14336 - **Vocabulary Size:** 32,010 (expanded from 32,000) - **Positional Encoding:** Rotary Position Embeddings (RoPE) ### Compute Infrastructure #### Hardware - CPU for embedding expansion - GPU for quantization and inference testing #### Software - **Transformers**: v4.31.0+ - **BitsAndBytes**: v0.41.0+ - **PyTorch**: v2.0.0+ - **Accelerate**: v0.20.0+ ## Model Card Authors - **Nassim JP** - Vocabulary expansion and quantization implementation - Original model: Mistral AI team ## Model Card Contact For questions or collaboration: - Hugging Face: [nassimjp](https://huggingface.co/nassimjp) ## Acknowledgments - Mistral AI for the base model - Stanford NLP for the vocabulary expansion method - Hugging Face for the transformers library - BitsAndBytes team for 4-bit quantization ## Glossary - **Vocabulary Expansion**: Adding new tokens to a model's tokenizer - **Mean Resizing**: Initializing new embeddings with the mean of existing embeddings - **4-bit NF4**: 4-bit Normal Float quantization method - **Fine-tuning**: Training a pre-trained model on domain-specific data ## Next Steps To make this model truly Pashto-capable, the next steps include: 1. **Collect Pashto corpus** (books, articles, web text) 2. **Pre-train on Pashto** data (continued pre-training) 3. **Fine-tune** for specific tasks (translation, QA, generation) 4. **Evaluate** on Pashto benchmarks 5. **Deploy** in production environments ## Citation If you use this model or the vocabulary expansion method, please cite: ```bibtex @misc{ghanam2024pashto, author = {Nassim JP}, title = {Ghanam-7B-Base-Pashto: Vocabulary Expansion for Pashto Language}, year = {2024}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/nassimjp/Ghanam-7B-Base-Pashto-v0.1}} } ``` ### Original Mistral-7B Citation ```bibtex @article{jiang2023mistral, title={Mistral 7B}, author={Jiang, Albert Q and others}, journal={arXiv preprint arXiv:2310.06825}, year={2023} } ``` --- ## Additional Information ### Usage Tips 1. **Tokenization**: The tokenizer now correctly handles all Pashto characters 2. **Memory**: 4-bit quantization allows running on consumer GPUs with 6-8GB VRAM 3. **Generation**: Use `do_sample=True` and `temperature=0.7` for diverse outputs 4. **Pad Token**: If using batch generation, set `pad_token=tokenizer.eos_token` ### Common Issues | Issue | Solution | |-------|----------| | Tokenizer missing characters | Ensure you're using the correct tokenizer from this repo | | Memory errors | Reduce `max_new_tokens` or use CPU offloading | | Poor generation quality | This is expected - the model needs fine-tuning | | Pashto text displayed as `[UNK]` | Tokenizer hasn't loaded properly - reload from the repo | --- **Model Status:** 🔬 Research Phase - Pre-fine-tuning checkpoint **Last Updated:** June 2026 ```