library_name, tags, language, license, pipeline_tag, base_model
library_name tags language license pipeline_tag base_model
transformers
pashto
mistral
4bit
nf4
vocabulary-expansion
afghanistan
pashto-language
ps
en
apache-2.0 text-generation mistralai/Mistral-7B-v0.1

Ghanam-7B-Base-Pashto-v0.1

Model Description

Ghanam-7B-Base-Pashto-v0.1 is an experimental base model that extends the Mistral-7B-v0.1 vocabulary to include 47 Pashto/Arabic characters. This is the first stage of creating a Pashto-capable language model, where the tokenizer has been modified and the embedding layer has been expanded to support Pashto script.

Key Features

  • Vocabulary Size: 32,010 tokens (original 32,000 + 47 Pashto characters)
  • Quantization: 4-bit NF4 with double quantization for efficient inference
  • Memory Efficient: ~4GB VRAM usage for inference
  • Pashto Script Support: Now recognizes all Pashto-specific characters including ښ، ږ، څ، ډ، ړ، ځ، ګ، ۍ، ې

Pashto Characters Added

ا ب پ ت ث ج چ ح خ د ذ ر ز ژ س ش ص ض ط ظ ع غ ف ق ک ګ گ ل م ن ڼ و ه ی ډ ړ ځ څ ښ ږ ۍ ې ء آ أ ؤ ئ

⚠️ Important Note

This model does not yet understand Pashto language semantics. It can:

  • Tokenize Pashto text correctly
  • Generate Pashto characters (if prompted in Pashto)
  • Understand Pashto meaning or context
  • Respond meaningfully in Pashto

The model still generates English text as it hasn't been fine-tuned on Pashto data.

Model Details

Model Description

This model is the first step in creating a Pashto language model based on Mistral-7B. The vocabulary was expanded using the Stanford vocabulary expansion method (mean resizing), where new embeddings were initialized using the mean of existing embeddings to minimize disruption to the model's performance.

  • Developed by: Nassim JP (Community Project)
  • Model type: Causal Language Model (Decoder-only Transformer)
  • Language(s): English (original) + Pashto script support (tokenization only)
  • License: Apache 2.0
  • Base Model: Mistral-7B-v0.1
  • Quantization: 4-bit NF4 (BitsAndBytes)
  • Vocabulary Expansion Method: Stanford mean resizing

Model Sources

Uses

Direct Use

You can use this model for:

  • Pashto tokenization and text preprocessing
  • Inference experiments with Pashto prompts (though responses will be in English)
  • As a starting point for fine-tuning on Pashto datasets
  • Research on vocabulary expansion for low-resource languages

Downstream Use

This model is intended to be fine-tuned on Pashto text data for:

  • Pashto language modeling
  • Pashto text generation
  • Pashto translation tasks
  • Pashto question answering

Out-of-Scope Use

  • Do not use this model for production Pashto applications without fine-tuning
  • Do not expect Pashto understanding or generation in Pashto
  • Not suitable for tasks requiring Pashto semantic comprehension
  • Not tested for bias, toxicity, or safety in Pashto context

Bias, Risks, and Limitations

Technical Limitations

  1. No Pashto Understanding: The model has only expanded vocabulary but has not learned Pashto semantics
  2. English-Only Knowledge: All pre-training knowledge remains in English
  3. Potential Performance Degradation: Vocabulary expansion may slightly impact English performance
  4. Untested on Pashto Tasks: No evaluation has been conducted on Pashto benchmarks

Recommendations

  • Use this model only as a foundation for Pashto fine-tuning
  • Validate carefully before using in any production environment
  • Expect English responses when prompting in Pashto (until fine-tuned)
  • Consider this a research checkpoint rather than a production model

How to Get Started with the Model

Loading the Model (4-bit)

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig

# Model path (local or from Hub)
model_path = "nassimjp/Ghanam-7B-Base-Pashto-v0.1"

# 4-bit quantization config
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_use_double_quant=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16
)

# Load model and tokenizer
model = AutoModelForCausalLM.from_pretrained(
    model_path,
    quantization_config=bnb_config,
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_path)

# Set padding token (optional)
tokenizer.pad_token = tokenizer.eos_token

# Test tokenization
pashto_text = "پښتونخوا کې ډېر ښکلي ښارونه او ځنګلونه شته"
tokens = tokenizer.encode(pashto_text)
print(f"Tokens: {tokens}")
print(f"Decoded: {tokenizer.decode(tokens)}")

Generating Text

# Prompt in Pashto (model will respond in English)
prompt = "سلام، څنګه یاست؟"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=50,
        temperature=0.7,
        do_sample=True,
        repetition_penalty=1.2
    )

response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)

Inspecting New Tokens

# Check if Pashto characters are recognized
pashto_chars = ["ښ", "ږ", "څ", "ډ", "ړ", "ځ", "ګ"]
for char in pashto_chars:
    token_id = tokenizer.encode(char)[0]
    token_str = tokenizer.decode([token_id])
    print(f"Character: {char} -> Token ID: {token_id} -> Decoded: {token_str}")

Training Details

Training Data

No training data was used for this model. The vocabulary was expanded without any fine-tuning on Pashto text. The model retains only the original Mistral-7B pre-training knowledge.

Training Procedure

Vocabulary Expansion Method

The model underwent a two-stage process:

Stage 1: Tokenizer Modification

  • Added 47 Pashto/Arabic characters to the original tokenizer
  • Preserved all original tokens (no replacements)
  • New vocabulary size: 32,010

Stage 2: Embedding Resizing

  • Used Stanford vocabulary expansion method (mean resizing)
  • New embeddings initialized with the mean of existing embeddings
  • Both lm_head and embedding layers were resized
  • No gradient updates or training performed

Training Hyperparameters

  • Quantization: 4-bit NF4 with double quantization
  • Compute dtype: bfloat16
  • Method: Vocabulary expansion only (no fine-tuning)
  • Embedding Initialization: Mean resizing

Evaluation

Testing Data, Factors & Metrics

Testing Data

Initial testing was performed using:

  • Pashto sample texts to verify tokenization
  • English prompts to ensure base functionality
  • Character-level inspection to confirm tokenizer additions

Metrics

  • Vocabulary Size: Successfully expanded from 32,000 to 32,010
  • Tokenization Accuracy: All 47 characters are properly tokenized
  • Generation Quality: English generation preserved
  • Memory Usage: ~4GB VRAM for inference

Results

Test Result Status
Pashto Tokenization All 47 characters recognized PASS
English Generation Original performance preserved PASS
Pashto Understanding No semantic understanding EXPECTED
Pashto Response Generates English EXPECTED
Model Loading (4-bit) Successfully loads PASS
Memory Efficiency ~4GB VRAM PASS

Environmental Impact

Carbon emissions were minimal as this model:

  • Did not require any training

  • Only performed vocabulary expansion and quantization

  • Used a single GPU for a few minutes

  • Hardware Type: NVIDIA A100 (or similar)

  • Hours used: < 1 hour

  • Cloud Provider: N/A (local)

  • Compute Region: N/A

  • Carbon Emitted: Negligible

Technical Specifications

Model Architecture and Objective

  • Architecture: Transformer decoder (Mistral-7B)
  • Number of Layers: 32
  • Hidden Size: 4096
  • Attention Heads: 32
  • Intermediate Size: 14336
  • Vocabulary Size: 32,010 (expanded from 32,000)
  • Positional Encoding: Rotary Position Embeddings (RoPE)

Compute Infrastructure

Hardware

  • CPU for embedding expansion
  • GPU for quantization and inference testing

Software

  • Transformers: v4.31.0+
  • BitsAndBytes: v0.41.0+
  • PyTorch: v2.0.0+
  • Accelerate: v0.20.0+

Model Card Authors

  • Nassim JP - Vocabulary expansion and quantization implementation
  • Original model: Mistral AI team

Model Card Contact

For questions or collaboration:

Acknowledgments

  • Mistral AI for the base model
  • Stanford NLP for the vocabulary expansion method
  • Hugging Face for the transformers library
  • BitsAndBytes team for 4-bit quantization

Glossary

  • Vocabulary Expansion: Adding new tokens to a model's tokenizer
  • Mean Resizing: Initializing new embeddings with the mean of existing embeddings
  • 4-bit NF4: 4-bit Normal Float quantization method
  • Fine-tuning: Training a pre-trained model on domain-specific data

Next Steps

To make this model truly Pashto-capable, the next steps include:

  1. Collect Pashto corpus (books, articles, web text)
  2. Pre-train on Pashto data (continued pre-training)
  3. Fine-tune for specific tasks (translation, QA, generation)
  4. Evaluate on Pashto benchmarks
  5. Deploy in production environments

Citation

If you use this model or the vocabulary expansion method, please cite:

@misc{ghanam2024pashto,
  author = {Nassim JP},
  title = {Ghanam-7B-Base-Pashto: Vocabulary Expansion for Pashto Language},
  year = {2024},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/nassimjp/Ghanam-7B-Base-Pashto-v0.1}}
}

Original Mistral-7B Citation

@article{jiang2023mistral,
  title={Mistral 7B},
  author={Jiang, Albert Q and others},
  journal={arXiv preprint arXiv:2310.06825},
  year={2023}
}

Additional Information

Usage Tips

  1. Tokenization: The tokenizer now correctly handles all Pashto characters
  2. Memory: 4-bit quantization allows running on consumer GPUs with 6-8GB VRAM
  3. Generation: Use do_sample=True and temperature=0.7 for diverse outputs
  4. Pad Token: If using batch generation, set pad_token=tokenizer.eos_token

Common Issues

Issue Solution
Tokenizer missing characters Ensure you're using the correct tokenizer from this repo
Memory errors Reduce max_new_tokens or use CPU offloading
Poor generation quality This is expected - the model needs fine-tuning
Pashto text displayed as [UNK] Tokenizer hasn't loaded properly - reload from the repo

Model Status: 🔬 Research Phase - Pre-fine-tuning checkpoint

Last Updated: June 2026

Description
Model synced from source: nassimjp/Ghanam-7B-Base-Pashto-v0.1
Readme 639 KiB