Files
KoshurAI_Tarjuma_v2/README.md
ModelHub XC 322d2af967 初始化项目,由ModelHub XC社区提供模型
Model: Faizaniqbal/KoshurAI_Tarjuma_v2
Source: Original Platform
2026-09-23 08:53:17 +08:00

135 lines
3.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
language:
- ks
- en
license: apache-2.0
base_model:
- Omarrran/koshur-kouter-ks-en_v1
library_name: transformers
pipeline_tag: text-generation
tags:
- kashmiri
- gemma3
- continual-pretraining
- low-resource
- language-model
---
# KoshurAI_Tarjuma_v2 — Kashmiri Continual Pretraining Base
> ⚠️ **This is a base language model, not a translation model.**
> For Kashmiri ↔ English translation, use the fine-tuned adapter:
> [`Faizaniqbal/KoshurAI_Tarjuma_v3`](https://huggingface.co/Faizaniqbal/KoshurAI_Tarjuma_v3)
---
## What is this?
`KoshurAI_Tarjuma_v2` is a **5B-parameter Gemma 3** model that has been
continually pretrained on **2.8 million tokens of native Kashmiri text**,
giving it deep knowledge of the Kashmiri language (Perso-Arabic script).
It serves as the **Stage 1 base** in the KoshurAI two-stage pipeline:
```
Omarrran/koshur-kouter-ks-en_v1 ← fine-tuned on Kashmiri–English pairs
↓ continual pretraining on 2.8M Kashmiri tokens
Faizaniqbal/KoshurAI_Tarjuma_v2 ← this model (language knowledge)
↓ SFT on 16,637 EN↔KS pairs (LoRA adapter)
Faizaniqbal/KoshurAI_Tarjuma_v3 ← final translation model
```
---
## Model Details
|
|
|
|
---
|
---
|
|
**
Author
**
|
Faizan Iqbal (
[
@Faizaniqbal
](
https://huggingface.co/Faizaniqbal
)
)
|
|
**
Base model
**
|
`Omarrran/koshur-kouter-ks-en_v1`
|
|
**
Architecture
**
|
Gemma3ForCausalLM (5B parameters)
|
|
**
Tensor type
**
|
BF16
|
|
**
Pretraining data
**
|
2.8M tokens of Kashmiri text
|
|
**
Languages
**
|
Kashmiri (ks · kas_Arab), English (en)
|
|
**
License
**
|
Apache-2.0
|
---
## Pretraining Corpus
The 2.8M token corpus was assembled from two sources:
- **InPage documents** — professionally published Kashmiri literature,
journalism, academic scholarship, and religious texts spanning multiple
decades, converted to Unicode via a custom InPage converter.
- **Native speaker translations** — texts translated into Kashmiri by
native speakers, providing natural human-authored language coverage.
---
## Full Model Lineage
```
google/gemma-3-4b-it (Google)
└─ sarvamai/sarvam-translate (Sarvam AI)
└─ Omarrran/koshur-kouter-ks-en_v1 (Malik & Nissar, 2026)
└─ Faizaniqbal/KoshurAI_Tarjuma_v2 ← this model
└─ Faizaniqbal/KoshurAI_Tarjuma_v3 (translation adapter)
```
---
## Citation
```bibtex
@misc{iqbal2026koshurai,
title = {KoshurAI v3: A Fine-Tuned Neural Machine Translation System
for Kashmiri--English Bidirectional Translation},
author = {Iqbal, Faizan},
year = {2026},
howpublished = {\url{https://huggingface.co/Faizaniqbal/KoshurAI_Tarjuma_v3}}
}
```
Please also cite the base model:
```bibtex
@misc{malik2026koshurkouter,
title = {Koshur Kouter KS-EN v1: A Merged QLoRA Kashmiri--English Translation Model},
author = {Malik, Haq Nawaz and Nissar, Nahfid},
year = {2026},
howpublished = {\url{https://huggingface.co/Omarrran/koshur-kouter-ks-en_v1}}
}
```