---
language:
- en
- th
license: apache-2.0
pipeline_tag: translation
tags:
- translation
- thai
- english
- gemma-3
- typhoon
- machine-translation
- bilingual
- sft
base_model: typhoon-ai/typhoon2.1-gemma3-4b
---
# πΉππΊπΈ Typhoon 2.1 Gemma 3 KordTranslate EN-TH 4B
**Typhoon 2.1 Gemma 3 KordTranslate EN-TH 4B** is a bilingual English β Thai translation model fine-tuned from **`typhoon-ai/typhoon2.1-gemma3-4b`** for high-quality machine translation between Thai and English.
The model is instruction-tuned specifically for translation tasks using bilingual translation datasets while preserving the strong language understanding inherited from the Typhoon base model.
---
# β¨ Features
- πΊπΈ English β Thai translation
- πΉπ Thai β English translation
- Instruction-tuned for translation
- Built on **Typhoon2.1 Gemma 3 4B**
- Lightweight and suitable for local inference
---
| Item | Value |
|------|-------|
| Model | `KordAI/Typhoon-Gemma3-KordTranslate-EN-TH-4B` |
| Base Model | `typhoon-ai/typhoon2.1-gemma3-4b` |
| Architecture | Gemma 3 (4B) |
| Task | Machine Translation |
| Languages | English, Thai |
| Fine-tuning | Supervised Fine-Tuning (SFT) |
---
This model is optimized for:
- English β Thai translation
- Thai β English translation
- Localization
- Educational applications
- Chatbot translation
- API integration
- Batch document translation
For the best performance, use the prompt templates below.
---
# π Prompt Templates
For optimal translation quality, structure your inputs exactly as shown below depending on the translation direction.
## English β Thai
```text
Translate the following English text to Thai.
### English:
[Your English text here]
```
## Thai β English
```text
Translate the following Thai text to English.
### Thai:
[Your Thai text here]
```
---
# π Inference Code (Python / Transformers)
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
# Hugging Face repository
model_name = "KordAI/Typhoon-Gemma3-KordTranslate-EN-TH-4B"
# Load tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
prompt = """
Translate the following English text to Thai.
### English:
But friends, I am so gay, that if I had a wife, I would encourage her to cheat on me.
"""
messages = [
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
generated_ids = model.generate(
**model_inputs,
max_new_tokens=512
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):]
translation = tokenizer.decode(
output_ids,
skip_special_tokens=True
).strip()
print(translation)
```
---
# ποΈ Training
This model is fine-tuned from **`typhoon-ai/typhoon2.1-gemma3-4b`** using supervised instruction tuning on bilingual EnglishβThai translation datasets.
The model is trained to perform direct translation between:
- πΊπΈ English β Thai
- πΉπ Thai β English
while maintaining natural, fluent, and context-aware outputs.
---
# β οΈ Limitations
- Translation quality may decrease on highly specialized domains such as medical, legal, or scientific texts.
- Proper nouns and newly coined terminology may require manual verification.
- The model may occasionally hallucinate or paraphrase ambiguous inputs.
- Human review is recommended for production-critical translations.
---
Special thanks to:
- **Typhoon AI** for the excellent **`typhoon2.1-gemma3-4b`** base model.
- **KordAI** for developing and fine-tuning the translation model.
- **Unsloth** for enabling efficient supervised fine-tuning.
- **llama.cpp** for the GGUF format and inference ecosystem.
- The open-source AI community for advancing multilingual language models.
---
# π Citation
```bibtex
@misc{kordtranslate2026,
title={Typhoon Gemma3 KordTranslate EN-TH 4B},
author={KordAI},
year={2026},
publisher={Hugging Face},
howpublished={https://huggingface.co/KordAI/Typhoon-Gemma3-KordTranslate-EN-TH-4B}
}
```