Files
wongwian-micro-instruct/README.md
ModelHub XC c7cc9fd4a7 初始化项目,由ModelHub XC社区提供模型
Model: wongwian-org/wongwian-micro-instruct
Source: Original Platform
2026-06-13 17:32:40 +08:00

243 lines
9.4 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: mit
language:
- th
- en
pipeline_tag: text-generation
library_name: transformers
tags:
- thai
- instruct
- causal-lm
- llama-architecture
- sentencepiece
- wongwian
base_model: wongwian-org/wongwian-micro-instruct
model_type: llama
---
<p align="center">
<img src="./ProfilePic.png" width="20%"/>
</p>
# Wongwian Micro Instruct — 272M
**The first model in the Wongwian series.** A compact 272M-parameter Thai-centric language model trained entirely from scratch on 40B tokens of Thai-dominant data and then instruction-fine-tuned. No base weights from existing open models were used at any stage.
> _"Cultural & Localization AI for the world"_ — Wongwian
---
# Wongwian Micro Instruct — 272M (ภาษาไทย)
**โมเดลตัวแรกในซีรีส์ Wongwian** เป็นโมเดลภาษาขนาดเล็ก 272 ล้านพารามิเตอร์ที่เน้นภาษาไทยเป็นหลัก **ฝึกขึ้นมาจากศูนย์ (train from scratch) บนข้อมูล 40 พันล้าน token** แล้วผ่านการ instruction fine-tuning โดยไม่ได้นำ open weights ของโมเดลอื่นมาต่อยอดแต่อย่างใด
> _"Cultural & Localization AI for the world"_ — Wongwian
---
## ภาพรวมโมเดล
| คุณสมบัติ | ค่า |
|---|---|
| **ตระกูล** | Wongwian |
| **ซีรีส์** | Micro |
| **เวอร์ชัน** | Instruct v1 (step 450) |
| **สถาปัตยกรรม** | LlamaForCausalLM |
| **จำนวนพารามิเตอร์** | ~272 ล้าน |
| **Context length** | 2,048 tokens |
| **Precision** | bfloat16 |
| **Token ที่ใช้ pre-train** | ~40 พันล้าน |
| **ภาษาหลัก** | ไทย 🇹🇭 |
| **ภาษารอง** | อังกฤษ 🇬🇧 |
| **License** | MIT |
---
## เกี่ยวกับ Wongwian
Wongwian คือโครงการวิจัย AI ภาษาไทยที่มุ่งสร้าง **โมเดลภาษา AI ที่มีประสิทธิภาพสูงและเข้าใจบริบทเชิงวัฒนธรรม** สำหรับภาษาไทยและภาษาอื่น ๆ ที่ขาดแคลนทรัพยากร โครงการนี้แสดงให้เห็นว่าโมเดล AI ที่มีคุณภาพสูงไม่จำเป็นต้องมีพารามิเตอร์มหาศาล โมเดลขนาดเล็กที่ออกแบบอย่างพิถีพิถันและใช้ข้อมูลที่เหมาะสม สามารถตอบโจทย์การใช้งานจริงได้อย่างมีประสิทธิผล
**วิสัยทัศน์:** _Cultural & Localization AI for the world_ — สร้าง AI ที่เข้าใจความละเอียดอ่อนทางภาษา วัฒนธรรม และบริบทของแต่ละชุมชนอย่างลึกซึ้ง
**จุดแข็งของโมเดลนี้:**
- **ฝึกจากศูนย์ (from scratch)** บนข้อมูลภาษาไทยเป็นหลัก ไม่ได้ต่อยอดจากโมเดลสาธารณะใดทั้งสิ้น
- **SentencePiece Unigram tokenizer** (32K vocab) ออกแบบมาเฉพาะสำหรับโครงสร้างภาษาไทยและข้อความผสมไทย-อังกฤษ
- สร้างบนเฟรมเวิร์ค [OLMo-core](https://github.com/allenai/OLMo-core) แบบ open-source
- ผ่านการ instruction fine-tuning (SFT) ด้วยข้อมูลบทสนทนาภาษาไทย
- รูปแบบ prompt สะท้อนการโต้ตอบผู้ช่วยภาษาไทยอย่างเป็นธรรมชาติ
---
## Model Overview
| Property | Value |
|---|---|
| **Model family** | Wongwian |
| **Series** | Micro |
| **Version** | Instruct v1 (step 450) |
| **Architecture** | LlamaForCausalLM |
| **Parameters** | ~272 M |
| **Context length** | 2 048 tokens |
| **Precision** | bfloat16 |
| **Pre-train tokens** | ~40 B |
| **Primary language** | Thai 🇹🇭 |
| **Secondary language** | English 🇬🇧 |
| **License** | MIT |
---
## About Wongwian
Wongwian is a Thai AI research initiative focused on building **efficient, culturally-grounded language models** for Thai and other under-resourced languages. The project demonstrates that high-quality language AI does not require massive parameter counts — a carefully designed small model, trained on the right data, can serve real-world use cases effectively.
**Vision:** _Cultural & Localization AI for the world_ — delivering language models that deeply understand the cultural, linguistic, and contextual nuances of each community they serve.
**What makes this model different:**
- Trained **from scratch** on Thai-dominant data — not a fine-tune of any existing public model
- Custom **SentencePiece Unigram** tokenizer (32K vocab) designed for Thai morphology and mixed Thai-English text
- Built on the open [OLMo-core](https://github.com/allenai/OLMo-core) training framework
- Full instruction fine-tuning (SFT) with Thai conversation data
- Chat prompt format mirrors natural Thai assistant interaction
---
## Quick Start
### pip install
```bash
pip install transformers sentencepiece
```
### Chat generation (recommended)
```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "wongwian-org/wongwian-micro-instruct"
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model.eval()
messages = [
{"role": "user", "content": "สวัสดี อธิบาย AI ให้ฟังหน่อย"},
]
input_ids = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
)
with torch.no_grad():
output = model.generate(
input_ids.to(model.device),
max_new_tokens=200,
do_sample=True,
temperature=0.7,
top_p=0.9,
repetition_penalty=1.3,
no_repeat_ngram_size=4,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.pad_token_id,
)
new_tokens = output[0, input_ids.shape[1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=True))
```
### Multi-turn conversation
```python
history = []
def chat(user_message: str, system_prompt: str | None = None) -> str:
messages = []
if system_prompt:
messages.append({"role": "system", "content": system_prompt})
messages.extend(history)
messages.append({"role": "user", "content": user_message})
input_ids = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
)
with torch.no_grad():
output = model.generate(
input_ids.to(model.device),
max_new_tokens=200,
do_sample=True,
temperature=0.7,
top_p=0.9,
repetition_penalty=1.3,
no_repeat_ngram_size=4,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.pad_token_id,
)
reply = tokenizer.decode(output[0, input_ids.shape[1]:], skip_special_tokens=True).strip()
history.append({"role": "user", "content": user_message})
history.append({"role": "assistant", "content": reply})
return reply
print(chat("สวัสดี คุณคือใคร?"))
print(chat("แล้วคุณทำอะไรได้บ้าง?"))
```
### Chat template format
The model uses the following plain-text prompt format (applied automatically by `apply_chat_template`):
```
<s>system: <system_prompt>
ผู้ใช้: <user_turn>
ผู้ช่วย: <assistant_turn></s>
ผู้ใช้: <user_turn>
ผู้ช่วย:
```
---
## Prompt Engineering Tips
| Tip | Detail |
|---|---|
| **Temperature** | 0.6 0.8 for natural Thai conversation |
| **Repetition penalty** | 1.2 1.4 — recommended; prevents looping |
| **no_repeat_ngram_size** | 4 — prevents phrase repetition |
| **max_new_tokens** | 150 300 for chat; 512 for long-form |
| **System prompt** | Optional but improves role adherence |
---
## Limitations
- **Size:** At 272M parameters, this model is suited for assistive tasks and demonstrations, not advanced reasoning or complex long-form analysis.
- **Hallucination:** Like all language models, the model may produce inaccurate information. Always verify critical outputs.
- **Context length:** Maximum 2 048 tokens per call; performance may degrade near the limit.
- **Bias:** The model may reflect biases present in the training corpus.
---
## Citation
```bibtex
@misc{wongwian2026micro,
title = {Wongwian Micro Instruct: A Thai-Centric Small Language Model Trained from Scratch},
author = {Wongwian AI Research},
year = {2026},
url = {https://huggingface.co/wongwian-org/wongwian-micro-instruct}
}
```
---