3.2 KiB
language, license, pipeline_tag, tags, base_model
| language | license | pipeline_tag | tags | base_model | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
|
apache-2.0 | text-generation |
|
|
Gaon-1.7B v2 Instruct (가온)
A bilingual (Korean + English) 1.7B chat model trained entirely from scratch — architecture, data pipeline, pretraining, and instruction tuning — by one person on borrowed idle GPUs, for $0. Gaon is pure Korean for "center/core."
- Architecture: Qwen3-1.7B-compatible (28L, hidden 2048, GQA 16Q/8KV, QK-Norm,
SwiGLU, tied embeddings, vocab 151,936). Loads with
Qwen3ForCausalLM. - Pretraining: 34B tokens (FineWeb-Edu EN + FineWeb-2 KO + ~20% Python code), 4× B200 FSDP, final loss 1.96.
- Instruction tuning: sequence-level distillation from Qwen2.5-7B-Instruct (Apache-2.0), ~9k KO/EN instructions, SFT loss 1.26.
- Code & full tech report: https://github.com/k08200/gaon
Run locally: ollama run hf.co/k08200/gaon-1.7b-v2-instruct-GGUF — or community
GGUF quants (12 sizes, 516 MB – 3.4 GB, by @mradermacher):
static ·
imatrix
Live demo: KO↔EN translator in your browser (WebGPU)
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
m = "k08200/gaon-1.7b-v2-instruct"
tok = AutoTokenizer.from_pretrained(m)
model = AutoModelForCausalLM.from_pretrained(m, torch_dtype=torch.bfloat16).eval()
msgs = [{"role": "user", "content": "한국의 수도는 어디인가요?"}]
prompt = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
enc = tok(prompt, return_tensors="pt")
out = model.generate(**enc, max_new_tokens=200, do_sample=True, temperature=0.7)
print(tok.decode(out[0][enc.input_ids.shape[1]:], skip_special_tokens=True))
Benchmarks
Base model (v2), lm-eval-harness, 0-shot, same settings for both:
| Benchmark | Gaon-1.7B v2 | Qwen3-1.7B-Base | Random |
|---|---|---|---|
| MMLU (English knowledge) | 25.1 | 62.6 | 25.0 |
| KMMLU (Korean knowledge) | 22.3 | 35.5 | 25.0 |
| HAERAE (Korean culture/lexis) | 19.9 | 46.8 | ~20 |
| KoBEST (Korean understanding) | 51.5 | — | ~50 |
This is the most honest number in the project: Gaon is fluent in two languages, follows instructions, and translates, yet scores at chance on knowledge — because Qwen3-1.7B (same architecture, same size) trained on ~36T tokens, roughly 1000× our 34B. Linguistic competence emerges in tens of billions of tokens; world knowledge needs trillions. A continued-pretraining run (96B Korean-heavy tokens) lifts KMMLU to ~27–29 then plateaus — see the tech report §6.
Honest limitations
A from-scratch credential and research artifact, not a frontier competitor. Knowledge benchmarks at chance level (above); coding and strict-format tasks are weak. Korean/English chat, explanation, and simple translation work well. Best used as a reproducible small-LLM pipeline and a base for vertical fine-tuning on domain data.