A style-preserving grammar correction model based on SmolLM-135M, trained with SFT + DPO to make minimal, targeted corrections while preserving your original writing style.
Why This Model?
Unlike large language models (GPT, Claude, etc.) that tend to rewrite entire sentences, this model makes minimal, targeted corrections - fixing only grammatical errors while preserving your vocabulary, tone, and voice. Perfect for:
ESL/EFL education: Help learners without changing their ideas
Professional communications: Keep your authentic voice
Key Features
Minimal corrections: Fixes only grammatical errors, doesn't rewrite your sentences
Style preservation: Maintains your vocabulary, tone, and voice
Small & efficient: Only 135M parameters (~500MB) - runs on CPU!
BLEU score: ~0.50 on grammar correction benchmarks
Quick Start
fromtransformersimportAutoModelForCausalLM,AutoTokenizermodel=AutoModelForCausalLM.from_pretrained("DanJZY/SmolLM-135M-GEC-SFT-DPO")tokenizer=AutoTokenizer.from_pretrained("DanJZY/SmolLM-135M-GEC-SFT-DPO")text="As the number of people grows, the need of habitable environment is essential."inputs=tokenizer(f"Fix grammar: {text}",return_tensors="pt")outputs=model.generate(**inputs,max_new_tokens=100)print(tokenizer.decode(outputs[0],skip_special_tokens=True))
Example: Style-Preserving vs Over-Correction
Original (with error):
"As the number of people grows, the need of habitable environment is essential."
✅ Our Model (Style-Preserving):
"As the number of people grows, the need for a habitable environment is essential."
↑
Only fixes "of" → "for a"
❌ Typical Model (Over-Correction):
"As population growth continues, the necessity for a habitable environment becomes essential."
↑
Completely rewrites: changes vocabulary, structure, and tone