GRM-1.5b is a general-purpose reasoning-focused 1.5B model fine-tuned to improve multi-domain reasoning (math, logic, coding, and broad problem-solving). It is designed to be a strong, lightweight “daily driver” for general reasoning tasks and as a solid base for further fine-tuning.
Key features
Dedicated reasoning behavior for general tasks (stepwise problem solving, better consistency).
Small & efficient (1.5B) — practical for local inference and experimentation.
Multi-domain mixture: reasoning + code + math + (some) medical reasoning data.
Fine-tune friendly: intended as a good starting point for your own SFT/GRPO/DPO pipelines.