base_model, library_name, tags
base_model library_name tags
Qwen/Qwen2.5-3B-Instruct transformers
horus-llm
dpo
qwen2.5-3b

qwen2.5-3b — DPO

Merged full-precision model after the DPO phase of the HorusLLM sequential training pipeline (SFT → DPO → Safety-GRPO).

Field Value
Base model Qwen/Qwen2.5-3B-Instruct
Phase DPO
Short name qwen2.5-3b
Generated 2026-07-01 19:29 UTC
Description
Model synced from source: Phantomcloak19/qwen2.5-3b-dpo
Readme 35 KiB
Languages
Jinja 100%