Model: Phantomcloak19/qwen2.5-3b-dpo Source: Original Platform
base_model, library_name, tags
| base_model | library_name | tags | |||
|---|---|---|---|---|---|
| Qwen/Qwen2.5-3B-Instruct | transformers |
|
qwen2.5-3b — DPO
Merged full-precision model after the DPO phase of the HorusLLM sequential training pipeline (SFT → DPO → Safety-GRPO).
| Field | Value |
|---|---|
| Base model | Qwen/Qwen2.5-3B-Instruct |
| Phase | DPO |
| Short name | qwen2.5-3b |
| Generated | 2026-07-01 19:29 UTC |
Description
Languages
Jinja
100%