21 lines
420 B
Markdown
21 lines
420 B
Markdown
---
|
|
base_model: Qwen/Qwen2.5-3B-Instruct
|
|
library_name: transformers
|
|
tags:
|
|
- horus-llm
|
|
- dpo
|
|
- qwen2.5-3b
|
|
---
|
|
|
|
# qwen2.5-3b — DPO
|
|
|
|
Merged full-precision model after the **DPO** phase of the
|
|
HorusLLM sequential training pipeline (SFT → DPO → Safety-GRPO).
|
|
|
|
| Field | Value |
|
|
|---|---|
|
|
| Base model | `Qwen/Qwen2.5-3B-Instruct` |
|
|
| Phase | DPO |
|
|
| Short name | qwen2.5-3b |
|
|
| Generated | 2026-07-01 19:29 UTC |
|