初始化项目,由ModelHub XC社区提供模型
Model: Phantomcloak19/qwen2.5-3b-dpo Source: Original Platform
This commit is contained in:
20
README.md
Normal file
20
README.md
Normal file
@@ -0,0 +1,20 @@
|
||||
---
|
||||
base_model: Qwen/Qwen2.5-3B-Instruct
|
||||
library_name: transformers
|
||||
tags:
|
||||
- horus-llm
|
||||
- dpo
|
||||
- qwen2.5-3b
|
||||
---
|
||||
|
||||
# qwen2.5-3b — DPO
|
||||
|
||||
Merged full-precision model after the **DPO** phase of the
|
||||
HorusLLM sequential training pipeline (SFT → DPO → Safety-GRPO).
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Base model | `Qwen/Qwen2.5-3B-Instruct` |
|
||||
| Phase | DPO |
|
||||
| Short name | qwen2.5-3b |
|
||||
| Generated | 2026-07-01 19:29 UTC |
|
||||
Reference in New Issue
Block a user