Model: Phantomcloak19/qwen2.5-3b-dpo-grpo Source: Original Platform
base_model, library_name, tags
| base_model | library_name | tags | ||
|---|---|---|---|---|
| Qwen/Qwen2.5-3B-Instruct | transformers |
|
qwen2.5-3b — DPO-GRPO
Merged full-precision model after the DPO-GRPO phase of the LLMPR sequential training pipeline (SFT → DPO → Safety-GRPO).
| Field | Value |
|---|---|
| Base model | Qwen/Qwen2.5-3B-Instruct |
| Phase | DPO-GRPO |
| Short name | qwen2.5-3b |
| Generated | 2026-06-30 03:11 UTC |
Description
Languages
Jinja
100%