base_model, library_name, tags
base_model library_name tags
Qwen/Qwen2.5-3B-Instruct transformers
dpo-grpo
qwen2.5-3b

qwen2.5-3b — DPO-GRPO

Merged full-precision model after the DPO-GRPO phase of the LLMPR sequential training pipeline (SFT → DPO → Safety-GRPO).

Field Value
Base model Qwen/Qwen2.5-3B-Instruct
Phase DPO-GRPO
Short name qwen2.5-3b
Generated 2026-06-30 03:11 UTC
Description
Model synced from source: Phantomcloak19/qwen2.5-3b-dpo-grpo
Readme 26 KiB
Languages
Jinja 100%