Model: BillyWang1/qwen2.5-3b-base-tool-n1-grpo Source: Original Platform
base_model, library_name, pipeline_tag, tags
| base_model | library_name | pipeline_tag | tags | |||||
|---|---|---|---|---|---|---|---|---|
| BillyWang1/qwen2.5-3b-base-tool-n1-sft | transformers | text-generation |
|
qwen2.5-3b-base-tool-n1-grpo
This is a Hugging Face export of the GRPO-trained Tool-N1 Qwen2.5-3B checkpoint:
- Base model:
BillyWang1/qwen2.5-3b-base-tool-n1-sft - Training script:
gflowrl/tool-call/scripts/qwen_3b_grpo.sh - Checkpoint:
gflowrl/tool-call/ckpt/qwen_grpo_3b_seed54321/global_step_280/actor - Seed:
54321 - Step:
280 - Export format: verl FSDP shards merged to Transformers safetensors
The model is intended for Tool-N1-style function/tool calling experiments.
Description
Languages
Jinja
100%