Files

25 lines
657 B
Markdown
Raw Permalink Normal View History

---
base_model: BillyWang1/qwen2.5-3b-base-tool-n1-sft
library_name: transformers
pipeline_tag: text-generation
tags:
- qwen2
- tool-use
- grpo
- verl
- tool-n1
---
# qwen2.5-3b-base-tool-n1-grpo
This is a Hugging Face export of the GRPO-trained Tool-N1 Qwen2.5-3B checkpoint:
- Base model: `BillyWang1/qwen2.5-3b-base-tool-n1-sft`
- Training script: `gflowrl/tool-call/scripts/qwen_3b_grpo.sh`
- Checkpoint: `gflowrl/tool-call/ckpt/qwen_grpo_3b_seed54321/global_step_280/actor`
- Seed: `54321`
- Step: `280`
- Export format: verl FSDP shards merged to Transformers safetensors
The model is intended for Tool-N1-style function/tool calling experiments.