25 lines
657 B
Markdown
25 lines
657 B
Markdown
---
|
|
base_model: BillyWang1/qwen2.5-3b-base-tool-n1-sft
|
|
library_name: transformers
|
|
pipeline_tag: text-generation
|
|
tags:
|
|
- qwen2
|
|
- tool-use
|
|
- grpo
|
|
- verl
|
|
- tool-n1
|
|
---
|
|
|
|
# qwen2.5-3b-base-tool-n1-grpo
|
|
|
|
This is a Hugging Face export of the GRPO-trained Tool-N1 Qwen2.5-3B checkpoint:
|
|
|
|
- Base model: `BillyWang1/qwen2.5-3b-base-tool-n1-sft`
|
|
- Training script: `gflowrl/tool-call/scripts/qwen_3b_grpo.sh`
|
|
- Checkpoint: `gflowrl/tool-call/ckpt/qwen_grpo_3b_seed54321/global_step_280/actor`
|
|
- Seed: `54321`
|
|
- Step: `280`
|
|
- Export format: verl FSDP shards merged to Transformers safetensors
|
|
|
|
The model is intended for Tool-N1-style function/tool calling experiments.
|