初始化项目,由ModelHub XC社区提供模型
Model: BillyWang1/qwen2.5-3b-base-tool-n1-grpo Source: Original Platform
This commit is contained in:
24
README.md
Normal file
24
README.md
Normal file
@@ -0,0 +1,24 @@
|
||||
---
|
||||
base_model: BillyWang1/qwen2.5-3b-base-tool-n1-sft
|
||||
library_name: transformers
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- qwen2
|
||||
- tool-use
|
||||
- grpo
|
||||
- verl
|
||||
- tool-n1
|
||||
---
|
||||
|
||||
# qwen2.5-3b-base-tool-n1-grpo
|
||||
|
||||
This is a Hugging Face export of the GRPO-trained Tool-N1 Qwen2.5-3B checkpoint:
|
||||
|
||||
- Base model: `BillyWang1/qwen2.5-3b-base-tool-n1-sft`
|
||||
- Training script: `gflowrl/tool-call/scripts/qwen_3b_grpo.sh`
|
||||
- Checkpoint: `gflowrl/tool-call/ckpt/qwen_grpo_3b_seed54321/global_step_280/actor`
|
||||
- Seed: `54321`
|
||||
- Step: `280`
|
||||
- Export format: verl FSDP shards merged to Transformers safetensors
|
||||
|
||||
The model is intended for Tool-N1-style function/tool calling experiments.
|
||||
Reference in New Issue
Block a user