初始化项目,由ModelHub XC社区提供模型

Model: OhhMoo/qwen05b-gsm8k-sft-instruct
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-14 20:46:40 +08:00
commit cc0def4d11
52 changed files with 758994 additions and 0 deletions

58
README.md Normal file
View File

@@ -0,0 +1,58 @@
---
license: apache-2.0
base_model: Qwen/Qwen2.5-0.5B-Instruct
datasets:
- openai/gsm8k
language:
- en
pipeline_tag: text-generation
library_name: transformers
tags:
- gsm8k
- math
- sft
- qwen2.5
---
# Qwen2.5-0.5B-Instruct — GSM8k SFT
Full supervised fine-tune of [`Qwen/Qwen2.5-0.5B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct)
on the [GSM8k](https://huggingface.co/datasets/openai/gsm8k) training split (7,473 worked
solutions), trained with [verl](https://github.com/volcengine/verl)'s `sft_trainer`.
## Results
| Metric | Value |
|---|---|
| GSM8k test accuracy (greedy, 1319 examples) | **36.2%** (477/1319) |
## Training
- Base model: `Qwen/Qwen2.5-0.5B-Instruct`
- Method: full fine-tune (no LoRA), bf16 mixed precision
- Data: GSM8k train, chat format (system + user); assistant target is the full worked solution ending in `#### N`
- Epochs: 3 (348 steps); AdamW, lr 1e-5 cosine, warmup ratio 0.1, weight decay 0.1, grad clip 1.0
- Max sequence length: 1024
## Prompt format
System prompt used during training; the model ends its answer with `#### <number>`:
> You are a math problem solver. Think step by step. You MUST end your response with '#### <number>' where <number> is the final numerical answer (digits only, no units or markdown).
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("OhhMoo/qwen05b-gsm8k-sft-instruct")
tok = AutoTokenizer.from_pretrained("OhhMoo/qwen05b-gsm8k-sft-instruct")
messages = [
{"role": "system", "content": "You are a math problem solver. Think step by step. You MUST end your response with '#### <number>' where <number> is the final numerical answer (digits only, no units or markdown)."},
{"role": "user", "content": "Natalia sold clips to 48 of her friends in April, and then she sold half as many clips in May. How many clips did Natalia sell altogether in April and May? Let's think step by step."},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
out = m.generate(ids, max_new_tokens=512, do_sample=False)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))
```