Files
aime-gen-qwen3-4b-v3/README.md
ModelHub XC f30a0cb94c 初始化项目,由ModelHub XC社区提供模型
Model: William2390401/aime-gen-qwen3-4b-v3
Source: Original Platform
2026-07-31 10:54:16 +08:00

67 lines
2.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
base_model: Qwen/Qwen3-4B-Instruct-2507
pipeline_tag: text-generation
language:
- en
tags:
- math
- competition-math
- aime
- problem-generation
---
# AIME-Style Problem Generator — Qwen3-4B (v3, merged)
Standalone 16-bit merge of [`aime-gen-qwen3-4b-lora-v3`](https://huggingface.co/William2390401/aime-gen-qwen3-4b-lora-v3)
into `Qwen/Qwen3-4B-Instruct-2507`. Generates **novel, difficulty-calibrated AIME-style problems**
from a bare one-line prompt (no system prompt, no few-shot). Trained on the
[companion SFT dataset](https://huggingface.co/datasets/William2390401/aime-gen-sft-v1).
**Thesis:** problem-*posing* failure in LLMs is a diversity deficit, not a reasoning deficit;
fine-tuning on a curated dataset instills the behavior a prompt can't reliably buy.
## Usage
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO = "William2390401/aime-gen-qwen3-4b-v3"
tok = AutoTokenizer.from_pretrained(REPO)
model = AutoModelForCausalLM.from_pretrained(REPO, torch_dtype="auto", device_map="auto")
msgs = [{"role": "user", "content":
"Write an AIME-style problem. Difficulty: late (problems 11-15). Topic: number theory."}]
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
out = model.generate(**tok(text, return_tensors="pt").to(model.device),
max_new_tokens=1600, do_sample=True, temperature=0.8, top_p=0.95)
print(tok.decode(out[0], skip_special_tokens=True))
```
Output format: `<problem>…</problem><solution>…</solution><answer>N</answer>` (integer 0–999).
## Results — tuned (bare prompt) vs base (full engineered prompt)
| Metric | Base (engineered) | Tuned (bare) | Δ |
|---|---|---|---|
| Format adherence | 28.9% | **63.9%** | +35.0 |
| Self-duplication (lower=better) | 70.6% | **18.3%** | −52.3 |
| Band accuracy | 60.0% | **64.3%** | +4.3 |
| Novelty vs corpus+train | 87.8% | 71.1% | −16.7 |
| Validity (solver consensus) | 47.2% | 12.2% | −35.0 |
**Win:** a fine-tuned 4B on a one-liner beats a fully-prompted base on format, diversity, and
calibration — the properties a dataset can encode.
**Honest limitation:** validity is **12%** — the model is a strong problem *stylist* but a weak
*verifier*; a 4B can't reliably solve the problems it poses (v2's higher 33% was inflated by
degenerate answer-0 collusion). Not fixable by data.
## Notes
- Merged from adapters trained against a 4-bit base (QLoRA). The **most faithful serving** is the
4-bit base + `-lora-v3` adapters (matches training); this merged 16-bit model is for convenience.
- Research/education use. Not a solver; answers are not verified beyond the strong-solver gate.
- Full analysis: `report_v3_analysis.md` / `BRAINLIFT_RESULTS.md` in the project repo.