2.0 KiB
2.0 KiB
language, license, library_name, pipeline_tag, base_model, base_model_relation, tags, datasets
| language | license | library_name | pipeline_tag | base_model | base_model_relation | tags | datasets | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
other | transformers | text-generation | Qwen/Qwen3-4B-Thinking-2507 | finetune |
|
|
Qwen3-4B-MegaR3ASONER-v1
A full merged reasoning model created by merging the MegaR3ASONER LoRA into
Qwen/Qwen3-4B-Thinking-2507.
Model type
This repository contains the complete merged Transformers model, not only the
PEFT adapter. It can be loaded directly without PeftModel.
- Base model:
Qwen/Qwen3-4B-Thinking-2507 - Merge dtype: BF16
- Adapter source:
RexTRO111/Qwen3-4B-MegaR3ASONER-LoRA-v1 - Training hardware: NVIDIA A10G on Modal
- Merge hardware: NVIDIA A10G on Modal
Preliminary evaluation
On the first 100 examples selected by EleutherAI's gsm8k_cot task:
- Flexible extraction exact match: 88%
- Strict match: 83%
This was a limited 100-question run, not a full GSM8K score and not a controlled base-versus-fine-tune comparison.
Loading
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "RexTRO111/Qwen3-4B-MegaR3ASONER-v1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model.eval()
Limitations
- It may produce an incorrect intermediate thought before correcting itself.
- It can overthink simple prompts.
- Long reasoning traces increase latency and cost.
- Benchmark contamination has not been exhaustively ruled out.
- Verify answers before high-stakes use.
Licensing note
The Qwen base model and every training dataset retain their own licenses and upstream terms. Review all applicable terms before redistribution or commercial use.