Model: RexTRO111/Qwen3-4B-MegaR3ASONER-v1 Source: Original Platform
language, license, library_name, pipeline_tag, base_model, base_model_relation, tags, datasets
| language | license | library_name | pipeline_tag | base_model | base_model_relation | tags | datasets | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
other | transformers | text-generation | Qwen/Qwen3-4B-Thinking-2507 | finetune |
|
|
Qwen3-4B-MegaR3ASONER-v1
A full merged reasoning model created by merging the MegaR3ASONER LoRA into
Qwen/Qwen3-4B-Thinking-2507.
Model type
This repository contains the complete merged Transformers model, not only the
PEFT adapter. It can be loaded directly without PeftModel.
- Base model:
Qwen/Qwen3-4B-Thinking-2507 - Merge dtype: BF16
- Adapter source:
RexTRO111/Qwen3-4B-MegaR3ASONER-LoRA-v1 - Training hardware: NVIDIA A10G on Modal
- Merge hardware: NVIDIA A10G on Modal
Preliminary evaluation
On the first 100 examples selected by EleutherAI's gsm8k_cot task:
- Flexible extraction exact match: 88%
- Strict match: 83%
This was a limited 100-question run, not a full GSM8K score and not a controlled base-versus-fine-tune comparison.
Loading
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "RexTRO111/Qwen3-4B-MegaR3ASONER-v1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model.eval()
Limitations
- It may produce an incorrect intermediate thought before correcting itself.
- It can overthink simple prompts.
- Long reasoning traces increase latency and cost.
- Benchmark contamination has not been exhaustively ruled out.
- Verify answers before high-stakes use.
Licensing note
The Qwen base model and every training dataset retain their own licenses and upstream terms. Review all applicable terms before redistribution or commercial use.
Description
Languages
Jinja
100%