Files
Qwen3-4B-MegaR3ASONER-v1/README.md
ModelHub XC 9dcfb4449f 初始化项目,由ModelHub XC社区提供模型
Model: RexTRO111/Qwen3-4B-MegaR3ASONER-v1
Source: Original Platform
2026-07-23 11:32:15 +08:00

2.0 KiB

language, license, library_name, pipeline_tag, base_model, base_model_relation, tags, datasets
language license library_name pipeline_tag base_model base_model_relation tags datasets
en
other transformers text-generation Qwen/Qwen3-4B-Thinking-2507 finetune
qwen
qwen3
reasoning
chain-of-thought
text-generation
bf16
merged
HuggingFaceH4/Bespoke-Stratos-17k
nvidia/OpenMathReasoning
nvidia/OpenCodeReasoning
nvidia/OpenScienceReasoning-2
lordx64/reasoning-distill-opus-4-7-max-sft

Qwen3-4B-MegaR3ASONER-v1

A full merged reasoning model created by merging the MegaR3ASONER LoRA into Qwen/Qwen3-4B-Thinking-2507.

Model type

This repository contains the complete merged Transformers model, not only the PEFT adapter. It can be loaded directly without PeftModel.

  • Base model: Qwen/Qwen3-4B-Thinking-2507
  • Merge dtype: BF16
  • Adapter source: RexTRO111/Qwen3-4B-MegaR3ASONER-LoRA-v1
  • Training hardware: NVIDIA A10G on Modal
  • Merge hardware: NVIDIA A10G on Modal

Preliminary evaluation

On the first 100 examples selected by EleutherAI's gsm8k_cot task:

  • Flexible extraction exact match: 88%
  • Strict match: 83%

This was a limited 100-question run, not a full GSM8K score and not a controlled base-versus-fine-tune comparison.

Loading

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "RexTRO111/Qwen3-4B-MegaR3ASONER-v1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
model.eval()

Limitations

  • It may produce an incorrect intermediate thought before correcting itself.
  • It can overthink simple prompts.
  • Long reasoning traces increase latency and cost.
  • Benchmark contamination has not been exhaustively ruled out.
  • Verify answers before high-stakes use.

Licensing note

The Qwen base model and every training dataset retain their own licenses and upstream terms. Review all applicable terms before redistribution or commercial use.