--- license: other license_name: qwen-research license_link: https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE language: - en base_model: - Qwen/Qwen2.5-VL-3B-Instruct datasets: - maveryn/trace pipeline_tag: image-text-to-text library_name: transformers tags: - multimodal - visual-reasoning - reinforcement-learning - grpo - rlvr - trace --- # TRACE Qwen2.5-VL 3B TRACE Qwen2.5-VL 3B is a GRPO checkpoint derived from [`Qwen/Qwen2.5-VL-3B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct) and trained on 64,000 grounded visual-reasoning examples spanning 1,000 tasks from [`maveryn/trace`](https://huggingface.co/datasets/maveryn/trace). This repository contains the merged step-500 inference checkpoint. [Paper](https://arxiv.org/abs/2607.19790) · [Project page](https://maveryn.github.io/trace/) · [GitHub](https://github.com/maveryn/trace) · [Collection](https://huggingface.co/collections/maveryn/trace-6a604291b4be4ed6399b9f24) · [Training configuration](https://github.com/maveryn/trace/blob/rlvr/rlvr/configs/trace-qwen2.5-vl-3b.yaml) · [Evaluation suite](https://github.com/maveryn/trace/tree/rlvr/rlvr/evaluation/trace_eval) · [Run artifacts](https://huggingface.co/datasets/maveryn/trace-eval-runs) ## Training and provenance | Field | Value | | --- | --- | | Base model | `Qwen/Qwen2.5-VL-3B-Instruct@66285546d2b821cf421d4f5eb2576359d3770cd3` | | Dataset | `maveryn/trace@4e5b54361360296a855542b40cfd8b7f81b355fe` | | Method | GRPO, 500 steps, global/rollout batch 128, 8 responses per prompt | | Optimization | AdamW, learning rate `1e-6`, constant schedule, KL disabled | | Reward | `0.95 × exact JSON answer + 0.05 × valid JSON format`; annotation reward `0` | | Reference run | 8 NVIDIA H100 80GB GPUs | | Released artifact | Merged step-500 inference weights | The released training profile reads `prompt_answer`, scores `answer_gt`, and does not use the advisory `trace_supervision_mode` column. The consumed fields, embedded image bytes, and row order in the current dataset release were verified identical to the original training input in the public [equivalence receipt](https://github.com/maveryn/trace/blob/rlvr/rlvr/dataset_equivalence.v1.json). The canonical output ends with `{"answer": ...}`. Exact hashes, source revision, and run provenance are recorded in [`trace_training_provenance.json`](trace_training_provenance.json) and [`.trace_model_revision.json`](.trace_model_revision.json). The repository does not include optimizer, scheduler, trainer, or FSDP state for continuation. ## Evaluation `trace_eval_v1` evaluates 24 external benchmarks and 32,805 examples per model and decoding seed. Scores below are the unweighted macro mean of the 24 benchmark percentages, reported as mean ± sample standard deviation across seeds 42, 43, and 44. | Model | Overall score | | --- | ---: | | Qwen2.5-VL-3B-Instruct | 39.34 ± 0.63 | | TRACE Qwen2.5-VL 3B | **42.85 ± 0.39** | | Paired improvement | **+3.51 ± 0.25** | Per-benchmark scores, model revisions, aggregation, and benchmark provenance are available in the public [`results.json`](https://github.com/maveryn/trace/blob/rlvr/rlvr/evaluation/trace_eval/results.json) and [evaluation documentation](https://github.com/maveryn/trace/blob/rlvr/rlvr/evaluation/trace_eval/README.md). ## Usage ```python from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration model_id = "maveryn/trace-qwen2.5-vl-3b" revision = "2ec2374d5c219e6b12e26bda93d3b3adeb1e30c5" image_url = "https://raw.githubusercontent.com/maveryn/trace/main/docs/assets/paper-domain-montage/trace-paper-domain-montage.png" processor = AutoProcessor.from_pretrained(model_id, revision=revision) model = Qwen2_5_VLForConditionalGeneration.from_pretrained( model_id, revision=revision, dtype="auto", device_map="auto", ) messages = [ { "role": "user", "content": [ {"type": "image", "url": image_url}, {"type": "text", "text": "How many visual domains are shown with example images? Respond with only a JSON object using the key \"answer\"."}, ], } ] inputs = processor.apply_chat_template( messages, tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt", ).to(model.device) output_ids = model.generate(**inputs, max_new_tokens=128) generated_ids = [ output[len(prompt) :] for prompt, output in zip(inputs.input_ids, output_ids) ] print(processor.batch_decode(generated_ids, skip_special_tokens=True)[0]) ``` ### Verified reference answer For the published montage used above, the expected output is: ```json {"answer": 11} ``` The value is backed by the committed [montage manifest](https://github.com/maveryn/trace/blob/main/docs/gallery/paper-domain-montage.v1.json) (`layout.panel_count`). This is reference ground truth, not a claim about a particular decoding run; publish qualitative model outputs only with their recorded inference settings. The pinned revision is the checkpoint used by the canonical evaluation; later repository heads may update documentation without changing the model weights. ## Intended use and limitations This checkpoint is intended for research on multimodal reasoning and verifiable-reward post-training. It can produce incorrect answers, unsupported reasoning, or unreliable grounding. It was trained on synthetic tasks and has not been validated for safety-critical, medical, legal, or autonomous decision-making uses. Results depend on the recorded prompts, decoding, parsers, and scorers. ## License This checkpoint is subject to the upstream [Qwen Research License](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE).