Files
trace-qwen2.5-vl-7b/README.md
ModelHub XC e0e07fbfad 初始化项目,由ModelHub XC社区提供模型
Model: maveryn/trace-qwen2.5-vl-7b
Source: Original Platform
2026-09-19 20:20:07 +08:00

154 lines
5.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
language:
- en
base_model:
- Qwen/Qwen2.5-VL-7B-Instruct
datasets:
- maveryn/trace
pipeline_tag: image-text-to-text
library_name: transformers
tags:
- multimodal
- visual-reasoning
- reinforcement-learning
- grpo
- rlvr
- trace
---
# TRACE Qwen2.5-VL 7B
TRACE Qwen2.5-VL 7B is a GRPO checkpoint derived from
[`Qwen/Qwen2.5-VL-7B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct)
and trained on 64,000 grounded visual-reasoning examples spanning 1,000 tasks
from
[`maveryn/trace`](https://huggingface.co/datasets/maveryn/trace). This
repository contains the merged step-500 inference checkpoint.
[Paper](https://arxiv.org/abs/2607.19790) ·
[Project page](https://maveryn.github.io/trace/) ·
[GitHub](https://github.com/maveryn/trace) ·
[Collection](https://huggingface.co/collections/maveryn/trace-6a604291b4be4ed6399b9f24) ·
[Training configuration](https://github.com/maveryn/trace/blob/rlvr/rlvr/configs/trace-qwen2.5-vl-7b.yaml) ·
[Evaluation suite](https://github.com/maveryn/trace/tree/rlvr/rlvr/evaluation/trace_eval) ·
[Run artifacts](https://huggingface.co/datasets/maveryn/trace-eval-runs)
## Training and provenance
| Field | Value |
| --- | --- |
| Base model | `Qwen/Qwen2.5-VL-7B-Instruct@cc594898137f460bfe9f0759e9844b3ce807cfb5` |
| Dataset | `maveryn/trace@4e5b54361360296a855542b40cfd8b7f81b355fe` |
| Method | GRPO, 500 steps, global/rollout batch 128, 8 responses per prompt |
| Optimization | AdamW, learning rate `1e-6`, constant schedule, KL disabled |
| Reward | `0.95 × exact JSON answer + 0.05 × valid JSON format`; annotation reward `0` |
| Reference run | 8 NVIDIA H100 80GB GPUs |
| Released artifact | Merged step-500 inference weights |
The released training profile reads `prompt_answer`, scores `answer_gt`, and
does not use the advisory `trace_supervision_mode` column. The consumed fields,
embedded image bytes, and row order in the current dataset release were
verified identical to the original training input in the public
[equivalence receipt](https://github.com/maveryn/trace/blob/rlvr/rlvr/dataset_equivalence.v1.json).
The canonical output ends with `{"answer": ...}`. Exact hashes, source
revision, and run provenance are recorded in
[`trace_training_provenance.json`](trace_training_provenance.json) and
[`.trace_model_revision.json`](.trace_model_revision.json). The repository
does not include optimizer, scheduler, trainer, or FSDP state for continuation.
## Evaluation
`trace_eval_v1` evaluates 24 external benchmarks and 32,805 examples per model
and decoding seed. Scores below are the unweighted macro mean of the 24
benchmark percentages, reported as mean ± sample standard deviation across
seeds 42, 43, and 44.
| Model | Overall score |
| --- | ---: |
| Qwen2.5-VL-7B-Instruct | 47.93 ± 0.30 |
| TRACE Qwen2.5-VL 7B | **51.99 ± 0.17** |
| VERO Qwen2.5-VL 7B | 52.22 ± 0.07 |
| Game-RL Qwen2.5-VL 7B | 48.05 ± 0.45 |
| Sphinx Qwen2.5-VL 7B | 49.34 ± 0.15 |
| PCGRPO Qwen2.5-VL 7B | 48.80 ± 0.28 |
| Paired improvement | **+4.06 ± 0.41** |
Per-benchmark scores, model revisions, aggregation, and benchmark provenance
are available in the public
[`results.json`](https://github.com/maveryn/trace/blob/rlvr/rlvr/evaluation/trace_eval/results.json)
and [evaluation documentation](https://github.com/maveryn/trace/blob/rlvr/rlvr/evaluation/trace_eval/README.md).
## Usage
```python
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
model_id = "maveryn/trace-qwen2.5-vl-7b"
revision = "4d0f1ae8ee25022058090dbdbff61957ece7331d"
image_url = "https://raw.githubusercontent.com/maveryn/trace/main/docs/assets/paper-domain-montage/trace-paper-domain-montage.png"
processor = AutoProcessor.from_pretrained(model_id, revision=revision)
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
model_id,
revision=revision,
dtype="auto",
device_map="auto",
)
messages = [
{
"role": "user",
"content": [
{"type": "image", "url": image_url},
{"type": "text", "text": "How many visual domains are shown with example images? Respond with only a JSON object using the key \"answer\"."},
],
}
]
inputs = processor.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
output_ids = model.generate(**inputs, max_new_tokens=128)
generated_ids = [
output[len(prompt) :] for prompt, output in zip(inputs.input_ids, output_ids)
]
print(processor.batch_decode(generated_ids, skip_special_tokens=True)[0])
```
### Verified reference answer
For the published montage used above, the expected output is:
```json
{"answer": 11}
```
The value is backed by the committed
[montage manifest](https://github.com/maveryn/trace/blob/main/docs/gallery/paper-domain-montage.v1.json)
(`layout.panel_count`). This is reference ground truth, not a claim about a
particular decoding run; publish qualitative model outputs only with their
recorded inference settings.
The pinned revision is the checkpoint used by the canonical evaluation; later
repository heads may update documentation without changing the model weights.
## Intended use and limitations
This checkpoint is intended for research on multimodal reasoning and
verifiable-reward post-training. It can produce incorrect answers,
unsupported reasoning, or unreliable grounding. It was trained on synthetic
tasks and has not been validated for safety-critical, medical, legal, or
autonomous decision-making uses. Results depend on the recorded prompts,
decoding, parsers, and scorers.
## License
This checkpoint is released under
[Apache-2.0](https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct/blob/main/LICENSE),
consistent with the upstream base model.