Files
MiniArt-2.0/benchmark_results.txt
ModelHub XC 8ffc55bc71 初始化项目,由ModelHub XC社区提供模型
Model: Dev4285/MiniArt-2.0
Source: Original Platform
2026-08-28 05:40:26 +08:00

17 lines
707 B
Plaintext

=====================================================
MINIART 2.0 COMPREHENSIVE BENCHMARK SCORES REPORT
=====================================================
1. GPQA DIAMOND (PhD Expert Domain Reasoning):
- GPQA Diamond Overall: 34.8% (+6.4% over MiniArt 1.0 baseline 28.4%)
- Physics Domain: 35.4%
- Chemistry Domain: 33.8%
- Biology Domain: 35.3%
2. GOLD STANDARD VLM & REASONING BENCHMARKS:
- GSM8K (Math Reasoning): 79.8% (+3.4% boost)
- VQA v2 (Visual QA): 64.2% (New Modality)
- ScienceQA (Multimodal): 72.5% (+30.4% boost)
- Logical Deduction: 76.2% (+2.4% boost)
- Code Reasoning: 71.4% (+2.5% boost)