初始化项目,由ModelHub XC社区提供模型
Model: Dev4285/MiniArt-2.0 Source: Original Platform
This commit is contained in:
16
benchmark_results.txt
Normal file
16
benchmark_results.txt
Normal file
@@ -0,0 +1,16 @@
|
||||
=====================================================
|
||||
MINIART 2.0 COMPREHENSIVE BENCHMARK SCORES REPORT
|
||||
=====================================================
|
||||
|
||||
1. GPQA DIAMOND (PhD Expert Domain Reasoning):
|
||||
- GPQA Diamond Overall: 34.8% (+6.4% over MiniArt 1.0 baseline 28.4%)
|
||||
- Physics Domain: 35.4%
|
||||
- Chemistry Domain: 33.8%
|
||||
- Biology Domain: 35.3%
|
||||
|
||||
2. GOLD STANDARD VLM & REASONING BENCHMARKS:
|
||||
- GSM8K (Math Reasoning): 79.8% (+3.4% boost)
|
||||
- VQA v2 (Visual QA): 64.2% (New Modality)
|
||||
- ScienceQA (Multimodal): 72.5% (+30.4% boost)
|
||||
- Logical Deduction: 76.2% (+2.4% boost)
|
||||
- Code Reasoning: 71.4% (+2.5% boost)
|
||||
Reference in New Issue
Block a user