Files
submmit/docs/gpu-memory-capacity-2026-08-10.md

3.8 KiB
Raw Permalink Blame History

ModelHub GPU memory boundaries — 2026-08-10

Scope and method

The live ModelHub machine-info endpoint returned 18 GPU types on 2026-08-10. Four were disabled (canVerify=false): Mthreads_s5000, Kunlunxin_r-200-8f, Cambricon_mlu-590, and Ascend_950. They are not current submission targets. The remaining 14 are covered below.

ModelHub PREFLIGHT_OOM logs reveal the platform's actual rule:

required memory = complete recursive repository size × 1.20
maximum repository size = allocated GPU memory ÷ 1.20

This was cross-checked against Kwaipilot/KAT-Coder-V2.5-Dev: the recursive ModelScope tree was 64.5916 GiB and ModelHub reported 64.6 GiB on disk, then required 77.5 GiB after the 20% multiplier. Consequently, parameter count or a single selected weight format is not a valid substitute.

Current boundaries

ModelHub GPU type Allocated memory Maximum complete repository Primary evidence
Cambricon_mlu-370-x4 24 GiB 20.000 GiB 49 ModelHub OOM logs; manufacturer says 24 GB
Ascend_910-b4 32 GiB 26.667 GiB 8 ModelHub OOM logs
Iluvatar_bi-100 32 GiB 26.667 GiB 48 ModelHub OOM logs
Iluvatar_bi-150 32 GiB 26.667 GiB 24 ModelHub OOM logs
Iluvatar_mrv-100 32 GiB 26.667 GiB 50 ModelHub OOM logs
Vastai_va16 32 GiB 26.667 GiB 40 ModelHub OOM logs
Cambricon_mlu-370-x8 48 GiB 40.000 GiB manufacturer specification
Mthreads_s4000 48 GiB 40.000 GiB manufacturer specification
Ascend_910-b3 64 GiB 53.333 GiB published deployment specification
Biren_166m 64 GiB 53.333 GiB manufacturer publication
hygon_k100-ai 64 GiB 53.333 GiB 15 ModelHub OOM logs
MetaX_c-500 64 GiB 53.333 GiB 29 ModelHub OOM logs; manufacturer says 64 GB
Sunrise_pt-200-x1 64 GiB 53.333 GiB 36 ModelHub OOM logs
Kunlunxin_p-800 96 GiB 80.000 GiB public procurement specification

Published sources:

Runtime behavior

The code uses the evidence hierarchy local ModelHub OOM > explicit environment override > historical ModelHub OOM > published specification. A newly downloaded structured OOM log records the actual allocation for that ModelHub GPU type and replaces the published fallback on subsequent planning cycles.

For a known GPU, missing even one file size causes preflight_model_size_unknown; the candidate is deferred rather than estimated. An exact boundary is accepted, while anything larger is rejected locally as preflight_predicted_oom before consuming a platform queue slot. If ModelHub later introduces another GPU, it is deferred as preflight_gpu_memory_unknown until a platform observation, published capacity, or explicit override supplies evidence.

This boundary reproduces ModelHub's repository preflight. It does not promise that every model below the boundary will run: framework support, operators, quantization, context length, and runtime KV-cache memory remain separate checks.