Files
submmit/modelhub_submmit_api
..
2026-07-10 00:37:43 +08:00
2026-08-02 16:59:44 +08:00
2026-07-10 00:37:43 +08:00
2026-07-10 00:37:43 +08:00

ModelHub Submission Runner

This package automates ModelScope model discovery and ModelHub submission. It currently supports:

  • one-shot submission planning via main.py
  • daily batch execution via run_daily.sh
  • continuous queue refill via run_poll.sh
  • multiple ModelHub tokens read from KEY.md and KEYS.md
  • automatic task/framework/template selection across the supported GPU catalog
  • adaptive long-term/recent GPU exploitation with a persistent local snapshot
  • live queue/throughput-aware GPU weighting and confidence-ranked framework selection

Layout

  • main.py: core discovery, scoring, dedup, and submission
  • daily_runner.py: daily wave orchestration
  • poll_runner.py: long-running queue refiller
  • queue_cleanup.py: fail-closed cleanup for deterministic OOM and architecture policies
  • runner_common.py: shared token / key file loading
  • hf_discovery.py: ModelScope model discovery and inspection (keeps the legacy module name)
  • modelhub_client.py: ModelHub API client and token-pool routing
  • history_stats.py: online history aggregation, ranking, and warnings
  • candidate_preflight.py: deterministic repository, memory, context, and compatibility gates
  • failure_taxonomy.py: deterministic/platform/semantic failure routing
  • architecture_compatibility.py: exact architecture identities and learned compatibility keys
  • llm_classifier.py: offline-only experimental ambiguity-analysis helper
  • template_selector.py: template lookup and GPU normalization
  • task_registry.py: task-type and framework selection rules
  • tests/: unit tests and regression coverage

Key Files

  • KEY.md: primary ModelScope and ModelHub tokens
  • KEYS.md: optional supplemental ModelHub tokens
  • templates/public_submit/adapt_task_templates.jsonl: public submit templates

The runner reads both files automatically. Add more accounts by appending XC_TOKEN3, XC_TOKEN4, and so on to KEYS.md.

Template lookup is also relative. The selector searches from the current working directory and the module directory. The primary project layout is:

  • templates/public_submit/adapt_task_templates.jsonl

It still accepts the legacy fallback path below for compatibility with older deployments:

  • model adaptation/templates/public_submit/adapt_task_templates.jsonl

If your Space keeps templates in another location, set MODELHUB_TEMPLATE_FILE to the exact JSONL path.

Quick Start

Run a single daily batch:

cd /path/to/submmit
# testing: one run defaults to 3 targets if daily-target is not specified
bash run_daily.sh --rounds 1

Run the continuous queue refiller:

cd /path/to/submmit
bash run_poll.sh

Dry-run either entrypoint to inspect candidate selection without submitting:

cd /path/to/submmit
bash run_daily.sh --dry-run
bash run_poll.sh --dry-run

Behavior

  • The runner auto-discovers all safe GPU/template combinations from the public submit catalog.
  • Automatic GPU selection uses exact 70/30 accepted-task scheduling: long-term Wilson-ranked top 3 GPUs and the top GPUs from the latest 1,000 terminal tasks. There is no all-GPU exploration category.
  • Within each category, weighted-fair scheduling uses estimated queue backlog hours, recent public throughput/success, machine availability, and worker concurrency. Unavailable or stalled GPU pools are circuit-broken instead of continuing to absorb work.
  • Compatible frameworks are ranked by ModelHub public aggregate success statistics plus capped local GPU+framework evidence, with a 300-sample public minimum and Wilson confidence bounds. Missing or undersized public evidence receives zero traffic rather than falling back to exploration.
  • New frameworks are discovered from the live catalog but get no novelty bonus. They are eligible only with a complete official build config that passes local validation and a confidence score at least 10% above the best incumbent.
  • Five consecutive local failures pause a GPU/framework pair for 12 hours; a sub-20% rate over the latest 20 terminal tasks pauses it for 6 hours.
  • An explicit "framework does not support this model/architecture" failure learns a 30-day GPU + framework + task + architecture block. Architecture identity comes from candidate config.json plus exact unsupported model_type or architectures strings in the runtime log, never from repository names. A newer success clears the block, and generic unsupported backend/operator messages cannot create one.
  • Blacklist additions are persisted and detected every three poll cycles. A new rule immediately launches a lightweight architecture-only queue scan. Exact matching waiting tasks are stopped after two state checks; running tasks and tasks without local framework/task metadata are protected.
  • Cold start performs no blocking history seed. After the first submission pass, a durable cursor completes the full history scan as one streaming phase using about 1,000 task rows per metadata batch and 200 logs per classification batch. Aggregate checkpoints are synchronized about every 2,000 rows; raw task rows and failure logs are never synchronized. Restarts resume the cursor, temporary deduplication IDs disappear at completion, and later cycles process only new outcome changes.
  • A strategy generation lasts exactly 200 platform-accepted submissions. Rejected API calls and duplicates do not advance it. The next cycle refreshes platform history before submitting again.
  • Strategy state is stored in .modelhub_state/gpu_strategy.json; a generation never recalculates during candidate submission.
  • Candidate discovery starts with the configured recent window, then automatically expands to 7 days, 30 days, and older history (up to 3,000 models) when the recent pool is exhausted.
  • Model verification responses are cached across poll cycles for 15 minutes. Local model/GPU failures cool down after 24 hours instead of remaining permanently blocked.
  • Community deduplication is model/GPU-specific: another GPU's adaptation does not block the current GPU. Every actual submission performs a fresh uncached check for its exact GPU.
  • If the community lookup is unavailable, submission is deferred. A platform model-uniqueness rejection permanently excludes only that model/GPU combination from future local retries.
  • Each model can be submitted at most once per GPU.
  • Multiple ModelHub tokens are pooled and used to route submissions to the account with available async capacity.
  • Concurrent submissions reserve account slots locally, and an account-capacity race automatically falls through to another account.
  • Concurrent local processes claim model/GPU pairs in .modelhub_state/submission_claims.jsonl; shared ledger, history, and outcome files use process locks and atomic replacement.
  • ModelScope list pages are paced and cached for 15 minutes. HTTP 429 responses use exponential backoff and Retry-After; pages already downloaded remain usable and the failed page is retried on the next cycle.
  • Every third poll cycle, a full account gets one controlled capacity probe. A successful probe raises that account's persisted known limit; a capacity rejection enters cooldown.
  • Each [scan] log records the discovery stage and candidate yield. The final [daily] wave_done log includes skip_reasons, making empty candidate pools distinguishable from API failures.
  • At startup and every 120 poll cycles, active tasks are checked with the same recursive-size memory rule as new submissions. Only tasks whose own repository size times 1.20 exceeds their selected GPU capacity are stopped, after a fresh account-scoped active-state check. Incomplete size/capacity evidence is never used for cancellation.
  • Models older than seven days may occupy only the current account capacity minus its final 5 positions. The account pool enforces this per-account admission boundary atomically and updates it when capacity probing discovers a higher limit. A failed active-count read makes that account ineligible for older models while other readable accounts are still tried. Recent models may use all available positions. Age policy never cancels an existing waiting or running task; startup and scheduled cleanup remain limited to deterministic OOM and learned architecture evidence. Date-deferred candidates are skipped, not failed or permanently excluded.

Important Flags

Common flags:

  • --daily-target: total target submissions for the day; 0 means unlimited
  • --min-downloads: ModelScope download floor
  • --history-stats-threshold: local ledger threshold before using online history stats
  • --max-scan-models: hard cap on scanned HF models for a run (0 = auto)
  • --scan-multiplier: multiplier used for auto scan cap derivation from quota/queue capacity
  • --read-concurrency: concurrent HTTP reads while scanning model candidates (default 4)
  • --max-submits-per-run: max tasks to submit per run cycle (0 = unlimited)
  • --submit-concurrency: concurrent task submissions (default auto, uses 0)
  • --skip-outcome-sync: skip outcome sync before scanning
  • --skip-history-archive: skip history archive download for this run
  • --dry-run: plan only, do not submit
  • --gpu-strategy-refresh-submissions: accepted tasks per strategy generation (default 200)
  • --disable-gpu-strategy: restore legacy ordering; explicit --gpu/--gpus also bypasses adaptive selection
  • --disable-market-intelligence: disable live queue/throughput and framework-stat weighting
  • run_daily.sh injects --daily-target 3 when no daily-target flag is provided. Set SUBMIT_DAILY_TARGET or pass --daily-target explicitly for a different target.

run_poll.sh adds:

  • --poll-interval-seconds: sleep when all accounts are saturated (default 15)
  • --idle-interval-seconds: sleep when a cycle submits nothing (default 60)
  • --max-scan-models: hard cap on scanned HF models for this cycle (0 = auto)
  • --scan-multiplier: multiplier used for auto scan cap derivation from quota/queue capacity
  • --max-submits-per-run: max tasks to submit per poll cycle (0 = unlimited)
  • --skip-outcome-sync: skip outcome sync before scanning
  • --skip-history-archive: skip history archive download for this cycle
  • --submit-concurrency: concurrent task submission calls used by each cycle (0 = auto)
  • --post-cycle-cooldown-seconds: pause after a successful cycle before next cycle (default 2)
  • --max-cycles: optional hard stop for testing or batch windows
  • --disable-queue-cleanup: disable automatic deterministic OOM/architecture queue cleanup

Model age is admission-only. The default MODELHUB_RECENT_MODEL_RESERVE_SLOTS=5 reserves each account's final five known-capacity positions for models updated within seven days. Queue cleanup reports ageCleanupMode=admission_only and the compatibility field oldOverflowCount=0; no restart migration or periodic date-based stop is performed.

Every successful worker-initiated stop is persisted as policy_cancelled in the outcome store. It is excluded from GPU/framework success rates, local failure cooldowns, and circuit breakers. If a task races to a real success before the stop takes effect, that success remains authoritative.

Unclassified failures and failures with failureScope=unknown are reported as unresolvedFailureCount. They do not penalize GPU/framework decision success, open attributable-failure circuits, or create model/GPU cooldowns. Explicitly classified non-platform failures remain strategy evidence, while platform failures retain their separate short-circuit handling.

Failure-informed preflight is enabled by default. It rejects deterministic missing-file and predicted-OOM cases, clamps unsafe context-length arguments, and records its decisions in candidatePreflight and each candidate's preflightMetadata. Use --disable-candidate-preflight only for diagnosis.

The memory gate totals the complete recursive repository and applies the same 20% overhead observed in ModelHub PREFLIGHT_OOM reports. All 14 currently verifiable GPU types have evidence-backed capacities; an incomplete repository size is deferred instead of estimated. See ../docs/gpu-memory-capacity-2026-08-10.md.

The online runner never constructs an LLM client. A DashScope key or any MODELHUB_QWEN_*/MODELHUB_LLM_CLASSIFIER_* environment variable cannot enable inference. llm_classifier.py remains available only for deliberately invoked, offline experiments whose output is reviewed before being converted into a deterministic rule.

Outcome sync also classifies a bounded set of this worker's failed-task ZIP logs. Hard error signatures run first; ambiguous runtime roots remain explicitly unclassified and are never sent to an LLM. Platform faults are excluded from long-term compatibility scores and use a short 30-minute breaker after three consecutive failures. Failed log downloads are persisted and stop after three attempts.

Output

Run artifacts are written under:

  • runs/: one-shot submission runs
  • daily_runs/: batch orchestration runs
  • poll_runs/: poller cycles

Each run typically includes:

  • summary.json
  • pre_submit_report.json
  • candidates.jsonl
  • submitted.jsonl
  • skipped.jsonl
  • failed.jsonl

Persistent local scheduler state is written under .modelhub_state/:

  • gpu_strategy.json: GPU ranks, generation progress, and 70/30 accepted counters
  • market_intelligence.json: cached public queue, throughput, health, and framework statistics
  • account_capacity.json: learned per-account active-task limits
  • submission_exclusions.jsonl: non-retryable model/GPU uniqueness rejections
  • queue_cleanup_latest.json: latest active-task sizing evidence and cancellation result
  • architecture_compatibility_blacklist.json: current dynamic compatibility blocks and evidence
  • architecture_history_backfill.json: resumable full-history cursor and temporary deduplication IDs

Verification

Run the full test suite:

cd /path/to/submmit
python3 -m unittest discover -s tests -v

Notes

  • This is a submission automation tool, not a scheduler daemon. Use screen, tmux, nohup, or systemd if you want it to keep running in the background.
  • The platform still enforces per-account async capacity limits, so the poller can keep the queue close to full but cannot override the platform cap.
  • bash run_poll.sh now defaults to unlimited mode and keeps refilling until you stop the process manually.
  • Queue polling defaults to 15 seconds, successful-cycle cooldown to 2 seconds, and per-cycle submissions to all available slots. Override these values when the platform requires a lower request rate.