Files
submmit/modelhub_submmit_api/README.md
2026-08-11 01:29:54 +08:00

12 KiB
Raw Blame History

ModelHub Submission Runner

This package automates ModelScope model discovery and ModelHub submission. It currently supports:

  • one-shot submission planning via main.py
  • daily batch execution via run_daily.sh
  • continuous queue refill via run_poll.sh
  • multiple ModelHub tokens read from KEY.md and KEYS.md
  • automatic task/framework/template selection across the supported GPU catalog
  • adaptive long-term/recent GPU exploitation with a persistent local snapshot
  • live queue/throughput-aware GPU weighting and confidence-ranked framework selection

Layout

  • main.py: core discovery, scoring, dedup, and submission
  • daily_runner.py: daily wave orchestration
  • poll_runner.py: long-running queue refiller
  • queue_cleanup.py: fail-closed cleanup for active tasks that are certain to exceed GPU memory
  • runner_common.py: shared token / key file loading
  • hf_discovery.py: ModelScope model discovery and inspection (keeps the legacy module name)
  • modelhub_client.py: ModelHub API client and token-pool routing
  • history_stats.py: online history aggregation, ranking, and warnings
  • candidate_preflight.py: deterministic repository, memory, context, and compatibility gates
  • failure_taxonomy.py: deterministic/platform/semantic failure routing
  • llm_classifier.py: offline-only experimental ambiguity-analysis helper
  • template_selector.py: template lookup and GPU normalization
  • task_registry.py: task-type and framework selection rules
  • tests/: unit tests and regression coverage

Key Files

  • KEY.md: primary ModelScope and ModelHub tokens
  • KEYS.md: optional supplemental ModelHub tokens
  • templates/public_submit/adapt_task_templates.jsonl: public submit templates

The runner reads both files automatically. Add more accounts by appending XC_TOKEN3, XC_TOKEN4, and so on to KEYS.md.

Template lookup is also relative. The selector searches from the current working directory and the module directory. The primary project layout is:

  • templates/public_submit/adapt_task_templates.jsonl

It still accepts the legacy fallback path below for compatibility with older deployments:

  • model adaptation/templates/public_submit/adapt_task_templates.jsonl

If your Space keeps templates in another location, set MODELHUB_TEMPLATE_FILE to the exact JSONL path.

Quick Start

Run a single daily batch:

cd /path/to/submmit
# testing: one run defaults to 3 targets if daily-target is not specified
bash run_daily.sh --rounds 1

Run the continuous queue refiller:

cd /path/to/submmit
bash run_poll.sh

Dry-run either entrypoint to inspect candidate selection without submitting:

cd /path/to/submmit
bash run_daily.sh --dry-run
bash run_poll.sh --dry-run

Behavior

  • The runner auto-discovers all safe GPU/template combinations from the public submit catalog.
  • Automatic GPU selection uses exact 70/30 accepted-task scheduling: long-term Wilson-ranked top 3 GPUs and the top GPUs from the latest 1,000 terminal tasks. There is no all-GPU exploration category.
  • Within each category, weighted-fair scheduling uses estimated queue backlog hours, recent public throughput/success, machine availability, and worker concurrency. Unavailable or stalled GPU pools are circuit-broken instead of continuing to absorb work.
  • Compatible frameworks are ranked by ModelHub public aggregate success statistics plus capped local GPU+framework evidence, with a 300-sample public minimum and Wilson confidence bounds. Missing or undersized public evidence receives zero traffic rather than falling back to exploration.
  • New frameworks are discovered from the live catalog but get no novelty bonus. They are eligible only with a complete official build config that passes local validation and a confidence score at least 10% above the best incumbent.
  • Five consecutive local failures pause a GPU/framework pair for 12 hours; a sub-20% rate over the latest 20 terminal tasks pauses it for 6 hours.
  • A strategy generation lasts exactly 200 platform-accepted submissions. Rejected API calls and duplicates do not advance it. The next cycle refreshes platform history before submitting again.
  • Strategy state is stored in .modelhub_state/gpu_strategy.json; a generation never recalculates during candidate submission.
  • Candidate discovery starts with the configured recent window, then automatically expands to 7 days, 30 days, and older history (up to 3,000 models) when the recent pool is exhausted.
  • Model verification responses are cached across poll cycles for 15 minutes. Local model/GPU failures cool down after 24 hours instead of remaining permanently blocked.
  • Community deduplication is model/GPU-specific: another GPU's adaptation does not block the current GPU. Every actual submission performs a fresh uncached check for its exact GPU.
  • If the community lookup is unavailable, submission is deferred. A platform model-uniqueness rejection permanently excludes only that model/GPU combination from future local retries.
  • Each model can be submitted at most once per GPU.
  • Multiple ModelHub tokens are pooled and used to route submissions to the account with available async capacity.
  • Concurrent submissions reserve account slots locally, and an account-capacity race automatically falls through to another account.
  • Concurrent local processes claim model/GPU pairs in .modelhub_state/submission_claims.jsonl; shared ledger, history, and outcome files use process locks and atomic replacement.
  • ModelScope list pages are paced and cached for 15 minutes. HTTP 429 responses use exponential backoff and Retry-After; pages already downloaded remain usable and the failed page is retried on the next cycle.
  • Every third poll cycle, a full account gets one controlled capacity probe. A successful probe raises that account's persisted known limit; a capacity rejection enters cooldown.
  • Each [scan] log records the discovery stage and candidate yield. The final [daily] wave_done log includes skip_reasons, making empty candidate pools distinguishable from API failures.
  • At startup and every 120 poll cycles, active tasks are checked with the same recursive-size memory rule as new submissions. Only tasks whose own repository size times 1.20 exceeds their selected GPU capacity are stopped, after a fresh account-scoped active-state check. Incomplete size/capacity evidence is never used for cancellation.
  • Models older than seven days may occupy only positions 180 of each account's active queue. The account pool enforces the boundary atomically. Once all accounts reach 80, discovery no longer expands beyond seven days. Cleanup stops OOM tasks first, recalculates the surviving queue order, and on startup stops old tasks still ranked 81 or later. Scheduled cleanup then retains positions 8195 and stops only old tasks ranked 96 or later. Recent overflow tasks stay. Unknown ModelScope timestamps never authorize a cancellation.

Important Flags

Common flags:

  • --daily-target: total target submissions for the day; 0 means unlimited
  • --min-downloads: ModelScope download floor
  • --history-stats-threshold: local ledger threshold before using online history stats
  • --max-scan-models: hard cap on scanned HF models for a run (0 = auto)
  • --scan-multiplier: multiplier used for auto scan cap derivation from quota/queue capacity
  • --read-concurrency: concurrent HTTP reads while scanning model candidates (default 4)
  • --max-submits-per-run: max tasks to submit per run cycle (0 = unlimited)
  • --submit-concurrency: concurrent task submissions (default auto, uses 0)
  • --skip-outcome-sync: skip outcome sync before scanning
  • --skip-history-archive: skip history archive download for this run
  • --dry-run: plan only, do not submit
  • --gpu-strategy-refresh-submissions: accepted tasks per strategy generation (default 200)
  • --disable-gpu-strategy: restore legacy ordering; explicit --gpu/--gpus also bypasses adaptive selection
  • --disable-market-intelligence: disable live queue/throughput and framework-stat weighting
  • run_daily.sh injects --daily-target 3 when no daily-target flag is provided. Set SUBMIT_DAILY_TARGET or pass --daily-target explicitly for a different target.

run_poll.sh adds:

  • --poll-interval-seconds: sleep when all accounts are saturated (default 15)
  • --idle-interval-seconds: sleep when a cycle submits nothing (default 60)
  • --max-scan-models: hard cap on scanned HF models for this cycle (0 = auto)
  • --scan-multiplier: multiplier used for auto scan cap derivation from quota/queue capacity
  • --max-submits-per-run: max tasks to submit per poll cycle (0 = unlimited)
  • --skip-outcome-sync: skip outcome sync before scanning
  • --skip-history-archive: skip history archive download for this cycle
  • --submit-concurrency: concurrent task submission calls used by each cycle (0 = auto)
  • --post-cycle-cooldown-seconds: pause after a successful cycle before next cycle (default 2)
  • --max-cycles: optional hard stop for testing or batch windows
  • --disable-queue-cleanup: disable automatic OOM and old-overflow queue cleanup

The first automatic cleanup removes models older than seven days after each account's position 80. Later scheduled cleanups use position 95, while admission continues to reserve positions 81-100 for recent models. Override the later threshold with MODELHUB_DYNAMIC_OLD_MODEL_CLEANUP_THRESHOLD.

Failure-informed preflight is enabled by default. It rejects deterministic missing-file and predicted-OOM cases, clamps unsafe context-length arguments, and records its decisions in candidatePreflight and each candidate's preflightMetadata. Use --disable-candidate-preflight only for diagnosis.

The memory gate totals the complete recursive repository and applies the same 20% overhead observed in ModelHub PREFLIGHT_OOM reports. All 14 currently verifiable GPU types have evidence-backed capacities; an incomplete repository size is deferred instead of estimated. See ../docs/gpu-memory-capacity-2026-08-10.md.

The online runner never constructs an LLM client. A DashScope key or any MODELHUB_QWEN_*/MODELHUB_LLM_CLASSIFIER_* environment variable cannot enable inference. llm_classifier.py remains available only for deliberately invoked, offline experiments whose output is reviewed before being converted into a deterministic rule.

Outcome sync also classifies a bounded set of this worker's failed-task ZIP logs. Hard error signatures run first; ambiguous runtime roots remain explicitly unclassified and are never sent to an LLM. Platform faults are excluded from long-term compatibility scores and use a short 30-minute breaker after three consecutive failures. Failed log downloads are persisted and stop after three attempts.

Output

Run artifacts are written under:

  • runs/: one-shot submission runs
  • daily_runs/: batch orchestration runs
  • poll_runs/: poller cycles

Each run typically includes:

  • summary.json
  • pre_submit_report.json
  • candidates.jsonl
  • submitted.jsonl
  • skipped.jsonl
  • failed.jsonl

Persistent local scheduler state is written under .modelhub_state/:

  • gpu_strategy.json: GPU ranks, generation progress, and 70/30 accepted counters
  • market_intelligence.json: cached public queue, throughput, health, and framework statistics
  • account_capacity.json: learned per-account active-task limits
  • submission_exclusions.jsonl: non-retryable model/GPU uniqueness rejections
  • queue_cleanup_latest.json: latest active-task sizing evidence and cancellation result

Verification

Run the full test suite:

cd /path/to/submmit
python3 -m unittest discover -s tests -v

Notes

  • This is a submission automation tool, not a scheduler daemon. Use screen, tmux, nohup, or systemd if you want it to keep running in the background.
  • The platform still enforces per-account async capacity limits, so the poller can keep the queue close to full but cannot override the platform cap.
  • bash run_poll.sh now defaults to unlimited mode and keeps refilling until you stop the process manually.
  • Queue polling defaults to 15 seconds, successful-cycle cooldown to 2 seconds, and per-cycle submissions to all available slots. Override these values when the platform requires a lower request rate.