298 lines
17 KiB
Markdown
298 lines
17 KiB
Markdown
# ModelHub Adaptation Agent
|
|
|
|
This repository is packaged for the ModelHub XC agent platform.
|
|
|
|
## Platform Contract
|
|
|
|
- Root-level `Dockerfile`
|
|
- Listens on port `8080`
|
|
- Exposes `GET /health`
|
|
- Handles `SIGTERM`
|
|
- Reads platform-provided `STRATEGY_ID` and attaches it to task submissions as `strategyId`
|
|
|
|
The root `main.py` starts a lightweight health server and runs the existing
|
|
submission poller in a child process.
|
|
|
|
## Runtime Environment
|
|
|
|
The image includes multi-account ModelHub and ModelScope token fallbacks for
|
|
the agent platform. Environment variables can override them without rebuilding
|
|
the image.
|
|
|
|
- `MODELHUB_XC_TOKEN`, `XC_TOKEN`, `XC_TOKEN2...`, or `MODELHUB_XC_TOKENS` for ModelHub API authentication
|
|
- `MODELHUB_JWT_TOKEN` or `JWT_TOKEN` can be used instead when the platform provides a JWT
|
|
- `MODELSCOPE_API_TOKEN` or `MODELSCOPE_TOKEN` optional override for the embedded ModelScope fallback token
|
|
- `STRATEGY_ID` is expected to be injected by the ModelHub agent platform and is attached to submissions for strategy attribution; it is not an API authentication token
|
|
|
|
Optional tuning:
|
|
|
|
- `MODELHUB_AGENT_POLL_INTERVAL_SECONDS` default `15`
|
|
- `MODELHUB_AGENT_IDLE_INTERVAL_SECONDS` default `60`
|
|
- `MODELHUB_AGENT_POST_CYCLE_COOLDOWN_SECONDS` default `2`
|
|
- The hosted entrypoint always uses `--max-submits-per-run 0` so stale deployment
|
|
settings cannot restrict a refill cycle to five submissions.
|
|
- `MODELHUB_AGENT_ACTIVE_TASK_CAP` default `100` per account
|
|
- `MODELHUB_CAPACITY_PROBE_INTERVAL_CYCLES` default `3`
|
|
- `MODELHUB_CAPACITY_STATE_PATH` default `.modelhub_state/account_capacity.json`
|
|
- `MODELHUB_AGENT_ACTIVE_COUNTS_TTL_SECONDS` default `15`
|
|
- `MODELHUB_AGENT_RESERVATION_TTL_SECONDS` default `120`
|
|
- `MODELHUB_AGENT_INSTANCE_ID` optional stable worker identity used to spread concurrent agents across accounts and candidates
|
|
- `MODELHUB_AGENT_CLAIMS_PATH` default `.modelhub_state/submission_claims.jsonl`
|
|
- `MODELHUB_SUBMISSION_EXCLUSIONS_PATH` default `.modelhub_state/submission_exclusions.jsonl`
|
|
- `MODELHUB_AGENT_DAILY_TARGET`
|
|
- `MODELHUB_AGENT_MIN_DOWNLOADS`
|
|
- `MODELHUB_AGENT_GPUS`
|
|
- `MODELHUB_AGENT_EXTRA_ARGS`
|
|
- `MODELHUB_GPU_STRATEGY_STATE_PATH` default `.modelhub_state/gpu_strategy.json`
|
|
- `MODELHUB_MARKET_INTELLIGENCE_PATH` default `.modelhub_state/market_intelligence.json`
|
|
- `MODELHUB_MARKET_QUEUE_REFRESH_SECONDS` default `600`
|
|
- `MODELHUB_MARKET_FRAMEWORK_REFRESH_SECONDS` default `21600`
|
|
- `MODELHUB_MARKET_THROUGHPUT_WINDOW_HOURS` default `6`
|
|
- `MODELHUB_MARKET_FRAMEWORK_MIN_SAMPLES` default `300`
|
|
- `MODELSCOPE_PAGE_INTERVAL_SECONDS` default `0.25`
|
|
- `MODELSCOPE_PAGE_CACHE_TTL_SECONDS` default `900`
|
|
- `MODELHUB_AGENT_VERIFY_CACHE_TTL_SECONDS` default `900`
|
|
- `MODELHUB_QUEUE_CLEANUP_INTERVAL_CYCLES` default `120`; cleanup also runs once at startup
|
|
- `MODELHUB_QUEUE_CLEANUP_READ_CONCURRENCY` default `6`
|
|
- `MODELHUB_QUEUE_CLEANUP_REPORT_PATH` default `.modelhub_state/queue_cleanup_latest.json`
|
|
- `MODELHUB_ARCHITECTURE_BLACKLIST_PATH` default `.modelhub_state/architecture_compatibility_blacklist.json`
|
|
- `MODELHUB_ARCHITECTURE_BLOCK_TTL_DAYS` default `30`
|
|
- `MODELHUB_RECENT_MODEL_RESERVE_SLOTS` default `10` per account
|
|
- `MODELHUB_DYNAMIC_OLD_MODEL_CLEANUP_RESERVE_SLOTS` default `5` per account
|
|
- `MODELHUB_RECENT_MODEL_DAYS` default `7`
|
|
|
|
## Adaptive GPU Strategy
|
|
|
|
When no explicit GPU override is supplied, the worker uses a success-first 70/30
|
|
strategy generation with no self-funded exploration:
|
|
|
|
- 70%: the three long-term GPUs with the best Wilson lower confidence score and at least 100 terminal samples
|
|
- 30%: the top recent GPUs among the latest 1,000 terminal tasks
|
|
- 0%: unvetted/all-GPU exploration; community-wide results provide the exploration signal
|
|
|
|
The 70/30 category ratio remains exact across accepted tasks. Inside each
|
|
category, weighted fair scheduling combines the category's historical rank with
|
|
live public market data:
|
|
|
|
- recent public success quality, scored with a strongly weighted Wilson lower confidence bound
|
|
- estimated backlog hours (`waiting / recent completions per hour`) as a bounded tie-breaker
|
|
- machine availability, running workers, and advertised concurrency
|
|
- a circuit breaker for unavailable or apparently stalled GPU pools
|
|
|
|
This optimizes expected successful completions rather than blindly selecting the
|
|
smallest queue. Queue/throughput data is refreshed every 10 minutes and persisted
|
|
in `.modelhub_state/market_intelligence.json`. A failed refresh keeps the last good
|
|
snapshot, uses a retry backoff, and never blocks normal submissions.
|
|
|
|
For each compatible model/GPU pair, the worker also ranks the GPU's supported
|
|
frameworks using ModelHub's public aggregate `modelCount` and `successCount` data,
|
|
then blends in the worker's own GPU+framework outcomes with a capped weight.
|
|
Only frameworks with at least 300 public samples and a safe Wilson lower bound
|
|
are eligible. Framework statistics refresh
|
|
every 6 hours, so they do not add per-model API traffic. Newly published
|
|
frameworks are discovered automatically, but receive no novelty bonus: they can
|
|
win only when their confidence-adjusted success score beats the best incumbent by
|
|
at least 10%. A new framework
|
|
is eligible only after the authenticated official build-config endpoint returns
|
|
a complete config that passes local structure, placeholder, framework-name, and
|
|
GPU-parallelism validation. Valid official configs are cached and refreshed with
|
|
the framework snapshot; local templates remain the fail-safe fallback.
|
|
|
|
Only platform-accepted tasks count. After exactly 200 accepted tasks, the next
|
|
poll cycle reloads all account history, generates a new immutable strategy snapshot,
|
|
and resets the generation counters to 140/60 targets. The active snapshot and
|
|
progress are stored in `.modelhub_state/gpu_strategy.json`.
|
|
|
|
Five consecutive local failures open a 12-hour GPU/framework circuit breaker.
|
|
A sub-20% success rate over the latest 20 terminal tasks opens a 6-hour breaker.
|
|
Platform/infrastructure failures are excluded from long-term compatibility rates
|
|
and model/profile breakers. Three consecutive platform failures on a GPU/framework
|
|
instead open a short 30-minute breaker, so a temporary broken runner or lack of an
|
|
idle card does not permanently poison otherwise successful evidence.
|
|
Candidate shortages expand the model search window; they never unlock an
|
|
unvetted GPU or framework.
|
|
|
|
Failed-task archives are also classified conservatively. When ModelHub explicitly
|
|
says that the selected framework does not support the model or architecture, the
|
|
runner learns an exact GPU + framework + task type + architecture block from the
|
|
candidate repository's `config.json`. Repository names are never used as
|
|
architecture evidence. Exact `architectures` values take priority and
|
|
`model_type` is also retained when the runtime explicitly says Transformers does
|
|
not recognize that type; missing metadata does not create a block. Generic unsupported operators, attention backends, GPU types,
|
|
quantization failures, and OOMs cannot enter this blacklist. A newer success for
|
|
the same exact combination clears the block, and otherwise it expires after 30
|
|
days. Set `MODELHUB_ARCHITECTURE_BLOCK_TTL_DAYS` to a value from 1 to 365 to
|
|
change that window. The stats report exposes `architectureCompatibilityBlocks`
|
|
and per-GPU/framework block counts. The live snapshot is written to
|
|
`.modelhub_state/architecture_compatibility_blacklist.json`.
|
|
|
|
Before a candidate reaches the submit queue, failure-informed preflight checks
|
|
the actual ModelScope repository structure and file sizes. Non-GGUF text
|
|
frameworks require root-level config, weights, and tokenizer assets. The memory
|
|
gate recursively totals the entire repository—including duplicate weight formats
|
|
and nested shards—and applies ModelHub's observed 20% loading overhead. It covers
|
|
all 14 GPU types currently marked `canVerify=true`; nine capacities come directly
|
|
from structured ModelHub OOM reports and five from published specifications until
|
|
ModelHub supplies a stronger observation. If a known GPU's repository file sizes
|
|
are incomplete, the candidate is deferred rather than guessed. Template context
|
|
length is also clamped to the model's advertised limit. A newly introduced GPU
|
|
with no capacity evidence is likewise deferred. Override or extend known
|
|
capacities with
|
|
`MODELHUB_GPU_MEMORY_GIB_JSON`, for example
|
|
`{"New_gpu": 64}`.
|
|
|
|
At poller startup, the same deterministic memory and learned architecture gates are applied to existing
|
|
`waiting` and `running` tasks across every configured account. A task is stopped
|
|
through `PUT /api/async/task/stop-create-contest-task` only when its own current
|
|
recursive repository size, multiplied by ModelHub's observed `1.20` overhead,
|
|
exceeds the known capacity of its selected GPU. The task ID is fetched and
|
|
stopped with the token belonging to that account, and its active state is
|
|
rechecked immediately before the mutation. Missing file sizes, unknown GPU
|
|
capacities, listing failures, and tasks that have already changed state all fail
|
|
closed and are never stopped. This does not match against another task from the
|
|
same model or infer failure from historical similarity. The cleanup repeats
|
|
every 120 poll cycles by default and writes its full evidence report to
|
|
`.modelhub_state/queue_cleanup_latest.json`.
|
|
|
|
Architecture cleanup joins each active task to the locally recorded submission
|
|
or ledger entry to recover its exact framework and task type, then reads the
|
|
model's `config.json`. Only an exact GPU + framework + task type + architecture
|
|
blacklist hit can authorize cancellation. Matching waiting tasks are stopped;
|
|
running tasks remain protected and their state is rechecked again immediately
|
|
before the stop call. Failed outcomes are synchronized every three poll cycles.
|
|
When fixed `MODEL_NOT_SUPPORTED` text adds a new blacklist entry, a lightweight
|
|
architecture-only cleanup runs immediately without repeating repository-size or
|
|
model-age scans.
|
|
|
|
Each account dynamically reserves its last 10 known-capacity positions for
|
|
models updated within seven days. If an account's discovered limit is 100, 200,
|
|
or 500, older models stop at positions 90, 190, or 490 respectively. Old-model
|
|
reservations are made under the same account lock as capacity reservations, so
|
|
concurrent submissions cannot enter the reserved suffix. Once every account's
|
|
old-model allowance is exhausted, discovery is capped at the seven-day window.
|
|
Unknown model timestamps are treated as old for new submissions. When capacity
|
|
probing raises an account's known limit, its old-model boundary moves with it.
|
|
|
|
On startup, the worker first stops deterministic OOM tasks, recalculates each
|
|
account's surviving task order by numeric task ID, and removes
|
|
older-than-seven-days tasks beyond that account's current limit minus 10. Later
|
|
scheduled cleanup uses limit minus 5, retaining a small hysteresis buffer that
|
|
avoids repeatedly over-pruning valid work. For a 100-task account the two
|
|
boundaries are 90 and 95; for a 500-task account they are 490 and 495. Recent
|
|
overflow tasks are always kept. Age-based cleanup applies only to tasks that are
|
|
still waiting; running validation tasks are protected even beyond the boundary.
|
|
ModelScope metadata failures fail closed and never trigger cancellation.
|
|
Immediately before the age-only stop batch, task ownership, active status, and
|
|
post-OOM queue position are checked again. Deterministic OOM cleanup may still
|
|
stop a running task because it cannot fit the selected GPU.
|
|
|
|
The verified capacities, safe repository-size boundaries, evidence hierarchy,
|
|
and source links are recorded in
|
|
`docs/gpu-memory-capacity-2026-08-10.md`.
|
|
|
|
The online worker makes zero LLM calls. Candidate admission, GPU/framework
|
|
selection, queue cleanup, and failure feedback are deterministic and based on
|
|
repository metadata, platform capabilities, public outcomes, and explicit error
|
|
signatures. Merely storing a DashScope key in `.env` or setting a Qwen/LLM
|
|
environment variable does not activate inference. The standalone classifier
|
|
module remains only as an offline research helper for human-reviewed batches of
|
|
previously unseen errors; it is not wired into submission or outcome sync.
|
|
|
|
Outcome synchronization downloads a bounded set of failure archives for
|
|
submissions created by this worker (at most 40 per sync, four workers, three
|
|
download attempts). Deterministic signatures classify memory, repository layout,
|
|
context-length, storage, and platform faults. Unresolved runtime errors remain
|
|
explicitly ambiguous for later rule development instead of being sent to an LLM.
|
|
Signed log URLs remain in the ignored local outcome store and are removed after
|
|
classification.
|
|
|
|
The 12-account failure study and routing rationale are recorded in
|
|
`docs/failure-analysis-2026-08-10.md`.
|
|
|
|
ModelScope HTTP 429 responses use exponential backoff and `Retry-After`. Successful
|
|
pages remain cached, so a later cycle retries the failed page instead of restarting
|
|
the whole pagination scan.
|
|
|
|
## Adaptive Candidate Discovery
|
|
|
|
The configured recent window remains the fast path. If it contains no usable
|
|
model/GPU combinations, the same run progressively expands discovery to the last
|
|
7 days, the last 30 days, and finally older history (up to 3,000 models). Scanning
|
|
stops as soon as enough replacement candidates have been found.
|
|
|
|
Model verification results are reused for 15 minutes across poll cycles, and a
|
|
locally failed model/GPU pair cools down for 24 hours instead of being excluded
|
|
forever. The `[scan]` lines show every expansion stage, while `[daily] wave_done`
|
|
includes `skip_reasons` so an empty candidate pool is directly diagnosable.
|
|
|
|
Community deduplication is scoped to the exact model/GPU combination. A model
|
|
adapted on one GPU remains eligible for another GPU. Before each submission the
|
|
worker performs an uncached community check; a lookup failure defers the task
|
|
instead of failing open. Platform uniqueness rejections are persisted per
|
|
model/GPU in `.modelhub_state/submission_exclusions.jsonl` and are not retried.
|
|
|
|
## Concurrent Agents
|
|
|
|
The token pool keeps a local reservation for every in-flight submission, so a
|
|
lagging platform count cannot send all concurrent requests to the same account.
|
|
If another process fills an account first, the submission is retried immediately
|
|
against another account with capacity.
|
|
|
|
Workers that share a filesystem also coordinate model/GPU claims through
|
|
`.modelhub_state/submission_claims.jsonl`. Workers in isolated containers use
|
|
different candidate ordering (derived from `STRATEGY_ID`, instance ID, or
|
|
hostname), which reduces duplicate work while the platform remains the final
|
|
authority for account capacity and model/GPU uniqueness.
|
|
|
|
If the platform reports that a model/GPU is already being validated, the claim
|
|
is retained and the runner immediately draws replacement candidates from the
|
|
same scan instead of retrying the duplicate every cycle. Startup logs and the
|
|
health response expose `agent_version`; version `2026.08.02.3` or newer includes
|
|
duplicate replacement behavior, while version `2026.08.02.4` adds adaptive
|
|
candidate-window expansion and skip-reason reporting. Version `2026.08.02.5`
|
|
adds fail-closed model/GPU prechecks and persistent uniqueness exclusions.
|
|
Version `2026.08.04.1` adds queue/throughput intelligence, GPU health circuit
|
|
breaking, weighted-fair scheduling, live framework/config discovery, and
|
|
confidence-ranked public-plus-local framework selection.
|
|
Version `2026.08.05.1` removes self-funded GPU exploration, switches accepted
|
|
traffic to 70/30 long-term/recent exploitation, raises the public framework gate
|
|
to 300 samples, makes success dominate queue pressure, and adds recent local
|
|
GPU/framework circuit breakers.
|
|
Version `2026.08.10.2` adds evidence-backed sizing for every currently verifiable
|
|
GPU, recursive repository-size checks, deterministic failure-aware preflight,
|
|
and rate-limited lazy Qwen review for unresolved semantic cases.
|
|
Version `2026.08.10.3` selects `qwen3.7-flash` by default and recognizes the
|
|
repository root `.env` key named `dashscope` without logging its value.
|
|
Version `2026.08.11.1` adds account-owned cancellation of queued tasks that are
|
|
deterministically over the selected GPU's ModelHub memory boundary, with a
|
|
second active-state check and fail-closed handling for incomplete evidence.
|
|
Version `2026.08.11.2` removes LLM inference from every online path; credentials
|
|
alone cannot activate it, and unresolved cases remain available for offline,
|
|
human-reviewed rule development.
|
|
Version `2026.08.11.3` adds atomic per-account 80/7-day admission, prevents
|
|
history expansion once old-model positions are exhausted, performs a strict
|
|
position-80 startup cleanup, and relaxes scheduled age cleanup to position 95
|
|
after OOM cleanup and a second queue check.
|
|
Version `2026.08.11.4` replaces fixed queue positions with per-account dynamic
|
|
boundaries derived from each discovered capacity: limit minus 10 for admission
|
|
and startup cleanup, then limit minus 5 for scheduled dynamic cleanup.
|
|
Version `2026.08.12.1` protects running tasks from age-based cleanup and records
|
|
worker-initiated stops as `policy_cancelled`, excluding them from success-rate,
|
|
failure-cooldown, and circuit-breaker evidence.
|
|
Version `2026.08.12.2` learns conservative, expiring GPU/framework/architecture
|
|
compatibility blocks only from explicit ModelHub failure text, matches candidate
|
|
`config.json` metadata instead of repository names, and lets newer success
|
|
evidence clear stale blocks.
|
|
Version `2026.08.12.3` extracts unsupported `model_type`/`architectures` from the
|
|
platform's fixed failure wording, persists a dynamically growing blacklist, and
|
|
immediately removes exact-matching waiting tasks with two active-state checks.
|
|
|
|
## Deploy
|
|
|
|
Create a tag and submit the repository URL plus tag in "我的适配智能体".
|
|
|
|
```bash
|
|
git tag agent-v21
|
|
git push origin agent-v21
|
|
```
|