fix: protect running validation tasks from queue cleanup

This commit is contained in:
CoolBoy
2026-08-12 01:04:34 +08:00
parent 2065ad6abc
commit d908706f9a
8 changed files with 343 additions and 20 deletions

View File

@@ -124,8 +124,10 @@ bash run_poll.sh --dry-run
positions. The account pool enforces this per-account boundary atomically and updates it when
capacity probing discovers a higher limit. Cleanup stops OOM tasks first, recalculates the
surviving queue order, and applies the same limit-minus-10 boundary on startup. Scheduled
cleanup relaxes to limit minus 5 to avoid excessive pruning. Recent overflow tasks stay, and
unknown ModelScope timestamps never authorize a cancellation.
cleanup relaxes to limit minus 5 to avoid excessive pruning. Age cleanup stops waiting tasks
only; running tasks are protected and rechecked after OOM cleanup, immediately before the
age-only stop batch. Recent
overflow tasks stay, and unknown ModelScope timestamps never authorize a cancellation.
## Important Flags
@@ -167,6 +169,11 @@ minus 5, while admission continues to reserve the final 10 slots for recent
models. Override these suffix sizes with `MODELHUB_RECENT_MODEL_RESERVE_SLOTS`
and `MODELHUB_DYNAMIC_OLD_MODEL_CLEANUP_RESERVE_SLOTS`.
Every successful worker-initiated stop is persisted as `policy_cancelled` in the
outcome store. It is excluded from GPU/framework success rates, local failure
cooldowns, and circuit breakers. If a task races to a real success before the
stop takes effect, that success remains authoritative.
Failure-informed preflight is enabled by default. It rejects deterministic
missing-file and predicted-OOM cases, clamps unsafe context-length arguments,
and records its decisions in `candidatePreflight` and each candidate's