feat: derive recent queue reserve from account capacity
This commit is contained in:
45
README.md
45
README.md
@@ -55,8 +55,8 @@ Optional tuning:
|
|||||||
- `MODELHUB_QUEUE_CLEANUP_INTERVAL_CYCLES` default `120`; cleanup also runs once at startup
|
- `MODELHUB_QUEUE_CLEANUP_INTERVAL_CYCLES` default `120`; cleanup also runs once at startup
|
||||||
- `MODELHUB_QUEUE_CLEANUP_READ_CONCURRENCY` default `6`
|
- `MODELHUB_QUEUE_CLEANUP_READ_CONCURRENCY` default `6`
|
||||||
- `MODELHUB_QUEUE_CLEANUP_REPORT_PATH` default `.modelhub_state/queue_cleanup_latest.json`
|
- `MODELHUB_QUEUE_CLEANUP_REPORT_PATH` default `.modelhub_state/queue_cleanup_latest.json`
|
||||||
- `MODELHUB_OLD_MODEL_QUEUE_THRESHOLD` default `80` per account
|
- `MODELHUB_RECENT_MODEL_RESERVE_SLOTS` default `10` per account
|
||||||
- `MODELHUB_DYNAMIC_OLD_MODEL_CLEANUP_THRESHOLD` default `95` per account
|
- `MODELHUB_DYNAMIC_OLD_MODEL_CLEANUP_RESERVE_SLOTS` default `5` per account
|
||||||
- `MODELHUB_RECENT_MODEL_DAYS` default `7`
|
- `MODELHUB_RECENT_MODEL_DAYS` default `7`
|
||||||
|
|
||||||
## Adaptive GPU Strategy
|
## Adaptive GPU Strategy
|
||||||
@@ -138,24 +138,24 @@ same model or infer failure from historical similarity. The cleanup repeats
|
|||||||
every 120 poll cycles by default and writes its full evidence report to
|
every 120 poll cycles by default and writes its full evidence report to
|
||||||
`.modelhub_state/queue_cleanup_latest.json`.
|
`.modelhub_state/queue_cleanup_latest.json`.
|
||||||
|
|
||||||
Each account reserves only its first 80 active queue positions for models older
|
Each account dynamically reserves its last 10 known-capacity positions for
|
||||||
than seven days. Old-model reservations are made under the same account lock as
|
models updated within seven days. If an account's discovered limit is 100, 200,
|
||||||
capacity reservations, so concurrent submissions cannot cross position 80.
|
or 500, older models stop at positions 90, 190, or 490 respectively. Old-model
|
||||||
Once every account has at least 80 active tasks, discovery is capped at the
|
reservations are made under the same account lock as capacity reservations, so
|
||||||
seven-day window and only recent models can fill positions 81–100. Unknown model
|
concurrent submissions cannot enter the reserved suffix. Once every account's
|
||||||
timestamps are treated as old for new submissions.
|
old-model allowance is exhausted, discovery is capped at the seven-day window.
|
||||||
|
Unknown model timestamps are treated as old for new submissions. When capacity
|
||||||
|
probing raises an account's known limit, its old-model boundary moves with it.
|
||||||
|
|
||||||
Queue cleanup applies a two-level version of the policy. On poller cycle 1 it
|
On startup, the worker first stops deterministic OOM tasks, recalculates each
|
||||||
first stops deterministic OOM tasks, recalculates each account's surviving task
|
account's surviving task order by numeric task ID, and removes
|
||||||
order by numeric task ID, and removes older-than-seven-days tasks after position
|
older-than-seven-days tasks beyond that account's current limit minus 10. Later
|
||||||
80. On later scheduled cleanup cycles, the age threshold relaxes to 95: old
|
scheduled cleanup uses limit minus 5, retaining a small hysteresis buffer that
|
||||||
tasks in positions 81-95 are retained and only old tasks at position 96 or later
|
avoids repeatedly over-pruning valid work. For a 100-task account the two
|
||||||
are removed. New old-model submissions are still blocked at position 80, so the
|
boundaries are 90 and 95; for a 500-task account they are 490 and 495. Recent
|
||||||
upper queue remains available to recent models without repeatedly over-pruning
|
overflow tasks are always kept. ModelScope metadata failures fail closed and
|
||||||
existing work. Recent overflow tasks are always kept. ModelScope metadata
|
never trigger cancellation. Immediately before mutation, task ownership, active
|
||||||
failures fail closed and never trigger cancellation. Immediately before
|
status, and post-OOM queue position are checked again.
|
||||||
mutation, task ownership, active status, and post-OOM queue position are checked
|
|
||||||
again.
|
|
||||||
|
|
||||||
The verified capacities, safe repository-size boundaries, evidence hierarchy,
|
The verified capacities, safe repository-size boundaries, evidence hierarchy,
|
||||||
and source links are recorded in
|
and source links are recorded in
|
||||||
@@ -244,12 +244,15 @@ Version `2026.08.11.3` adds atomic per-account 80/7-day admission, prevents
|
|||||||
history expansion once old-model positions are exhausted, performs a strict
|
history expansion once old-model positions are exhausted, performs a strict
|
||||||
position-80 startup cleanup, and relaxes scheduled age cleanup to position 95
|
position-80 startup cleanup, and relaxes scheduled age cleanup to position 95
|
||||||
after OOM cleanup and a second queue check.
|
after OOM cleanup and a second queue check.
|
||||||
|
Version `2026.08.11.4` replaces fixed queue positions with per-account dynamic
|
||||||
|
boundaries derived from each discovered capacity: limit minus 10 for admission
|
||||||
|
and startup cleanup, then limit minus 5 for scheduled dynamic cleanup.
|
||||||
|
|
||||||
## Deploy
|
## Deploy
|
||||||
|
|
||||||
Create a tag and submit the repository URL plus tag in "我的适配智能体".
|
Create a tag and submit the repository URL plus tag in "我的适配智能体".
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
git tag agent-v17
|
git tag agent-v18
|
||||||
git push origin agent-v17
|
git push origin agent-v18
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -120,12 +120,12 @@ bash run_poll.sh --dry-run
|
|||||||
memory rule as new submissions. Only tasks whose own repository size times `1.20` exceeds
|
memory rule as new submissions. Only tasks whose own repository size times `1.20` exceeds
|
||||||
their selected GPU capacity are stopped, after a fresh account-scoped active-state check.
|
their selected GPU capacity are stopped, after a fresh account-scoped active-state check.
|
||||||
Incomplete size/capacity evidence is never used for cancellation.
|
Incomplete size/capacity evidence is never used for cancellation.
|
||||||
- Models older than seven days may occupy only positions 1–80 of each account's active queue.
|
- Models older than seven days may occupy only the current account capacity minus its final 10
|
||||||
The account pool enforces the boundary atomically. Once all accounts reach 80, discovery no
|
positions. The account pool enforces this per-account boundary atomically and updates it when
|
||||||
longer expands beyond seven days. Cleanup stops OOM tasks first, recalculates the surviving
|
capacity probing discovers a higher limit. Cleanup stops OOM tasks first, recalculates the
|
||||||
queue order, and on startup stops old tasks still ranked 81 or later. Scheduled cleanup then
|
surviving queue order, and applies the same limit-minus-10 boundary on startup. Scheduled
|
||||||
retains positions 81–95 and stops only old tasks ranked 96 or later. Recent overflow tasks stay.
|
cleanup relaxes to limit minus 5 to avoid excessive pruning. Recent overflow tasks stay, and
|
||||||
Unknown ModelScope timestamps never authorize a cancellation.
|
unknown ModelScope timestamps never authorize a cancellation.
|
||||||
|
|
||||||
## Important Flags
|
## Important Flags
|
||||||
|
|
||||||
@@ -161,10 +161,11 @@ Common flags:
|
|||||||
- `--max-cycles`: optional hard stop for testing or batch windows
|
- `--max-cycles`: optional hard stop for testing or batch windows
|
||||||
- `--disable-queue-cleanup`: disable automatic OOM and old-overflow queue cleanup
|
- `--disable-queue-cleanup`: disable automatic OOM and old-overflow queue cleanup
|
||||||
|
|
||||||
The first automatic cleanup removes models older than seven days after each
|
The first automatic cleanup removes models older than seven days beyond each
|
||||||
account's position 80. Later scheduled cleanups use position 95, while admission
|
account's discovered capacity minus 10. Later scheduled cleanups use capacity
|
||||||
continues to reserve positions 81-100 for recent models. Override the later
|
minus 5, while admission continues to reserve the final 10 slots for recent
|
||||||
threshold with `MODELHUB_DYNAMIC_OLD_MODEL_CLEANUP_THRESHOLD`.
|
models. Override these suffix sizes with `MODELHUB_RECENT_MODEL_RESERVE_SLOTS`
|
||||||
|
and `MODELHUB_DYNAMIC_OLD_MODEL_CLEANUP_RESERVE_SLOTS`.
|
||||||
|
|
||||||
Failure-informed preflight is enabled by default. It rejects deterministic
|
Failure-informed preflight is enabled by default. It rejects deterministic
|
||||||
missing-file and predicted-OOM cases, clamps unsafe context-length arguments,
|
missing-file and predicted-OOM cases, clamps unsafe context-length arguments,
|
||||||
|
|||||||
@@ -86,9 +86,9 @@ def build_parser() -> argparse.ArgumentParser:
|
|||||||
help="Maximum tasks to submit in one run (0 means unlimited)",
|
help="Maximum tasks to submit in one run (0 means unlimited)",
|
||||||
)
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--old-model-queue-threshold",
|
"--recent-model-reserve-slots",
|
||||||
type=int,
|
type=int,
|
||||||
default=int(os.getenv("MODELHUB_OLD_MODEL_QUEUE_THRESHOLD", "80")),
|
default=int(os.getenv("MODELHUB_RECENT_MODEL_RESERVE_SLOTS", "10")),
|
||||||
help=argparse.SUPPRESS,
|
help=argparse.SUPPRESS,
|
||||||
)
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
@@ -252,7 +252,7 @@ def make_wave_namespace(base_args: argparse.Namespace, wave: WaveSpec) -> argpar
|
|||||||
capacity_probe_interval_cycles=getattr(base_args, "capacity_probe_interval_cycles", 3),
|
capacity_probe_interval_cycles=getattr(base_args, "capacity_probe_interval_cycles", 3),
|
||||||
submit_concurrency=getattr(base_args, "submit_concurrency", 1),
|
submit_concurrency=getattr(base_args, "submit_concurrency", 1),
|
||||||
max_submits_per_run=getattr(base_args, "max_submits_per_run", 0),
|
max_submits_per_run=getattr(base_args, "max_submits_per_run", 0),
|
||||||
old_model_queue_threshold=getattr(base_args, "old_model_queue_threshold", 80),
|
recent_model_reserve_slots=getattr(base_args, "recent_model_reserve_slots", 10),
|
||||||
recent_model_days=getattr(base_args, "recent_model_days", 7),
|
recent_model_days=getattr(base_args, "recent_model_days", 7),
|
||||||
disable_candidate_preflight=getattr(base_args, "disable_candidate_preflight", False),
|
disable_candidate_preflight=getattr(base_args, "disable_candidate_preflight", False),
|
||||||
llm_classifier_endpoint=getattr(base_args, "llm_classifier_endpoint", None),
|
llm_classifier_endpoint=getattr(base_args, "llm_classifier_endpoint", None),
|
||||||
|
|||||||
@@ -109,9 +109,9 @@ def build_parser() -> argparse.ArgumentParser:
|
|||||||
help="Maximum tasks to submit in one run (0 means unlimited)",
|
help="Maximum tasks to submit in one run (0 means unlimited)",
|
||||||
)
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--old-model-queue-threshold",
|
"--recent-model-reserve-slots",
|
||||||
type=int,
|
type=int,
|
||||||
default=int(os.getenv("MODELHUB_OLD_MODEL_QUEUE_THRESHOLD", "80")),
|
default=int(os.getenv("MODELHUB_RECENT_MODEL_RESERVE_SLOTS", "10")),
|
||||||
help=argparse.SUPPRESS,
|
help=argparse.SUPPRESS,
|
||||||
)
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
@@ -839,7 +839,7 @@ def run_submission(
|
|||||||
clients,
|
clients,
|
||||||
capacity_probe_interval_cycles=max(0, int(getattr(args, "capacity_probe_interval_cycles", 3) or 0)),
|
capacity_probe_interval_cycles=max(0, int(getattr(args, "capacity_probe_interval_cycles", 3) or 0)),
|
||||||
capacity_state_path=Path(getattr(args, "capacity_state_path", DEFAULT_CAPACITY_STATE_PATH)),
|
capacity_state_path=Path(getattr(args, "capacity_state_path", DEFAULT_CAPACITY_STATE_PATH)),
|
||||||
old_model_queue_threshold=max(1, int(getattr(args, "old_model_queue_threshold", 80) or 80)),
|
recent_model_reserve_slots=max(0, int(getattr(args, "recent_model_reserve_slots", 10) or 0)),
|
||||||
recent_model_days=max(1, int(getattr(args, "recent_model_days", 7) or 7)),
|
recent_model_days=max(1, int(getattr(args, "recent_model_days", 7) or 7)),
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -1108,13 +1108,17 @@ def run_submission(
|
|||||||
pipeline_tags = pipeline_tags_for_task_types(selected_task_types)
|
pipeline_tags = pipeline_tags_for_task_types(selected_task_types)
|
||||||
explicit_scan_cap = max(0, int(getattr(args, "max_scan_models", 0) or 0))
|
explicit_scan_cap = max(0, int(getattr(args, "max_scan_models", 0) or 0))
|
||||||
recent_model_days = max(1, int(getattr(args, "recent_model_days", 7) or 7))
|
recent_model_days = max(1, int(getattr(args, "recent_model_days", 7) or 7))
|
||||||
old_model_queue_threshold = max(1, int(getattr(args, "old_model_queue_threshold", 80) or 80))
|
recent_model_reserve_slots = max(0, int(getattr(args, "recent_model_reserve_slots", 10) or 0))
|
||||||
old_model_slots_before_scan: int | None = None
|
old_model_slots_before_scan: int | None = None
|
||||||
if hasattr(modelhub_client, "old_model_submit_slots"):
|
if hasattr(modelhub_client, "old_model_submit_slots"):
|
||||||
old_model_slots_before_scan = int(modelhub_client.old_model_submit_slots())
|
old_model_slots_before_scan = int(modelhub_client.old_model_submit_slots())
|
||||||
|
old_model_queue_thresholds: list[int] | None = None
|
||||||
|
if hasattr(modelhub_client, "old_model_queue_thresholds"):
|
||||||
|
old_model_queue_thresholds = list(modelhub_client.old_model_queue_thresholds())
|
||||||
allow_older_models_for_scan = old_model_slots_before_scan is None or old_model_slots_before_scan > 0
|
allow_older_models_for_scan = old_model_slots_before_scan is None or old_model_slots_before_scan > 0
|
||||||
print(
|
print(
|
||||||
f"[age-policy] recent_days={recent_model_days} account_threshold={old_model_queue_threshold} "
|
f"[age-policy] recent_days={recent_model_days} reserve_recent_slots={recent_model_reserve_slots} "
|
||||||
|
f"account_thresholds={','.join(str(value) for value in old_model_queue_thresholds) if old_model_queue_thresholds is not None else 'n/a'} "
|
||||||
f"old_model_slots={old_model_slots_before_scan if old_model_slots_before_scan is not None else 'n/a'} "
|
f"old_model_slots={old_model_slots_before_scan if old_model_slots_before_scan is not None else 'n/a'} "
|
||||||
f"scan_older={'on' if allow_older_models_for_scan else 'off'}",
|
f"scan_older={'on' if allow_older_models_for_scan else 'off'}",
|
||||||
flush=True,
|
flush=True,
|
||||||
@@ -1425,7 +1429,8 @@ def run_submission(
|
|||||||
"candidatePreflight": preflight_summary,
|
"candidatePreflight": preflight_summary,
|
||||||
"agePolicy": {
|
"agePolicy": {
|
||||||
"recentModelDays": recent_model_days,
|
"recentModelDays": recent_model_days,
|
||||||
"oldModelSubmitThreshold": old_model_queue_threshold,
|
"recentModelReserveSlots": recent_model_reserve_slots,
|
||||||
|
"oldModelSubmitThresholds": old_model_queue_thresholds,
|
||||||
"oldModelSlotsBeforeScan": old_model_slots_before_scan,
|
"oldModelSlotsBeforeScan": old_model_slots_before_scan,
|
||||||
"olderHistoryScanEnabled": allow_older_models_for_scan,
|
"olderHistoryScanEnabled": allow_older_models_for_scan,
|
||||||
},
|
},
|
||||||
|
|||||||
@@ -389,7 +389,7 @@ class ModelHubClientPool:
|
|||||||
capacity_probe_interval_cycles: int = 3,
|
capacity_probe_interval_cycles: int = 3,
|
||||||
capacity_probe_cooldown_cycles: int = 3,
|
capacity_probe_cooldown_cycles: int = 3,
|
||||||
capacity_state_path: Path | str | None = None,
|
capacity_state_path: Path | str | None = None,
|
||||||
old_model_queue_threshold: int | None = None,
|
recent_model_reserve_slots: int | None = None,
|
||||||
recent_model_days: int | None = None,
|
recent_model_days: int | None = None,
|
||||||
) -> None:
|
) -> None:
|
||||||
if not clients:
|
if not clients:
|
||||||
@@ -397,17 +397,17 @@ class ModelHubClientPool:
|
|||||||
self.clients = clients
|
self.clients = clients
|
||||||
configured_cap = active_task_cap if active_task_cap is not None else os.getenv("MODELHUB_AGENT_ACTIVE_TASK_CAP", "100")
|
configured_cap = active_task_cap if active_task_cap is not None else os.getenv("MODELHUB_AGENT_ACTIVE_TASK_CAP", "100")
|
||||||
self.active_task_cap = max(1, int(configured_cap))
|
self.active_task_cap = max(1, int(configured_cap))
|
||||||
configured_old_threshold = (
|
configured_recent_reserve = (
|
||||||
old_model_queue_threshold
|
recent_model_reserve_slots
|
||||||
if old_model_queue_threshold is not None
|
if recent_model_reserve_slots is not None
|
||||||
else os.getenv("MODELHUB_OLD_MODEL_QUEUE_THRESHOLD", "80")
|
else os.getenv("MODELHUB_RECENT_MODEL_RESERVE_SLOTS", "10")
|
||||||
)
|
)
|
||||||
configured_recent_days = (
|
configured_recent_days = (
|
||||||
recent_model_days
|
recent_model_days
|
||||||
if recent_model_days is not None
|
if recent_model_days is not None
|
||||||
else os.getenv("MODELHUB_RECENT_MODEL_DAYS", "7")
|
else os.getenv("MODELHUB_RECENT_MODEL_DAYS", "7")
|
||||||
)
|
)
|
||||||
self.old_model_queue_threshold = max(1, int(configured_old_threshold))
|
self.recent_model_reserve_slots = max(0, int(configured_recent_reserve))
|
||||||
self.recent_model_days = max(1, int(configured_recent_days))
|
self.recent_model_days = max(1, int(configured_recent_days))
|
||||||
self._capacity_state_path = Path(capacity_state_path) if capacity_state_path else None
|
self._capacity_state_path = Path(capacity_state_path) if capacity_state_path else None
|
||||||
self._account_keys = [self._account_key(client, index) for index, client in enumerate(clients)]
|
self._account_keys = [self._account_key(client, index) for index, client in enumerate(clients)]
|
||||||
@@ -491,6 +491,22 @@ class ModelHubClientPool:
|
|||||||
with self._state_lock:
|
with self._state_lock:
|
||||||
return list(self._account_caps)
|
return list(self._account_caps)
|
||||||
|
|
||||||
|
def observe_capacity_lower_bounds(self, active_counts: list[int | None]) -> list[int]:
|
||||||
|
"""Promote known caps from complete account listings without guessing an upper bound."""
|
||||||
|
caps_changed = False
|
||||||
|
with self._state_lock:
|
||||||
|
for index, value in enumerate(active_counts[: len(self._account_caps)]):
|
||||||
|
if value is None:
|
||||||
|
continue
|
||||||
|
observed_count = max(0, int(value))
|
||||||
|
if observed_count > self._account_caps[index]:
|
||||||
|
self._account_caps[index] = observed_count
|
||||||
|
caps_changed = True
|
||||||
|
result = list(self._account_caps)
|
||||||
|
if caps_changed:
|
||||||
|
self._persist_account_caps()
|
||||||
|
return result
|
||||||
|
|
||||||
def capacity_probe_enabled(self) -> bool:
|
def capacity_probe_enabled(self) -> bool:
|
||||||
with self._state_lock:
|
with self._state_lock:
|
||||||
return self._capacity_probe_enabled
|
return self._capacity_probe_enabled
|
||||||
@@ -693,7 +709,7 @@ class ModelHubClientPool:
|
|||||||
self,
|
self,
|
||||||
excluded: set[int],
|
excluded: set[int],
|
||||||
*,
|
*,
|
||||||
effective_count_below: int | None = None,
|
reserve_capacity_slots: int | None = None,
|
||||||
) -> tuple[int, int] | None:
|
) -> tuple[int, int] | None:
|
||||||
with self._state_lock:
|
with self._state_lock:
|
||||||
remaining_by_index = {
|
remaining_by_index = {
|
||||||
@@ -701,8 +717,9 @@ class ModelHubClientPool:
|
|||||||
for index in range(len(self.clients))
|
for index in range(len(self.clients))
|
||||||
if index not in excluded
|
if index not in excluded
|
||||||
and (
|
and (
|
||||||
effective_count_below is None
|
reserve_capacity_slots is None
|
||||||
or self._effective_count_locked(index) < effective_count_below
|
or self._effective_count_locked(index)
|
||||||
|
< max(0, self._account_caps[index] - reserve_capacity_slots)
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
usable = {index: remaining for index, remaining in remaining_by_index.items() if remaining > 0}
|
usable = {index: remaining for index, remaining in remaining_by_index.items() if remaining > 0}
|
||||||
@@ -727,7 +744,7 @@ class ModelHubClientPool:
|
|||||||
self,
|
self,
|
||||||
excluded: set[int],
|
excluded: set[int],
|
||||||
*,
|
*,
|
||||||
effective_count_below: int | None = None,
|
reserve_capacity_slots: int | None = None,
|
||||||
) -> tuple[int, int] | None:
|
) -> tuple[int, int] | None:
|
||||||
with self._state_lock:
|
with self._state_lock:
|
||||||
if not self._capacity_probe_enabled:
|
if not self._capacity_probe_enabled:
|
||||||
@@ -740,8 +757,9 @@ class ModelHubClientPool:
|
|||||||
and self._capacity_probe_cycle >= self._capacity_probe_cooldown_until[index]
|
and self._capacity_probe_cycle >= self._capacity_probe_cooldown_until[index]
|
||||||
and self._effective_count_locked(index) >= self._account_caps[index]
|
and self._effective_count_locked(index) >= self._account_caps[index]
|
||||||
and (
|
and (
|
||||||
effective_count_below is None
|
reserve_capacity_slots is None
|
||||||
or self._effective_count_locked(index) < effective_count_below
|
or self._effective_count_locked(index)
|
||||||
|
< max(0, self._account_caps[index] - reserve_capacity_slots)
|
||||||
)
|
)
|
||||||
]
|
]
|
||||||
if not eligible:
|
if not eligible:
|
||||||
@@ -806,12 +824,19 @@ class ModelHubClientPool:
|
|||||||
return sum(
|
return sum(
|
||||||
max(
|
max(
|
||||||
0,
|
0,
|
||||||
min(self._account_caps[index], self.old_model_queue_threshold)
|
self._old_model_count_limit_locked(index) - self._effective_count_locked(index),
|
||||||
- self._effective_count_locked(index),
|
|
||||||
)
|
)
|
||||||
for index in range(len(self.clients))
|
for index in range(len(self.clients))
|
||||||
)
|
)
|
||||||
|
|
||||||
|
def _old_model_count_limit_locked(self, index: int) -> int:
|
||||||
|
return max(0, self._account_caps[index] - self.recent_model_reserve_slots)
|
||||||
|
|
||||||
|
def old_model_queue_thresholds(self) -> list[int]:
|
||||||
|
"""Return each account's current cap minus its recent-model reserve."""
|
||||||
|
with self._state_lock:
|
||||||
|
return [self._old_model_count_limit_locked(index) for index in range(len(self.clients))]
|
||||||
|
|
||||||
def can_submit_old_models(self) -> bool:
|
def can_submit_old_models(self) -> bool:
|
||||||
return self.old_model_submit_slots() > 0
|
return self.old_model_submit_slots() > 0
|
||||||
|
|
||||||
@@ -838,18 +863,18 @@ class ModelHubClientPool:
|
|||||||
attempted_accounts: set[int] = set()
|
attempted_accounts: set[int] = set()
|
||||||
forced_refresh_done = False
|
forced_refresh_done = False
|
||||||
last_capacity_error: ModelHubAPIError | None = None
|
last_capacity_error: ModelHubAPIError | None = None
|
||||||
account_count_limit = self.old_model_queue_threshold if old_model_only else None
|
reserved_capacity_slots = self.recent_model_reserve_slots if old_model_only else None
|
||||||
|
|
||||||
while True:
|
while True:
|
||||||
reservation = self._reserve_account(
|
reservation = self._reserve_account(
|
||||||
attempted_accounts,
|
attempted_accounts,
|
||||||
effective_count_below=account_count_limit,
|
reserve_capacity_slots=reserved_capacity_slots,
|
||||||
)
|
)
|
||||||
capacity_probe = False
|
capacity_probe = False
|
||||||
if reservation is None:
|
if reservation is None:
|
||||||
reservation = self._reserve_probe_account(
|
reservation = self._reserve_probe_account(
|
||||||
attempted_accounts,
|
attempted_accounts,
|
||||||
effective_count_below=account_count_limit,
|
reserve_capacity_slots=reserved_capacity_slots,
|
||||||
)
|
)
|
||||||
capacity_probe = reservation is not None
|
capacity_probe = reservation is not None
|
||||||
if reservation is None:
|
if reservation is None:
|
||||||
@@ -860,8 +885,10 @@ class ModelHubClientPool:
|
|||||||
if last_capacity_error is not None:
|
if last_capacity_error is not None:
|
||||||
raise last_capacity_error
|
raise last_capacity_error
|
||||||
if old_model_only:
|
if old_model_only:
|
||||||
|
thresholds = self.old_model_queue_thresholds()
|
||||||
raise OldModelQueuePolicyError(
|
raise OldModelQueuePolicyError(
|
||||||
f"所有账号活动队列均已达到 {self.old_model_queue_threshold},"
|
f"所有账号均已达到动态旧模型阈值 {thresholds}(账号上限减 "
|
||||||
|
f"{self.recent_model_reserve_slots}),"
|
||||||
f"仅允许提交最近 {self.recent_model_days} 天内更新的模型"
|
f"仅允许提交最近 {self.recent_model_days} 天内更新的模型"
|
||||||
)
|
)
|
||||||
raise ModelHubAPIError(
|
raise ModelHubAPIError(
|
||||||
|
|||||||
@@ -59,9 +59,9 @@ def build_parser() -> argparse.ArgumentParser:
|
|||||||
help="Maximum tasks to submit in one cycle (0 means unlimited)",
|
help="Maximum tasks to submit in one cycle (0 means unlimited)",
|
||||||
)
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--old-model-queue-threshold",
|
"--recent-model-reserve-slots",
|
||||||
type=int,
|
type=int,
|
||||||
default=int(os.getenv("MODELHUB_OLD_MODEL_QUEUE_THRESHOLD", "80")),
|
default=int(os.getenv("MODELHUB_RECENT_MODEL_RESERVE_SLOTS", "10")),
|
||||||
help=argparse.SUPPRESS,
|
help=argparse.SUPPRESS,
|
||||||
)
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
@@ -71,9 +71,9 @@ def build_parser() -> argparse.ArgumentParser:
|
|||||||
help=argparse.SUPPRESS,
|
help=argparse.SUPPRESS,
|
||||||
)
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--dynamic-old-model-cleanup-threshold",
|
"--dynamic-old-model-cleanup-reserve-slots",
|
||||||
type=int,
|
type=int,
|
||||||
default=int(os.getenv("MODELHUB_DYNAMIC_OLD_MODEL_CLEANUP_THRESHOLD", "95")),
|
default=int(os.getenv("MODELHUB_DYNAMIC_OLD_MODEL_CLEANUP_RESERVE_SLOTS", "5")),
|
||||||
help=argparse.SUPPRESS,
|
help=argparse.SUPPRESS,
|
||||||
)
|
)
|
||||||
parser.add_argument("--disable-candidate-preflight", action="store_true", help=argparse.SUPPRESS)
|
parser.add_argument("--disable-candidate-preflight", action="store_true", help=argparse.SUPPRESS)
|
||||||
@@ -224,7 +224,7 @@ def _build_modelhub_client(base_args: argparse.Namespace) -> ModelHubClientPool:
|
|||||||
clients,
|
clients,
|
||||||
capacity_probe_interval_cycles=max(0, int(getattr(base_args, "capacity_probe_interval_cycles", 3) or 0)),
|
capacity_probe_interval_cycles=max(0, int(getattr(base_args, "capacity_probe_interval_cycles", 3) or 0)),
|
||||||
capacity_state_path=Path(getattr(base_args, "capacity_state_path", DEFAULT_CAPACITY_STATE_PATH)),
|
capacity_state_path=Path(getattr(base_args, "capacity_state_path", DEFAULT_CAPACITY_STATE_PATH)),
|
||||||
old_model_queue_threshold=max(1, int(getattr(base_args, "old_model_queue_threshold", 80) or 80)),
|
recent_model_reserve_slots=max(0, int(getattr(base_args, "recent_model_reserve_slots", 10) or 0)),
|
||||||
recent_model_days=max(1, int(getattr(base_args, "recent_model_days", 7) or 7)),
|
recent_model_days=max(1, int(getattr(base_args, "recent_model_days", 7) or 7)),
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -234,18 +234,21 @@ def resolve_age_cleanup_policy(
|
|||||||
*,
|
*,
|
||||||
initial_cleanup_pending: bool,
|
initial_cleanup_pending: bool,
|
||||||
) -> tuple[str, int]:
|
) -> tuple[str, int]:
|
||||||
"""Use the strict admission boundary once, then retain a 15-slot buffer."""
|
"""Reserve ten slots initially, then use a five-slot cleanup hysteresis."""
|
||||||
initial_threshold = max(
|
initial_reserve_slots = max(
|
||||||
1,
|
0,
|
||||||
int(getattr(base_args, "old_model_queue_threshold", 80) or 80),
|
int(getattr(base_args, "recent_model_reserve_slots", 10) or 0),
|
||||||
)
|
)
|
||||||
dynamic_threshold = max(
|
dynamic_reserve_slots = min(
|
||||||
initial_threshold,
|
initial_reserve_slots,
|
||||||
int(getattr(base_args, "dynamic_old_model_cleanup_threshold", 95) or 95),
|
max(
|
||||||
|
0,
|
||||||
|
int(getattr(base_args, "dynamic_old_model_cleanup_reserve_slots", 5) or 0),
|
||||||
|
),
|
||||||
)
|
)
|
||||||
if initial_cleanup_pending:
|
if initial_cleanup_pending:
|
||||||
return "initial", initial_threshold
|
return "initial", initial_reserve_slots
|
||||||
return "dynamic", dynamic_threshold
|
return "dynamic", dynamic_reserve_slots
|
||||||
|
|
||||||
|
|
||||||
def run_poll_loop(
|
def run_poll_loop(
|
||||||
@@ -303,13 +306,13 @@ def run_poll_loop(
|
|||||||
try:
|
try:
|
||||||
cleanup_feedback = outcome_tracker.get_stats_report()
|
cleanup_feedback = outcome_tracker.get_stats_report()
|
||||||
cleanup_gpu_memory = cleanup_feedback.get("observedGpuMemoryGiB") or {}
|
cleanup_gpu_memory = cleanup_feedback.get("observedGpuMemoryGiB") or {}
|
||||||
age_cleanup_mode, age_cleanup_threshold = resolve_age_cleanup_policy(
|
age_cleanup_mode, age_cleanup_reserve_slots = resolve_age_cleanup_policy(
|
||||||
base_args,
|
base_args,
|
||||||
initial_cleanup_pending=initial_age_cleanup_pending,
|
initial_cleanup_pending=initial_age_cleanup_pending,
|
||||||
)
|
)
|
||||||
log(
|
log(
|
||||||
f"[queue-cleanup] mode={age_cleanup_mode} "
|
f"[queue-cleanup] mode={age_cleanup_mode} "
|
||||||
f"old_model_threshold={age_cleanup_threshold} "
|
f"reserve_recent_slots={age_cleanup_reserve_slots} "
|
||||||
f"recent_days={max(1, int(getattr(base_args, 'recent_model_days', 7) or 7))}"
|
f"recent_days={max(1, int(getattr(base_args, 'recent_model_days', 7) or 7))}"
|
||||||
)
|
)
|
||||||
cleanup_summary = cleanup_certain_oom_tasks(
|
cleanup_summary = cleanup_certain_oom_tasks(
|
||||||
@@ -318,7 +321,7 @@ def run_poll_loop(
|
|||||||
dry_run=bool(base_args.dry_run),
|
dry_run=bool(base_args.dry_run),
|
||||||
read_concurrency=max(1, int(getattr(base_args, "queue_cleanup_read_concurrency", 6) or 6)),
|
read_concurrency=max(1, int(getattr(base_args, "queue_cleanup_read_concurrency", 6) or 6)),
|
||||||
gpu_memory_gib=cleanup_gpu_memory if isinstance(cleanup_gpu_memory, dict) else None,
|
gpu_memory_gib=cleanup_gpu_memory if isinstance(cleanup_gpu_memory, dict) else None,
|
||||||
age_queue_threshold=age_cleanup_threshold,
|
age_reserved_slots=age_cleanup_reserve_slots,
|
||||||
log=log,
|
log=log,
|
||||||
)
|
)
|
||||||
write_json(
|
write_json(
|
||||||
@@ -335,7 +338,8 @@ def run_poll_loop(
|
|||||||
{
|
{
|
||||||
"cycle": cycles,
|
"cycle": cycles,
|
||||||
"mode": age_cleanup_mode,
|
"mode": age_cleanup_mode,
|
||||||
"ageQueueThreshold": age_cleanup_threshold,
|
"ageReservedSlots": age_cleanup_reserve_slots,
|
||||||
|
"ageQueueThresholds": cleanup_summary["oldModelQueueThresholds"],
|
||||||
"activeScanned": cleanup_summary["activeScanned"],
|
"activeScanned": cleanup_summary["activeScanned"],
|
||||||
"certainOomCount": cleanup_summary["certainOomCount"],
|
"certainOomCount": cleanup_summary["certainOomCount"],
|
||||||
"oldOverflowCount": cleanup_summary["oldOverflowCount"],
|
"oldOverflowCount": cleanup_summary["oldOverflowCount"],
|
||||||
|
|||||||
@@ -242,12 +242,11 @@ def find_old_overflow_tasks(
|
|||||||
tasks: list[OwnedTask],
|
tasks: list[OwnedTask],
|
||||||
*,
|
*,
|
||||||
model_last_modified: dict[str, datetime],
|
model_last_modified: dict[str, datetime],
|
||||||
queue_threshold: int = 80,
|
queue_threshold: int | dict[int, int] = 80,
|
||||||
recent_model_days: int = 7,
|
recent_model_days: int = 7,
|
||||||
reference_time: datetime | None = None,
|
reference_time: datetime | None = None,
|
||||||
incomplete_accounts: set[int] | None = None,
|
incomplete_accounts: set[int] | None = None,
|
||||||
) -> tuple[list[dict[str, Any]], dict[str, int]]:
|
) -> tuple[list[dict[str, Any]], dict[str, int]]:
|
||||||
threshold = max(1, int(queue_threshold))
|
|
||||||
recent_days = max(1, int(recent_model_days))
|
recent_days = max(1, int(recent_model_days))
|
||||||
cutoff = (reference_time or utc_now()) - timedelta(days=recent_days)
|
cutoff = (reference_time or utc_now()) - timedelta(days=recent_days)
|
||||||
incomplete_accounts = incomplete_accounts or set()
|
incomplete_accounts = incomplete_accounts or set()
|
||||||
@@ -261,10 +260,19 @@ def find_old_overflow_tasks(
|
|||||||
"withinFirstQueuePositions": 0,
|
"withinFirstQueuePositions": 0,
|
||||||
"recentOverflowTasks": 0,
|
"recentOverflowTasks": 0,
|
||||||
"modelAgeUnknown": 0,
|
"modelAgeUnknown": 0,
|
||||||
|
"accountThresholdUnknown": 0,
|
||||||
}
|
}
|
||||||
for account_index, account_tasks in sorted(grouped.items()):
|
for account_index, account_tasks in sorted(grouped.items()):
|
||||||
if account_index in incomplete_accounts:
|
if account_index in incomplete_accounts:
|
||||||
continue
|
continue
|
||||||
|
if isinstance(queue_threshold, dict):
|
||||||
|
configured_threshold = queue_threshold.get(account_index)
|
||||||
|
if configured_threshold is None:
|
||||||
|
skipped["accountThresholdUnknown"] += 1
|
||||||
|
continue
|
||||||
|
else:
|
||||||
|
configured_threshold = queue_threshold
|
||||||
|
threshold = max(0, int(configured_threshold))
|
||||||
ordered = sorted(account_tasks, key=lambda item: item.task_id)
|
ordered = sorted(account_tasks, key=lambda item: item.task_id)
|
||||||
skipped["withinFirstQueuePositions"] += min(threshold, len(ordered))
|
skipped["withinFirstQueuePositions"] += min(threshold, len(ordered))
|
||||||
for position, task in enumerate(ordered, start=1):
|
for position, task in enumerate(ordered, start=1):
|
||||||
@@ -307,19 +315,19 @@ def cleanup_certain_oom_tasks(
|
|||||||
read_concurrency: int = 6,
|
read_concurrency: int = 6,
|
||||||
stop_batch_size: int = DEFAULT_STOP_BATCH_SIZE,
|
stop_batch_size: int = DEFAULT_STOP_BATCH_SIZE,
|
||||||
gpu_memory_gib: dict[str, float] | None = None,
|
gpu_memory_gib: dict[str, float] | None = None,
|
||||||
age_queue_threshold: int | None = None,
|
age_reserved_slots: int | None = None,
|
||||||
reference_time: datetime | None = None,
|
reference_time: datetime | None = None,
|
||||||
log: Callable[[str], None] = print,
|
log: Callable[[str], None] = print,
|
||||||
) -> dict[str, Any]:
|
) -> dict[str, Any]:
|
||||||
"""Stop deterministic OOM tasks and old tasks beyond each account's protected prefix."""
|
"""Stop deterministic OOM tasks and old tasks beyond each account's protected prefix."""
|
||||||
clients = list(modelhub.clients)
|
clients = list(modelhub.clients)
|
||||||
reference_time = reference_time or utc_now()
|
reference_time = reference_time or utc_now()
|
||||||
configured_queue_threshold = (
|
configured_reserved_slots = (
|
||||||
age_queue_threshold
|
age_reserved_slots
|
||||||
if age_queue_threshold is not None
|
if age_reserved_slots is not None
|
||||||
else getattr(modelhub, "old_model_queue_threshold", 80)
|
else getattr(modelhub, "recent_model_reserve_slots", 10)
|
||||||
)
|
)
|
||||||
queue_threshold = max(1, int(configured_queue_threshold or 80))
|
reserved_slots = max(0, int(configured_reserved_slots or 0))
|
||||||
recent_model_days = max(1, int(getattr(modelhub, "recent_model_days", 7) or 7))
|
recent_model_days = max(1, int(getattr(modelhub, "recent_model_days", 7) or 7))
|
||||||
tasks, listing_errors = collect_active_tasks(clients, read_concurrency=read_concurrency)
|
tasks, listing_errors = collect_active_tasks(clients, read_concurrency=read_concurrency)
|
||||||
log(
|
log(
|
||||||
@@ -327,6 +335,25 @@ def cleanup_certain_oom_tasks(
|
|||||||
f"listing_errors={sum(len(items) for items in listing_errors.values())}"
|
f"listing_errors={sum(len(items) for items in listing_errors.values())}"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
observed_active_counts: list[int | None] = [0 for _ in clients]
|
||||||
|
for task in tasks:
|
||||||
|
current_count = observed_active_counts[task.account_index]
|
||||||
|
if current_count is not None:
|
||||||
|
observed_active_counts[task.account_index] = current_count + 1
|
||||||
|
for account_index in listing_errors:
|
||||||
|
observed_active_counts[account_index] = None
|
||||||
|
if hasattr(modelhub, "observe_capacity_lower_bounds"):
|
||||||
|
account_caps = list(modelhub.observe_capacity_lower_bounds(observed_active_counts))
|
||||||
|
elif hasattr(modelhub, "account_capacity_limits"):
|
||||||
|
account_caps = list(modelhub.account_capacity_limits())
|
||||||
|
else:
|
||||||
|
default_cap = max(1, int(getattr(modelhub, "active_task_cap", 100) or 100))
|
||||||
|
account_caps = [default_cap for _ in clients]
|
||||||
|
queue_thresholds = {
|
||||||
|
index: max(0, int(account_caps[index]) - reserved_slots)
|
||||||
|
for index in range(min(len(clients), len(account_caps)))
|
||||||
|
}
|
||||||
|
|
||||||
model_ids = {task.model_id for task in tasks}
|
model_ids = {task.model_id for task in tasks}
|
||||||
repository_sizes, size_errors = _load_repository_sizes(
|
repository_sizes, size_errors = _load_repository_sizes(
|
||||||
model_ids,
|
model_ids,
|
||||||
@@ -362,7 +389,7 @@ def cleanup_certain_oom_tasks(
|
|||||||
for task in sorted(
|
for task in sorted(
|
||||||
(item for item in age_rank_tasks if item.account_index == account_index),
|
(item for item in age_rank_tasks if item.account_index == account_index),
|
||||||
key=lambda item: item.task_id,
|
key=lambda item: item.task_id,
|
||||||
)[queue_threshold:]
|
)[queue_thresholds.get(account_index, len(age_rank_tasks)):]
|
||||||
}
|
}
|
||||||
model_last_modified, age_errors = _load_model_last_modified(
|
model_last_modified, age_errors = _load_model_last_modified(
|
||||||
overflow_model_ids,
|
overflow_model_ids,
|
||||||
@@ -373,14 +400,15 @@ def cleanup_certain_oom_tasks(
|
|||||||
old_overflow_decisions, age_skipped = find_old_overflow_tasks(
|
old_overflow_decisions, age_skipped = find_old_overflow_tasks(
|
||||||
age_rank_tasks,
|
age_rank_tasks,
|
||||||
model_last_modified=model_last_modified,
|
model_last_modified=model_last_modified,
|
||||||
queue_threshold=queue_threshold,
|
queue_threshold=queue_thresholds,
|
||||||
recent_model_days=recent_model_days,
|
recent_model_days=recent_model_days,
|
||||||
reference_time=reference_time,
|
reference_time=reference_time,
|
||||||
incomplete_accounts=set(listing_errors),
|
incomplete_accounts=set(listing_errors),
|
||||||
)
|
)
|
||||||
log(
|
log(
|
||||||
f"[queue-cleanup] old_overflow={len(old_overflow_decisions)} "
|
f"[queue-cleanup] old_overflow={len(old_overflow_decisions)} "
|
||||||
f"threshold={queue_threshold} recent_days={recent_model_days} "
|
f"thresholds={','.join(str(queue_thresholds[index]) for index in sorted(queue_thresholds))} "
|
||||||
|
f"reserve_recent_slots={reserved_slots} recent_days={recent_model_days} "
|
||||||
f"recent_overflow={age_skipped['recentOverflowTasks']} "
|
f"recent_overflow={age_skipped['recentOverflowTasks']} "
|
||||||
f"age_unknown={age_skipped['modelAgeUnknown']}"
|
f"age_unknown={age_skipped['modelAgeUnknown']}"
|
||||||
)
|
)
|
||||||
@@ -453,7 +481,12 @@ def cleanup_certain_oom_tasks(
|
|||||||
cleanup_reasons = set(decision.get("cleanupReasons") or [decision.get("reason")])
|
cleanup_reasons = set(decision.get("cleanupReasons") or [decision.get("reason")])
|
||||||
age_only = cleanup_reasons == {"old_model_beyond_account_queue_threshold"}
|
age_only = cleanup_reasons == {"old_model_beyond_account_queue_threshold"}
|
||||||
current_position = active_positions_by_account.get(account_index, {}).get(int(decision["taskId"]))
|
current_position = active_positions_by_account.get(account_index, {}).get(int(decision["taskId"]))
|
||||||
if age_only and (current_position is None or current_position <= queue_threshold):
|
account_queue_threshold = queue_thresholds.get(account_index)
|
||||||
|
if age_only and (
|
||||||
|
current_position is None
|
||||||
|
or account_queue_threshold is None
|
||||||
|
or current_position <= account_queue_threshold
|
||||||
|
):
|
||||||
policy_no_longer_applies.append(
|
policy_no_longer_applies.append(
|
||||||
{**decision, "recheckedQueuePosition": current_position}
|
{**decision, "recheckedQueuePosition": current_position}
|
||||||
)
|
)
|
||||||
@@ -516,7 +549,14 @@ def cleanup_certain_oom_tasks(
|
|||||||
"certainOomTasks": oom_decisions,
|
"certainOomTasks": oom_decisions,
|
||||||
"oldOverflowCount": len(old_overflow_decisions),
|
"oldOverflowCount": len(old_overflow_decisions),
|
||||||
"oldOverflowTasks": old_overflow_decisions,
|
"oldOverflowTasks": old_overflow_decisions,
|
||||||
"oldModelQueueThreshold": queue_threshold,
|
"oldModelQueueThresholds": [
|
||||||
|
queue_thresholds.get(index) for index in range(len(clients))
|
||||||
|
],
|
||||||
|
"accountCapacityLimits": [
|
||||||
|
account_caps[index] if index < len(account_caps) else None
|
||||||
|
for index in range(len(clients))
|
||||||
|
],
|
||||||
|
"recentModelReserveSlots": reserved_slots,
|
||||||
"recentModelDays": recent_model_days,
|
"recentModelDays": recent_model_days,
|
||||||
"agePolicySkipped": age_skipped,
|
"agePolicySkipped": age_skipped,
|
||||||
"cleanupCandidateCount": len(decisions),
|
"cleanupCandidateCount": len(decisions),
|
||||||
|
|||||||
@@ -1 +1 @@
|
|||||||
AGENT_VERSION = "2026.08.11.3"
|
AGENT_VERSION = "2026.08.11.4"
|
||||||
|
|||||||
@@ -215,21 +215,21 @@ def make_candidate(index: int) -> dict:
|
|||||||
|
|
||||||
|
|
||||||
class ClientPoolConcurrencyTests(unittest.TestCase):
|
class ClientPoolConcurrencyTests(unittest.TestCase):
|
||||||
def test_old_models_use_only_each_accounts_first_eighty_queue_positions(self) -> None:
|
def test_old_models_reserve_each_accounts_last_ten_queue_positions(self) -> None:
|
||||||
below_threshold = FakeClient(active_count=79)
|
below_threshold = FakeClient(active_count=89)
|
||||||
at_threshold = FakeClient(active_count=80)
|
at_threshold = FakeClient(active_count=90)
|
||||||
pool = ModelHubClientPool(
|
pool = ModelHubClientPool(
|
||||||
[below_threshold, at_threshold], # type: ignore[list-item]
|
[below_threshold, at_threshold], # type: ignore[list-item]
|
||||||
active_task_cap=100,
|
active_task_cap=100,
|
||||||
active_counts_ttl=60,
|
active_counts_ttl=60,
|
||||||
old_model_queue_threshold=80,
|
recent_model_reserve_slots=10,
|
||||||
recent_model_days=7,
|
recent_model_days=7,
|
||||||
instance_id="old-model-threshold-test",
|
instance_id="old-model-threshold-test",
|
||||||
)
|
)
|
||||||
submitted_at = datetime(2026, 8, 11, tzinfo=timezone.utc)
|
submitted_at = datetime(2026, 8, 11, tzinfo=timezone.utc)
|
||||||
|
|
||||||
pool.add_task_for_model(
|
pool.add_task_for_model(
|
||||||
{"model": "old-allowed-as-position-80"},
|
{"model": "old-allowed-as-position-90"},
|
||||||
model_last_modified=datetime(2026, 7, 1, tzinfo=timezone.utc),
|
model_last_modified=datetime(2026, 7, 1, tzinfo=timezone.utc),
|
||||||
submitted_at=submitted_at,
|
submitted_at=submitted_at,
|
||||||
)
|
)
|
||||||
@@ -251,13 +251,13 @@ class ClientPoolConcurrencyTests(unittest.TestCase):
|
|||||||
)
|
)
|
||||||
self.assertEqual(2, len(below_threshold.submitted) + len(at_threshold.submitted))
|
self.assertEqual(2, len(below_threshold.submitted) + len(at_threshold.submitted))
|
||||||
|
|
||||||
def test_concurrent_old_model_submissions_cannot_cross_eighty(self) -> None:
|
def test_concurrent_old_model_submissions_cannot_enter_reserved_ten_slots(self) -> None:
|
||||||
client = FakeClient(active_count=78)
|
client = FakeClient(active_count=88)
|
||||||
pool = ModelHubClientPool(
|
pool = ModelHubClientPool(
|
||||||
[client], # type: ignore[list-item]
|
[client], # type: ignore[list-item]
|
||||||
active_task_cap=100,
|
active_task_cap=100,
|
||||||
active_counts_ttl=60,
|
active_counts_ttl=60,
|
||||||
old_model_queue_threshold=80,
|
recent_model_reserve_slots=10,
|
||||||
recent_model_days=7,
|
recent_model_days=7,
|
||||||
)
|
)
|
||||||
submitted_at = datetime(2026, 8, 11, tzinfo=timezone.utc)
|
submitted_at = datetime(2026, 8, 11, tzinfo=timezone.utc)
|
||||||
@@ -294,12 +294,12 @@ class ClientPoolConcurrencyTests(unittest.TestCase):
|
|||||||
self.assertEqual(["configured_window", "last_7_days"], [stage["name"] for stage in stages])
|
self.assertEqual(["configured_window", "last_7_days"], [stage["name"] for stage in stages])
|
||||||
self.assertTrue(all(stage["updatedAfter"] >= now - timedelta(days=7) for stage in stages))
|
self.assertTrue(all(stage["updatedAfter"] >= now - timedelta(days=7) for stage in stages))
|
||||||
|
|
||||||
def test_submit_candidate_reports_old_model_policy_skip_at_eighty(self) -> None:
|
def test_submit_candidate_reports_old_model_policy_skip_at_dynamic_threshold(self) -> None:
|
||||||
client = FakeClient(active_count=80)
|
client = FakeClient(active_count=90)
|
||||||
pool = ModelHubClientPool(
|
pool = ModelHubClientPool(
|
||||||
[client], # type: ignore[list-item]
|
[client], # type: ignore[list-item]
|
||||||
active_task_cap=100,
|
active_task_cap=100,
|
||||||
old_model_queue_threshold=80,
|
recent_model_reserve_slots=10,
|
||||||
recent_model_days=7,
|
recent_model_days=7,
|
||||||
)
|
)
|
||||||
result = submit_candidate(
|
result = submit_candidate(
|
||||||
@@ -318,6 +318,23 @@ class ClientPoolConcurrencyTests(unittest.TestCase):
|
|||||||
self.assertEqual("old_model_policy_skipped", result["outcome"])
|
self.assertEqual("old_model_policy_skipped", result["outcome"])
|
||||||
self.assertEqual([], client.submitted)
|
self.assertEqual([], client.submitted)
|
||||||
|
|
||||||
|
def test_old_model_threshold_tracks_a_discovered_capacity_increase(self) -> None:
|
||||||
|
client = DynamicCapacityClient(active=100, limit=101)
|
||||||
|
pool = ModelHubClientPool(
|
||||||
|
[client], # type: ignore[list-item]
|
||||||
|
active_task_cap=100,
|
||||||
|
recent_model_reserve_slots=10,
|
||||||
|
capacity_probe_interval_cycles=1,
|
||||||
|
capacity_state_path=None,
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual([90], pool.old_model_queue_thresholds())
|
||||||
|
pool.configure_capacity_probe(1)
|
||||||
|
pool.add_task({"model": "recent-capacity-probe"})
|
||||||
|
|
||||||
|
self.assertEqual([101], pool.account_capacity_limits())
|
||||||
|
self.assertEqual([91], pool.old_model_queue_thresholds())
|
||||||
|
|
||||||
def test_online_submission_does_not_construct_llm_even_when_key_is_present(self) -> None:
|
def test_online_submission_does_not_construct_llm_even_when_key_is_present(self) -> None:
|
||||||
with tempfile.TemporaryDirectory() as temporary_dir:
|
with tempfile.TemporaryDirectory() as temporary_dir:
|
||||||
root = Path(temporary_dir)
|
root = Path(temporary_dir)
|
||||||
|
|||||||
@@ -15,29 +15,29 @@ from poll_runner import resolve_age_cleanup_policy # noqa: E402
|
|||||||
|
|
||||||
|
|
||||||
class PollPolicyTests(unittest.TestCase):
|
class PollPolicyTests(unittest.TestCase):
|
||||||
def test_age_cleanup_uses_eighty_once_then_ninety_five(self) -> None:
|
def test_age_cleanup_uses_minus_ten_once_then_minus_five(self) -> None:
|
||||||
args = argparse.Namespace(
|
args = argparse.Namespace(
|
||||||
old_model_queue_threshold=80,
|
recent_model_reserve_slots=10,
|
||||||
dynamic_old_model_cleanup_threshold=95,
|
dynamic_old_model_cleanup_reserve_slots=5,
|
||||||
)
|
)
|
||||||
|
|
||||||
self.assertEqual(
|
self.assertEqual(
|
||||||
("initial", 80),
|
("initial", 10),
|
||||||
resolve_age_cleanup_policy(args, initial_cleanup_pending=True),
|
resolve_age_cleanup_policy(args, initial_cleanup_pending=True),
|
||||||
)
|
)
|
||||||
self.assertEqual(
|
self.assertEqual(
|
||||||
("dynamic", 95),
|
("dynamic", 5),
|
||||||
resolve_age_cleanup_policy(args, initial_cleanup_pending=False),
|
resolve_age_cleanup_policy(args, initial_cleanup_pending=False),
|
||||||
)
|
)
|
||||||
|
|
||||||
def test_dynamic_cleanup_cannot_be_stricter_than_admission(self) -> None:
|
def test_dynamic_cleanup_cannot_be_stricter_than_initial_cleanup(self) -> None:
|
||||||
args = argparse.Namespace(
|
args = argparse.Namespace(
|
||||||
old_model_queue_threshold=80,
|
recent_model_reserve_slots=10,
|
||||||
dynamic_old_model_cleanup_threshold=70,
|
dynamic_old_model_cleanup_reserve_slots=20,
|
||||||
)
|
)
|
||||||
|
|
||||||
self.assertEqual(
|
self.assertEqual(
|
||||||
("dynamic", 80),
|
("dynamic", 10),
|
||||||
resolve_age_cleanup_policy(args, initial_cleanup_pending=False),
|
resolve_age_cleanup_policy(args, initial_cleanup_pending=False),
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|||||||
@@ -97,13 +97,18 @@ class RecordingHttpClient:
|
|||||||
|
|
||||||
|
|
||||||
class QueueCleanupTests(unittest.TestCase):
|
class QueueCleanupTests(unittest.TestCase):
|
||||||
def test_old_models_are_selected_only_after_each_accounts_first_eighty_tasks(self) -> None:
|
def test_old_models_use_each_accounts_own_capacity_minus_ten_threshold(self) -> None:
|
||||||
now = datetime(2026, 8, 11, tzinfo=timezone.utc)
|
now = datetime(2026, 8, 11, tzinfo=timezone.utc)
|
||||||
tasks = [
|
tasks = [
|
||||||
OwnedTask(0, index, "owner/old", "Iluvatar_bi-100", "waiting")
|
OwnedTask(0, index, "owner/old", "Iluvatar_bi-100", "waiting")
|
||||||
for index in range(1, 82)
|
for index in range(1, 92)
|
||||||
]
|
]
|
||||||
tasks.append(OwnedTask(0, 82, "owner/recent", "Iluvatar_bi-100", "waiting"))
|
tasks.extend(
|
||||||
|
OwnedTask(1, 1000 + index, "owner/old", "Iluvatar_bi-100", "waiting")
|
||||||
|
for index in range(1, 192)
|
||||||
|
)
|
||||||
|
tasks.append(OwnedTask(0, 92, "owner/recent", "Iluvatar_bi-100", "waiting"))
|
||||||
|
tasks.append(OwnedTask(1, 1192, "owner/recent", "Iluvatar_bi-100", "waiting"))
|
||||||
|
|
||||||
selected, skipped = find_old_overflow_tasks(
|
selected, skipped = find_old_overflow_tasks(
|
||||||
tasks,
|
tasks,
|
||||||
@@ -111,16 +116,16 @@ class QueueCleanupTests(unittest.TestCase):
|
|||||||
"owner/old": datetime(2026, 7, 1, tzinfo=timezone.utc),
|
"owner/old": datetime(2026, 7, 1, tzinfo=timezone.utc),
|
||||||
"owner/recent": datetime(2026, 8, 10, tzinfo=timezone.utc),
|
"owner/recent": datetime(2026, 8, 10, tzinfo=timezone.utc),
|
||||||
},
|
},
|
||||||
queue_threshold=80,
|
queue_threshold={0: 90, 1: 190},
|
||||||
recent_model_days=7,
|
recent_model_days=7,
|
||||||
reference_time=now,
|
reference_time=now,
|
||||||
)
|
)
|
||||||
|
|
||||||
self.assertEqual([81], [item["taskId"] for item in selected])
|
self.assertEqual([91, 1191], [item["taskId"] for item in selected])
|
||||||
self.assertEqual(81, selected[0]["queuePosition"])
|
self.assertEqual([91, 191], [item["queuePosition"] for item in selected])
|
||||||
self.assertEqual(1, skipped["recentOverflowTasks"])
|
self.assertEqual(2, skipped["recentOverflowTasks"])
|
||||||
|
|
||||||
def test_old_overflow_task_is_not_stopped_if_it_moves_into_first_eighty(self) -> None:
|
def test_old_overflow_task_is_not_stopped_if_it_moves_inside_dynamic_threshold(self) -> None:
|
||||||
records = [
|
records = [
|
||||||
{
|
{
|
||||||
"taskId": index,
|
"taskId": index,
|
||||||
@@ -128,13 +133,13 @@ class QueueCleanupTests(unittest.TestCase):
|
|||||||
"gpuType": "Iluvatar_bi-100",
|
"gpuType": "Iluvatar_bi-100",
|
||||||
"status": "waiting",
|
"status": "waiting",
|
||||||
}
|
}
|
||||||
for index in range(1, 82)
|
for index in range(1, 92)
|
||||||
]
|
]
|
||||||
client = FakeQueueClient(records, drop_first_on_recheck=True)
|
client = FakeQueueClient(records, drop_first_on_recheck=True)
|
||||||
pool = ModelHubClientPool(
|
pool = ModelHubClientPool(
|
||||||
[client], # type: ignore[list-item]
|
[client], # type: ignore[list-item]
|
||||||
active_task_cap=100,
|
active_task_cap=100,
|
||||||
old_model_queue_threshold=80,
|
recent_model_reserve_slots=10,
|
||||||
)
|
)
|
||||||
summary = cleanup_certain_oom_tasks(
|
summary = cleanup_certain_oom_tasks(
|
||||||
pool,
|
pool,
|
||||||
@@ -151,7 +156,7 @@ class QueueCleanupTests(unittest.TestCase):
|
|||||||
self.assertEqual(1, summary["policyNoLongerAppliesCount"])
|
self.assertEqual(1, summary["policyNoLongerAppliesCount"])
|
||||||
self.assertEqual([], client.stopped)
|
self.assertEqual([], client.stopped)
|
||||||
|
|
||||||
def test_cleanup_stops_old_task_beyond_eightieth_position(self) -> None:
|
def test_cleanup_promotes_capacity_from_complete_active_listing(self) -> None:
|
||||||
records = [
|
records = [
|
||||||
{
|
{
|
||||||
"taskId": index,
|
"taskId": index,
|
||||||
@@ -159,13 +164,44 @@ class QueueCleanupTests(unittest.TestCase):
|
|||||||
"gpuType": "Iluvatar_bi-100",
|
"gpuType": "Iluvatar_bi-100",
|
||||||
"status": "waiting",
|
"status": "waiting",
|
||||||
}
|
}
|
||||||
for index in range(1, 82)
|
for index in range(1, 151)
|
||||||
]
|
]
|
||||||
client = FakeQueueClient(records)
|
client = FakeQueueClient(records)
|
||||||
pool = ModelHubClientPool(
|
pool = ModelHubClientPool(
|
||||||
[client], # type: ignore[list-item]
|
[client], # type: ignore[list-item]
|
||||||
active_task_cap=100,
|
active_task_cap=100,
|
||||||
old_model_queue_threshold=80,
|
recent_model_reserve_slots=10,
|
||||||
|
capacity_state_path=None,
|
||||||
|
)
|
||||||
|
summary = cleanup_certain_oom_tasks(
|
||||||
|
pool,
|
||||||
|
FakeDiscovery(
|
||||||
|
{"owner/old": 1 * GIB},
|
||||||
|
{"owner/old": datetime(2026, 7, 1, tzinfo=timezone.utc)},
|
||||||
|
), # type: ignore[arg-type]
|
||||||
|
reference_time=datetime(2026, 8, 11, tzinfo=timezone.utc),
|
||||||
|
log=lambda _message: None,
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual([150], summary["accountCapacityLimits"])
|
||||||
|
self.assertEqual([140], summary["oldModelQueueThresholds"])
|
||||||
|
self.assertEqual(list(range(141, 151)), [item["taskId"] for item in summary["oldOverflowTasks"]])
|
||||||
|
|
||||||
|
def test_initial_cleanup_stops_old_task_beyond_capacity_minus_ten(self) -> None:
|
||||||
|
records = [
|
||||||
|
{
|
||||||
|
"taskId": index,
|
||||||
|
"modelId": "owner/old",
|
||||||
|
"gpuType": "Iluvatar_bi-100",
|
||||||
|
"status": "waiting",
|
||||||
|
}
|
||||||
|
for index in range(1, 92)
|
||||||
|
]
|
||||||
|
client = FakeQueueClient(records)
|
||||||
|
pool = ModelHubClientPool(
|
||||||
|
[client], # type: ignore[list-item]
|
||||||
|
active_task_cap=100,
|
||||||
|
recent_model_reserve_slots=10,
|
||||||
)
|
)
|
||||||
summary = cleanup_certain_oom_tasks(
|
summary = cleanup_certain_oom_tasks(
|
||||||
pool,
|
pool,
|
||||||
@@ -179,9 +215,10 @@ class QueueCleanupTests(unittest.TestCase):
|
|||||||
|
|
||||||
self.assertEqual(1, summary["oldOverflowCount"])
|
self.assertEqual(1, summary["oldOverflowCount"])
|
||||||
self.assertEqual(1, summary["cancelledCount"])
|
self.assertEqual(1, summary["cancelledCount"])
|
||||||
self.assertEqual([[81]], client.stopped)
|
self.assertEqual([90], summary["oldModelQueueThresholds"])
|
||||||
|
self.assertEqual([[91]], client.stopped)
|
||||||
|
|
||||||
def test_dynamic_cleanup_keeps_positions_through_ninety_five(self) -> None:
|
def test_scheduled_cleanup_uses_capacity_minus_five(self) -> None:
|
||||||
records = [
|
records = [
|
||||||
{
|
{
|
||||||
"taskId": index,
|
"taskId": index,
|
||||||
@@ -195,7 +232,7 @@ class QueueCleanupTests(unittest.TestCase):
|
|||||||
pool = ModelHubClientPool(
|
pool = ModelHubClientPool(
|
||||||
[client], # type: ignore[list-item]
|
[client], # type: ignore[list-item]
|
||||||
active_task_cap=100,
|
active_task_cap=100,
|
||||||
old_model_queue_threshold=80,
|
recent_model_reserve_slots=10,
|
||||||
)
|
)
|
||||||
summary = cleanup_certain_oom_tasks(
|
summary = cleanup_certain_oom_tasks(
|
||||||
pool,
|
pool,
|
||||||
@@ -203,12 +240,12 @@ class QueueCleanupTests(unittest.TestCase):
|
|||||||
{"owner/old": 1 * GIB},
|
{"owner/old": 1 * GIB},
|
||||||
{"owner/old": datetime(2026, 7, 1, tzinfo=timezone.utc)},
|
{"owner/old": datetime(2026, 7, 1, tzinfo=timezone.utc)},
|
||||||
), # type: ignore[arg-type]
|
), # type: ignore[arg-type]
|
||||||
age_queue_threshold=95,
|
age_reserved_slots=5,
|
||||||
reference_time=datetime(2026, 8, 11, tzinfo=timezone.utc),
|
reference_time=datetime(2026, 8, 11, tzinfo=timezone.utc),
|
||||||
log=lambda _message: None,
|
log=lambda _message: None,
|
||||||
)
|
)
|
||||||
|
|
||||||
self.assertEqual(95, summary["oldModelQueueThreshold"])
|
self.assertEqual([95], summary["oldModelQueueThresholds"])
|
||||||
self.assertEqual([96], [item["taskId"] for item in summary["oldOverflowTasks"]])
|
self.assertEqual([96], [item["taskId"] for item in summary["oldOverflowTasks"]])
|
||||||
self.assertEqual([[96]], client.stopped)
|
self.assertEqual([[96]], client.stopped)
|
||||||
|
|
||||||
@@ -220,13 +257,13 @@ class QueueCleanupTests(unittest.TestCase):
|
|||||||
"gpuType": "Iluvatar_bi-100",
|
"gpuType": "Iluvatar_bi-100",
|
||||||
"status": "waiting",
|
"status": "waiting",
|
||||||
}
|
}
|
||||||
for index in range(1, 83)
|
for index in range(1, 93)
|
||||||
]
|
]
|
||||||
client = FakeQueueClient(records)
|
client = FakeQueueClient(records)
|
||||||
pool = ModelHubClientPool(
|
pool = ModelHubClientPool(
|
||||||
[client], # type: ignore[list-item]
|
[client], # type: ignore[list-item]
|
||||||
active_task_cap=100,
|
active_task_cap=100,
|
||||||
old_model_queue_threshold=80,
|
recent_model_reserve_slots=10,
|
||||||
)
|
)
|
||||||
summary = cleanup_certain_oom_tasks(
|
summary = cleanup_certain_oom_tasks(
|
||||||
pool,
|
pool,
|
||||||
@@ -239,8 +276,8 @@ class QueueCleanupTests(unittest.TestCase):
|
|||||||
)
|
)
|
||||||
|
|
||||||
self.assertEqual(1, summary["certainOomCount"])
|
self.assertEqual(1, summary["certainOomCount"])
|
||||||
self.assertEqual([82], [item["taskId"] for item in summary["oldOverflowTasks"]])
|
self.assertEqual([92], [item["taskId"] for item in summary["oldOverflowTasks"]])
|
||||||
self.assertEqual([[1], [82]], client.stopped)
|
self.assertEqual([[1], [92]], client.stopped)
|
||||||
|
|
||||||
def test_stop_tasks_uses_documented_put_endpoint_and_integer_ids(self) -> None:
|
def test_stop_tasks_uses_documented_put_endpoint_and_integer_ids(self) -> None:
|
||||||
http = RecordingHttpClient()
|
http = RecordingHttpClient()
|
||||||
|
|||||||
Reference in New Issue
Block a user