perf: complete history bootstrap in one streaming phase
This commit is contained in:
39
README.md
39
README.md
@@ -59,15 +59,11 @@ Optional tuning:
|
|||||||
- `MODELHUB_QUEUE_CLEANUP_REPORT_PATH` default `.modelhub_state/queue_cleanup_latest.json`
|
- `MODELHUB_QUEUE_CLEANUP_REPORT_PATH` default `.modelhub_state/queue_cleanup_latest.json`
|
||||||
- `MODELHUB_ARCHITECTURE_BLACKLIST_PATH` default `.modelhub_state/architecture_compatibility_blacklist.json`
|
- `MODELHUB_ARCHITECTURE_BLACKLIST_PATH` default `.modelhub_state/architecture_compatibility_blacklist.json`
|
||||||
- `MODELHUB_ARCHITECTURE_BLOCK_TTL_DAYS` default `30`
|
- `MODELHUB_ARCHITECTURE_BLOCK_TTL_DAYS` default `30`
|
||||||
- `MODELHUB_ARCHITECTURE_COMMUNITY_PROBE_SIZE` default `50`
|
|
||||||
- `MODELHUB_ARCHITECTURE_COMMUNITY_LOOKBACK_DAYS` default `30`
|
|
||||||
- `MODELHUB_ARCHITECTURE_COMMUNITY_LATEST_LIMIT` default `100` (synchronous seed only)
|
|
||||||
- `MODELHUB_ARCHITECTURE_BOOTSTRAP_WORKERS` default `8`
|
- `MODELHUB_ARCHITECTURE_BOOTSTRAP_WORKERS` default `8`
|
||||||
- `MODELHUB_ARCHITECTURE_BOOTSTRAP_TASKS_PER_ACCOUNT` default `10` (synchronous seed only)
|
- `MODELHUB_ARCHITECTURE_BACKFILL_PAGE_SIZE` default `100`
|
||||||
- `MODELHUB_ARCHITECTURE_BOOTSTRAP_MAX_LOGS` default `120` (synchronous seed only)
|
- `MODELHUB_ARCHITECTURE_BACKFILL_PAGES_PER_BATCH` default `10` (about 1,000 task rows)
|
||||||
- `MODELHUB_ARCHITECTURE_BACKFILL_PAGE_SIZE` default `50`
|
- `MODELHUB_ARCHITECTURE_BACKFILL_LOG_BATCH_SIZE` default `200`
|
||||||
- `MODELHUB_ARCHITECTURE_BACKFILL_PAGES_PER_CYCLE` default `2`
|
- `MODELHUB_ARCHITECTURE_BACKFILL_CHECKPOINT_RECORDS` default `2000`
|
||||||
- `MODELHUB_ARCHITECTURE_BACKFILL_LOGS_PER_CYCLE` default `100`
|
|
||||||
- `MODELHUB_RECENT_MODEL_RESERVE_SLOTS` default `5` per account
|
- `MODELHUB_RECENT_MODEL_RESERVE_SLOTS` default `5` per account
|
||||||
- `MODELHUB_RECENT_MODEL_DAYS` default `7`
|
- `MODELHUB_RECENT_MODEL_DAYS` default `7`
|
||||||
- `MODELHUB_STATE_SYNC_REMOTE` default `https://dev.modelhub.org.cn/CoolBoy/submmit.git`
|
- `MODELHUB_STATE_SYNC_REMOTE` default `https://dev.modelhub.org.cn/CoolBoy/submmit.git`
|
||||||
@@ -173,17 +169,18 @@ change that window. The stats report exposes `architectureCompatibilityBlocks`
|
|||||||
and per-GPU/framework block counts. The live snapshot is written to
|
and per-GPU/framework block counts. The live snapshot is written to
|
||||||
`.modelhub_state/architecture_compatibility_blacklist.json`.
|
`.modelhub_state/architecture_compatibility_blacklist.json`.
|
||||||
|
|
||||||
Before the first cleanup/submission cycle, the poller imports only a small recent
|
Cold start does not scan historical tasks before submission. The first normal
|
||||||
seed so cold start cannot be dominated by historical API and log downloads. The
|
submission pass runs with official/public live routing evidence, then a one-time
|
||||||
submission pass then runs first on every cycle. Afterwards a resumable backfill
|
resumable initialization walks every configured account's complete history in
|
||||||
walks every configured account's complete history in bounded pages, classifies
|
bounded batches of about 1,000 rows. It classifies every available failure log
|
||||||
each available failure log, and immediately reduces the page to cumulative
|
in internal batches of 200 and immediately reduces each metadata batch to cumulative
|
||||||
success/failure statistics, memory observations, safe configurations, and exact
|
success/failure statistics, memory observations, safe configurations, and exact
|
||||||
compatibility blocks. Only the cursor, temporary deduplication task IDs, and
|
compatibility blocks. Only the cursor, temporary deduplication task IDs, and
|
||||||
derived decision checkpoint are synchronized; raw task rows and logs are not.
|
derived decision checkpoint are synchronized; raw task rows and logs are not.
|
||||||
The task-ID set is deleted when backfill completes. A restart resumes the cursor
|
The task-ID set is deleted when backfill completes. A restart resumes the cursor
|
||||||
instead of rescanning prior pages. Empty public evidence is a fallback signal,
|
instead of rescanning prior pages. Checkpoints are pushed about every 2,000 rows,
|
||||||
never proof that no incompatibility exists.
|
not after each API page. Once complete, full-history work is permanently disabled
|
||||||
|
and normal outcome synchronization processes only newer task changes.
|
||||||
|
|
||||||
Before a candidate reaches the submit queue, failure-informed preflight checks
|
Before a candidate reaches the submit queue, failure-informed preflight checks
|
||||||
the actual ModelScope repository structure and file sizes. Non-GGUF text
|
the actual ModelScope repository structure and file sizes. Non-GGUF text
|
||||||
@@ -439,11 +436,19 @@ processed, and each page is compacted immediately into decision state. It also
|
|||||||
fixes large-batch classification so checkpoint compaction cannot discard an
|
fixes large-batch classification so checkpoint compaction cannot discard an
|
||||||
uninspected failure row.
|
uninspected failure row.
|
||||||
|
|
||||||
|
Version `2026.09.04.4` removes the long-lived per-cycle backfill. With no durable
|
||||||
|
checkpoint, the worker first fills the queue and then completes the entire
|
||||||
|
historical initialization as one streaming phase: 1,000 task rows per bounded
|
||||||
|
metadata batch, 200 failure logs per internal classification batch, and one
|
||||||
|
durable Git checkpoint per roughly 2,000 scanned rows. Raw rows are discarded
|
||||||
|
after every aggregate merge. Transient API failure pauses at the current cursor
|
||||||
|
for the next poll cycle; successful completion disables the scan permanently.
|
||||||
|
|
||||||
## Deploy
|
## Deploy
|
||||||
|
|
||||||
Create a tag and submit the repository URL plus tag in "我的适配智能体".
|
Create a tag and submit the repository URL plus tag in "我的适配智能体".
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
git tag -a agent-v32 -m "ModelHub agent 2026.09.04.3"
|
git tag -a agent-v33 -m "ModelHub agent 2026.09.04.4"
|
||||||
git push origin main agent-v32
|
git push origin main agent-v33
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -103,12 +103,13 @@ bash run_poll.sh --dry-run
|
|||||||
rule immediately launches a lightweight architecture-only queue scan. Exact
|
rule immediately launches a lightweight architecture-only queue scan. Exact
|
||||||
matching waiting tasks are stopped after two state checks; running tasks and
|
matching waiting tasks are stopped after two state checks; running tasks and
|
||||||
tasks without local framework/task metadata are protected.
|
tasks without local framework/task metadata are protected.
|
||||||
- Startup imports a small community/owned seed for immediate decisions. After
|
- Cold start performs no blocking history seed. After the first submission pass,
|
||||||
each submission pass, a durable cursor incrementally scans complete terminal
|
a durable cursor completes the full history scan as one streaming phase using
|
||||||
history across all configured accounts. Each bounded page is classified and
|
about 1,000 task rows per metadata batch and 200 logs per classification batch.
|
||||||
folded into aggregate decision state immediately; raw task rows and failure
|
Aggregate checkpoints are synchronized about every 2,000 rows; raw task rows
|
||||||
logs are never synchronized. Restarts resume the cursor, and the temporary
|
and failure logs are never synchronized. Restarts resume the cursor, temporary
|
||||||
deduplication IDs are removed when the exhaustive backfill completes.
|
deduplication IDs disappear at completion, and later cycles process only new
|
||||||
|
outcome changes.
|
||||||
- A strategy generation lasts exactly 200 platform-accepted submissions. Rejected API calls and
|
- A strategy generation lasts exactly 200 platform-accepted submissions. Rejected API calls and
|
||||||
duplicates do not advance it. The next cycle refreshes platform history before submitting again.
|
duplicates do not advance it. The next cycle refreshes platform history before submitting again.
|
||||||
- Strategy state is stored in `.modelhub_state/gpu_strategy.json`; a generation never recalculates
|
- Strategy state is stored in `.modelhub_state/gpu_strategy.json`; a generation never recalculates
|
||||||
|
|||||||
@@ -485,6 +485,7 @@ class OutcomeTracker:
|
|||||||
task_contexts: dict[str, dict[str, Any]] | None = None,
|
task_contexts: dict[str, dict[str, Any]] | None = None,
|
||||||
enrichment_limit: int = 0,
|
enrichment_limit: int = 0,
|
||||||
enrichment_workers: int = 8,
|
enrichment_workers: int = 8,
|
||||||
|
enrichment_batch_size: int = 200,
|
||||||
log: Any = None,
|
log: Any = None,
|
||||||
) -> dict[str, int]:
|
) -> dict[str, int]:
|
||||||
"""Import terminal history and classify every usable historical failure."""
|
"""Import terminal history and classify every usable historical failure."""
|
||||||
@@ -556,7 +557,7 @@ class OutcomeTracker:
|
|||||||
enrichment_attempts = 0
|
enrichment_attempts = 0
|
||||||
classified_total = 0
|
classified_total = 0
|
||||||
errors_total = 0
|
errors_total = 0
|
||||||
batch_size = 200
|
batch_size = max(1, int(enrichment_batch_size))
|
||||||
total_candidates = len(candidates)
|
total_candidates = len(candidates)
|
||||||
for offset in range(0, total_candidates, batch_size):
|
for offset in range(0, total_candidates, batch_size):
|
||||||
# Keep direct references until classification completes. Calling
|
# Keep direct references until classification completes. Calling
|
||||||
|
|||||||
@@ -6,8 +6,6 @@ import os
|
|||||||
import sys
|
import sys
|
||||||
import time
|
import time
|
||||||
from collections import deque
|
from collections import deque
|
||||||
from concurrent.futures import ThreadPoolExecutor, as_completed
|
|
||||||
from datetime import timedelta
|
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Any, Callable
|
from typing import Any, Callable
|
||||||
|
|
||||||
@@ -334,200 +332,6 @@ def _load_task_compatibility_contexts(
|
|||||||
return contexts
|
return contexts
|
||||||
|
|
||||||
|
|
||||||
def _load_bounded_owned_history(
|
|
||||||
modelhub_client: ModelHubClient | ModelHubClientPool,
|
|
||||||
*,
|
|
||||||
max_records_per_account: int,
|
|
||||||
) -> tuple[list[dict[str, Any]], list[int]]:
|
|
||||||
clients = (
|
|
||||||
list(modelhub_client.clients)
|
|
||||||
if isinstance(modelhub_client, ModelHubClientPool)
|
|
||||||
else [modelhub_client]
|
|
||||||
)
|
|
||||||
by_account: dict[int, list[dict[str, Any]]] = {}
|
|
||||||
errors: list[int] = []
|
|
||||||
with ThreadPoolExecutor(max_workers=min(12, max(1, len(clients)))) as executor:
|
|
||||||
futures = {
|
|
||||||
executor.submit(
|
|
||||||
client.list_tasks,
|
|
||||||
page_size=100,
|
|
||||||
only_mine=True,
|
|
||||||
max_records=max_records_per_account,
|
|
||||||
): index
|
|
||||||
for index, client in enumerate(clients, start=1)
|
|
||||||
}
|
|
||||||
for future in as_completed(futures):
|
|
||||||
account_index = futures[future]
|
|
||||||
try:
|
|
||||||
by_account[account_index] = future.result()
|
|
||||||
except Exception:
|
|
||||||
errors.append(account_index)
|
|
||||||
|
|
||||||
deduped: dict[str, dict[str, Any]] = {}
|
|
||||||
anonymous: list[dict[str, Any]] = []
|
|
||||||
for account_index in sorted(by_account):
|
|
||||||
for task in by_account[account_index]:
|
|
||||||
if not isinstance(task, dict):
|
|
||||||
continue
|
|
||||||
task_id = task.get("taskId")
|
|
||||||
if task_id is None:
|
|
||||||
anonymous.append(task)
|
|
||||||
continue
|
|
||||||
deduped[str(task_id)] = task
|
|
||||||
return [*deduped.values(), *anonymous], sorted(errors)
|
|
||||||
|
|
||||||
|
|
||||||
def _bootstrap_architecture_history(
|
|
||||||
*,
|
|
||||||
modelhub_client: ModelHubClient | ModelHubClientPool,
|
|
||||||
outcome_tracker: OutcomeTracker,
|
|
||||||
ledger_path: Path,
|
|
||||||
now,
|
|
||||||
) -> dict[str, Any]:
|
|
||||||
"""Prefer bounded recent public evidence, then bounded owned history."""
|
|
||||||
probe_size = _env_int(
|
|
||||||
"MODELHUB_ARCHITECTURE_COMMUNITY_PROBE_SIZE",
|
|
||||||
50,
|
|
||||||
minimum=1,
|
|
||||||
maximum=100,
|
|
||||||
)
|
|
||||||
community_records: list[dict[str, Any]] = []
|
|
||||||
community_probe_error: str | None = None
|
|
||||||
try:
|
|
||||||
probe_payload = modelhub_client.list_tasks_page(
|
|
||||||
current=1,
|
|
||||||
page_size=probe_size,
|
|
||||||
only_mine=False,
|
|
||||||
status="success",
|
|
||||||
verify_result=-1,
|
|
||||||
)
|
|
||||||
probe_records = (probe_payload.get("data") or {}).get("records") or []
|
|
||||||
community_records = [record for record in probe_records if isinstance(record, dict)]
|
|
||||||
except Exception as exc:
|
|
||||||
community_probe_error = f"{type(exc).__name__}: {exc}"
|
|
||||||
|
|
||||||
community_usable = [
|
|
||||||
record
|
|
||||||
for record in community_records
|
|
||||||
if record.get("logCosUrl")
|
|
||||||
and record.get("modelId")
|
|
||||||
and record.get("gpuType")
|
|
||||||
]
|
|
||||||
log(
|
|
||||||
f"[architecture-bootstrap] community_probe={len(community_records)} "
|
|
||||||
f"usable_failure_details={len(community_usable)} "
|
|
||||||
f"error={'none' if community_probe_error is None else community_probe_error}"
|
|
||||||
)
|
|
||||||
|
|
||||||
source = "community_latest"
|
|
||||||
listing_errors: list[int] = []
|
|
||||||
if community_usable:
|
|
||||||
lookback_days = _env_int(
|
|
||||||
"MODELHUB_ARCHITECTURE_COMMUNITY_LOOKBACK_DAYS",
|
|
||||||
30,
|
|
||||||
minimum=1,
|
|
||||||
maximum=365,
|
|
||||||
)
|
|
||||||
latest_limit = _env_int(
|
|
||||||
"MODELHUB_ARCHITECTURE_COMMUNITY_LATEST_LIMIT",
|
|
||||||
100,
|
|
||||||
minimum=1,
|
|
||||||
maximum=50_000,
|
|
||||||
)
|
|
||||||
try:
|
|
||||||
history_tasks = modelhub_client.list_tasks(
|
|
||||||
page_size=100,
|
|
||||||
only_mine=False,
|
|
||||||
begin_time=now - timedelta(days=lookback_days),
|
|
||||||
end_time=now,
|
|
||||||
status="success",
|
|
||||||
verify_result=-1,
|
|
||||||
max_records=latest_limit,
|
|
||||||
)
|
|
||||||
history_tasks = [task for task in history_tasks if task.get("logCosUrl")]
|
|
||||||
log(
|
|
||||||
f"[architecture-bootstrap] source=community_latest "
|
|
||||||
f"lookback_days={lookback_days} records={len(history_tasks)} limit={latest_limit}"
|
|
||||||
)
|
|
||||||
except Exception as exc:
|
|
||||||
source = "owned_recent_bounded"
|
|
||||||
log(
|
|
||||||
f"[architecture-bootstrap] community_history_error={type(exc).__name__}: {exc} "
|
|
||||||
"fallback=owned_recent_bounded"
|
|
||||||
)
|
|
||||||
history_tasks, listing_errors = _load_bounded_owned_history(
|
|
||||||
modelhub_client,
|
|
||||||
max_records_per_account=_env_int(
|
|
||||||
"MODELHUB_ARCHITECTURE_BOOTSTRAP_TASKS_PER_ACCOUNT",
|
|
||||||
10,
|
|
||||||
minimum=10,
|
|
||||||
maximum=500,
|
|
||||||
),
|
|
||||||
)
|
|
||||||
else:
|
|
||||||
source = "owned_recent_bounded"
|
|
||||||
per_account_limit = _env_int(
|
|
||||||
"MODELHUB_ARCHITECTURE_BOOTSTRAP_TASKS_PER_ACCOUNT",
|
|
||||||
10,
|
|
||||||
minimum=10,
|
|
||||||
maximum=500,
|
|
||||||
)
|
|
||||||
history_tasks, listing_errors = _load_bounded_owned_history(
|
|
||||||
modelhub_client,
|
|
||||||
max_records_per_account=per_account_limit,
|
|
||||||
)
|
|
||||||
account_count = (
|
|
||||||
len(modelhub_client.clients)
|
|
||||||
if isinstance(modelhub_client, ModelHubClientPool)
|
|
||||||
else 1
|
|
||||||
)
|
|
||||||
log(
|
|
||||||
f"[architecture-bootstrap] source=owned_recent_bounded records={len(history_tasks)} "
|
|
||||||
f"accounts={account_count} per_account_limit={per_account_limit} "
|
|
||||||
f"listing_errors={','.join(map(str, listing_errors)) or 'none'}"
|
|
||||||
)
|
|
||||||
|
|
||||||
contexts = _load_task_compatibility_contexts(
|
|
||||||
outcome_tracker,
|
|
||||||
ledger_path=ledger_path,
|
|
||||||
)
|
|
||||||
summary = outcome_tracker.bootstrap_from_history_tasks(
|
|
||||||
history_tasks,
|
|
||||||
task_contexts=contexts,
|
|
||||||
enrichment_limit=_env_int(
|
|
||||||
"MODELHUB_ARCHITECTURE_BOOTSTRAP_MAX_LOGS",
|
|
||||||
120,
|
|
||||||
minimum=0,
|
|
||||||
maximum=100_000,
|
|
||||||
),
|
|
||||||
enrichment_workers=_env_int(
|
|
||||||
"MODELHUB_ARCHITECTURE_BOOTSTRAP_WORKERS",
|
|
||||||
8,
|
|
||||||
minimum=1,
|
|
||||||
maximum=16,
|
|
||||||
),
|
|
||||||
log=log,
|
|
||||||
)
|
|
||||||
feedback = outcome_tracker.get_stats_report()
|
|
||||||
block_count = len(feedback.get("architectureCompatibilityBlocks") or {})
|
|
||||||
summary.update(
|
|
||||||
{
|
|
||||||
"source": source,
|
|
||||||
"communityProbeRecords": len(community_records),
|
|
||||||
"communityUsableFailureDetails": len(community_usable),
|
|
||||||
"listingErrorAccounts": listing_errors,
|
|
||||||
"architectureBlocks": block_count,
|
|
||||||
}
|
|
||||||
)
|
|
||||||
log(
|
|
||||||
f"[architecture-bootstrap] finished source={source} "
|
|
||||||
f"terminal={summary['terminalRecords']} failure_logs={summary['eligibleFailureLogs']} "
|
|
||||||
f"explicit_architecture_failures={summary['explicitArchitectureFailures']} "
|
|
||||||
f"recovered_frameworks={summary['recoveredFrameworks']} blocks={block_count}"
|
|
||||||
)
|
|
||||||
return summary
|
|
||||||
|
|
||||||
|
|
||||||
def _architecture_backfill_clients(
|
def _architecture_backfill_clients(
|
||||||
modelhub_client: ModelHubClient | ModelHubClientPool,
|
modelhub_client: ModelHubClient | ModelHubClientPool,
|
||||||
) -> list[ModelHubClient]:
|
) -> list[ModelHubClient]:
|
||||||
@@ -620,20 +424,20 @@ def _advance_architecture_history_backfill(
|
|||||||
|
|
||||||
page_size = _env_int(
|
page_size = _env_int(
|
||||||
"MODELHUB_ARCHITECTURE_BACKFILL_PAGE_SIZE",
|
"MODELHUB_ARCHITECTURE_BACKFILL_PAGE_SIZE",
|
||||||
50,
|
100,
|
||||||
minimum=10,
|
minimum=10,
|
||||||
maximum=100,
|
maximum=100,
|
||||||
)
|
)
|
||||||
pages_per_cycle = _env_int(
|
pages_per_batch = _env_int(
|
||||||
"MODELHUB_ARCHITECTURE_BACKFILL_PAGES_PER_CYCLE",
|
"MODELHUB_ARCHITECTURE_BACKFILL_PAGES_PER_BATCH",
|
||||||
2,
|
10,
|
||||||
minimum=1,
|
minimum=1,
|
||||||
maximum=12,
|
maximum=20,
|
||||||
)
|
)
|
||||||
max_logs = _env_int(
|
log_batch_size = _env_int(
|
||||||
"MODELHUB_ARCHITECTURE_BACKFILL_LOGS_PER_CYCLE",
|
"MODELHUB_ARCHITECTURE_BACKFILL_LOG_BATCH_SIZE",
|
||||||
page_size * pages_per_cycle,
|
200,
|
||||||
minimum=1,
|
minimum=10,
|
||||||
maximum=500,
|
maximum=500,
|
||||||
)
|
)
|
||||||
accounts = progress.get("accounts") or {}
|
accounts = progress.get("accounts") or {}
|
||||||
@@ -642,7 +446,7 @@ def _advance_architecture_history_backfill(
|
|||||||
page_calls = 0
|
page_calls = 0
|
||||||
page_errors = 0
|
page_errors = 0
|
||||||
|
|
||||||
for _ in range(pages_per_cycle):
|
for _ in range(pages_per_batch):
|
||||||
incomplete = [
|
incomplete = [
|
||||||
index
|
index
|
||||||
for index in range(len(clients))
|
for index in range(len(clients))
|
||||||
@@ -669,6 +473,7 @@ def _advance_architecture_history_backfill(
|
|||||||
data = payload.get("data") or {}
|
data = payload.get("data") or {}
|
||||||
records = [item for item in (data.get("records") or []) if isinstance(item, dict)]
|
records = [item for item in (data.get("records") or []) if isinstance(item, dict)]
|
||||||
pages = max(0, int(data.get("pages") or 0))
|
pages = max(0, int(data.get("pages") or 0))
|
||||||
|
state.pop("lastError", None)
|
||||||
state["recordsScanned"] = int(state.get("recordsScanned") or 0) + len(records)
|
state["recordsScanned"] = int(state.get("recordsScanned") or 0) + len(records)
|
||||||
progress["recordsScanned"] = int(progress.get("recordsScanned") or 0) + len(records)
|
progress["recordsScanned"] = int(progress.get("recordsScanned") or 0) + len(records)
|
||||||
new_count = 0
|
new_count = 0
|
||||||
@@ -695,6 +500,7 @@ def _advance_architecture_history_backfill(
|
|||||||
page_errors += 1
|
page_errors += 1
|
||||||
state["listingErrors"] = int(state.get("listingErrors") or 0) + 1
|
state["listingErrors"] = int(state.get("listingErrors") or 0) + 1
|
||||||
state["lastError"] = f"{type(exc).__name__}: {exc}"[:500]
|
state["lastError"] = f"{type(exc).__name__}: {exc}"[:500]
|
||||||
|
break
|
||||||
|
|
||||||
batch_summary: dict[str, Any] = {
|
batch_summary: dict[str, Any] = {
|
||||||
"terminalRecords": 0,
|
"terminalRecords": 0,
|
||||||
@@ -708,13 +514,17 @@ def _advance_architecture_history_backfill(
|
|||||||
outcome_tracker,
|
outcome_tracker,
|
||||||
ledger_path=ledger_path,
|
ledger_path=ledger_path,
|
||||||
),
|
),
|
||||||
enrichment_limit=max_logs,
|
# Every usable failure in this metadata batch is inspected. The
|
||||||
|
# batch boundary bounds memory; enrichment_batch_size bounds the
|
||||||
|
# number of retained log jobs at one time.
|
||||||
|
enrichment_limit=0,
|
||||||
enrichment_workers=_env_int(
|
enrichment_workers=_env_int(
|
||||||
"MODELHUB_ARCHITECTURE_BOOTSTRAP_WORKERS",
|
"MODELHUB_ARCHITECTURE_BOOTSTRAP_WORKERS",
|
||||||
8,
|
8,
|
||||||
minimum=1,
|
minimum=1,
|
||||||
maximum=16,
|
maximum=16,
|
||||||
),
|
),
|
||||||
|
enrichment_batch_size=log_batch_size,
|
||||||
log=log,
|
log=log,
|
||||||
)
|
)
|
||||||
# Each slice is immediately reduced to cumulative counters, routing
|
# Each slice is immediately reduced to cumulative counters, routing
|
||||||
@@ -753,6 +563,66 @@ def _advance_architecture_history_backfill(
|
|||||||
return progress
|
return progress
|
||||||
|
|
||||||
|
|
||||||
|
def _complete_architecture_history_backfill(
|
||||||
|
*,
|
||||||
|
modelhub_client: ModelHubClient | ModelHubClientPool,
|
||||||
|
outcome_tracker: OutcomeTracker,
|
||||||
|
ledger_path: Path,
|
||||||
|
progress_path: Path = DEFAULT_ARCHITECTURE_BACKFILL_PATH,
|
||||||
|
sync_callback: Callable[[str], bool] | None = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
"""Finish the one-time scan in bounded batches after the first submit pass."""
|
||||||
|
checkpoint_interval = _env_int(
|
||||||
|
"MODELHUB_ARCHITECTURE_BACKFILL_CHECKPOINT_RECORDS",
|
||||||
|
2000,
|
||||||
|
minimum=500,
|
||||||
|
maximum=10_000,
|
||||||
|
)
|
||||||
|
scanned_since_sync = 0
|
||||||
|
try:
|
||||||
|
progress = read_json(progress_path)
|
||||||
|
except (FileNotFoundError, ValueError, TypeError):
|
||||||
|
progress = _new_architecture_backfill_progress(
|
||||||
|
modelhub_client,
|
||||||
|
now=utc_now(),
|
||||||
|
)
|
||||||
|
write_json(progress_path, progress)
|
||||||
|
|
||||||
|
while not bool(progress.get("complete")):
|
||||||
|
before_scanned = int(progress.get("recordsScanned") or 0)
|
||||||
|
progress = _advance_architecture_history_backfill(
|
||||||
|
modelhub_client=modelhub_client,
|
||||||
|
outcome_tracker=outcome_tracker,
|
||||||
|
ledger_path=ledger_path,
|
||||||
|
progress_path=progress_path,
|
||||||
|
)
|
||||||
|
scanned_now = int(progress.get("recordsScanned") or 0)
|
||||||
|
scanned_delta = max(0, scanned_now - before_scanned)
|
||||||
|
scanned_since_sync += scanned_delta
|
||||||
|
|
||||||
|
should_sync = bool(progress.get("complete")) or scanned_since_sync >= checkpoint_interval
|
||||||
|
if should_sync and sync_callback is not None:
|
||||||
|
if not sync_callback("history_backfill"):
|
||||||
|
log(
|
||||||
|
f"[architecture-backfill] status=paused reason=state_sync_failed "
|
||||||
|
f"records_since_checkpoint={scanned_since_sync}"
|
||||||
|
)
|
||||||
|
break
|
||||||
|
log(
|
||||||
|
f"[architecture-backfill] checkpoint=durable "
|
||||||
|
f"records_total={scanned_now} complete={str(bool(progress.get('complete'))).lower()}"
|
||||||
|
)
|
||||||
|
scanned_since_sync = 0
|
||||||
|
|
||||||
|
# A page/API error must not create a hot loop. The normal poll loop will
|
||||||
|
# retry the same durable cursor on its next cycle.
|
||||||
|
if scanned_delta <= 0 and not bool(progress.get("complete")):
|
||||||
|
log("[architecture-backfill] status=deferred reason=no_scan_progress retry_next_cycle=true")
|
||||||
|
break
|
||||||
|
|
||||||
|
return progress
|
||||||
|
|
||||||
|
|
||||||
def run_poll_loop(
|
def run_poll_loop(
|
||||||
*,
|
*,
|
||||||
base_args: argparse.Namespace,
|
base_args: argparse.Namespace,
|
||||||
@@ -868,17 +738,15 @@ def run_poll_loop(
|
|||||||
f"history_backfill={architecture_bootstrap_summary['historyBackfill']}"
|
f"history_backfill={architecture_bootstrap_summary['historyBackfill']}"
|
||||||
)
|
)
|
||||||
else:
|
else:
|
||||||
architecture_bootstrap_summary = _bootstrap_architecture_history(
|
architecture_bootstrap_summary = {
|
||||||
modelhub_client=modelhub_client,
|
"source": "cold_start_streaming_backfill",
|
||||||
outcome_tracker=outcome_tracker,
|
"terminalRecords": 0,
|
||||||
ledger_path=Path(base_args.ledger_path),
|
"architectureBlocks": 0,
|
||||||
now=now,
|
"historyBackfill": "pending_after_first_submission_pass",
|
||||||
)
|
}
|
||||||
outcome_tracker.compact_decision_state()
|
# Progress without its aggregate checkpoint is not usable:
|
||||||
# A missing decision checkpoint means the prior aggregate was
|
# skipping those pages would under-count history. Reset both
|
||||||
# unavailable. Start a resumable full owned-history backfill;
|
# sides of the resumable scan whenever the checkpoint is gone.
|
||||||
# the small synchronous seed above only makes the first routing
|
|
||||||
# decisions useful without delaying submissions.
|
|
||||||
write_json(
|
write_json(
|
||||||
architecture_backfill_path,
|
architecture_backfill_path,
|
||||||
_new_architecture_backfill_progress(
|
_new_architecture_backfill_progress(
|
||||||
@@ -886,6 +754,10 @@ def run_poll_loop(
|
|||||||
now=now,
|
now=now,
|
||||||
),
|
),
|
||||||
)
|
)
|
||||||
|
log(
|
||||||
|
"[architecture-bootstrap] source=cold_start_streaming_backfill "
|
||||||
|
"startup_history_scan=disabled backfill=after_first_submission_pass"
|
||||||
|
)
|
||||||
architecture_bootstrap_summary["enabled"] = True
|
architecture_bootstrap_summary["enabled"] = True
|
||||||
_persist_architecture_blacklist(
|
_persist_architecture_blacklist(
|
||||||
outcome_tracker.get_stats_report(),
|
outcome_tracker.get_stats_report(),
|
||||||
@@ -927,7 +799,7 @@ def run_poll_loop(
|
|||||||
pending_architecture_cleanup = False
|
pending_architecture_cleanup = False
|
||||||
last_cleaned_architecture_blocks: set[str] = set()
|
last_cleaned_architecture_blocks: set[str] = set()
|
||||||
|
|
||||||
def advance_architecture_backfill() -> None:
|
def complete_architecture_backfill() -> None:
|
||||||
nonlocal pending_architecture_cleanup
|
nonlocal pending_architecture_cleanup
|
||||||
if getattr(base_args, "skip_outcome_sync", False) or not architecture_backfill_path.is_file():
|
if getattr(base_args, "skip_outcome_sync", False) or not architecture_backfill_path.is_file():
|
||||||
return
|
return
|
||||||
@@ -938,11 +810,16 @@ def run_poll_loop(
|
|||||||
previous_blocks = set(
|
previous_blocks = set(
|
||||||
(outcome_tracker.get_stats_report().get("architectureCompatibilityBlocks") or {}).keys()
|
(outcome_tracker.get_stats_report().get("architectureCompatibilityBlocks") or {}).keys()
|
||||||
)
|
)
|
||||||
_advance_architecture_history_backfill(
|
_complete_architecture_history_backfill(
|
||||||
modelhub_client=modelhub_client,
|
modelhub_client=modelhub_client,
|
||||||
outcome_tracker=outcome_tracker,
|
outcome_tracker=outcome_tracker,
|
||||||
ledger_path=Path(base_args.ledger_path),
|
ledger_path=Path(base_args.ledger_path),
|
||||||
progress_path=architecture_backfill_path,
|
progress_path=architecture_backfill_path,
|
||||||
|
sync_callback=(
|
||||||
|
(lambda phase: state_sync.sync(phase))
|
||||||
|
if state_sync is not None
|
||||||
|
else None
|
||||||
|
),
|
||||||
)
|
)
|
||||||
current_feedback = outcome_tracker.get_stats_report()
|
current_feedback = outcome_tracker.get_stats_report()
|
||||||
current_blocks = _persist_architecture_blacklist(
|
current_blocks = _persist_architecture_blacklist(
|
||||||
@@ -1165,7 +1042,7 @@ def run_poll_loop(
|
|||||||
# Historical learning is deliberately behind queue maintenance and
|
# Historical learning is deliberately behind queue maintenance and
|
||||||
# capacity checks. It advances even while full, but never blocks
|
# capacity checks. It advances even while full, but never blocks
|
||||||
# the worker's initial recovery or first submission attempt.
|
# the worker's initial recovery or first submission attempt.
|
||||||
advance_architecture_backfill()
|
complete_architecture_backfill()
|
||||||
if state_sync is not None:
|
if state_sync is not None:
|
||||||
try:
|
try:
|
||||||
if cycles % 3 == 0 and hasattr(modelhub_client, "list_active_tasks_by_account"):
|
if cycles % 3 == 0 and hasattr(modelhub_client, "list_active_tasks_by_account"):
|
||||||
@@ -1205,9 +1082,10 @@ def run_poll_loop(
|
|||||||
f"stop={cycle_summary['stoppedReason']}"
|
f"stop={cycle_summary['stoppedReason']}"
|
||||||
)
|
)
|
||||||
|
|
||||||
# Do this only after the submission pass. Each cycle consumes at most a
|
# The first submission pass gets priority. Then the one-time cold-start
|
||||||
# small configured slice and immediately reduces it to decision state.
|
# scan runs continuously in bounded memory and disappears permanently
|
||||||
advance_architecture_backfill()
|
# once its compact decision checkpoint is complete.
|
||||||
|
complete_architecture_backfill()
|
||||||
|
|
||||||
if state_sync is not None:
|
if state_sync is not None:
|
||||||
try:
|
try:
|
||||||
|
|||||||
@@ -1 +1 @@
|
|||||||
AGENT_VERSION = "2026.09.04.3"
|
AGENT_VERSION = "2026.09.04.4"
|
||||||
|
|||||||
@@ -17,60 +17,13 @@ sys.path.insert(0, str(PACKAGE_DIR))
|
|||||||
from outcome_tracker import OutcomeTracker # noqa: E402
|
from outcome_tracker import OutcomeTracker # noqa: E402
|
||||||
from poll_runner import ( # noqa: E402
|
from poll_runner import ( # noqa: E402
|
||||||
_advance_architecture_history_backfill,
|
_advance_architecture_history_backfill,
|
||||||
_bootstrap_architecture_history,
|
_complete_architecture_history_backfill,
|
||||||
_load_task_compatibility_contexts,
|
_load_task_compatibility_contexts,
|
||||||
build_parser,
|
build_parser,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
class PollPolicyTests(unittest.TestCase):
|
class PollPolicyTests(unittest.TestCase):
|
||||||
def test_architecture_bootstrap_falls_back_when_public_logs_are_hidden(self) -> None:
|
|
||||||
class HistoryClient:
|
|
||||||
@staticmethod
|
|
||||||
def list_tasks_page(**_kwargs): # noqa: ANN003
|
|
||||||
return {
|
|
||||||
"data": {
|
|
||||||
"records": [
|
|
||||||
{
|
|
||||||
"taskId": "public-failure",
|
|
||||||
"modelId": "public/model",
|
|
||||||
"gpuType": "gpu",
|
|
||||||
"status": "success",
|
|
||||||
"verifyResult": -1,
|
|
||||||
"logCosUrl": None,
|
|
||||||
}
|
|
||||||
]
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
@staticmethod
|
|
||||||
def list_tasks(**_kwargs): # noqa: ANN003
|
|
||||||
return [
|
|
||||||
{
|
|
||||||
"taskId": "owned-success",
|
|
||||||
"modelId": "owner/model",
|
|
||||||
"gpuType": "gpu",
|
|
||||||
"modelTaskLevelId": 23,
|
|
||||||
"status": "success",
|
|
||||||
"verifyResult": 1,
|
|
||||||
"updateTime": datetime.now(timezone.utc).isoformat(),
|
|
||||||
}
|
|
||||||
]
|
|
||||||
|
|
||||||
with tempfile.TemporaryDirectory() as temporary_dir:
|
|
||||||
root = Path(temporary_dir)
|
|
||||||
tracker = OutcomeTracker(root / "outcomes.jsonl")
|
|
||||||
summary = _bootstrap_architecture_history(
|
|
||||||
modelhub_client=HistoryClient(), # type: ignore[arg-type]
|
|
||||||
outcome_tracker=tracker,
|
|
||||||
ledger_path=root / "ledger.jsonl",
|
|
||||||
now=datetime.now(timezone.utc),
|
|
||||||
)
|
|
||||||
|
|
||||||
self.assertEqual("owned_recent_bounded", summary["source"])
|
|
||||||
self.assertEqual(1, summary["terminalRecords"])
|
|
||||||
self.assertEqual(0, summary["communityUsableFailureDetails"])
|
|
||||||
|
|
||||||
def test_architecture_history_backfill_resumes_and_keeps_only_decision_state(self) -> None:
|
def test_architecture_history_backfill_resumes_and_keeps_only_decision_state(self) -> None:
|
||||||
class PagedHistoryClient:
|
class PagedHistoryClient:
|
||||||
calls: list[int] = []
|
calls: list[int] = []
|
||||||
@@ -109,7 +62,7 @@ class PollPolicyTests(unittest.TestCase):
|
|||||||
|
|
||||||
with tempfile.TemporaryDirectory() as temporary_dir, patch.dict(
|
with tempfile.TemporaryDirectory() as temporary_dir, patch.dict(
|
||||||
"os.environ",
|
"os.environ",
|
||||||
{"MODELHUB_ARCHITECTURE_BACKFILL_PAGES_PER_CYCLE": "1"},
|
{"MODELHUB_ARCHITECTURE_BACKFILL_PAGES_PER_BATCH": "1"},
|
||||||
):
|
):
|
||||||
root = Path(temporary_dir)
|
root = Path(temporary_dir)
|
||||||
outcomes = root / "outcomes.jsonl"
|
outcomes = root / "outcomes.jsonl"
|
||||||
@@ -154,6 +107,53 @@ class PollPolicyTests(unittest.TestCase):
|
|||||||
self.assertNotIn("owner/model", persisted_progress)
|
self.assertNotIn("owner/model", persisted_progress)
|
||||||
self.assertNotIn("logs.invalid", persisted_progress)
|
self.assertNotIn("logs.invalid", persisted_progress)
|
||||||
|
|
||||||
|
def test_cold_start_backfill_finishes_all_pages_in_one_phase(self) -> None:
|
||||||
|
class ThreePageClient:
|
||||||
|
calls: list[int] = []
|
||||||
|
|
||||||
|
def list_tasks_page(self, **kwargs): # noqa: ANN003, ANN201
|
||||||
|
current = int(kwargs["current"])
|
||||||
|
self.calls.append(current)
|
||||||
|
return {
|
||||||
|
"data": {
|
||||||
|
"records": [
|
||||||
|
{
|
||||||
|
"taskId": f"task-{current}",
|
||||||
|
"modelId": f"owner/model-{current}",
|
||||||
|
"gpuType": "gpu-a",
|
||||||
|
"status": "success",
|
||||||
|
"verifyResult": 1,
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"pages": 3,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
with tempfile.TemporaryDirectory() as temporary_dir, patch.dict(
|
||||||
|
"os.environ",
|
||||||
|
{"MODELHUB_ARCHITECTURE_BACKFILL_PAGES_PER_BATCH": "1"},
|
||||||
|
):
|
||||||
|
root = Path(temporary_dir)
|
||||||
|
tracker = OutcomeTracker(
|
||||||
|
root / "outcomes.jsonl",
|
||||||
|
checkpoint_path=root / "checkpoint.json",
|
||||||
|
recent_path=root / "recent.jsonl",
|
||||||
|
)
|
||||||
|
client = ThreePageClient()
|
||||||
|
durable_phases: list[str] = []
|
||||||
|
progress = _complete_architecture_history_backfill(
|
||||||
|
modelhub_client=client, # type: ignore[arg-type]
|
||||||
|
outcome_tracker=tracker,
|
||||||
|
ledger_path=root / "ledger.jsonl",
|
||||||
|
progress_path=root / "backfill.json",
|
||||||
|
sync_callback=lambda phase: durable_phases.append(phase) or True,
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertTrue(progress["complete"])
|
||||||
|
self.assertEqual([1, 2, 3], client.calls)
|
||||||
|
self.assertEqual(["history_backfill"], durable_phases)
|
||||||
|
self.assertEqual(3, tracker.get_stats_report()["terminalRecords"])
|
||||||
|
|
||||||
def test_cleanup_contexts_merge_outcomes_with_older_ledger_entries(self) -> None:
|
def test_cleanup_contexts_merge_outcomes_with_older_ledger_entries(self) -> None:
|
||||||
with tempfile.TemporaryDirectory() as temporary_dir:
|
with tempfile.TemporaryDirectory() as temporary_dir:
|
||||||
root = Path(temporary_dir)
|
root = Path(temporary_dir)
|
||||||
|
|||||||
Reference in New Issue
Block a user