88 lines
4.5 KiB
Markdown
88 lines
4.5 KiB
Markdown
# Durable log analysis — 2026-09-04
|
|
|
|
This report was produced from the live `agent-state` snapshot and the monthly
|
|
`agent-archive-2026-08` branch. Records were deduplicated by task identity before
|
|
classification; intent rows and state-sync generations were not counted as
|
|
independent model submissions.
|
|
|
|
## Data quality
|
|
|
|
- 15,178 raw outcome rows became 14,178 unique outcomes after deduplication.
|
|
- 13,318 unique terminal outcomes contained 733 successes and 12,585 failures.
|
|
- The hot intent WAL contained 9,213 rows, but only 1,299 unique model/GPU
|
|
routes. Capacity-full responses alone appeared 5,887 times.
|
|
- The ledger contained 1,028 platform tasks. A single model can legitimately
|
|
produce several tasks because GPU/framework routes and retries differ.
|
|
- Raw terminal success rate is biased downward: failures often terminate much
|
|
faster, while successful tasks can remain waiting/running for a long time.
|
|
Routing therefore uses attributable outcomes and confidence bounds rather
|
|
than this raw percentage.
|
|
|
|
## Deduplicated failure categories
|
|
|
|
| Category | Count | Share |
|
|
| --- | ---: | ---: |
|
|
| Ambiguous runtime | 3,067 | 24.37% |
|
|
| Parameters/template | 2,935 | 23.32% |
|
|
| Framework/architecture unsupported | 1,428 | 11.35% |
|
|
| Memory capacity | 1,169 | 9.29% |
|
|
| Platform infrastructure | 911 | 7.24% |
|
|
| Repository structure | 725 | 5.76% |
|
|
| Missing log evidence | 719 | 5.71% |
|
|
| Tokenizer compatibility | 576 | 4.58% |
|
|
| Model load | 382 | 3.04% |
|
|
| Context length | 318 | 2.53% |
|
|
| Architecture compatibility | 212 | 1.68% |
|
|
| Runtime memory | 55 | 0.44% |
|
|
| Backend operator | 44 | 0.35% |
|
|
| Generic validation failure | 43 | 0.34% |
|
|
| Attention backend | 1 | 0.01% |
|
|
|
|
The most useful structured codes were `MODEL_NOT_SUPPORTED` (1,700),
|
|
`EXECUTE_EMPTY_RESULT` (1,392), `PREFLIGHT_OOM` (1,166), `TOKENIZER_FAILED`
|
|
(696), `MODEL_FILE_NOT_FOUND` (590), `MODEL_LOAD_FAILED` (382), and
|
|
`CONTEXT_LENGTH_ERROR` (268).
|
|
|
|
## Representative evidence and algorithm changes
|
|
|
|
1. Capacity saturation repeatedly returned the account-limit message. This is
|
|
platform flow control, not model/framework failure. It is now recorded as
|
|
`capacity_deferred`, ends the current submission cycle after the pool has
|
|
tried every account, and does not enter success statistics.
|
|
2. The platform repeatedly required a failed non-`transformers` verification
|
|
before accepting `transformers`. Routing now uses a vetted non-transformers
|
|
route when one exists and treats the prerequisite response as a neutral
|
|
deferred result. Once durable history contains a non-transformers failure
|
|
for that model, transformers is unlocked and competes normally by success
|
|
confidence.
|
|
3. `PREFLIGHT_OOM` logs explicitly compare repository load size (including the
|
|
platform's 20% loading allowance) with available device memory. These remain
|
|
deterministic pre-submit blocks and deterministic queue cleanup evidence.
|
|
4. Explicit “framework does not support model/architecture” messages continue
|
|
to create exact GPU + framework + task + architecture blocks. Generic image,
|
|
driver, storage and connection errors remain platform-neutral and do not
|
|
poison compatibility scores.
|
|
5. Platform history often omits `framework`. Outcome sync now joins task IDs
|
|
with the local durable ledger and restores framework/task/profile metadata,
|
|
improving per-framework statistics without guessing from model names.
|
|
|
|
## Durable-state corrections
|
|
|
|
- Hot intent history retains unresolved intents plus 300 compact recent
|
|
terminal intents; older terminal attempts are discarded after extraction.
|
|
- Community raw samples are capped locally at 200 and excluded from the hot Git
|
|
snapshot. The aggregated GPU/framework statistics remain durable.
|
|
- Model/GPU official-capability cache is bounded to the 750 newest entries.
|
|
- Outcome compaction starts at 500 rows and retains 300 compact recent samples.
|
|
- An unchanged snapshot produces no Git commit. A failed push retries the exact
|
|
same commit and generation instead of creating a new generation every minute.
|
|
- Submission intent batches default to 100, reducing Git transactions while
|
|
preserving write-ahead recovery.
|
|
|
|
After this audit was completed, version `2026.09.04.2` changed the ongoing
|
|
retention model to `decision_state_only`. The extracted aggregate statistics,
|
|
compatibility rules, memory observations, active recovery state, and compact
|
|
recent samples remain durable, but new full outcome/intent archives and the
|
|
duplicated event stream are no longer produced. Previously created archive
|
|
branches remain untouched and are not needed during startup.
|