5 Commits

Author SHA1 Message Date
fb5e1079a8 Refresh GGUF download batch and fail closed on expired token 2026-10-08 19:13:45 +08:00
9ebc1f3666 update GGUF download batch to 370 models 2026-10-01 18:05:49 +08:00
78f8b5ad82 requeue downloads rejected for platform concurrency instead of discarding them
v2.3.0 burned through all 498 models in minutes with download_failed=498: the
platform already had 8 downloads running, every create returned "用户最多同时运行
8 个下载任务", and the worker treated a create failure as terminal (CREATE_FAILED)
and moved on.

Now a concurrency rejection puts the model back at the head of the queue and
waits for the next poll, so only genuine per-model failures (uniqueness check,
etc.) are recorded as failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 16:01:42 +08:00
251f21ac8f add DOWNLOAD_ONLY mode: run zhoushasha downloads without submitting any validation task
All 498 models in the current list already had their hygon and bi150 slots
reserved in earlier rounds, so this round only needs the downloads. Without a
switch the download list would be empty, since phase 2 only downloads models
reserved in the same run.

DOWNLOAD_ONLY skips the reservation phase entirely and downloads all of
ALL_MODEL_IDS. Also splits the worker into _reserve_phase / _download_phase and
gates bi150 behind ENABLE_BI150_SUBMIT for symmetry with hygon.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 15:56:45 +08:00
fae53f5a2a add a dedicated bi150 fallback account and skip hygon this round
bi150 coverage lagged hygon: the zero-credit accounts each take only one model,
so once all 34 were used the remaining models got no bi150 at all. Add
zhangyuanxi as the fallback so bi150 can catch up to hygon's coverage, using a
separate account from zhoukaile so the two cards don't compete for one quota.

hygon is already submitted for every model in the current 498-model list, so
ENABLE_HYGON_SUBMIT is off this round; flip it back on with a new model list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 15:39:26 +08:00
2 changed files with 1541 additions and 1450 deletions

View File

@@ -1,13 +1,12 @@
# xc_validation_strategy_gguf
GGUF 模型下载 + 验证任务提交流水线策略服务:批量创建 GGUF 模型下载任务(最大并发 8),
每个模型下载成功后立即提交 hygon / bi150 两个验证任务,之后保持 HTTP 服务存活供平台探活。
本版本为纯下载策略:使用 `zhoushasha` 账号批量创建 `main.py` 中未注释的 GGUF 模型下载任务(最大并发 8),不提交验证任务;之后保持 HTTP 服务存活供平台探活。
## 功能
- 自动登录 ModelHub 获取 Token(失败时回退到预设 Token)
- 启动时登录 ModelHub 获取新 Token,并检查账号身份及有效期;登录失败时停止,不回退到过期 Token
- 按流水线批量创建 GGUF 模型下载任务(HuggingFace 源,最大并发 8)
- 每个模型下载成功后,立即提交 hygon_k100-ai 与 Iluvatar_bi-150 两个验证任务(llamacpp 框架)
- 当前 `DOWNLOAD_ONLY=True`,不调用验证任务提交接口;验证任务由其他策略负责
- 下载成功的模型 ID 写入 `downloaded_success_models.txt`
- 暴露 `/health` 和 `/status` 接口满足平台运行时契约
@@ -15,7 +14,7 @@ GGUF 模型下载 + 验证任务提交流水线策略服务:批量创建 GGUF
```
.
├── main.py # 主入口:HTTP 服务 + 下载/提交流水线
├── main.py # 主入口:HTTP 服务 + 纯下载流程
├── Dockerfile # 平台镜像构建配置
├── requirements.txt # Python 依赖
└── downloaded_success_models.txt # 运行后自动生成,记录下载成功的模型

2976
main.py

File diff suppressed because it is too large Load Diff