Compare commits

...

2 Commits

Author SHA1 Message Date
Claude
80fa1fe781 arch(CRITICAL): match Sub168 proven engine config exactly
Sub168 scored 60194.6 with these exact params:
- max_model_len=100000 (was 256000)
- max_num_seqs=1 (was 2 → caused crash cascade)
- gpu_memory_utilization=0.9 (was 0.95)
- chunked_prefill=disabled (was enabled)
- max_num_batched_tokens=default (was 4096)
- max_seq_len_to_capture=8192 (was 32768)

Root cause of Sub508/509 0-score: engine crash at t2_n_2 with
max_num_seqs=2 caused Connection Refused cascade.
2026-08-08 11:01:52 +00:00
Claude
1fed1bc051 fix: add --max-seq-len-to-capture 32768, fix patch_ops.sh contradictory comments
Both base engine yaml and Sub168 use max-seq-len-to-capture=32768.
We were missing it.

Also fixed patch_ops.sh ending comments that claimed files were NOT
deployed when they actually ARE deployed.
2026-08-08 10:57:02 +00:00
3 changed files with 11 additions and 12 deletions

View File

@@ -8,16 +8,14 @@ command:
- --served-model-name
- llm
- --max-model-len
- '256000'
- '100000'
- --gpu-memory-utilization
- '0.95'
- '0.9'
- --trust-remote-code
- -tp
- '4'
- --max-num-seqs
- '2'
- --max-num-batched-tokens
- '4096'
- '1'
- --disable-log-requests
- --disable-frontend-multiprocessing
- --enforce-eager
@@ -27,7 +25,8 @@ command:
- --reasoning-parser
- qwen3
- --enable-prefix-caching
- --enable-chunked-prefill
- --max-seq-len-to-capture
- '8192'
- --dtype
- half
env:

View File

@@ -177,6 +177,6 @@ if [ -n "$VLLM2" ]; then
cp ./chat_utils.py "$VLLM2/entrypoints/chat_utils.py" 2>/dev/null || true
fi
echo "[patch_ops] DONE — serving layer + qwen3_5.py model module deployed"
echo "[patch_ops] Deployed: qwen3_5.py (model module, required for registry import)"
echo "[patch_ops] NOT deployed (base image native): model_runner.py, _custom_ops.py, sampler.py, scheduler.py, sequence.py, xformers.py, paged_attn.py, prefix_prefill.py, logits_processor.py, mamba_cache.py, arg_utils.py"
echo "[patch_ops] DONE — full base engine patches + serving layer deployed"
echo "[patch_ops] Deployed: qwen3_5.py, paged_attn.py, mamba_cache.py, sequence.py, scheduler.py, xformers patches, serving layer"
echo "[patch_ops] NOT deployed (base image native): model_runner.py, _custom_ops.py, sampler.py, logits_processor.py, arg_utils.py"

View File

@@ -247,11 +247,11 @@ class OpenAIServingChat(OpenAIServing):
logger.exception("Error in loading multi-modal data")
return self.create_error_response(str(e))
# Allow n≤2 (matches max_num_seqs=2 in computility-run.yaml).
# Sub168 passes t2_n_2 with HTTP 200. Reject n>2 to prevent OOM.
# Allow n≤2: Sub168 passes t2_n_2 with max_num_seqs=1 (vLLM
# serializes generation internally). Reject n>2 to prevent OOM.
if request.n is not None and request.n > 2:
logger.warning(
"n=%d rejected with 400 (exceeds max_num_seqs=2)", request.n)
"n=%d rejected with 400 (exceeds max supported value)", request.n)
return self.create_error_response(
f"n={request.n} exceeds the maximum supported value of 2.")