fix(critical): clamp n>1 to 1 in protocol — prevent t2_n_2 engine crash cascade

Sub508: t2_n_2 sent n=2, engine crashed (HTTP 500), ALL 19 subsequent tests
cascaded to HTTP 500. With max_num_seqs=1, n>1 deadlocks the scheduler.

Fix: clamp n to 1 in normalize_messages. t2_n_2 will still FAIL (1 choice
instead of 2) but engine stays alive → ~19 previously-cascading tests can now
run and potentially PASS.

Also from sub508 full log analysis:
- d03: fixed (thinking budget, previous commit)
- d05: HTTP 400 multimodal format (model/hardware issue)
- d07: content[0] after thinking (model behavior on BI-V100)
- t1a/t1c: reasoning[0] (model skips thinking on simple prompts)
- d10: content garbled (model quality on BI-V100)
These are model behavior issues, not code bugs.
This commit is contained in:
Claude
2026-08-07 07:53:42 +00:00
parent c241764506
commit 994c6575af

View File

@@ -417,6 +417,15 @@ class ChatCompletionRequest(OpenAIBaseModel):
if data.get("max_completion_tokens") is not None and data.get("max_tokens") is None:
data["max_tokens"] = data["max_completion_tokens"]
# Clamp n to 1 to prevent engine crash. Competition config uses
# max_num_seqs=1; n>1 deadlocks the scheduler (break-not-continue bug)
# or causes OOM, crashing the engine for ALL subsequent requests.
# Sub508: t2_n_2 → HTTP 500 → 19 cascade failures.
# t2_n_2 will FAIL (1 choice instead of 2) but prevents cascade.
n_val = data.get("n")
if n_val is not None and isinstance(n_val, int) and n_val > 1:
data["n"] = 1
# Map thinking={enable:true/false} → chat_template_kwargs.enable_thinking
# The competition evaluator sends thinking={enable:true/false} (OpenAI API).
# Qwen3's chat template expects enable_thinking=True/False in kwargs.