fix(critical): match Sub168 proven config — max_model_len=100K, max_num_seqs=1, gpu_mem=0.9
Root cause analysis of Sub508 (41.2% score): 1. max_model_len=256000 → 100000 (Sub168 value) - Reduces KV cache preallocation by 2.56x - d01: should drop from 95.87s to ~8-10s - Frees GPU memory for stable inference 2. max_num_seqs=2 → 1 - Eliminates t2_n_2 OOM crash that killed engine - Sub508 lost 23 tests + 881 replay to this single crash 3. gpu_memory_utilization=0.95 → 0.9 (matches Sub168 docker log) 4. serving_chat.py content fallback improved for d07
This commit is contained in:
@@ -8,14 +8,14 @@ command:
|
||||
- --served-model-name
|
||||
- llm
|
||||
- --max-model-len
|
||||
- '256000'
|
||||
- '100000'
|
||||
- --gpu-memory-utilization
|
||||
- '0.95'
|
||||
- '0.9'
|
||||
- --trust-remote-code
|
||||
- -tp
|
||||
- '4'
|
||||
- --max-num-seqs
|
||||
- '2'
|
||||
- '1'
|
||||
- --max-num-batched-tokens
|
||||
- '4096'
|
||||
- --disable-log-requests
|
||||
|
||||
@@ -955,14 +955,26 @@ class OpenAIServingChat(OpenAIServing):
|
||||
# all output as reasoning with no content.
|
||||
content_for_message = output_text
|
||||
if not content_for_message and reasoning_text:
|
||||
# For tool-call paths, skip fallback (output must be raw XML)
|
||||
if request.tools and request.tool_choice in ("auto", None):
|
||||
# For tool-call paths with active tool_choice, skip fallback
|
||||
# (output must be raw XML)
|
||||
_is_active_tool_path = (
|
||||
request.tools
|
||||
and request.tool_choice in ("auto", "required")
|
||||
and self.enable_auto_tools and self.tool_parser)
|
||||
if _is_active_tool_path:
|
||||
pass
|
||||
else:
|
||||
# Use the last paragraph of reasoning as content
|
||||
lines = [l for l in reasoning_text.strip().split('\n') if l.strip()]
|
||||
if lines:
|
||||
content_for_message = lines[-1]
|
||||
# Use the last non-empty paragraph of reasoning as content.
|
||||
# Split on double-newline first (paragraphs), fall back to
|
||||
# lines. This produces more coherent content than a single
|
||||
# line when the model wrote a multi-paragraph reasoning block.
|
||||
paras = [p.strip() for p in reasoning_text.strip().split('\n\n') if p.strip()]
|
||||
if paras:
|
||||
content_for_message = paras[-1]
|
||||
else:
|
||||
lines = [l for l in reasoning_text.strip().split('\n') if l.strip()]
|
||||
if lines:
|
||||
content_for_message = lines[-1]
|
||||
if not content_for_message:
|
||||
content_for_message = reasoning_text[:500]
|
||||
|
||||
|
||||
Reference in New Issue
Block a user