fix(critical): disable thinking for tool_call requests — fixes d03_tool_call FAIL
Root cause: When tool_choice=auto + tools present, the model enters <think>...</think> mode by default. On BI-V100 hardware, decode is slow enough that thinking consumes the entire max_tokens budget, and the model finishes (finish=stop) before ever emitting <tool_call> XML. Sub168 reference: d03 in 2.12s with tools=1, finish=tool_calls Our sub509: d03 in 49.04s with tools=0, finish=stop — FAIL Fix: Two-layer defense: 1. protocol.py normalize_messages: when tools active + tool_choice=auto and thinking not explicitly set, auto-set enable_thinking=False 2. qwen3coder_tool_parser.py adjust_request: same logic as defense-in-depth 3. baseline.muh synced with actual computility-run.yaml
This commit is contained in:
@@ -19,13 +19,10 @@ vllm:
|
||||
max_model_len: 100000
|
||||
gpu_memory_utilization: 0.90
|
||||
tensor_parallel: 4
|
||||
max_num_seqs: 2
|
||||
max_num_batched_tokens: 4096
|
||||
max_seq_len_to_capture: 32768
|
||||
max_num_seqs: 1
|
||||
trust_remote_code: true
|
||||
disable_log_requests: true
|
||||
disable_frontend_multiprocessing: true
|
||||
enable_chunked_prefill: true
|
||||
enable_auto_tool_choice: true
|
||||
tool_call_parser: qwen3_coder
|
||||
reasoning_parser: qwen3
|
||||
|
||||
Reference in New Issue
Block a user