Compare commits

...

7 Commits

Author SHA1 Message Date
Claude
e0344b1730 fix(critical): disable thinking for tool_call requests — fixes d03_tool_call FAIL
Root cause: When tool_choice=auto + tools present, the model enters
<think>...</think> mode by default. On BI-V100 hardware, decode is slow
enough that thinking consumes the entire max_tokens budget, and the model
finishes (finish=stop) before ever emitting <tool_call> XML.

Sub168 reference: d03 in 2.12s with tools=1, finish=tool_calls
Our sub509: d03 in 49.04s with tools=0, finish=stop — FAIL

Fix: Two-layer defense:
1. protocol.py normalize_messages: when tools active + tool_choice=auto
   and thinking not explicitly set, auto-set enable_thinking=False
2. qwen3coder_tool_parser.py adjust_request: same logic as defense-in-depth
3. baseline.muh synced with actual computility-run.yaml
2026-08-07 07:45:28 +00:00
dylanyunlon
812c374f7a fix(critical): sync baseline.muh max_model_len=100000 gpu_mem=0.90 — match computility-run.yaml
Root cause of job 105 scoring 0.0:
- baseline.muh had max_model_len=256000 + gpu_memory_utilization=0.95
- computility-run.yaml had the safe values (100000 + 0.90)
- Platform scheduler sent baseline.muh values to docker run command
- Result: OOM on KV cache allocation → service crash → 881/881 Connection refused

Diagnosis from submit日志:
- benchmark-agent marked success (model loaded OK)
- But service crashed before evaluation started
- All 881 replay requests → Connection refused
- All 5 opencompass benchmarks → 0.0 (aime, gpqa, hle, simpleqa, longbench)

Fix: sync baseline.muh to match computility-run.yaml safe values
2026-08-07 07:22:52 +00:00
dylanyunlon
19879ccaae docs: pipeline ground truth — CCCL parity audit, SM100 benchmark data, injection status
Key findings:
- scale_mem_bound: 11/11 FULL PARITY with CCCL
- 27/27 tuning headers complete
- All vllm injection points already deployed via Python modifications
- gen_patch.py role: verification tool (enginex has no .cu source)
- CCCL SM100 benchmark values extracted to JSON for reference
2026-08-07 07:20:34 +00:00
dylanyunlon
611c491b0f audit: CCCL vs muh parity check — scale_mem_bound 11/11 PASS, SM100 benchmark values extracted
- scale_mem_bound: FULL PARITY with CCCL (all 11 test cases match)
- Extracted all SM100 benchmark annotations from 26 tuning headers
- reduce: 7 SM100 tunings + 3 deterministic (SM90/SM86)
- scan: 7 SM100 lookback tunings with delay policies
- Generated machine-readable JSON with benchmark runner params
- Identified 5 pending verification items for BI-V100 hardware
2026-08-07 07:19:00 +00:00
dylan-claude
95d03147e7 fix(deploy): reduce max-model-len 131072→100000, remove chunked-prefill, single-seq — prevent OOM crash
CCCL design reference: block_topk_air.cuh tile_items = threads * items must fit hardware.
max-model-len is vllm's tile size — 131072 overflows BI-V100 VRAM budget.
Submission 508 failed with 100% Connection refused = service never started.
2026-08-07 07:16:28 +00:00
Claude
2c353da28b fix(protocol): reduce HTTP 400 errors for replay — accept tool_choice=none + extra fields
1. tool_choice='none' now accepted per OpenAI spec (strip and continue).
   Previously raised ValueError, causing 400 on replay requests.

2. Pydantic extra='forbid' → extra='ignore'. Real-world replay requests
   from Tencent API contain fields like service_tier, store, metadata,
   reasoning_effort etc. that our model doesn't declare. forbid rejects
   them all; ignore silently drops them.

Sub 168 had 77 http_400 errors in replay — these two fixes should
eliminate most of them, improving successful request count and score.

CCCL tuning_transform.cuh pattern: accept all valid input configurations
gracefully (policy_selector handles unknown cc values with fallback).
2026-08-07 07:12:35 +00:00
Claude
c2bca49aa8 fix(serving_chat): shallow copy bug — [[]] * n and [parser] * n share references
When n>=2, all_previous_token_ids entries pointed to the SAME list,
so appending tokens for choice 0 corrupted choice 1's history.
Same for tool_parsers: all choices shared one stateful parser instance.

Changed to list comprehensions that create independent objects.

Found via CCCL result_policy.cuh read: distributed result delivery
requires isolated per-rank state — same principle applies to
per-choice token tracking in vLLM streaming.
2026-08-07 07:11:11 +00:00
8 changed files with 519 additions and 16 deletions

134
CCCL_MUH_PARITY_AUDIT.md Normal file
View File

@@ -0,0 +1,134 @@
================================================================================
CCCL vs muh 精确比对审计报告
================================================================================
### 1. scale_mem_bound 函数 parity check
------------------------------------------------------------
float32 (CCCL SM100 reduce) CCCL=( 16i, 512t,tile= 32768B) muh=( 16i, 512t,tile= 32768B) ✓
float64 (CCCL SM100 reduce) CCCL=( 8i, 640t,tile= 40960B) muh=( 8i, 640t,tile= 40960B) ✓
accum8 (CCCL SM100 reduce) CCCL=( 7i, 512t,tile= 28672B) muh=( 7i, 512t,tile= 28672B) ✓
scan 4B (CCCL SM100 scan) CCCL=( 22i, 384t,tile= 33792B) muh=( 22i, 384t,tile= 33792B) ✓
scan 8B (CCCL SM100 scan) CCCL=( 11i, 416t,tile= 36608B) muh=( 11i, 416t,tile= 36608B) ✓
det float32 SM90 CCCL=( 13i, 224t,tile= 11648B) muh=( 13i, 224t,tile= 11648B) ✓
det float64 SM86 CCCL=( 5i, 128t,tile= 5120B) muh=( 5i, 128t,tile= 5120B) ✓
1-byte type CCCL=( 32i, 256t,tile= 8192B) muh=( 32i, 256t,tile= 8192B) ✓
2-byte type CCCL=( 32i, 256t,tile= 16384B) muh=( 32i, 256t,tile= 16384B) ✓
16-byte type (int128) CCCL=( 4i, 256t,tile= 16384B) muh=( 4i, 256t,tile= 16384B) ✓
SMEM cap test (should trigger) CCCL=( 8i, 768t,tile= 49152B) muh=( 8i, 768t,tile= 49152B) ✓
→ scale_mem_bound: FULL PARITY ✓
### 2. reduce tuning: CCCL SM100值 → BI-V100 scale_mem_bound适配后
------------------------------------------------------------
CCCL benchmarked on SM100 → muh should use scale_mem_bound for BI-V100
Key: reduce loads to REGISTERS not SMEM → SMEM cap rarely triggers
float32_plus_o4 @4B: scaled=(16i, 512t) tile= 32768B (66.7%)
float32_plus_o4 @8B: scaled=( 8i, 512t) tile= 32768B (66.7%)
float64_plus_o4 @4B: scaled=(16i, 640t) tile= 40960B (83.3%)
float64_plus_o4 @8B: scaled=( 8i, 640t) tile= 40960B (83.3%)
accum8_plus_o4 @4B: scaled=(15i, 512t) tile= 30720B (62.5%)
accum8_plus_o4 @8B: scaled=( 7i, 512t) tile= 28672B (58.3%)
accum8_plus_o8 @4B: scaled=(15i, 512t) tile= 30720B (62.5%)
accum8_plus_o8 @8B: scaled=( 7i, 512t) tile= 28672B (58.3%)
det_float32_sm90 @4B: scaled=(13i, 224t) tile= 11648B (23.7%)
det_float32_sm90 @8B: scaled=( 6i, 224t) tile= 10752B (21.9%)
det_float32_sm86 @4B: scaled=( 6i, 224t) tile= 5376B (10.9%)
det_float32_sm86 @8B: scaled=( 3i, 224t) tile= 5376B (10.9%)
det_float64_sm86 @4B: scaled=(11i, 128t) tile= 5632B (11.5%)
det_float64_sm86 @8B: scaled=( 5i, 128t) tile= 5120B (10.4%)
default_fallback @4B: scaled=(16i, 256t) tile= 16384B (33.3%)
default_fallback @8B: scaled=( 8i, 256t) tile= 16384B (33.3%)
### 3. muh bi100 reduce当前值 vs CCCL参考
------------------------------------------------------------
muh改用了更大的items (24 vs SM100的16)来补偿16 SMs
这是对的——reduce加载到寄存器,SMEM不是瓶颈
★ float32 plus (paged_attention score reduction — 83% weight):
CCCL SM100: items=16, threads=512, vec=2
muh BI-V100: items=24, threads=512, vec=2
理由: 16 SMs vs 148 SMs, 每个CTA需要处理更多数据
tile对比: SM100=512*16*4=32768B | BI-V100=512*24*4=49152B (exactly 48KB)
→ items=24 用满了SMEM → 合理但有风险,如果BlockReduce实际占SMEM则溢出
→ 但注释说reduce不用BlockLoad(loads to registers) → 安全
### 4. scan tuning: CCCL SM100 → BI-V100 SMEM约束
------------------------------------------------------------
Scan DOES use BlockLoad staging in SMEM → tile_bytes ≤ 49152 is HARD
lookback_1B_o4 @1B: tpb= 512 ipt=18 tile= 9216B ✓
lookback_1B_o4 @2B: tpb= 512 ipt=18 tile= 18432B ✓
lookback_1B_o4 @4B: tpb= 512 ipt=18 tile= 36864B ✓
lookback_1B_o4 @8B: tpb= 512 ipt=18 tile= 73728B ✗ OVERFLOW → max_items=12
lookback_2B_o4 @1B: tpb= 512 ipt=13 tile= 6656B ✓
lookback_2B_o4 @2B: tpb= 512 ipt=13 tile= 13312B ✓
lookback_2B_o4 @4B: tpb= 512 ipt=13 tile= 26624B ✓
lookback_2B_o4 @8B: tpb= 512 ipt=13 tile= 53248B ✗ OVERFLOW → max_items=12
lookback_4B_o4 @1B: tpb= 384 ipt=22 tile= 8448B ✓
lookback_4B_o4 @2B: tpb= 384 ipt=22 tile= 16896B ✓
lookback_4B_o4 @4B: tpb= 384 ipt=22 tile= 33792B ✓
lookback_4B_o4 @8B: tpb= 384 ipt=22 tile= 67584B ✗ OVERFLOW → max_items=16
lookback_8B_o4 @1B: tpb= 416 ipt=23 tile= 9568B ✓
lookback_8B_o4 @2B: tpb= 416 ipt=23 tile= 19136B ✓
lookback_8B_o4 @4B: tpb= 416 ipt=23 tile= 38272B ✓
lookback_8B_o4 @8B: tpb= 416 ipt=23 tile= 76544B ✗ OVERFLOW → max_items=14
lookback_1B_o8 @1B: tpb= 384 ipt=14 tile= 5376B ✓
lookback_1B_o8 @2B: tpb= 384 ipt=14 tile= 10752B ✓
lookback_1B_o8 @4B: tpb= 384 ipt=14 tile= 21504B ✓
lookback_1B_o8 @8B: tpb= 384 ipt=14 tile= 43008B ✓
lookback_4B_o8 @1B: tpb= 416 ipt=19 tile= 7904B ✓
lookback_4B_o8 @2B: tpb= 416 ipt=19 tile= 15808B ✓
lookback_4B_o8 @4B: tpb= 416 ipt=19 tile= 31616B ✓
lookback_4B_o8 @8B: tpb= 416 ipt=19 tile= 63232B ✗ OVERFLOW → max_items=14
lookback_8B_o8 @1B: tpb= 320 ipt=22 tile= 7040B ✓
lookback_8B_o8 @2B: tpb= 320 ipt=22 tile= 14080B ✓
lookback_8B_o8 @4B: tpb= 320 ipt=22 tile= 28160B ✓
lookback_8B_o8 @8B: tpb= 320 ipt=22 tile= 56320B ✗ OVERFLOW → max_items=19
关键发现:
- scan lookback_4B_o4: items=22, threads=384 → tile@4B=33792 ✓ tile@8B=67584 ✗
- scan lookback_8B_o4: items=23, threads=416 → tile@8B=76544 ✗
- 这些值在SM100上是安全的(228KB SMEM),但在BI-V100(48KB)上必须降级
- muh已经做了降级(用scale_mem_bound),但需要验证降级后的值是否正确
### 5. CCCL benchmark format解析
------------------------------------------------------------
NVIDIA的benchmark注释格式:
ipt_<items>.tpb_<threads>.ns_<delay>.dcid_<algo>.l2w_<latency>.trp_<transpose>.ld_<load>
后跟4个浮点数: 在[2^16, 2^20, 2^24, 2^28]四个problem size下的speedup
dcid映射:
0 = no_delay
1 = fixed_delay
2 = exp_backoff
3 = exp_backoff_jitter
4 = exp_backoff_jitter_window
5 = exp_backon_jitter_window
6 = exp_backon_jitter
7 = exp_backon
### 6. 竞赛关键路径优先级
------------------------------------------------------------
Token吞吐加权值 = Output_TPS × 16.796 + Input_TPS × 2.799 + Cache_TPS × 0.56
→ Output_TPS权重83%, Input_TPS权重14%, Cache_TPS权重3%
decode热路径 (Output TPS):
1. paged_attention score reduction → reduce (DONE: muh tuned)
2. softmax denominator prefix-sum → scan (DONE: muh tuned)
3. top-k/top-p sampling → topk/radix_sort (DONE: muh tuned)
4. RMSNorm/SiLU/RoPE element-wise → transform (DONE: muh tuned)
prefill热路径 (Input TPS):
5. flash_attention → scan + reduce
6. MoE expert routing → select_if + reduce_by_key
cache热路径 (Cache TPS):
7. KV cache block copy → batch_memcpy (DONE: muh tuned)
### 7. 待验证的关键问题
------------------------------------------------------------
1. reduce items=24: 虽然loads to registers, 但实际BlockReduce<WARP_REDUCTIONS>的SMEM用量需要确认
2. scan delay参数: 0.5x/0.6x缩放是启发式, 需要BI-V100实测L2 write latency
3. LOAD_LDG vs LOAD_DEFAULT: topk bench显示BI-V100上LOAD_DEFAULT更快, reduce/scan可能同理
4. SM count=16 → wave efficiency: 所有tuning都需要重新算occupancy
5. transform bytes_in_flight: 从18GB/s改为56GB/s后items需要相应增大

50
PIPELINE_GROUND_TRUTH.md Normal file
View File

@@ -0,0 +1,50 @@
# muh Pipeline Ground Truth — 2026-08-07
## 管道实际状态(不是设计稿,是已部署代码的真实描述)
### scale_mem_bound: FULL PARITY ✓
11/11测试用例与CCCL `cub::detail::scale_mem_bound` 完全匹配。
返回值顺序 `{items_per_thread, threads_per_block}` — items-first,与CCCL一致。
### C++ Tuning Headers: 27/27 ✓
所有26个算法(+common)都有bi100 header,`policy_selector::operator()` 接受
`hardware_capability` 参数。SMEM overflow保护覆盖所有type_size。
### Injection现状(enginex没有.cu源码)
| 注入位置 | 状态 | 值 | commit |
|---------|------|-----|--------|
| prefix_prefill.py BLOCK | ✓ 已手动修改 | BLOCK=64, WARPS=4 | 多个commit |
| paged_attn.py _PARTITION_SIZE | ✓ 保持默认 | 512 | — |
| paged_attn.py V1/V2 dispatch | ✓ 已手动修改 | use_v1 threshold | cbd1f08 |
| _custom_ops.py SMEM | ✓ 已手动修改 | 48KB | 16f0b30 |
| triton_flash_attention.py | ✓ 已添加BI-V100 configs | BLOCK=32/64 | 多个commit |
| protocol.py 兼容性 | ✓ 已修复 | max_completion_tokens等 | 2c353da |
### gen_patch.py 角色
设计时期望: C++ header → unified diff → vllm .cu文件
实际情况: enginex只有Python + .so, 没有.cu源码
当前角色: 文档工具 + 验证(确认header值与已部署Python代码一致)
### CCCL SM100 Benchmark数据(从源码提取,已存入cccl_sm100_benchmark_values.json)
**Reduce** (paged_attention score reduction, Output TPS 83%权重):
- float32+plus: items=16, threads=512, vec=2, speedup=[1.061, 1.000, 1.065, 1.167]
- float64+plus: items=16, threads=640, vec=1, speedup=[1.018, 1.000, 1.016, 1.057]
**Scan** (softmax prefix-sum):
- 4B lookback: items=22, threads=384, delay=1904ns/dcid=6/l2w=830, speedup=[1.148, 0.997, 1.140, 1.463]
- 8B lookback: items=23, threads=416, delay=772ns/dcid=5/l2w=710, speedup=[1.089, 1.016, 1.086, 1.265]
**muh BI-V100适配**:
- reduce float32: items=24(+50%), threads=512(=), vec=2(=) → 补偿16 SMs
- scan 4B: 通过scale_mem_bound自动适配(items=22 @4B安全, @8B降级到16)
- delay参数: ns×0.5, l2w×0.6 (启发式, 待实测)
### 竞赛门槛
- 功能测试: 50+ TC, 项目看板14个FEA item覆盖
- 效果测试: benchmark偏差 ≤ ±4%
- 性能测试: Token吞吐加权值 ≥ 8000
- Output TPS × 16.796 (83%) → reduce/scan/topk
- Input TPS × 2.799 (14%) → scan/transform
- Cache TPS × 0.56 (3%) → batch_memcpy

View File

@@ -16,16 +16,13 @@
vllm:
model_path: /model
served_model_name: llm
max_model_len: 256000
gpu_memory_utilization: 0.95
max_model_len: 100000
gpu_memory_utilization: 0.90
tensor_parallel: 4
max_num_seqs: 2
max_num_batched_tokens: 4096
max_seq_len_to_capture: 32768
max_num_seqs: 1
trust_remote_code: true
disable_log_requests: true
disable_frontend_multiprocessing: true
enable_chunked_prefill: true
enable_auto_tool_choice: true
tool_call_parser: qwen3_coder
reasoning_parser: qwen3

View File

@@ -0,0 +1,278 @@
{
"source": "cccl_upstream/cub/cub/device/dispatch/tuning/tuning_*.cuh",
"extracted_by": "automated audit from CCCL source code",
"reduce": {
"sm100_float32_plus_o4": {
"items": 16,
"threads": 512,
"vec": 2,
"benchmark": "ipt_16.tpb_512.ipv_2",
"speedup": [
1.061295,
1.0,
1.065478,
1.167139
]
},
"sm100_float64_plus_o4": {
"items": 16,
"threads": 640,
"vec": 1,
"benchmark": "ipt_16.tpb_640.ipv_1",
"speedup": [
1.017834,
1.0,
1.015835,
1.057092
]
},
"sm100_accum8_plus_o4": {
"items": 15,
"threads": 512,
"vec": 2,
"benchmark": "ipt_15.tpb_512.ipv_2",
"speedup": [
1.019887,
1.0,
1.017636,
1.058036
]
},
"sm100_accum8_plus_o8": {
"items": 15,
"threads": 512,
"vec": 1,
"benchmark": "ipt_15.tpb_512.ipv_1",
"speedup": [
1.019414,
1.0,
1.017218,
1.057143
]
},
"sm90_det_float32": {
"items": 13,
"threads": 224,
"benchmark": "ipt_13.tpb_224",
"speedup": [
1.107188,
1.009709,
1.097114,
1.31682
]
},
"sm86_det_float32": {
"items": 6,
"threads": 224,
"benchmark": "ipt_6.tpb_224",
"speedup": [
1.034383,
1.0,
1.032097,
1.090909
]
},
"sm86_det_float64": {
"items": 11,
"threads": 128,
"benchmark": "ipt_11.tpb_128",
"speedup": [
1.232089,
1.002124,
1.245336,
1.582279
]
}
},
"scan": {
"sm100_lookback_1B_o4": {
"items": 18,
"threads": 512,
"delay": {
"ns": 768,
"dcid": 7,
"l2w": 820
},
"load": {
"transpose": 1,
"modifier": 0
},
"benchmark": "ipt_18.tpb_512.ns_768.dcid_7.l2w_820.trp_1.ld_0",
"speedup": [
1.188818,
1.005682,
1.173041,
1.305288
]
},
"sm100_lookback_2B_o4": {
"items": 13,
"threads": 512,
"delay": {
"ns": 1384,
"dcid": 7,
"l2w": 720
},
"load": {
"transpose": 1,
"modifier": 0
},
"benchmark": "ipt_13.tpb_512.ns_1384.dcid_7.l2w_720.trp_1.ld_0",
"speedup": [
1.128443,
1.002841,
1.119688,
1.307692
]
},
"sm100_lookback_4B_o4": {
"items": 22,
"threads": 384,
"delay": {
"ns": 1904,
"dcid": 6,
"l2w": 830
},
"load": {
"transpose": 1,
"modifier": 0
},
"benchmark": "ipt_22.tpb_384.ns_1904.dcid_6.l2w_830.trp_1.ld_0",
"speedup": [
1.148442,
0.997167,
1.139902,
1.462651
]
},
"sm100_lookback_8B_o4": {
"items": 23,
"threads": 416,
"delay": {
"ns": 772,
"dcid": 5,
"l2w": 710
},
"load": {
"transpose": 1,
"modifier": 0
},
"benchmark": "ipt_23.tpb_416.ns_772.dcid_5.l2w_710.trp_1.ld_0",
"speedup": [
1.089468,
1.015581,
1.08563,
1.264583
]
},
"sm100_lookback_1B_o8": {
"items": 14,
"threads": 384,
"delay": {
"ns": 228,
"dcid": 7,
"l2w": 775
},
"load": {
"transpose": 1,
"modifier": 1
},
"benchmark": "ipt_14.tpb_384.ns_228.dcid_7.l2w_775.trp_1.ld_1",
"speedup": [
1.10721,
1.0,
1.100637,
1.307692
]
},
"sm100_lookback_4B_o8": {
"items": 19,
"threads": 416,
"delay": {
"ns": 956,
"dcid": 7,
"l2w": 550
},
"load": {
"transpose": 1,
"modifier": 1
},
"benchmark": "ipt_19.tpb_416.ns_956.dcid_7.l2w_550.trp_1.ld_1",
"speedup": [
1.146142,
0.99435,
1.137459,
1.455636
]
},
"sm100_lookback_8B_o8": {
"items": 22,
"threads": 320,
"delay": {
"ns": 328,
"dcid": 2,
"l2w": 965
},
"load": {
"transpose": 1,
"modifier": 0
},
"benchmark": "ipt_22.tpb_320.ns_328.dcid_2.l2w_965.trp_1.ld_0",
"speedup": [
1.080133,
1.0,
1.075577,
1.248963
]
}
},
"benchmark_runner_params": {
"reduce": {
"items_range": "7:24:1",
"threads_range": "128:1024:32",
"vec_pow2_range": "1:2:1",
"problem_sizes": [
"2^16",
"2^20",
"2^24",
"2^28"
]
},
"scan_lookback": {
"items_range": "7:24:1",
"threads_range": "128:1024:32",
"delay_ns_range": "0:2048:4",
"delay_algo_range": "0:7:1",
"l2w_range": "0:1200:5",
"transpose_range": "0:1:1",
"load_range": "0:1:1",
"problem_sizes": [
"2^16",
"2^20",
"2^24",
"2^28",
"2^32"
]
},
"topk": {
"items_range": "7:24:1",
"threads_range": "128:1024:32",
"load_algo_range": "0:2:1"
},
"radix_sort": {
"items_range": "7:24:1",
"threads_range": "128:1024:32",
"radix_bits": 8
}
},
"dcid_mapping": {
"0": "no_delay",
"1": "fixed_delay",
"2": "exponential_backoff",
"3": "exponential_backoff_jitter",
"4": "exponential_backoff_jitter_window",
"5": "exponential_backon_jitter_window",
"6": "exponential_backon_jitter",
"7": "exponential_backon"
}
}

View File

@@ -8,21 +8,16 @@ command:
- --served-model-name
- llm
- --max-model-len
- '131072'
- '100000'
- --gpu-memory-utilization
- '0.90'
- --trust-remote-code
- -tp
- '4'
- --max-num-seqs
- '2'
- '1'
- --disable-log-requests
- --disable-frontend-multiprocessing
- --max-num-batched-tokens
- '8192'
- --enable-chunked-prefill
- --max-seq-len-to-capture
- '32768'
- --enforce-eager
- --enable-auto-tool-choice
- --tool-call-parser

View File

@@ -57,7 +57,10 @@ class CustomChatCompletionMessageParam(TypedDict, total=False):
class OpenAIBaseModel(BaseModel):
# OpenAI API does not allow extra fields
model_config = ConfigDict(extra="forbid")
# Real-world clients (replay, third-party SDKs) may send extra fields
# like service_tier, store, metadata, reasoning_effort, etc.
# "ignore" accepts the request and silently drops unknown fields.
model_config = ConfigDict(extra="ignore")
class ErrorResponse(OpenAIBaseModel):
@@ -418,13 +421,30 @@ class ChatCompletionRequest(OpenAIBaseModel):
# The competition evaluator sends thinking={enable:true/false} (OpenAI API).
# Qwen3's chat template expects enable_thinking=True/False in kwargs.
thinking = data.get("thinking")
thinking_explicitly_set = False
if isinstance(thinking, dict):
enable = thinking.get("enable")
if enable is not None:
thinking_explicitly_set = True
ctk = data.get("chat_template_kwargs") or {}
ctk["enable_thinking"] = bool(enable)
data["chat_template_kwargs"] = ctk
# CRITICAL: When tools are present with tool_choice=auto and thinking
# is NOT explicitly requested, disable thinking to preserve token budget
# for tool call XML generation. Without this, the model spends all
# tokens on <think>...</think> and finishes before emitting <tool_call>.
# This matches the competition reference (sub168: d03 in 2.12s).
if not thinking_explicitly_set:
has_tools = data.get("tools") is not None and len(data.get("tools", [])) > 0
tc = data.get("tool_choice")
tool_choice_active = (tc == "auto" or (tc is None and has_tools)
or isinstance(tc, dict))
if has_tools and tool_choice_active:
ctk = data.get("chat_template_kwargs") or {}
ctk["enable_thinking"] = False
data["chat_template_kwargs"] = ctk
messages = data.get("messages")
if not isinstance(messages, list):
return data
@@ -517,6 +537,12 @@ class ChatCompletionRequest(OpenAIBaseModel):
# if "tool_choice" is specified -- validation
if "tool_choice" in data:
# "none" means don't use any tools — valid per OpenAI spec,
# just strip tool_choice and let vLLM ignore tools.
if data["tool_choice"] == "none":
del data["tool_choice"]
return data
# ensure that if "tool choice" is specified, tools are present
if "tools" not in data or data["tools"] is None:
raise ValueError(

View File

@@ -77,6 +77,28 @@ class Qwen3CoderToolParser(ToolParser):
logger.debug("vLLM Successfully imported tool parser %s !",
self.__class__.__name__)
def adjust_request(
self, request: "ChatCompletionRequest") -> "ChatCompletionRequest":
"""Disable thinking when tools are active with auto choice.
On BI-V100 hardware, the model's <think>...</think> phase can consume
the entire max_tokens budget, leaving no room for the <tool_call> XML.
Competition reference (sub168) completes d03_tool_call in 2.12s with
tools=1; our sub509 took 49s with tools=0 because thinking ate the
budget. Disabling thinking for tool-call requests ensures the model
emits tool XML within the token budget.
"""
if (request.tools and request.tool_choice in ("auto", None)
and not isinstance(request.tool_choice,
type(None).__class__)):
# Only override if thinking was not explicitly requested
ctk = request.chat_template_kwargs or {}
if "enable_thinking" not in ctk:
ctk = dict(ctk) # shallow copy
ctk["enable_thinking"] = False
request.chat_template_kwargs = ctk
return request
def _generate_tool_call_id(self) -> str:
return f"call_{uuid.uuid4().hex[:24]}"

View File

@@ -348,7 +348,7 @@ class OpenAIServingChat(OpenAIServing):
# parsing and reasoning parsing (both require full-history context).
if tool_choice_auto or use_reasoning:
previous_texts = [""] * num_choices
all_previous_token_ids = [[]] * num_choices
all_previous_token_ids = [[] for _ in range(num_choices)]
else:
previous_texts, all_previous_token_ids = None, None
@@ -357,7 +357,8 @@ class OpenAIServingChat(OpenAIServing):
if tool_choice_auto and self.tool_parser:
tool_parsers: List[Optional[ToolParser]] = [
self.tool_parser(tokenizer)
] * num_choices
for _ in range(num_choices)
]
else:
tool_parsers = [None] * num_choices
except RuntimeError as e: