Compare commits
7 Commits
cbd1f08a3e
...
e0344b1730
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
e0344b1730 | ||
|
|
812c374f7a | ||
|
|
19879ccaae | ||
|
|
611c491b0f | ||
|
|
95d03147e7 | ||
|
|
2c353da28b | ||
|
|
c2bca49aa8 |
134
CCCL_MUH_PARITY_AUDIT.md
Normal file
134
CCCL_MUH_PARITY_AUDIT.md
Normal file
@@ -0,0 +1,134 @@
|
||||
================================================================================
|
||||
CCCL vs muh 精确比对审计报告
|
||||
================================================================================
|
||||
|
||||
### 1. scale_mem_bound 函数 parity check
|
||||
------------------------------------------------------------
|
||||
float32 (CCCL SM100 reduce) CCCL=( 16i, 512t,tile= 32768B) muh=( 16i, 512t,tile= 32768B) ✓
|
||||
float64 (CCCL SM100 reduce) CCCL=( 8i, 640t,tile= 40960B) muh=( 8i, 640t,tile= 40960B) ✓
|
||||
accum8 (CCCL SM100 reduce) CCCL=( 7i, 512t,tile= 28672B) muh=( 7i, 512t,tile= 28672B) ✓
|
||||
scan 4B (CCCL SM100 scan) CCCL=( 22i, 384t,tile= 33792B) muh=( 22i, 384t,tile= 33792B) ✓
|
||||
scan 8B (CCCL SM100 scan) CCCL=( 11i, 416t,tile= 36608B) muh=( 11i, 416t,tile= 36608B) ✓
|
||||
det float32 SM90 CCCL=( 13i, 224t,tile= 11648B) muh=( 13i, 224t,tile= 11648B) ✓
|
||||
det float64 SM86 CCCL=( 5i, 128t,tile= 5120B) muh=( 5i, 128t,tile= 5120B) ✓
|
||||
1-byte type CCCL=( 32i, 256t,tile= 8192B) muh=( 32i, 256t,tile= 8192B) ✓
|
||||
2-byte type CCCL=( 32i, 256t,tile= 16384B) muh=( 32i, 256t,tile= 16384B) ✓
|
||||
16-byte type (int128) CCCL=( 4i, 256t,tile= 16384B) muh=( 4i, 256t,tile= 16384B) ✓
|
||||
SMEM cap test (should trigger) CCCL=( 8i, 768t,tile= 49152B) muh=( 8i, 768t,tile= 49152B) ✓
|
||||
→ scale_mem_bound: FULL PARITY ✓
|
||||
|
||||
### 2. reduce tuning: CCCL SM100值 → BI-V100 scale_mem_bound适配后
|
||||
------------------------------------------------------------
|
||||
CCCL benchmarked on SM100 → muh should use scale_mem_bound for BI-V100
|
||||
Key: reduce loads to REGISTERS not SMEM → SMEM cap rarely triggers
|
||||
|
||||
float32_plus_o4 @4B: scaled=(16i, 512t) tile= 32768B (66.7%)
|
||||
float32_plus_o4 @8B: scaled=( 8i, 512t) tile= 32768B (66.7%)
|
||||
float64_plus_o4 @4B: scaled=(16i, 640t) tile= 40960B (83.3%)
|
||||
float64_plus_o4 @8B: scaled=( 8i, 640t) tile= 40960B (83.3%)
|
||||
accum8_plus_o4 @4B: scaled=(15i, 512t) tile= 30720B (62.5%)
|
||||
accum8_plus_o4 @8B: scaled=( 7i, 512t) tile= 28672B (58.3%)
|
||||
accum8_plus_o8 @4B: scaled=(15i, 512t) tile= 30720B (62.5%)
|
||||
accum8_plus_o8 @8B: scaled=( 7i, 512t) tile= 28672B (58.3%)
|
||||
det_float32_sm90 @4B: scaled=(13i, 224t) tile= 11648B (23.7%)
|
||||
det_float32_sm90 @8B: scaled=( 6i, 224t) tile= 10752B (21.9%)
|
||||
det_float32_sm86 @4B: scaled=( 6i, 224t) tile= 5376B (10.9%)
|
||||
det_float32_sm86 @8B: scaled=( 3i, 224t) tile= 5376B (10.9%)
|
||||
det_float64_sm86 @4B: scaled=(11i, 128t) tile= 5632B (11.5%)
|
||||
det_float64_sm86 @8B: scaled=( 5i, 128t) tile= 5120B (10.4%)
|
||||
default_fallback @4B: scaled=(16i, 256t) tile= 16384B (33.3%)
|
||||
default_fallback @8B: scaled=( 8i, 256t) tile= 16384B (33.3%)
|
||||
|
||||
### 3. muh bi100 reduce当前值 vs CCCL参考
|
||||
------------------------------------------------------------
|
||||
muh改用了更大的items (24 vs SM100的16)来补偿16 SMs
|
||||
这是对的——reduce加载到寄存器,SMEM不是瓶颈
|
||||
|
||||
★ float32 plus (paged_attention score reduction — 83% weight):
|
||||
CCCL SM100: items=16, threads=512, vec=2
|
||||
muh BI-V100: items=24, threads=512, vec=2
|
||||
理由: 16 SMs vs 148 SMs, 每个CTA需要处理更多数据
|
||||
tile对比: SM100=512*16*4=32768B | BI-V100=512*24*4=49152B (exactly 48KB)
|
||||
→ items=24 用满了SMEM → 合理但有风险,如果BlockReduce实际占SMEM则溢出
|
||||
→ 但注释说reduce不用BlockLoad(loads to registers) → 安全
|
||||
|
||||
### 4. scan tuning: CCCL SM100 → BI-V100 SMEM约束
|
||||
------------------------------------------------------------
|
||||
Scan DOES use BlockLoad staging in SMEM → tile_bytes ≤ 49152 is HARD
|
||||
|
||||
lookback_1B_o4 @1B: tpb= 512 ipt=18 tile= 9216B ✓
|
||||
lookback_1B_o4 @2B: tpb= 512 ipt=18 tile= 18432B ✓
|
||||
lookback_1B_o4 @4B: tpb= 512 ipt=18 tile= 36864B ✓
|
||||
lookback_1B_o4 @8B: tpb= 512 ipt=18 tile= 73728B ✗ OVERFLOW → max_items=12
|
||||
lookback_2B_o4 @1B: tpb= 512 ipt=13 tile= 6656B ✓
|
||||
lookback_2B_o4 @2B: tpb= 512 ipt=13 tile= 13312B ✓
|
||||
lookback_2B_o4 @4B: tpb= 512 ipt=13 tile= 26624B ✓
|
||||
lookback_2B_o4 @8B: tpb= 512 ipt=13 tile= 53248B ✗ OVERFLOW → max_items=12
|
||||
lookback_4B_o4 @1B: tpb= 384 ipt=22 tile= 8448B ✓
|
||||
lookback_4B_o4 @2B: tpb= 384 ipt=22 tile= 16896B ✓
|
||||
lookback_4B_o4 @4B: tpb= 384 ipt=22 tile= 33792B ✓
|
||||
lookback_4B_o4 @8B: tpb= 384 ipt=22 tile= 67584B ✗ OVERFLOW → max_items=16
|
||||
lookback_8B_o4 @1B: tpb= 416 ipt=23 tile= 9568B ✓
|
||||
lookback_8B_o4 @2B: tpb= 416 ipt=23 tile= 19136B ✓
|
||||
lookback_8B_o4 @4B: tpb= 416 ipt=23 tile= 38272B ✓
|
||||
lookback_8B_o4 @8B: tpb= 416 ipt=23 tile= 76544B ✗ OVERFLOW → max_items=14
|
||||
lookback_1B_o8 @1B: tpb= 384 ipt=14 tile= 5376B ✓
|
||||
lookback_1B_o8 @2B: tpb= 384 ipt=14 tile= 10752B ✓
|
||||
lookback_1B_o8 @4B: tpb= 384 ipt=14 tile= 21504B ✓
|
||||
lookback_1B_o8 @8B: tpb= 384 ipt=14 tile= 43008B ✓
|
||||
lookback_4B_o8 @1B: tpb= 416 ipt=19 tile= 7904B ✓
|
||||
lookback_4B_o8 @2B: tpb= 416 ipt=19 tile= 15808B ✓
|
||||
lookback_4B_o8 @4B: tpb= 416 ipt=19 tile= 31616B ✓
|
||||
lookback_4B_o8 @8B: tpb= 416 ipt=19 tile= 63232B ✗ OVERFLOW → max_items=14
|
||||
lookback_8B_o8 @1B: tpb= 320 ipt=22 tile= 7040B ✓
|
||||
lookback_8B_o8 @2B: tpb= 320 ipt=22 tile= 14080B ✓
|
||||
lookback_8B_o8 @4B: tpb= 320 ipt=22 tile= 28160B ✓
|
||||
lookback_8B_o8 @8B: tpb= 320 ipt=22 tile= 56320B ✗ OVERFLOW → max_items=19
|
||||
|
||||
关键发现:
|
||||
- scan lookback_4B_o4: items=22, threads=384 → tile@4B=33792 ✓ tile@8B=67584 ✗
|
||||
- scan lookback_8B_o4: items=23, threads=416 → tile@8B=76544 ✗
|
||||
- 这些值在SM100上是安全的(228KB SMEM),但在BI-V100(48KB)上必须降级
|
||||
- muh已经做了降级(用scale_mem_bound),但需要验证降级后的值是否正确
|
||||
|
||||
### 5. CCCL benchmark format解析
|
||||
------------------------------------------------------------
|
||||
NVIDIA的benchmark注释格式:
|
||||
ipt_<items>.tpb_<threads>.ns_<delay>.dcid_<algo>.l2w_<latency>.trp_<transpose>.ld_<load>
|
||||
后跟4个浮点数: 在[2^16, 2^20, 2^24, 2^28]四个problem size下的speedup
|
||||
|
||||
dcid映射:
|
||||
0 = no_delay
|
||||
1 = fixed_delay
|
||||
2 = exp_backoff
|
||||
3 = exp_backoff_jitter
|
||||
4 = exp_backoff_jitter_window
|
||||
5 = exp_backon_jitter_window
|
||||
6 = exp_backon_jitter
|
||||
7 = exp_backon
|
||||
|
||||
### 6. 竞赛关键路径优先级
|
||||
------------------------------------------------------------
|
||||
Token吞吐加权值 = Output_TPS × 16.796 + Input_TPS × 2.799 + Cache_TPS × 0.56
|
||||
→ Output_TPS权重83%, Input_TPS权重14%, Cache_TPS权重3%
|
||||
|
||||
decode热路径 (Output TPS):
|
||||
1. paged_attention score reduction → reduce (DONE: muh tuned)
|
||||
2. softmax denominator prefix-sum → scan (DONE: muh tuned)
|
||||
3. top-k/top-p sampling → topk/radix_sort (DONE: muh tuned)
|
||||
4. RMSNorm/SiLU/RoPE element-wise → transform (DONE: muh tuned)
|
||||
|
||||
prefill热路径 (Input TPS):
|
||||
5. flash_attention → scan + reduce
|
||||
6. MoE expert routing → select_if + reduce_by_key
|
||||
|
||||
cache热路径 (Cache TPS):
|
||||
7. KV cache block copy → batch_memcpy (DONE: muh tuned)
|
||||
|
||||
### 7. 待验证的关键问题
|
||||
------------------------------------------------------------
|
||||
1. reduce items=24: 虽然loads to registers, 但实际BlockReduce<WARP_REDUCTIONS>的SMEM用量需要确认
|
||||
2. scan delay参数: 0.5x/0.6x缩放是启发式, 需要BI-V100实测L2 write latency
|
||||
3. LOAD_LDG vs LOAD_DEFAULT: topk bench显示BI-V100上LOAD_DEFAULT更快, reduce/scan可能同理
|
||||
4. SM count=16 → wave efficiency: 所有tuning都需要重新算occupancy
|
||||
5. transform bytes_in_flight: 从18GB/s改为56GB/s后items需要相应增大
|
||||
50
PIPELINE_GROUND_TRUTH.md
Normal file
50
PIPELINE_GROUND_TRUTH.md
Normal file
@@ -0,0 +1,50 @@
|
||||
# muh Pipeline Ground Truth — 2026-08-07
|
||||
|
||||
## 管道实际状态(不是设计稿,是已部署代码的真实描述)
|
||||
|
||||
### scale_mem_bound: FULL PARITY ✓
|
||||
11/11测试用例与CCCL `cub::detail::scale_mem_bound` 完全匹配。
|
||||
返回值顺序 `{items_per_thread, threads_per_block}` — items-first,与CCCL一致。
|
||||
|
||||
### C++ Tuning Headers: 27/27 ✓
|
||||
所有26个算法(+common)都有bi100 header,`policy_selector::operator()` 接受
|
||||
`hardware_capability` 参数。SMEM overflow保护覆盖所有type_size。
|
||||
|
||||
### Injection现状(enginex没有.cu源码)
|
||||
|
||||
| 注入位置 | 状态 | 值 | commit |
|
||||
|---------|------|-----|--------|
|
||||
| prefix_prefill.py BLOCK | ✓ 已手动修改 | BLOCK=64, WARPS=4 | 多个commit |
|
||||
| paged_attn.py _PARTITION_SIZE | ✓ 保持默认 | 512 | — |
|
||||
| paged_attn.py V1/V2 dispatch | ✓ 已手动修改 | use_v1 threshold | cbd1f08 |
|
||||
| _custom_ops.py SMEM | ✓ 已手动修改 | 48KB | 16f0b30 |
|
||||
| triton_flash_attention.py | ✓ 已添加BI-V100 configs | BLOCK=32/64 | 多个commit |
|
||||
| protocol.py 兼容性 | ✓ 已修复 | max_completion_tokens等 | 2c353da |
|
||||
|
||||
### gen_patch.py 角色
|
||||
设计时期望: C++ header → unified diff → vllm .cu文件
|
||||
实际情况: enginex只有Python + .so, 没有.cu源码
|
||||
当前角色: 文档工具 + 验证(确认header值与已部署Python代码一致)
|
||||
|
||||
### CCCL SM100 Benchmark数据(从源码提取,已存入cccl_sm100_benchmark_values.json)
|
||||
|
||||
**Reduce** (paged_attention score reduction, Output TPS 83%权重):
|
||||
- float32+plus: items=16, threads=512, vec=2, speedup=[1.061, 1.000, 1.065, 1.167]
|
||||
- float64+plus: items=16, threads=640, vec=1, speedup=[1.018, 1.000, 1.016, 1.057]
|
||||
|
||||
**Scan** (softmax prefix-sum):
|
||||
- 4B lookback: items=22, threads=384, delay=1904ns/dcid=6/l2w=830, speedup=[1.148, 0.997, 1.140, 1.463]
|
||||
- 8B lookback: items=23, threads=416, delay=772ns/dcid=5/l2w=710, speedup=[1.089, 1.016, 1.086, 1.265]
|
||||
|
||||
**muh BI-V100适配**:
|
||||
- reduce float32: items=24(+50%), threads=512(=), vec=2(=) → 补偿16 SMs
|
||||
- scan 4B: 通过scale_mem_bound自动适配(items=22 @4B安全, @8B降级到16)
|
||||
- delay参数: ns×0.5, l2w×0.6 (启发式, 待实测)
|
||||
|
||||
### 竞赛门槛
|
||||
- 功能测试: 50+ TC, 项目看板14个FEA item覆盖
|
||||
- 效果测试: benchmark偏差 ≤ ±4%
|
||||
- 性能测试: Token吞吐加权值 ≥ 8000
|
||||
- Output TPS × 16.796 (83%) → reduce/scan/topk
|
||||
- Input TPS × 2.799 (14%) → scan/transform
|
||||
- Cache TPS × 0.56 (3%) → batch_memcpy
|
||||
@@ -16,16 +16,13 @@
|
||||
vllm:
|
||||
model_path: /model
|
||||
served_model_name: llm
|
||||
max_model_len: 256000
|
||||
gpu_memory_utilization: 0.95
|
||||
max_model_len: 100000
|
||||
gpu_memory_utilization: 0.90
|
||||
tensor_parallel: 4
|
||||
max_num_seqs: 2
|
||||
max_num_batched_tokens: 4096
|
||||
max_seq_len_to_capture: 32768
|
||||
max_num_seqs: 1
|
||||
trust_remote_code: true
|
||||
disable_log_requests: true
|
||||
disable_frontend_multiprocessing: true
|
||||
enable_chunked_prefill: true
|
||||
enable_auto_tool_choice: true
|
||||
tool_call_parser: qwen3_coder
|
||||
reasoning_parser: qwen3
|
||||
|
||||
278
cccl_sm100_benchmark_values.json
Normal file
278
cccl_sm100_benchmark_values.json
Normal file
@@ -0,0 +1,278 @@
|
||||
{
|
||||
"source": "cccl_upstream/cub/cub/device/dispatch/tuning/tuning_*.cuh",
|
||||
"extracted_by": "automated audit from CCCL source code",
|
||||
"reduce": {
|
||||
"sm100_float32_plus_o4": {
|
||||
"items": 16,
|
||||
"threads": 512,
|
||||
"vec": 2,
|
||||
"benchmark": "ipt_16.tpb_512.ipv_2",
|
||||
"speedup": [
|
||||
1.061295,
|
||||
1.0,
|
||||
1.065478,
|
||||
1.167139
|
||||
]
|
||||
},
|
||||
"sm100_float64_plus_o4": {
|
||||
"items": 16,
|
||||
"threads": 640,
|
||||
"vec": 1,
|
||||
"benchmark": "ipt_16.tpb_640.ipv_1",
|
||||
"speedup": [
|
||||
1.017834,
|
||||
1.0,
|
||||
1.015835,
|
||||
1.057092
|
||||
]
|
||||
},
|
||||
"sm100_accum8_plus_o4": {
|
||||
"items": 15,
|
||||
"threads": 512,
|
||||
"vec": 2,
|
||||
"benchmark": "ipt_15.tpb_512.ipv_2",
|
||||
"speedup": [
|
||||
1.019887,
|
||||
1.0,
|
||||
1.017636,
|
||||
1.058036
|
||||
]
|
||||
},
|
||||
"sm100_accum8_plus_o8": {
|
||||
"items": 15,
|
||||
"threads": 512,
|
||||
"vec": 1,
|
||||
"benchmark": "ipt_15.tpb_512.ipv_1",
|
||||
"speedup": [
|
||||
1.019414,
|
||||
1.0,
|
||||
1.017218,
|
||||
1.057143
|
||||
]
|
||||
},
|
||||
"sm90_det_float32": {
|
||||
"items": 13,
|
||||
"threads": 224,
|
||||
"benchmark": "ipt_13.tpb_224",
|
||||
"speedup": [
|
||||
1.107188,
|
||||
1.009709,
|
||||
1.097114,
|
||||
1.31682
|
||||
]
|
||||
},
|
||||
"sm86_det_float32": {
|
||||
"items": 6,
|
||||
"threads": 224,
|
||||
"benchmark": "ipt_6.tpb_224",
|
||||
"speedup": [
|
||||
1.034383,
|
||||
1.0,
|
||||
1.032097,
|
||||
1.090909
|
||||
]
|
||||
},
|
||||
"sm86_det_float64": {
|
||||
"items": 11,
|
||||
"threads": 128,
|
||||
"benchmark": "ipt_11.tpb_128",
|
||||
"speedup": [
|
||||
1.232089,
|
||||
1.002124,
|
||||
1.245336,
|
||||
1.582279
|
||||
]
|
||||
}
|
||||
},
|
||||
"scan": {
|
||||
"sm100_lookback_1B_o4": {
|
||||
"items": 18,
|
||||
"threads": 512,
|
||||
"delay": {
|
||||
"ns": 768,
|
||||
"dcid": 7,
|
||||
"l2w": 820
|
||||
},
|
||||
"load": {
|
||||
"transpose": 1,
|
||||
"modifier": 0
|
||||
},
|
||||
"benchmark": "ipt_18.tpb_512.ns_768.dcid_7.l2w_820.trp_1.ld_0",
|
||||
"speedup": [
|
||||
1.188818,
|
||||
1.005682,
|
||||
1.173041,
|
||||
1.305288
|
||||
]
|
||||
},
|
||||
"sm100_lookback_2B_o4": {
|
||||
"items": 13,
|
||||
"threads": 512,
|
||||
"delay": {
|
||||
"ns": 1384,
|
||||
"dcid": 7,
|
||||
"l2w": 720
|
||||
},
|
||||
"load": {
|
||||
"transpose": 1,
|
||||
"modifier": 0
|
||||
},
|
||||
"benchmark": "ipt_13.tpb_512.ns_1384.dcid_7.l2w_720.trp_1.ld_0",
|
||||
"speedup": [
|
||||
1.128443,
|
||||
1.002841,
|
||||
1.119688,
|
||||
1.307692
|
||||
]
|
||||
},
|
||||
"sm100_lookback_4B_o4": {
|
||||
"items": 22,
|
||||
"threads": 384,
|
||||
"delay": {
|
||||
"ns": 1904,
|
||||
"dcid": 6,
|
||||
"l2w": 830
|
||||
},
|
||||
"load": {
|
||||
"transpose": 1,
|
||||
"modifier": 0
|
||||
},
|
||||
"benchmark": "ipt_22.tpb_384.ns_1904.dcid_6.l2w_830.trp_1.ld_0",
|
||||
"speedup": [
|
||||
1.148442,
|
||||
0.997167,
|
||||
1.139902,
|
||||
1.462651
|
||||
]
|
||||
},
|
||||
"sm100_lookback_8B_o4": {
|
||||
"items": 23,
|
||||
"threads": 416,
|
||||
"delay": {
|
||||
"ns": 772,
|
||||
"dcid": 5,
|
||||
"l2w": 710
|
||||
},
|
||||
"load": {
|
||||
"transpose": 1,
|
||||
"modifier": 0
|
||||
},
|
||||
"benchmark": "ipt_23.tpb_416.ns_772.dcid_5.l2w_710.trp_1.ld_0",
|
||||
"speedup": [
|
||||
1.089468,
|
||||
1.015581,
|
||||
1.08563,
|
||||
1.264583
|
||||
]
|
||||
},
|
||||
"sm100_lookback_1B_o8": {
|
||||
"items": 14,
|
||||
"threads": 384,
|
||||
"delay": {
|
||||
"ns": 228,
|
||||
"dcid": 7,
|
||||
"l2w": 775
|
||||
},
|
||||
"load": {
|
||||
"transpose": 1,
|
||||
"modifier": 1
|
||||
},
|
||||
"benchmark": "ipt_14.tpb_384.ns_228.dcid_7.l2w_775.trp_1.ld_1",
|
||||
"speedup": [
|
||||
1.10721,
|
||||
1.0,
|
||||
1.100637,
|
||||
1.307692
|
||||
]
|
||||
},
|
||||
"sm100_lookback_4B_o8": {
|
||||
"items": 19,
|
||||
"threads": 416,
|
||||
"delay": {
|
||||
"ns": 956,
|
||||
"dcid": 7,
|
||||
"l2w": 550
|
||||
},
|
||||
"load": {
|
||||
"transpose": 1,
|
||||
"modifier": 1
|
||||
},
|
||||
"benchmark": "ipt_19.tpb_416.ns_956.dcid_7.l2w_550.trp_1.ld_1",
|
||||
"speedup": [
|
||||
1.146142,
|
||||
0.99435,
|
||||
1.137459,
|
||||
1.455636
|
||||
]
|
||||
},
|
||||
"sm100_lookback_8B_o8": {
|
||||
"items": 22,
|
||||
"threads": 320,
|
||||
"delay": {
|
||||
"ns": 328,
|
||||
"dcid": 2,
|
||||
"l2w": 965
|
||||
},
|
||||
"load": {
|
||||
"transpose": 1,
|
||||
"modifier": 0
|
||||
},
|
||||
"benchmark": "ipt_22.tpb_320.ns_328.dcid_2.l2w_965.trp_1.ld_0",
|
||||
"speedup": [
|
||||
1.080133,
|
||||
1.0,
|
||||
1.075577,
|
||||
1.248963
|
||||
]
|
||||
}
|
||||
},
|
||||
"benchmark_runner_params": {
|
||||
"reduce": {
|
||||
"items_range": "7:24:1",
|
||||
"threads_range": "128:1024:32",
|
||||
"vec_pow2_range": "1:2:1",
|
||||
"problem_sizes": [
|
||||
"2^16",
|
||||
"2^20",
|
||||
"2^24",
|
||||
"2^28"
|
||||
]
|
||||
},
|
||||
"scan_lookback": {
|
||||
"items_range": "7:24:1",
|
||||
"threads_range": "128:1024:32",
|
||||
"delay_ns_range": "0:2048:4",
|
||||
"delay_algo_range": "0:7:1",
|
||||
"l2w_range": "0:1200:5",
|
||||
"transpose_range": "0:1:1",
|
||||
"load_range": "0:1:1",
|
||||
"problem_sizes": [
|
||||
"2^16",
|
||||
"2^20",
|
||||
"2^24",
|
||||
"2^28",
|
||||
"2^32"
|
||||
]
|
||||
},
|
||||
"topk": {
|
||||
"items_range": "7:24:1",
|
||||
"threads_range": "128:1024:32",
|
||||
"load_algo_range": "0:2:1"
|
||||
},
|
||||
"radix_sort": {
|
||||
"items_range": "7:24:1",
|
||||
"threads_range": "128:1024:32",
|
||||
"radix_bits": 8
|
||||
}
|
||||
},
|
||||
"dcid_mapping": {
|
||||
"0": "no_delay",
|
||||
"1": "fixed_delay",
|
||||
"2": "exponential_backoff",
|
||||
"3": "exponential_backoff_jitter",
|
||||
"4": "exponential_backoff_jitter_window",
|
||||
"5": "exponential_backon_jitter_window",
|
||||
"6": "exponential_backon_jitter",
|
||||
"7": "exponential_backon"
|
||||
}
|
||||
}
|
||||
@@ -8,21 +8,16 @@ command:
|
||||
- --served-model-name
|
||||
- llm
|
||||
- --max-model-len
|
||||
- '131072'
|
||||
- '100000'
|
||||
- --gpu-memory-utilization
|
||||
- '0.90'
|
||||
- --trust-remote-code
|
||||
- -tp
|
||||
- '4'
|
||||
- --max-num-seqs
|
||||
- '2'
|
||||
- '1'
|
||||
- --disable-log-requests
|
||||
- --disable-frontend-multiprocessing
|
||||
- --max-num-batched-tokens
|
||||
- '8192'
|
||||
- --enable-chunked-prefill
|
||||
- --max-seq-len-to-capture
|
||||
- '32768'
|
||||
- --enforce-eager
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
|
||||
@@ -57,7 +57,10 @@ class CustomChatCompletionMessageParam(TypedDict, total=False):
|
||||
|
||||
class OpenAIBaseModel(BaseModel):
|
||||
# OpenAI API does not allow extra fields
|
||||
model_config = ConfigDict(extra="forbid")
|
||||
# Real-world clients (replay, third-party SDKs) may send extra fields
|
||||
# like service_tier, store, metadata, reasoning_effort, etc.
|
||||
# "ignore" accepts the request and silently drops unknown fields.
|
||||
model_config = ConfigDict(extra="ignore")
|
||||
|
||||
|
||||
class ErrorResponse(OpenAIBaseModel):
|
||||
@@ -418,13 +421,30 @@ class ChatCompletionRequest(OpenAIBaseModel):
|
||||
# The competition evaluator sends thinking={enable:true/false} (OpenAI API).
|
||||
# Qwen3's chat template expects enable_thinking=True/False in kwargs.
|
||||
thinking = data.get("thinking")
|
||||
thinking_explicitly_set = False
|
||||
if isinstance(thinking, dict):
|
||||
enable = thinking.get("enable")
|
||||
if enable is not None:
|
||||
thinking_explicitly_set = True
|
||||
ctk = data.get("chat_template_kwargs") or {}
|
||||
ctk["enable_thinking"] = bool(enable)
|
||||
data["chat_template_kwargs"] = ctk
|
||||
|
||||
# CRITICAL: When tools are present with tool_choice=auto and thinking
|
||||
# is NOT explicitly requested, disable thinking to preserve token budget
|
||||
# for tool call XML generation. Without this, the model spends all
|
||||
# tokens on <think>...</think> and finishes before emitting <tool_call>.
|
||||
# This matches the competition reference (sub168: d03 in 2.12s).
|
||||
if not thinking_explicitly_set:
|
||||
has_tools = data.get("tools") is not None and len(data.get("tools", [])) > 0
|
||||
tc = data.get("tool_choice")
|
||||
tool_choice_active = (tc == "auto" or (tc is None and has_tools)
|
||||
or isinstance(tc, dict))
|
||||
if has_tools and tool_choice_active:
|
||||
ctk = data.get("chat_template_kwargs") or {}
|
||||
ctk["enable_thinking"] = False
|
||||
data["chat_template_kwargs"] = ctk
|
||||
|
||||
messages = data.get("messages")
|
||||
if not isinstance(messages, list):
|
||||
return data
|
||||
@@ -517,6 +537,12 @@ class ChatCompletionRequest(OpenAIBaseModel):
|
||||
# if "tool_choice" is specified -- validation
|
||||
if "tool_choice" in data:
|
||||
|
||||
# "none" means don't use any tools — valid per OpenAI spec,
|
||||
# just strip tool_choice and let vLLM ignore tools.
|
||||
if data["tool_choice"] == "none":
|
||||
del data["tool_choice"]
|
||||
return data
|
||||
|
||||
# ensure that if "tool choice" is specified, tools are present
|
||||
if "tools" not in data or data["tools"] is None:
|
||||
raise ValueError(
|
||||
|
||||
@@ -77,6 +77,28 @@ class Qwen3CoderToolParser(ToolParser):
|
||||
logger.debug("vLLM Successfully imported tool parser %s !",
|
||||
self.__class__.__name__)
|
||||
|
||||
def adjust_request(
|
||||
self, request: "ChatCompletionRequest") -> "ChatCompletionRequest":
|
||||
"""Disable thinking when tools are active with auto choice.
|
||||
|
||||
On BI-V100 hardware, the model's <think>...</think> phase can consume
|
||||
the entire max_tokens budget, leaving no room for the <tool_call> XML.
|
||||
Competition reference (sub168) completes d03_tool_call in 2.12s with
|
||||
tools=1; our sub509 took 49s with tools=0 because thinking ate the
|
||||
budget. Disabling thinking for tool-call requests ensures the model
|
||||
emits tool XML within the token budget.
|
||||
"""
|
||||
if (request.tools and request.tool_choice in ("auto", None)
|
||||
and not isinstance(request.tool_choice,
|
||||
type(None).__class__)):
|
||||
# Only override if thinking was not explicitly requested
|
||||
ctk = request.chat_template_kwargs or {}
|
||||
if "enable_thinking" not in ctk:
|
||||
ctk = dict(ctk) # shallow copy
|
||||
ctk["enable_thinking"] = False
|
||||
request.chat_template_kwargs = ctk
|
||||
return request
|
||||
|
||||
|
||||
def _generate_tool_call_id(self) -> str:
|
||||
return f"call_{uuid.uuid4().hex[:24]}"
|
||||
|
||||
@@ -348,7 +348,7 @@ class OpenAIServingChat(OpenAIServing):
|
||||
# parsing and reasoning parsing (both require full-history context).
|
||||
if tool_choice_auto or use_reasoning:
|
||||
previous_texts = [""] * num_choices
|
||||
all_previous_token_ids = [[]] * num_choices
|
||||
all_previous_token_ids = [[] for _ in range(num_choices)]
|
||||
else:
|
||||
previous_texts, all_previous_token_ids = None, None
|
||||
|
||||
@@ -357,7 +357,8 @@ class OpenAIServingChat(OpenAIServing):
|
||||
if tool_choice_auto and self.tool_parser:
|
||||
tool_parsers: List[Optional[ToolParser]] = [
|
||||
self.tool_parser(tokenizer)
|
||||
] * num_choices
|
||||
for _ in range(num_choices)
|
||||
]
|
||||
else:
|
||||
tool_parsers = [None] * num_choices
|
||||
except RuntimeError as e:
|
||||
|
||||
Reference in New Issue
Block a user