POTENTIAL BUG FIX in BASE file:
vllm/worker/model_runner.py line ~833
_get_cuda_graph_pad_size was called with:
max_decode_seq_len=max_encoder_seq_len (WRONG)
should be:
max_decode_seq_len=max_decode_seq_len (FIXED)
For decoder-only Qwen3.6, max_encoder_seq_len=0 always.
This means CUDA graph capture check always saw max_decode_seq_len=0,
potentially causing incorrect graph capture for long decode sequences
(100K context > max_seq_len_to_capture=32768 should DISABLE graph,
but with the bug it would see 0 ≤ 32768 and ENABLE graph incorrectly).
CCCL insight from thrust/examples/bounding_box.cu:
bbox compound reduce tracks lower_left.x/y and upper_right.x/y
as INDEPENDENT dimensions. Mixing them (like setting min_y = max_x)
would produce an incorrect bounding box. Same principle applies to
max_decode_seq_len vs max_encoder_seq_len.
CCCL file: thrust/examples/bounding_box.cu