Source: cccl_upstream/cub/cub/detail/temporary_storage.cuh
Target: vllm/worker/cache_engine.py
CCCL system design applied:
- temporary_storage::layout<SlotsCount>: Phase 1 get_size() computes
total bytes, Phase 2 map_to_buffer() allocates one blob and aliases
into per-slot views
- Applied to _allocate_kv_cache: compute total numel for all layers,
allocate one contiguous torch.zeros, slice into per-layer views
- Reduces cudaMalloc calls from num_attention_layers to 1
- Guarantees cross-layer memory contiguity (better L2 locality)
- slot.create_alias<T>() → layer_flat.view(kv_cache_shape)
POTENTIAL BUG FIX in BASE file:
vllm/worker/model_runner.py line ~833
_get_cuda_graph_pad_size was called with:
max_decode_seq_len=max_encoder_seq_len (WRONG)
should be:
max_decode_seq_len=max_decode_seq_len (FIXED)
For decoder-only Qwen3.6, max_encoder_seq_len=0 always.
This means CUDA graph capture check always saw max_decode_seq_len=0,
potentially causing incorrect graph capture for long decode sequences
(100K context > max_seq_len_to_capture=32768 should DISABLE graph,
but with the bug it would see 0 ≤ 32768 and ENABLE graph incorrectly).
CCCL insight from thrust/examples/bounding_box.cu:
bbox compound reduce tracks lower_left.x/y and upper_right.x/y
as INDEPENDENT dimensions. Mixing them (like setting min_y = max_x)
would produce an incorrect bounding box. Same principle applies to
max_decode_seq_len vs max_encoder_seq_len.
CCCL file: thrust/examples/bounding_box.cu