Source: cccl_upstream/cub/cub/detail/temporary_storage.cuh
Target: vllm/worker/cache_engine.py
CCCL system design applied:
- temporary_storage::layout<SlotsCount>: Phase 1 get_size() computes
total bytes, Phase 2 map_to_buffer() allocates one blob and aliases
into per-slot views
- Applied to _allocate_kv_cache: compute total numel for all layers,
allocate one contiguous torch.zeros, slice into per-layer views
- Reduces cudaMalloc calls from num_attention_layers to 1
- Guarantees cross-layer memory contiguity (better L2 locality)
- slot.create_alias<T>() → layer_flat.view(kv_cache_shape)