44bdf49cae8342013a5791486ddb46f2dc33a9b9
Source: cccl_upstream/cub/benchmarks/bench/partition/flagged.cu (random pick) CCCL partition benchmark shows DevicePartition::Flagged uses lookback scan with tunable ipt/tpb/ns/dcid/l2w — same architecture as top-k radix select. Key insight: radix select is O(N × bits_per_pass) vs full sort O(N log N). For Qwen3.6 vocab_size=152064: topk: ~11 radix passes sort: ~17 comparison-based passes = 1.5x more kernel cycles Applied: _apply_top_k_top_p fast path when all sequences have top_p=1.0 - Skips: sort(152K) + softmax + cumsum + scatter - Uses: torch.topk (radix select internally) + threshold mask - This was already in vllm/sampler.py but NEVER DEPLOYED to base image Also adds sampler.py to patch_ops.sh cp list for Docker deployment.
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%