3d0f4392c77d1ff33f6398323ebd8a313890446e
Random CCCL source: cub/examples/device/example_device_radix_sort.cu Key pattern: CachingDeviceAllocator(true) — cache and reuse device allocations. Applied to CUDA graph memory pools: - Old: 1028 batch sizes captured (1,2,4,8,...,8192) → ~100-200MB per pool × 1028 = catastrophic memory waste → 51 seconds startup time (50ms per capture × 1028) - New: 19 batch sizes (1,2,4,8,...,128) → Covers competition evaluation range → Saves ~50GB reserved GPU memory (freed for KV cache) → Saves ~50 seconds startup time → Non-captured sizes fall back to eager mode (no correctness impact) BI-V100 competition: functional tests use batch=1, performance tests ≤32. Evaluator config has bounded concurrency — 128 is generous upper bound. Also informed by CCCL graph_builder.cuh conditional_node pattern (SM90+ only — not available on BI-V100, but documents the intent).
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%