1c9ac93feef069014f343a666ea55be2d11d4d46
1. model_runner.py: CCCL CachingDeviceAllocator (example_device_radix_sort.cu) → CUDA graph capture 1028→19 sizes, saves ~50GB memory + 50s startup 2. verify_functional.py: TC-22→TC-30 from CCCL dispatch_segmented_reduce.cuh - TC-22/23: Unicode fidelity (Chinese/Japanese exact repeat) - TC-24: n=2 multiple choices (segmented output) - TC-25/26: Error handling (empty body, missing role) - TC-27/28: Sampling boundary (top_k=1, temperature=2.0) - TC-29/30: Endpoint health (/v1/models, /health) Total: 21→30 test cases (target: 50+ for competition) CCCL sources read this round: cub/examples/device/example_device_radix_sort.cu → DoubleBuffer + CachingDeviceAllocator cudax/test/multi_gpu/concepts/has_gather_v.cu → TP gather pattern cub/cub/device/dispatch/dispatch_segmented_reduce.cuh → 3-tier policy (large/medium/small)
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%