dylanyunlon
1c9ac93fee
[ENGINE+TEST] 2 changes from CCCL random source reading
1. model_runner.py: CCCL CachingDeviceAllocator (example_device_radix_sort.cu)
→ CUDA graph capture 1028→19 sizes, saves ~50GB memory + 50s startup
2. verify_functional.py: TC-22→TC-30 from CCCL dispatch_segmented_reduce.cuh
- TC-22/23: Unicode fidelity (Chinese/Japanese exact repeat)
- TC-24: n=2 multiple choices (segmented output)
- TC-25/26: Error handling (empty body, missing role)
- TC-27/28: Sampling boundary (top_k=1, temperature=2.0)
- TC-29/30: Endpoint health (/v1/models, /health)
Total: 21→30 test cases (target: 50+ for competition)
CCCL sources read this round:
cub/examples/device/example_device_radix_sort.cu → DoubleBuffer + CachingDeviceAllocator
cudax/test/multi_gpu/concepts/has_gather_v.cu → TP gather pattern
cub/cub/device/dispatch/dispatch_segmented_reduce.cuh → 3-tier policy (large/medium/small)
2026-08-07 02:01:06 +00:00
..
2026-07-30 16:06:20 +00:00
2026-07-30 16:06:20 +00:00
2026-07-30 16:06:20 +00:00
2026-08-06 07:01:14 +00:00
2026-07-30 16:06:20 +00:00
2026-08-05 08:24:43 +00:00
2026-07-30 16:06:20 +00:00
2026-07-30 16:06:20 +00:00
2026-08-05 08:24:43 +00:00
2026-08-06 06:33:26 +00:00
2026-08-07 01:59:18 +00:00
2026-08-06 07:01:14 +00:00
2026-08-06 06:10:56 +00:00
2026-08-06 07:01:14 +00:00
2026-07-30 16:06:20 +00:00
2026-08-06 06:10:56 +00:00
2026-07-30 16:06:20 +00:00
2026-08-06 03:02:29 +00:00
2026-07-30 16:06:20 +00:00
2026-08-05 08:36:52 +00:00
2026-08-06 02:55:51 +00:00
2026-07-30 16:06:20 +00:00
2026-07-30 16:06:20 +00:00
2026-07-30 16:06:20 +00:00
2026-08-05 08:36:52 +00:00
2026-08-07 02:01:06 +00:00
2026-08-07 01:54:52 +00:00