Files
project_6/qwen3_6_scripts
dylanyunlon 1c9ac93fee [ENGINE+TEST] 2 changes from CCCL random source reading
1. model_runner.py: CCCL CachingDeviceAllocator (example_device_radix_sort.cu)
   → CUDA graph capture 1028→19 sizes, saves ~50GB memory + 50s startup

2. verify_functional.py: TC-22→TC-30 from CCCL dispatch_segmented_reduce.cuh
   - TC-22/23: Unicode fidelity (Chinese/Japanese exact repeat)
   - TC-24: n=2 multiple choices (segmented output)
   - TC-25/26: Error handling (empty body, missing role)
   - TC-27/28: Sampling boundary (top_k=1, temperature=2.0)
   - TC-29/30: Endpoint health (/v1/models, /health)
   Total: 21→30 test cases (target: 50+ for competition)

CCCL sources read this round:
  cub/examples/device/example_device_radix_sort.cu → DoubleBuffer + CachingDeviceAllocator
  cudax/test/multi_gpu/concepts/has_gather_v.cu → TP gather pattern
  cub/cub/device/dispatch/dispatch_segmented_reduce.cuh → 3-tier policy (large/medium/small)
2026-08-07 02:01:06 +00:00
..