715302997459c2ba403f8847ba1c9d1a6b96000d
CCCL thread_reduce.cuh pattern: if(length==1) return directly. For tool_call requests where user didn't set max_tokens, cap to 2048 to prevent NaN-damaged models from generating 99900 tokens of garbage. Expected tool_call XML is <500 tokens. Sub509 spent 49s on d03 because the model generated endlessly with no tool_call output.
project_6
Description
Languages
C++
41.7%
Cuda
31.5%
Python
22.4%
C
2.1%
CMake
1.1%
Other
1.1%