451bdc8204b9b91bedf352a346bfd9c06684b69b
CCCL CachingDeviceAllocator reserves some memory for its bin cache. With 0.90 utilization + max-model-len=80000, profiling stage OOMs. 0.85 leaves ~1.6GB headroom per GPU for profiling + allocator cache.
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%