bed1fc4d544c0b18a5d5c9fbfa5b157d6b21748d
flash_attn_varlen_func allocates O(n²) temp memory for 4096 dummy tokens during profile_run, causing OOM at gpu_memory_utilization=0.80. BI100_IN_STARTUP_PROFILE=1 env var is already set by patch_worker_startup_profile_guard.py during the synthetic forward pass. Real inference requests still use flash_attn_varlen (much faster).
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%