9f02200edecd1246190c6585159a823433b74e1c
flash_attn_func WORKS with head_dim=256 on BI-V100! This is the path to 10-50x attention speedup. Tests: correctness vs ref, GQA, long seq, varlen, paged decode, perf.
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%