flash_attn_func WORKS with head_dim=256 on BI-V100! This is the path to 10-50x attention speedup. Tests: correctness vs ref, GQA, long seq, varlen, paged decode, perf.