a8b16da5dae794b8e1bdfb1f0aee4da7d25c5397
Root cause of Sub 520 output_tps=2.6 (vs Sub 168 output_tps=11.9): - patch_xformers_sdpa_seq.py replaces ixformer flash attention with pure PyTorch O(L^2) matmul+softmax serial implementation - 32 full attention layers x every token = 4.6x slower Sub 168 (base image) proof: - output_tps_avg=11.9, output_tps_p50=13.0, output_tps_p90=18.1 - XFormers backend used WITHOUT any patches - ixformer flash_attn works correctly on BI-V100 This commit: skip xformers patches in patch_ops.sh Expected: output_tps should recover to ~11.9 (Sub 168 level)
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%