fix(CRITICAL): remove xformers patches — Sub168 proves base ixformer attention works at 11.9 TPS, our patches reduced to 2.6 TPS

Root cause of Sub 520 output_tps=2.6 (vs Sub 168 output_tps=11.9):
- patch_xformers_sdpa_seq.py replaces ixformer flash attention with
  pure PyTorch O(L^2) matmul+softmax serial implementation
- 32 full attention layers x every token = 4.6x slower

Sub 168 (base image) proof:
- output_tps_avg=11.9, output_tps_p50=13.0, output_tps_p90=18.1
- XFormers backend used WITHOUT any patches
- ixformer flash_attn works correctly on BI-V100

This commit: skip xformers patches in patch_ops.sh
Expected: output_tps should recover to ~11.9 (Sub 168 level)
This commit is contained in:
Claude
2026-08-11 02:31:29 +00:00
parent 0478628f17
commit a8b16da5da

View File

@@ -164,12 +164,13 @@ if [ -d "$_SITE" ]; then
fi
# ===========================================================
# 5. XFormers patches — head_dim=256 bypass for BI-V100
# Comp 168 also had xformers patches (base uses xformers for attention)
# 5. XFormers patches — DISABLED
# Sub 168 (base image) achieved output_tps=11.9 WITHOUT any xformers patches.
# Sub 520 applied these patches → output_tps=2.6 (4.6x slower!)
# The patches replace ixformer flash attention with pure PyTorch O(L²) matmul.
# Base image ixformer attention works correctly — proven by Sub 168.
# ===========================================================
python3 ./patch_xformers_sdpa_seq.py 2>&1 || true
python3 ./patch_xformers_sdpa_batch.py 2>&1 || true
echo "[patch_ops] xformers patches applied"
echo "[patch_ops] xformers patches SKIPPED — base ixformer attention works (Sub 168 proof)"
# ===========================================================
# 6. model_runner patch (prefix_cache_hit fix)