336f3349ca9be995d74a542faa0c259c6684717c
38eca5c2 improvements
Keeps ALL infrastructure from the last 70 commits: - paged_attn.py: ixformer native v1/v2 decode dispatch (Output TPS impact) - protocol.py: extra='allow' (fixes ~180 rejected replay requests) - qwen3_5.py: .float() router_logits, chunk_recurrent, index_combine - 14 prebuilt .so (including corex_gdn_chunk_recurrent) - patch_xformers: flash_attn_varlen_func + profiling guard (>32K→Q-tiling) yaml: max-num-seqs=2, TOPK=1, gpu-mem=0.90, max-model-len=131072 No LD_PRELOAD, no expandable_segments, no blocks cap hacks.
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%