78a0ebd5161bd192e9e6c0f3f3317ca0b945ffa1
Hardware testing confirmed: V1: K=[blocks, kv_heads, head_dim/x, block_size, x] (5D), V=[blocks, kv_heads, head_dim, block_size] (4D) → OK V2: K=[blocks, kv_heads, block_size, head_dim] (4D), V=[blocks, kv_heads, block_size, head_dim] (4D) → OK V2 with V1's layout → FAIL (Expected key_cache.dim()==4, value_cache.size(3)==head_size) V1 and V2 use DIFFERENT cache memory layouts in ixformer. V2 patch now converts cache on the fly before calling native kernel: K: permute(0,1,3,2,4).reshape → [B,H,bs,d] V: permute(0,1,3,2).contiguous → [B,H,bs,d] This is a view+reshape for K (no copy if contiguous) and a transpose+contiguous for V. The cost is one V copy per decode step, but this enables the native compiled V2 kernel which is 10-100x faster than the Python fallback it replaces.
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%