045ea5df792574ea64bb81cd8cf239d6613c86c9
From corex-samples batched_gemm.cu, changed: float → half_t, OpClassSimt → OpClassTensorOp, Sm61 → Cu10 Uses __ivcorex_matrix_mad_f32x4_f16x4 via mma_cu10.h Default config: TB<128,128,32> Warp<32,32,32> Inst<16,16,16> Standalone test: correctness + perf for MoE decode (8 × 1x4096@4096x11008)
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%