From corex-samples batched_gemm.cu, changed: float → half_t, OpClassSimt → OpClassTensorOp, Sm61 → Cu10 Uses __ivcorex_matrix_mad_f32x4_f16x4 via mma_cu10.h Default config: TB<128,128,32> Warp<32,32,32> Inst<16,16,16> Standalone test: correctness + perf for MoE decode (8 × 1x4096@4096x11008)