This website requires JavaScript.
Explore
Help
Register
Sign In
dylanyunlong
/
project_6
Watch
1
Star
0
Fork
0
You've already forked project_6
Code
Issues
Pull Requests
Actions
Projects
Releases
Wiki
Activity
Files
b922d694dc89fa431de8167e56103dcf8e27fca2
project_6
/
ex_engine
/
xllm_kernels
History
Claude
395b3e4042
test: clean rebuild + debug output for kernel 10 correctness
2026-08-15 05:14:18 +00:00
..
cuda
perf: hgemm_warptiling Config B — beats cublas on MoE-sized GEMM (0.7x)
2026-08-15 05:11:32 +00:00
ilu
feat: import CUDA kernels from xllm/CCCL/FLA upstream repos
2026-08-14 07:48:52 +00:00
npu
fix(critical): fold max_completion_tokens + max_num_seqs=2 + max_model_len=80000 + xllm_latest layer import
2026-08-13 03:19:39 +00:00
build_test_hgemm_warp.sh
feat: hgemm_warptiling.cu — siboehm kernel 10 ported to WARPSIZE=64 FP16
2026-08-14 17:05:59 +00:00
build_test_hgemm.sh
feat: hgemm_blocktiling.cu — FP16 GEMM kernel for MoE expert dispatch on BI-V100
2026-08-14 16:22:00 +00:00
rebuild_test_k10.sh
test: clean rebuild + debug output for kernel 10 correctness
2026-08-15 05:14:18 +00:00