sglang

Author	SHA1	Message	Date
Yineng Zhang	de15d1405a	Revert "Fix flashinfer version in sgl-kernel (#10135 )" (#10310 )	2025-09-11 01:27:58 -07:00
Yi Zhang	8cbe1538ef	Add mamba kernel (#10234 )	2025-09-09 12:58:43 -07:00
Yineng Zhang	94fb4e9e54	feat: support fa cute in sgl-kernel (#10205 ) Co-authored-by: cicirori <32845984+cicirori@users.noreply.github.com>	2025-09-09 00:14:39 -07:00
fzyzcjy	0096798ed6	[1/2] Speed up prefill mla attention (#10156 )	2025-09-08 09:00:33 -07:00
Rain Jiang	6049ca209e	move compile threads to an option to avoid OOM on low memory host (#10123 )	2025-09-07 21:36:14 -07:00
Lianmin Zheng	76a2c86b88	Fix flashinfer version in sgl-kernel (#10135 )	2025-09-07 12:54:07 -07:00
hlu1	5f1eb20484	[chore] Remove unused ep_moe cuda kernels (#9956 )	2025-09-06 01:35:50 -07:00
hlu1	039cef76aa	Remove non-accelerated targets(100 and up) from cmake (#10041 ) Signed-off-by: Hao Lu <14827759+hlu1@users.noreply.github.com>	2025-09-06 01:35:28 -07:00
fzyzcjy	bd7f882142	Support copying tensor from cpu to gpu without using copy engines (#10007 )	2025-09-05 20:07:19 +08:00
Lianmin Zheng	d631290e32	Remove annoying warnings in sgl kernel build (#9905 )	2025-09-02 20:18:25 -07:00
PGFLMG	7fe89f7cdb	[sgl-kernel] fix: fix missing FetchContent_Populate for fmt (#9826 )	2025-08-30 12:57:42 -07:00
Rain Jiang	6b39f9cf8c	Support compile sgl-kernel on cuda 13.0 (#9721 )	2025-08-28 10:18:03 -07:00
PGFLMG	aa3eba8eb4	[sgl-kernel] misc: update deepgemm version for sgl-kernel (#9340 ) Co-authored-by: Yineng Zhang <me@zhyncs.com> Co-authored-by: fzyzcjy <ch271828n@outlook.com>	2025-08-27 12:01:30 -07:00
Rain Jiang	79e6a8a6ac	support cuda 13.0 and trtllm kernel by Aug 25 2025 (#9495 )	2025-08-26 23:13:27 -07:00
Qi Yuhang	fda4792620	Update CUTLASS 4.2 & Enable K-Major Scale Factor for SM90 FP8 Blockwise Group GEMM (#9559 )	2025-08-24 23:24:43 -07:00
EduardDurech	720cd308ba	Add `CMakeLists.txt` binary_dir (#7019 )	2025-08-18 18:36:33 -07:00
Lianmin Zheng	c480a3f6ea	Minor style fixes for sgl-kernel (#9289 )	2025-08-18 09:38:35 -07:00
Liangsheng Yin	4d98e48649	Revert "[Misc] feat: Deepgemm update for sgl-kernel (#8790 )" to fix kernel CI (#9260 )	2025-08-17 22:59:50 +08:00
Liangsheng Yin	0c8594e67d	Optional extension for green context (#9231 )	2025-08-15 21:33:52 +08:00
PGFLMG	a3d99d6dcd	[Misc] feat: Deepgemm update for sgl-kernel (#8790 )	2025-08-15 01:05:27 -07:00
Yineng Zhang	9d54c6e6dd	feat: remove sm75 (#9207 )	2025-08-14 22:27:14 -07:00
strgrb	1f9d65f57d	use fast math for per_token_group_quant_8bit. (#9177 ) Co-authored-by: Zhang Kaihong <zhangkaihong.zkh@alibaba-inc.com>	2025-08-14 22:19:56 -07:00
Peng Zhang	5aa1ebd242	[2/n]decouple quantization implementation from vLLM dependency (#8112 ) Co-authored-by: walker-ai <yiyun.wyt@antgroup.com> Co-authored-by: leoneo <1320612015@qq.com>	2025-08-14 03:19:03 -07:00
DarkSharpness	86a0be65d8	[Feature] Support custom set kv buffer kernel (#8884 )	2025-08-12 16:56:51 -07:00
Lianmin Zheng	2c7f01bc89	Reorganize CI and test files (#9027 )	2025-08-10 12:30:06 -07:00
Yineng Zhang	8e8545caf6	fix: update cmake (#8817 )	2025-08-05 09:38:30 -07:00
Qiaolin Yu	fc8c8e5041	Integrate triton_kernels in sgl-kernel (#8762 )	2025-08-04 12:12:14 -07:00
Baizhou Zhang	91e3d1542e	Update Cutlass in sgl-kernel to v4.1 (#8392 )	2025-07-27 00:36:15 -07:00
Yineng Zhang	4c605235aa	fix: workaround for deepgemm warmup issue (#8302 )	2025-07-23 12:01:51 -07:00
Baizhou Zhang	282eb59ff3	Add bf16 output option for dsv3_router_gemm kernel (#7999 )	2025-07-20 09:49:37 +08:00
ykcombat	1ebec1a8b0	[Feature] CUDA Green Context Support (#7649 )	2025-07-15 02:49:16 +08:00
SijiaYang	da3890e82a	[1/n]: add cutlass W4A8 moe kernel for hopper architecture (#7772 ) Signed-off-by: yangsijia.614 <yangsijia.614@bytedance.com> Co-authored-by: yicwang <yichen.wang@bytedance.com>	2025-07-04 20:50:12 -07:00
AniZpZ	8e03b641ba	[1/n] apply wna16marlin kernel in moe weight only quantization (#7683 ) Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com> Co-authored-by: yych0745 <1398089567@qq.com> Co-authored-by: HandH1998 <1335248067@qq.com> Co-authored-by: 弋云 <yiyun.wyt@antgroup.com> Co-authored-by: walker-ai <2398833647@qq.com>	2025-07-01 23:21:25 -07:00
Baizhou Zhang	7248272ccc	Add dsv3 router gemm kernel (#7627 )	2025-06-29 23:31:55 -07:00
Ke Bao	04b35190e2	Add dsv3 fused a gemm to sgl-kernel (#7630 )	2025-06-29 02:52:24 -07:00
Ruihang Lai	16d76b9f23	[CMake] Fix sgl-kernel CMakeLists for Blackwell (#7543 )	2025-06-25 19:00:46 -07:00
Zhiqiang Xie	34c3f9b2d3	kvcache io kernels and test case (#7382 )	2025-06-23 11:58:59 -07:00
Lianmin Zheng	55e03b10c4	Fix a bug in BatchTokenIDOut & Misc style and dependency updates (#7457 )	2025-06-23 06:20:39 -07:00
Yineng Zhang	7046e0fab7	feat: update blackwell setup (#7119 )	2025-06-12 01:54:40 -07:00
Yuan Luo	84727a5139	[sgl-kernel] Add cuda kernel for moe_ep_silu_and_mul (#6919 ) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>	2025-06-11 20:43:08 -07:00
JieXin Liang	22fe787852	[sgl-kernel] update deepgemm (#6942 )	2025-06-06 23:24:41 -07:00
zyksir	8e3797be1c	support 1 shot allreduce in 1-node and 2-node using mscclpp (#6277 )	2025-06-04 22:11:24 -07:00
Pavani Majety	eb38c7d1ca	[1/2] Add Kernel support for Cutlass based Fused FP4 MoE (#6093 ) Signed-off-by: Pavani Majety <pmajety@nvidia.com>	2025-06-02 13:48:03 -07:00
Yuan Luo	55444ed667	[EP] Add cuda kernel for moe_ep_pre_reorder (#6699 ) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>	2025-06-01 20:49:01 -07:00
Qiaolin Yu	0b9557fcd7	Disable compiling arch below sm_90 in aarch64 by default (#6380 )	2025-05-27 15:50:02 -07:00
HandH1998	4d643f6c7a	[1/2] Support Qserve (#6457 ) Co-authored-by: yych0745 <1398089567@qq.com> Co-authored-by: sleepcoo <sleepcoo@gmail.com>	2025-05-21 19:48:59 -07:00
Elfie Guo	6fc9357503	[2/2] Add python wrapper for CUTLASS FP8 Blockscale MoE Kernel. (#5694 )	2025-05-16 13:14:07 -07:00
Elfie Guo	c23a7072b6	Upgrade CUTLASS 4.0 (#6336 ) Co-authored-by: zhyncs <me@zhyncs.com>	2025-05-15 17:42:23 -07:00
Yineng Zhang	213e8c7dd5	chore: upgrade deepgemm (#6073 )	2025-05-11 02:17:24 -07:00
applesaucethebun	2ce8793519	Add typo checker in pre-commit (#6179 ) Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>	2025-05-11 12:55:00 +08:00

1 2

89 Commits