sglang

Author	SHA1	Message	Date
Zhihao Zhang	24f7cb1ece	[speculative decoding] rename lookahead to ngram (#11010 ) Co-authored-by: a4zhangfei <a4zhangfei@qq.com>	2025-09-28 21:06:59 -07:00
Yuhao Yao	fe531d6f4e	[Bug] Fix Issue#10215 (#10572 )	2025-09-25 09:51:50 +08:00
Xiaoyu Zhang	c4e314f986	Restruct sgl-kernel benchmark (#10861 )	2025-09-25 07:45:25 +08:00
Zhihao Zhang	e7bc600304	[Feature] Speculative decoding support lookahead (#9873 ) Co-authored-by: a4zhangfei <a4zhangfei@qq.com> Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>	2025-09-18 16:42:41 -07:00
Yineng Zhang	6d55f60e77	Revert "[1/2] Optimizations and refactors about quant kernel (#9534 )" (#10292 )	2025-09-10 18:24:23 -07:00
Rain Jiang	2286e85e77	pass a_scale from fp8 quant result instead of hard code to 1.0f (#10241 ) Co-authored-by: Yichen Wang <yichen.wang@bytedance.com> Co-authored-by: Jinwu Guo <641876696@qq.com>	2025-09-10 12:56:05 -07:00
huangtingwei	5be8c2f7f7	Page first direct IO kernel (#10060 ) Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>	2025-09-10 13:35:34 +08:00
Yi Zhang	8cbe1538ef	Add mamba kernel (#10234 )	2025-09-09 12:58:43 -07:00
Yineng Zhang	94fb4e9e54	feat: support fa cute in sgl-kernel (#10205 ) Co-authored-by: cicirori <32845984+cicirori@users.noreply.github.com>	2025-09-09 00:14:39 -07:00
Yineng Zhang	cdc56ef6c1	feat: use sgl-kernel cu129 as default (#10188 )	2025-09-08 22:01:17 -07:00
Yuhao Yao	ee0b3c5bad	[1/N][Bug] Fix w4afp8 MoE NaN issue (sgl-kernel, fixed) (#10108 )	2025-09-07 21:39:07 -07:00
hlu1	5f1eb20484	[chore] Remove unused ep_moe cuda kernels (#9956 )	2025-09-06 01:35:50 -07:00
hlu1	4c22ebe2e8	Disable kernel cutlass_mla_decode on SM103 (#10058 ) Signed-off-by: Hao Lu <14827759+hlu1@users.noreply.github.com>	2025-09-06 01:35:18 -07:00
fzyzcjy	339f8eef09	[1/2] Optimizations and refactors about quant kernel (#9534 )	2025-09-05 18:45:08 +08:00
chenxj	d4a938417d	[feat] Support tp mode for DeepSeek-R1-W4AFP8 (#8118 ) Co-authored-by: yuhyao <827623970@qq.com>	2025-09-01 22:17:26 -07:00
Kaixi Hou	5c34b4f1c7	[NVIDIA] [2/N] Optimize `silu_and_mul_scaled_fp4_grouped_quant` perf (#9556 )	2025-08-29 17:17:03 -07:00
hlu1	a7d825fccc	Skip some tests on Blackwell (#9777 ) Signed-off-by: Hao Lu <14827759+hlu1@users.noreply.github.com>	2025-08-28 20:00:32 -07:00
Qi Yuhang	fda4792620	Update CUTLASS 4.2 & Enable K-Major Scale Factor for SM90 FP8 Blockwise Group GEMM (#9559 )	2025-08-24 23:24:43 -07:00
Kaixi Hou	e5638573c1	[NVIDA] [1/N] Nvfp4 Masked Gemm: Add quant op for the flashinfer grouped gemm (#9200 )	2025-08-22 12:19:45 -07:00
fzyzcjy	e85cb1ce9d	Fix quant kernel test errors and benchmark wrong output speeds (#7604 )	2025-08-21 03:48:41 -07:00
Trevor Morris	a91e90d9a3	[2/2] Fuse routed scaling factor into select_experts (#8690 )	2025-08-20 15:10:16 -07:00
Yichen Yan	c9bf3877a0	Reduce overhead for fa by not calling heavy CUDA property check (#7375 )	2025-08-20 16:26:28 +08:00
Yuan Luo	53dcc750b6	[sgl-kernel] Support FlashInfer top_k_top_p_sampling_from_logits (#9060 ) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>	2025-08-14 10:56:36 -07:00
Peng Zhang	5aa1ebd242	[2/n]decouple quantization implementation from vLLM dependency (#8112 ) Co-authored-by: walker-ai <yiyun.wyt@antgroup.com> Co-authored-by: leoneo <1320612015@qq.com>	2025-08-14 03:19:03 -07:00
Ke Bao	94f44b88d1	Update fa3 interface and add unit test (#9150 )	2025-08-13 20:05:02 +08:00
fzyzcjy	9aea255522	Fuse writing KV buffer into rope kernel (part 1: sgl-kernel) (#9077 )	2025-08-12 01:46:40 -07:00
Yineng Zhang	dd949ace23	Revert "[1/2][resubmit] sgl-kernel: Fuse routed scaling factor into m… (#9035 )	2025-08-10 17:34:54 -07:00
Trevor Morris	591c232f7c	[1/2][resubmit] sgl-kernel: Fuse routed scaling factor into moe_fused_gate (select_experts) (#8770 )	2025-08-08 17:55:06 -07:00
Hongbo Xu	39fd178831	refactor: Move scalar_types.py to sgl-kernel to avoid circular import (#8720 )	2025-08-07 19:22:16 -07:00
Liangsheng Yin	f9f0138f80	Revert "[1/2] sgl-kernel: Fuse routed scaling factor into select_experts" (#8706 )	2025-08-02 20:14:30 +08:00
Trevor Morris	f642524fd9	[1/2] sgl-kernel: Fuse routed scaling factor into select_experts (#8364 )	2025-08-01 18:14:24 -07:00
Stefan He	db7343c992	fix per token cuda kernel hidden dim cannot divide by 16 (#8543 )	2025-08-01 09:27:18 -07:00
Peter Pan	6bdd27861b	[Kimi K2] dsv3_router_gemm supports NUM_EXPERTS == 384 (#8013 )	2025-08-01 22:01:24 +08:00
Cheng Wan	a5f5ab4030	update sgl-kernel for EP: kernel part (#8514 ) Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com> Co-authored-by: Ke Bao <ispobaoke@gmail.com>	2025-07-30 22:19:55 -07:00
Ke Bao	5973675bc3	Fix moe align kernel test (#8531 )	2025-07-29 11:03:02 -07:00
Ke Bao	8af145b7dc	Fix test_moe_fused_gate_combined sgl-kernel ci test (#8374 )	2025-07-26 09:30:12 +08:00
Zhiqiang Xie	d40846d456	breakdown kernel update (#8334 )	2025-07-25 08:33:17 +08:00
Zhiqiang Xie	b43263307f	Hicache IO kernel refactoring (#8264 )	2025-07-23 16:49:03 +08:00
Ke Bao	e885bfdc6a	Fix sgl-kernel ci test (#8284 )	2025-07-23 14:01:47 +08:00
Baizhou Zhang	282eb59ff3	Add bf16 output option for dsv3_router_gemm kernel (#7999 )	2025-07-20 09:49:37 +08:00
Peng Zhang	c28ad1990d	[1/n] chore: decouple quantization implementation from vLLM dependency (#7992 )	2025-07-16 15:56:26 -07:00
Qi Yuhang	c268c11c71	[feat]Support fusion kernel for constructing quant input and scale factor for fp8_blockwise_scaled_grouped_mm (#8023 )	2025-07-15 00:02:44 -07:00
ykcombat	1ebec1a8b0	[Feature] CUDA Green Context Support (#7649 )	2025-07-15 02:49:16 +08:00
Qi Yuhang	26118a133d	[fix]Update unitest for fp8_blockwise_scaled_grouped_mm kernel (#7932 )	2025-07-11 14:29:13 -07:00
SijiaYang	da3890e82a	[1/n]: add cutlass W4A8 moe kernel for hopper architecture (#7772 ) Signed-off-by: yangsijia.614 <yangsijia.614@bytedance.com> Co-authored-by: yicwang <yichen.wang@bytedance.com>	2025-07-04 20:50:12 -07:00
Yi Zhang	2998c4bdf4	[optimize] fuse renormalize into moe_topk_softmax (#7744 ) Co-authored-by: ispobock <ispobaoke@gmail.com>	2025-07-03 12:42:44 -07:00
ayrnb	2c4feaf308	Add CUTLASS FP8 Blockscale MoE kernel for Hopper architecture (#7278 ) Co-authored-by: HydraQYH <QYH820@Outlook.com> Co-authored-by: TianQiLin666666 <1834987979@qq.com>	2025-07-02 23:27:03 -07:00
AniZpZ	8e03b641ba	[1/n] apply wna16marlin kernel in moe weight only quantization (#7683 ) Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com> Co-authored-by: yych0745 <1398089567@qq.com> Co-authored-by: HandH1998 <1335248067@qq.com> Co-authored-by: 弋云 <yiyun.wyt@antgroup.com> Co-authored-by: walker-ai <2398833647@qq.com>	2025-07-01 23:21:25 -07:00
Baizhou Zhang	7248272ccc	Add dsv3 router gemm kernel (#7627 )	2025-06-29 23:31:55 -07:00
Ke Bao	04b35190e2	Add dsv3 fused a gemm to sgl-kernel (#7630 )	2025-06-29 02:52:24 -07:00

1 2 3 4

161 Commits