sglang

Author	SHA1	Message	Date
Chunyuan WU	ac80f4da57	[CPU] [FP8] set SGLANG_CPU_FP8_CVT_FTZ in CMakeLists.txt (#7885 )	2025-07-09 01:53:53 -07:00
Chunyuan WU	128f16a817	[CPU]convert topk_weights to fp32 for INT8 and FP8 paths (for llama4) and fix LmHead weight pack (#7818 )	2025-07-08 19:27:24 -07:00
Ke Bao	a3398d8478	Optimize moe align block size kernel (#7794 )	2025-07-07 09:20:30 +08:00
Yineng Zhang	f200af0d8c	chore: bump sgl-kernel v0.2.4 (#7800 )	2025-07-05 15:03:31 -07:00
Lianmin Zheng	5589b75024	Add treemask mode to build_eagle_tree & release sgl-kernel 0.2.3 (#7756 ) Co-authored-by: Pranjal Shankhdhar <pranjal.ssh@gmail.com>	2025-07-05 12:17:05 -07:00
Yineng Zhang	4fece12be9	chore: bump sgl-kernel v0.2.3 (#7784 )	2025-07-05 00:05:45 -07:00
Mick	c797322280	fix: fix apply_shuffle_mul_sum (#7444 )	2025-07-04 23:23:30 -07:00
Qi Yuhang	8e9fb43d82	Optimize Hopper CUTLASS FP8 Blockwise Grouped GEMM Kernel in Small K Scenario (#7782 )	2025-07-04 22:25:49 -07:00
SijiaYang	da3890e82a	[1/n]: add cutlass W4A8 moe kernel for hopper architecture (#7772 ) Signed-off-by: yangsijia.614 <yangsijia.614@bytedance.com> Co-authored-by: yicwang <yichen.wang@bytedance.com>	2025-07-04 20:50:12 -07:00
Yineng Zhang	aca1101a13	chore: bump sgl-kernel 0.2.2 (#7755 )	2025-07-03 12:49:10 -07:00
Yi Zhang	2998c4bdf4	[optimize] fuse renormalize into moe_topk_softmax (#7744 ) Co-authored-by: ispobock <ispobaoke@gmail.com>	2025-07-03 12:42:44 -07:00
ayrnb	2c4feaf308	Add CUTLASS FP8 Blockscale MoE kernel for Hopper architecture (#7278 ) Co-authored-by: HydraQYH <QYH820@Outlook.com> Co-authored-by: TianQiLin666666 <1834987979@qq.com>	2025-07-02 23:27:03 -07:00
Chunyuan WU	36cc3ffdc7	[CPU] [sgl-kernel] set dispatch key of initialize to CatchAll (#7734 )	2025-07-02 22:39:24 -07:00
YanbingJiang	b044400dd3	Support non-contiguous query input for extend/decode attention (#7462 )	2025-07-02 19:59:45 -07:00
AniZpZ	8e03b641ba	[1/n] apply wna16marlin kernel in moe weight only quantization (#7683 ) Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com> Co-authored-by: yych0745 <1398089567@qq.com> Co-authored-by: HandH1998 <1335248067@qq.com> Co-authored-by: 弋云 <yiyun.wyt@antgroup.com> Co-authored-by: walker-ai <2398833647@qq.com>	2025-07-01 23:21:25 -07:00
Yineng Zhang	637bfee448	chore: bump sgl-kernel v0.2.1 (#7675 )	2025-06-30 22:12:33 -07:00
Chunyuan WU	6005eceee3	[CPU] remove process_group from inputs of shm_allreduce and shm_allgather (#7486 )	2025-06-30 21:54:11 -07:00
Baizhou Zhang	7248272ccc	Add dsv3 router gemm kernel (#7627 )	2025-06-29 23:31:55 -07:00
Chunyuan WU	c5131f7a2f	[CPU] add c++ kernel to bind CPU cores and memory node (#7524 )	2025-06-29 19:45:25 -07:00
Ke Bao	04b35190e2	Add dsv3 fused a gemm to sgl-kernel (#7630 )	2025-06-29 02:52:24 -07:00
Ruihang Lai	16d76b9f23	[CMake] Fix sgl-kernel CMakeLists for Blackwell (#7543 )	2025-06-25 19:00:46 -07:00
Chunyuan WU	7eb47b0f3d	[CPU] [BF16] Call fused_experts_cpu, weight_packed_linear and bmm_cpu kernel in DeepSeek model (#6641 ) Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg>	2025-06-25 01:43:33 -07:00
Ke Bao	57ab776910	Fuse sorted_token_ids padding to moe_align_block_size kernel (#7437 )	2025-06-24 17:44:27 -07:00
Yineng Zhang	e846d95ef6	chore: bump sgl-kernel v0.2.0 (#7490 )	2025-06-23 22:29:50 -07:00
Zhiqiang Xie	34c3f9b2d3	kvcache io kernels and test case (#7382 )	2025-06-23 11:58:59 -07:00
Lianmin Zheng	55e03b10c4	Fix a bug in BatchTokenIDOut & Misc style and dependency updates (#7457 )	2025-06-23 06:20:39 -07:00
kk	8aa68ed5c4	Solve docker build failed in the virtual machine (#7290 ) Co-authored-by: wunhuang <wunhuang@amd.com> Co-authored-by: Sai Enduri <saimanas.enduri@amd.com> Co-authored-by: HAI <hixiao@gmail.com>	2025-06-23 09:10:30 +00:00
xutizhou	506c4928f5	feat: integrate deepgemm into EPMoE (#6821 ) Co-authored-by: tianqilin.99 <tianqilin.99@bytedance.com> Co-authored-by: TianQiLin666666 <1834987979@qq.com> Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>	2025-06-23 01:38:58 -07:00
linzhuo	1de4db9bef	update invalid link in doc (#7297 )	2025-06-18 01:37:36 -07:00
Yineng Zhang	0650e5176f	fix: only enable flash_attn test on sm80 sm90 (#7289 )	2025-06-17 16:56:41 -07:00
AniZpZ	3eb4a800e8	Fix AWQ Dequant and Weight Loading of deepseek v2 (#6842 )	2025-06-17 13:45:10 -07:00
Lianmin Zheng	8321f8e45e	Release sgl-kernel 0.1.9 (#7232 )	2025-06-16 03:37:40 -07:00
Lianmin Zheng	cfceb83d05	Fix sampling for speculative decoding & simplify kernels (#7207 )	2025-06-16 03:28:30 -07:00
Yineng Zhang	4473320380	chore: bump v0.1.8.post2 (#7189 )	2025-06-14 17:01:48 -07:00
JieXin Liang	ab1a4fa5cb	[fix] fix cutlass_mla_backend with cuda_graph and add sm_scale for sgl-kernel cutlass_mla (#7184 )	2025-06-14 12:45:41 -07:00
Yineng Zhang	8ab7d93c2e	chore: bump v0.1.8.post1 (#7152 )	2025-06-13 03:14:26 -07:00
fzyzcjy	5c66c4424f	Support new DeepGEMM format in per token group quant (#7146 )	2025-06-13 02:00:22 -07:00
fzyzcjy	aa46ed34d2	Remove 200us slow concat kernel (part 1: kernel) (#7145 )	2025-06-13 01:58:29 -07:00
sogalin	4b9971e401	Add gfx950 support for sgl-kernel. (#7092 ) Co-authored-by: HAI <hixiao@gmail.com> Co-authored-by: Yineng Zhang <me@zhyncs.com>	2025-06-12 11:07:48 -07:00
Yineng Zhang	7046e0fab7	feat: update blackwell setup (#7119 )	2025-06-12 01:54:40 -07:00
Yuan Luo	84727a5139	[sgl-kernel] Add cuda kernel for moe_ep_silu_and_mul (#6919 ) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>	2025-06-11 20:43:08 -07:00
Yineng Zhang	344adb00ec	fix arm sgl-kernel link issue (#7066 )	2025-06-10 15:23:22 -07:00
fzyzcjy	19995dd78e	Tiny fix cutlass_mla_get_workspace_size stub incorrect signature (#7057 )	2025-06-10 12:27:57 -07:00
YanbingJiang	fcde67b016	CPU: map changes from developing branch in sgl-kernel (#6833 ) Co-authored-by: mingfeima <mingfei.ma@intel.com>	2025-06-10 01:08:15 -07:00
JieXin Liang	18efb5e8e0	[perf][sgl-kernel] extend cutlass_mla_decode to support num_head < 128 (#6929 )	2025-06-08 19:37:34 -07:00
Yineng Zhang	6c0a48282a	chore: bump sgl-kernel v0.1.7 (#6963 )	2025-06-08 02:43:15 -07:00
Yineng Zhang	8db3ac55a9	chore: bump sgl-kernel v0.1.6.post1 (#6955 )	2025-06-07 15:25:46 -07:00
Elfie Guo	3e56f557fd	Add a CUDA kernel for fusing mapping and weighted sum for MoE. (#6916 ) Co-authored-by: Elfie Guo <elfiegxf@gmail.com>	2025-06-07 15:24:39 -07:00
Xiaoyu Zhang	8b5f83ed3b	reduce torch.zeros overhead in moe align block size kernel (#6369 )	2025-06-07 02:47:36 -07:00
Yineng Zhang	d664ca18f2	chore: bump sgl-kernel v0.1.6 (#6943 )	2025-06-07 00:25:22 -07:00

1 2 3 4 5 ...

372 Commits