sglang

Author	SHA1	Message	Date
Yineng Zhang	a043f7f2ab	chore: use torch 2.6 for sgl-kernel build (#5898 )	2025-04-29 17:51:18 -07:00
Johnny	2c7dbb7cc2	[FEATURE] Enhance platform compatibility for ARM (#5746 )	2025-04-29 15:06:16 -07:00
Xiaoyu Zhang	5bb0accbcf	cutlass 3.9 supported to improve fp8_blockwise_gemm (#5820 )	2025-04-28 21:52:36 -07:00
PGFLMG	ee71ed8a41	[Feat] QWen-1M context support[1/2]: Update block sparse attention backend utils kernel (#5847 ) Co-authored-by: sighingnow <sighingnow@gmail.com>	2025-04-28 11:03:17 -07:00
Xiaoyu Zhang	ef15dcda26	Add a doc to fix sgl-kernel build link error in py39 with ccache (#5809 )	2025-04-27 21:34:27 -07:00
Trevor Morris	84810da4ae	Add Cutlass MLA attention backend (#5390 )	2025-04-27 20:58:53 -07:00
Yineng Zhang	7d0edf3cae	chore: bump sgl-kernel 0.1.0 (#5688 )	2025-04-23 14:23:59 -07:00
Yineng Zhang	ce4ecba477	fix: only compile ApplyTokenBitmaskInplace cu124+ (#5686 )	2025-04-23 14:17:42 -07:00
Yineng Zhang	15fabcc07f	fix sgl-kernel unit tests (#5666 )	2025-04-23 01:18:30 -07:00
Elfie Guo	e62c49557d	[1/2] Add FP8 Blockscale MoE CUTLASS kernel for Blackwell (#5281 )	2025-04-22 22:28:20 -07:00
Yubo Wang	20f1c8e374	Fix sampler nan check when calling top_k_top_p_sampling_from_probs (#5546 )	2025-04-19 21:47:23 -07:00
Yineng Zhang	f28d82997a	chore: bump sgl-kernel 0.0.9.post2 (#5518 )	2025-04-17 23:42:39 -07:00
Xiaoyu Zhang	8e09b37077	Sgl kernel fused_moe_gate support n_shared_experts (#5440 )	2025-04-17 23:05:15 -07:00
PGFLMG	c08a717c77	[Feat] Update sgl-kernel flashinfer to latest main version (#5500 ) Co-authored-by: zhyncs <me@zhyncs.com>	2025-04-17 12:43:23 -07:00
Michael Yao	22c2a79dc5	Fix a link in sgl-kernel/README.md (#5493 )	2025-04-17 02:25:28 -07:00
Baizhou Zhang	81c891111f	Add test for flash_attn_varlen_func kernel (#5484 )	2025-04-17 01:42:56 -07:00
Elfie Guo	85ec0440a5	Update cutlass dependency. (#5447 )	2025-04-15 23:28:04 -07:00
Trevor Morris	e8f62b20ca	BLackwell cutlass mla: Add check for bad page size/block num combinations (#5431 )	2025-04-15 14:07:42 -07:00
Yineng Zhang	88defc4d89	fix: solve release issue (#5434 )	2025-04-15 12:58:11 -07:00
Yineng Zhang	6f509d5503	chore: bump sgl-kernel v0.0.9.post1 (#5430 )	2025-04-15 11:00:21 -07:00
DefTruth	12ef7e3bc3	bugfix: fix merge_state_v2 cuda graph (#5419 )	2025-04-15 10:18:47 -07:00
Lianmin Zheng	838fa0f218	[minor] cleanup cmakelists.txt (#5420 )	2025-04-15 07:07:07 -07:00
Yineng Zhang	e940dc4f06	chore: bump sgl-kernel 0.0.9 (#5400 )	2025-04-14 21:34:04 -07:00
DefTruth	388e15c0db	kernel: support slightly faster merge_state_v2 cuda kernel (#5381 )	2025-04-14 21:28:23 -07:00
Yineng Zhang	6c41fcf0e4	chore: upgrade DeepGEMM (#5395 )	2025-04-14 20:32:46 -07:00
Lianmin Zheng	dae7944440	minor clean up of sgl-kernel/CMakeLists.txt (#5393 )	2025-04-14 18:38:44 -07:00
Yineng Zhang	b62e7e99b8	feat: adapt merge_state (#5337 )	2025-04-12 21:14:04 -07:00
Yineng Zhang	b371f7cd36	chore: bump sgl-kernel v0.0.8.post3 (#5332 )	2025-04-12 12:53:37 -07:00
Yineng Zhang	812e82f35e	fix: solve cu118 issue for cutlass mla (#5331 )	2025-04-12 12:51:09 -07:00
PGFLMG	4879e50c6d	[Feat] Add sparse attn to sgl-kernel (#5327 )	2025-04-12 11:36:36 -07:00
Yineng Zhang	115ae2e728	chore: bump sgl-kernel v0.0.8.post2 (#5317 )	2025-04-11 23:42:03 -07:00
Baizhou Zhang	e4155e96d0	Add flash_attn_varlen_func to sgl-kernel (#5315 )	2025-04-11 23:36:36 -07:00
Zhaoyi Li	3c9740d200	update variable naming and comments for rocm (#5299 )	2025-04-11 23:15:05 -07:00
Yineng Zhang	2eb55770f9	misc: cleanup 3rdparty (#5311 )	2025-04-11 22:53:50 -07:00
Trevor Morris	f65b8d5c89	Blackwell Cutlass MLA kernel (#5142 )	2025-04-11 22:16:51 -07:00
Yineng Zhang	4f288113ce	fix: update flash attn (#5308 )	2025-04-11 16:23:09 -07:00
Yineng Zhang	136b8e6afb	fix: remove cublas_grouped_gemm (#5307 )	2025-04-11 16:22:37 -07:00
Yineng Zhang	c1dd773c19	fix: use fa3 unit test on hopper only (#5304 )	2025-04-11 15:10:49 -07:00
Yineng Zhang	c163bf4ff1	chore: bump sgl-kernel v0.0.8.post1 (#5289 )	2025-04-11 02:11:53 -07:00
Yineng Zhang	5598634326	chore: relax the torch version restriction for sgl-kernel compilation (#5288 )	2025-04-11 02:05:53 -07:00
Yineng Zhang	b75275b6f2	feat: add cu128 identifier for sgl-kernel (#5287 )	2025-04-11 01:58:46 -07:00
Yineng Zhang	7074e9ca20	fix: enable fp4 compilation on cu128 (#5286 )	2025-04-11 01:43:44 -07:00
Elfie Guo	a222945df2	Update Makefile / build script to avoid installing incompatible torch dependency (#5245 )	2025-04-10 22:21:02 +00:00
PGFLMG	ed01b4515e	[Misc] Clean sgl-kernel test (#5216 )	2025-04-10 11:28:41 -07:00
HAI	d050df368c	ROCm sgl-kernel: compatible to later torch (#5167 )	2025-04-10 09:18:36 -07:00
Richard Zou	76f44c2a8d	Fix deepseek-v3 with torch.compile in PyTorch 2.6. (#5213 )	2025-04-10 09:14:38 -07:00
Xiaoyu Zhang	f730362ee2	reduce moe_align_block_size_kernel small batch mode overhead (#5086 )	2025-04-09 17:59:35 -07:00
Yi Zhang	ebf495f013	sgl-kernel use cutlass latest version for fp8 blockwise gemm (#5207 )	2025-04-09 11:47:04 -07:00
yinfan98	d2e507df3c	[Misc] clean up vllm in sgl-kernel test (#5189 )	2025-04-09 01:22:13 -07:00
Trevor Morris	11d760d56a	FP4 weight loading and inference (2/2) (#3972 )	2025-04-08 17:26:21 -07:00

1 2 3 4 5 ...

274 Commits