sglang

Author	SHA1	Message	Date
Yineng Zhang	87dab54824	Revert "chore: bump sgl-kernel v0.3.6 (#9220 )" (#9247 )	2025-08-15 17:24:36 -07:00
Yineng Zhang	5121af4627	Revert "chore(docker): update sgl_kernel version to 0.3.6 in Dockerfi… (#9246 )	2025-08-15 17:19:38 -07:00
kk	983aa4967b	Fix nan value generated after custom all reduce (#8663 ) Co-authored-by: wunhuang <wunhuang@amd.com>	2025-08-15 12:33:54 -07:00
Hubert Lu	9c3e95d98b	[AMD] Expand test coverage for AMD CI and enable apply_token_bitmask_inplace_cuda in sgl-kernel (#8268 )	2025-08-15 12:32:51 -07:00
ishandhanani	e52c3866eb	chore(docker): update sgl_kernel version to 0.3.6 in Dockerfile.gb200 (#9243 )	2025-08-15 12:06:52 -07:00
Simo Lin	da53e13cbb	[router] preserve original worker response header in router (#9236 )	2025-08-15 11:01:47 -07:00
Jeff Nettleton	d7e38b2f6d	[router] clean up lint warnings with clippy execution (#9201 )	2025-08-15 11:01:21 -07:00
Simo Lin	21b8846066	[router] allow more health check configuration (#9198 )	2025-08-15 08:07:45 -07:00
Liangsheng Yin	0c8594e67d	Optional extension for green context (#9231 )	2025-08-15 21:33:52 +08:00
Yineng Zhang	c186feed7f	chore: bump sgl-kernel v0.3.6 (#9220 )	2025-08-15 02:50:50 -07:00
Cheng Wan	84b006b278	Cleanup MoE Refactor (#9223 )	2025-08-15 02:28:33 -07:00
Shangming Cai	8ca07bd948	[CI] Fix sgl-router disaggregation test (#9222 ) Signed-off-by: Shangming Cai <csmthu@gmail.com>	2025-08-15 02:24:44 -07:00
jy-song-hub	4fc09e0df0	Fp4 MOE quant kernel optimization (#8777 ) Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com>	2025-08-15 01:46:16 -07:00
PGFLMG	a3d99d6dcd	[Misc] feat: Deepgemm update for sgl-kernel (#8790 )	2025-08-15 01:05:27 -07:00
Xuchun Shang	189af90896	[Eagle Warning fix] replace the deprecated 'and' with & (#9215 ) Signed-off-by: Xuchun Shang <xuchun.shang@linux.alibaba.com>	2025-08-15 15:43:36 +08:00
fzyzcjy	f8644a5632	Tiny update tmux history limit on dev container (#9218 )	2025-08-15 00:22:08 -07:00
Cheng Wan	e3e75a786a	Fix the deprecation warning for enable_flashinfer_mxfp4_moe (#9214 )	2025-08-14 23:59:35 -07:00
shilinlee	d4db9b028b	fix: the store_dtype typo for ascend mla (#9208 ) Signed-off-by: shilinlee_com <836160610@qq.com>	2025-08-14 23:58:42 -07:00
hzh0425	f7dd651dbd	feat(hicache-3fs): 3FS-SGLang Hierarchical Cache Deployment Guide (#9213 )	2025-08-14 23:54:31 -07:00
Yineng Zhang	9d54c6e6dd	feat: remove sm75 (#9207 )	2025-08-14 22:27:14 -07:00
strgrb	1f9d65f57d	use fast math for per_token_group_quant_8bit. (#9177 ) Co-authored-by: Zhang Kaihong <zhangkaihong.zkh@alibaba-inc.com>	2025-08-14 22:19:56 -07:00
Cheng Wan	295895120d	[6/N] MoE Refactor: Cleanup MoE-related configs (#8849 )	2025-08-14 21:14:53 -07:00
Mick	584e1ab2d0	fix: fix unsupported palette mode of images in bench_serving for mmmu (#9206 )	2025-08-14 18:44:46 -07:00
fzyzcjy	392de007cb	Minor fix docker container DeepEP on multi platforms (#9205 )	2025-08-14 17:41:49 -07:00
Philo	004f7f1972	[typo fix] Fix a typo in communicator.py (#9183 ) Signed-off-by: Philo <lul16@foxmail.com>	2025-08-14 17:29:38 -07:00
zixuanzhang226	d2fbf2de0c	feat: add fused moe config for Qwen3-235B-A22B-FP8 on B200 (#9204 )	2025-08-14 17:21:30 -07:00
Yineng Zhang	fab0f6e77d	chore: bump v0.5.0rc2 (#9203 )	2025-08-14 16:11:16 -07:00
Yineng Zhang	27985c27aa	feat: update model config (#9202 )	2025-08-14 15:15:27 -07:00
Yineng Zhang	ac474869d4	chore: upgrade transformers 4.55.2 (#9197 )	2025-08-14 13:51:02 -07:00
Adarsh Shirawalmath	0b1e04f083	[VLM] Improving multimodal tensor hash kernel (#9008 )	2025-08-14 13:45:55 -07:00
Chengxing Xie	c1c7dc4534	feat: Add model version tracking with API endpoints and response metadata (#8795 )	2025-08-14 12:13:46 -07:00
Hongbo Xu	2cc9eeab01	[4/n]decouple quantization implementation from vLLM dependency (#9191 ) Co-authored-by: AniZpZ <aniz1905@gmail.com> Co-authored-by: Yineng Zhang <me@zhyncs.com>	2025-08-14 12:05:46 -07:00
Xiaoyu Zhang	63d82a776a	refine mxfp4 shuffling log (#9194 )	2025-08-14 10:57:29 -07:00
Yuan Luo	53dcc750b6	[sgl-kernel] Support FlashInfer top_k_top_p_sampling_from_logits (#9060 ) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>	2025-08-14 10:56:36 -07:00
Yuan Luo	432f2053dd	[sgl-kernel] 1/N Refactor sglang cutlass 3x - gemm fp8 blockwise sm90 (#8913 ) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>	2025-08-14 10:55:54 -07:00
Yineng Zhang	1fea998a45	chore: bump sgl-kernel v0.3.5 (#9185 )	2025-08-14 03:20:48 -07:00
Peng Zhang	5aa1ebd242	[2/n]decouple quantization implementation from vLLM dependency (#8112 ) Co-authored-by: walker-ai <yiyun.wyt@antgroup.com> Co-authored-by: leoneo <1320612015@qq.com>	2025-08-14 03:19:03 -07:00
eigen	4dbf43601d	fix: zero_init buffer (#9065 ) Co-authored-by: Yineng Zhang <me@zhyncs.com>	2025-08-14 02:39:09 -07:00
lukec	3d6be1fbce	add w8a8-fp8-block-wise H20-3e triton config (#8018 )	2025-08-13 23:15:09 -07:00
Jun Liu	4063234c1a	Add H200 fused MoE kernel configs for DeepSeek-V3 in triton 3.3.1 (#7687 )	2025-08-13 23:14:09 -07:00
Tommy Yang	83feef5b2c	Add H20 fused MoE kernel configs for Dpsk & Qwen3 (#7631 )	2025-08-13 23:13:22 -07:00
Brayden Zhong	2871eacc05	Add Triton Fused MoE kernel config for E=16 on B200 (#7004 )	2025-08-13 23:12:27 -07:00
forestlee95	ac15bdc194	Add H200 fused MoE kernel tuning configs for Qwen3-Coder-480B-A35B-Instruct (#8852 )	2025-08-13 23:11:11 -07:00
Li Hui	d6451c3f65	Add A800 fused MoE kernel tuning configs for GLM4.5 and GLM4.5-Air (#8808 )	2025-08-13 23:03:17 -07:00
henryg	841810f227	[Perf] Tunings for SM100 FP8 CUTLASS kernel (#8818 )	2025-08-13 21:59:22 -07:00
pansicheng	733446dd36	fix io group (#9154 ) Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>	2025-08-14 12:46:42 +08:00
wxzhoucs	4c22897a66	Feature: support qwen and llama4 reducescatter for dp attention padding (#9101 )	2025-08-13 21:10:29 -07:00
Alex Yang	1bc183c6de	Faster weight processing (trtllm-gen moe nvfp4) (#9162 )	2025-08-13 21:09:34 -07:00
Cheng Wan	b87aacb5c5	[DP Attention] Refactor: adding some utility functions (#9136 )	2025-08-13 21:08:06 -07:00
fzyzcjy	b3363cc1aa	Fix docker container DeepEP error on Blackwell (#9171 )	2025-08-13 21:06:48 -07:00

1 2 3 4 5 ...

4690 Commits