sglang

Author	SHA1	Message	Date
Ximingwang-09	f2a75a66c4	update doc (#7046 ) Co-authored-by: ximing.wxm <ximing.wxm@antgroup.com>	2025-06-11 10:02:01 +08:00
Lianmin Zheng	6b12d6a8d5	Simplify the heuristics for setting --mem-fraction-static (#7054 )	2025-06-10 19:01:39 -07:00
Lianmin Zheng	0f218731e3	Do not run frontend_reasoning.ipynb to reduce the CI load (#7073 )	2025-06-10 17:15:31 -07:00
Yudi Xue	14c18d25df	Frontend language separate reasoning support (#6031 )	2025-06-10 17:11:29 -07:00
Lianmin Zheng	90bd3e32d6	Improve perf tuning docs (#7071 )	2025-06-10 16:55:04 -07:00
Brayden Zhong	ca9291181d	[Feature] Add Logit Bias (#6579 ) Co-authored-by: Cinjon Resnick <cinjon.resnick@gmail.com>	2025-06-10 15:39:25 -07:00
Yineng Zhang	344adb00ec	fix arm sgl-kernel link issue (#7066 )	2025-06-10 15:23:22 -07:00
kyle-pena-kuzco	b56de8f943	Open AI API hidden states (#6716 )	2025-06-10 14:37:29 -07:00
Yineng Zhang	ce5ee3bdf0	chore: update dev docker (#7064 )	2025-06-10 12:57:58 -07:00
Xu Wenqing	a0e4d4eb53	Fix missing tool call id if tool call index >0 in streaming tool call output. (#7049 ) Signed-off-by: 许文卿 <xwq391974@alibaba-inc.com>	2025-06-10 12:56:43 -07:00
Yineng Zhang	2f58445531	Revert "Add sanity checks when a test file is not added to CI (#6947 )" (#7063 )	2025-06-10 12:43:25 -07:00
fzyzcjy	fe55947acd	Add sanity checks when a test file is not added to CI (#6947 )	2025-06-10 12:34:57 -07:00
fzyzcjy	19995dd78e	Tiny fix cutlass_mla_get_workspace_size stub incorrect signature (#7057 )	2025-06-10 12:27:57 -07:00
Baizhou Zhang	3b014bc13d	Fix test_lora.py CI (#7061 )	2025-06-10 12:24:46 -07:00
Hank Han	d7c3e8e93d	[fix] libmlx5.so already in base image (#7060 )	2025-06-10 12:22:24 -07:00
kk	8ea7df6114	[WA] fix output data is nan in CI test "test_moe_eval_accuracy_large.py" (#7021 ) Co-authored-by: wunhuang <wunhuang@amd.com> Co-authored-by: HAI <hixiao@gmail.com>	2025-06-10 16:08:10 +00:00
Lianmin Zheng	4a102a2b02	Minor style fix in cuda_graph_runner.py (#7053 )	2025-06-10 06:32:41 -07:00
Lianmin Zheng	6406408a70	Clean up server_args.py (#7037 )	2025-06-10 05:34:29 -07:00
Lianmin Zheng	019851d099	Fix eagle on AMD (#7051 )	2025-06-10 05:22:40 -07:00
Lianmin Zheng	2dae104dca	Minor cleanup of fa3 backend (#6999 )	2025-06-10 03:58:44 -07:00
Yineng Zhang	cef6655b26	fix 24.12 docker (#7045 )	2025-06-10 02:43:22 -07:00
Swipe4057	27196d4148	[Docker] Upgrading base image from 24.04 to 24.12 (#7043 )	2025-06-10 02:23:05 -07:00
Lianmin Zheng	bb185b0e92	Update README.md (#7040 )	2025-06-10 01:59:14 -07:00
Yineng Zhang	4f723edd3b	chore: bump v0.4.7 (#7038 )	2025-06-10 01:56:20 -07:00
YanbingJiang	fcde67b016	CPU: map changes from developing branch in sgl-kernel (#6833 ) Co-authored-by: mingfeima <mingfei.ma@intel.com>	2025-06-10 01:08:15 -07:00
yudian0504	81372f3bef	Fix fused_moe triton configs (#7029 )	2025-06-09 23:23:03 -07:00
Arthur Cheng	baa6624d7c	[CI] Add CI workflow for sgl-router docker build (#7027 )	2025-06-09 23:16:44 -07:00
Byron Hsu	c2b16795b5	Add decode req pool (#6980 )	2025-06-09 21:23:36 -07:00
fzyzcjy	f6ebba537a	Support both approximate and exact expert distribution collection (#6964 )	2025-06-09 20:56:17 -07:00
Baizhou Zhang	6716b41786	Update default settings for blackwell (#7023 )	2025-06-09 20:37:47 -07:00
Yineng Zhang	1c8b42c84c	chore: update pr test xeon (#7018 )	2025-06-09 17:36:25 -07:00
Qiaolin Yu	f20f70003d	Fix torch version in blackwell dockerfile (#7017 )	2025-06-09 17:06:50 -07:00
Emmanuel Ferdman	f40942ad63	Migrate to assertEqual (#6741 ) Signed-off-by: Emmanuel Ferdman <emmanuelferdman@gmail.com>	2025-06-09 16:47:39 -07:00
Lianmin Zheng	dc0705a504	Simplify prepare_extend_after_decode (#6987 )	2025-06-09 16:39:21 -07:00
Wenxuan Tan	a968c888c0	Fix torchvision version for Blackwell (#7015 )	2025-06-09 15:50:19 -07:00
Baizhou Zhang	a979daac3b	Fallback to lower triton version for unfound fused moe configs (#7013 )	2025-06-09 15:41:03 -07:00
ishandhanani	f1569876d5	feat: add direct routing strategy to DP worker (#6884 )	2025-06-09 11:44:05 -07:00
Sai Enduri	3465d7ae78	Update amd nightly models CI. (#6992 )	2025-06-09 10:54:08 -07:00
fzyzcjy	e58423b2b9	Fix cutlass MLA gets almost zero accuracy (#6998 )	2025-06-09 10:16:29 -07:00
Yineng Zhang	7059ae16fb	chore: update pr test xeon (#7008 )	2025-06-09 10:08:44 -07:00
Yineng Zhang	51d9a597f9	cleanup tmp dir (#7007 )	2025-06-09 09:26:04 -07:00
Yineng Zhang	56ccd3c22c	chore: upgrade flashinfer v0.2.6.post1 jit (#6958 ) Co-authored-by: alcanderian <alcanderian@gmail.com> Co-authored-by: Qiaolin Yu <qy254@cornell.edu> Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com> Co-authored-by: Mick <mickjagger19@icloud.com> Co-authored-by: ispobock <ispobaoke@gmail.com>	2025-06-09 09:22:39 -07:00
Yueyang Pan	98c00a2df1	Fix torch profiler bugs for bench_offline_throughput.py (#6557 )	2025-06-09 20:33:41 +08:00
Pan Lyu	451ffe74d9	support qwen3 emebedding (#6990 )	2025-06-09 01:32:49 -07:00
Lifu Huang	b1e5a33ae3	Eliminate stream sync to speed up LoRA batch init (#6960 )	2025-06-09 00:22:45 -07:00
Lianmin Zheng	9d5fa68b90	Use torch.compile to fuse flash attention decode metadata preparation (#6973 )	2025-06-08 23:05:40 -07:00
Sai Enduri	2c18642502	Enable more unit tests for AMD CI. (#6983 )	2025-06-08 19:41:55 -07:00
JieXin Liang	18efb5e8e0	[perf][sgl-kernel] extend cutlass_mla_decode to support num_head < 128 (#6929 )	2025-06-08 19:37:34 -07:00
fzyzcjy	de1350ea20	Minor remove one kernel for DeepSeek (#6977 )	2025-06-08 17:41:35 -07:00
fzyzcjy	86fe943bc3	Fix expert distribution dumping causes OOM (#6967 )	2025-06-08 17:41:14 -07:00

1 2 3 4 5 ...

3646 Commits