sglang

Author	SHA1	Message	Date
Liangsheng Yin	08ebdf79d0	Fix the `--allow-auto-truncate` argument in tokenizer manager. (#9391 ) Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-08-20 16:56:47 +08:00
fzyzcjy	42c8704560	Add PDL support for quant kernel and rope kernel (#9106 )	2025-08-20 01:56:29 -07:00
Yichen Yan	c9bf3877a0	Reduce overhead for fa by not calling heavy CUDA property check (#7375 )	2025-08-20 16:26:28 +08:00
Even Zhou	de2dd73831	Revert "[feature] Rework Ascend NPU graph support" (#9385 )	2025-08-20 00:35:10 -07:00
Lianmin Zheng	1ec9769753	[Docs] Update contribution guide (#9383 )	2025-08-19 23:37:45 -07:00
Shangming Cai	d8ed60f254	[CI] Fix disaggregation failure tolerance CI (#9378 ) Signed-off-by: Shangming Cai <csmthu@gmail.com>	2025-08-19 23:31:08 -07:00
Mingyi	f1b0eda55c	[readme] Add SGLang x AMD SF meetup information (#9380 )	2025-08-19 22:25:09 -07:00
Lianmin Zheng	f20b6a3f2b	[minor] Sync style changes (#9376 )	2025-08-19 21:35:01 -07:00
Even Zhou	3680d6f88b	[feature] Rework Ascend NPU graph support (#9350 ) Co-authored-by: ronnie_zheng <zl19940307@163.com> Co-authored-by: yezhifeng (D) <y00897525@china.huawei.com> Co-authored-by: anon189Ty <Stari_Falcon@outlook.com> Co-authored-by: Maksim <makcum888e@mail.ru> Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>	2025-08-19 20:32:27 -07:00
Keyang Ru	f515449582	Fix gpt-oss response api streaming issue (#9368 )	2025-08-19 20:19:42 -07:00
Ke Bao	e0ce171d79	Fix triton backend eagle illegal memory access (#9344 )	2025-08-19 20:16:26 -07:00
fzyzcjy	fe43e889f8	Fix mini lb timeout issue (#9369 )	2025-08-19 20:15:16 -07:00
Keyang Ru	5ae5ecaa15	[router] Implement OpenAI Responses API specification (#9367 )	2025-08-19 20:14:47 -07:00
Simo Lin	5fbad308cd	[router] add tokenizer chat template support (#9370 ) Co-authored-by: Chang Su <chang.s.su@oracle.com>	2025-08-19 20:14:02 -07:00
Chang Su	7638f5e44e	[router] Implement gRPC SGLangSchedulerClient (#9364 )	2025-08-19 16:44:11 -07:00
Simo Lin	b45f753cba	[router] adds reasoning parser pooling and thread-safe (#9360 )	2025-08-19 13:35:39 -07:00
Keyang Ru	c5057262fa	[Router] Add validation module for API parameters (#9335 )	2025-08-19 13:25:53 -07:00
Chang Su	46fe8b8cb2	[CI] Fix lint issues (#9361 )	2025-08-19 13:05:36 -07:00
Simo Lin	0b95a01a8f	[router] add tiktokenizer and sequence in router (#9354 ) Co-authored-by: Chang Su <chang.s.su@oracle.com>	2025-08-19 10:46:28 -07:00
mpashkovskiy	a3b810ebdb	fix: enable multi-GPU Triton fused MoE tuning (#6295 )	2025-08-19 10:16:58 -07:00
Simo Lin	94959237bf	[router] add dsr1, kimi, and qwen reasoning parser (#9353 )	2025-08-19 10:15:24 -07:00
Even Zhou	f4fafacc5d	Revert "[feature] Ascend NPU graph support (#8027 )" (#9348 )	2025-08-19 10:11:23 -07:00
chenxu140	01d47a27b6	[Bugfix] fix kv buffer register & dp attention & deepepmoe (#9327 )	2025-08-19 10:09:48 -07:00
Lianmin Zheng	ecc9f3e47a	[Minor] Fix the style of sgl-kernel (#9332 )	2025-08-18 23:45:00 -07:00
Yineng Zhang	7e8187e004	docs: fix spec (#9326 )	2025-08-18 19:35:46 -07:00
Enrique Shockwave	e483ab6d20	enable marlin fp8 blockwise (#8990 )	2025-08-18 18:53:15 -07:00
EduardDurech	720cd308ba	Add `CMakeLists.txt` binary_dir (#7019 )	2025-08-18 18:36:33 -07:00
Keyang Ru	ce67b2d586	[router]restructure protocol modules for better organization (#9321 )	2025-08-19 01:07:58 +00:00
Jiaqi Gu	3c2c9f6c9e	[Bug] Fix input arguments of flashinfer_trtllm_moe (#9317 )	2025-08-18 18:03:19 -07:00
zxy	a31ea44824	support for interns1-mini (#9299 )	2025-08-18 17:56:04 -07:00
Chang Su	439df4548a	[router] Add spec for sglang scheduler (#9322 )	2025-08-18 17:20:20 -07:00
fzyzcjy	5626e20b2b	Tiny fix CI (#9306 )	2025-08-18 16:54:36 -07:00
Hubert Lu	c6c379ab31	[AMD] Reorganize hip-related header files in sgl-kernel (#9320 )	2025-08-18 16:53:44 -07:00
Binyao Jiang	c2fbf60f39	[GLM4.1V and GLM4.5V] Add vision transformer num_dummy_head support: max tp=4 -> max tp=8 (#9059 )	2025-08-18 14:40:13 -07:00
datdo-msft	98b44e9e56	[PD] Propagate internal server errors from aborted requests to clients instead of blindly returning 200's (#8936 )	2025-08-18 14:23:46 -07:00
Swipe4057	6805f6da40	upgrade xgrammar 0.1.23 and openai-harmony 0.0.4 (#9284 )	2025-08-18 14:02:00 -07:00
江家瑋	ca533580f2	[Docs] Correct and clarify notes in Engine docstring (#9313 ) Signed-off-by: JiangJiaWei1103 <waynechuang97@gmail.com>	2025-08-18 13:24:19 -07:00
Keyang Ru	886454e8e7	[MISC] use dynamic choices for tool-call-parser argument (#9316 )	2025-08-18 13:02:10 -07:00
gongwei-130	0cf3fbeb18	should return invalide request for empty prompt (#9315 )	2025-08-18 11:44:11 -07:00
Zhiyu	2256d62d36	Modelopt quant config adaptation (#8829 )	2025-08-18 11:27:30 -07:00
JieXin Liang	6cdcbcc674	[fix] fix enable_pdl for blackwell (#9011 )	2025-08-19 01:16:08 +08:00
Lianmin Zheng	c480a3f6ea	Minor style fixes for sgl-kernel (#9289 )	2025-08-18 09:38:35 -07:00
Simo Lin	6e316588f8	[router] add reasoning parser base structure (#9310 ) Co-authored-by: Chang Su <chang.s.su@oracle.com>	2025-08-18 09:26:09 -07:00
Simo Lin	24247b4168	[router] add tokenizer metrics (#9307 ) Co-authored-by: Chang Su <chang.s.su@oracle.com>	2025-08-18 09:25:51 -07:00
fzyzcjy	4c0bb411e5	Further fix memory pool leak error (#9298 )	2025-08-18 00:58:06 -07:00
Yuan Luo	968e181826	Fix triton_fused_moe unit test and benchmark (#9276 ) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>	2025-08-18 00:54:33 -07:00
Simo Lin	d08663eec1	[router] tokenizer factory, hf tokenizer, and stop sequence detector (#9293 ) Co-authored-by: Chang Su <chang.s.su@oracle.com>	2025-08-17 22:38:38 -07:00
b8zhong	716e682721	[Fix] Add undefined `update_tensor_inplace` function (#6307 )	2025-08-18 11:11:00 +08:00
zifeitong	84b30d9e00	Set the default attention backend for GLM-4.5v to fa3 (#9245 )	2025-08-17 16:34:19 -07:00
Simo Lin	ff0cf51c8e	[router] introducing tokenizer trait (#9287 )	2025-08-17 16:30:01 -07:00

1 2 3 4 5 ...

4765 Commits