sglang

Author	SHA1	Message	Date
Liangsheng Yin	08ebdf79d0	Fix the `--allow-auto-truncate` argument in tokenizer manager. (#9391 ) Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-08-20 16:56:47 +08:00
fzyzcjy	42c8704560	Add PDL support for quant kernel and rope kernel (#9106 )	2025-08-20 01:56:29 -07:00
Even Zhou	de2dd73831	Revert "[feature] Rework Ascend NPU graph support" (#9385 )	2025-08-20 00:35:10 -07:00
Lianmin Zheng	1ec9769753	[Docs] Update contribution guide (#9383 )	2025-08-19 23:37:45 -07:00
Lianmin Zheng	f20b6a3f2b	[minor] Sync style changes (#9376 )	2025-08-19 21:35:01 -07:00
Even Zhou	3680d6f88b	[feature] Rework Ascend NPU graph support (#9350 ) Co-authored-by: ronnie_zheng <zl19940307@163.com> Co-authored-by: yezhifeng (D) <y00897525@china.huawei.com> Co-authored-by: anon189Ty <Stari_Falcon@outlook.com> Co-authored-by: Maksim <makcum888e@mail.ru> Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>	2025-08-19 20:32:27 -07:00
Keyang Ru	f515449582	Fix gpt-oss response api streaming issue (#9368 )	2025-08-19 20:19:42 -07:00
Ke Bao	e0ce171d79	Fix triton backend eagle illegal memory access (#9344 )	2025-08-19 20:16:26 -07:00
fzyzcjy	fe43e889f8	Fix mini lb timeout issue (#9369 )	2025-08-19 20:15:16 -07:00
Even Zhou	f4fafacc5d	Revert "[feature] Ascend NPU graph support (#8027 )" (#9348 )	2025-08-19 10:11:23 -07:00
chenxu140	01d47a27b6	[Bugfix] fix kv buffer register & dp attention & deepepmoe (#9327 )	2025-08-19 10:09:48 -07:00
Enrique Shockwave	e483ab6d20	enable marlin fp8 blockwise (#8990 )	2025-08-18 18:53:15 -07:00
Jiaqi Gu	3c2c9f6c9e	[Bug] Fix input arguments of flashinfer_trtllm_moe (#9317 )	2025-08-18 18:03:19 -07:00
zxy	a31ea44824	support for interns1-mini (#9299 )	2025-08-18 17:56:04 -07:00
fzyzcjy	5626e20b2b	Tiny fix CI (#9306 )	2025-08-18 16:54:36 -07:00
Binyao Jiang	c2fbf60f39	[GLM4.1V and GLM4.5V] Add vision transformer num_dummy_head support: max tp=4 -> max tp=8 (#9059 )	2025-08-18 14:40:13 -07:00
datdo-msft	98b44e9e56	[PD] Propagate internal server errors from aborted requests to clients instead of blindly returning 200's (#8936 )	2025-08-18 14:23:46 -07:00
Swipe4057	6805f6da40	upgrade xgrammar 0.1.23 and openai-harmony 0.0.4 (#9284 )	2025-08-18 14:02:00 -07:00
江家瑋	ca533580f2	[Docs] Correct and clarify notes in Engine docstring (#9313 ) Signed-off-by: JiangJiaWei1103 <waynechuang97@gmail.com>	2025-08-18 13:24:19 -07:00
Keyang Ru	886454e8e7	[MISC] use dynamic choices for tool-call-parser argument (#9316 )	2025-08-18 13:02:10 -07:00
gongwei-130	0cf3fbeb18	should return invalide request for empty prompt (#9315 )	2025-08-18 11:44:11 -07:00
Zhiyu	2256d62d36	Modelopt quant config adaptation (#8829 )	2025-08-18 11:27:30 -07:00
Lianmin Zheng	c480a3f6ea	Minor style fixes for sgl-kernel (#9289 )	2025-08-18 09:38:35 -07:00
fzyzcjy	4c0bb411e5	Further fix memory pool leak error (#9298 )	2025-08-18 00:58:06 -07:00
b8zhong	716e682721	[Fix] Add undefined `update_tensor_inplace` function (#6307 )	2025-08-18 11:11:00 +08:00
zifeitong	84b30d9e00	Set the default attention backend for GLM-4.5v to fa3 (#9245 )	2025-08-17 16:34:19 -07:00
blzheng	ebbb75e917	[CPU] Fix TP padding issue on Phi-4 (#8289 )	2025-08-17 16:25:26 -07:00
fzyzcjy	b498cd21d7	Tiny make fp4 moe method parameters more static (#8520 )	2025-08-17 13:26:02 -07:00
kousakawang	0fc54b971e	[fix]: fix cutlass moe ut and and Opt H20 cutlass groupGemm performance (#9272 ) Co-authored-by: wanghanpei <wanghanpei@bytedance.com>	2025-08-17 13:09:49 -07:00
fzyzcjy	b3c1f2e4f2	Fix memory pool leak error (#9271 )	2025-08-17 12:53:34 -07:00
Ke Bao	be1a3cd9b4	Fix swa eagle verify accuracy for Triton backend (#9279 )	2025-08-17 12:52:02 -07:00
Lifu Huang	4b74c3fcca	[chore] Clean up redundant lora_weight_names concept to simplify code (#9131 )	2025-08-17 12:36:58 -07:00
Netanel Haber	3d77a31885	`from python.sglang.srt` -> `from sglang.srt` (#9268 )	2025-08-17 02:45:45 -07:00
Netanel Haber	845d12a979	model: support nvidia/Llama-3_3-Nemotron-Super-49B-v1 (#9067 ) Co-authored-by: Kyle Huang <kylhuang@nvidia.com>	2025-08-17 01:48:15 -07:00
Stefan He	e47800e176	Quick Fix GLM (#9264 )	2025-08-16 23:43:41 -07:00
Mick	1df84ff414	ci: simplify multi-modality tests by using mixins (#9006 )	2025-08-16 22:25:02 -07:00
Binyao Jiang	66d6be0874	Bug fix: use correct mm_items in embed_mm_inputs (#8893 )	2025-08-16 19:55:56 -07:00
kk	1c1f8a118e	Combine fp4.py and mxfp4.py into one file and support dynamic mxfp4 quantization in mxfp4.py (#9049 ) Co-authored-by: wunhuang <wunhuang@amd.com>	2025-08-16 19:01:54 -07:00
Shangming Cai	384f8ab5ce	[PD] Support PD disaggregation with Prefill PP (#8846 ) Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com> Signed-off-by: Shangming Cai <csmthu@gmail.com> Co-authored-by: root <huzhiyuan@xiaohongshu.com> Co-authored-by: Ying Sheng <sqy1415@gmail.com> Co-authored-by: Francis <38564764+ssssnow@users.noreply.github.com> Co-authored-by: zitto <zhjc1124@gmail.com>	2025-08-16 18:31:31 -07:00
zyksir	6a9d6ca33c	fix unexcepted answer in EAGLE mode (#9252 )	2025-08-16 17:45:36 -07:00
VDV1985	94371dbbd6	[feature] Ascend NPU graph support (#8027 ) Co-authored-by: ronnie_zheng <zl19940307@163.com> Co-authored-by: yezhifeng (D) <y00897525@china.huawei.com> Co-authored-by: anon189Ty <Stari_Falcon@outlook.com> Co-authored-by: Maksim <makcum888e@mail.ru> Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>	2025-08-16 17:25:17 -07:00
Hank Han	81da16f6d3	[CI] add deepseek w4a8 test on h20 ci (#7758 )	2025-08-16 01:54:13 -07:00
Brayden Zhong	bc938ea13f	Fix DP load for embedding (#9165 )	2025-08-15 23:58:44 -07:00
Trevor Morris	eff4eb3fdd	Add fp4 quantize before all-gather for Flashinfer cutlass MoE DP (max throughput) (#7667 )	2025-08-15 22:08:11 -07:00
kk	983aa4967b	Fix nan value generated after custom all reduce (#8663 ) Co-authored-by: wunhuang <wunhuang@amd.com>	2025-08-15 12:33:54 -07:00
Hubert Lu	9c3e95d98b	[AMD] Expand test coverage for AMD CI and enable apply_token_bitmask_inplace_cuda in sgl-kernel (#8268 )	2025-08-15 12:32:51 -07:00
Cheng Wan	84b006b278	Cleanup MoE Refactor (#9223 )	2025-08-15 02:28:33 -07:00
Xuchun Shang	189af90896	[Eagle Warning fix] replace the deprecated 'and' with & (#9215 ) Signed-off-by: Xuchun Shang <xuchun.shang@linux.alibaba.com>	2025-08-15 15:43:36 +08:00
Cheng Wan	e3e75a786a	Fix the deprecation warning for enable_flashinfer_mxfp4_moe (#9214 )	2025-08-14 23:59:35 -07:00
shilinlee	d4db9b028b	fix: the store_dtype typo for ascend mla (#9208 ) Signed-off-by: shilinlee_com <836160610@qq.com>	2025-08-14 23:58:42 -07:00

1 2 3 4 5 ...

3215 Commits