sglang

Author	SHA1	Message	Date
narutolhy	839c93bd2d	feat: add original logprobs to response (#8375 ) Co-authored-by: Chayenne <zhaochen20@outlook.com> Co-authored-by: luhongyu.4869 <luhongyu.4869@bytedance.com>	2025-08-29 11:43:57 -07:00
Hubert Lu	711390a971	[AMD] Support Hierarchical Caching on AMD GPUs (#8236 )	2025-08-28 15:27:07 -07:00
ZhengdQin	f92b729d52	[new feat] ascend backend support fia fusion kernel (#8328 ) Co-authored-by: Even Zhou <even.y.zhou@outlook.com>	2025-08-25 23:13:08 -07:00
Jonas	a0a77d937b	Fix Harmony reasoning parser for and auto-separation for gpt-oss models (#9190 ) Co-authored-by: Chang Su <chang.s.su@oracle.com> Co-authored-by: Chayenne <zhaochen20@outlook.com> Co-authored-by: zhaochenyang20 <zhaochenyang20@gmail.com> Co-authored-by: minleminzui <2969413251@qq.com> Co-authored-by: maocheng23 <maocheng@berkeley.edu> Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>	2025-08-25 15:26:26 -07:00
VDV1985	2c4b4b786b	[feature] Ascend NPU graph support (#9399 ) Co-authored-by: ronnie_zheng <zl19940307@163.com> Co-authored-by: yezhifeng (D) <y00897525@china.huawei.com> Co-authored-by: anon189Ty <Stari_Falcon@outlook.com> Co-authored-by: Maksim <makcum888e@mail.ru> Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>	2025-08-20 21:13:27 -07:00
Even Zhou	de2dd73831	Revert "[feature] Rework Ascend NPU graph support" (#9385 )	2025-08-20 00:35:10 -07:00
Even Zhou	3680d6f88b	[feature] Rework Ascend NPU graph support (#9350 ) Co-authored-by: ronnie_zheng <zl19940307@163.com> Co-authored-by: yezhifeng (D) <y00897525@china.huawei.com> Co-authored-by: anon189Ty <Stari_Falcon@outlook.com> Co-authored-by: Maksim <makcum888e@mail.ru> Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>	2025-08-19 20:32:27 -07:00
Even Zhou	f4fafacc5d	Revert "[feature] Ascend NPU graph support (#8027 )" (#9348 )	2025-08-19 10:11:23 -07:00
Mick	1df84ff414	ci: simplify multi-modality tests by using mixins (#9006 )	2025-08-16 22:25:02 -07:00
VDV1985	94371dbbd6	[feature] Ascend NPU graph support (#8027 ) Co-authored-by: ronnie_zheng <zl19940307@163.com> Co-authored-by: yezhifeng (D) <y00897525@china.huawei.com> Co-authored-by: anon189Ty <Stari_Falcon@outlook.com> Co-authored-by: Maksim <makcum888e@mail.ru> Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>	2025-08-16 17:25:17 -07:00
Hank Han	81da16f6d3	[CI] add deepseek w4a8 test on h20 ci (#7758 )	2025-08-16 01:54:13 -07:00
Hubert Lu	9c3e95d98b	[AMD] Expand test coverage for AMD CI and enable apply_token_bitmask_inplace_cuda in sgl-kernel (#8268 )	2025-08-15 12:32:51 -07:00
Hongbo Xu	2cc9eeab01	[4/n]decouple quantization implementation from vLLM dependency (#9191 ) Co-authored-by: AniZpZ <aniz1905@gmail.com> Co-authored-by: Yineng Zhang <me@zhyncs.com>	2025-08-14 12:05:46 -07:00
Stefan He	930fe467bd	Support Triton FP8 Gemm can handle hidden_dim not divisible by 16 (#9093 ) Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>	2025-08-12 21:21:55 -07:00
jacky.cheng	25caa7a8a9	[AMD] Support Wave attention backend with AMD GPU optimizations (#8660 ) Signed-off-by: Stanley Winata <stanley.winata@amd.com> Signed-off-by: Harsh Menon <harsh@nod-labs.com> Signed-off-by: nithinsubbiah <nithinsubbiah@gmail.com> Signed-off-by: Ivan Butygin <ivan.butygin@gmail.com> Signed-off-by: xintin <gaurav.verma@amd.com> Co-authored-by: Harsh Menon <harsh@nod-labs.com> Co-authored-by: Stanley Winata <stanley.winata@amd.com> Co-authored-by: Stanley Winata <68087699+raikonenfnu@users.noreply.github.com> Co-authored-by: Stanley Winata <stanley@nod-labs.com> Co-authored-by: Ivan Butygin <ivan.butygin@gmail.com> Co-authored-by: nithinsubbiah <nithinsubbiah@gmail.com> Co-authored-by: Nithin Meganathan <18070964+nithinsubbiah@users.noreply.github.com> Co-authored-by: Ivan Butygin <ibutygin@amd.com>	2025-08-12 13:49:11 -07:00
Baizhou Zhang	75e6a7cde1	Support radix cache for Lora feature (#7216 )	2025-08-11 10:14:11 -07:00
Lifu Huang	e322a94d1f	Reduce CI duration of test_lora_update. (#9024 )	2025-08-10 15:34:04 -07:00
Lianmin Zheng	2c7f01bc89	Reorganize CI and test files (#9027 )	2025-08-10 12:30:06 -07:00
Lianmin Zheng	ef48d5547e	Fix CI (#9013 )	2025-08-09 16:00:10 -07:00
fzyzcjy	442534aa44	Add CI for gpt-oss model on hopper (#8851 )	2025-08-09 00:34:23 -07:00
Lianmin Zheng	706bd69cc5	Clean up server_args.py to have a dedicated function for model specific adjustments (#8983 )	2025-08-08 19:56:50 -07:00
Lianmin Zheng	a947154286	Revert "Support Multi Process Tokenizer Manager" (#8960 )	2025-08-08 02:28:27 -07:00
ybyang	7490e3f67d	Support Multi Process Tokenizer Manager (#6555 ) Signed-off-by: ybyang <ybyang7@iflytek.com> Signed-off-by: huanglong <huanglong@linux.alibaba.com> Co-authored-by: lw9527 <952799980@qq.com> Co-authored-by: huanglong <huanglong@linux.alibaba.com> Co-authored-by: Huang Long <121648372+LLLL114@users.noreply.github.com>	2025-08-08 01:45:50 -07:00
fzyzcjy	b114a8105b	Support B200 in CI (#8861 )	2025-08-06 21:42:44 +08:00
Even Zhou	fee0ab0fba	[CI] Ascend NPU CI enhancement (#8294 ) Co-authored-by: ronnie_zheng <zl19940307@163.com>	2025-08-03 22:16:38 -07:00
harrisonlimh	747dd45077	feat: throttle requests at scheduler based on --max_queued_requests (#7565 )	2025-07-28 22:32:33 +08:00
Qiaolin Yu	2810338401	[feat] Support different attention backends for prefill and decode (#6338 ) Co-authored-by: tianqilin.99 <tianqilin.99@bytedance.com> Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>	2025-07-28 11:42:29 +08:00
Stefan He	ce32bc2ba9	Extract update_weights from RL Engine to SGLang to keep simplicity and fix torch reduce (#8267 ) Co-authored-by: CuiBo 82354186+SuperCB@users.noreply.github.com Co-authored-by: GeLee 865038696@qq.com Co-authored-by: 杨睿 yangruipis@163.com	2025-07-26 02:00:59 -07:00
Xinyuan Tong	38000a5f44	Fix gemma3n with hybrid swa (#8240 ) Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com> Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>	2025-07-23 13:29:18 -07:00
Lifu Huang	8abd3e77fe	Introduce Stable LoRA ID System for Overlapped Updates and Prefix Caching (#8261 )	2025-07-23 00:32:16 -07:00
Pavel Logachev	877e35d775	Add get_hidden_dim to qwen3.py for correct lora (#7312 )	2025-07-19 19:31:16 -07:00
Lifu Huang	4e3defe5a7	Support start up LoRA server without initial adapters (#8019 )	2025-07-19 15:38:09 -07:00
Lifu Huang	3de617a75b	Fix LoRA buffer contamination during adapter eviction (#8103 )	2025-07-19 13:14:08 -07:00
Lianmin Zheng	9c7a46180c	[Doc] Steps to add a new attention backend (#8155 )	2025-07-18 16:38:26 -07:00
Hubert Lu	7750b91ca8	[AMD] Add triton awq_dequantize kernel to support AWQ on ROCm (#7661 )	2025-07-18 14:27:25 -07:00
Zhiqiang Xie	9d33fcfb8e	Hicache Storage Layer Prototype (#7704 )	2025-07-18 15:20:19 +08:00
Lifu Huang	e2ed9d049a	Refactor dynamic LoRA update to fix incorrect handling of variant weight shapes (#7844 )	2025-07-13 18:36:01 -07:00
kyleliang-nv	dd445a41f5	[feature] Add start step profile argument in /start_profile (#7608 )	2025-07-09 18:42:15 -07:00
Cheng Wan	d487555f84	[CI] Add deepep tests to CI (#7872 )	2025-07-09 01:49:47 -07:00
Xinyuan Tong	136c6e0431	fix: Handles input_embeds in GenerateReqInput when n>1 (#7830 ) Signed-off-by: Xinyuan Tong <justinning0323@outlook.com>	2025-07-08 14:00:42 -07:00
Lifu Huang	2b0e1d1ce0	[Minor] Fix sporadic CI timeout caused by underestimated tests. (#7850 )	2025-07-08 01:01:49 -07:00
Hubert Lu	e00715eb66	[AMD] Add test_fused_moe.py and test_rope_rocm.py to AMD CI (#5246 )	2025-07-06 01:47:16 -07:00
YanbingJiang	4de0395343	Add V2-lite model test (#7390 ) Co-authored-by: DiweiSun <105627594+DiweiSun@users.noreply.github.com>	2025-07-03 22:25:50 -07:00
ronnie_zheng	1e0e549766	Ascend attention backend(PA&MLA) (#7722 ) Co-authored-by: Maksim <makcum888e@mail.ru> Co-authored-by: VDV1985 <vladdv85@mail.ru>	2025-07-03 09:23:19 -07:00
Hubert Lu	b116b21a46	[AMD] Temporarily disable test_no_overlap_scheduler and test_vision_chunked_prefill (#7717 )	2025-07-02 12:39:18 -07:00
Lianmin Zheng	22352d47a9	Improve streaming, log_level, memory report, weight loading, and benchmark script (#7632 ) Co-authored-by: Kan Wu <wukanustc@gmail.com>	2025-06-29 23:16:19 -07:00
Chunyuan WU	c5131f7a2f	[CPU] add c++ kernel to bind CPU cores and memory node (#7524 )	2025-06-29 19:45:25 -07:00
Xinyuan Tong	8f335b5bd6	Fix stream reasoning parser and Adds Kimi reasoning parser (#7432 ) Signed-off-by: Xinyuan Tong <justinning0323@outlook.com>	2025-06-29 14:39:05 -07:00
Yineng Zhang	a8c10aeeee	fix unit tests (#7618 )	2025-06-28 00:32:41 -07:00
Lifu Huang	49538d111b	Support dynamic LoRA loading / unloading in engine/server API (#7446 )	2025-06-27 21:00:27 -07:00

1 2 3 4 5

236 Commits