Garry Fang
|
60468da4e2
|
bugfix: fix sglang crash in NVIDIA MIG container (#8167)
Signed-off-by: Garrybest <garrybest@foxmail.com>
|
2025-07-19 14:41:27 -07:00 |
|
Simo Lin
|
41d33e4736
|
[router] add ut for worker and errors (#8170)
|
2025-07-19 14:38:33 -07:00 |
|
kyleliang-nv
|
bfdd226f35
|
Fix Dockerfile.gb200 (#8169)
|
2025-07-19 14:37:53 -07:00 |
|
Lifu Huang
|
3de617a75b
|
Fix LoRA buffer contamination during adapter eviction (#8103)
|
2025-07-19 13:14:08 -07:00 |
|
Lianmin Zheng
|
bb0e8a32b5
|
Clean up server args (#8161)
|
2025-07-19 11:32:52 -07:00 |
|
Lianmin Zheng
|
1b427dae02
|
Update README.md (#8171)
|
2025-07-19 11:04:19 -07:00 |
|
Charles Chen
|
f3d9736156
|
Fix suffix mismatch for the metrics. (#8168)
Signed-off-by: Charles Chen <chenliqian@chenliqian.cn>
|
2025-07-19 10:11:24 -07:00 |
|
Yineng Zhang
|
561dd7b2ce
|
chore: upgrade sgl-kernel 0.2.6 (#8166)
|
2025-07-19 03:17:08 -07:00 |
|
Yineng Zhang
|
f98e88b9fb
|
chore: bump sgl-kernel v0.2.6 (#8165)
|
2025-07-19 00:56:18 -07:00 |
|
Cheng Wan
|
15ad6c9086
|
[1/N] MoE Refactor: refactor select_experts (#7966)
|
2025-07-19 00:51:15 -07:00 |
|
kyleliang-nv
|
cfab0ff6e2
|
Add GB200 wide-EP docker (#8157)
|
2025-07-18 22:34:29 -07:00 |
|
Simo Lin
|
b763cf7e8e
|
[router] allow router to have empty workers (#8160)
|
2025-07-18 22:09:54 -07:00 |
|
Simo Lin
|
8fcc55cfa1
|
[router] router metrics cleanup (#8158)
|
2025-07-18 22:09:17 -07:00 |
|
Yingchun Lai
|
610381b75e
|
[health_generate] fix: fix the /health_generate always success bug (#8028)
|
2025-07-18 22:08:46 -07:00 |
|
Shangming Cai
|
1403ea5694
|
[PD] Support non-MLA models PD different TP with DP attention (#7931)
Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com>
|
2025-07-18 22:00:49 -07:00 |
|
Binyao Jiang
|
b7e951a6db
|
Feat: Support audio in Phi4-mm model (#8048)
|
2025-07-18 21:03:53 -07:00 |
|
Haohui Mai
|
d918ab7985
|
Support NVFP4 quantized dense models on AMD CDNA2/CDNA3 GPUs (#7302)
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: Sai Enduri <saimanas.enduri@amd.com>
|
2025-07-18 19:59:39 -07:00 |
|
Mick
|
3964b352c3
|
chore: tune mem fraction static for vlm (#6881)
|
2025-07-18 17:19:27 -07:00 |
|
Lianmin Zheng
|
9c7a46180c
|
[Doc] Steps to add a new attention backend (#8155)
|
2025-07-18 16:38:26 -07:00 |
|
Hubert Lu
|
7750b91ca8
|
[AMD] Add triton awq_dequantize kernel to support AWQ on ROCm (#7661)
|
2025-07-18 14:27:25 -07:00 |
|
Simo Lin
|
c8f31042a8
|
[router] Refactor router and policy traits with dependency injection (#7987)
Co-authored-by: Jin Pan <jpan236@wisc.edu>
Co-authored-by: Keru Yang <rukeyang@gmail.com>
Co-authored-by: Yingyi Huang <yingyihuang2000@outlook.com>
Co-authored-by: Philip Zhu <phlipzhux@gmail.com>
|
2025-07-18 14:24:24 -07:00 |
|
Hongbo Xu
|
1f76fc8747
|
[3/n] chore: decouple AWQ implementation from vLLM dependency (#8113)
Co-authored-by: AniZpZ <zhuangsen.zp@antgroup.com>
|
2025-07-18 11:45:22 -07:00 |
|
Even Zhou
|
6737671c82
|
[Bugfix] Fix w8a8_int8 import error on NPU (#8147)
|
2025-07-18 11:34:55 -07:00 |
|
Enrique Shockwave
|
fd63b62eaa
|
fix compressed tensors WNA16 imports (#8142)
|
2025-07-18 11:34:14 -07:00 |
|
Peng Zhang
|
719b29f218
|
feat: enchance green context stream creation robust with backward compatibility (#8136)
|
2025-07-18 02:45:03 -07:00 |
|
Sai Enduri
|
d0510f08fe
|
Revert "Fix different device type adjustment in PP" (#8141)
|
2025-07-18 01:12:11 -07:00 |
|
Zhiqiang Xie
|
9d33fcfb8e
|
Hicache Storage Layer Prototype (#7704)
|
2025-07-18 15:20:19 +08:00 |
|
jianan-gu
|
7891bac16b
|
[Quantization][w8a8_int8] Fix weight loading issue for w8a8_int8 path with "ignore" layer list in quantization config (#7820)
|
2025-07-17 22:03:56 -07:00 |
|
jianan-gu
|
48c1fa7bb6
|
[CPU][Llama4] Fix Llama4 MoE inputs with "apply_router_weight_on_input" (#7889)
|
2025-07-17 21:43:25 -07:00 |
|
yilian49
|
8aa5ae6b04
|
load draft model fix (#7506)
|
2025-07-17 21:41:32 -07:00 |
|
Minglei Zhu
|
8a32355704
|
Feat: Support Granite 3.0 MoE in SGLang (#7959)
|
2025-07-17 20:56:03 -07:00 |
|
Qi Yuhang
|
6e92da8fca
|
[Fix][Ready]Fix register spilling in cutlass nvfp4 gemm kernel on Blackwell (#8127)
|
2025-07-17 20:49:36 -07:00 |
|
Mick
|
e1020dc588
|
refactor: simply MultimodalTokens logic (#7924)
|
2025-07-17 17:59:15 -07:00 |
|
Zhao Chen
|
3586b4cef2
|
feat: add production metric for retracted requests due to insufficient kvcache (#7030)
Signed-off-by: Zhao Chen <zhaochen.zju@gmail.com>
|
2025-07-17 11:59:05 -07:00 |
|
Asher
|
4296021499
|
[Hunyuan]: Fix Dense Model Support (#8117)
Signed-off-by: Asher Zhang <asherszhang@tencent.com>
|
2025-07-17 10:00:11 -07:00 |
|
Ziqi Fan
|
01857fab61
|
fix: update HostKVCache init to report correct msg when available memory is not enough (#8102)
|
2025-07-17 21:24:34 +08:00 |
|
fzyzcjy
|
519ff5c8e6
|
Super tiny fix typo (#8046)
|
2025-07-17 21:15:51 +08:00 |
|
Yuan Luo
|
af1cc8fe2d
|
[kernel] opt moe align block kernel by block/warp scan algorithm (#7884)
|
2025-07-17 19:33:02 +08:00 |
|
Cheng Wan
|
49b8777460
|
Refactor: move all quantization-related code to srt/layer/quantization (#7989)
|
2025-07-17 00:47:07 -07:00 |
|
Cheng Wan
|
02404a1e35
|
[ci] recover 8-gpu deepep test (#8105)
|
2025-07-17 00:46:40 -07:00 |
|
hzh0425
|
5c08a36cbf
|
[Fix] ensure DeepGEMM is only enabled for FP8_W8A8 models (#8110)
|
2025-07-16 21:33:29 -07:00 |
|
Cheng Wan
|
9069884b51
|
[ci] disable memory imbalance check for draft worker (#8108)
|
2025-07-16 20:41:47 -07:00 |
|
Simo Lin
|
8a7a7770e5
|
[ci] limit cmake build nproc (#8100)
|
2025-07-16 18:09:28 -07:00 |
|
Yingchun Lai
|
795668dc73
|
feat: add tp_rank, pp_rank and dp_rank labels for scheduler metrics (#7597)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
|
2025-07-16 17:55:59 -07:00 |
|
Mick
|
4395c87a9b
|
refactor: unify names of the feature field of MultimodalDataItem (#8075)
|
2025-07-16 17:52:38 -07:00 |
|
Peng Zhang
|
c28ad1990d
|
[1/n] chore: decouple quantization implementation from vLLM dependency (#7992)
|
2025-07-16 15:56:26 -07:00 |
|
Xiaoze Fan
|
570d33437b
|
[Feature] Layer-wise Prefill (#7634)
Signed-off-by: jason-fxz <jason341132@qq.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-07-17 01:57:46 +08:00 |
|
Simo Lin
|
d9eb5efc71
|
[misc] update nvshmem and pin deepEP commit hash (#8098)
|
2025-07-16 08:54:55 -07:00 |
|
Peng Zhang
|
6dc4af4937
|
fix greenctx stream compability (#8090)
|
2025-07-16 07:08:46 -07:00 |
|
YanbingJiang
|
b188a89a5d
|
Fix CI xeon test with triton 3.3.1 (#8086)
|
2025-07-16 02:12:23 -07:00 |
|