Shangming Cai
|
9c339d6b47
|
[PD] Extract the PP transfer layer calculate logic from Mooncake to Common backend (#10565)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2025-09-28 00:10:41 +08:00 |
|
Shangming Cai
|
e23e280e16
|
Add support for topk metadata transferring for PD (#10616)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2025-09-28 00:09:38 +08:00 |
|
Muqi Li
|
51f7c6bd3c
|
Add auth to get server info (#10751)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2025-09-27 02:54:39 -07:00 |
|
Xinyuan Tong
|
62e2e99db6
|
fix: make inference deterministic for large TP (#10930)
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Yuan Luo <yuan.luo@hotmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-09-27 02:46:45 -07:00 |
|
kk
|
8ebf72fef3
|
[Fix] RuntimeError: get_cfg Unsupported input_type:Float4_e2m1fn_x2 in using aiter-mxfp4-moe (#10981)
Co-authored-by: wunhuang <wunhuang@amd.com>
|
2025-09-26 22:12:22 -07:00 |
|
Yueyang Pan
|
8260574729
|
fix: fp8 quantization failure of qwen 2.5 VL 7B model (#10112)
Signed-off-by: PanJason <pyyjason@gmail.com>
|
2025-09-27 05:05:23 +00:00 |
|
Chang Su
|
37f3325b06
|
[router][grpc] Support E2E non-stream chat completions (#10980)
|
2025-09-26 22:02:06 -07:00 |
|
Muqi Li
|
bd95944cf6
|
[Bugfix][Minor][Benchmark] Fix some bugs due to PR #10495 (#10982)
|
2025-09-26 22:01:05 -07:00 |
|
hzh0425
|
c8a5d12abe
|
[HiCache]: Support dynamic loading backends for hicache (#10551)
Co-authored-by: Teng Ma <sima.mt@alibaba-inc.com>
|
2025-09-26 18:34:11 -07:00 |
|
Xiaoyu Zhang
|
2387c22b56
|
Ci monitor support performance (#10965)
|
2025-09-27 09:11:21 +08:00 |
|
hlu1
|
592ddf374f
|
Add simple docker file for B300 (#10944)
Signed-off-by: Hao Lu <14827759+hlu1@users.noreply.github.com>
|
2025-09-26 17:26:57 -07:00 |
|
Chang Su
|
0c3db88978
|
[router][grpc] Add helpfer functions for decoder in router.rs and fix specs (#10971)
|
2025-09-26 20:10:45 -04:00 |
|
amysaq2023
|
2bdaf482f9
|
refactor loading weights from remote instance coding format (#10941)
Signed-off-by: Anqi Shen <amy.saq@antgroup.com>
|
2025-09-26 15:25:39 -07:00 |
|
Mick
|
777eb53897
|
ci: refactor nightly test (#10495)
|
2025-09-26 15:24:30 -07:00 |
|
Xiaoyu Zhang
|
05a3526654
|
Restruct gpu_memory_settings in a unify function and relax max_cuda_graph_bs (#10372)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: sglang-bot <sglangbot@gmail.com>
|
2025-09-26 15:10:49 -07:00 |
|
Lianmin Zheng
|
e56c64bfaf
|
Update label field comment to indicate deprecation (#10970)
|
2025-09-26 12:59:59 -07:00 |
|
Mick
|
fff7fbabe6
|
ci: fix rate-limit of huggingface with hf auth login (#10947)
|
2025-09-26 11:02:44 -07:00 |
|
Simo Lin
|
aae7ead2d0
|
[router] remove old/oudated/useless comments across code base (#10968)
|
2025-09-26 10:48:50 -07:00 |
|
Simo Lin
|
a7fe6e10a1
|
[router] remove old/oudated/useless comments (#10967)
|
2025-09-26 09:45:15 -07:00 |
|
Simo Lin
|
be059b83d6
|
[router] grpc router regular mode import cleanup (#10963)
|
2025-09-26 04:06:59 -07:00 |
|
Simo Lin
|
5d4fe1ceee
|
[router] add move grpc worker management from router to worker manager (#10960)
|
2025-09-26 03:57:57 -07:00 |
|
Simo Lin
|
1b011e68dc
|
[router] move grpc client from router to worker and builder (#10958)
|
2025-09-26 03:13:47 -07:00 |
|
Simo Lin
|
5c0efa562b
|
[router]fix code owner syntax error (#10956)
|
2025-09-26 03:07:18 -07:00 |
|
Simo Lin
|
1e57b9472d
|
[router] add grpc client get and set (#10955)
|
2025-09-26 03:07:05 -07:00 |
|
Yuan Luo
|
a5095d6262
|
Fuse write kv buffer into rope for qwen3 moe & bailing moe (#10749)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-09-26 15:18:41 +08:00 |
|
Simo Lin
|
6c2c467d77
|
[router] update owners for router components (#10927)
|
2025-09-25 23:46:46 -07:00 |
|
Sahithi Chigurupati
|
c3d2ad4ee6
|
CI: Fix docker manifest build (#10936)
|
2025-09-25 23:22:55 -07:00 |
|
hzh0425
|
7ec5b4e89c
|
[PD-HiCache]: Support Async Offloading KVCache In Decode Side (#10192)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2025-09-25 23:20:49 -07:00 |
|
Xinyuan Tong
|
6088548216
|
Update CODEOWNERS to include JustinTong0323 in FC (#10939)
|
2025-09-25 22:55:56 -07:00 |
|
Yineng Zhang
|
172bcf0152
|
Revert "Refactor kv_cache_scheme handling for quantization (#10132)" (#10935)
|
2025-09-25 20:14:15 -07:00 |
|
Chang Su
|
37158f2018
|
router: Support parallel sampling num > 1 in grpc_server and non-stream handling (#10929)
|
2025-09-25 20:03:35 -07:00 |
|
Lianmin Zheng
|
3e95aa1a09
|
Remove pull_request trigger from CI monitor workflow (#10932)
|
2025-09-25 19:40:38 -07:00 |
|
Xiaoyu Zhang
|
c4197e99bb
|
[ci] add ci-monitor workflow (#10898)
|
2025-09-25 19:29:47 -07:00 |
|
eraser00
|
0ac6114694
|
Replace the Kimi-K2 generated tool call idx with history tool call count (#10612)
Co-authored-by: eraser00 <eraser00@github.com>
|
2025-09-25 18:47:40 -07:00 |
|
Chang Su
|
7dcd689b47
|
[router][refactor] Clean up protobuf fields (#10923)
|
2025-09-25 17:48:47 -07:00 |
|
Simo Lin
|
f7bab41a29
|
[router] change log level to warning (#10926)
|
2025-09-25 17:32:59 -07:00 |
|
Lianmin Zheng
|
f68dd998b9
|
Rename customer label -> custom label (#10899)
Co-authored-by: Yingchun Lai <laiyingchun@apache.org>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-09-25 16:19:53 -07:00 |
|
Lianmin Zheng
|
35ec2a45a8
|
[minor] Remove deprecated function get_ip (#10883)
|
2025-09-25 16:18:04 -07:00 |
|
Swipe4057
|
0035f1cefa
|
fix env flashinfer (#10910)
|
2025-09-25 15:44:48 -07:00 |
|
Chang Su
|
5e21d6aec0
|
refactor: Move grpc/client.rs to grpc_client/sglang_scheduler.rs (#10924)
|
2025-09-25 17:21:22 -04:00 |
|
Mohammad Miadh Angkad
|
cd4da1f19b
|
Refactor kv_cache_scheme handling for quantization (#10132)
|
2025-09-25 10:32:15 -07:00 |
|
Chang Su
|
916784746b
|
router: Fix constraint proto and build_constraint in grpc router (#10881)
|
2025-09-25 11:12:06 -04:00 |
|
Simo Lin
|
d511b2d905
|
[router] consolidate worker load monitoring (#10894)
|
2025-09-25 09:59:30 -04:00 |
|
lukec
|
77830a265e
|
Add fuse_moe per-channel tune (#10915)
|
2025-09-25 21:12:09 +08:00 |
|
yi wang
|
fce170480a
|
integrate AIBrix KVcache (#10376)
|
2025-09-25 14:47:09 +08:00 |
|
Zhiqiang Xie
|
3d40794fcf
|
[HiCache] Cleaning the deprecated host memory state (#10778)
|
2025-09-25 14:43:53 +08:00 |
|
Xiaoyu Zhang
|
c1f39013b7
|
[ci feature] add ci monitor (#10872)
|
2025-09-24 23:16:29 -07:00 |
|
Lianmin Zheng
|
3e43eb137b
|
[Auto Sync] Update model_config.py (20250925) (#10885)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
|
2025-09-24 22:59:16 -07:00 |
|
Simo Lin
|
458c0219a6
|
[router] simplify tokenizer dev doc (#10895)
|
2025-09-24 22:15:56 -07:00 |
|
Keyang Ru
|
a73eb8cd20
|
[router] Support Oracle DB(ATP) Data Connector (#10845)
|
2025-09-24 23:59:32 -04:00 |
|