hzh0425
|
7ec5b4e89c
|
[PD-HiCache]: Support Async Offloading KVCache In Decode Side (#10192)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2025-09-25 23:20:49 -07:00 |
|
Xinyuan Tong
|
6088548216
|
Update CODEOWNERS to include JustinTong0323 in FC (#10939)
|
2025-09-25 22:55:56 -07:00 |
|
Yineng Zhang
|
172bcf0152
|
Revert "Refactor kv_cache_scheme handling for quantization (#10132)" (#10935)
|
2025-09-25 20:14:15 -07:00 |
|
Chang Su
|
37158f2018
|
router: Support parallel sampling num > 1 in grpc_server and non-stream handling (#10929)
|
2025-09-25 20:03:35 -07:00 |
|
Lianmin Zheng
|
3e95aa1a09
|
Remove pull_request trigger from CI monitor workflow (#10932)
|
2025-09-25 19:40:38 -07:00 |
|
Xiaoyu Zhang
|
c4197e99bb
|
[ci] add ci-monitor workflow (#10898)
|
2025-09-25 19:29:47 -07:00 |
|
eraser00
|
0ac6114694
|
Replace the Kimi-K2 generated tool call idx with history tool call count (#10612)
Co-authored-by: eraser00 <eraser00@github.com>
|
2025-09-25 18:47:40 -07:00 |
|
Chang Su
|
7dcd689b47
|
[router][refactor] Clean up protobuf fields (#10923)
|
2025-09-25 17:48:47 -07:00 |
|
Simo Lin
|
f7bab41a29
|
[router] change log level to warning (#10926)
|
2025-09-25 17:32:59 -07:00 |
|
Lianmin Zheng
|
f68dd998b9
|
Rename customer label -> custom label (#10899)
Co-authored-by: Yingchun Lai <laiyingchun@apache.org>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-09-25 16:19:53 -07:00 |
|
Lianmin Zheng
|
35ec2a45a8
|
[minor] Remove deprecated function get_ip (#10883)
|
2025-09-25 16:18:04 -07:00 |
|
Swipe4057
|
0035f1cefa
|
fix env flashinfer (#10910)
|
2025-09-25 15:44:48 -07:00 |
|
Chang Su
|
5e21d6aec0
|
refactor: Move grpc/client.rs to grpc_client/sglang_scheduler.rs (#10924)
|
2025-09-25 17:21:22 -04:00 |
|
Mohammad Miadh Angkad
|
cd4da1f19b
|
Refactor kv_cache_scheme handling for quantization (#10132)
|
2025-09-25 10:32:15 -07:00 |
|
Chang Su
|
916784746b
|
router: Fix constraint proto and build_constraint in grpc router (#10881)
|
2025-09-25 11:12:06 -04:00 |
|
Simo Lin
|
d511b2d905
|
[router] consolidate worker load monitoring (#10894)
|
2025-09-25 09:59:30 -04:00 |
|
lukec
|
77830a265e
|
Add fuse_moe per-channel tune (#10915)
|
2025-09-25 21:12:09 +08:00 |
|
yi wang
|
fce170480a
|
integrate AIBrix KVcache (#10376)
|
2025-09-25 14:47:09 +08:00 |
|
Zhiqiang Xie
|
3d40794fcf
|
[HiCache] Cleaning the deprecated host memory state (#10778)
|
2025-09-25 14:43:53 +08:00 |
|
Xiaoyu Zhang
|
c1f39013b7
|
[ci feature] add ci monitor (#10872)
|
2025-09-24 23:16:29 -07:00 |
|
Lianmin Zheng
|
3e43eb137b
|
[Auto Sync] Update model_config.py (20250925) (#10885)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
|
2025-09-24 22:59:16 -07:00 |
|
Simo Lin
|
458c0219a6
|
[router] simplify tokenizer dev doc (#10895)
|
2025-09-24 22:15:56 -07:00 |
|
Keyang Ru
|
a73eb8cd20
|
[router] Support Oracle DB(ATP) Data Connector (#10845)
|
2025-09-24 23:59:32 -04:00 |
|
Simo Lin
|
e738703547
|
[router] consolidate worker get loads (#10880)
|
2025-09-24 22:13:31 -04:00 |
|
Yuhao Yao
|
fe531d6f4e
|
[Bug] Fix Issue#10215 (#10572)
|
2025-09-25 09:51:50 +08:00 |
|
Xiaoyu Zhang
|
c4e314f986
|
Restruct sgl-kernel benchmark (#10861)
|
2025-09-25 07:45:25 +08:00 |
|
Simo Lin
|
7a06ef984d
|
[router] consolidate health endpoints and flush cache (#10876)
|
2025-09-24 15:23:21 -07:00 |
|
Chang Su
|
4a87ba217f
|
router-grpc: Add tools processing and other paramters for apply_chat_template (#10877)
|
2025-09-24 15:23:06 -07:00 |
|
kushanam
|
d7b20dd65d
|
chore: Initial support for input config files (#10534)
Co-authored-by: root <root@umbriel-b200-017.ipp4a1.colossus.nvidia.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
|
2025-09-24 14:45:52 -07:00 |
|
luna
|
c3faf2d6e6
|
[router] select first healthy worker on proxied get requests (#10827)
|
2025-09-24 11:45:41 -07:00 |
|
Chang Su
|
9209b209be
|
router-grpc: Support jinja chat template content format detection (#10832)
|
2025-09-24 11:45:01 -07:00 |
|
ishandhanani
|
adba172fd1
|
ci: free space on workers for build (#10786)
Co-authored-by: zhyncs <me@zhyncs.com>
|
2025-09-24 02:58:22 -07:00 |
|
GuoweiWangU
|
cd641a995c
|
fix bailing_moe with enable_dp_attention (#10860)
|
2025-09-24 02:29:32 -07:00 |
|
Xinyuan Tong
|
71f24ef8f6
|
feat: add cache_salt support to request (#10718)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-09-23 23:30:25 -07:00 |
|
Lianmin Zheng
|
b1f0fc1c0b
|
Add CI timeout guidelines (#10829)
|
2025-09-23 22:08:02 -07:00 |
|
Lianmin Zheng
|
32d893730f
|
Revert "[fix][pd-disag]no need set next batch sampling info done in prefill" (#10828)
|
2025-09-23 17:02:01 -07:00 |
|
Lianmin Zheng
|
f47a2c67e6
|
[Auto Sync] Update load_config.py, model_config.py, configu... (20250923) (#10825)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2025-09-23 16:48:12 -07:00 |
|
Chang Su
|
ee704e6265
|
[router] add auth middleware for api key auth (#10826)
|
2025-09-23 16:07:34 -07:00 |
|
Keyang Ru
|
f4e3ebeb05
|
[router] Support streaming for Openai Router Response api (#10822)
|
2025-09-23 14:56:28 -07:00 |
|
Lianmin Zheng
|
312bfc4c95
|
[Auto Sync] Update simple_eval_common.py (20250923) (#10824)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: cctry <shiyang@x.ai>
|
2025-09-23 13:50:47 -07:00 |
|
Lianmin Zheng
|
e290303ea1
|
[Auto Sync] Update elementwise.py (20250923) (#10823)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2025-09-23 13:50:22 -07:00 |
|
Xinyuan Tong
|
aab35bccb4
|
fix: draft model IMA by overide max_positional_embeddings (#10787)
Co-authored-by: Qiaolin Yu <qy254@cornell.edu>
|
2025-09-23 12:56:16 -07:00 |
|
Yineng Zhang
|
42aedb02af
|
[Auto Sync] Update protocol.py (20250923) (#10820)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: cicirori <32845984+cicirori@users.noreply.github.com>
|
2025-09-23 12:49:56 -07:00 |
|
Yiakwy
|
984730b732
|
add tunning files for QWEN-3-NEXT (#10794)
|
2025-09-23 12:46:30 -07:00 |
|
Shangming Cai
|
23632d350c
|
Fix latest main ci (#10799)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2025-09-23 12:46:13 -07:00 |
|
Chang Su
|
08b8c0c3cd
|
[router] fix axum default body limit (#10818)
|
2025-09-23 12:44:17 -07:00 |
|
Lzhang-hub
|
d42975c641
|
Remove duplicate code in qwen2 model (#10540)
|
2025-09-24 02:40:51 +08:00 |
|
ZhengHSI
|
adc24a3a0c
|
fix ceval (#10504)
Co-authored-by: lukec <118525388+sleepcoo@users.noreply.github.com>
|
2025-09-24 02:35:25 +08:00 |
|
Chang Su
|
7ff93e613f
|
router(grpc): Implement route for chat_cmpl endpoint (#10761)
|
2025-09-23 11:26:33 -07:00 |
|
Simo Lin
|
b24b2e7ed7
|
[router] use dashmap for radix tree instead of hash for multi model (#10814)
|
2025-09-23 11:25:53 -07:00 |
|