Commit Graph

3127 Commits

Author SHA1 Message Date
Chayenne
73dcf2b326 Remove token in token out in Native API (#5967) 2025-05-01 21:59:43 -07:00
Chang Su
170d1f218a feat: Refactor DeepSeekV3 function call (#5908) 2025-05-01 21:28:57 -07:00
Sai Enduri
73bc1d00fc Add 1 gpu perf and 2 gpu accuracy tests for AMD MI300x CI. (#5960) 2025-05-01 20:56:59 -07:00
XinyuanTong
c5645e928f feat: add concurrency evaluation logic in mmmu benchmark (#5782) 2025-05-01 18:20:08 -07:00
KCFindstr
d33955d28a Properly return error response in vertex_generate HTTP endpoint (#5956) 2025-05-01 11:48:58 -07:00
Stefan He
6fc175968c Optimize a pad operation to accelerate 25us (#5945) 2025-05-01 10:48:55 -07:00
江家瑋
ad506a4e6b docs: Fix Qwen model typo (#5944)
Signed-off-by: JiangJiaWei1103 <waynechuang97@gmail.com>
2025-05-01 10:23:00 -07:00
Ke Bao
ebaba85655 Update ci test and doc for MTP api change (#5952) 2025-05-01 09:30:27 -07:00
Ke Bao
de2faef97e Remove extra contiguous (#5953) 2025-05-01 09:28:46 -07:00
Yuan Luo
67b7d5b1df [PD] Vectorise group_concurrent_contiguous in NumPy (#5834)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2025-05-01 22:42:37 +08:00
ryang
4322c31e24 Support XiaomiMiMo/MiMo model inference (#5921) 2025-05-01 07:41:13 -07:00
Yineng Zhang
9858113c33 chore: bump v0.4.6.post2 (#5939) 2025-04-30 22:04:40 -07:00
Yineng Zhang
8441baad6e fix: update model runner (#5934) 2025-04-30 19:49:26 -07:00
mlmz
256c4c2519 fix: correct stream response when enable_thinking is set to false (#5881) 2025-04-30 19:44:37 -07:00
Johnny
9f21e75453 add Thor & Spark (#5915) 2025-04-30 19:43:40 -07:00
Qiaolin Yu
7bcd8b1cb2 Fix lora batch processing when input lora_path contains None (#5930) 2025-04-30 19:42:42 -07:00
Ying Sheng
11383cec3c [PP] Add pipeline parallelism (#5724) 2025-04-30 18:18:07 -07:00
XinyuanTong
e97e57e699 Remove unused method calculate_num_image_tokens from qwen2_vl.py (#5783) 2025-04-30 17:46:59 -07:00
Yineng Zhang
9a6ad8916d chore: upgrade sgl-kernel 0.1.1 (#5933) 2025-04-30 16:13:30 -07:00
Yineng Zhang
d353d08b4e chore: bump sgl-kernel 0.1.1 (#5932) 2025-04-30 14:01:49 -07:00
PGFLMG
08acdb5c3d [Feat] Scale up fa3 kernel to sm8x arch (#5912)
Co-authored-by: zhyncs <me@zhyncs.com>
2025-04-30 13:59:36 -07:00
Sai Enduri
2afba1b1c1 Add TP2 MOE benchmarks for AMD. (#5909) 2025-04-30 11:38:20 -07:00
laixin
e330f2b86c [qwen3] support qwen3 ep moe (#5917)
Co-authored-by: sleepcoo <sleepcoo@gmail.com>
2025-04-30 09:15:21 -07:00
PGFLMG
3ddf5b9d61 [Misc] use parallel build for cmake in sgl-kernel (#5919) 2025-04-30 08:56:46 -07:00
JieXin Liang
3cff963335 [fix] kimi-vl test in test_vision_openai_server.py (#5910) 2025-04-29 23:59:10 -07:00
Yi Zhang
d50e36a79d support vlm benchmark profile (#5905) 2025-04-29 23:48:27 -07:00
liwenju0
8fefdd32c7 [Feature] add support kimi vl model (#5383)
Co-authored-by: wenju.li <wenju.li@deepctr.cn>
2025-04-29 21:31:19 -07:00
zhjunqin
403b855a22 Add sm_120 for blackwell (#5903) 2025-04-29 20:45:24 -07:00
lambert0312
1698e94e67 Add A800 fused moe config for qwen3 235b (#5900) 2025-04-29 20:18:11 -07:00
Qiaolin Yu
58195dd588 [Fix] Unload lora in HF_Runner if needed (#5899) 2025-04-29 20:17:42 -07:00
Baizhou Zhang
799789afed Bump Flashinfer to 0.2.5 (#5870)
Co-authored-by: Yuhao Chen <yxckeis8@gmail.com>
2025-04-29 19:50:57 -07:00
ybyang
cc4a80caf6 [PD] Fix Assertion failed: /DeepEP/csrc/kernels/internode.cu:483, condition: ibgda_get_state()->num_rc_per_pe >= num_channels #134 (#5830) 2025-04-29 19:38:54 -07:00
lambert0312
3c8a52311a Fix check_env script (#5901) 2025-04-29 18:54:54 -07:00
Yineng Zhang
a043f7f2ab chore: use torch 2.6 for sgl-kernel build (#5898) 2025-04-29 17:51:18 -07:00
saienduri
e3a5304475 Add AMD MI300x Nightly Testing. (#5861) 2025-04-29 17:34:32 -07:00
Chang Su
28b26dbf48 [Bugfix]: fix missing queue_time_start for requests from grammar_queue (#5696) 2025-04-29 17:31:44 -07:00
Chang Su
2b06484bd1 feat: support pythonic tool call and index in tool call streaming (#5725) 2025-04-29 17:30:44 -07:00
JieXin Liang
e4b6133b78 [fix] relax mem_fraction_static for h200 (#5893)
Co-authored-by: alcanerian <alcanerian@gmail.com>
2025-04-29 17:01:12 -07:00
Ke Bao
dd408ee481 Auto set draft model path for MTP (#5793) 2025-04-29 16:25:40 -07:00
Chang Su
9419e75d60 [CI] Add test_function_calling.py to run_suite.py (#5896) 2025-04-29 15:54:53 -07:00
Johnny
2c7dbb7cc2 [FEATURE] Enhance platform compatibility for ARM (#5746) 2025-04-29 15:06:16 -07:00
Yineng Zhang
9a62191ba7 chore: update CODEOWNERS (#5895) 2025-04-29 14:12:04 -07:00
simveit
ae523675e5 [Doc] Tables instead of bulletpoints for sampling doc (#5841) 2025-04-29 13:49:39 -07:00
Adarsh Shirawalmath
5c08aa4958 [Docs] Update docs for Qwen3 and Qwen3MoE (#5836) 2025-04-29 13:48:30 -07:00
Yineng Zhang
f4c191a712 chore: update Dockerfile (#5894) 2025-04-29 12:55:13 -07:00
Simo Lin
771669cbe0 [fix]: PyO3 macOS linking and consolidate on tracing for logging 2025-04-29 11:26:38 -07:00
Simo Lin
1468769bde [Misc] add service discovery for sgl router 2025-04-29 10:21:19 -07:00
lambert0312
91dda4cd06 Add A800 fused moe config for qwen3 30b (#5880) 2025-04-29 02:02:24 -07:00
pengcuo
8e5a6d3441 [Fix] Fix a bug for flashmla to run R1 model (#5875)
Co-authored-by: pengcuo <dgpengcuo@gmail.com>
2025-04-29 01:03:13 -07:00
XinyuanTong
8465f035d1 Add qwen3 30b fused moe config (#5859) 2025-04-29 00:24:00 -07:00