Ximingwang-09
|
f2a75a66c4
|
update doc (#7046)
Co-authored-by: ximing.wxm <ximing.wxm@antgroup.com>
|
2025-06-11 10:02:01 +08:00 |
|
Lianmin Zheng
|
6b12d6a8d5
|
Simplify the heuristics for setting --mem-fraction-static (#7054)
|
2025-06-10 19:01:39 -07:00 |
|
Lianmin Zheng
|
0f218731e3
|
Do not run frontend_reasoning.ipynb to reduce the CI load (#7073)
|
2025-06-10 17:15:31 -07:00 |
|
Yudi Xue
|
14c18d25df
|
Frontend language separate reasoning support (#6031)
|
2025-06-10 17:11:29 -07:00 |
|
Lianmin Zheng
|
90bd3e32d6
|
Improve perf tuning docs (#7071)
|
2025-06-10 16:55:04 -07:00 |
|
Brayden Zhong
|
ca9291181d
|
[Feature] Add Logit Bias (#6579)
Co-authored-by: Cinjon Resnick <cinjon.resnick@gmail.com>
|
2025-06-10 15:39:25 -07:00 |
|
Yineng Zhang
|
344adb00ec
|
fix arm sgl-kernel link issue (#7066)
|
2025-06-10 15:23:22 -07:00 |
|
kyle-pena-kuzco
|
b56de8f943
|
Open AI API hidden states (#6716)
|
2025-06-10 14:37:29 -07:00 |
|
Yineng Zhang
|
ce5ee3bdf0
|
chore: update dev docker (#7064)
|
2025-06-10 12:57:58 -07:00 |
|
Xu Wenqing
|
a0e4d4eb53
|
Fix missing tool call id if tool call index >0 in streaming tool call output. (#7049)
Signed-off-by: 许文卿 <xwq391974@alibaba-inc.com>
|
2025-06-10 12:56:43 -07:00 |
|
Yineng Zhang
|
2f58445531
|
Revert "Add sanity checks when a test file is not added to CI (#6947)" (#7063)
|
2025-06-10 12:43:25 -07:00 |
|
fzyzcjy
|
fe55947acd
|
Add sanity checks when a test file is not added to CI (#6947)
|
2025-06-10 12:34:57 -07:00 |
|
fzyzcjy
|
19995dd78e
|
Tiny fix cutlass_mla_get_workspace_size stub incorrect signature (#7057)
|
2025-06-10 12:27:57 -07:00 |
|
Baizhou Zhang
|
3b014bc13d
|
Fix test_lora.py CI (#7061)
|
2025-06-10 12:24:46 -07:00 |
|
Hank Han
|
d7c3e8e93d
|
[fix] libmlx5.so already in base image (#7060)
|
2025-06-10 12:22:24 -07:00 |
|
kk
|
8ea7df6114
|
[WA] fix output data is nan in CI test "test_moe_eval_accuracy_large.py" (#7021)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2025-06-10 16:08:10 +00:00 |
|
Lianmin Zheng
|
4a102a2b02
|
Minor style fix in cuda_graph_runner.py (#7053)
|
2025-06-10 06:32:41 -07:00 |
|
Lianmin Zheng
|
6406408a70
|
Clean up server_args.py (#7037)
|
2025-06-10 05:34:29 -07:00 |
|
Lianmin Zheng
|
019851d099
|
Fix eagle on AMD (#7051)
|
2025-06-10 05:22:40 -07:00 |
|
Lianmin Zheng
|
2dae104dca
|
Minor cleanup of fa3 backend (#6999)
|
2025-06-10 03:58:44 -07:00 |
|
Yineng Zhang
|
cef6655b26
|
fix 24.12 docker (#7045)
|
2025-06-10 02:43:22 -07:00 |
|
Swipe4057
|
27196d4148
|
[Docker] Upgrading base image from 24.04 to 24.12 (#7043)
|
2025-06-10 02:23:05 -07:00 |
|
Lianmin Zheng
|
bb185b0e92
|
Update README.md (#7040)
|
2025-06-10 01:59:14 -07:00 |
|
Yineng Zhang
|
4f723edd3b
|
chore: bump v0.4.7 (#7038)
|
2025-06-10 01:56:20 -07:00 |
|
YanbingJiang
|
fcde67b016
|
CPU: map changes from developing branch in sgl-kernel (#6833)
Co-authored-by: mingfeima <mingfei.ma@intel.com>
|
2025-06-10 01:08:15 -07:00 |
|
yudian0504
|
81372f3bef
|
Fix fused_moe triton configs (#7029)
|
2025-06-09 23:23:03 -07:00 |
|
Arthur Cheng
|
baa6624d7c
|
[CI] Add CI workflow for sgl-router docker build (#7027)
|
2025-06-09 23:16:44 -07:00 |
|
Byron Hsu
|
c2b16795b5
|
Add decode req pool (#6980)
|
2025-06-09 21:23:36 -07:00 |
|
fzyzcjy
|
f6ebba537a
|
Support both approximate and exact expert distribution collection (#6964)
|
2025-06-09 20:56:17 -07:00 |
|
Baizhou Zhang
|
6716b41786
|
Update default settings for blackwell (#7023)
|
2025-06-09 20:37:47 -07:00 |
|
Yineng Zhang
|
1c8b42c84c
|
chore: update pr test xeon (#7018)
|
2025-06-09 17:36:25 -07:00 |
|
Qiaolin Yu
|
f20f70003d
|
Fix torch version in blackwell dockerfile (#7017)
|
2025-06-09 17:06:50 -07:00 |
|
Emmanuel Ferdman
|
f40942ad63
|
Migrate to assertEqual (#6741)
Signed-off-by: Emmanuel Ferdman <emmanuelferdman@gmail.com>
|
2025-06-09 16:47:39 -07:00 |
|
Lianmin Zheng
|
dc0705a504
|
Simplify prepare_extend_after_decode (#6987)
|
2025-06-09 16:39:21 -07:00 |
|
Wenxuan Tan
|
a968c888c0
|
Fix torchvision version for Blackwell (#7015)
|
2025-06-09 15:50:19 -07:00 |
|
Baizhou Zhang
|
a979daac3b
|
Fallback to lower triton version for unfound fused moe configs (#7013)
|
2025-06-09 15:41:03 -07:00 |
|
ishandhanani
|
f1569876d5
|
feat: add direct routing strategy to DP worker (#6884)
|
2025-06-09 11:44:05 -07:00 |
|
Sai Enduri
|
3465d7ae78
|
Update amd nightly models CI. (#6992)
|
2025-06-09 10:54:08 -07:00 |
|
fzyzcjy
|
e58423b2b9
|
Fix cutlass MLA gets almost zero accuracy (#6998)
|
2025-06-09 10:16:29 -07:00 |
|
Yineng Zhang
|
7059ae16fb
|
chore: update pr test xeon (#7008)
|
2025-06-09 10:08:44 -07:00 |
|
Yineng Zhang
|
51d9a597f9
|
cleanup tmp dir (#7007)
|
2025-06-09 09:26:04 -07:00 |
|
Yineng Zhang
|
56ccd3c22c
|
chore: upgrade flashinfer v0.2.6.post1 jit (#6958)
Co-authored-by: alcanderian <alcanderian@gmail.com>
Co-authored-by: Qiaolin Yu <qy254@cornell.edu>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
|
2025-06-09 09:22:39 -07:00 |
|
Yueyang Pan
|
98c00a2df1
|
Fix torch profiler bugs for bench_offline_throughput.py (#6557)
|
2025-06-09 20:33:41 +08:00 |
|
Pan Lyu
|
451ffe74d9
|
support qwen3 emebedding (#6990)
|
2025-06-09 01:32:49 -07:00 |
|
Lifu Huang
|
b1e5a33ae3
|
Eliminate stream sync to speed up LoRA batch init (#6960)
|
2025-06-09 00:22:45 -07:00 |
|
Lianmin Zheng
|
9d5fa68b90
|
Use torch.compile to fuse flash attention decode metadata preparation (#6973)
|
2025-06-08 23:05:40 -07:00 |
|
Sai Enduri
|
2c18642502
|
Enable more unit tests for AMD CI. (#6983)
|
2025-06-08 19:41:55 -07:00 |
|
JieXin Liang
|
18efb5e8e0
|
[perf][sgl-kernel] extend cutlass_mla_decode to support num_head < 128 (#6929)
|
2025-06-08 19:37:34 -07:00 |
|
fzyzcjy
|
de1350ea20
|
Minor remove one kernel for DeepSeek (#6977)
|
2025-06-08 17:41:35 -07:00 |
|
fzyzcjy
|
86fe943bc3
|
Fix expert distribution dumping causes OOM (#6967)
|
2025-06-08 17:41:14 -07:00 |
|