Lianmin Zheng
|
be2d985df8
|
Minor style change of triton backend (#7165)
|
2025-06-13 16:01:23 -07:00 |
|
fzyzcjy
|
5b1afa7814
|
Re-quantize DeepSeek model weights to support DeepGEMM new input format (#7156)
|
2025-06-13 15:57:45 -07:00 |
|
fzyzcjy
|
c49c1d9226
|
Remove 200us slow concat kernel (part 2: srt) (#7020)
|
2025-06-13 15:19:31 -07:00 |
|
fzyzcjy
|
0f1dfa1efe
|
Tiny add sanity checks for DeepGEMM inputs (#7157)
|
2025-06-13 14:36:27 -07:00 |
|
fzyzcjy
|
e3ec6bf4b6
|
Minor speed up block_quant_dequant (#6814)
|
2025-06-13 14:32:46 -07:00 |
|
fzyzcjy
|
b04df75acd
|
Fix DeepEP error in some environments (#7154)
|
2025-06-13 14:27:25 -07:00 |
|
fzyzcjy
|
bec3e48402
|
Support new DeepGEMM format in per token group quant (part 2: srt) (#7155)
|
2025-06-13 14:25:40 -07:00 |
|
Yineng Zhang
|
8ab7d93c2e
|
chore: bump v0.1.8.post1 (#7152)
|
2025-06-13 03:14:26 -07:00 |
|
fzyzcjy
|
5c66c4424f
|
Support new DeepGEMM format in per token group quant (#7146)
|
2025-06-13 02:00:22 -07:00 |
|
fzyzcjy
|
aa46ed34d2
|
Remove 200us slow concat kernel (part 1: kernel) (#7145)
|
2025-06-13 01:58:29 -07:00 |
|
pansicheng
|
2f4ec752bc
|
filter by num_hidden_layers (#7056)
|
2025-06-13 00:53:09 -07:00 |
|
fzyzcjy
|
da47621ccc
|
Minor speedup topk postprocessing (#7058)
|
2025-06-13 00:50:18 -07:00 |
|
Binyao Jiang
|
22a6b9fc05
|
Remove unnecessary metadata_expand.max_seq_len_k operations in fa3 to… (#7140)
|
2025-06-12 23:25:52 -07:00 |
|
fzyzcjy
|
b02df20a8d
|
Minor add utility to read expert distribution recorder output (#7134)
|
2025-06-12 16:50:05 -07:00 |
|
Povilas Kanapickas
|
bd7cfbd2f8
|
[Fix] Reduce busy polling when scheduler is idle (#6026)
|
2025-06-12 14:58:22 -07:00 |
|
sogalin
|
4b9971e401
|
Add gfx950 support for sgl-kernel. (#7092)
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
|
2025-06-12 11:07:48 -07:00 |
|
Lianmin Zheng
|
dcc79d325b
|
Update CODEOWNERS (#7126)
|
2025-06-12 02:12:49 -07:00 |
|
Yineng Zhang
|
7046e0fab7
|
feat: update blackwell setup (#7119)
|
2025-06-12 01:54:40 -07:00 |
|
Liangsheng Yin
|
930746d93c
|
Improve log status (#7115)
|
2025-06-12 14:38:24 +08:00 |
|
Yuan Luo
|
84727a5139
|
[sgl-kernel] Add cuda kernel for moe_ep_silu_and_mul (#6919)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-06-11 20:43:08 -07:00 |
|
Quanfeng Li
|
ef32677444
|
Fix positional argument (#7093)
|
2025-06-11 18:31:13 -07:00 |
|
Lifu Huang
|
021f76e4f4
|
[Perf] Refactor LoRAManager to eliminate stream syncs and redundant computations (#6994)
|
2025-06-11 16:18:57 -07:00 |
|
Faradawn Yang
|
777688b892
|
[feat]: Emit fixed-size KV blocks events (#6824)
|
2025-06-11 13:07:58 -07:00 |
|
Neo
|
0ca594eda9
|
[FIX]remove redundant code in logits_processor.py (#7079)
|
2025-06-11 11:49:30 -07:00 |
|
Zijian
|
31d6dee5c4
|
Support VILA models (#6106)
|
2025-06-11 11:47:25 -07:00 |
|
sogalin
|
02543b545c
|
Fix misusing the "_is_cuda". (#7091)
|
2025-06-11 11:21:31 -07:00 |
|
Baizhou Zhang
|
25a6a9aa22
|
Fix circular import in test_prefix_chunk_info.py (#7097)
|
2025-06-11 10:57:45 -07:00 |
|
Mick
|
83d87685c5
|
vlm: adapt internvl to VisionAttention (#6870)
|
2025-06-11 01:16:04 -07:00 |
|
Baizhou Zhang
|
2a5f0100e0
|
Fix GGuf and add back test_gguf.py (#7067)
|
2025-06-10 21:07:20 -07:00 |
|
Lianmin Zheng
|
dbdf76ca98
|
Clean up docs for server args and sampling parameters (generated by grok) (#7076)
|
2025-06-10 19:55:42 -07:00 |
|
Ximingwang-09
|
f2a75a66c4
|
update doc (#7046)
Co-authored-by: ximing.wxm <ximing.wxm@antgroup.com>
|
2025-06-11 10:02:01 +08:00 |
|
Lianmin Zheng
|
6b12d6a8d5
|
Simplify the heuristics for setting --mem-fraction-static (#7054)
|
2025-06-10 19:01:39 -07:00 |
|
Lianmin Zheng
|
0f218731e3
|
Do not run frontend_reasoning.ipynb to reduce the CI load (#7073)
|
2025-06-10 17:15:31 -07:00 |
|
Yudi Xue
|
14c18d25df
|
Frontend language separate reasoning support (#6031)
|
2025-06-10 17:11:29 -07:00 |
|
Lianmin Zheng
|
90bd3e32d6
|
Improve perf tuning docs (#7071)
|
2025-06-10 16:55:04 -07:00 |
|
Brayden Zhong
|
ca9291181d
|
[Feature] Add Logit Bias (#6579)
Co-authored-by: Cinjon Resnick <cinjon.resnick@gmail.com>
|
2025-06-10 15:39:25 -07:00 |
|
Yineng Zhang
|
344adb00ec
|
fix arm sgl-kernel link issue (#7066)
|
2025-06-10 15:23:22 -07:00 |
|
kyle-pena-kuzco
|
b56de8f943
|
Open AI API hidden states (#6716)
|
2025-06-10 14:37:29 -07:00 |
|
Yineng Zhang
|
ce5ee3bdf0
|
chore: update dev docker (#7064)
|
2025-06-10 12:57:58 -07:00 |
|
Xu Wenqing
|
a0e4d4eb53
|
Fix missing tool call id if tool call index >0 in streaming tool call output. (#7049)
Signed-off-by: 许文卿 <xwq391974@alibaba-inc.com>
|
2025-06-10 12:56:43 -07:00 |
|
Yineng Zhang
|
2f58445531
|
Revert "Add sanity checks when a test file is not added to CI (#6947)" (#7063)
|
2025-06-10 12:43:25 -07:00 |
|
fzyzcjy
|
fe55947acd
|
Add sanity checks when a test file is not added to CI (#6947)
|
2025-06-10 12:34:57 -07:00 |
|
fzyzcjy
|
19995dd78e
|
Tiny fix cutlass_mla_get_workspace_size stub incorrect signature (#7057)
|
2025-06-10 12:27:57 -07:00 |
|
Baizhou Zhang
|
3b014bc13d
|
Fix test_lora.py CI (#7061)
|
2025-06-10 12:24:46 -07:00 |
|
Hank Han
|
d7c3e8e93d
|
[fix] libmlx5.so already in base image (#7060)
|
2025-06-10 12:22:24 -07:00 |
|
kk
|
8ea7df6114
|
[WA] fix output data is nan in CI test "test_moe_eval_accuracy_large.py" (#7021)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2025-06-10 16:08:10 +00:00 |
|
Lianmin Zheng
|
4a102a2b02
|
Minor style fix in cuda_graph_runner.py (#7053)
|
2025-06-10 06:32:41 -07:00 |
|
Lianmin Zheng
|
6406408a70
|
Clean up server_args.py (#7037)
|
2025-06-10 05:34:29 -07:00 |
|
Lianmin Zheng
|
019851d099
|
Fix eagle on AMD (#7051)
|
2025-06-10 05:22:40 -07:00 |
|
Lianmin Zheng
|
2dae104dca
|
Minor cleanup of fa3 backend (#6999)
|
2025-06-10 03:58:44 -07:00 |
|