yhang
|
14f1f1514b
|
H20 tune config for Kimi (#8047)
|
2025-07-15 13:48:31 -07:00 |
|
Albert
|
38216cf049
|
concurrently load weights of DeepseekV2ForCausalLM (#7943)
Signed-off-by: Tianyu Zhou <albert.zty@antgroup.com>
|
2025-07-15 13:41:19 -07:00 |
|
Yineng Zhang
|
4a8837950a
|
fix: resolve arm build issue (#8052)
|
2025-07-15 01:19:34 -07:00 |
|
jiawei
|
f1f1d1d40d
|
Fix the input tools format and history tool_calls in OpenAI API (#6556)
|
2025-07-15 00:58:55 -07:00 |
|
Xinyuan Tong
|
9120e83d03
|
fix: remove redundant rotary embedding cache recomputation in MiniCPM (#8022)
|
2025-07-15 00:12:45 -07:00 |
|
Xinyuan Tong
|
6e923dbd30
|
feat: update multimodal data handling in engine entrypoint (#8002)
Signed-off-by: Xinyuan Tong <justinning0323@outlook.com>
|
2025-07-15 00:12:22 -07:00 |
|
Qi Yuhang
|
c268c11c71
|
[feat]Support fusion kernel for constructing quant input and scale factor for fp8_blockwise_scaled_grouped_mm (#8023)
|
2025-07-15 00:02:44 -07:00 |
|
Chang Su
|
e6d5988442
|
Update CODEOWNERS (#8044)
|
2025-07-15 14:38:15 +08:00 |
|
杨睿
|
9b560c3e1c
|
fix: modality length mismatch with image_data (#7887)
|
2025-07-15 14:27:54 +08:00 |
|
Sai Enduri
|
5dc5866e8e
|
Setup workflow for releasing mi300x and mi350x dockers. (#8035)
|
2025-07-14 21:51:43 -07:00 |
|
ehuaa
|
64e78bb31b
|
prevent server crash from potential invalid grammar (#7897)
|
2025-07-15 11:21:45 +08:00 |
|
hzh0425
|
7c39e8a198
|
Fix Bug 'get_cpu_copy not Implemented' in pd offloading mode (#7982)
|
2025-07-14 14:57:10 -07:00 |
|
Lifu Huang
|
d969504d9a
|
Fix flaky CI: test_vlm_models (#8006)
|
2025-07-14 14:56:41 -07:00 |
|
ykcombat
|
1ebec1a8b0
|
[Feature] CUDA Green Context Support (#7649)
|
2025-07-15 02:49:16 +08:00 |
|
ykcombat
|
d4d0c7c367
|
[Feature]TP Group Switching for PD-Multiplexing (#7653)
|
2025-07-15 02:35:46 +08:00 |
|
Lianmin Zheng
|
8d2cf38c79
|
[Minor] Remove redundant print (#8005)
|
2025-07-14 10:55:13 -07:00 |
|
Hank Han
|
2117f82def
|
[ci] CI supports use cached models (#7874)
|
2025-07-14 11:42:21 +00:00 |
|
Yusong Gao
|
c07f647c9f
|
perf: add kimi k2 fused_moe tuning config for h30_3e (#8021)
Co-authored-by: yudian0504 <yudian.zy@antgroup.com>
|
2025-07-14 02:56:11 -07:00 |
|
Chunyuan WU
|
07452cbe8e
|
[CPU] fix no attribute 'can_fuse_mlp_allreduce' error (#8010)
|
2025-07-14 01:32:43 -07:00 |
|
mqhc2020
|
a562c8a35c
|
[Dockerfile] Multi-arch support for ROCm (#7902)
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2025-07-14 06:13:09 +00:00 |
|
Praneth Paruchuri
|
cb736df854
|
Support for Phi-1.5 & Phi-2 models (#7862)
|
2025-07-13 18:43:40 -07:00 |
|
Lifu Huang
|
e2ed9d049a
|
Refactor dynamic LoRA update to fix incorrect handling of variant weight shapes (#7844)
|
2025-07-13 18:36:01 -07:00 |
|
Peng Zhang
|
b5dd5e8741
|
chore: remove unnecessary limits on quantization methods in test script (#7997)
|
2025-07-13 16:11:49 -07:00 |
|
Hanming Lu
|
9379da77de
|
SWA Prefix Cache (#7367)
Co-authored-by: Ying Sheng <sqy1415@gmail.com>
|
2025-07-13 12:31:07 -07:00 |
|
ehuaa
|
0c55cbcfc5
|
[BugFix] add verify logit_bias to avoid crash because of IndexError (#7749)
|
2025-07-14 02:44:12 +08:00 |
|
fzyzcjy
|
c46e069d34
|
Tiny fix mooncake log warning wrong output (#7952)
|
2025-07-12 21:22:44 -07:00 |
|
Ying Sheng
|
42fc44100a
|
[minor] Add server_args check for Llama4 with hybrid (#7988)
|
2025-07-12 20:13:40 -07:00 |
|
Morpheus Guo
|
5f6756b038
|
[BugFix] fix pre_reorder_triton_kernel default int32 issue (#7814)
|
2025-07-12 13:42:36 -07:00 |
|
Cheng Wan
|
98aa836bbf
|
Overlap the gating function with shared experts in DeepSeek (#7978)
|
2025-07-12 13:41:50 -07:00 |
|
Yineng Zhang
|
22bd857cb5
|
docs: update README (#7985)
|
2025-07-12 13:31:11 -07:00 |
|
Ying Sheng
|
ccfa084125
|
[script] update loogle test (#7975)
|
2025-07-12 00:06:17 -07:00 |
|
Ying Sheng
|
bcc5ba94b4
|
[minor fix] SWA missing methods (#7972)
|
2025-07-11 23:57:02 -07:00 |
|
Ying Sheng
|
cee9f329c4
|
[minor fix] llama4 hybrid memory (#7950)
|
2025-07-11 23:11:36 -07:00 |
|
Yineng Zhang
|
eb118d88c4
|
chore: bump v0.4.9.post2 (#7963)
|
2025-07-11 21:11:20 -07:00 |
|
Yineng Zhang
|
732fc8e405
|
chore: upgrade sgl-kernel 0.2.5 (#7971)
|
2025-07-11 20:35:06 -07:00 |
|
Simo Lin
|
f2d5c4920e
|
[router] add worker abstraction (#7960)
|
2025-07-11 20:17:48 -07:00 |
|
fzyzcjy
|
2a2d3478af
|
Fix wrong gemm branch cause 250us slower (#7969)
|
2025-07-11 19:45:09 -07:00 |
|
Xiaoyu Zhang
|
aa2056091a
|
delete uselese code caused by fuse allreduce+add_rmsnorm pr (#7970)
|
2025-07-11 19:43:38 -07:00 |
|
Yineng Zhang
|
61bb285827
|
chore: upgrade xgrammar 0.1.21 (#7962)
|
2025-07-11 19:26:52 -07:00 |
|
fzyzcjy
|
880221bd3b
|
Revert "[PD Disaggregation] replace transfer with batch transfer for better performance (#7236)" (#7968)
|
2025-07-11 19:03:01 -07:00 |
|
Yineng Zhang
|
8f3173d0b0
|
chore: bump sgl-kernel v0.2.5 (#7964)
|
2025-07-11 18:24:20 -07:00 |
|
Qi Yuhang
|
26118a133d
|
[fix]Update unitest for fp8_blockwise_scaled_grouped_mm kernel (#7932)
|
2025-07-11 14:29:13 -07:00 |
|
Cheng Wan
|
475a249bb8
|
temporarily disable deepep-8-gpu and activate two small tests (#7961)
|
2025-07-11 14:22:05 -07:00 |
|
Peng Zhang
|
191d836ff6
|
fix: minor fix for modelopt weight load compatibility (#7953)
|
2025-07-11 14:20:58 -07:00 |
|
ronnie_zheng
|
86044712c6
|
[feature] kv transfer support of ascend npu (#7795)
Co-authored-by: liupeng <liupeng374@huawei.com>
|
2025-07-11 00:07:51 -07:00 |
|
Atream
|
615553079d
|
Support Kimi K2 (#7940)
|
2025-07-11 00:02:21 -07:00 |
|
Xiaoyu Zhang
|
49a5915f53
|
[ready b200] fuse allreduce+add_rmsnorm in prepare_attention + mlp module (#7775)
|
2025-07-10 15:12:39 -07:00 |
|
ronnie_zheng
|
766392c6bd
|
[feature]Ascend quantization support (#7791)
Co-authored-by: ichernob <ichernobnn@gmail.com>
Co-authored-by: liupeng <liupeng374@huawei.com>
|
2025-07-10 09:17:37 -07:00 |
|
likesen-alibaba
|
4a0d19198b
|
Fix bug of deepseek-v3 under DP+EP mode with large batchsize/seqlen (#6449)
|
2025-07-10 01:19:56 -07:00 |
|
Zaili Wang
|
5748241549
|
add sentencepiece as dependency explicitly (#7922)
|
2025-07-10 01:06:27 -07:00 |
|