ishandhanani
|
9f24dfefd1
|
chore(gb200): remove ToT flashinfer installation (#9079)
|
2025-08-11 11:02:15 -07:00 |
|
Yi Zhang
|
89f1d4f536
|
update deepep commit to support qwen3-coder (#9066)
|
2025-08-11 10:42:33 -07:00 |
|
Yineng Zhang
|
48b8b4c124
|
fix nvshmem cu126 (#9001)
|
2025-08-09 03:34:54 -07:00 |
|
ishandhanani
|
4e7f025219
|
chore(gb200): update to CUDA 12.9 and improve build process (#8772)
|
2025-08-08 13:42:47 -07:00 |
|
Yineng Zhang
|
41357e511b
|
chore: update flashinfer (#8958)
|
2025-08-08 02:15:22 -07:00 |
|
Yineng Zhang
|
54ea57f245
|
chore: bump sgl-kernel v0.3.3 (#8957)
|
2025-08-08 01:35:37 -07:00 |
|
Yineng Zhang
|
cbbd685a46
|
chore: use torch 2.8 stable (#8880)
|
2025-08-06 15:51:40 -07:00 |
|
Mick
|
01c99a9959
|
chore: update Dockerfile (#8872)
Co-authored-by: zhyncs <me@zhyncs.com>
|
2025-08-06 09:30:33 -07:00 |
|
Yineng Zhang
|
4ef47839ae
|
feat: use py312 (#8832)
|
2025-08-05 13:38:22 -07:00 |
|
Yineng Zhang
|
0a56b721d5
|
chore: bump sgl-kernel v0.2.9 (#8713)
|
2025-08-02 16:21:56 -07:00 |
|
ishandhanani
|
b27b11919c
|
chore(gb200): update dockerfile to handle fp4 disaggregation (#8694)
|
2025-08-01 18:58:00 -07:00 |
|
Cheng Wan
|
6c88f6c8d9
|
[5/N] MoE Refactor: Update MoE parallelism arguments (#8658)
|
2025-08-01 01:20:03 -07:00 |
|
Yineng Zhang
|
43118f5f2a
|
chore: bump sgl-kernel v0.2.8 (#8599)
|
2025-07-30 22:23:52 -07:00 |
|
Charles Chen
|
659bfd1023
|
Add GKE's default CUDA runtime lib location to PATH and LD_LIBRARY_PATH. (#8544)
|
2025-07-30 20:28:07 -07:00 |
|
TimWang
|
2e1d2d7e66
|
Add PVC and update resource limits in k8s config (#8489)
|
2025-07-28 23:15:31 -07:00 |
|
kyleliang-nv
|
bb81daefb8
|
Fix docker buildx push error (#8425)
|
2025-07-27 17:59:38 -07:00 |
|
kyleliang-nv
|
e6312d271d
|
Uodate Dockerfile.gb200 to latest sglang (#8356)
|
2025-07-26 00:22:06 -07:00 |
|
Yineng Zhang
|
7181ec8cfc
|
fix: upgrade nccl version (#8359)
|
2025-07-25 14:59:02 -07:00 |
|
Shangming Cai
|
70e37b97bf
|
chore: upgrade mooncake 0.3.5 (#8341)
Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com>
|
2025-07-25 01:17:26 -07:00 |
|
Yineng Zhang
|
4c605235aa
|
fix: workaround for deepgemm warmup issue (#8302)
|
2025-07-23 12:01:51 -07:00 |
|
Yineng Zhang
|
429bb0efa2
|
chore: bump sgl-kernel v0.2.6.post1 (#8200)
|
2025-07-20 19:50:28 -07:00 |
|
Yineng Zhang
|
2db6719cc5
|
feat: update nccl 2.27.6 (#8182)
|
2025-07-19 22:55:45 -07:00 |
|
kyleliang-nv
|
bfdd226f35
|
Fix Dockerfile.gb200 (#8169)
|
2025-07-19 14:37:53 -07:00 |
|
Yineng Zhang
|
f98e88b9fb
|
chore: bump sgl-kernel v0.2.6 (#8165)
|
2025-07-19 00:56:18 -07:00 |
|
kyleliang-nv
|
cfab0ff6e2
|
Add GB200 wide-EP docker (#8157)
|
2025-07-18 22:34:29 -07:00 |
|
Simo Lin
|
8a7a7770e5
|
[ci] limit cmake build nproc (#8100)
|
2025-07-16 18:09:28 -07:00 |
|
Simo Lin
|
d9eb5efc71
|
[misc] update nvshmem and pin deepEP commit hash (#8098)
|
2025-07-16 08:54:55 -07:00 |
|
mqhc2020
|
a562c8a35c
|
[Dockerfile] Multi-arch support for ROCm (#7902)
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2025-07-14 06:13:09 +00:00 |
|
Yineng Zhang
|
eb118d88c4
|
chore: bump v0.4.9.post2 (#7963)
|
2025-07-11 21:11:20 -07:00 |
|
Yineng Zhang
|
8f3173d0b0
|
chore: bump sgl-kernel v0.2.5 (#7964)
|
2025-07-11 18:24:20 -07:00 |
|
Yineng Zhang
|
066f4ec91f
|
chore: bump v0.4.9.post1 (#7882)
|
2025-07-09 00:28:17 -07:00 |
|
Yineng Zhang
|
ec5f9c6269
|
chore: bump v0.4.9 (#7802)
|
2025-07-05 17:40:29 -07:00 |
|
Yineng Zhang
|
f200af0d8c
|
chore: bump sgl-kernel v0.2.4 (#7800)
|
2025-07-05 15:03:31 -07:00 |
|
Yineng Zhang
|
75354d9ae9
|
fix: use nvidia-nccl-cu12 2.27.5 (#7787)
|
2025-07-05 01:28:21 -07:00 |
|
Yineng Zhang
|
4fece12be9
|
chore: bump sgl-kernel v0.2.3 (#7784)
|
2025-07-05 00:05:45 -07:00 |
|
Yineng Zhang
|
aca1101a13
|
chore: bump sgl-kernel 0.2.2 (#7755)
|
2025-07-03 12:49:10 -07:00 |
|
Yineng Zhang
|
637bfee448
|
chore: bump sgl-kernel v0.2.1 (#7675)
|
2025-06-30 22:12:33 -07:00 |
|
Xiaoyu Zhang
|
ff2e9c9479
|
Add small requirements for benchmark/parse_result tools (#7671)
|
2025-06-30 21:52:20 -07:00 |
|
Chunyuan WU
|
c5131f7a2f
|
[CPU] add c++ kernel to bind CPU cores and memory node (#7524)
|
2025-06-29 19:45:25 -07:00 |
|
Yineng Zhang
|
69183f8808
|
chore: bump v0.4.8.post1 (#7559)
|
2025-06-26 02:21:12 -07:00 |
|
Shangming Cai
|
a07f8ae4b7
|
[CI] Upgrade mooncake to v0.3.4.post2 to fix potential slice failed bug (#7522)
Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com>
|
2025-06-25 01:49:22 -07:00 |
|
Yineng Zhang
|
7c3a12c000
|
chore: bump v0.4.8 (#7493)
|
2025-06-23 23:14:22 -07:00 |
|
Yineng Zhang
|
e846d95ef6
|
chore: bump sgl-kernel v0.2.0 (#7490)
|
2025-06-23 22:29:50 -07:00 |
|
Liangsheng Yin
|
76139bfba0
|
update mooncake in dockerfile (#7480)
|
2025-06-24 02:29:30 +08:00 |
|
kk
|
bd4f581896
|
Fix torch compile run (#7391)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Sai Enduri <saimanas.enduri@amd.com>
|
2025-06-22 15:33:09 -07:00 |
|
Yineng Zhang
|
4d8d9b8efd
|
chore: upgrade mooncake-transfer-engine 0.3.4 (#7401)
|
2025-06-20 16:38:54 -07:00 |
|
ybyang
|
906dbc34f1
|
[Docker] optimize dockerfile remove deepep and blackwell merge it to… (#7343)
Co-authored-by: Yineng Zhang <me@zhyncs.com>
|
2025-06-19 17:42:40 -07:00 |
|
Yineng Zhang
|
20a503c7d1
|
fix: resolve blackwell deepep image issue (#7331)
|
2025-06-18 17:04:02 -07:00 |
|
ybyang
|
712bf9ec9b
|
[pd] optimize dockerfile for pd disaggregation (#7319)
Co-authored-by: zhyncs <me@zhyncs.com>
|
2025-06-18 11:26:41 -07:00 |
|
YanbingJiang
|
094c116f7d
|
Update python API of activation, topk, norm and rope and remove vllm dependency (#6614)
Co-authored-by: Wu, Chunyuan <chunyuan.wu@intel.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
Co-authored-by: sdp <sdp@gnr799219.jf.intel.com>
|
2025-06-17 22:11:50 -07:00 |
|