Yineng Zhang
|
3289c1207d
|
Update the retry count (#5051)
|
2025-04-03 17:07:38 -07:00 |
|
renxin
|
cccfc10e9c
|
Feature/revise docs ci (#5009)
|
2025-04-02 20:08:56 -07:00 |
|
Yuhong Guo
|
87fafa0105
|
Revert PR 4764 & 4813 related to R1 RoPE (#4959)
|
2025-03-31 20:56:58 -07:00 |
|
Lianmin Zheng
|
f842853a40
|
Fix the timeout for unit-test-2-gpu in pr-test.yml (#4927)
|
2025-03-30 12:15:40 -07:00 |
|
Adarsh Shirawalmath
|
9fccda3111
|
[Feature] use pytest for sgl-kernel (#4896)
|
2025-03-30 10:36:52 -07:00 |
|
Lianmin Zheng
|
4ede6770cd
|
Fix retract for page size > 1 (#4914)
|
2025-03-30 02:57:15 -07:00 |
|
Yineng Zhang
|
72549263c6
|
update sgl-kernel test ci (#4866)
|
2025-03-28 11:42:41 -07:00 |
|
Lianmin Zheng
|
74e0ac1dbd
|
Clean up import vllm in quantization/__init__.py (#4834)
|
2025-03-28 10:34:10 -07:00 |
|
warjiang
|
18317ddc13
|
ci: add condition for daily docker build (#4487)
|
2025-03-27 21:44:37 -07:00 |
|
fzyzcjy
|
0d3e3072ee
|
Fix CI of test_patch_torch (#4844)
|
2025-03-27 21:22:45 -07:00 |
|
Yineng Zhang
|
5fa3058f01
|
fix the release doc dependency issue (#4828)
|
2025-03-27 13:28:12 -07:00 |
|
strgrb
|
668ecc6c5b
|
Fix ut mla-test-1-gpu-amd (#4813)
Co-authored-by: Zhang Kaihong <zhangkaihong.zkh@alibaba-inc.com>
|
2025-03-27 08:27:51 -07:00 |
|
Yineng Zhang
|
8bf6d7f406
|
support cmake for sgl-kernel (#4706)
Co-authored-by: hebiao064 <hebiaobuaa@gmail.com>
Co-authored-by: yinfan98 <1106310035@qq.com>
|
2025-03-27 01:42:28 -07:00 |
|
Xiaoyu Zhang
|
04e3ff6975
|
Support compressed tensors fp8w8a8 (#4743)
|
2025-03-26 13:21:25 -07:00 |
|
fzyzcjy
|
26f07294f1
|
Warn users when release_memory_occupation is called without memory saver enabled (#4566)
|
2025-03-26 00:18:14 -07:00 |
|
fzyzcjy
|
15ddd84322
|
Add retry for flaky tests in CI (#4755)
|
2025-03-25 16:53:12 -07:00 |
|
fzyzcjy
|
e45ae444db
|
Revert "Add DeepEP tests into CI (#4737)" (#4751)
|
2025-03-25 00:44:01 -07:00 |
|
Yineng Zhang
|
9b7cf9ee6c
|
support cu128 sgl-kernel (#4744)
|
2025-03-24 20:53:23 -07:00 |
|
fzyzcjy
|
64129fa632
|
Add DeepEP tests into CI (#4737)
|
2025-03-24 19:54:31 -07:00 |
|
aoshen524
|
588865f0e0
|
[Feature] Support Tensor Parallelism and Weight Slicing for Lora (#4274)
Co-authored-by: ShenAo1111 <1377693092@qq.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-03-18 20:33:07 -07:00 |
|
Yineng Zhang
|
c787298547
|
use sgl custom all reduce (#4441)
|
2025-03-18 00:46:41 -07:00 |
|
Lianmin Zheng
|
82dec1f70b
|
Remove redundant type conversion (#4513)
|
2025-03-17 05:57:35 -07:00 |
|
Lianmin Zheng
|
5493c3343e
|
Fix data parallel + tensor parallel (#4499)
|
2025-03-17 05:13:16 -07:00 |
|
Lianmin Zheng
|
06d12b39d3
|
Remove filter for pr-tests (#4468)
|
2025-03-16 00:57:26 -07:00 |
|
Lianmin Zheng
|
c30976fb41
|
Fix finish step for pr tests and notebook tests (#4467)
|
2025-03-16 00:52:06 -07:00 |
|
Yineng Zhang
|
ad1ae7f7cd
|
use topk_softmax with sgl-kernel (#4439)
|
2025-03-14 15:59:06 -07:00 |
|
Yineng Zhang
|
977d7cd26a
|
cleanup deps 1/n (#4400)
Co-authored-by: sleepcoo <sleepcoo@gmail.com>
|
2025-03-14 00:00:33 -07:00 |
|
Lianmin Zheng
|
a5a892ffd3
|
Fix auto merge & add back get_flat_data_by_layer (#4393)
|
2025-03-13 08:46:25 -07:00 |
|
HandH1998
|
2ac189edc8
|
Amd test fp8 (#4261)
|
2025-03-10 10:12:09 -07:00 |
|
Lianmin Zheng
|
5a6400eec5
|
Test no vllm custom allreduce (#4256)
|
2025-03-10 10:08:25 -07:00 |
|
Lianmin Zheng
|
3d56585a97
|
increase the timeout of nightly-test.yml (#4262)
|
2025-03-10 05:07:03 -07:00 |
|
Lianmin Zheng
|
aa957102a9
|
Simplify tests & Fix trtllm custom allreduce registration (#4252)
|
2025-03-10 01:24:22 -07:00 |
|
Lianmin Zheng
|
e8a69e4d0c
|
Clean up fp8 support (#4230)
|
2025-03-09 21:46:35 -07:00 |
|
Lianmin Zheng
|
fbd560028a
|
Auto balance CI tests (#4238)
|
2025-03-09 21:05:55 -07:00 |
|
Lianmin Zheng
|
8abf74e3c9
|
Rename files in sgl kernel to avoid nested folder structure (#4213)
Co-authored-by: zhyncs <me@zhyncs.com>
|
2025-03-08 22:54:51 -08:00 |
|
Yineng Zhang
|
ee132a4515
|
use latest sgl-kernel for mla test (#4222)
|
2025-03-08 22:27:47 -08:00 |
|
Lianmin Zheng
|
48473684cc
|
Split test_mla.py into two files (#4216)
|
2025-03-08 15:40:49 -08:00 |
|
Lianmin Zheng
|
2cadd51d11
|
Test no vllm custom allreduce (#4210)
|
2025-03-08 05:23:06 -08:00 |
|
Lianmin Zheng
|
8d323e95e4
|
Use clang format 18 in pr-test-sgl-kernel.yml (#4203)
|
2025-03-08 01:28:10 -08:00 |
|
saienduri
|
e1aaa79ac9
|
Update amd ci docker image to v0.4.3.post4-rocm630. (#4189)
|
2025-03-07 13:02:02 -08:00 |
|
Yineng Zhang
|
7e3bb52705
|
update release-pypi-kernel
|
2025-03-07 01:48:47 -08:00 |
|
Chayenne
|
9854a18a51
|
Hot fix small vocal eagle in docs (#4154)
Co-authored-by: ybyang <ybyang7@iflytek.com>
|
2025-03-06 15:13:26 -08:00 |
|
Lianmin Zheng
|
bc1534ff32
|
Fix a draft model accuracy bug in eagle; support step=1; return logprob in eagle (#4134)
Co-authored-by: Sehoon Kim <kssteven418@gmail.com>
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
Co-authored-by: Sehoon Kim <sehoon@x.ai>
|
2025-03-06 06:13:59 -08:00 |
|
saienduri
|
55dc8e4d52
|
Add tag suffix to nightly docker builds. (#4129)
|
2025-03-05 23:22:36 -08:00 |
|
saienduri
|
44d7646371
|
remove testing on PR workflow change (#4110)
|
2025-03-05 16:03:18 -08:00 |
|
saienduri
|
cd85b78f94
|
Create release-docker-amd-nightly.yml (#4105)
|
2025-03-05 14:46:26 -08:00 |
|
Ke Bao
|
d3fe9bae56
|
Add accuracy test for TP torch compile (#3994)
|
2025-03-02 13:18:18 -08:00 |
|
fzyzcjy
|
e3e0bc50a9
|
[Feature] SPMD for SGLang + Verl (#3852)
|
2025-02-28 09:53:10 -08:00 |
|
Qing
|
0519269d20
|
[Docs] Disable notebook CI when merge to main (#3905)
|
2025-02-26 22:13:33 -08:00 |
|
Lianmin Zheng
|
d7934cde45
|
Fix CI and install docs (#3821)
|
2025-02-24 16:17:38 -08:00 |
|