Xiaoyu Zhang
|
027e65248f
|
support echo=true and logprobs in openai api when logprobs=1 in lm-evaluation-harness (#1998)
|
2024-11-11 23:21:20 -08:00 |
|
Ke Bao
|
b808a38365
|
Filter empty prompt in random bench serving (#2011)
|
2024-11-12 14:53:41 +08:00 |
|
Byron Hsu
|
602ebc661d
|
remove sglang folder in rust (#2010)
|
2024-11-11 20:45:52 -08:00 |
|
Lianmin Zheng
|
530ae1bdc8
|
Fix weight loading for tied word embedding when TP > 1 (#2009)
|
2024-11-11 17:52:42 -08:00 |
|
Lianmin Zheng
|
befc6beb86
|
Fix a typo in io_struct.py (#2008)
|
2024-11-11 16:34:10 -08:00 |
|
Lianmin Zheng
|
59a5ba9be0
|
[Minor] Remove unused imports (#2006)
|
2024-11-11 15:36:14 -08:00 |
|
Byron Hsu
|
86c37d010a
|
fix sglang_router not found (#2005)
|
2024-11-11 15:20:14 -08:00 |
|
RangiLyu
|
f18b9c7252
|
support internlm2-reward (#1994)
|
2024-11-11 15:09:58 -08:00 |
|
Byron Hsu
|
3e33574374
|
run rust test on ubuntu instead of 1-gpu-runner (#2003)
|
2024-11-11 14:46:08 -08:00 |
|
Byron Hsu
|
0d94f1dd03
|
Bump router to 0.0.3 (#2004)
|
2024-11-11 14:42:22 -08:00 |
|
Byron Hsu
|
e728258d34
|
release router from py38 to py312 (#2002)
|
2024-11-11 14:30:25 -08:00 |
|
Byron Hsu
|
239eafbd2e
|
Fix rust unit test and pypi token (#2001)
|
2024-11-11 14:18:21 -08:00 |
|
James Xu
|
9d427265fd
|
Add Engine::encode example (#2000)
|
2024-11-11 13:43:35 -08:00 |
|
Byron Hsu
|
00ffde206f
|
setup router python binding ci (#1999)
|
2024-11-11 12:19:32 -08:00 |
|
James Xu
|
ddeb9d42de
|
Add engine encode (#1995)
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
|
2024-11-11 11:48:17 -08:00 |
|
Yineng Zhang
|
aaf0a3156e
|
docs: add slides link in README (#1997)
|
2024-11-11 05:03:16 -08:00 |
|
Byron Hsu
|
f9633fa9b9
|
[rust] cache-aware DP - approx tree (#1934)
|
2024-11-10 21:57:32 -08:00 |
|
HAI
|
087ab83223
|
[Performance, Triton] Optimize over mask compute to tl.load in fused_moe_kernel (#1980)
|
2024-11-10 18:54:43 -08:00 |
|
Byron Hsu
|
8169c6f4ef
|
Add gen-shared-prefix dataset in bench_serving (#1990)
|
2024-11-11 08:39:56 +08:00 |
|
Lianmin Zheng
|
3d043319aa
|
[CI] Balance unit tests (#1988)
|
2024-11-10 11:45:01 -08:00 |
|
yizhang2077
|
a8aad9357d
|
qwen2vl fix bug for #1971 #1897 (#1984)
|
2024-11-10 08:10:45 -08:00 |
|
Yineng Zhang
|
47ffe7af81
|
docs: add shm size for docker run (#1986)
|
2024-11-10 22:14:48 +08:00 |
|
Yineng Zhang
|
b3523af8eb
|
fix: update pyzmq version (#1983)
|
2024-11-10 21:33:23 +08:00 |
|
Lianmin Zheng
|
1929c06762
|
Simplify prometheus metrics (#1981)
Co-authored-by: Mohit Reddy <mohitreddy1996@users.noreply.github.com>
|
2024-11-10 04:39:32 -08:00 |
|
Huanzhi (Hans) Mao
|
ed53ac84b4
|
Specify zmq Version Requirement (#1982)
|
2024-11-10 01:32:07 -08:00 |
|
Lianmin Zheng
|
520f0094e4
|
[CI] balance unit tests (#1977)
|
2024-11-09 16:46:14 -08:00 |
|
Lianmin Zheng
|
9c939a3d8b
|
Clean up metrics code (#1972)
|
2024-11-09 15:43:20 -08:00 |
|
Lianmin Zheng
|
549e8b8366
|
[Minor] Fix a typo in test_torchao.py (#1976)
|
2024-11-09 15:07:27 -08:00 |
|
Lianmin Zheng
|
a1f32867ca
|
Update pr-test-rust.yml to add a "finish" step (#1975)
|
2024-11-09 13:53:35 -08:00 |
|
Lianmin Zheng
|
760552e068
|
Update README.md (#1974)
|
2024-11-09 11:32:13 -08:00 |
|
Kursat Aktas
|
d9aada9db1
|
Introducing SGLang Guru on Gurubase.io (#1745)
|
2024-11-09 11:29:26 -08:00 |
|
Enrique Shockwave
|
f11eb90fe4
|
Initialize model_worker_batch variable (#1973)
|
2024-11-09 11:28:02 -08:00 |
|
Yudi Xue
|
95a4ed129a
|
Fix metrics (#1963)
|
2024-11-08 23:21:11 -08:00 |
|
leishaoSC
|
d1150e9a00
|
Updated Instructions on Profiling SGLang Infer System with AMD GPUs (#1966)
Co-authored-by: wunhuang <wunhuang@amd.com>
|
2024-11-08 23:19:03 -08:00 |
|
Chayenne
|
e3126e3c5f
|
Update README.md's Slack invitation link (#1962)
|
2024-11-08 11:46:25 -08:00 |
|
Lianmin Zheng
|
a509552087
|
[minor] Improve code style and compatibility (#1961)
|
2024-11-08 02:19:41 -08:00 |
|
Lianmin Zheng
|
7ef0084b0d
|
Add sentence_transformers to CI dependency (#1958)
|
2024-11-08 01:21:29 -08:00 |
|
HAI
|
f9a377f650
|
[Release, ROCm] release ROCm docker build for AMD MI GPUs (#1957)
|
2024-11-08 00:14:15 -08:00 |
|
aqweteddy
|
4ade15dd32
|
Adjust reward model's score module and pooler module order for reducing computation (#1956)
|
2024-11-08 00:10:54 -08:00 |
|
Lianmin Zheng
|
8dc84da084
|
Remove the useless to_srt_kwargs (#1955)
|
2024-11-07 23:15:08 -08:00 |
|
aqweteddy
|
f16eb15d0d
|
Gemma2 reward model support (#1954)
|
2024-11-07 22:42:27 -08:00 |
|
Yudi Xue
|
5bc2508b80
|
Monitoring documentation (#1933)
|
2024-11-07 22:14:16 -08:00 |
|
Lianmin Zheng
|
a71a44f203
|
Update setup_github_runner.md (#1952)
|
2024-11-07 19:20:47 -08:00 |
|
Lianmin Zheng
|
691808d587
|
Add a timeout for execute-notebook.yml (#1951)
|
2024-11-08 10:28:29 +08:00 |
|
HAI
|
d32fba2a4d
|
[ENV, ROCm] update environment settings (#1939)
|
2024-11-07 18:24:36 -08:00 |
|
HAI
|
67c424cce3
|
[Performance, Triton Kernel Args] extend_attention, optimize kern args to _fwd_kernel (#1941)
|
2024-11-07 18:24:02 -08:00 |
|
Lianmin Zheng
|
1ae270c5d0
|
[Doc] fix docs (#1949)
|
2024-11-07 18:20:41 -08:00 |
|
Chayenne
|
c77c1e05ba
|
fix black in pre-commit (#1940)
|
2024-11-08 07:42:47 +08:00 |
|
HAI
|
dca87ec348
|
[Docs] fix 404 - Contributor Guide (#1942)
|
2024-11-07 16:50:45 +08:00 |
|
Austin Liu
|
4b1d7a2583
|
Add Rust Router Python Binding (#1891)
Signed-off-by: Austin Liu <austin362667@gmail.com>
Co-authored-by: ByronHsu <byronhsu1230@gmail.com>
|
2024-11-06 18:08:30 -08:00 |
|