Zhiqiang Xie
|
aee30630d8
|
Add a pointer to the real KV cache pool (#4113)
|
2025-03-05 21:39:07 -08:00 |
|
Lianmin Zheng
|
286e6540a6
|
Remove prefill-only-one-req (#4117)
|
2025-03-05 20:58:48 -08:00 |
|
Wenxuan Tan
|
718c391fd7
|
[Hoxfix] Fix incomplete token_to_kv_pool refactor (#4121)
|
2025-03-05 19:32:42 -08:00 |
|
Yineng Zhang
|
fc671f66c1
|
chore: bump v0.4.3.post3 (#4114)
|
2025-03-05 17:26:10 -08:00 |
|
Yueyang Pan
|
25482edb5c
|
Online serving benchmarks of real datasets for hierarchical KV caching (#3211)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
|
2025-03-05 16:16:43 -08:00 |
|
luzengxiangcn
|
62b362b1f1
|
Debug radixcache: refactor recursive helper methods (#3029)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
|
2025-03-05 16:11:42 -08:00 |
|
Jhin
|
70b3c6eeb1
|
Add update_weights_from_disk endpoint to Engine (#4102)
Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>
|
2025-03-05 12:25:18 -08:00 |
|
Ke Bao
|
ef9d3b3c2c
|
Fix triton kernel illegal memory issue for eagle (#4100)
|
2025-03-05 11:23:53 -08:00 |
|
Baizhou Zhang
|
fc91d08a8f
|
[Revision] Add fast decode plan for flashinfer mla (#4012)
|
2025-03-05 11:20:41 -08:00 |
|
HAI
|
71ab0dabe0
|
Fix the moe padding conditional logic (#4081)
|
2025-03-05 10:56:51 -08:00 |
|
Ying Sheng
|
d3d4d76758
|
[Eagle] Refactor eagle speculative decoding (#3986)
Co-authored-by: Ke Bao <ISPObaoke@163.com>
|
2025-03-05 08:06:07 -08:00 |
|
yigex
|
5be8f1ed98
|
ROCM: AITER BLOCK GEMM (#4075)
|
2025-03-05 03:10:49 -08:00 |
|
Qubitium-ModelCloud
|
56a724eba3
|
[QUANT] Add GPTQModel Dynamic Quantization + lm_head Quantization (#3790)
Signed-off-by: ZX-ModelCloud <zx@modelcloud.ai>
Co-authored-by: ZX-ModelCloud <zx@modelcloud.ai>
|
2025-03-05 01:11:00 -08:00 |
|
Mick
|
583d6af71b
|
example: add vlm to token in & out example (#3941)
Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>
|
2025-03-04 22:18:26 -08:00 |
|
Lianmin Zheng
|
e074d84e5b
|
[Minor] more code cleanup (#4077)
|
2025-03-04 21:23:47 -08:00 |
|
Lianmin Zheng
|
77a3954bf7
|
Simplify eagle tests and TP sync in grammar backend (#4066)
|
2025-03-04 13:40:40 -08:00 |
|
Lianmin Zheng
|
2dd7d0c533
|
Revert "Fix nightly-test CI" (#4065)
|
2025-03-04 05:38:24 -08:00 |
|
William
|
0d4e3228cf
|
[Feature] Add test for speculative_token_map (#4016)
|
2025-03-04 04:26:24 -08:00 |
|
Xiuyu Li
|
9545bfb28a
|
fix: support gelu_new activation function in gpt2 (#3712)
|
2025-03-04 04:09:52 -08:00 |
|
Chen Shengzhi
|
61261b3996
|
[XCCL] Use xccl for xpu backend since xccl is ready in latest PyTorch. (#3954)
|
2025-03-04 04:05:56 -08:00 |
|
DarkSharpness
|
19120f71f3
|
[Fix & Style] Refactor the grammar backend to reduce human errors and improve readability (#4030)
|
2025-03-04 03:56:45 -08:00 |
|
Qubitium-ModelCloud
|
87f671ab58
|
Fix debug_tensor_dump_output_folder optional key missing (#4046)
|
2025-03-04 03:42:48 -08:00 |
|
HAI
|
51d25405a7
|
ROCm: update aiter and its usage to fused moe (bloat16, fp8, fp8 block-quant) (#4053)
|
2025-03-04 03:00:46 -08:00 |
|
kk
|
e0a2c96308
|
Fix breakage problem when using custom_ar (#4052)
|
2025-03-04 02:59:03 -08:00 |
|
Xihuai Wang
|
12f2e6c3f1
|
Fix: #3988 using blockwise_int8 (#4023)
|
2025-03-03 23:49:58 -08:00 |
|
Xihuai Wang
|
95575aa76a
|
Reasoning parser (#4000)
Co-authored-by: Lucas Pickup <lupickup@microsoft.com>
|
2025-03-03 21:16:36 -08:00 |
|
kk
|
11eea69e70
|
Fix assert options.num_stages != 0 error in the latest ROCm build image (#4049)
Co-authored-by: wunhuang <wunhuang@amd.com>
|
2025-03-03 20:37:03 -08:00 |
|
Ke Bao
|
9fafa62db7
|
Share target model embed and head weights for nextn (#4033)
|
2025-03-03 13:30:04 -08:00 |
|
Qiaolin Yu
|
57a404fd55
|
Remove outdated test utils and fix links for the doc of sampling params (#3999)
|
2025-03-03 09:41:38 -08:00 |
|
Lianmin Zheng
|
935cda944b
|
Misc clean up; Remove the support of jump forward (#4032)
|
2025-03-03 07:02:14 -08:00 |
|
Lianmin Zheng
|
1a8f995c46
|
remove cache configs in model definitions (#4031)
|
2025-03-03 05:00:50 -08:00 |
|
Lianmin Zheng
|
a3ab768a2b
|
Clean up custom allreduce (#4029)
|
2025-03-03 04:59:53 -08:00 |
|
Lianmin Zheng
|
66301e124f
|
Improve code styles (#4021)
|
2025-03-03 03:20:23 -08:00 |
|
Lianmin Zheng
|
ac2387279e
|
Support penalty in overlap mode; return logprob with chunked prefill; improve benchmark scripts (#3988)
Co-authored-by: SangBin Cho <rkooo567@gmail.com>
Co-authored-by: dhou-xai <dhou@x.ai>
Co-authored-by: Hanming Lu <hanming_lu@berkeley.edu>
|
2025-03-03 00:12:04 -08:00 |
|
yinfan98
|
b4d34cd35d
|
Fix nightly-test CI (#3826)
|
2025-03-02 23:14:45 -08:00 |
|
Lianmin Zheng
|
9e1014cf99
|
Revert "Add fast decode plan for flashinfer mla" (#4008)
|
2025-03-02 19:29:10 -08:00 |
|
Baizhou Zhang
|
fa56106731
|
Add fast decode plan for flashinfer mla (#3987)
|
2025-03-02 19:16:37 -08:00 |
|
Zhousx
|
7fbab730bd
|
[feat] add small vocab table for eagle's draft model[1]. (#3822)
Co-authored-by: Achazwl <323163497@qq.com>
Co-authored-by: Chayenne <zhaochen20@outlook.com>
|
2025-03-02 18:58:45 -08:00 |
|
Hubert Lu
|
9cf4077294
|
Enable custom AR for AMD GPUs and maintain it in sgl-kernel (#3406)
|
2025-03-02 15:19:06 -08:00 |
|
Ke Bao
|
00ce7e311c
|
Fix all gather torch compile (#3992)
Co-authored-by: yizhang2077 <1109276519@qq.com>
|
2025-03-02 00:41:38 -08:00 |
|
Qiaolin Yu
|
40782f05d7
|
Refactor: Move return_hidden_states to the generate input (#3985)
Co-authored-by: Beichen-Ma <mabeichen12@gmail.com>
|
2025-03-01 17:51:29 -08:00 |
|
Chayenne
|
930da877c4
|
rename FunctionCallReqInput to ParseFunctionCallReq (#3976)
|
2025-02-28 18:46:25 -08:00 |
|
Baizhou Zhang
|
90a4b7d98a
|
[Feature]Support ragged prefill in flashinfer mla backend (#3967)
Co-authored-by: Yineng Zhang <me@zhyncs.com>
Co-authored-by: pankajroark <pankajroark@users.noreply.github.com>
|
2025-02-28 18:13:56 -08:00 |
|
Yineng Zhang
|
f3b99f73b3
|
update flashinfer-python version
|
2025-02-28 16:31:59 -08:00 |
|
Chaitanya Sri Krishna Lolla
|
77a6c9d229
|
Remove unused imports from rocm mla kernel. (#3963)
|
2025-02-28 10:01:08 -08:00 |
|
fzyzcjy
|
e3e0bc50a9
|
[Feature] SPMD for SGLang + Verl (#3852)
|
2025-02-28 09:53:10 -08:00 |
|
mlmz
|
bac414ab53
|
[Feature] integrate Structural Tag in xgrammar backend for function calling (#3566)
Co-authored-by: shuaills <shishuaiuoe@gmail.com>
|
2025-02-27 23:33:41 -08:00 |
|
Chang Su
|
eec3f6d1eb
|
[Bugfix] Fix tokenizer_manager not getting 400 when req is too long (#3678)
Co-authored-by: voidxb <unkown>
|
2025-02-27 22:59:43 -08:00 |
|
Chayenne
|
90bc26a813
|
set a strict sgl-kernel version (#3950)
|
2025-02-27 22:44:57 -08:00 |
|
Kebe
|
ec0a72c2d9
|
Fix bench_serving not recognizing OPENAI_API_KEY (#3870)
Signed-off-by: Kebe <mail@kebe7jun.com>
|
2025-02-27 20:18:53 -08:00 |
|