sglang

Author	SHA1	Message	Date
Yineng Zhang	0aaccbbfec	revert deepseek docs (#4109 )	2025-03-05 13:23:11 -08:00
Qiaolin Yu	357671e216	Add examples for server token-in-token-out (#4103 ) Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>	2025-03-05 13:16:31 -08:00
Chayenne	e70fa279bc	Docs: reorganize dpsk docs (#4108 )	2025-03-05 13:01:03 -08:00
Tommy Yang	abe74b7b59	Docs: Add DeepSeek optimization ablations documentation (#4107 ) Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>	2025-03-05 12:25:51 -08:00
Jhin	70b3c6eeb1	Add update_weights_from_disk endpoint to Engine (#4102 ) Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>	2025-03-05 12:25:18 -08:00
Ke Bao	ef9d3b3c2c	Fix triton kernel illegal memory issue for eagle (#4100 )	2025-03-05 11:23:53 -08:00
Baizhou Zhang	fc91d08a8f	[Revision] Add fast decode plan for flashinfer mla (#4012 )	2025-03-05 11:20:41 -08:00
HAI	71ab0dabe0	Fix the moe padding conditional logic (#4081 )	2025-03-05 10:56:51 -08:00
Ying Sheng	d3d4d76758	[Eagle] Refactor eagle speculative decoding (#3986 ) Co-authored-by: Ke Bao <ISPObaoke@163.com>	2025-03-05 08:06:07 -08:00
yigex	5be8f1ed98	ROCM: AITER BLOCK GEMM (#4075 )	2025-03-05 03:10:49 -08:00
Lu Changqi	e5760bc40a	bench: add dataset param for bench_multiturn (#3990 )	2025-03-05 01:21:37 -08:00
Qubitium-ModelCloud	56a724eba3	[QUANT] Add GPTQModel Dynamic Quantization + `lm_head` Quantization (#3790 ) Signed-off-by: ZX-ModelCloud <zx@modelcloud.ai> Co-authored-by: ZX-ModelCloud <zx@modelcloud.ai>	2025-03-05 01:11:00 -08:00
Mick	583d6af71b	example: add vlm to token in & out example (#3941 ) Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>	2025-03-04 22:18:26 -08:00
Lianmin Zheng	e074d84e5b	[Minor] more code cleanup (#4077 )	2025-03-04 21:23:47 -08:00
Qiaolin Yu	4725e3f652	Add examples for returning hidden states when using the server (#4074 )	2025-03-04 19:31:50 -08:00
Lianmin Zheng	77a3954bf7	Simplify eagle tests and TP sync in grammar backend (#4066 )	2025-03-04 13:40:40 -08:00
Ke Bao	03b0364f76	Update nextn ci test (#4071 )	2025-03-04 13:01:24 -08:00
Lianmin Zheng	2dd7d0c533	Revert "Fix nightly-test CI" (#4065 )	2025-03-04 05:38:24 -08:00
William	0d4e3228cf	[Feature] Add test for speculative_token_map (#4016 )	2025-03-04 04:26:24 -08:00
Liu Jinjie	926f8efc0c	remove unused max_jobs (#3607 ) Signed-off-by: Jinjie Liu <jinjie.liu@usc.edu>	2025-03-04 04:23:39 -08:00
Xiuyu Li	9545bfb28a	fix: support gelu_new activation function in gpt2 (#3712 )	2025-03-04 04:09:52 -08:00
Michael Feil	37373ef2bb	sgl-router - issues on routing and project build. (#3870 ) (#3948 )	2025-03-04 04:06:30 -08:00
Chen Shengzhi	61261b3996	[XCCL] Use xccl for xpu backend since xccl is ready in latest PyTorch. (#3954 )	2025-03-04 04:05:56 -08:00
DarkSharpness	19120f71f3	[Fix & Style] Refactor the grammar backend to reduce human errors and improve readability (#4030 )	2025-03-04 03:56:45 -08:00
Kebe	2415ec3896	Remove grafana dashboard's datasource uid (#4051 )	2025-03-04 03:44:51 -08:00
Qubitium-ModelCloud	87f671ab58	Fix `debug_tensor_dump_output_folder` optional key missing (#4046 )	2025-03-04 03:42:48 -08:00
HAI	51d25405a7	ROCm: update aiter and its usage to fused moe (bloat16, fp8, fp8 block-quant) (#4053 )	2025-03-04 03:00:46 -08:00
kk	e0a2c96308	Fix breakage problem when using custom_ar (#4052 )	2025-03-04 02:59:03 -08:00
Xihuai Wang	12f2e6c3f1	Fix: #3988 using blockwise_int8 (#4023 )	2025-03-03 23:49:58 -08:00
Xihuai Wang	95575aa76a	Reasoning parser (#4000 ) Co-authored-by: Lucas Pickup <lupickup@microsoft.com>	2025-03-03 21:16:36 -08:00
kk	11eea69e70	Fix assert options.num_stages != 0 error in the latest ROCm build image (#4049 ) Co-authored-by: wunhuang <wunhuang@amd.com>	2025-03-03 20:37:03 -08:00
Yineng Zhang	1baa9e6cf9	docs: update README (#4044 )	2025-03-03 17:09:18 -08:00
Lianmin Zheng	911fcd0910	Update README.md (#4043 )	2025-03-03 16:29:46 -08:00
Ke Bao	9fafa62db7	Share target model embed and head weights for nextn (#4033 )	2025-03-03 13:30:04 -08:00
Chayenne	146ac8df07	Add examples in sampling parameters (#4039 )	2025-03-03 13:04:32 -08:00
Qiaolin Yu	57a404fd55	Remove outdated test utils and fix links for the doc of sampling params (#3999 )	2025-03-03 09:41:38 -08:00
Chayenne	2796fbb53d	Docs: Fix sampling parameter (#4034 )	2025-03-03 09:32:36 -08:00
Lianmin Zheng	935cda944b	Misc clean up; Remove the support of jump forward (#4032 )	2025-03-03 07:02:14 -08:00
Lianmin Zheng	110e006673	Reorganize python source files in sgl-kernel with multiple files (#4027 )	2025-03-03 06:36:40 -08:00
Lianmin Zheng	6b45a21d16	Reorganize c++ source files in sgl-kernel with multiple folders (#4025 )	2025-03-03 05:32:30 -08:00
Yudi Xue	a7000a7650	Update metrics documentation (#3264 )	2025-03-03 05:03:58 -08:00
Lianmin Zheng	1a8f995c46	remove cache configs in model definitions (#4031 )	2025-03-03 05:00:50 -08:00
Lianmin Zheng	a3ab768a2b	Clean up custom allreduce (#4029 )	2025-03-03 04:59:53 -08:00
Lianmin Zheng	66301e124f	Improve code styles (#4021 )	2025-03-03 03:20:23 -08:00
Lianmin Zheng	ac2387279e	Support penalty in overlap mode; return logprob with chunked prefill; improve benchmark scripts (#3988 ) Co-authored-by: SangBin Cho <rkooo567@gmail.com> Co-authored-by: dhou-xai <dhou@x.ai> Co-authored-by: Hanming Lu <hanming_lu@berkeley.edu>	2025-03-03 00:12:04 -08:00
Stefan He	0194948fd9	Optimize Triton Kernel of Group GEMM in DeepGEMM Benchmark (#4014 )	2025-03-02 23:29:55 -08:00
yinfan98	b4d34cd35d	Fix nightly-test CI (#3826 )	2025-03-02 23:14:45 -08:00
Chayenne	728e175fc4	Add examples to token-in-token-out for LLM (#4010 )	2025-03-02 21:03:49 -08:00
Lianmin Zheng	9e1014cf99	Revert "Add fast decode plan for flashinfer mla" (#4008 )	2025-03-02 19:29:10 -08:00
Baizhou Zhang	fa56106731	Add fast decode plan for flashinfer mla (#3987 )	2025-03-02 19:16:37 -08:00

1 2 3 4 5 ...

2244 Commits