sglang

Author	SHA1	Message	Date
Mick	3d93f84a00	[Feature] Support minicpmv v2.6 (#2785 ) Co-authored-by: Chayenne <zhaochen20@outlook.com> Co-authored-by: yizhang2077 <1109276519@qq.com>	2025-01-18 14:14:19 -08:00
bjmsong	d3024f4fc8	support e4m3 kvcache in qwen2 & add kv scaling facotr json (#2894 ) Co-authored-by: bjmsong <bjmsong@126.com>	2025-01-18 11:43:22 +08:00
Lianmin Zheng	8f2c522aba	Improve benchmark scripts and error message printing (#2922 )	2025-01-16 06:24:31 -08:00
Lianmin Zheng	b22f3f6475	Fix nightly accuracy tests (#2780 )	2025-01-07 21:02:35 -08:00
Lianmin Zheng	bdc1acf6cd	Misc fix for min_p_sampling, --cuda-graph-bs (#2761 )	2025-01-07 02:52:53 -08:00
Lianmin Zheng	a6ca736c8e	Simplify stream_output (#2398 )	2024-12-08 12:27:13 -08:00
Lianmin Zheng	a2486eb58f	Fix a bug with logprob streaming + chunked prefill (#2403 )	2024-12-08 03:55:27 -08:00
Lianmin Zheng	ccaf1f997c	[CI] Print summary on github actions (#2274 )	2024-11-29 23:48:54 -08:00
Chayenne	7d5d1d3d29	udate weights from disk (#2265 )	2024-11-30 01:17:00 +00:00
bjmsong	01017d4c20	Support LoRA in Completion API (#2243 ) Co-authored-by: root <bjmsong@126.com>	2024-11-29 16:13:38 -08:00
Lianmin Zheng	b2ccf36d4d	Fix memory leak during abort (#2238 )	2024-11-28 02:22:15 -08:00
Lianmin Zheng	d4fc1a70e3	Crash the server correctly during error (#2231 )	2024-11-28 00:22:39 -08:00
Lianmin Zheng	fb6e04a0c2	Use an env var SGLANG_SET_CPU_AFFINITY to set cpu affinity; turn it off by default (#2222 )	2024-11-27 02:52:46 -08:00
Lianmin Zheng	6997e28f6e	Revert "Use an env var SGLANG_SET_CPU_AFFINITY to set cpu affinity; turn it off by default" (#2221 )	2024-11-27 02:02:01 -08:00
Lianmin Zheng	a0e58740a8	Use an env var SGLANG_SET_CPU_AFFINITY to set cpu affinity; turn it off by default (#2217 )	2024-11-27 01:13:41 -08:00
Lianmin Zheng	8e1adb8441	Allow overwrite flashinfer use_tensorcore (#2169 )	2024-11-24 20:58:17 -08:00
Yineng Zhang	4f8c3aeafc	minor: update gsm8k threshold (#2125 )	2024-11-22 19:23:58 +08:00
bjmsong	ad30d5cf9a	Benchmark with Pytorch Profiler easily (#2110 ) Co-authored-by: root <bjmsong@126.com>	2024-11-21 23:29:50 -08:00
Lianmin Zheng	dfec7fca06	Rename sglang.bench_latency to sglang.bench_one_batch (#2118 )	2024-11-21 20:07:48 -08:00
Lianmin Zheng	7d671e4ad2	Enable overlap by default (#2067 )	2024-11-19 22:07:58 -08:00
Yineng Zhang	766192610e	feat: update torch 2.5.1 (#2069 )	2024-11-18 21:29:13 +08:00
ws	29ebe3dff4	fix: align enable_overlap_scheduler naming between code and docs (#2038 )	2024-11-15 03:39:10 -08:00
Lianmin Zheng	aae5434bdf	Fix unit tests (#2034 )	2024-11-14 11:08:37 -08:00
Lianmin Zheng	c3eac1b010	Fix torch.compile for MoE (#2033 )	2024-11-14 01:30:24 -08:00
James Xu	ddeb9d42de	Add engine encode (#1995 ) Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>	2024-11-11 11:48:17 -08:00
Lianmin Zheng	520f0094e4	[CI] balance unit tests (#1977 )	2024-11-09 16:46:14 -08:00
Lianmin Zheng	9c939a3d8b	Clean up metrics code (#1972 )	2024-11-09 15:43:20 -08:00
Yudi Xue	95a4ed129a	Fix metrics (#1963 )	2024-11-08 23:21:11 -08:00
Chayenne	c77c1e05ba	fix black in pre-commit (#1940 )	2024-11-08 07:42:47 +08:00
Lianmin Zheng	f7102fbd2b	Fix mixed chunked prefill (#1850 )	2024-10-30 21:20:41 -07:00
Lianmin Zheng	86fc0d79d0	Add a watch dog thread (#1816 )	2024-10-27 02:00:50 -07:00
Lianmin Zheng	c555ce2ca2	Revert "Fix memory leak when doing chunked prefill" (#1797 )	2024-10-25 10:24:44 -07:00
Lianmin Zheng	40900baea7	[Fix] Fix the log parsing in chunked prefill uni tests (#1794 )	2024-10-25 08:31:08 -07:00
Liangsheng Yin	a2f5e7555f	Fix memory leak when doing chunked prefill (#1787 )	2024-10-25 08:01:17 -07:00
Lianmin Zheng	2148914e1b	Fix log parsing in the chunked prefill unit tests (#1793 )	2024-10-25 08:00:55 -07:00
Lianmin Zheng	1701b0db31	Enhance the test case for chunked prefill (#1785 )	2024-10-24 21:23:09 -07:00
Byron Hsu	dde8bb16fe	default sampling param should be deepcopied (#1581 )	2024-10-05 17:27:43 -07:00
Ying Sheng	04b262cd91	[Fix] Fix major performance bug in certain cases (#1563 ) Co-authored-by: hnyls2002 <hnyls2002@gmail.com>	2024-10-04 08:51:11 +00:00
Theresa Barton	2c7d0a5b8b	[Fix] Fix all the Huggingface paths (#1553 )	2024-10-02 10:12:07 -07:00
Yineng Zhang	42a2d82ba7	minor: add mla fp8 test (#1494 )	2024-09-23 20:40:17 +08:00
Lianmin Zheng	167591e864	Better unit tests for adding a new model (#1488 )	2024-09-22 01:50:37 -07:00
Ke Bao	b8ccaf4d73	Add MLA gsm8k eval (#1484 )	2024-09-21 11:16:13 +08:00
Ke Bao	a68cb201dd	Fix triton head num (#1482 )	2024-09-21 10:25:20 +08:00
Yineng Zhang	a6db88626e	minor: add quant eval compared with base (#1475 )	2024-09-20 01:57:19 +08:00
Lianmin Zheng	1acccb364a	Fix oom issues with fp8 for llama (#1454 )	2024-09-18 03:45:19 -07:00
Lianmin Zheng	899cf5c438	Remove deprecated configs (#1431 )	2024-09-15 08:52:18 -07:00
Lianmin Zheng	9463bc1385	Enable torch.compile for triton backend (#1422 )	2024-09-14 15:38:37 -07:00
Lianmin Zheng	68be2f6d3b	[CI] Include triton backend and online serving benchmark into CI (#1408 )	2024-09-12 21:36:41 -07:00
Yineng Zhang	2561ed012c	feat: update nightly gsm8k eval (#1304 )	2024-09-03 01:18:41 +10:00
Mingyi	97589a60a2	[CI] Parallelize unit tests in CI (#1219 )	2024-08-26 04:54:02 +00:00

1 2 3

130 Commits