sglang

Author	SHA1	Message	Date
Lianmin Zheng	048685430d	Improve process creation (#1534 )	2024-09-29 02:36:12 -07:00
Ying Sheng	9aa6553d2a	[Feature] Support reward model LxzGordon/URM-LLaMa-3.1-8B (#1525 )	2024-09-27 23:32:11 -07:00
Lianmin Zheng	bc068e9618	[CI] Move AMD test to a separate file (#1500 )	2024-09-24 02:06:28 -07:00
Yineng Zhang	42a2d82ba7	minor: add mla fp8 test (#1494 )	2024-09-23 20:40:17 +08:00
Ying Sheng	6f3cf1297e	[CI, AMD] Add AMD tests to CI (#1491 )	2024-09-22 04:45:10 -07:00
Lianmin Zheng	13f1357ef0	Add a unit test for data parallelism (#1489 )	2024-09-22 02:21:05 -07:00
Ke Bao	b8ccaf4d73	Add MLA gsm8k eval (#1484 )	2024-09-21 11:16:13 +08:00
Ke Bao	a68cb201dd	Fix triton head num (#1482 )	2024-09-21 10:25:20 +08:00
Lianmin Zheng	1acccb364a	Fix oom issues with fp8 for llama (#1454 )	2024-09-18 03:45:19 -07:00
Lianmin Zheng	9ba1f09760	[Fix] Fix logprob and normalized_logprob (#1428 )	2024-09-15 06:36:06 -07:00
Yineng Zhang	f3d32f888a	ci: fix finish (#1414 )	2024-09-14 01:01:30 +10:00
Lianmin Zheng	8779da95d6	Update pr-test.yml (#1412 )	2024-09-13 00:37:13 -07:00
Lianmin Zheng	ad0ff62a4c	Balance test in CI (#1411 )	2024-09-12 23:29:44 -07:00
Lianmin Zheng	68be2f6d3b	[CI] Include triton backend and online serving benchmark into CI (#1408 )	2024-09-12 21:36:41 -07:00
Lianmin Zheng	f64eae3a29	[Fix] Reduce memory usage for loading llava model & Remove EntryClassRemapping (#1308 )	2024-09-02 21:44:45 -07:00
Lianmin Zheng	761b2cebd6	[CI] merge all ci tests into one file (#1289 )	2024-09-01 02:36:56 -07:00