sglang

Author	SHA1	Message	Date
Chayenne	73dcf2b326	Remove token in token out in Native API (#5967 )	2025-05-01 21:59:43 -07:00
Chang Su	170d1f218a	feat: Refactor DeepSeekV3 function call (#5908 )	2025-05-01 21:28:57 -07:00
Sai Enduri	73bc1d00fc	Add 1 gpu perf and 2 gpu accuracy tests for AMD MI300x CI. (#5960 )	2025-05-01 20:56:59 -07:00
XinyuanTong	c5645e928f	feat: add concurrency evaluation logic in mmmu benchmark (#5782 )	2025-05-01 18:20:08 -07:00
KCFindstr	d33955d28a	Properly return error response in vertex_generate HTTP endpoint (#5956 )	2025-05-01 11:48:58 -07:00
Stefan He	6fc175968c	Optimize a pad operation to accelerate 25us (#5945 )	2025-05-01 10:48:55 -07:00
江家瑋	ad506a4e6b	docs: Fix Qwen model typo (#5944 ) Signed-off-by: JiangJiaWei1103 <waynechuang97@gmail.com>	2025-05-01 10:23:00 -07:00
Ke Bao	ebaba85655	Update ci test and doc for MTP api change (#5952 )	2025-05-01 09:30:27 -07:00
Ke Bao	de2faef97e	Remove extra contiguous (#5953 )	2025-05-01 09:28:46 -07:00
Yuan Luo	67b7d5b1df	[PD] Vectorise group_concurrent_contiguous in NumPy (#5834 ) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>	2025-05-01 22:42:37 +08:00
ryang	4322c31e24	Support XiaomiMiMo/MiMo model inference (#5921 )	2025-05-01 07:41:13 -07:00
Yineng Zhang	9858113c33	chore: bump v0.4.6.post2 (#5939 )	2025-04-30 22:04:40 -07:00
Yineng Zhang	8441baad6e	fix: update model runner (#5934 )	2025-04-30 19:49:26 -07:00
mlmz	256c4c2519	fix: correct stream response when enable_thinking is set to false (#5881 )	2025-04-30 19:44:37 -07:00
Johnny	9f21e75453	add Thor & Spark (#5915 )	2025-04-30 19:43:40 -07:00
Qiaolin Yu	7bcd8b1cb2	Fix lora batch processing when input lora_path contains None (#5930 )	2025-04-30 19:42:42 -07:00
Ying Sheng	11383cec3c	[PP] Add pipeline parallelism (#5724 )	2025-04-30 18:18:07 -07:00
XinyuanTong	e97e57e699	Remove unused method `calculate_num_image_tokens` from qwen2_vl.py (#5783 )	2025-04-30 17:46:59 -07:00
Yineng Zhang	9a6ad8916d	chore: upgrade sgl-kernel 0.1.1 (#5933 )	2025-04-30 16:13:30 -07:00
Yineng Zhang	d353d08b4e	chore: bump sgl-kernel 0.1.1 (#5932 )	2025-04-30 14:01:49 -07:00
PGFLMG	08acdb5c3d	[Feat] Scale up fa3 kernel to sm8x arch (#5912 ) Co-authored-by: zhyncs <me@zhyncs.com>	2025-04-30 13:59:36 -07:00
Sai Enduri	2afba1b1c1	Add TP2 MOE benchmarks for AMD. (#5909 )	2025-04-30 11:38:20 -07:00
laixin	e330f2b86c	[qwen3] support qwen3 ep moe (#5917 ) Co-authored-by: sleepcoo <sleepcoo@gmail.com>	2025-04-30 09:15:21 -07:00
PGFLMG	3ddf5b9d61	[Misc] use parallel build for cmake in sgl-kernel (#5919 )	2025-04-30 08:56:46 -07:00
JieXin Liang	3cff963335	[fix] kimi-vl test in test_vision_openai_server.py (#5910 )	2025-04-29 23:59:10 -07:00
Yi Zhang	d50e36a79d	support vlm benchmark profile (#5905 )	2025-04-29 23:48:27 -07:00
liwenju0	8fefdd32c7	[Feature] add support kimi vl model (#5383 ) Co-authored-by: wenju.li <wenju.li@deepctr.cn>	2025-04-29 21:31:19 -07:00
zhjunqin	403b855a22	Add sm_120 for blackwell (#5903 )	2025-04-29 20:45:24 -07:00
lambert0312	1698e94e67	Add A800 fused moe config for qwen3 235b (#5900 )	2025-04-29 20:18:11 -07:00
Qiaolin Yu	58195dd588	[Fix] Unload lora in HF_Runner if needed (#5899 )	2025-04-29 20:17:42 -07:00
Baizhou Zhang	799789afed	Bump Flashinfer to 0.2.5 (#5870 ) Co-authored-by: Yuhao Chen <yxckeis8@gmail.com>	2025-04-29 19:50:57 -07:00
ybyang	cc4a80caf6	[PD] Fix Assertion failed: /DeepEP/csrc/kernels/internode.cu:483, condition: ibgda_get_state()->num_rc_per_pe >= num_channels #134 (#5830 )	2025-04-29 19:38:54 -07:00
lambert0312	3c8a52311a	Fix check_env script (#5901 )	2025-04-29 18:54:54 -07:00
Yineng Zhang	a043f7f2ab	chore: use torch 2.6 for sgl-kernel build (#5898 )	2025-04-29 17:51:18 -07:00
saienduri	e3a5304475	Add AMD MI300x Nightly Testing. (#5861 )	2025-04-29 17:34:32 -07:00
Chang Su	28b26dbf48	[Bugfix]: fix missing queue_time_start for requests from grammar_queue (#5696 )	2025-04-29 17:31:44 -07:00
Chang Su	2b06484bd1	feat: support pythonic tool call and index in tool call streaming (#5725 )	2025-04-29 17:30:44 -07:00
JieXin Liang	e4b6133b78	[fix] relax mem_fraction_static for h200 (#5893 ) Co-authored-by: alcanerian <alcanerian@gmail.com>	2025-04-29 17:01:12 -07:00
Ke Bao	dd408ee481	Auto set draft model path for MTP (#5793 )	2025-04-29 16:25:40 -07:00
Chang Su	9419e75d60	[CI] Add test_function_calling.py to run_suite.py (#5896 )	2025-04-29 15:54:53 -07:00
Johnny	2c7dbb7cc2	[FEATURE] Enhance platform compatibility for ARM (#5746 )	2025-04-29 15:06:16 -07:00
Yineng Zhang	9a62191ba7	chore: update CODEOWNERS (#5895 )	2025-04-29 14:12:04 -07:00
simveit	ae523675e5	[Doc] Tables instead of bulletpoints for sampling doc (#5841 )	2025-04-29 13:49:39 -07:00
Adarsh Shirawalmath	5c08aa4958	[Docs] Update docs for Qwen3 and Qwen3MoE (#5836 )	2025-04-29 13:48:30 -07:00
Yineng Zhang	f4c191a712	chore: update Dockerfile (#5894 )	2025-04-29 12:55:13 -07:00
Simo Lin	771669cbe0	[fix]: PyO3 macOS linking and consolidate on tracing for logging	2025-04-29 11:26:38 -07:00
Simo Lin	1468769bde	[Misc] add service discovery for sgl router	2025-04-29 10:21:19 -07:00
lambert0312	91dda4cd06	Add A800 fused moe config for qwen3 30b (#5880 )	2025-04-29 02:02:24 -07:00
pengcuo	8e5a6d3441	[Fix] Fix a bug for flashmla to run R1 model (#5875 ) Co-authored-by: pengcuo <dgpengcuo@gmail.com>	2025-04-29 01:03:13 -07:00
XinyuanTong	8465f035d1	Add qwen3 30b fused moe config (#5859 )	2025-04-29 00:24:00 -07:00

1 2 3 4 5 ...

3127 Commits