Chang Su
|
ca240eefb4
|
[router][grpc] Support parallel queue puts in grpc_request_manager and remove mutex for grpc_client (#11798)
|
2025-10-17 20:49:43 -07:00 |
|
Chang Su
|
d1984e218c
|
[router][grpc] Remove timeout for connections and remove max_tokens deprecation warning log (#11775)
|
2025-10-17 12:36:36 -07:00 |
|
Simo Lin
|
a5978a20f0
|
[router] fix grpc client time out to 1h (#11768)
|
2025-10-17 10:26:12 -07:00 |
|
Simo Lin
|
e483c1eae5
|
[router] Fix UTF-8 Boundary Panic in Stop Sequence Decoder (#11766)
|
2025-10-17 10:21:00 -07:00 |
|
Keyang Ru
|
7780230a15
|
Revert "[router] fix get_models endpoint for openai router (#11687)" (#11740)
|
2025-10-16 18:36:53 -07:00 |
|
Chang Su
|
dc01313da1
|
[router] Add rustfmt and set group imports by default (#11732)
|
2025-10-16 17:33:29 -07:00 |
|
Chang Su
|
c7962868c1
|
[router] Fix tool_choice normalization in ChatCompletionRequest and fix ut (#11731)
|
2025-10-16 14:20:13 -07:00 |
|
Simo Lin
|
64affab495
|
[router] fix p and d worker filtering and bootstrap port handling (#11729)
|
2025-10-16 14:19:39 -07:00 |
|
Keyang Ru
|
4c9bcb9d56
|
[Router] Refactor protocol definitions: split spec.rs into modular files (#11677)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
|
2025-10-16 13:44:44 -07:00 |
|
Keyang Ru
|
0975ba99bc
|
[router] fix get_models endpoint for openai router (#11687)
|
2025-10-16 09:00:08 -07:00 |
|
Simo Lin
|
f5d30dae89
|
[router] Refactor StopSequenceDecoder to Use Sequence for Incremental Decoding (#11676)
|
2025-10-15 16:31:03 -07:00 |
|
Chang Su
|
2479b89405
|
[router][grpc] Simplify model_id determination (#11684)
|
2025-10-15 15:56:58 -07:00 |
|
Keyang Ru
|
d2478cd4ff
|
[router] Fix response api related spec (#11621)
|
2025-10-15 09:59:38 -07:00 |
|
Simo Lin
|
40e0082d8d
|
[router] add worker self discovery for metadata (#11638)
|
2025-10-14 22:07:25 -04:00 |
|
Simo Lin
|
3962e39d7c
|
[router] cleanup app context and move to startup (#11617)
|
2025-10-14 10:19:28 -07:00 |
|
Keyang Ru
|
eb8cac6fe2
|
[router] add py binding and readme for openai router and history backend (#11453)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-10-14 09:42:34 -07:00 |
|
Simo Lin
|
a04efc4933
|
[router] when given both local tokenizer and chat template, log all (#11601)
|
2025-10-14 02:22:58 -07:00 |
|
Simo Lin
|
28ad2297a0
|
[router] delete useless table content comment in spec (#11597)
|
2025-10-14 01:08:18 -07:00 |
|
Simo Lin
|
4b62af92ef
|
[router] change worker api to async instead of sync (#11566)
|
2025-10-14 00:32:21 -07:00 |
|
Simo Lin
|
0b9915c132
|
[router] update generate spec to align with sgl io struct (#11591)
|
2025-10-14 02:51:33 -04:00 |
|
Chang Su
|
27ef1459e6
|
[router][protocols] Add Axum validate extractor and use it for /v1/chat/completions endpoint (#11588)
|
2025-10-13 22:51:15 -07:00 |
|
Chang Su
|
887c2b4575
|
[router][grpc] Add serve_grpc to launch_server and log id for HealthCheck (#11564)
|
2025-10-13 16:07:19 -07:00 |
|
Chang Su
|
4b694e7d5a
|
[router][grpc] Add error handling to generate_tool_constraints (#11562)
|
2025-10-13 12:26:09 -07:00 |
|
Jonah Bernard
|
f4aa78801e
|
[router] Add Rust CLI flags for queue size, timeout, and rate limit for token bucket rate limiter (#11483)
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
|
2025-10-13 11:08:48 -07:00 |
|
Simo Lin
|
728af88781
|
[router] allow user to specify chat template path (#11549)
|
2025-10-13 10:47:57 -07:00 |
|
Chang Su
|
7b59b0b8b0
|
[router][grpc] Further delegate non-stream processing to processing.rs (#11553)
|
2025-10-13 10:36:27 -07:00 |
|
Simo Lin
|
7c94eaeeb0
|
[router] allow tokenizer path to be dir (#11530)
|
2025-10-13 09:30:09 -04:00 |
|
Keyang Ru
|
63e84352b7
|
[router] openai router: support grok model (#11511)
|
2025-10-12 22:44:43 -04:00 |
|
Antoine Roux
|
ec1cd90ac9
|
Fix the GPT function calling regex to allow dash in the name (#10577)
|
2025-10-12 20:34:58 +08:00 |
|
Wenyi Xu
|
9b5efe3464
|
[Router]: Small Typo in a comment within tree.rs (#11489)
|
2025-10-11 21:59:48 -07:00 |
|
fzyzcjy
|
d957177a22
|
Super tiny delete unused openai router in sgl-router (#11448)
|
2025-10-11 15:59:30 +08:00 |
|
Chang Su
|
92777135a0
|
[router][grpc] Consolidate parser checks for chat completions (#11439)
|
2025-10-10 20:44:29 -04:00 |
|
Simo Lin
|
c495833186
|
[router] leverage RAII to actively cancel request during client disconnect (#11399)
|
2025-10-10 20:43:38 -04:00 |
|
Simo Lin
|
2eeb27515a
|
[router] disable rate limiter by default (#11435)
|
2025-10-10 20:43:07 -04:00 |
|
Keyang Ru
|
eb7d9261c0
|
[router] conversation item API: create, retrieve and delete (#11369)
|
2025-10-09 17:43:16 -04:00 |
|
Simo Lin
|
88bb627d0d
|
[router] change grpc client from mutable to clone (#11394)
|
2025-10-09 11:00:24 -07:00 |
|
Chang Su
|
ab926dd697
|
[router][grpc] Fix streaming bugs: empty tool names, state pollution, and panics (#11373)
|
2025-10-09 06:53:23 -04:00 |
|
Chang Su
|
a0557642ea
|
[router][lint] Add unused_qualifications to cargo lint warnings (#11366)
|
2025-10-08 22:17:11 -07:00 |
|
Keyang Ru
|
84768d1017
|
[router] Refactor OpenAI router: split monolithic file and move location (#11359)
|
2025-10-09 00:46:39 -04:00 |
|
Simo Lin
|
368fd20622
|
[router][grpc] disable health check generation and increase timeout (#11353)
|
2025-10-08 19:23:08 -07:00 |
|
Chang Su
|
fccac7d126
|
[router][grpc] Add dependencies in Cargo.toml to support chat template rendering (#11342)
|
2025-10-08 15:38:37 -07:00 |
|
Keyang Ru
|
7ac6b900f4
|
[router] Support history management using conversation (#11339)
|
2025-10-08 15:24:02 -07:00 |
|
Chang Su
|
a1080b72a0
|
[router] Fix all unused_qualifications (#11341)
|
2025-10-08 13:55:27 -07:00 |
|
Chang Su
|
a65ca73911
|
[router][grpc] Cleanup debug logs in grpc_server and grpc_router (#11340)
|
2025-10-08 13:26:19 -07:00 |
|
Simo Lin
|
677aa0e25f
|
[router] improve reasoning parser lock and reduce req cloning (#11336)
|
2025-10-08 11:18:15 -07:00 |
|
Simo Lin
|
01c9ee1ab4
|
[router] refactor generate to use new pipeline arch (#11323)
|
2025-10-08 09:38:50 -07:00 |
|
Chang Su
|
edd86b8853
|
[router][grpc] Refactor chat handler in grpc/ to use centralized orchestrator (#11314)
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
|
2025-10-07 20:50:20 -07:00 |
|
Simo Lin
|
fde9b96392
|
[router] cleanup worker health check to return early (#11310)
|
2025-10-07 16:53:10 -07:00 |
|
Keyang Ru
|
4ed67c27e3
|
[router] support Openai router conversation API CRUD (#11297)
|
2025-10-07 15:31:35 -07:00 |
|
Chang Su
|
420c99acfe
|
[router][grpc] Fix error message format in grpc chat handler (#11307)
|
2025-10-07 13:54:02 -07:00 |
|