Dynamic Bucket Load Balancer
A dynamic bucketing-based hybrid load balance proxy for vLLM.
The proxy fronts multiple vLLM backend servers and distributes OpenAI-compatible requests across them. It can run in two modes:
- Plain load balancing (default): for each request, estimate a load score and forward it to the least-loaded backend instance.
- Dynamic bucket load balancing (
--enable-dynamic-bucket): split the backend pool into a short-request group and a long-request group, route requests to a group by their length, and dynamically rebalance across groups based on the load gap and length affinity.
Files
dynamic_bucket_load_balancer.py— the core algorithm (pure standard library). Buckets requests by length, then dynamically adjusts bucket assignment using bucket load and length affinity.hybrid_proxy_server.py— the FastAPI proxy server that uses the algorithm to route requests to the backend servers.
How It Works
-
Static bucketing by length. Each request is first mapped to its standard bucket by request length. With dynamic bucketing enabled the proxy uses two buckets: short
[0, --server-group-threshold)and long[--server-group-threshold, --max-request-tokens). -
Server groups. The ordered backend list is split into the same number of groups as buckets, in order: the first instances form the short group, the last instances form the long group. With 4 backends and 2 buckets, backends 0 and 1 serve the short bucket, backends 2 and 3 serve the long bucket.
-
Tip: configure the first two instances for short sequences and the last two for long sequences (e.g. smaller
max-model-len/ KV cache for the short group, larger for the long group) to get the best throughput.
-
-
Dynamic rebalancing. For a new request, the balancer looks at neighbor buckets with a lighter load and computes a redirect probability
(load-gap probability) × (length-affinity factor). If it exceeds the threshold (0.12), the request is redirected to the neighbor bucket. This means a large load gap is suppressed when the request length is far from the neighbor bucket, while a modest gap can still trigger a redirect when the length is close to the boundary. -
Within a group, the least-loaded server (smallest active token count) is picked via a min-heap, the load is accumulated for the duration of the request, and released when streaming completes.
Prerequisites
-
Python 3.10+
-
Install dependencies:
pip install "fastapi<0.124.0" httpx uvicorn
Step 1: Start Your Backend Servers
Start at least two vLLM servers, each as a separate process on its own port. The proxy also works with a single backend, but load balancing is only meaningful with two or more.
vllm serve --host 0.0.0.0 --port 8100 ... # vLLM Server 0
vllm serve --host 0.0.0.0 --port 8101 ... # vLLM Server 1
Step 2: Start the Proxy Server
From examples/dynamic_bucket_load_balancer/, point the proxy at each backend
with --server-hosts / --server-ports:
python hybrid_proxy_server.py \
--host 0.0.0.0 --port 8000 \
--server-hosts 127.0.0.1 127.0.0.1 \
--server-ports 8100 8101
This starts the proxy on port 8000 and load balances across the two backends.
Enable Dynamic Bucket Load Balancing
Add --enable-dynamic-bucket to split the pool into short/long groups. The
server count must be >= 2 so each bucket has at least one instance. With 4
servers the first two form the short group and the last two the long group:
python hybrid_proxy_server.py \
--host 0.0.0.0 --port 8000 \
--server-hosts 127.0.0.1 127.0.0.1 127.0.0.1 127.0.0.1 \
--server-ports 8100 8101 8102 8103 \
--enable-dynamic-bucket \
--server-group-threshold 32768
Step 3: Send a Request to the Proxy
Send OpenAI-compatible requests to the proxy. For example:
curl -X POST http://localhost:8000/v1/completions \
-H "Content-Type: application/json" \
-d '{
"model": "your-model",
"prompt": "The quick brown fox jumps over the lazy dog",
"max_tokens": 16
}'
Or for chat completions:
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "your-model",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 16
}'
Step 4: Health Check
Check that the proxy is running and how many backends it fronts:
curl http://localhost:8000/healthcheck
Returns a JSON object, e.g.:
{"status": "ok", "server_instances": 2}
Configuration
| Argument | Default | Description |
|---|---|---|
--host |
localhost |
Proxy listen host. |
--port |
8000 |
Proxy listen port. |
--server-hosts |
localhost |
Hosts of the backend vLLM servers (one per server, in order). |
--server-ports |
8001 |
Ports of the backend vLLM servers (one per server, in order). |
--enable-dynamic-bucket |
False |
Enable dynamic bucket load balancing. |
--server-group-threshold |
32768 |
Length boundary between the short and long buckets. |
--max-request-tokens |
131072 |
Upper bound of the long bucket (max request length). |
--max-retries |
3 |
Max retries for a backend HTTP request. |
--retry-delay |
0.001 |
Base delay (seconds) for exponential backoff retries. |
The number of --server-hosts must equal the number of --server-ports.