ref(upstream): FULL TREE — Deep-Spark xllm (1470) + ds_vllm csrc/models (703)
Replaces cherry-picked upstream_ref with complete source trees. xllm/ — Iluvatar official C++ inference engine (15MB, 1470 files) Complete: kernels → layers → models → runtime → scheduler → api Excluded: .git, binary images, third_party submodule checkouts ds_vllm/ — Iluvatar official vllm fork (8MB, 703 files) Included: csrc/ (ALL CUDA kernels), fused_moe/, qwen3_5 model, _custom_ops Excluded: tests, benchmarks, docs, examples (not needed for reference) Critical call chains now fully traceable: MoE: moe_topk_softmax_kernels.cuh → ixformer.h → fused_moe.cpp → layer GDN: qwen3_gated_delta_net_base.cpp → qwen3_5_gated_delta_net.cpp Attention: ixformer.h → xllm_paged_attention → attention.cpp
This commit is contained in:
83
upstream_ref/xllm/docs/en/features/disagg_pd.md
Normal file
83
upstream_ref/xllm/docs/en/features/disagg_pd.md
Normal file
@@ -0,0 +1,83 @@
|
||||
# Disaggregated PD
|
||||
## Background
|
||||
LLM online inference services typically need to meet two performance metrics: TTFT and TPOT. Traditional Contiguous Batching scheduling strategies mix Prefill and Decode requests during scheduling, causing Prefill and Decode phases to compete for computational resources. This prevents maximized utilization of computing resources and impacts performance metrics. To resolve this conflict, the Prefill and Decode phases are split to run on independent computational resources, enabling parallel execution. This simultaneously reduces TTFT and TPOT while improving throughput.
|
||||
|
||||
## Introduction
|
||||
The xLLM PD Separation feature is primarily implemented through the following three modules:
|
||||
|
||||
- **etcd**: Stores metadata such as instance information.
|
||||
- **xLLM Service**: Schedules requests and manages all computing instances.
|
||||
- **xLLM**: Handles request computation instances.
|
||||
|
||||
The overall architecture is shown below:
|
||||

|
||||
|
||||
## Usage
|
||||
### Preparation
|
||||
#### Install Dependencies
|
||||
- **xLLM**: Refer to [Installation && Compilation](../getting_started/quick_start.md)
|
||||
- **xLLM Service**: Refer to [PD disaggregation](../getting_started/disagg_pd.md)
|
||||
|
||||
#### Obtain Environment Information
|
||||
Deploying Disaggregated PD Service requires obtaining the Device IP of the machine to create communication resources. Execute the command `cat /etc/hccn.conf | grep address` on the current AI Server to get the Device IP, for example:
|
||||
```
|
||||
address_0=xx.xx.xx.xx
|
||||
address_1=xx.xx.xx.xx
|
||||
```
|
||||
`address_xx` represents the Device IP.
|
||||
|
||||
### Start Disaggregated PD Service
|
||||
1. Start etcd
|
||||
```bash
|
||||
./etcd
|
||||
```
|
||||
2. Start xLLM Service
|
||||
```bash
|
||||
ENABLE_DECODE_RESPONSE_TO_SERVICE=true ./xllm_master_serving --etcd_addr="127.0.0.1:12389" --http_server_port 28888 --rpc_server_port 28889 --tokenizer_path=/path/to/tokenizer_config_dir/
|
||||
```
|
||||
3. Start xLLM
|
||||
- Taking Qwen2-7B as an example
|
||||
- Start Prefill Instance
|
||||
```bash
|
||||
/path/to/xllm --model=Qwen2-7B-Instruct \
|
||||
--port=8010 \
|
||||
--devices="npu:0" \
|
||||
--master_node_addr="127.0.0.1:18888" \
|
||||
--enable_prefix_cache=false \
|
||||
--enable_chunked_prefill=false \
|
||||
--enable_disagg_pd=true \
|
||||
--instance_role=PREFILL \
|
||||
--etcd_addr="127.0.0.1:12389" \
|
||||
--transfer_listen_port=26000 \
|
||||
--disagg_pd_port=7777 \
|
||||
--node_rank=0 \
|
||||
--nnodes=1
|
||||
```
|
||||
- Start Decode Instance
|
||||
```bash
|
||||
/path/to/xllm --model=Qwen2-7B-Instruct \
|
||||
--port=8020 \
|
||||
--devices="npu:1" \
|
||||
--master_node_addr="127.0.0.1:18898" \
|
||||
--enable_prefix_cache=false \
|
||||
--enable_chunked_prefill=false \
|
||||
--enable_disagg_pd=true \
|
||||
--instance_role=DECODE \
|
||||
--etcd_addr="127.0.0.1:12389" \
|
||||
--transfer_listen_port=26100 \
|
||||
--disagg_pd_port=7787 \
|
||||
--node_rank=0 \
|
||||
--nnodes=1
|
||||
```
|
||||
Important notes:
|
||||
|
||||
- PD disaggregation requires reading the `/etc/hccn.conf` file. Make sure this file on the physical machine is mapped into the container.
|
||||
|
||||
- `etcd_addr` must match the `etcd_addr` of `xllm_service`
|
||||
|
||||
## Notice
|
||||
Disaggregated PD **does not support** enabling prefix cache or chunked prefill. These features must be disabled using the following parameters:
|
||||
```shell
|
||||
--enable_prefix_cache=false
|
||||
--enable_chunked_prefill=false
|
||||
```
|
||||
Reference in New Issue
Block a user