Replaces cherry-picked upstream_ref with complete source trees. xllm/ — Iluvatar official C++ inference engine (15MB, 1470 files) Complete: kernels → layers → models → runtime → scheduler → api Excluded: .git, binary images, third_party submodule checkouts ds_vllm/ — Iluvatar official vllm fork (8MB, 703 files) Included: csrc/ (ALL CUDA kernels), fused_moe/, qwen3_5 model, _custom_ops Excluded: tests, benchmarks, docs, examples (not needed for reference) Critical call chains now fully traceable: MoE: moe_topk_softmax_kernels.cuh → ixformer.h → fused_moe.cpp → layer GDN: qwen3_gated_delta_net_base.cpp → qwen3_5_gated_delta_net.cpp Attention: ixformer.h → xllm_paged_attention → attention.cpp
1.0 KiB
Prefix Cache Optimization
Feature Introduction
xLLM supports prefix cache matching. The prefix cache is based on murmur_hash and uses an LRU eviction policy, delivering superior matching efficiency and increased prefix cache hit rates.
Additionally, the prefix cache has been optimized to support the continuous_scheduler, chunked_scheduler, and zero_evict_scheduler. The cache is updated immediately after prefill operations, enhancing matching timeliness. For the chunked_scheduler, multi-stage chunked prefill matching is supported, reducing computational overhead and minimizing KV cache usage as much as possible.
Usage
The prefix cache is implemented in xLLM and exposed through gflags parameters to control its functionality.
- Enable prefix cache with specific policy and settings:
--enable_prefix_cache=true
Performance Impact
After enabling prefix cache, on the Qwen3-8B model with a TPOT constraint of 50ms, the E2E latency decreased by 10%.
!!! warning "Note" PD separation scheduler is not currently supported.