Files
project_6/upstream_ref/xllm/docs/en/features/prefix_cache.md
EX Engine 002f9879b2 ref(upstream): FULL TREE — Deep-Spark xllm (1470) + ds_vllm csrc/models (703)
Replaces cherry-picked upstream_ref with complete source trees.

xllm/ — Iluvatar official C++ inference engine (15MB, 1470 files)
  Complete: kernels → layers → models → runtime → scheduler → api
  Excluded: .git, binary images, third_party submodule checkouts

ds_vllm/ — Iluvatar official vllm fork (8MB, 703 files)
  Included: csrc/ (ALL CUDA kernels), fused_moe/, qwen3_5 model, _custom_ops
  Excluded: tests, benchmarks, docs, examples (not needed for reference)

Critical call chains now fully traceable:
  MoE: moe_topk_softmax_kernels.cuh → ixformer.h → fused_moe.cpp → layer
  GDN: qwen3_gated_delta_net_base.cpp → qwen3_5_gated_delta_net.cpp
  Attention: ixformer.h → xllm_paged_attention → attention.cpp
2026-08-10 02:54:03 +00:00

1.0 KiB

Prefix Cache Optimization

Feature Introduction

xLLM supports prefix cache matching. The prefix cache is based on murmur_hash and uses an LRU eviction policy, delivering superior matching efficiency and increased prefix cache hit rates. Additionally, the prefix cache has been optimized to support the continuous_scheduler, chunked_scheduler, and zero_evict_scheduler. The cache is updated immediately after prefill operations, enhancing matching timeliness. For the chunked_scheduler, multi-stage chunked prefill matching is supported, reducing computational overhead and minimizing KV cache usage as much as possible.

Usage

The prefix cache is implemented in xLLM and exposed through gflags parameters to control its functionality.

  • Enable prefix cache with specific policy and settings:
--enable_prefix_cache=true

Performance Impact

After enabling prefix cache, on the Qwen3-8B model with a TPOT constraint of 50ms, the E2E latency decreased by 10%.

!!! warning "Note" PD separation scheduler is not currently supported.