Files
project_6/upstream_ref/xllm/docs/en/features/chunked_scheduler.md
EX Engine 002f9879b2 ref(upstream): FULL TREE — Deep-Spark xllm (1470) + ds_vllm csrc/models (703)
Replaces cherry-picked upstream_ref with complete source trees.

xllm/ — Iluvatar official C++ inference engine (15MB, 1470 files)
  Complete: kernels → layers → models → runtime → scheduler → api
  Excluded: .git, binary images, third_party submodule checkouts

ds_vllm/ — Iluvatar official vllm fork (8MB, 703 files)
  Included: csrc/ (ALL CUDA kernels), fused_moe/, qwen3_5 model, _custom_ops
  Excluded: tests, benchmarks, docs, examples (not needed for reference)

Critical call chains now fully traceable:
  MoE: moe_topk_softmax_kernels.cuh → ixformer.h → fused_moe.cpp → layer
  GDN: qwen3_gated_delta_net_base.cpp → qwen3_5_gated_delta_net.cpp
  Attention: ixformer.h → xllm_paged_attention → attention.cpp
2026-08-10 02:54:03 +00:00

983 B

ChunkedPrefill Scheduler

Feature Introduction

xLLM supports the chunked prefill scheduling strategy. Chunked prefill is a technique that optimizes large language model inference by splitting long prompts into smaller chunks for batch processing, rather than processing the entire prompt at once. This method can effectively reduce peak GPU memory usage, improve device utilization, and better schedule and mix processing with requests from the decode stage.

Usage

The aforementioned strategy has been implemented in xLLM and is exposed through gflags parameters to control the feature's on/off state.

  • Enable chunked prefill and set the chunked size, if not set chunked size, its default value is equal to max_tokens_per_batch.
--enable_chunked_prefill=true
--max_tokens_per_chunk_for_prefill=20480 # optional

Performance Impact

After enabling chunked prefill, on the Qwen3-8B model with a TPOT constraint of 50ms, the TTFT latency decreased by 46%.