Files
project_6/upstream_ref/xllm/docs/zh/features/multimodal.md
EX Engine 002f9879b2 ref(upstream): FULL TREE — Deep-Spark xllm (1470) + ds_vllm csrc/models (703)
Replaces cherry-picked upstream_ref with complete source trees.

xllm/ — Iluvatar official C++ inference engine (15MB, 1470 files)
  Complete: kernels → layers → models → runtime → scheduler → api
  Excluded: .git, binary images, third_party submodule checkouts

ds_vllm/ — Iluvatar official vllm fork (8MB, 703 files)
  Included: csrc/ (ALL CUDA kernels), fused_moe/, qwen3_5 model, _custom_ops
  Excluded: tests, benchmarks, docs, examples (not needed for reference)

Critical call chains now fully traceable:
  MoE: moe_topk_softmax_kernels.cuh → ixformer.h → fused_moe.cpp → layer
  GDN: qwen3_gated_delta_net_base.cpp → qwen3_5_gated_delta_net.cpp
  Attention: ixformer.h → xllm_paged_attention → attention.cpp
2026-08-10 02:54:03 +00:00

807 B
Executable File
Raw Blame History

多模态支持

本文档主要介绍xLLM推理引擎中多模态的支持进展包括支持模型及模态类型以及离在线接口等。

支持模型

  • Qwen2.5-VL: 包括7B/32B/72B。
  • Qwen3-VL: 包括2B/4B/8B/32B。
  • Qwen3-VL-MoE: 包括A3B/A22B。
  • MiniCPM-V-2_6: 7B。

模态类型

  • 图片: 支持单图、多图的输入,以及图片+Prompt组合、纯文本Promot等输入方式。

!!! warning "注意事项" - 目前多模态后端不支持prefix cache以及chunk prefill正在支持中。 - 目前xLLM统一基于JinJa渲染ChatTemplate部署MiniCPM-V-2_6模型目录需提供ChatTemplate文件。 - 图片支持Base64输入以及图片Url。 - 目前多模态模型主要支持了图片模态,视频、音频等模态正在推进中。