vllm 0.6.3 KeyError on qwen3_5_moe model type. Model is hybrid linear+full attention MoE with 256 experts (top-8). enginex-vllm-bi100-qwen36-main.zip in repo likely contains the fix.
Hardware: 4× Iluvatar BI-V100 32GB, Xeon Gold 6530, 503GB RAM Software: vllm 0.6.3+corex.3.2.3, torch 2.1.0+corex.3.2.3 Model: Qwen3.6-35B-A3B at /root/public-storage/models/Qwen/ Benchmark: benchmark_server_v0.5.0.py with automated sweeps