vllm 0.6.3 KeyError on qwen3_5_moe model type. Model is hybrid linear+full attention MoE with 256 experts (top-8). enginex-vllm-bi100-qwen36-main.zip in repo likely contains the fix.
vllm 0.6.3 KeyError on qwen3_5_moe model type. Model is hybrid linear+full attention MoE with 256 experts (top-8). enginex-vllm-bi100-qwen36-main.zip in repo likely contains the fix.