fee8f1b9e46209ef02024ef25bb83532eaba3d3b
vllm 0.6.3 KeyError on qwen3_5_moe model type. Model is hybrid linear+full attention MoE with 256 experts (top-8). enginex-vllm-bi100-qwen36-main.zip in repo likely contains the fix.
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%