add qwen3_moe

2026-02-10 18:55:35 +08:00
parent 6f6997bafb
commit 463fbf8cd1
1 changed files with 1 additions and 0 deletions
--- a/README.md
+++ b/README.md
@@ -175,3 +175,4 @@ curl http://localhost:80/v1/chat/completions \
 | v0.0.3.1 | 2026-02-06 | **CNNL Tensor 溢出修复**：解决极小模型在大显存设备上部署时 KV cache 元素数超过 int32 限制的问题，在 mlu_worker 和 cache_engine 中添加双重防护 |
 | v0.0.4 | 2026-02-10 | **Gemma3 模型支持**：新增 Gemma3ForCausalLM 模型实现（含 QK Normalization、per-layer rope 配置、layer_types 滑动窗口），修复 `patch_rope_scaling_dict` 在 rope_scaling 缺少 `rope_type` 键时崩溃的问题，更新模型注册表及 config.py 中 interleaved attention 和 dtype 自动处理逻辑 |
 | v0.0.4.1 | 2026-02-10 | **Gemma3 rope 兼容性修复**：修复新版 transformers `Gemma3TextConfig` 缺少 `rope_theta` 属性的问题，从 `rope_parameters` 字典兼容提取 rope 配置（支持 Transformers v4/v5）；修复 `rope_scaling` 嵌套字典导致 `get_rope` 缓存 unhashable 的问题；适配 MLU `forward_mlu` 接口，将 q/k 合并为单张量调用 rotary_emb 后再拆分 |
+| v0.0.5 | 2026-02-10 | **Qwen3MoE 模型支持**：新增 Qwen3MoeForCausalLM 模型实现（含 QK Normalization、ReplicatedLinear shared_expert_gate），修复 FusedMoE `forward_mlu` 签名缺少 `layer` 参数的已有 bug（影响所有 MLU 上的 MoE 模型），更新模型注册表 |