Logo
Explore Help
Register Sign In
dylanyunlong/project_6
1
0
Fork 0
You've already forked project_6
Code Issues Pull Requests Actions Projects Releases Wiki Activity
Files
8d2f30f06541c22a4d77ec9848b3a34a0bafd414
project_6/ixformer_sdk/contrib/tgi/__init__.py

1 line
32 B
Python
Raw Normal View History

feat(CRITICAL): 从 GitHub 扫描搬运 ixformer SDK + xllm 完整 GDN/MoE 代码 来源: 1. Chranos/ixformer (GitHub) → ixformer_sdk/ (230 files, 70K lines) - inference/functions/vllm.py: vllm_moe_topk_softmax 完整实现 (2033 lines) - inference/functions/moe.py: MoE ops 完整实现 (1380 lines) - contrib/vllm_flash_attn/: FA2 Python 接口 (1018 lines) - contrib/tgi/fused_moe.py: TGI fused MoE (429 lines) - csrc/include/ixformer/: C++ kernel headers + cmake 2. Deep-Spark/xllm (GitHub) → upstream_ref/xllm_latest/ (+15 files) - npu_torch/qwen3_5_decoder_layer_impl.cpp/.h - npu_torch/qwen3_5_gated_delta_net.cpp/.h - npu_torch/qwen3_next_*.cpp/.h (6 files) - npu_torch/attention.cpp/.h + fused_moe.cpp/.h + CMakeLists.txt - models/llm/qwen3_5.h + qwen3_5_mtp.h + qwen3_next.h - models/vlm/qwen3_5.h 调用链完整性: ixformer_sdk/inference/functions/vllm.py → ops.infer.moe_topk_softmax() (C++ 层) → 这就是 base 镜像 libixformer.so 里的实现 upstream_ref/xllm_latest/core/layers/ilu/fused_moe.cpp → ixformer::infer::topk_softmax() (直接 C++ 调用) → ixformer::infer::group_gemm() → 完整 7-step MoE pipeline
2026-08-11 02:31:56 +00:00
from .fused_moe import fused_moe
Reference in New Issue Copy Permalink
Powered by Gitea Version: 1.24.3 Page: 409ms Template: 2ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API