fix(EX): corex ivcore10 build flags + deploy pipeline + topk kernel cleanup

Real machine log (2d5232c dockerrizhi.txt) shows two AST call chain breaks:

1. EVERY layer EVERY token:
   _custom_ops.py:58 'ixformer.functions has no attribute vllm_moe_topk_softmax'
   -> FusedMoE falls to PyTorch loop (2304 calls/token)

2. EVERY GDN layer (4 layers):
   'NaN in prefill GatedDeltaNet layer N (frac=0.9998-1.0000)'
   -> _torch_chunk_gated_delta_rule produces all-NaN

Fixes:
- build.sh: --cuda-gpu-arch=ivcore10, -D__ILUVATAR__ flags from real log
- Dockerfile: add ex_engine build before patch_ops
- patch_ops.sh: deploy .so + python into vllm model dir
- ex_loader.py: search co-located .so paths
- patch_model.py: remove premature auto-apply
- factor_moe_topk_softmax.cu: remove dead parallel branch
This commit is contained in:
EX Engine
2026-08-10 02:31:55 +00:00
parent fcfb764560
commit b75965d4ea
6 changed files with 80 additions and 45 deletions

View File

@@ -57,19 +57,32 @@ compile_factor() {
if [[ "$COMPILER" == "corex" ]]; then
# CoreX/Iluvatar: clang-based CUDA compilation
# SM70 = BI-V100 architecture
# From real machine GDN compile log (dockerrizhi.txt):
# /usr/local/corex/bin/clang++ ... --cuda-gpu-arch=ivcore10
# --cuda-path=/usr/local/corex -std=c++17
# -D__ILUVATAR__ -D__ILUVATAR_WORKAROUND__
local OBJ="${BUILD_DIR}/$(basename ${cu_file} .cu).cuda.o"
"${COREX_ROOT}/bin/clang++" \
-x cuda \
--cuda-gpu-arch=sm_70 \
-std=c++17 \
-D__ILUVATAR__ \
-D__ILUVATAR_WORKAROUND__ \
-D__ILUVATAR_DIAG__ \
-fPIC \
-O2 \
-shared -fPIC \
--cuda-gpu-arch=ivcore10 \
--cuda-path="${COREX_ROOT}" \
-std=c++17 \
-I"${INCLUDE_DIR}" \
-I"${COREX_ROOT}/include" \
-isystem "${COREX_ROOT}/include" \
-c "${cu_file}" \
-o "${OBJ}"
# Link .o → .so (match real machine: c++ ... -shared -L ... -lcudart)
c++ "${OBJ}" -shared \
-L"${COREX_ROOT}/lib64" \
-lcudart \
-o "${so_path}" \
"${cu_file}"
-o "${so_path}"
rm -f "${OBJ}"
else
# Standard nvcc
nvcc \