Commit Graph

  • 36909bf964 [fix] prebuilt root 2026-08-17 10:11:24 +00:00
  • f10cccf9df [fix] baseline4 :ix_moe_bridge.so 。 v2 bridge 真机编译通过,13函数全导出 root 2026-08-17 09:18:32 +00:00
  • 7d6884d15a tool: verify_bridge.sh — 验证prebuilt .so是v1还是v2,缺MoE则自动重编 Claude 2026-08-17 08:49:14 +00:00
  • b81a071f84 merge modelhub: bridge_v2 + patch_vllm_ops fix root 2026-08-17 08:34:14 +00:00
  • b342eb6b98 [fix] baseline4 fix prebuilt ix_full_bridge.so 大概率是从 v2 编译的 root 2026-08-17 08:26:54 +00:00
  • c044d04509 merge modelhub: dockerignore fix + build_moe_bridge root 2026-08-17 07:54:23 +00:00
  • ba127fe66b delay baseline4 docker build root 2026-08-17 07:39:52 +00:00
  • 8211a45464 platform test baseline4 root 2026-08-17 07:19:44 +00:00
  • f7070b2075 platform test baseline4 root 2026-08-17 07:19:44 +00:00
  • be4d661191 revert: undo 2 premature pushes (ee516bd2, 9ea0a1d4) — code needs review first dev 2026-08-17 07:04:50 +00:00
  • 9ea0a1d4f4 fix: resolve all 8 deployment pipeline breaks dev 2026-08-17 07:01:41 +00:00
  • ee516bd206 fix: connect ex_engine to Docker build pipeline dev 2026-08-17 06:56:11 +00:00
  • 461addf428 Revert "feat: 10-file algorithm factor system — full 10-layer AST call chain" Claude 2026-08-17 05:32:19 +00:00
  • 79562c342d Revert "feat: build_all.sh — compile 3 .so from 10 algorithm factor files" project6-dev 2026-08-17 05:31:05 +00:00
  • 3581dd5435 feat: 10-file algorithm factor system — full 10-layer AST call chain Claude 2026-08-17 05:30:22 +00:00
  • 401d33ca6b feat: build_all.sh — compile 3 .so from 10 algorithm factor files project6-dev 2026-08-17 05:30:21 +00:00
  • 512f384a49 Merge branch 'main' of https://dev.modelhub.org.cn/dylanyunlong/project_6 into main root 2026-08-17 04:19:31 +00:00
  • cdcf115037 feat: CUTLASS Cu10 grouped GEMM — real device verified root 2026-08-17 04:18:29 +00:00
  • cb926707af Merge branch 'main' of https://github.com/dylanyunlon/project_6 root 2026-08-17 02:21:09 +00:00
  • beaa8dbb65 Merge branch 'main' of https://dev.modelhub.org.cn/dylanyunlong/project_6 root 2026-08-17 02:18:46 +00:00
  • 03be5f2b15 [feat] group gemm root 2026-08-17 02:16:58 +00:00
  • 330669b309 Revert "fix: ix_full_bridge_v2.cpp — align namespace+signatures to real nm -D symbol dump" Claude 2026-08-17 02:08:03 +00:00
  • 5c03156978 fix: ix_full_bridge_v2.cpp — align namespace+signatures to real nm -D symbol dump Claude 2026-08-17 02:04:47 +00:00
  • 34a8fbf27e revert: undo 2 premature pushes (c54923a1, 49034d1d) — code needs review first project_6 2026-08-16 17:48:17 +00:00
  • 49034d1d09 feat: 10-file MoE bridge pipeline — compile, dispatch, patch, test project_6 2026-08-16 17:46:53 +00:00
  • c54923a17e feat: implement 5 missing MoE ops — topk_softmax + token_index + expand + group_gemm + combine project_6 2026-08-16 17:35:08 +00:00
  • 3712c06861 data: full symbol dumps root 2026-08-16 17:15:11 +00:00
  • dec268d252 data: full symbol dumps root 2026-08-16 17:15:11 +00:00
  • 415ca12afc fix: group_gemm format "TN" + Layer 3 ops_api dispatch from xllm upstream project_6 2026-08-16 16:09:15 +00:00
  • 5172f94b1f Revert "feat: 3-tier ixformer flash prefill dispatch + OpenCompass max_tokens clamp + n>1 fanout + index sanitizer" Claude 2026-08-16 15:42:36 +00:00
  • 7a7ddf38db Revert "test: deploy_and_verify.sh — pull+patch+probe ixformer backends+clamp test" Claude 2026-08-16 15:42:36 +00:00
  • 587e18309b test: deploy_and_verify.sh — pull+patch+probe ixformer backends+clamp test Claude 2026-08-16 15:40:47 +00:00
  • cdec569977 feat: 3-tier ixformer flash prefill dispatch + OpenCompass max_tokens clamp + n>1 fanout + index sanitizer Claude 2026-08-16 15:37:23 +00:00
  • 522e8376b6 data: sub 694 (683分) 日志 + resolve merge root 2026-08-16 15:28:08 +00:00
  • 4189f44d27 test: dump ALL symbols from ALL ixformer/cuinfer .so — no grep filter, find what we missed Claude 2026-08-16 04:58:39 +00:00
  • 8eabbac857 test: probe_model_shapes.sh — 真机验证模型config和MoE权重shape Claude 2026-08-15 15:05:26 +00:00
  • b7149f810a fix: decode MoE路径对齐base — F.linear+bmm替换pre-transpose+bmm Claude 2026-08-15 14:55:01 +00:00
  • ea2c15f699 fix: decode MoE路径对齐base — F.linear+bmm替换pre-transpose+bmm Claude 2026-08-15 14:55:01 +00:00
  • b187f52ced data: so import chain probe root 2026-08-15 14:53:57 +00:00
  • ee62ea13ba data: so import chain probe root 2026-08-15 14:49:42 +00:00
  • 924e48b502 data: so import chain probe root 2026-08-15 14:49:42 +00:00
  • 1700c35bd7 test: probe_so_import_chain.sh — 验证.so部署路径+import链+flag值+shape匹配 Claude 2026-08-15 14:46:49 +00:00
  • c290278b35 test: probe_so_import_chain.sh — 验证.so部署路径+import链+flag值+shape匹配 Claude 2026-08-15 14:46:49 +00:00
  • a6cc233880 data: base MoE forward + corex_moe签名 root 2026-08-15 14:39:03 +00:00
  • 784dea96c0 data: base MoE forward + corex_moe签名 root 2026-08-15 14:39:03 +00:00
  • adf05b6bfb test: probe_base_moe_forward.sh — cat base qwen3_5.py的完整MoE forward + 所有corex_moe_*.so签名 Claude 2026-08-15 14:36:29 +00:00
  • 7a8545f7c2 test: push_probe_results.sh — 真机commit probe结果到modelhub Claude 2026-08-15 14:34:49 +00:00
  • b00429d81c test: probe_base_moe_forward.sh — cat base qwen3_5.py的完整MoE forward + 所有corex_moe_*.so签名 Claude 2026-08-15 14:36:29 +00:00
  • 6a4459d405 data: probe bridge output root 2026-08-15 14:35:07 +00:00
  • 437ad8aaa4 data: probe bridge output root 2026-08-15 14:35:07 +00:00
  • 6e22415a91 test: push_probe_results.sh — 真机commit probe结果到modelhub Claude 2026-08-15 14:34:49 +00:00
  • de1212c271 test: probe_ix_unified_bridge.sh — cat base镜像的ix_unified_bridge + corex_*.so + _custom_ops.py完整接口 Claude 2026-08-15 14:30:53 +00:00
  • a823bdf9ea test: probe_real_machine.sh — cat ixformer/vllm/cublas真机数据 Claude 2026-08-15 14:28:10 +00:00
  • 6415249693 data: port complete MoE + xllm layer call chains from upstream repos Claude 2026-08-15 14:26:18 +00:00
  • 7aa5054574 feat: ILU kernel pipeline — ix_full_bridge_v2 build + deploy + 7-step MoE dispatch dylan 2026-08-15 14:15:41 +00:00
  • 52e2ef31a8 feat: xllm_ops NO-FALLBACK kernel loader + 6 missing .so build targets + hot-path patcher Claude 2026-08-15 14:13:13 +00:00
  • e873e5f27b fix: eliminate 8x CUDA sync in MoE decode — tolist() once instead of .item() per expert dylan 2026-08-15 13:14:11 +00:00
  • 23fe535985 fix: add ex_engine/__init__.py for Python package import dylan 2026-08-15 13:08:24 +00:00
  • e18ece8f3a feat: port NaiveBatchedExperts from ds_vllm — view transpose + cublas transB dylan 2026-08-15 13:05:45 +00:00
  • 6f1904aa8c perf: MoE decode — pre-transposed bmm replaces F.linear (6.9ms vs 8.0ms, 14%) dylan 2026-08-15 12:52:55 +00:00
  • 9f265894cc test: MoE breakdown — F.linear vs torch.mm vs torch.bmm vs bmm pre-transposed Claude 2026-08-15 12:46:37 +00:00
  • e47b66e268 fix: module name in probe_moe_fused_breakdown.sh Claude 2026-08-15 12:43:55 +00:00
  • b0af7d54ff test: breakdown moe_decode_fused timing by step — find the real bottleneck Claude 2026-08-15 12:41:29 +00:00
  • 0795e064b2 fix: corex_batched_gemm TCU OpClassTensorOp + Cu10 + float accum (merge) dylan 2026-08-15 12:36:18 +00:00
  • 3b2a0bc4d3 fix: corex_batched_gemm TCU OpClassTensorOp + Cu10 + float accum dylan 2026-08-15 12:36:06 +00:00
  • e2fc3f270f fix: corex_batched_gemm use TCU OpClassTensorOp + Cu10 + float accum dylan 2026-08-15 12:35:56 +00:00
  • d8d241bf9f fix: corex_batched_gemm_kernel — use OpClassTensorOp + arch::Cu10 + FP32 accumulator Claude 2026-08-15 12:34:01 +00:00
  • 3481f2903f fix(build): cutlass.h lives under tensorflow/include on this image dylan 2026-08-15 12:03:12 +00:00
  • f41900c06b fix(build): auto-find cutlass/cutlass.h under COREX_ROOT dylan 2026-08-15 12:02:23 +00:00
  • bfa18cd5b4 fix(build): use CoreX clang++ instead of nvcc — match working build scripts dylan 2026-08-15 12:01:17 +00:00
  • 04cc9b88af fix(build): add CUDA include path to g++ step in build_corex_batched_gemm.sh dylan 2026-08-15 11:58:53 +00:00
  • ddcfbad431 feat: pybind wrapper for CUTLASS batched GEMM → MoE decode path dylan 2026-08-15 11:54:26 +00:00
  • a875fa5d4c Revert "data: cat SGEMM files from 3 repos into cat_files/" dylan 2026-08-15 11:48:10 +00:00
  • 36676f2d1b data: complete SGEMM upstream from 3 repos (siboehm+wangzyon+edtallison) + xllm fused_qknorm_rope + xattention kernels Claude 2026-08-15 07:00:04 +00:00
  • 7cfa87b5ac data: cat SGEMM files from 3 repos into cat_files/ dylan 2026-08-15 06:59:18 +00:00
  • 284804ac53 data: cat 3 SGEMM repos — siboehm, wangzyon, edtallison (full clone, no --depth) dylan 2026-08-15 06:58:07 +00:00
  • 854fb93a8e test: add test_ex_engine_cuda.py — test all 22 prebuilt .so on BI-V100 dylan 2026-08-15 06:28:34 +00:00
  • e8f0948fe1 feat: ix_ops integration layer — wire ix_full_bridge.so into vllm hot path dylan 2026-08-15 06:15:17 +00:00
  • 109d29fa60 Merge remote-tracking branch 'modelhub/main' root 2026-08-15 05:58:37 +00:00
  • 045ea5df79 feat: Cu10 TensorOp batched HGEMM via Iluvatar CUTLASS framework Claude 2026-08-15 05:44:15 +00:00
  • 30f98c0674 data: Cu10 CUTLASS part 2 — tensorop example, arch.h, cutlass.h root 2026-08-15 05:41:30 +00:00
  • f006ab1a01 test: cat tensorop GEMM example + arch.h + cutlass.h from corex-samples Claude 2026-08-15 05:41:03 +00:00
  • b922d694dc data: Cu10 CUTLASS headers from corex-samples root 2026-08-15 05:32:10 +00:00
  • 1d36754efc fix: cat_cutlass_cu10.sh writes to cat_files/ directory instead of stdout Claude 2026-08-15 05:31:33 +00:00
  • 4abb4df215 test: cat Cu10 CUTLASS files — mma_cu10.h, iluvatar_mma.hpp, batched_gemm.cu, default_mma_core_cu10.h Claude 2026-08-15 05:30:15 +00:00
  • 6b9086c3a9 test: probe Cu10 CUTLASS fork — find mma_cu10.h, tensor op files, batched_gemm example Claude 2026-08-15 05:27:43 +00:00
  • a465dd1d75 fix: use F.silu in test script for old corex torch Claude 2026-08-15 05:22:31 +00:00
  • a12d070d82 fix: replace torch::silu with x*sigmoid(x) for old corex torch Claude 2026-08-15 05:22:25 +00:00
  • c840c9159f feat: moe_tcu_dispatch.cpp — C++ MoE expert loop via torch::mm (TCU kernel) Claude 2026-08-15 05:20:05 +00:00
  • 9514092980 test: probe torch.matmul backend + ixformer.matmul/linear + Python loop overhead Claude 2026-08-15 05:17:13 +00:00
  • 395b3e4042 test: clean rebuild + debug output for kernel 10 correctness Claude 2026-08-15 05:14:18 +00:00
  • a8ca42b59c perf: hgemm_warptiling Config B — beats cublas on MoE-sized GEMM (0.7x) Claude 2026-08-15 05:11:32 +00:00
  • 21417319bc test: sweep 6 kernel 10 configs + cublas baseline — find best params for warp64 Claude 2026-08-15 05:08:11 +00:00
  • 27bb8d28df test: probe kernel 10 perf with CUDA events — isolate bottleneck Claude 2026-08-14 17:23:41 +00:00
  • 2b12fe687e feat: hgemm_warptiling.cu — siboehm kernel 10 ported to WARPSIZE=64 FP16 Claude 2026-08-14 17:05:59 +00:00
  • 11b8a98eea test: probe warp_size=64 behavior + kernel 10 warp tiling with WARPSIZE=64 on BI-V100 Claude 2026-08-14 17:02:55 +00:00
  • 1af7e7cf48 fix: use c10::cuda::getCurrentCUDAStream().stream() for corex torch Claude 2026-08-14 16:49:47 +00:00
  • 3bee73207e fix: add cuda_runtime.h to hgemm_bind.cpp for cudaStream_t Claude 2026-08-14 16:33:29 +00:00
  • 09e5261ba6 refactor: hgemm_blocktiling.cu — strict 1:1 from siboehm kernel 6 Claude 2026-08-14 16:24:22 +00:00
  • ab42fc1fd7 feat: hgemm_blocktiling.cu — FP16 GEMM kernel for MoE expert dispatch on BI-V100 Claude 2026-08-14 16:22:00 +00:00