-
36909bf964
[fix] prebuilt
root
2026-08-17 10:11:24 +00:00
-
f10cccf9df
[fix] baseline4 :ix_moe_bridge.so 。 v2 bridge 真机编译通过,13函数全导出
root
2026-08-17 09:18:32 +00:00
-
-
7d6884d15a
tool: verify_bridge.sh — 验证prebuilt .so是v1还是v2,缺MoE则自动重编
Claude
2026-08-17 08:49:14 +00:00
-
b81a071f84
merge modelhub: bridge_v2 + patch_vllm_ops fix
root
2026-08-17 08:34:14 +00:00
-
-
b342eb6b98
[fix] baseline4 fix prebuilt ix_full_bridge.so 大概率是从 v2 编译的
root
2026-08-17 08:26:54 +00:00
-
c044d04509
merge modelhub: dockerignore fix + build_moe_bridge
root
2026-08-17 07:54:23 +00:00
-
-
ba127fe66b
delay baseline4 docker build
root
2026-08-17 07:39:52 +00:00
-
8211a45464
platform test baseline4
root
2026-08-17 07:19:44 +00:00
-
f7070b2075
platform test baseline4
root
2026-08-17 07:19:44 +00:00
-
be4d661191
revert: undo 2 premature pushes (
ee516bd2, 9ea0a1d4) — code needs review first
dev
2026-08-17 07:04:50 +00:00
-
9ea0a1d4f4
fix: resolve all 8 deployment pipeline breaks
dev
2026-08-17 07:01:41 +00:00
-
ee516bd206
fix: connect ex_engine to Docker build pipeline
dev
2026-08-17 06:56:11 +00:00
-
461addf428
Revert "feat: 10-file algorithm factor system — full 10-layer AST call chain"
Claude
2026-08-17 05:32:19 +00:00
-
79562c342d
Revert "feat: build_all.sh — compile 3 .so from 10 algorithm factor files"
project6-dev
2026-08-17 05:31:05 +00:00
-
3581dd5435
feat: 10-file algorithm factor system — full 10-layer AST call chain
Claude
2026-08-17 05:30:22 +00:00
-
401d33ca6b
feat: build_all.sh — compile 3 .so from 10 algorithm factor files
project6-dev
2026-08-17 05:30:21 +00:00
-
-
512f384a49
Merge branch 'main' of https://dev.modelhub.org.cn/dylanyunlong/project_6 into main
root
2026-08-17 04:19:31 +00:00
-
-
cdcf115037
feat: CUTLASS Cu10 grouped GEMM — real device verified
root
2026-08-17 04:18:29 +00:00
-
cb926707af
Merge branch 'main' of https://github.com/dylanyunlon/project_6
root
2026-08-17 02:21:09 +00:00
-
-
beaa8dbb65
Merge branch 'main' of https://dev.modelhub.org.cn/dylanyunlong/project_6
root
2026-08-17 02:18:46 +00:00
-
-
-
-
03be5f2b15
[feat] group gemm
root
2026-08-17 02:16:58 +00:00
-
330669b309
Revert "fix: ix_full_bridge_v2.cpp — align namespace+signatures to real nm -D symbol dump"
Claude
2026-08-17 02:08:03 +00:00
-
5c03156978
fix: ix_full_bridge_v2.cpp — align namespace+signatures to real nm -D symbol dump
Claude
2026-08-17 02:04:47 +00:00
-
34a8fbf27e
revert: undo 2 premature pushes (
c54923a1, 49034d1d) — code needs review first
project_6
2026-08-16 17:48:17 +00:00
-
49034d1d09
feat: 10-file MoE bridge pipeline — compile, dispatch, patch, test
project_6
2026-08-16 17:46:53 +00:00
-
c54923a17e
feat: implement 5 missing MoE ops — topk_softmax + token_index + expand + group_gemm + combine
project_6
2026-08-16 17:35:08 +00:00
-
-
3712c06861
data: full symbol dumps
root
2026-08-16 17:15:11 +00:00
-
dec268d252
data: full symbol dumps
root
2026-08-16 17:15:11 +00:00
-
415ca12afc
fix: group_gemm format "TN" + Layer 3 ops_api dispatch from xllm upstream
project_6
2026-08-16 16:09:15 +00:00
-
5172f94b1f
Revert "feat: 3-tier ixformer flash prefill dispatch + OpenCompass max_tokens clamp + n>1 fanout + index sanitizer"
Claude
2026-08-16 15:42:36 +00:00
-
7a7ddf38db
Revert "test: deploy_and_verify.sh — pull+patch+probe ixformer backends+clamp test"
Claude
2026-08-16 15:42:36 +00:00
-
587e18309b
test: deploy_and_verify.sh — pull+patch+probe ixformer backends+clamp test
Claude
2026-08-16 15:40:47 +00:00
-
-
cdec569977
feat: 3-tier ixformer flash prefill dispatch + OpenCompass max_tokens clamp + n>1 fanout + index sanitizer
Claude
2026-08-16 15:37:23 +00:00
-
522e8376b6
data: sub 694 (683分) 日志 + resolve merge
root
2026-08-16 15:28:08 +00:00
-
-
4189f44d27
test: dump ALL symbols from ALL ixformer/cuinfer .so — no grep filter, find what we missed
Claude
2026-08-16 04:58:39 +00:00
-
8eabbac857
test: probe_model_shapes.sh — 真机验证模型config和MoE权重shape
Claude
2026-08-15 15:05:26 +00:00
-
b7149f810a
fix: decode MoE路径对齐base — F.linear+bmm替换pre-transpose+bmm
Claude
2026-08-15 14:55:01 +00:00
-
ea2c15f699
fix: decode MoE路径对齐base — F.linear+bmm替换pre-transpose+bmm
Claude
2026-08-15 14:55:01 +00:00
-
b187f52ced
data: so import chain probe
root
2026-08-15 14:53:57 +00:00
-
ee62ea13ba
data: so import chain probe
root
2026-08-15 14:49:42 +00:00
-
924e48b502
data: so import chain probe
root
2026-08-15 14:49:42 +00:00
-
1700c35bd7
test: probe_so_import_chain.sh — 验证.so部署路径+import链+flag值+shape匹配
Claude
2026-08-15 14:46:49 +00:00
-
c290278b35
test: probe_so_import_chain.sh — 验证.so部署路径+import链+flag值+shape匹配
Claude
2026-08-15 14:46:49 +00:00
-
a6cc233880
data: base MoE forward + corex_moe签名
root
2026-08-15 14:39:03 +00:00
-
784dea96c0
data: base MoE forward + corex_moe签名
root
2026-08-15 14:39:03 +00:00
-
adf05b6bfb
test: probe_base_moe_forward.sh — cat base qwen3_5.py的完整MoE forward + 所有corex_moe_*.so签名
Claude
2026-08-15 14:36:29 +00:00
-
7a8545f7c2
test: push_probe_results.sh — 真机commit probe结果到modelhub
Claude
2026-08-15 14:34:49 +00:00
-
b00429d81c
test: probe_base_moe_forward.sh — cat base qwen3_5.py的完整MoE forward + 所有corex_moe_*.so签名
Claude
2026-08-15 14:36:29 +00:00
-
6a4459d405
data: probe bridge output
root
2026-08-15 14:35:07 +00:00
-
437ad8aaa4
data: probe bridge output
root
2026-08-15 14:35:07 +00:00
-
6e22415a91
test: push_probe_results.sh — 真机commit probe结果到modelhub
Claude
2026-08-15 14:34:49 +00:00
-
-
de1212c271
test: probe_ix_unified_bridge.sh — cat base镜像的ix_unified_bridge + corex_*.so + _custom_ops.py完整接口
Claude
2026-08-15 14:30:53 +00:00
-
a823bdf9ea
test: probe_real_machine.sh — cat ixformer/vllm/cublas真机数据
Claude
2026-08-15 14:28:10 +00:00
-
6415249693
data: port complete MoE + xllm layer call chains from upstream repos
Claude
2026-08-15 14:26:18 +00:00
-
7aa5054574
feat: ILU kernel pipeline — ix_full_bridge_v2 build + deploy + 7-step MoE dispatch
dylan
2026-08-15 14:15:41 +00:00
-
52e2ef31a8
feat: xllm_ops NO-FALLBACK kernel loader + 6 missing .so build targets + hot-path patcher
Claude
2026-08-15 14:13:13 +00:00
-
e873e5f27b
fix: eliminate 8x CUDA sync in MoE decode — tolist() once instead of .item() per expert
dylan
2026-08-15 13:14:11 +00:00
-
23fe535985
fix: add ex_engine/__init__.py for Python package import
dylan
2026-08-15 13:08:24 +00:00
-
e18ece8f3a
feat: port NaiveBatchedExperts from ds_vllm — view transpose + cublas transB
dylan
2026-08-15 13:05:45 +00:00
-
6f1904aa8c
perf: MoE decode — pre-transposed bmm replaces F.linear (6.9ms vs 8.0ms, 14%)
dylan
2026-08-15 12:52:55 +00:00
-
9f265894cc
test: MoE breakdown — F.linear vs torch.mm vs torch.bmm vs bmm pre-transposed
Claude
2026-08-15 12:46:37 +00:00
-
e47b66e268
fix: module name in probe_moe_fused_breakdown.sh
Claude
2026-08-15 12:43:55 +00:00
-
b0af7d54ff
test: breakdown moe_decode_fused timing by step — find the real bottleneck
Claude
2026-08-15 12:41:29 +00:00
-
0795e064b2
fix: corex_batched_gemm TCU OpClassTensorOp + Cu10 + float accum (merge)
dylan
2026-08-15 12:36:18 +00:00
-
-
3b2a0bc4d3
fix: corex_batched_gemm TCU OpClassTensorOp + Cu10 + float accum
dylan
2026-08-15 12:36:06 +00:00
-
e2fc3f270f
fix: corex_batched_gemm use TCU OpClassTensorOp + Cu10 + float accum
dylan
2026-08-15 12:35:56 +00:00
-
d8d241bf9f
fix: corex_batched_gemm_kernel — use OpClassTensorOp + arch::Cu10 + FP32 accumulator
Claude
2026-08-15 12:34:01 +00:00
-
-
3481f2903f
fix(build): cutlass.h lives under tensorflow/include on this image
dylan
2026-08-15 12:03:12 +00:00
-
f41900c06b
fix(build): auto-find cutlass/cutlass.h under COREX_ROOT
dylan
2026-08-15 12:02:23 +00:00
-
bfa18cd5b4
fix(build): use CoreX clang++ instead of nvcc — match working build scripts
dylan
2026-08-15 12:01:17 +00:00
-
04cc9b88af
fix(build): add CUDA include path to g++ step in build_corex_batched_gemm.sh
dylan
2026-08-15 11:58:53 +00:00
-
ddcfbad431
feat: pybind wrapper for CUTLASS batched GEMM → MoE decode path
dylan
2026-08-15 11:54:26 +00:00
-
a875fa5d4c
Revert "data: cat SGEMM files from 3 repos into cat_files/"
dylan
2026-08-15 11:48:10 +00:00
-
36676f2d1b
data: complete SGEMM upstream from 3 repos (siboehm+wangzyon+edtallison) + xllm fused_qknorm_rope + xattention kernels
Claude
2026-08-15 07:00:04 +00:00
-
7cfa87b5ac
data: cat SGEMM files from 3 repos into cat_files/
dylan
2026-08-15 06:59:18 +00:00
-
284804ac53
data: cat 3 SGEMM repos — siboehm, wangzyon, edtallison (full clone, no --depth)
dylan
2026-08-15 06:58:07 +00:00
-
854fb93a8e
test: add test_ex_engine_cuda.py — test all 22 prebuilt .so on BI-V100
dylan
2026-08-15 06:28:34 +00:00
-
e8f0948fe1
feat: ix_ops integration layer — wire ix_full_bridge.so into vllm hot path
dylan
2026-08-15 06:15:17 +00:00
-
109d29fa60
Merge remote-tracking branch 'modelhub/main'
root
2026-08-15 05:58:37 +00:00
-
-
045ea5df79
feat: Cu10 TensorOp batched HGEMM via Iluvatar CUTLASS framework
Claude
2026-08-15 05:44:15 +00:00
-
30f98c0674
data: Cu10 CUTLASS part 2 — tensorop example, arch.h, cutlass.h
root
2026-08-15 05:41:30 +00:00
-
-
f006ab1a01
test: cat tensorop GEMM example + arch.h + cutlass.h from corex-samples
Claude
2026-08-15 05:41:03 +00:00
-
b922d694dc
data: Cu10 CUTLASS headers from corex-samples
root
2026-08-15 05:32:10 +00:00
-
1d36754efc
fix: cat_cutlass_cu10.sh writes to cat_files/ directory instead of stdout
Claude
2026-08-15 05:31:33 +00:00
-
4abb4df215
test: cat Cu10 CUTLASS files — mma_cu10.h, iluvatar_mma.hpp, batched_gemm.cu, default_mma_core_cu10.h
Claude
2026-08-15 05:30:15 +00:00
-
6b9086c3a9
test: probe Cu10 CUTLASS fork — find mma_cu10.h, tensor op files, batched_gemm example
Claude
2026-08-15 05:27:43 +00:00
-
a465dd1d75
fix: use F.silu in test script for old corex torch
Claude
2026-08-15 05:22:31 +00:00
-
a12d070d82
fix: replace torch::silu with x*sigmoid(x) for old corex torch
Claude
2026-08-15 05:22:25 +00:00
-
c840c9159f
feat: moe_tcu_dispatch.cpp — C++ MoE expert loop via torch::mm (TCU kernel)
Claude
2026-08-15 05:20:05 +00:00
-
9514092980
test: probe torch.matmul backend + ixformer.matmul/linear + Python loop overhead
Claude
2026-08-15 05:17:13 +00:00
-
395b3e4042
test: clean rebuild + debug output for kernel 10 correctness
Claude
2026-08-15 05:14:18 +00:00
-
a8ca42b59c
perf: hgemm_warptiling Config B — beats cublas on MoE-sized GEMM (0.7x)
Claude
2026-08-15 05:11:32 +00:00
-
21417319bc
test: sweep 6 kernel 10 configs + cublas baseline — find best params for warp64
Claude
2026-08-15 05:08:11 +00:00
-
27bb8d28df
test: probe kernel 10 perf with CUDA events — isolate bottleneck
Claude
2026-08-14 17:23:41 +00:00
-
2b12fe687e
feat: hgemm_warptiling.cu — siboehm kernel 10 ported to WARPSIZE=64 FP16
Claude
2026-08-14 17:05:59 +00:00
-
11b8a98eea
test: probe warp_size=64 behavior + kernel 10 warp tiling with WARPSIZE=64 on BI-V100
Claude
2026-08-14 17:02:55 +00:00
-
1af7e7cf48
fix: use c10::cuda::getCurrentCUDAStream().stream() for corex torch
Claude
2026-08-14 16:49:47 +00:00
-
3bee73207e
fix: add cuda_runtime.h to hgemm_bind.cpp for cudaStream_t
Claude
2026-08-14 16:33:29 +00:00
-
09e5261ba6
refactor: hgemm_blocktiling.cu — strict 1:1 from siboehm kernel 6
Claude
2026-08-14 16:24:22 +00:00
-
ab42fc1fd7
feat: hgemm_blocktiling.cu — FP16 GEMM kernel for MoE expert dispatch on BI-V100
Claude
2026-08-14 16:22:00 +00:00