This website requires JavaScript.
Explore
Help
Register
Sign In
dylanyunlong
/
project_6
Watch
1
Star
0
Fork
0
You've already forked project_6
Code
Issues
Pull Requests
Actions
Projects
Releases
Wiki
Activity
Files
52e2ef31a86aaedc58211a15b85c9936de8eb530
project_6
/
ex_engine
/
moe
History
dylan
e873e5f27b
fix: eliminate 8x CUDA sync in MoE decode — tolist() once instead of .item() per expert
2026-08-15 13:14:11 +00:00
..
__init__.py
feat: port NaiveBatchedExperts from ds_vllm — view transpose + cublas transB
2026-08-15 13:05:45 +00:00
activation.py
feat: port NaiveBatchedExperts from ds_vllm — view transpose + cublas transB
2026-08-15 13:05:45 +00:00
naive_batched_experts.py
fix: eliminate 8x CUDA sync in MoE decode — tolist() once instead of .item() per expert
2026-08-15 13:14:11 +00:00