[BufFix]Fix the error when using Ascend custom operators with rank=128 (#5394)

### What this PR does / why we need it? The customized ascend operator sgmv_expand and sgmv_shrink applies only to the scenario where rank is 8,16,32,64. When rank >= 128, the operator is out of range, causing the model to report an error. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? Depends on this commit https://github.com/vllm-project/vllm/pull/31408 - vLLM version: release/v0.13.0 - vLLM main: 254f6b9867 --------- Signed-off-by: ZT-AIA <1028681969@qq.com> Signed-off-by: ZT-AIA <63220130+ZT-AIA@users.noreply.github.com>
2026-01-09 15:57:43 +08:00
parent d36ca88cf4
commit e11ff8e535
4 changed files with 35 additions and 23 deletions
--- a/.github/workflows/_e2e_test.yaml
+++ b/.github/workflows/_e2e_test.yaml
@@ -105,7 +105,7 @@ jobs:
          # xgrammar has parameter mismatching bug, please follows: https://github.com/vllm-project/vllm-ascend/issues/5524
          # pytest -sv --durations=0 tests/e2e/singlecard/test_guided_decoding.py
          # torch 2.8 doesn't work with lora, fix me
-          #pytest -sv --durations=0 tests/e2e/singlecard/test_ilama_lora.py
+          pytest -sv --durations=0 tests/e2e/singlecard/test_ilama_lora.py
          pytest -sv --durations=0 tests/e2e/singlecard/test_models.py
          pytest -sv --durations=0 tests/e2e/singlecard/test_multistream_overlap_shared_expert.py
          pytest -sv --durations=0 tests/e2e/singlecard/test_profile_execute_duration.py
@@ -216,7 +216,7 @@ jobs:
          pytest -sv --durations=0 tests/e2e/multicard/2-cards/test_external_launcher.py
          pytest -sv --durations=0 tests/e2e/multicard/2-cards/test_full_graph_mode.py
          # torch 2.8 doesn't work with lora, fix me
-          #pytest -sv --durations=0 tests/e2e/multicard/2-cards/test_ilama_lora_tp2.py
+          pytest -sv --durations=0 tests/e2e/multicard/2-cards/test_ilama_lora_tp2.py
          

          # To avoid oom, we need to run the test in a single process.