Main2main upgrade to vllm 0317 afternoon (#7409)

### What this PR does / why we need it? 1.fix "TypeError: get_attn_backend() remove variable": [Refactor `check_and_update_config`](https://github.com/vllm-project/vllm/pull/35122) 2.fix [Rename `compile_ranges_split_points` to `compile_ranges_endpoints`](https://github.com/vllm-project/vllm/pull/36027) 3.fix "RuntimeError: device_allocator not a DeviceAllocator":[Replace memory related torch.cuda APIs"](https://github.com/vllm-project/vllm/pull/37031) 4.fix [Support multiple KV groups in OffloadingSpec ](https://github.com/vllm-project/vllm/pull/36610) removed self.offloaded_block_size and changed self.gpu_block_size from a scalar to a tuple of per-group block sizes, adding block_size_factor. 5.fix [Consolidate SupportsEagle](https://github.com/vllm-project/vllm/pull/36063) renamed get_eagle3_aux_hidden_state_layers() to get_eagle3_default_aux_hidden_state_layers() and added a supports_eagle3() guard before calling it. ### Does this PR introduce _any_ user-facing change? NA ### How was this patch tested? E2E - vLLM version: v0.17.0 - vLLM main: 8a680463fa --------- Signed-off-by: leo-pony <nengjunma@outlook.com> Co-authored-by: Claude Code <noreply@anthropic.com>
2026-03-18 23:24:27 +08:00
parent 305820f1a9
commit 8b79d4de52
13 changed files with 125 additions and 41 deletions
--- a/.github/workflows/pr_test_light.yaml
+++ b/.github/workflows/pr_test_light.yaml
@@ -41,7 +41,7 @@ jobs:
  lint:
    uses: ./.github/workflows/_pre_commit.yml
    with:
-      vllm: 4497431df654e46fb1fb5e64bf8611e762ae5d87
+      vllm: 8a680463fab3bc9e6760417cd5c0a6aa58283065
  changes:
    runs-on: linux-aarch64-a2b3-0
    outputs:
@@ -90,7 +90,7 @@ jobs:
    if: ${{ needs.lint.result == 'success' && (needs.changes.outputs.e2e_tracker == 'true' || needs.changes.outputs.ut_tracker == 'true') }}
    strategy:
      matrix:
-        vllm_version: [4497431df654e46fb1fb5e64bf8611e762ae5d87, v0.17.0]
+        vllm_version: [8a680463fab3bc9e6760417cd5c0a6aa58283065, v0.17.0]
    uses: ./.github/workflows/_unit_test.yaml
    with:
      vllm: ${{ matrix.vllm_version }}
@@ -102,7 +102,7 @@ jobs:
    name: e2e-light
    strategy:
      matrix:
-        vllm_version: [4497431df654e46fb1fb5e64bf8611e762ae5d87, v0.17.0]
+        vllm_version: [8a680463fab3bc9e6760417cd5c0a6aa58283065, v0.17.0]
    # Note (yikun): If CI resource are limited we can split job into two chain jobs
    needs: [lint, changes]
    # only trigger e2e test after lint passed and the change is e2e related with pull request.