xc-llm-ascend

Files

NeverRaR 71de52d3a9 feat: add kv cache memory cache and skip dynamo guard (#1549 )

### What this PR does / why we need it?

1、Sometimes loading torchair cache will fail because of the floating of
npu memory, so this pr add a new cache to save the old kv cache bytes to
avoid the possible crash while loading the torchair graph cache.
2、When caching is enabled and does not exist, the first compilation
introduces the overhead of Dynamo Gurad. So in this case, we will
compile them directly twice to skip them (This will bring 3-4 ms of tpot
optimization)

### Does this PR introduce _any_ user-facing change?
Add a new env `VLLM_ASCEND_KV_CACHE_MEGABYTES_FLOATING_TOLERANCE` to
control kv cache floating tolerance

### How was this patch tested?

- vLLM version: v0.9.1
- vLLM main:
1fd471e957

Signed-off-by: boying <897013703@qq.com>

2025-07-07 22:37:14 +08:00

e2e

[CI] Fix oom in chunk prefill (#1622 )

2025-07-07 10:14:40 +08:00

feat: add kv cache memory cache and skip dynamo guard (#1549 )

2025-07-07 22:37:14 +08:00

__init__.py

[SpecDecode] Add spec decode support (#500 )

2025-04-17 20:16:32 +08:00

conftest.py

[CI/UT] Unify model usage via ModelScope in CI (#1207 )

2025-07-04 10:52:17 +08:00

model_utils.py

[CI] Refactor CI (#952 )

2025-05-28 06:31:35 +08:00

utils.py

[V1][ModelRunner] Support pooling model for v1 engine (#1359 )

2025-06-30 16:31:12 +08:00