zhangxinyuehfad
a5b099c73d
[Bugfix] Reset incompatible config (#6005)
### What this PR does / why we need it?
This PR introduces compatibility fixes for running vLLM on Ascend NPU
hardware. The changes ensure that GPU-specific parameters are
automatically detected and reset to Ascend-compatible values with
appropriate warnings logged.
| Module | Parameter | Default Value |
|--------|-----------|---------------|
| Model Config | `disable_cascade_attn` | `False` |
| Parallel Config | `all2all_backend` | `"allgather_reducescatter"` |
| Cache Config | `cpu_kvcache_space_bytes` | `None` |
| MultiModal Config | `mm_encoder_attn_backend` | `None` |
| Observability Config | `enable_layerwise_nvtx_tracing` | `False` |
| Scheduler Config | `max_num_partial_prefills` | `1` |
| Speculative Config | `quantization` | `None` |
| KV Transfer Config | `kv_buffer_size` | `1e9` |
| KV Transfer Config | `enable_permute_local_kv` | `False` |
| Attention Config | `use_prefill_decode_attention` | `False` |
| Attention Config | `use_cudnn_prefill` | `False` |
| Attention Config | `use_trtllm_ragged_deepseek_prefill` | `False` |
| Attention Config | `use_trtllm_attention` | `False` |
| Attention Config | `disable_flashinfer_prefill` | `False` |
| Attention Config | `disable_flashinfer_q_quantization` | `False` |
| Attention Config | `flash_attn_version` | `None` |
| Attention Config | `backend` | `None` |
| Attention Config | `flash_attn_max_num_splits_for_cuda_graph` | `32` |
### Does this PR introduce _any_ user-facing change?
### How was this patch tested?
- vLLM version: v0.13.0
- vLLM main:
2c24bc6996
Signed-off-by: hfadzxy <starmoon_zhang@163.com>
2026-01-20 11:02:38 +08:00
..
2026-01-16 09:52:48 +00:00
2026-01-07 09:03:45 +08:00
2026-01-06 16:41:39 +08:00
2026-01-13 09:21:28 +08:00
2026-01-15 08:57:40 +08:00
2026-01-20 11:02:38 +08:00
2025-06-16 18:32:28 +08:00
2026-01-15 08:57:40 +08:00
2025-12-30 15:05:47 +08:00
2026-01-19 09:24:25 +08:00
2025-10-21 20:19:46 +08:00
2026-01-15 10:26:44 +08:00
2026-01-07 18:41:45 +08:00
2026-01-19 08:58:07 +08:00
2026-01-19 09:27:55 +08:00
2025-07-21 19:43:30 +08:00
2025-07-28 15:13:37 +08:00
2026-01-06 08:44:29 +08:00
2026-01-20 11:02:38 +08:00
2025-08-14 09:33:39 +08:00
2026-01-20 11:02:38 +08:00
2025-12-19 14:27:24 +08:00