### What this PR does / why we need it?
vllm model runner v2 use uva buffer to prepare input data, but npu
doesn't support uva yet, this pr implement a uvawrapper class to mimic
gpu's uva backend. what's more, this pr make some modifications to adapt
to the newer main branch.
### Does this PR introduce _any_ user-facing change?
no
### How was this patch tested?
- vLLM main:
13397841ab
---------
Signed-off-by: Ronald1995 <ronaldautomobile@163.com>
9 lines
262 B
Markdown
9 lines
262 B
Markdown
# [Experimental] Model Runner V2
|
|
|
|
This directory contains the new model runner which is under active development.
|
|
|
|
please see [Model Runner V2](https://github.com/vllm-project/vllm-ascend/issues/5208)
|
|
to get specific plans.
|
|
|
|
supported vllm version: main@1339784
|