Add NVIDIA vLLM image and ModelHub release workflow
Some checks failed
Docker Build and Push / docker (push) Failing after 2s
Some checks failed
Docker Build and Push / docker (push) Failing after 2s
This commit is contained in:
21
README.md
Normal file
21
README.md
Normal file
@@ -0,0 +1,21 @@
|
||||
# Step-3.5-Flash NVIDIA inference image
|
||||
|
||||
Rebuilds the text-inference environment validated on seven NVIDIA H800 GPUs:
|
||||
vLLM 0.25.0, CUDA 12.9, BF16, tensor parallelism 1 and pipeline parallelism 7.
|
||||
|
||||
TorchCodec is removed because the original image fails to import it with a
|
||||
missing `libnvrtc.so.13` dependency. The build verifies the API server import.
|
||||
Model weights are not included. Mount them and supply serving arguments at runtime.
|
||||
|
||||
## ModelHub release
|
||||
|
||||
The workflow is copied from https://dev.modelhub.org.cn/4pdadmin/cicd_demo.
|
||||
Push a new `v*` Git tag to trigger image build, push, and review submission.
|
||||
The runner supplies `DOCKER_REGISTRY`, `DOCKER_USERNAME`, `DOCKER_PASSWORD`,
|
||||
and `FIXED_TOKEN`. It must be able to pull the Harbor base image.
|
||||
ModelHub validates `GPU_TYPE="NVIDIA H800"` and `TASK_TYPE=text-generation`
|
||||
before building. Approval is required before selecting the image for evaluation.
|
||||
|
||||
The image inherits its base image's entrypoint; the validated deployment overrides
|
||||
it with `python3 -m vllm.entrypoints.openai.api_server`. Docker run options and
|
||||
host model paths are not embedded into this image.
|
||||
Reference in New Issue
Block a user