Step-3.5-Flash NVIDIA inference image
Rebuilds the text-inference environment validated on seven NVIDIA H800 GPUs: vLLM 0.25.0, CUDA 12.9, BF16, tensor parallelism 1 and pipeline parallelism 7.
TorchCodec is removed because the original image fails to import it with a
missing libnvrtc.so.13 dependency. The build verifies the API server import.
Model weights are not included. Mount them and supply serving arguments at runtime.
ModelHub release
The workflow is copied from https://dev.modelhub.org.cn/4pdadmin/cicd_demo.
Push a new v* Git tag to trigger image build, push, and review submission.
The runner supplies DOCKER_REGISTRY, DOCKER_USERNAME, DOCKER_PASSWORD,
and FIXED_TOKEN. It must be able to pull the Harbor base image.
ModelHub validates GPU_TYPE="Nvidia-H800" and TASK_TYPE=text-generation
before building. Approval is required before selecting the image for evaluation.
The image inherits its base image's entrypoint; the validated deployment overrides
it with python3 -m vllm.entrypoints.openai.api_server. Docker run options and
host model paths are not embedded into this image.