[v0.18.0][Doc] Translated Doc files 2026-04-14 (#8257)

## Auto-Translation Summary

Translated **102** file(s):

-
<code>docs/source/locale/zh_CN/LC_MESSAGES/community/contributors.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/community/governance.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/community/user_stories/index.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/community/user_stories/llamafactory.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/community/versioning_policy.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/patch.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/contribution/index.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/contribution/testing.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/evaluation/using_evalscope.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/evaluation/using_lm_eval.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/evaluation/using_opencompass.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/performance_and_debug/msprobe_guide.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/performance_and_debug/performance_benchmark.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/performance_and_debug/service_profiling_guide.po</code>
- <code>docs/source/locale/zh_CN/LC_MESSAGES/faqs.po</code>
- <code>docs/source/locale/zh_CN/LC_MESSAGES/index.po</code>
- <code>docs/source/locale/zh_CN/LC_MESSAGES/installation.po</code>
- <code>docs/source/locale/zh_CN/LC_MESSAGES/quick_start.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/configuration/additional_config.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/graph_mode.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/lora.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/quantization.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/sleep_mode.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/structured_output.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/release_notes.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/support_matrix/index.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/support_matrix/supported_features.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/support_matrix/supported_models.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/ACL_Graph.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/KV_Cache_Pool_Guide.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/ModelRunner_prepare_inputs.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/add_custom_aclnn_op.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/context_parallel.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/cpu_binding.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/disaggregated_prefill.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/eplb_swift_balancer.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/npugraph_ex.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/quantization.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/contribution/multi_node_test.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/evaluation/using_ais_bench.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/performance_and_debug/optimization_and_tuning.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/index.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/long_sequence_context_parallel_multi_node.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/long_sequence_context_parallel_single_node.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/pd_colocated_mooncake_multi_instance.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/pd_disaggregation_mooncake_multi_node.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/pd_disaggregation_mooncake_single_node.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/ray.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/suffix_speculative_decoding.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/hardwares/310p.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/hardwares/index.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/DeepSeek-R1.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/DeepSeek-V3.1.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/DeepSeek-V3.2.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/GLM4.x.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/GLM5.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Kimi-K2-Thinking.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Kimi-K2.5.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/MiniMax-M2.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/PaddleOCR-VL.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen-VL-Dense.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen2.5-7B.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen2.5-Omni.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-235B-A22B.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-30B-A3B.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-32B-W4A4.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-8B-W4A8.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-Coder-30B-A3B.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-Dense.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-Next.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-VL-235B-A22B-Instruct.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-VL-30B-A3B-Instruct.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-VL-Embedding.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-VL-Reranker.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3.5-27B.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3.5-397B-A17B.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3_embedding.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3_reranker.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/index.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/deployment_guide/index.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/deployment_guide/using_volcano_kthena.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/Fine_grained_TP.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/Multi_Token_Prediction.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/batch_invariance.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/context_parallel.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/cpu_binding.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/dynamic_batch.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/epd_disaggregation.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/eplb_swift_balancer.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/external_dp.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/kv_pool.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/large_scale_ep.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/layer_sharding.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/lmcache_ascend_deployment.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/netloader.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/npugraph_ex.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/rfork.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/sequence_parallelism.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/speculative_decoding.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/ucm_deployment.po</code>
-
<code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/weight_prefetch.po</code>

---

[Workflow
run](https://github.com/vllm-project/vllm-ascend/actions/runs/24390263284)

Signed-off-by: vllm-ascend-ci <vllm-ascend-ci@users.noreply.github.com>
Co-authored-by: vllm-ascend-ci <vllm-ascend-ci@users.noreply.github.com>
This commit is contained in:
vllm-ascend-ci
2026-04-15 15:27:09 +08:00
committed by GitHub
parent b6aa5bbdbf
commit 147b589f62
102 changed files with 41760 additions and 6023 deletions

View File

@@ -0,0 +1,29 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/features/index.md:1
#: ../../source/tutorials/features/index.md:5
msgid "Feature Tutorials"
msgstr "功能教程"
#: ../../source/tutorials/features/index.md:3
msgid "This section provides tutorials for different features of vLLM Ascend."
msgstr "本节提供 vLLM Ascend 不同功能的使用教程。"

View File

@@ -0,0 +1,447 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:1
msgid "Long-Sequence Context Parallel (Deepseek)"
msgstr "长序列上下文并行 (Deepseek)"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:3
msgid "Getting Started"
msgstr "快速开始"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:6
msgid ""
"Context parallel feature currently is only supported on Atlas A3 device, "
"and will be supported on Atlas A2 in the future."
msgstr "上下文并行特性目前仅在 Atlas A3 设备上受支持,未来将在 Atlas A2 上提供支持。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:9
msgid ""
"vLLM-Ascend now supports long sequence with context parallel options. "
"This guide takes one-by-one steps to verify these features with "
"constrained resources."
msgstr "vLLM-Ascend 现已支持长序列上下文并行选项。本指南将逐步引导您在有限资源下验证这些功能。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:11
msgid ""
"Take the Deepseek-V3.1-w8a8 model as an example, use 3 Atlas 800T A3 "
"servers to deploy the “1P1D” architecture. Node p is deployed across "
"multiple machines, while node d is deployed on a single machine. Assume "
"the IP of the prefiller server is 192.0.0.1 (prefill 1) and 192.0.0.2 "
"(prefill 2), and the decoder servers are 192.0.0.3 (decoder 1). On each "
"server, use 8 NPUs 16 chips to deploy one service instance. In the "
"current example, we will enable the context parallel feature on node p to"
" improve TTFT. Although enabling the DCP feature on node d can reduce "
"memory usage, it would introduce additional communication and small "
"operator overhead. Therefore, we will not enable the DCP feature on node "
"d."
msgstr "以 Deepseek-V3.1-w8a8 模型为例,使用 3 台 Atlas 800T A3 服务器部署“1P1D”架构。节点 p 跨多台机器部署,而节点 d 部署在单台机器上。假设预填充服务器的 IP 为 192.0.0.1(预填充 1和 192.0.0.2(预填充 2解码器服务器为 192.0.0.3(解码器 1。每台服务器使用 8 个 NPU16 个芯片)部署一个服务实例。在当前示例中,我们将在节点 p 上启用上下文并行特性以改善 TTFT。虽然在节点 d 上启用 DCP 特性可以减少内存使用,但会引入额外的通信和小算子开销。因此,我们不会在节点 d 上启用 DCP 特性。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:13
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:15
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:17
msgid ""
"`DeepSeek-V3.1_w8a8mix_mtp` (Quantized version with mix mtp): [Download "
"model weight](https://www.modelscope.cn/models/Eco-"
"Tech/DeepSeek-V3.1-w8a8). Please modify `torch_dtype` from `float16` to "
"`bfloat16` in `config.json`."
msgstr "`DeepSeek-V3.1_w8a8mix_mtp`(混合 MTP 量化版本):[下载模型权重](https://www.modelscope.cn/models/Eco-Tech/DeepSeek-V3.1-w8a8)。请在 `config.json` 中将 `torch_dtype` 从 `float16` 修改为 `bfloat16`。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:19
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`"
msgstr "建议将模型权重下载到多个节点的共享目录中,例如 `/root/.cache/`"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:21
msgid "Verify Multi-node Communication"
msgstr "验证多节点通信"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:23
msgid ""
"Refer to [verify multi-node communication "
"environment](../../installation.md#verify-multi-node-communication) to "
"verify multi-node communication."
msgstr "请参考[验证多节点通信环境](../../installation.md#verify-multi-node-communication)来验证多节点通信。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:25
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:27
msgid "You can use our official Docker image to run `DeepSeek-V3.1` directly."
msgstr "您可以使用我们的官方 Docker 镜像直接运行 `DeepSeek-V3.1`。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:29
msgid ""
"Select an image based on your machine type and start the Docker image on "
"your node, refer to [using Docker](../../installation.md#set-up-using-"
"docker)."
msgstr "根据您的机器类型选择镜像并在节点上启动 Docker 镜像,请参考[使用 Docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:64
msgid "You need to set up environment on each node."
msgstr "您需要在每个节点上设置环境。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:66
msgid "Prefiller/Decoder Deployment"
msgstr "预填充器/解码器部署"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:68
msgid ""
"We can run the following scripts to launch a server on the "
"prefiller/decoder node, respectively. Please note that each P/D node will"
" occupy ports ranging from kv_port to kv_port + num_chips to initialize "
"socket listeners. To avoid any issues, port conflicts should be "
"prevented. Additionally, ensure that each node's engine_id is uniquely "
"assigned to avoid conflicts."
msgstr "我们可以分别在预填充器/解码器节点上运行以下脚本来启动服务器。请注意,每个 P/D 节点将占用从 kv_port 到 kv_port + num_chips 的端口范围来初始化 socket 监听器。为避免任何问题,应防止端口冲突。此外,请确保每个节点的 engine_id 被唯一分配以避免冲突。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:70
msgid ""
"Run the following script to execute online 128k inference on three nodes "
"respectively."
msgstr "运行以下脚本,分别在三个节点上执行在线 128k 推理。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md
msgid "Prefiller node 1"
msgstr "预填充节点 1"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md
msgid "Prefiller node 2"
msgstr "预填充节点 2"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md
msgid "Decoder node 1"
msgstr "解码节点 1"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:276
msgid "Prefill master node `proxy.sh` script"
msgstr "预填充主节点 `proxy.sh` 脚本"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:292
msgid "Run proxy"
msgstr "运行代理"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:294
msgid ""
"Run a proxy server on the same node with the prefiller service instance. "
"You can get the proxy program in the repository's examples: "
"[load\\_balance\\_proxy\\_server\\_example.py](https://github.com/vllm-"
"project/vllm-"
"ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
msgstr "在与预填充服务实例相同的节点上运行代理服务器。您可以在仓库的示例中找到代理程序:[load_balance_proxy_server_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:301
msgid "**Notice:** The parameters are explained as follows:"
msgstr "**注意:** 参数解释如下:"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:304
msgid ""
"`--tensor-parallel-size` 16 are common settings for tensor parallelism "
"(TP) sizes."
msgstr "`--tensor-parallel-size` 16 是张量并行TP大小的常见设置。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:305
msgid ""
"`--prefill-context-parallel-size` 2 are common settings for prefill "
"context parallelism (PCP) sizes."
msgstr "`--prefill-context-parallel-size` 2 是预填充上下文并行PCP大小的常见设置。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:306
msgid ""
"`--decode-context-parallel-size` 8 are common settings for decode context"
" parallelism (DCP) sizes."
msgstr "`--decode-context-parallel-size` 8 是解码上下文并行DCP大小的常见设置。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:307
msgid ""
"`--max-model-len` represents the context length, which is the maximum "
"value of the input plus output for a single request."
msgstr "`--max-model-len` 表示上下文长度,即单个请求的输入加输出的最大值。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:308
msgid ""
"`--max-num-seqs` indicates the maximum number of requests that each DP "
"group is allowed to process. If the number of requests sent to the "
"service exceeds this limit, the excess requests will remain in a waiting "
"state and will not be scheduled. Note that the time spent in the waiting "
"state is also counted in metrics such as TTFT and TPOT. Therefore, when "
"testing performance, it is generally recommended that `--max-num-seqs` * "
"`--data-parallel-size` >= the actual total concurrency."
msgstr "`--max-num-seqs` 表示每个 DP 组允许处理的最大请求数。如果发送到服务的请求数量超过此限制,超出的请求将保持在等待状态,不会被调度。请注意,在等待状态所花费的时间也会计入 TTFT 和 TPOT 等指标。因此,在测试性能时,通常建议 `--max-num-seqs` * `--data-parallel-size` >= 实际总并发数。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:309
msgid ""
"`--max-num-batched-tokens` represents the maximum number of tokens that "
"the model can process in a single step. Currently, vLLM v1 scheduling "
"enables ChunkPrefill/SplitFuse by default, which means:"
msgstr "`--max-num-batched-tokens` 表示模型单步可以处理的最大 token 数。目前vLLM v1 调度默认启用 ChunkPrefill/SplitFuse这意味着"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:310
msgid ""
"(1) If the input length of a request is greater than `--max-num-batched-"
"tokens`, it will be divided into multiple rounds of computation according"
" to `--max-num-batched-tokens`;"
msgstr "1如果请求的输入长度大于 `--max-num-batched-tokens`,它将根据 `--max-num-batched-tokens` 被分成多轮计算;"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:311
msgid ""
"(2) Decode requests are prioritized for scheduling, and prefill requests "
"are scheduled only if there is available capacity."
msgstr "2解码请求优先调度预填充请求仅在有空闲容量时才会被调度。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:312
msgid ""
"Generally, if `--max-num-batched-tokens` is set to a larger value, the "
"overall latency will be lower, but the pressure on GPU memory (activation"
" value usage) will be greater."
msgstr "通常,如果 `--max-num-batched-tokens` 设置得较大,整体延迟会更低,但 GPU 内存(激活值使用)的压力会更大。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:313
msgid ""
"`--gpu-memory-utilization` represents the proportion of HBM that vLLM "
"will use for actual inference. Its essential function is to calculate the"
" available kv_cache size. During the warm-up phase (referred to as "
"profile run in vLLM), vLLM records the peak GPU memory usage during an "
"inference process with an input size of `--max-num-batched-tokens`. The "
"available kv_cache size is then calculated as: `--gpu-memory-utilization`"
" * HBM size - peak GPU memory usage. Therefore, the larger the value of "
"`--gpu-memory-utilization`, the more kv_cache can be used. However, since"
" the GPU memory usage during the warm-up phase may differ from that "
"during actual inference (e.g., due to uneven EP load), setting `--gpu-"
"memory-utilization` too high may lead to OOM (Out of Memory) issues "
"during actual inference. The default value is `0.9`."
msgstr "`--gpu-memory-utilization` 表示 vLLM 将用于实际推理的 HBM 比例。其核心功能是计算可用的 kv_cache 大小。在预热阶段vLLM 中称为 profile runvLLM 会记录输入大小为 `--max-num-batched-tokens` 的推理过程中的峰值 GPU 内存使用量。然后,可用的 kv_cache 大小计算为:`--gpu-memory-utilization` * HBM 大小 - 峰值 GPU 内存使用量。因此,`--gpu-memory-utilization` 的值越大,可用的 kv_cache 就越多。然而,由于预热阶段的 GPU 内存使用量可能与实际推理期间不同(例如,由于 EP 负载不均),将 `--gpu-memory-utilization` 设置得过高可能导致实际推理时出现 OOM内存不足问题。默认值为 `0.9`。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:314
msgid ""
"`--enable-expert-parallel` indicates that EP is enabled. Note that vLLM "
"does not support a mixed approach of ETP and EP; that is, MoE can either "
"use pure EP or pure TP."
msgstr "`--enable-expert-parallel` 表示启用了 EP。请注意vLLM 不支持 ETP 和 EP 的混合方法也就是说MoE 只能使用纯 EP 或纯 TP。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:315
msgid ""
"`--no-enable-prefix-caching` indicates that prefix caching is disabled. "
"To enable it, remove this option."
msgstr "`--no-enable-prefix-caching` 表示前缀缓存被禁用。要启用它,请移除此选项。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:316
msgid ""
"`--quantization` \"ascend\" indicates that quantization is used. To "
"disable quantization, remove this option."
msgstr "`--quantization` \"ascend\" 表示使用了量化。要禁用量化,请移除此选项。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:317
msgid ""
"`--compilation-config` contains configurations related to the aclgraph "
"graph mode. The most significant configurations are \"cudagraph_mode\" "
"and \"cudagraph_capture_sizes\", which have the following meanings: "
"\"cudagraph_mode\": represents the specific graph mode. Currently, "
"\"PIECEWISE\" and \"FULL_DECODE_ONLY\" are supported. The graph mode is "
"mainly used to reduce the cost of operator dispatch. Currently, "
"\"FULL_DECODE_ONLY\" is recommended."
msgstr "`--compilation-config` 包含与 aclgraph 图模式相关的配置。最重要的配置是 \"cudagraph_mode\" 和 \"cudagraph_capture_sizes\",其含义如下:\"cudagraph_mode\":表示特定的图模式。目前支持 \"PIECEWISE\" 和 \"FULL_DECODE_ONLY\"。图模式主要用于降低算子调度的开销。目前推荐使用 \"FULL_DECODE_ONLY\"。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:319
msgid ""
"\"cudagraph_capture_sizes\": represents different levels of graph modes. "
"The default value is [1, 2, 4, 8, 16, 24, 32, 40,..., `--max-num-seqs`]. "
"In the graph mode, the input for graphs at different levels is fixed, and"
" inputs between levels are automatically padded to the next level. "
"Currently, the default setting is recommended. Only in some scenarios is "
"it necessary to set this separately to achieve optimal performance."
msgstr "\"cudagraph_capture_sizes\":表示不同级别的图模式。默认值为 [1, 2, 4, 8, 16, 24, 32, 40,..., `--max-num-seqs`]。在图模式下,不同级别图的输入是固定的,级别之间的输入会自动填充到下一级别。目前推荐使用默认设置。仅在部分场景中,需要单独设置此参数以达到最佳性能。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:320
msgid ""
"`export VLLM_ASCEND_ENABLE_FLASHCOMM1=1` indicates that Flashcomm1 "
"optimization is enabled. Currently, this optimization is only supported "
"for MoE in scenarios where tensor-parallel-size > 1."
msgstr "`export VLLM_ASCEND_ENABLE_FLASHCOMM1=1` 表示启用了 Flashcomm1 优化。目前,此优化仅在 tensor-parallel-size > 1 的场景下对 MoE 提供支持。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:321
msgid ""
"`export VLLM_ASCEND_ENABLE_CONTEXT_PARALLEL=1` indicates that context "
"parallel is enabled. This environment variable is required in the PD "
"architecture but not needed in the PD co-locate deployment scenario. It "
"will be removed in the future."
msgstr "`export VLLM_ASCEND_ENABLE_CONTEXT_PARALLEL=1` 表示启用了上下文并行。此环境变量在 PD 架构中是必需的,但在 PD 共置部署场景中不需要。未来将被移除。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:323
msgid "**Notice:**"
msgstr "**注意:**"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:325
msgid ""
"tensor-parallel-size needs to be divisible by decode-context-parallel-"
"size."
msgstr "tensor-parallel-size 需要能被 decode-context-parallel-size 整除。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:326
msgid ""
"decode-context-parallel-size must be less than or equal to tensor-"
"parallel-size."
msgstr "decode-context-parallel-size 必须小于或等于 tensor-parallel-size。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:328
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:330
msgid "Here are two accuracy evaluation methods."
msgstr "以下是两种精度评估方法。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:332
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:344
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:334
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参考[使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:336
msgid ""
"After execution, you can get the result, here is the result of "
"`DeepSeek-V3.1-w8a8` for reference only."
msgstr "执行后,您可以获得结果,以下是 `DeepSeek-V3.1-w8a8` 的结果,仅供参考。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "version"
msgstr "版本"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "aime2024"
msgstr "aime2024"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "-"
msgstr "-"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "accuracy"
msgstr "准确率"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "gen"
msgstr "生成"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "86.67"
msgstr "86.67"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:342
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:346
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参阅[使用 AISBench 进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:348
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM 基准测试"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:350
msgid "Run performance evaluation of `DeepSeek-V3.1-w8a8` as an example."
msgstr "以运行 `DeepSeek-V3.1-w8a8` 的性能评估为例。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:352
msgid ""
"Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) "
"for more details."
msgstr "更多详情请参阅 [vllm 基准测试](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:354
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 包含三个子命令:"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:356
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批请求的延迟进行基准测试。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:357
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:358
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:360
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例,按如下方式运行代码。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:367
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您将获得性能评估结果。"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "ttft"
msgstr "首字元延迟"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "random"
msgstr "随机"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "performance"
msgstr "性能"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "perf"
msgstr "性能"
#: ../../source/tutorials/features/long_sequence_context_parallel_multi_node.md:211
msgid "20.7s"
msgstr "20.7秒"

View File

@@ -0,0 +1,386 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:1
msgid "Long-Sequence Context Parallel (Qwen3-235B-A22B)"
msgstr "长序列上下文并行 (Qwen3-235B-A22B)"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:3
msgid "Getting Started"
msgstr "快速开始"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:5
msgid ""
"vLLM-Ascend now supports long-sequence context parallel. This guide takes"
" one-by-one steps to verify these features with constrained resources."
msgstr "vLLM-Ascend 现已支持长序列上下文并行。本指南将引导您在使用有限资源的情况下,逐步验证这些功能。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:7
msgid ""
"Using the `Qwen3-235B-A22B-w8a8` (Quantized version) model as an example,"
" use 1 Atlas 800 A3 (64G × 16) server to deploy the single node \"pd co-"
"locate\" architecture."
msgstr "以 `Qwen3-235B-A22B-w8a8`(量化版本)模型为例,使用 1 台 Atlas 800 A364G × 16服务器部署单节点 \"pd co-locate\" 架构。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:9
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:11
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:13
msgid ""
"`Qwen3-235B-A22B-w8a8` (Quantized version): requires 1 Atlas 800 A3 (64G "
"× 16) node. [Download model weight](https://modelscope.cn/models/vllm-"
"ascend/Qwen3-235B-A22B-W8A8)"
msgstr "`Qwen3-235B-A22B-w8a8`(量化版本):需要 1 个 Atlas 800 A364G × 16节点。[下载模型权重](https://modelscope.cn/models/vllm-ascend/Qwen3-235B-A22B-W8A8)"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:15
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`"
msgstr "建议将模型权重下载到多节点的共享目录,例如 `/root/.cache/`"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:17
msgid "Run with Docker"
msgstr "使用 Docker 运行"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:19
msgid "Start a Docker container on each node."
msgstr "在每个节点上启动一个 Docker 容器。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:63
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:65
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:67
msgid ""
"`Qwen3-235B-A22B-w8a8` can be deployed on 1 Atlas 800 A364G*16. "
"Quantized version needs to start with parameter `--quantization ascend`."
msgstr "`Qwen3-235B-A22B-w8a8` 可以部署在 1 台 Atlas 800 A364G*16上。量化版本需要使用参数 `--quantization ascend` 启动。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:70
msgid "Run the following script to execute online 128k inference."
msgstr "运行以下脚本以执行在线 128k 推理。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:106
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:131
msgid "**Notice:**"
msgstr "**注意:**"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:108
#, python-brace-format
msgid ""
"for vllm version below `v0.12.0` use parameter: `--rope_scaling "
"'{\"rope_type\":\"yarn\",\"factor\":4,\"original_max_position_embeddings\":32768}'"
" \\`"
msgstr "对于 vllm 版本低于 `v0.12.0`,使用参数:`--rope_scaling '{\"rope_type\":\"yarn\",\"factor\":4,\"original_max_position_embeddings\":32768}' \\`"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:109
#, python-brace-format
msgid ""
"for vllm version `v0.12.0` use parameter: `--hf-overrides "
"'{\"rope_parameters\": "
"{\"rope_type\":\"yarn\",\"rope_theta\":1000000,\"factor\":4,\"original_max_position_embeddings\":32768}}'"
" \\`"
msgstr "对于 vllm 版本 `v0.12.0`,使用参数:`--hf-overrides '{\"rope_parameters\": {\"rope_type\":\"yarn\",\"rope_theta\":1000000,\"factor\":4,\"original_max_position_embeddings\":32768}}' \\`"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:111
msgid "The parameters are explained as follows:"
msgstr "参数解释如下:"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:113
msgid ""
"`--tensor-parallel-size` 8 are common settings for tensor parallelism "
"(TP) sizes."
msgstr "`--tensor-parallel-size` 8 是张量并行TP大小的常见设置。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:114
msgid ""
"`--prefill-context-parallel-size` 2 are common settings for prefill "
"context parallelism (PCP) sizes."
msgstr "`--prefill-context-parallel-size` 2 是预填充上下文并行PCP大小的常见设置。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:115
msgid ""
"`--decode-context-parallel-size` 2 are common settings for decode context"
" parallelism (DCP) sizes."
msgstr "`--decode-context-parallel-size` 2 是解码上下文并行DCP大小的常见设置。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:116
msgid ""
"`--max-model-len` represents the context length, which is the maximum "
"value of the input plus output for a single request."
msgstr "`--max-model-len` 表示上下文长度,即单个请求的输入加输出的最大值。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:117
msgid ""
"`--max-num-seqs` indicates the maximum number of requests that each DP "
"group is allowed to process. If the number of requests sent to the "
"service exceeds this limit, the excess requests will remain in a waiting "
"state and will not be scheduled. Note that the time spent in the waiting "
"state is also counted in metrics such as TTFT and TPOT. Therefore, when "
"testing performance, it is generally recommended that `--max-num-seqs` * "
"`--data-parallel-size` >= the actual total concurrency."
msgstr "`--max-num-seqs` 表示每个 DP 组允许处理的最大请求数。如果发送到服务的请求数量超过此限制,超出的请求将保持在等待状态,不会被调度。请注意,在等待状态所花费的时间也会计入 TTFT 和 TPOT 等指标。因此,在测试性能时,通常建议 `--max-num-seqs` * `--data-parallel-size` >= 实际总并发数。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:118
msgid ""
"`--max-num-batched-tokens` represents the maximum number of tokens that "
"the model can process in a single step. Currently, vLLM v1 scheduling "
"enables ChunkPrefill/SplitFuse by default, which means:"
msgstr "`--max-num-batched-tokens` 表示模型单步可以处理的最大 token 数。目前vLLM v1 调度默认启用 ChunkPrefill/SplitFuse这意味着"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:119
msgid ""
"(1) If the input length of a request is greater than `--max-num-batched-"
"tokens`, it will be divided into multiple rounds of computation according"
" to `--max-num-batched-tokens`;"
msgstr "1如果请求的输入长度大于 `--max-num-batched-tokens`,它将根据 `--max-num-batched-tokens` 被分成多轮计算;"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:120
msgid ""
"(2) Decode requests are prioritized for scheduling, and prefill requests "
"are scheduled only if there is available capacity."
msgstr "2解码请求优先调度预填充请求仅在有空闲容量时才会被调度。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:121
msgid ""
"Generally, if `--max-num-batched-tokens` is set to a larger value, the "
"overall latency will be lower, but the pressure on GPU memory (activation"
" value usage) will be greater."
msgstr "通常,如果 `--max-num-batched-tokens` 设置得较大,整体延迟会更低,但 GPU 内存(激活值使用)的压力会更大。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:122
msgid ""
"`--gpu-memory-utilization` represents the proportion of HBM that vLLM "
"will use for actual inference. Its essential function is to calculate the"
" available kv_cache size. During the warm-up phase (referred to as "
"profile run in vLLM), vLLM records the peak GPU memory usage during an "
"inference process with an input size of `--max-num-batched-tokens`. The "
"available kv_cache size is then calculated as: `--gpu-memory-utilization`"
" * HBM size - peak GPU memory usage. Therefore, the larger the value of "
"`--gpu-memory-utilization`, the more kv_cache can be used. However, since"
" the GPU memory usage during the warm-up phase may differ from that "
"during actual inference (e.g., due to uneven EP load), setting `--gpu-"
"memory-utilization` too high may lead to OOM (Out of Memory) issues "
"during actual inference. The default value is `0.9`."
msgstr "`--gpu-memory-utilization` 表示 vLLM 将用于实际推理的 HBM 比例。其核心功能是计算可用的 kv_cache 大小。在预热阶段vLLM 中称为 profile runvLLM 会记录输入大小为 `--max-num-batched-tokens` 的推理过程中的峰值 GPU 内存使用量。然后,可用的 kv_cache 大小计算为:`--gpu-memory-utilization` * HBM 大小 - 峰值 GPU 内存使用量。因此,`--gpu-memory-utilization` 的值越大,可用的 kv_cache 就越多。然而,由于预热阶段的 GPU 内存使用量可能与实际推理时不同(例如,由于 EP 负载不均),将 `--gpu-memory-utilization` 设置得过高可能导致实际推理时出现 OOM内存不足问题。默认值为 `0.9`。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:123
msgid ""
"`--enable-expert-parallel` indicates that EP is enabled. Note that vLLM "
"does not support a mixed approach of ETP and EP; that is, MoE can either "
"use pure EP or pure TP."
msgstr "`--enable-expert-parallel` 表示启用了 EP。请注意vLLM 不支持 ETP 和 EP 的混合方法也就是说MoE 要么使用纯 EP要么使用纯 TP。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:124
msgid ""
"`--no-enable-prefix-caching` indicates that prefix caching is disabled. "
"To enable it, remove this option."
msgstr "`--no-enable-prefix-caching` 表示前缀缓存被禁用。要启用它,请移除此选项。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:125
msgid ""
"`--quantization` \"ascend\" indicates that quantization is used. To "
"disable quantization, remove this option."
msgstr "`--quantization` \"ascend\" 表示使用了量化。要禁用量化,请移除此选项。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:126
msgid ""
"`--compilation-config` contains configurations related to the aclgraph "
"graph mode. The most significant configurations are \"cudagraph_mode\" "
"and \"cudagraph_capture_sizes\", which have the following meanings: "
"\"cudagraph_mode\": represents the specific graph mode. Currently, "
"\"PIECEWISE\" and \"FULL_DECODE_ONLY\" are supported. The graph mode is "
"mainly used to reduce the cost of operator dispatch. Currently, "
"\"FULL_DECODE_ONLY\" is recommended."
msgstr "`--compilation-config` 包含与 aclgraph 图模式相关的配置。最重要的配置是 \"cudagraph_mode\" 和 \"cudagraph_capture_sizes\",其含义如下:\"cudagraph_mode\":表示具体的图模式。目前支持 \"PIECEWISE\" 和 \"FULL_DECODE_ONLY\"。图模式主要用于降低算子调度的开销。目前推荐使用 \"FULL_DECODE_ONLY\"。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:128
msgid ""
"\"cudagraph_capture_sizes\": represents different levels of graph modes. "
"The default value is [1, 2, 4, 8, 16, 24, 32, 40,..., `--max-num-seqs`]. "
"In the graph mode, the input for graphs at different levels is fixed, and"
" inputs between levels are automatically padded to the next level. "
"Currently, the default setting is recommended. Only in some scenarios is "
"it necessary to set this separately to achieve optimal performance."
msgstr "\"cudagraph_capture_sizes\":表示不同级别的图模式。默认值为 [1, 2, 4, 8, 16, 24, 32, 40,..., `--max-num-seqs`]。在图模式下,不同级别图的输入是固定的,级别之间的输入会自动填充到下一个级别。目前推荐使用默认设置。仅在部分场景中,需要单独设置此参数以达到最佳性能。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:129
msgid ""
"`export VLLM_ASCEND_ENABLE_FLASHCOMM1=1` indicates that Flashcomm1 "
"optimization is enabled. Currently, this optimization is only supported "
"for MoE in scenarios where tp_size > 1."
msgstr "`export VLLM_ASCEND_ENABLE_FLASHCOMM1=1` 表示启用了 Flashcomm1 优化。目前,此优化仅在 tp_size > 1 的场景下对 MoE 支持。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:133
msgid "tp_size needs to be divisible by dcp_size"
msgstr "tp_size 需要能被 dcp_size 整除"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:134
msgid ""
"decode context parallel size must be less than or equal to max_dcp_size, "
"where max_dcp_size = tensor_parallel_size // total_num_kv_heads."
msgstr "解码上下文并行大小必须小于或等于 max_dcp_size其中 max_dcp_size = tensor_parallel_size // total_num_kv_heads。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:136
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:138
msgid "Here are two accuracy evaluation methods."
msgstr "以下是两种精度评估方法。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:140
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:152
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:142
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参阅[使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:144
msgid ""
"After execution, you can get the result, here is the result of `Qwen3"
"-235B-A22B-w8a8` for reference only."
msgstr "执行后,您可以获得结果,以下是 `Qwen3-235B-A22B-w8a8` 的结果,仅供参考。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "version"
msgstr "版本"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "aime2024"
msgstr "aime2024"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "-"
msgstr "-"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "accuracy"
msgstr "准确率"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "gen"
msgstr "生成"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "83.33"
msgstr "83.33"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:150
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:154
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参阅[使用 AISBench 进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:156
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM Benchmark"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:158
msgid "Run performance evaluation of `Qwen3-235B-A22B-w8a8` as an example."
msgstr "以运行 `Qwen3-235B-A22B-w8a8` 的性能评估为例。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:160
msgid ""
"Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) "
"for more details."
msgstr "更多详情请参阅 [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:162
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 有三个子命令:"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:164
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批请求的延迟进行基准测试。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:165
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:166
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:168
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。运行代码如下。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:175
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您将获得性能评估结果。"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "ttft"
msgstr "首词元时间"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "random"
msgstr "随机"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "performance"
msgstr "性能"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "perf"
msgstr "性能"
#: ../../source/tutorials/features/long_sequence_context_parallel_single_node.md:21
msgid "17.36s"
msgstr "17.36秒"

View File

@@ -0,0 +1,509 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:1
msgid "PD-Colocated with Mooncake Multi-Instance"
msgstr "PD 共置与 Mooncake 多实例"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:3
msgid "Getting Started"
msgstr "快速开始"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:5
msgid ""
"vLLM-Ascend now supports PD-colocated deployment with Mooncake features. "
"This guide provides step-by-step instructions to test these features with"
" constrained resources."
msgstr "vLLM-Ascend 现已支持结合 Mooncake 功能的 PD 共置部署。本指南提供了在有限资源下测试这些功能的逐步说明。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:9
msgid ""
"Using the Qwen2.5-72B-Instruct model as an example, this guide "
"demonstrates how to use vllm-ascend v0.11.0 (with vLLM v0.11.0) on two "
"Atlas 800T A2 nodes to deploy two vLLM instances. Each instance occupies "
"4 NPU cards and uses PD-colocated deployment."
msgstr "本指南以 Qwen2.5-72B-Instruct 模型为例,演示如何在两个 Atlas 800T A2 节点上使用 vllm-ascend v0.11.0(包含 vLLM v0.11.0)部署两个 vLLM 实例。每个实例占用 4 个 NPU 卡,并采用 PD 共置部署。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:14
msgid "Verify Multi-Node Communication Environment"
msgstr "验证多节点通信环境"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:16
msgid "Physical Layer Requirements"
msgstr "物理层要求"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:18
msgid ""
"The two Atlas 800T A2 nodes must be physically interconnected via a RoCE "
"network. Without RoCE interconnection, cross-node KV Cache access "
"performance will be significantly degraded."
msgstr "两个 Atlas 800T A2 节点必须通过 RoCE 网络进行物理互连。若无 RoCE 互连,跨节点 KV Cache 访问性能将显著下降。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:21
msgid ""
"All NPU cards must communicate properly. Intra-node communication uses "
"HCCS, while inter-node communication uses the RoCE network."
msgstr "所有 NPU 卡必须能够正常通信。节点内通信使用 HCCS节点间通信使用 RoCE 网络。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:24
msgid "Verification Process"
msgstr "验证流程"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:26
msgid ""
"The following process serves as a reference example. Please modify "
"parameters such as IP addresses according to your actual environment."
msgstr "以下流程作为参考示例。请根据您的实际环境修改 IP 地址等参数。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:29
msgid "Single Node Verification:"
msgstr "单节点验证:"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:31
msgid ""
"Execute the following commands sequentially. The results must all be "
"`success` and the status must be `UP`:"
msgstr "依次执行以下命令。结果必须全部为 `success` 且状态必须为 `UP`"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:47
msgid "Check NPU HCCN Configuration:"
msgstr "检查 NPU HCCN 配置:"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:49
msgid ""
"Ensure that the hccn.conf file exists in the environment. If using "
"Docker, mount it into the container."
msgstr "确保环境中存在 hccn.conf 文件。如果使用 Docker请将其挂载到容器中。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:56
msgid "Get NPU IP Addresses:"
msgstr "获取 NPU IP 地址:"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:62
msgid "Cross-Node PING Test:"
msgstr "跨节点 PING 测试:"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:70
msgid "Check NPU TLS Configuration"
msgstr "检查 NPU TLS 配置"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:77
msgid "Run with Docker"
msgstr "使用 Docker 运行"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:79
msgid "Start a Docker container on each node."
msgstr "在每个节点上启动一个 Docker 容器。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:112
msgid "(Optional) Install Mooncake"
msgstr "(可选)安装 Mooncake"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:114
msgid ""
"Mooncake is pre-installed and functional in the v0.11.0 image. The "
"following installation steps are optional."
msgstr "Mooncake 在 v0.11.0 镜像中已预安装且功能正常。以下安装步骤是可选的。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:117
msgid ""
"Mooncake is the serving platform for Kimi, a leading LLM service provided"
" by Moonshot AI. Installation and compilation guide: <https://github.com"
"/kvcache-ai/Mooncake?tab=readme-ov-file#build-and-use-binaries>."
msgstr "Mooncake 是 Kimi 的服务平台Kimi 是由 Moonshot AI 提供的领先 LLM 服务。安装和编译指南:<https://github.com/kvcache-ai/Mooncake?tab=readme-ov-file#build-and-use-binaries>。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:121
msgid "First, obtain the Mooncake project using the following command:"
msgstr "首先,使用以下命令获取 Mooncake 项目:"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:129
msgid "Install MPI:"
msgstr "安装 MPI"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:135
msgid "Install the relevant dependencies (Go installation is not required):"
msgstr "安装相关依赖(无需安装 Go"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:141
msgid "Compile and install:"
msgstr "编译并安装:"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:151
msgid "After installation, verify that Mooncake is installed correctly:"
msgstr "安装后,验证 Mooncake 是否正确安装:"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:160
msgid "Start Mooncake Master Service"
msgstr "启动 Mooncake Master 服务"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:162
msgid ""
"To start the Mooncake master service in one of the node containers, use "
"the following command:"
msgstr "要在其中一个节点容器中启动 Mooncake master 服务,请使用以下命令:"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Parameter"
msgstr "参数"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Value"
msgstr "值"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Explanation"
msgstr "说明"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "port"
msgstr "端口"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "50088"
msgstr "50088"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Port for the master service"
msgstr "Master 服务端口"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "eviction_high_watermark_ratio"
msgstr "驱逐高水位线比例"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "0.95"
msgstr "0.95"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "High watermark ratio (95% threshold)"
msgstr "高水位线比例95% 阈值)"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "eviction_ratio"
msgstr "驱逐比例"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "0.05"
msgstr "0.05"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Percentage to evict when full (5%)"
msgstr "缓存满时驱逐的百分比5%"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:179
msgid "Create a Mooncake Configuration File Named mooncake.json"
msgstr "创建名为 mooncake.json 的 Mooncake 配置文件"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:181
msgid "The template for the mooncake.json file is as follows:"
msgstr "mooncake.json 文件的模板如下:"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "metadata_server"
msgstr "元数据服务器"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "P2PHANDSHAKE"
msgstr "P2PHANDSHAKE"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Point-to-point handshake mode"
msgstr "点对点握手模式"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "protocol"
msgstr "协议"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "ascend"
msgstr "ascend"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Ascend proprietary protocol"
msgstr "Ascend 专有协议"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "master_server_address"
msgstr "主服务器地址"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "90.90.100.188:50088(for example)"
msgstr "90.90.100.188:50088示例"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Master server address"
msgstr "主服务器地址"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "global_segment_size"
msgstr "全局段大小"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "107374182400"
msgstr "107374182400"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Size per segment (100 GB)"
msgstr "每个段的大小100 GB"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:200
msgid "vLLM Instance Deployment"
msgstr "vLLM 实例部署"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:202
msgid ""
"Create containers on both Node 1 and Node 2, and launch the Qwen2.5-72B-"
"Instruct model service in each to test the reusability and performance of"
" cross-node, cross-instance KV Cache. Instance 1 utilizes NPU cards [0-3]"
" on the first Atlas 800T A2 server, while Instance 2 utilizes cards [0-3]"
" on the second server."
msgstr "在节点 1 和节点 2 上分别创建容器,并在每个容器中启动 Qwen2.5-72B-Instruct 模型服务,以测试跨节点、跨实例 KV Cache 的可重用性和性能。实例 1 使用第一个 Atlas 800T A2 服务器上的 NPU 卡 [0-3],而实例 2 使用第二个服务器上的卡 [0-3]。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:208
msgid "Deploy Instance 1"
msgstr "部署实例 1"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:210
msgid ""
"Replace file paths, host, and port parameters based on your actual "
"environment configuration."
msgstr "请根据您的实际环境配置替换文件路径、主机和端口参数。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:242
msgid "Deploy Instance 2"
msgstr "部署实例 2"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:244
msgid ""
"The deployment method for Instance 2 is identical to Instance 1. Simply "
"modify the `--host` and `--port` parameters according to your Instance 2 "
"configuration."
msgstr "实例 2 的部署方法与实例 1 相同。只需根据您的实例 2 配置修改 `--host` 和 `--port` 参数。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:248
msgid "Configuration Parameters"
msgstr "配置参数"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "kv_connector"
msgstr "kv_connector"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "MooncakeConnectorStoreV1"
msgstr "MooncakeConnectorStoreV1"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Use StoreV1 version"
msgstr "使用 StoreV1 版本"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "kv_role"
msgstr "kv_role"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "kv_both"
msgstr "kv_both"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Enable both produce and consume"
msgstr "同时启用生产和消费"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "use_layerwise"
msgstr "use_layerwise"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "false"
msgstr "false"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Transfer entire cache (see note)"
msgstr "传输整个缓存(参见备注)"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "mooncake_rpc_port"
msgstr "mooncake_rpc_port"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "0"
msgstr "0"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Automatic port assignment"
msgstr "自动端口分配"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "load_async"
msgstr "load_async"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "true"
msgstr "true"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Enable asynchronous loading"
msgstr "启用异步加载"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "register_buffer"
msgstr "register_buffer"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Required for PD-colocated mode"
msgstr "PD 共置模式必需"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:259
msgid "**Note on use_layerwise:**"
msgstr "**关于 use_layerwise 的说明:**"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:261
msgid ""
"`false`: Transfer entire KV Cache (suitable for cross-node with "
"sufficient bandwidth)"
msgstr "`false`: 传输整个KV缓存适用于跨节点且带宽充足的情况"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:263
msgid ""
"`true`: Layer-by-layer transfer (suitable for single-node memory "
"constraints)"
msgstr "`true`: 逐层传输(适用于单节点内存受限的情况)"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:266
msgid "Benchmark"
msgstr "性能基准测试"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:268
msgid ""
"We recommend using the **AISBench** tool to assess performance. The test "
"uses **Dataset A**, consisting of fully random data, with the following "
"configuration:"
msgstr "我们推荐使用 **AISBench** 工具进行性能评估。测试使用 **数据集A**,该数据集由完全随机的数据组成,配置如下:"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:272
msgid "Input/output tokens: 1024/10"
msgstr "输入/输出令牌数1024/10"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:273
msgid "Total requests: 100"
msgstr "总请求数100"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:274
msgid "Concurrency: 25"
msgstr "并发数25"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:276
msgid "The test procedure consists of three steps:"
msgstr "测试流程包含三个步骤:"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:278
msgid "Step 1: Baseline (No Cache)"
msgstr "步骤 1基准测试无缓存"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:280
msgid ""
"Send Dataset A to Instance 1 on Node 1 and record the Time to First Token"
" (TTFT) as **TTFT1**."
msgstr "将数据集A发送到节点1上的实例1并记录首令牌时间TTFT为 **TTFT1**。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:283
msgid "Preparation for Step 2"
msgstr "步骤 2 的准备工作"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:285
msgid ""
"Before Step 2, send a fully random Dataset B to Instance 1. Due to the "
"unified HBM/DRAM KV Cache with LRU (Least Recently Used) eviction policy,"
" Dataset B's cache evicts Dataset A's cache from HBM, leaving Dataset A's"
" cache only in Node 1's DRAM."
msgstr "在步骤2之前向实例1发送一个完全随机的数据集B。由于采用了具有LRU最近最少使用淘汰策略的统一HBM/DRAM KV缓存数据集B的缓存会将数据集A的缓存从HBM中淘汰使得数据集A的缓存仅保留在节点1的DRAM中。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:290
msgid "Step 2: Local DRAM Hit"
msgstr "步骤 2本地DRAM命中"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:292
msgid ""
"Send Dataset A to Instance 1 again to measure the performance when "
"hitting the KV Cache in local DRAM. Record the TTFT as **TTFT2**."
msgstr "再次将数据集A发送到实例1以测量命中本地DRAM中KV缓存时的性能。记录TTFT为 **TTFT2**。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:295
msgid "Step 3: Cross-Node DRAM Hit"
msgstr "步骤 3跨节点DRAM命中"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:297
msgid ""
"Send Dataset A to Instance 2. With the Mooncake KV Cache pool, this "
"results in a cross-node KV Cache hit from Node 1's DRAM. Record the TTFT "
"as **TTFT3**."
msgstr "将数据集A发送到实例2。借助Mooncake KV缓存池这将导致一次来自节点1 DRAM的跨节点KV缓存命中。记录TTFT为 **TTFT3**。"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:301
msgid "**Model Configuration**:"
msgstr "**模型配置**"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:329
msgid "**Performance Benchmarking Commands**:"
msgstr "**性能基准测试命令**"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md:337
msgid "Test Results"
msgstr "测试结果"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Requests"
msgstr "请求数"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "Concur"
msgstr "并发数"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "TTFT1 (ms)"
msgstr "TTFT1 (毫秒)"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "TTFT2 (ms)"
msgstr "TTFT2 (毫秒)"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "TTFT3 (ms)"
msgstr "TTFT3 (毫秒)"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "100"
msgstr "100"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "25"
msgstr "25"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "2322"
msgstr "2322"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "739"
msgstr "739"
#: ../../source/tutorials/features/pd_colocated_mooncake_multi_instance.md
msgid "948"
msgstr "948"

View File

@@ -0,0 +1,471 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:1
msgid "Prefill-Decode Disaggregation (Deepseek)"
msgstr "预填充-解码解耦部署 (Deepseek)"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:3
msgid "Getting Started"
msgstr "快速开始"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:5
msgid ""
"vLLM-Ascend now supports prefill-decode (PD) disaggregation with EP "
"(Expert Parallel) options. This guide takes one-by-one steps to verify "
"these features with constrained resources."
msgstr "vLLM-Ascend 现已支持结合专家并行EP选项的预填充-解码PD解耦部署。本指南将逐步引导您在有限资源下验证这些功能。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:7
msgid ""
"Take the Deepseek-r1-w8a8 model as an example, use 4 Atlas 800T A3 "
"servers to deploy the \"2P1D\" architecture. Assume the IP of the "
"prefiller server is 192.0.0.1 (prefill 1) and 192.0.0.2 (prefill 2), and "
"the decoder servers are 192.0.0.3 (decoder 1) and 192.0.0.4 (decoder 2). "
"On each server, use 8 NPUs 16 chips to deploy one service instance."
msgstr "以 Deepseek-r1-w8a8 模型为例,使用 4 台 Atlas 800T A3 服务器部署 \"2P1D\" 架构。假设预填充服务器 IP 为 192.0.0.1(预填充节点 1和 192.0.0.2(预填充节点 2解码服务器 IP 为 192.0.0.3(解码节点 1和 192.0.0.4(解码节点 2。每台服务器使用 8 个 NPU16 个芯片)部署一个服务实例。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:9
msgid "Verify Multi-Node Communication Environment"
msgstr "验证多节点通信环境"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:11
msgid "Physical Layer Requirements"
msgstr "物理层要求"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:13
msgid ""
"The physical machines must be located on the same WLAN, with network "
"connectivity."
msgstr "物理服务器必须位于同一局域网内,并具备网络连通性。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:14
msgid ""
"All NPUs must be interconnected. Intra-node connectivity is via HCCS, and"
" inter-node connectivity is via RDMA."
msgstr "所有 NPU 必须能够互联。节点内通过 HCCS 连接,节点间通过 RDMA 连接。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:16
msgid "Verification Process"
msgstr "验证流程"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:18
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:27
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:83
msgid ""
"Execute the following commands on each node in sequence. The results must"
" all be `success` and the status must be `UP`:"
msgstr "依次在每个节点上执行以下命令。所有结果必须为 `success` 且状态必须为 `UP`"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md
msgid "A3"
msgstr "A3"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:25
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:81
msgid "Single Node Verification:"
msgstr "单节点验证:"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:42
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:98
msgid "Check NPU HCCN Configuration:"
msgstr "检查 NPU HCCN 配置:"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:44
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:100
msgid ""
"Ensure that the hccn.conf file exists in the environment. If using "
"Docker, mount it into the container."
msgstr "确保环境中存在 hccn.conf 文件。如果使用 Docker请将其挂载到容器中。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:50
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:106
msgid "Get NPU IP Addresses"
msgstr "获取 NPU IP 地址"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:57
msgid "Get superpodid and SDID"
msgstr "获取 superpodid 和 SDID"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:63
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:112
msgid "Cross-Node PING Test"
msgstr "跨节点 PING 测试"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:70
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:119
msgid "Check NPU TLS Configuration"
msgstr "检查 NPU TLS 配置"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md
msgid "A2"
msgstr "A2"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:128
msgid "Run with Docker"
msgstr "使用 Docker 运行"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:130
msgid "Start a Docker container on each node."
msgstr "在每个节点上启动一个 Docker 容器。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:174
msgid "Install Mooncake"
msgstr "安装 Mooncake"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:176
msgid ""
"Mooncake is the serving platform for Kimi, a leading LLM service provided"
" by Moonshot AI.Installation and Compilation Guide: <https://github.com"
"/kvcache-ai/Mooncake?tab=readme-ov-file#build-and-use-binaries> First, we"
" need to obtain the Mooncake project. Refer to the following command:"
msgstr "Mooncake 是月之暗面Moonshot AI提供的领先 LLM 服务 Kimi 的推理平台。安装与编译指南:<https://github.com/kvcache-ai/Mooncake?tab=readme-ov-file#build-and-use-binaries> 首先,我们需要获取 Mooncake 项目。参考以下命令:"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:183
msgid "(Optional) Replace go install url if the network is poor"
msgstr "(可选)如果网络状况不佳,请替换 go install 的 URL"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:190
msgid "Install mpi"
msgstr "安装 mpi"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:196
msgid "Install the relevant dependencies. The installation of Go is not required."
msgstr "安装相关依赖。无需安装 Go。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:202
msgid "Compile and install"
msgstr "编译并安装"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:212
msgid "Set environment variables"
msgstr "设置环境变量"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:214
msgid "**Note:**"
msgstr "**注意:**"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:216
msgid "Adjust the Python path according to your specific Python installation"
msgstr "请根据您具体的 Python 安装路径进行调整"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:217
msgid ""
"Ensure `/usr/local/lib` and `/usr/local/lib64` are in your "
"`LD_LIBRARY_PATH`"
msgstr "确保 `/usr/local/lib` 和 `/usr/local/lib64` 在您的 `LD_LIBRARY_PATH` 中"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:223
msgid "Prefiller/Decoder Deployment"
msgstr "预填充器/解码器部署"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:225
msgid ""
"We can run the following scripts to launch a server on the "
"prefiller/decoder node, respectively. Please note that each P/D node will"
" occupy ports ranging from kv_port to kv_port + num_chips to initialize "
"socket listeners. To avoid any issues, port conflicts should be "
"prevented. Additionally, ensure that each node's engine_id is uniquely "
"assigned to avoid conflicts."
msgstr "我们可以分别运行以下脚本来在预填充器/解码器节点上启动服务器。请注意,每个 P/D 节点将占用从 kv_port 到 kv_port + num_chips 的端口范围来初始化 socket 监听器。为避免问题,应防止端口冲突。此外,请确保每个节点的 engine_id 被唯一分配,以避免冲突。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:227
msgid "kv_port Configuration Guide"
msgstr "kv_port 配置指南"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:229
msgid ""
"On Ascend NPU, Mooncake uses AscendDirectTransport for RDMA data "
"transfer, which randomly allocates ports within range `[20000, 20000 + "
"npu_per_node × 1000)`. If `kv_port` overlaps with this range, "
"intermittent port conflicts may occur. To avoid this, configure `kv_port`"
" according to the table below:"
msgstr "在 Ascend NPU 上Mooncake 使用 AscendDirectTransport 进行 RDMA 数据传输,它会在 `[20000, 20000 + npu_per_node × 1000)` 范围内随机分配端口。如果 `kv_port` 与此范围重叠,可能会发生间歇性端口冲突。为避免此问题,请根据下表配置 `kv_port`"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:132
msgid "NPUs per Node"
msgstr "每节点 NPU 数量"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:132
msgid "Reserved Port Range"
msgstr "保留端口范围"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:132
msgid "Recommended kv_port"
msgstr "推荐 kv_port"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:132
msgid "8"
msgstr "8"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:132
msgid "20000 - 27999"
msgstr "20000 - 27999"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:132
msgid ">= 28000"
msgstr ">= 28000"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:132
msgid "16"
msgstr "16"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:132
msgid "20000 - 35999"
msgstr "20000 - 35999"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:132
msgid ">= 36000"
msgstr ">= 36000"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:237
msgid ""
"If you occasionally see `zmq.error.ZMQError: Address already in use` "
"during startup, it may be caused by kv_port conflicting with randomly "
"allocated AscendDirectTransport ports. Increase your kv_port value to "
"avoid the reserved range."
msgstr "如果在启动时偶尔看到 `zmq.error.ZMQError: Address already in use`,可能是由于 kv_port 与随机分配的 AscendDirectTransport 端口冲突所致。请增加您的 kv_port 值以避开保留范围。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:240
msgid "launch_online_dp.py"
msgstr "launch_online_dp.py"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:242
msgid ""
"Use `launch_online_dp.py` to launch external dp vllm servers. "
"[launch\\_online\\_dp.py](https://github.com/vllm-project/vllm-"
"ascend/blob/main/examples/external_online_dp/launch_online_dp.py)"
msgstr "使用 `launch_online_dp.py` 启动外部解耦 vllm 服务器。[launch\\_online\\_dp.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/external_online_dp/launch_online_dp.py)"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:245
msgid "run_dp_template.sh"
msgstr "run_dp_template.sh"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:247
msgid ""
"Modify `run_dp_template.sh` on each node. "
"[run\\_dp\\_template.sh](https://github.com/vllm-project/vllm-"
"ascend/blob/main/examples/external_online_dp/run_dp_template.sh)"
msgstr "在每个节点上修改 `run_dp_template.sh`。[run\\_dp\\_template.sh](https://github.com/vllm-project/vllm-ascend/blob/main/examples/external_online_dp/run_dp_template.sh)"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:250
msgid "Layerwise"
msgstr "分层模式"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md
msgid "Prefiller node 1"
msgstr "预填充节点 1"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md
msgid "Prefiller node 2"
msgstr "预填充节点 2"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md
msgid "Decoder node 1"
msgstr "解码节点 1"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md
msgid "Decoder node 2"
msgstr "解码节点 2"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:493
msgid "Non-layerwise"
msgstr "非分层模式"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:735
msgid "Start the service"
msgstr "启动服务"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:748
msgid "Example Proxy for Deployment"
msgstr "部署示例代理"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:750
msgid ""
"Run a proxy server on the same node where your prefiller service instance"
" is deployed. You can find the proxy implementation in the repository's "
"examples directory."
msgstr "在部署了预填充器服务实例的同一节点上运行一个代理服务器。您可以在仓库的 examples 目录中找到代理实现。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:752
msgid ""
"We provide two different proxy implementations with distinct request "
"routing behaviors:"
msgstr "我们提供两种具有不同请求路由行为的代理实现:"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:754
msgid ""
"**`load_balance_proxy_layerwise_server_example.py`**: Requests are first "
"routed to the D nodes, which then forward to the P nodes as needed.This "
"proxy is designed for use with the "
"MooncakeLayerwiseConnector.[load\\_balance\\_proxy\\_layerwise\\_server\\_example.py](https://github.com"
"/vllm-project/vllm-"
"ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_layerwise_server_example.py)"
msgstr "**`load_balance_proxy_layerwise_server_example.py`**:请求首先被路由到 D 节点,然后根据需要转发到 P 节点。此代理设计用于与 MooncakeLayerwiseConnector 配合使用。[load\\_balance\\_proxy\\_layerwise\\_server\\_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_layerwise_server_example.py)"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:756
msgid ""
"**`load_balance_proxy_server_example.py`**: Requests are first routed to "
"the P nodes, which then forward to the D nodes for subsequent "
"processing.This proxy is designed for use with the "
"MooncakeConnector.[load\\_balance\\_proxy\\_server\\_example.py](https://github.com"
"/vllm-project/vllm-"
"ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
msgstr "**`load_balance_proxy_server_example.py`**:请求首先被路由到 P 节点,然后转发到 D 节点进行后续处理。此代理设计用于与 MooncakeConnector 配合使用。[load\\_balance\\_proxy\\_server\\_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "Parameter"
msgstr "参数"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "meaning"
msgstr "含义"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "--port"
msgstr "--port"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "Proxy service Port"
msgstr "代理服务端口"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "--host"
msgstr "--host"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "Proxy service Host IP"
msgstr "代理服务主机 IP"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "--prefiller-hosts"
msgstr "--prefiller-hosts"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "Hosts of prefiller nodes"
msgstr "预填充节点主机列表"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "--prefiller-ports"
msgstr "--prefiller-ports"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "Ports of prefiller nodes"
msgstr "预填充节点的端口"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "--decoder-hosts"
msgstr "--decoder-hosts"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "Hosts of decoder nodes"
msgstr "解码器节点的主机地址"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "--decoder-ports"
msgstr "--decoder-ports"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:814
msgid "Ports of decoder nodes"
msgstr "解码器节点的端口"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:877
msgid ""
"You can get the proxy program in the repository's examples, "
"[load\\_balance\\_proxy\\_server\\_example.py](https://github.com/vllm-"
"project/vllm-"
"ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
msgstr ""
"您可以在代码仓库的示例中找到代理程序,"
"[load\\_balance\\_proxy\\_server\\_example.py](https://github.com/vllm-"
"project/vllm-"
"ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:879
msgid "Benchmark"
msgstr "基准测试"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:881
msgid ""
"We recommend use aisbench tool to assess performance. "
"[aisbench](https://gitee.com/aisbench/benchmark) Execute the following "
"commands to install aisbench"
msgstr ""
"我们推荐使用 aisbench 工具进行性能评估。"
"[aisbench](https://gitee.com/aisbench/benchmark) 执行以下命令安装 aisbench"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:889
msgid ""
"You need to cancel the http proxy before assessing performance, as "
"following"
msgstr "在评估性能前,您需要取消 HTTP 代理设置,如下所示"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:897
msgid "You can place your datasets in the dir: `benchmark/ais_bench/datasets`"
msgstr "您可以将数据集放置在目录:`benchmark/ais_bench/datasets` 中"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:898
msgid ""
"You can change the configuration in the dir "
":`benchmark/ais_bench/benchmark/configs/models/vllm_api` Take the "
"``vllm_api_stream_chat.py`` for example"
msgstr ""
"您可以在目录 `benchmark/ais_bench/benchmark/configs/models/vllm_api` 中修改配置。以 "
"`vllm_api_stream_chat.py` 为例"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:924
msgid ""
"Take gsm8k dataset for example, execute the following commands to assess"
" performance."
msgstr "以 gsm8k 数据集为例,执行以下命令来评估性能。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:930
msgid ""
"For more details for commands and parameters for aisbench, refer to "
"[aisbench](https://gitee.com/aisbench/benchmark)"
msgstr "有关 aisbench 命令和参数的更多详细信息,请参考 [aisbench](https://gitee.com/aisbench/benchmark)"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:932
msgid "FAQ"
msgstr "常见问题"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:934
msgid "1. Prefiller nodes need to warmup"
msgstr "1. 预填充节点需要预热"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:936
msgid ""
"Since the computation of some NPU operators requires several rounds of "
"warm-up to achieve best performance, we recommend preheating the service "
"with some requests before conducting performance tests to achieve the "
"best end-to-end throughput."
msgstr ""
"由于部分 NPU 算子的计算需要经过多轮预热才能达到最佳性能,我们建议在进行性能测试前,先用一些请求预热服务,以获得最佳的端到端吞吐量。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:938
msgid "Verification"
msgstr "验证"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_multi_node.md:940
msgid "Check service health using the proxy server endpoint."
msgstr "使用代理服务器端点检查服务健康状况。"

View File

@@ -0,0 +1,213 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:1
msgid "Prefill-Decode Disaggregation (Qwen2.5-VL)"
msgstr "预填充-解码解耦架构 (Qwen2.5-VL)"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:3
msgid "Getting Start"
msgstr "开始使用"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:5
msgid ""
"vLLM-Ascend now supports prefill-decode (PD) disaggregation. This guide "
"takes one-by-one steps to verify these features with constrained "
"resources."
msgstr "vLLM-Ascend 现已支持预填充-解码 (PD) 解耦架构。本指南将逐步引导您在有限资源下验证这些功能。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:7
msgid ""
"Using the Qwen2.5-VL-7B-Instruct model as an example, use vllm-ascend "
"v0.11.0rc1 (with vLLM v0.11.0) on 1 Atlas 800T A2 server to deploy the "
"\"1P1D\" architecture. Assume the IP address is 192.0.0.1."
msgstr "以 Qwen2.5-VL-7B-Instruct 模型为例,在 1 台 Atlas 800T A2 服务器上使用 vllm-ascend v0.11.0rc1 (包含 vLLM v0.11.0) 部署 \"1P1D\" 架构。假设 IP 地址为 192.0.0.1。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:9
msgid "Verify Communication Environment"
msgstr "验证通信环境"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:11
msgid "Verification Process"
msgstr "验证流程"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:13
msgid "Single Node Verification:"
msgstr "单节点验证:"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:15
msgid ""
"Execute the following commands in sequence. The results must all be "
"`success` and the status must be `UP`:"
msgstr "依次执行以下命令。结果必须均为 `success` 且状态必须为 `UP`"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:30
msgid "Check NPU HCCN Configuration:"
msgstr "检查 NPU HCCN 配置:"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:32
msgid ""
"Ensure that the hccn.conf file exists in the environment. If using "
"Docker, mount it into the container."
msgstr "确保环境中存在 hccn.conf 文件。如果使用 Docker请将其挂载到容器中。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:38
msgid "Get NPU IP Addresses"
msgstr "获取 NPU IP 地址"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:44
msgid "Cross-Node PING Test"
msgstr "跨节点 PING 测试"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:51
msgid "Check NPU TLS Configuration"
msgstr "检查 NPU TLS 配置"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:58
msgid "Run with Docker"
msgstr "使用 Docker 运行"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:60
msgid "Start a Docker container."
msgstr "启动一个 Docker 容器。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:95
msgid "Install Mooncake"
msgstr "安装 Mooncake"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:97
msgid ""
"Mooncake is the serving platform for Kimi, a leading LLM service provided"
" by Moonshot AI. Installation and Compilation Guide: <https://github.com"
"/kvcache-ai/Mooncake?tab=readme-ov-file#build-and-use-binaries>. First, "
"we need to obtain the Mooncake project. Refer to the following command:"
msgstr "Mooncake 是 Kimi 的服务平台Kimi 是由 Moonshot AI 提供的领先 LLM 服务。安装与编译指南:<https://github.com/kvcache-ai/Mooncake?tab=readme-ov-file#build-and-use-binaries>。首先,我们需要获取 Mooncake 项目。参考以下命令:"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:104
msgid "(Optional) Replace go install url if the network is poor."
msgstr "(可选)如果网络状况不佳,请替换 go install 的 URL。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:111
msgid "Install mpi."
msgstr "安装 mpi。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:117
msgid "Install the relevant dependencies. The installation of Go is not required."
msgstr "安装相关依赖。无需安装 Go。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:123
msgid "Compile and install."
msgstr "编译并安装。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:133
msgid "Set environment variables."
msgstr "设置环境变量。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:135
msgid "**Note:**"
msgstr "**注意:**"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:137
msgid "Adjust the Python path according to your specific Python installation"
msgstr "根据您具体的 Python 安装情况调整 Python 路径"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:138
msgid ""
"Ensure `/usr/local/lib` and `/usr/local/lib64` are in your "
"`LD_LIBRARY_PATH`"
msgstr "确保 `/usr/local/lib` 和 `/usr/local/lib64` 在您的 `LD_LIBRARY_PATH` 中"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:144
msgid "Prefiller/Decoder Deployment"
msgstr "预填充器/解码器部署"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:146
msgid ""
"We can run the following scripts to launch a server on the "
"prefiller/decoder NPU, respectively."
msgstr "我们可以分别运行以下脚本来在预填充器/解码器 NPU 上启动服务器。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md
msgid "Prefiller"
msgstr "预填充器"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md
msgid "Decoder"
msgstr "解码器"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:236
msgid ""
"If you want to run \"2P1D\", please set ASCEND_RT_VISIBLE_DEVICES and "
"port to different values for each P process."
msgstr "如果您想运行 \"2P1D\",请为每个 P 进程将 ASCEND_RT_VISIBLE_DEVICES 和 port 设置为不同的值。"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:238
msgid "Example Proxy for Deployment"
msgstr "部署示例代理"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:240
msgid ""
"Run a proxy server on the same node with the prefiller service instance. "
"You can get the proxy program in the repository's examples: "
"[load\\_balance\\_proxy\\_server\\_example.py](https://github.com/vllm-"
"project/vllm-"
"ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
msgstr "在与预填充器服务实例相同的节点上运行一个代理服务器。您可以在仓库的示例中找到该代理程序:[load\\_balance\\_proxy\\_server\\_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:193
msgid "Parameter"
msgstr "参数"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:193
msgid "Meaning"
msgstr "含义"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:193
msgid "--port"
msgstr "--port"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:193
msgid "Port of proxy"
msgstr "代理端口"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:193
msgid "--prefiller-port"
msgstr "--prefiller-port"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:193
msgid "All ports of prefill"
msgstr "所有预填充端口"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:193
msgid "--decoder-ports"
msgstr "--decoder-ports"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:193
msgid "All ports of decoder"
msgstr "所有解码器端口"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:258
msgid "Verification"
msgstr "验证"
#: ../../source/tutorials/features/pd_disaggregation_mooncake_single_node.md:260
msgid "Check service health using the proxy server endpoint."
msgstr "使用代理服务器端点检查服务健康状态。"

View File

@@ -0,0 +1,219 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/features/ray.md:1
msgid "Ray Distributed (Qwen3-235B-A22B)"
msgstr "Ray 分布式部署 (Qwen3-235B-A22B)"
#: ../../source/tutorials/features/ray.md:3
msgid ""
"Multi-node inference is suitable for scenarios where the model cannot be "
"deployed on a single machine. In such cases, the model can be distributed"
" using tensor parallelism or pipeline parallelism. The specific "
"parallelism strategies will be covered in the following sections. To "
"successfully deploy multi-node inference, the following three steps need "
"to be completed:"
msgstr ""
"多节点推理适用于模型无法在单机上部署的场景。在这种情况下,可以使用张量并行或流水线并行来分布模型。具体的并行策略将在后续章节中介绍。要成功部署多节点推理,需要完成以下三个步骤:"
#: ../../source/tutorials/features/ray.md:5
msgid "**Verify Multi-Node Communication Environment**"
msgstr "**验证多节点通信环境**"
#: ../../source/tutorials/features/ray.md:6
msgid "**Set Up and Start the Ray Cluster**"
msgstr "**设置并启动 Ray 集群**"
#: ../../source/tutorials/features/ray.md:7
msgid "**Start the Online Inference Service on Multi-node**"
msgstr "**在多节点上启动在线推理服务**"
#: ../../source/tutorials/features/ray.md:9
msgid "Verify Multi-Node Communication Environment"
msgstr "验证多节点通信环境"
#: ../../source/tutorials/features/ray.md:11
msgid "Physical Layer Requirements"
msgstr "物理层要求"
#: ../../source/tutorials/features/ray.md:13
msgid ""
"The physical machines must be located on the same LAN, with network "
"connectivity."
msgstr "物理机必须位于同一局域网内,并具备网络连通性。"
#: ../../source/tutorials/features/ray.md:14
msgid ""
"All NPUs are connected with optical modules, and the connection status "
"must be normal."
msgstr "所有 NPU 均通过光模块连接,且连接状态必须正常。"
#: ../../source/tutorials/features/ray.md:16
msgid "Verification Process"
msgstr "验证流程"
#: ../../source/tutorials/features/ray.md:18
msgid ""
"Execute the following commands on each node in sequence. The results must"
" all be `success` and the status must be `UP`:"
msgstr "依次在每个节点上执行以下命令。结果必须均为 `success`,状态必须为 `UP`"
#: ../../source/tutorials/features/ray.md:35
msgid "NPU Interconnect Verification"
msgstr "NPU 互联验证"
#: ../../source/tutorials/features/ray.md:37
msgid "1. Get NPU IP Addresses"
msgstr "1. 获取 NPU IP 地址"
#: ../../source/tutorials/features/ray.md:43
msgid "2. Cross-Node PING Test"
msgstr "2. 跨节点 PING 测试"
#: ../../source/tutorials/features/ray.md:50
msgid "Set Up and Start the Ray Cluster"
msgstr "设置并启动 Ray 集群"
#: ../../source/tutorials/features/ray.md:52
msgid "Setting Up the Basic Container"
msgstr "设置基础容器"
#: ../../source/tutorials/features/ray.md:54
msgid ""
"To ensure a consistent execution environment across all nodes, including "
"the model path and Python environment, it is advised to use Docker "
"images."
msgstr "为确保所有节点(包括模型路径和 Python 环境)的执行环境一致,建议使用 Docker 镜像。"
#: ../../source/tutorials/features/ray.md:56
msgid ""
"For setting up a multi-node inference cluster with Ray, **containerized "
"deployment** is the preferred approach. Containers should be started on "
"both the primary and secondary nodes, with the `--net=host` option to "
"enable proper network connectivity."
msgstr "对于使用 Ray 设置多节点推理集群,**容器化部署**是首选方法。应在主节点和从节点上都启动容器,并使用 `--net=host` 选项以确保正确的网络连接。"
#: ../../source/tutorials/features/ray.md:58
msgid ""
"Below is the example container setup command, which should be executed on"
" **all nodes** :"
msgstr "以下是容器设置命令示例,应在 **所有节点** 上执行:"
#: ../../source/tutorials/features/ray.md:94
msgid "Start Ray Cluster"
msgstr "启动 Ray 集群"
#: ../../source/tutorials/features/ray.md:96
msgid ""
"After setting up the containers and installing vllm-ascend on each node, "
"follow the steps below to start the Ray cluster and execute inference "
"tasks."
msgstr "在每个节点上设置好容器并安装 vllm-ascend 后,按照以下步骤启动 Ray 集群并执行推理任务。"
#: ../../source/tutorials/features/ray.md:98
msgid ""
"Choose one machine as the primary node and the others as secondary nodes."
" Before proceeding, use `ip addr` to check your `nic_name` (network "
"interface name)."
msgstr "选择一台机器作为主节点,其他作为从节点。在继续之前,使用 `ip addr` 检查您的 `nic_name`(网络接口名称)。"
#: ../../source/tutorials/features/ray.md:100
msgid ""
"Set the `ASCEND_RT_VISIBLE_DEVICES` environment variable to specify the "
"NPU devices to use. For Ray versions above 2.1, also set the "
"`RAY_EXPERIMENTAL_NOSET_ASCEND_RT_VISIBLE_DEVICES` variable to avoid "
"device recognition issues."
msgstr "设置 `ASCEND_RT_VISIBLE_DEVICES` 环境变量以指定要使用的 NPU 设备。对于 Ray 2.1 以上版本,还需设置 `RAY_EXPERIMENTAL_NOSET_ASCEND_RT_VISIBLE_DEVICES` 变量以避免设备识别问题。"
#: ../../source/tutorials/features/ray.md:102
msgid "Below are the commands for the primary and secondary nodes:"
msgstr "以下是主节点和从节点的命令:"
#: ../../source/tutorials/features/ray.md:104
msgid "**Primary node**:"
msgstr "**主节点**"
#: ../../source/tutorials/features/ray.md:107
#: ../../source/tutorials/features/ray.md:124
msgid ""
"When starting a Ray cluster for multi-node inference, the environment "
"variables on each node must be set **before** starting the Ray cluster "
"for them to take effect. Updating the environment variables requires "
"restarting the Ray cluster."
msgstr "在为多节点推理启动 Ray 集群时,必须在启动 Ray 集群 **之前** 设置每个节点上的环境变量,它们才会生效。更新环境变量需要重启 Ray 集群。"
#: ../../source/tutorials/features/ray.md:121
msgid "**Secondary node**:"
msgstr "**从节点**"
#: ../../source/tutorials/features/ray.md:137
msgid ""
"Once the cluster is started on multiple nodes, execute `ray status` and "
"`ray list nodes` to verify the Ray cluster's status. You should see the "
"correct number of nodes and NPUs listed."
msgstr "在多个节点上启动集群后,执行 `ray status` 和 `ray list nodes` 以验证 Ray 集群的状态。您应该看到列出的正确节点数和 NPU 数。"
#: ../../source/tutorials/features/ray.md:139
msgid ""
"After Ray is successfully started, the following content will appear: A "
"local Ray instance has started successfully. Dashboard URL: The access "
"address for the Ray Dashboard (default: <http://localhost:8265>); Node "
"status (CPU/memory resources, number of healthy nodes); Cluster "
"connection address (used for adding multiple nodes)."
msgstr "Ray 成功启动后,将出现以下内容:本地 Ray 实例已成功启动。仪表板 URLRay 仪表板的访问地址(默认:<http://localhost:8265>节点状态CPU/内存资源、健康节点数);集群连接地址(用于添加多个节点)。"
#: ../../source/tutorials/features/ray.md:143
msgid "Start the Online Inference Service on Multi-node scenario"
msgstr "在多节点场景下启动在线推理服务"
#: ../../source/tutorials/features/ray.md:145
msgid ""
"In the container, you can use vLLM as if all NPUs were on a single node. "
"vLLM will utilize NPU resources across all nodes in the Ray cluster."
msgstr "在容器中,您可以像所有 NPU 都在单个节点上一样使用 vLLM。vLLM 将利用 Ray 集群中所有节点的 NPU 资源。"
#: ../../source/tutorials/features/ray.md:147
msgid "**You only need to run the vllm command on one node.**"
msgstr "**您只需在一个节点上运行 vllm 命令。**"
#: ../../source/tutorials/features/ray.md:149
msgid ""
"To set up parallelism, the common practice is to set the `tensor-"
"parallel-size` to the number of NPUs per node, and the `pipeline-"
"parallel-size` to the number of nodes."
msgstr "要设置并行,通常的做法是将 `tensor-parallel-size` 设置为每个节点的 NPU 数量,将 `pipeline-parallel-size` 设置为节点数量。"
#: ../../source/tutorials/features/ray.md:151
msgid ""
"For example, with 16 NPUs across 2 nodes (8 NPUs per node), set the "
"tensor parallel size to 8 and the pipeline parallel size to 2:"
msgstr "例如,对于分布在 2 个节点上的 16 个 NPU每个节点 8 个 NPU将张量并行大小设置为 8流水线并行大小设置为 2"
#: ../../source/tutorials/features/ray.md:167
msgid ""
"Alternatively, if you want to use only tensor parallelism, set the tensor"
" parallel size to the total number of NPUs in the cluster. For example, "
"with 16 NPUs across 2 nodes, set the tensor parallel size to 16:"
msgstr "或者,如果您只想使用张量并行,请将张量并行大小设置为集群中 NPU 的总数。例如,对于分布在 2 个节点上的 16 个 NPU将张量并行大小设置为 16"
#: ../../source/tutorials/features/ray.md:182
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "服务器启动后,您可以使用输入提示词查询模型:"

View File

@@ -0,0 +1,854 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:1
msgid "Suffix Speculative Decoding"
msgstr "后缀推测解码"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:3
msgid "**Introduction**"
msgstr "**简介**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:5
msgid ""
"Suffix Decoding is an optimization technique for speculative decoding "
"based on pattern matching. It simultaneously retrieves repetitive "
"sequences from both the prompt and the generated content, using frequency"
" statistics to predict the most likely token continuations. Unlike "
"traditional speculative decoding methods, Suffix Decoding runs entirely "
"on the CPU, eliminating the need for additional GPU resources or draft "
"models, which results in superior acceleration for repetitive tasks such "
"as AI agents and code generation."
msgstr ""
"后缀解码是一种基于模式匹配的推测解码优化技术。它同时从提示词和已生成内容中检索重复序列利用频率统计来预测最可能的后续标记。与传统的推测解码方法不同后缀解码完全在CPU上运行无需额外的GPU资源或草稿模型从而在AI智能体和代码生成等重复性任务上实现卓越的加速效果。"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:7
msgid ""
"This document provides step-by-step guidance on how to deploy and "
"benchmark the Suffix Decoding speculative inference technology supported "
"by `vllm-ascend` on Atlas A2 hardware. The setup utilizes a single Atlas "
"800T A2 node with a 4-card deployment of the Qwen3-32B model instance. "
"Benchmarking is conducted using authentic open-source datasets covering "
"the following categories:"
msgstr ""
"本文档提供了在Atlas A2硬件上部署和基准测试`vllm-ascend`支持的后缀解码推测推理技术的分步指南。该设置使用单个Atlas 800T A2节点部署了4卡的Qwen3-32B模型实例。基准测试使用涵盖以下类别的真实开源数据集进行"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**Dataset Category**"
msgstr "**数据集类别**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**Dataset Name**"
msgstr "**数据集名称**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Code Generation"
msgstr "代码生成"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "HumanEval"
msgstr "HumanEval"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Common Sense Reasoning"
msgstr "常识推理"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "ARC"
msgstr "ARC"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Mathematical Reasoning"
msgstr "数学推理"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "gsm8k"
msgstr "gsm8k"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Natural Language Understanding"
msgstr "自然语言理解"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "SuperGLUE_BoolQ"
msgstr "SuperGLUE_BoolQ"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Comprehensive Examination"
msgstr "综合评测"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "AGIEval"
msgstr "AGIEval"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Multi-turn Dialogue"
msgstr "多轮对话"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "ShareGPT"
msgstr "ShareGPT"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:18
#, python-format
msgid ""
"The benchmarking tool used in this tutorial is AISBench, which supports "
"performance testing for all the datasets listed above. The final section "
"of this tutorial presents a performance comparison between enabling and "
"disabling Suffix Decoding under the condition of satisfying an SLO TPOT <"
" 50ms across different datasets and concurrency levels. Validations "
"demonstrate that the Qwen3-32B model achieves a throughput improvement of"
" approximately 20% to 80% on various real-world datasets when Suffix "
"Decoding is enabled."
msgstr ""
"本教程使用的基准测试工具是AISBench它支持对上述所有数据集进行性能测试。本教程最后一节展示了在不同数据集和并发级别下满足SLO TPOT < 50ms条件时启用与禁用后缀解码的性能对比。验证表明启用后缀解码后Qwen3-32B模型在各种真实数据集上实现了约20%至80%的吞吐量提升。"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:20
msgid "**Download vllm-ascend Image**"
msgstr "**下载 vllm-ascend 镜像**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:22
msgid ""
"This tutorial uses the official image, version v0.13.0rc1. Use the "
"following command to download:"
msgstr "本教程使用官方镜像版本为v0.13.0rc1。使用以下命令下载:"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:28
msgid "**Run with Docker**"
msgstr "**使用 Docker 运行**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:30
msgid "Container startup command:"
msgstr "容器启动命令:"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:64
msgid "**Install arctic-inference**"
msgstr "**安装 arctic-inference**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:66
msgid ""
"Before enabling Suffix Decoding speculative inference on Ascend, the "
"Arctic Inference plugin must be installed. Arctic Inference is an open-"
"source plugin launched by Snowflake specifically to optimize LLM "
"inference speed. For detailed technical principles, please refer to the "
"following article: [Fastest Speculative Decoding in vLLM with Arctic "
"Inference and Arctic Training](https://www.snowflake.com/en/engineering-"
"blog/fast-speculative-decoding-vllm-arctic/). Install it within the "
"container using the following command:"
msgstr ""
"在Ascend上启用后缀解码推测推理之前必须安装Arctic Inference插件。Arctic Inference是Snowflake推出的一个开源插件专门用于优化LLM推理速度。详细技术原理请参考以下文章[Fastest Speculative Decoding in vLLM with Arctic Inference and Arctic Training](https://www.snowflake.com/en/engineering-blog/fast-speculative-decoding-vllm-arctic/)。在容器内使用以下命令安装:"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:72
msgid "**vLLM Instance Deployment**"
msgstr "**vLLM 实例部署**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:74
msgid ""
"Use the following command to start the container service instance. "
"Speculative inference is enabled via the `--speculative-config` "
"parameter, where `method` is set to `suffix`. For this test, "
"`num_speculative_tokens` is uniformly set to `3`."
msgstr ""
"使用以下命令启动容器服务实例。通过`--speculative-config`参数启用推测推理,其中`method`设置为`suffix`。本次测试中,`num_speculative_tokens`统一设置为`3`。"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:99
msgid "**AISbench Benchmark Testing**"
msgstr "**AISbench 基准测试**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:101
msgid ""
"Performance for all open-source datasets is tested using AISbench. For "
"specific instructions, refer to [Using AISBench for performance "
"evaluation](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/evaluation/using_ais_bench.html"
"#execute-performance-evaluation)."
msgstr ""
"所有开源数据集的性能均使用AISbench进行测试。具体操作说明请参考[使用AISBench进行性能评估](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/evaluation/using_ais_bench.html#execute-performance-evaluation)。"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:103
msgid "**Model Configuration**:"
msgstr "**模型配置**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:132
msgid "**Performance Benchmarking Commands**:"
msgstr "**性能基准测试命令**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:141
msgid "**Test Results**"
msgstr "**测试结果**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:143
msgid ""
"Below are the detailed test results of the six open-source datasets in "
"this evaluation. Compared to the baseline performance, the improvement in"
" TPOT and throughput performance at different concurrency levels after "
"enabling Suffix Decoding varies across datasets. The extent of "
"improvement after enabling Suffix Decoding differs among the datasets. "
"Below is a summary of the results:"
msgstr ""
"以下是本次评估中六个开源数据集的详细测试结果。与基线性能相比启用后缀解码后不同并发级别下的TPOT和吞吐量性能提升程度因数据集而异。启用后缀解码后的提升幅度在不同数据集间存在差异。以下是结果总结"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**Typical Representative**"
msgstr "**典型代表**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**Throughput Improvement (BS=1-10)**"
msgstr "**吞吐量提升 (BS=1-10)**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**SLO TPOT**"
msgstr "**SLO TPOT**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**High Gain**"
msgstr "**高增益**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "AGIEval, GSM8K"
msgstr "AGIEval, GSM8K"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**> 50%**"
msgstr "**> 50%**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "< 50ms"
msgstr "< 50ms"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**Medium-Low Gain**"
msgstr "**中低增益**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "ARC, ShareGPT"
msgstr "ARC, ShareGPT"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**20% ~ 30%**"
msgstr "**20% ~ 30%**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md:150
msgid "Below is the raw detailed test results:"
msgstr "以下是原始详细测试结果:"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Concurrency"
msgstr "并发数"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Avg Input"
msgstr "平均输入长度"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Avg Output"
msgstr "平均输出长度"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Requests"
msgstr "请求数"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Base TPOT(ms)"
msgstr "基线 TPOT(ms)"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Base Throughput(TPS)"
msgstr "基线吞吐量(TPS)"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Suffix TPOT(ms)"
msgstr "后缀解码 TPOT(ms)"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Suffix Throughput(TPS)"
msgstr "后缀解码吞吐量(TPS)"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "Accept Rate"
msgstr "接受率"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "TPOT Gain"
msgstr "TPOT 增益"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "TPS Gain"
msgstr "TPS 增益"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**Humaneval**"
msgstr "**Humaneval**"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "1"
msgstr "1"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "150"
msgstr "150"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "2700"
msgstr "2700"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "100"
msgstr "100"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "55.1"
msgstr "55.1"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "18.1"
msgstr "18.1"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "37.9"
msgstr "37.9"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "26.3"
msgstr "26.3"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "27.0%"
msgstr "27.0%"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "45.2%"
msgstr "45.2%"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "45.1%"
msgstr "45.1%"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "15"
msgstr "15"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "61.6"
msgstr "61.6"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "233.8"
msgstr "233.8"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "45.8"
msgstr "45.8"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "318.2"
msgstr "318.2"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "34.6%"
msgstr "34.6%"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "36.1%"
msgstr "36.1%"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "26"
msgstr "26"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "64.7"
msgstr "64.7"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "403.8"
msgstr "403.8"
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "50.9"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "519.2"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "27.2%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "28.6%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**ARC**"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "76"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "960"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "52.8"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "18.9"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "39.5"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "25.4"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "23.9%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "33.7%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "8"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "59.1"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "125.4"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "47.0"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "163.1"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "25.7%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "30.0%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "59.8"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "245.8"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "48.9"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "311.7"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "22.3%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "26.8%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**GSM8K**"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "67"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "1570"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "55.5"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "18.0"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "35.7"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "28.5"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "31.1%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "55.6%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "58.4%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "17"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "61.5"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "279.8"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "45.4"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "403.0"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "35.6%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "44.0%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "63.9"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "396.4"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "50.0"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "527.6"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "27.8%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "33.1%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**ShareGPT**"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "666"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "231"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "327"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "54.1"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "18.3"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "39.2"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "24.1"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "37.9%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "31.5%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "58.8"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "125.0"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "46.2"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "153.2"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "27.1%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "22.5%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "14"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "61.8"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "227.0"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "49.9"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "273.9"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "23.8%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "20.7%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**SuperGLUE_BoolQ**"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "207"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "314"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "18.4"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "36.1"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "26.8"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "33.4%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "49.8%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "45.6%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "16"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "60.0"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "229.7"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "43.5"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "303.9"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "38.0%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "32.3%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "32"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "62.7"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "47.8"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "507.5"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "31.3%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "28.0%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "**AGIEval**"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "735"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "1880"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "53.1"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "18.7"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "31.8"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "34.1"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "50.3%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "66.8%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "81.9%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "24"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "64.0"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "381.2"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "43.3"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "629.0"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "47.8%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "65.0%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "34"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "70.0"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "494.6"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "50.2"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "768.4"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "39.4%"
msgstr ""
#: ../../source/tutorials/features/suffix_speculative_decoding.md
msgid "55.3%"
msgstr ""

View File

@@ -0,0 +1,142 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/hardwares/310p.md:1
msgid "Atlas 300I"
msgstr "Atlas 300I"
#: ../../source/tutorials/hardwares/310p.md:4
msgid ""
"This Atlas 300I series is currently experimental. In future versions, "
"there may be behavioral changes related to model coverage and performance"
" improvement."
msgstr ""
"Atlas 300I 系列目前处于实验阶段。在未来的版本中,可能会发生与模型覆盖范围和性能改进相关的行为变更。"
#: ../../source/tutorials/hardwares/310p.md:5
msgid ""
"Currently, the Atlas 300I series only supports eager mode and the float16"
" data type."
msgstr "目前Atlas 300I 系列仅支持 eager 模式和 float16 数据类型。"
#: ../../source/tutorials/hardwares/310p.md:8
msgid "Run vLLM on Atlas 300I Series"
msgstr "在 Atlas 300I 系列上运行 vLLM"
#: ../../source/tutorials/hardwares/310p.md:10
msgid "Run docker container:"
msgstr "运行 docker 容器:"
#: ../../source/tutorials/hardwares/310p.md:40
msgid "Set up environment variables:"
msgstr "设置环境变量:"
#: ../../source/tutorials/hardwares/310p.md:50
msgid "Online Inference on NPU"
msgstr "在 NPU 上进行在线推理"
#: ../../source/tutorials/hardwares/310p.md:53
msgid ""
"For Atlas 300I (310P), do not rely on `max-model-len` auto detection "
"(omit `--max-model-len`), because it may cause OOM."
msgstr "对于 Atlas 300I (310P),不要依赖 `max-model-len` 的自动检测(省略 `--max-model-len`),因为这可能导致 OOM。"
#: ../../source/tutorials/hardwares/310p.md:56
msgid "Reason (current 310P attention path):"
msgstr "原因(当前 310P 注意力路径):"
#: ../../source/tutorials/hardwares/310p.md:57
msgid ""
"`AscendAttentionMetadataBuilder310` passes `model_config.max_model_len` "
"to `AttentionMaskBuilder310`."
msgstr "`AscendAttentionMetadataBuilder310` 将 `model_config.max_model_len` 传递给 `AttentionMaskBuilder310`。"
#: ../../source/tutorials/hardwares/310p.md:59
msgid ""
"`AttentionMaskBuilder310` builds a full causal mask with shape "
"`[max_model_len, max_model_len]` in float16, then casts it to FRACTAL_NZ."
msgstr "`AttentionMaskBuilder310` 构建一个形状为 `[max_model_len, max_model_len]` 的完整因果掩码float16 类型),然后将其转换为 FRACTAL_NZ 格式。"
#: ../../source/tutorials/hardwares/310p.md:61
msgid ""
"In 310P `attention_v1` prefill/chunked-prefill (`_npu_flash_attention` / "
"`_npu_paged_attention_splitfuse`), this explicit mask tensor is consumed "
"directly, and there is no compressed-mask path."
msgstr "在 310P 的 `attention_v1` prefill/chunked-prefill (`_npu_flash_attention` / `_npu_paged_attention_splitfuse`) 中,这个显式的掩码张量被直接使用,不存在压缩掩码路径。"
#: ../../source/tutorials/hardwares/310p.md:66
msgid ""
"So if auto resolves to a large context length, the mask allocation "
"(`O(max_model_len^2)`) can exceed NPU memory and trigger OOM. Always set "
"a conservative explicit value, for example `--max-model-len 4096`."
msgstr "因此,如果自动解析到一个很大的上下文长度,掩码分配(`O(max_model_len^2)`)可能会超出 NPU 内存并触发 OOM。请始终设置一个保守的显式值例如 `--max-model-len 4096`。"
#: ../../source/tutorials/hardwares/310p.md:71
msgid ""
"Run the following script to start the vLLM server on NPU (Qwen3-0.6B:1 "
"card, Qwen2.5-7B-Instruct:2 cards, Pangu-Pro-MoE-72B: 8 cards):"
msgstr "运行以下脚本在 NPU 上启动 vLLM 服务器Qwen3-0.6B1卡Qwen2.5-7B-Instruct2卡Pangu-Pro-MoE-72B8卡"
#: ../../source/tutorials/hardwares/310p.md
msgid "Qwen3-0.6B"
msgstr "Qwen3-0.6B"
#: ../../source/tutorials/hardwares/310p.md:81
#: ../../source/tutorials/hardwares/310p.md:111
#: ../../source/tutorials/hardwares/310p.md:141
msgid "Run the following command to start the vLLM server:"
msgstr "运行以下命令启动 vLLM 服务器:"
#: ../../source/tutorials/hardwares/310p.md:92
#: ../../source/tutorials/hardwares/310p.md:122
#: ../../source/tutorials/hardwares/310p.md:152
msgid "Once your server is started, you can query the model with input prompts."
msgstr "服务器启动后,您可以使用输入提示词查询模型。"
#: ../../source/tutorials/hardwares/310p.md
msgid "Qwen2.5-7B-Instruct"
msgstr "Qwen2.5-7B-Instruct"
#: ../../source/tutorials/hardwares/310p.md
msgid "Qwen2.5-VL-3B-Instruct"
msgstr "Qwen2.5-VL-3B-Instruct"
#: ../../source/tutorials/hardwares/310p.md:168
msgid "If you run this script successfully, you can see the results."
msgstr "如果此脚本运行成功,您将看到结果。"
#: ../../source/tutorials/hardwares/310p.md:170
msgid "Offline Inference"
msgstr "离线推理"
#: ../../source/tutorials/hardwares/310p.md:172
msgid ""
"Run the following script (`example.py`) to execute offline inference on "
"NPU:"
msgstr "运行以下脚本 (`example.py`) 在 NPU 上执行离线推理:"
#: ../../source/tutorials/hardwares/310p.md:307
msgid "Run script:"
msgstr "运行脚本:"
#: ../../source/tutorials/hardwares/310p.md:313
msgid "If you run this script successfully, you can see the info shown below:"
msgstr "如果此脚本运行成功,您将看到如下信息:"

View File

@@ -0,0 +1,29 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/hardwares/index.md:1
#: ../../source/tutorials/hardwares/index.md:5
msgid "Hardware Tutorials"
msgstr "硬件教程"
#: ../../source/tutorials/hardwares/index.md:3
msgid "This section provides tutorials on different hardware of vLLM Ascend."
msgstr "本节提供关于 vLLM Ascend 不同硬件的教程。"

View File

@@ -0,0 +1,364 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/DeepSeek-R1.md:1
msgid "DeepSeek-R1"
msgstr "DeepSeek-R1"
#: ../../source/tutorials/models/DeepSeek-R1.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/DeepSeek-R1.md:5
msgid ""
"DeepSeek-R1 is a high-performance Mixture-of-Experts (MoE) large language"
" model developed by DeepSeek Company. It excels in complex logical "
"reasoning, mathematical problem-solving, and code generation. By "
"dynamically activating its expert networks, it delivers exceptional "
"performance while maintaining computational efficiency. Building upon R1,"
" DeepSeek-R1-W8A8 is a fully quantized version of the model. It employs "
"8-bit integer (INT8) quantization for both weights and activations, which"
" significantly reduces the model's memory footprint and computational "
"requirements, enabling more efficient deployment and application in "
"resource-constrained environments. This article takes the "
"`DeepSeek-R1-W8A8` version as an example to introduce the deployment of "
"the R1 series models."
msgstr ""
"DeepSeek-R1 是由深度求索公司开发的高性能混合专家MoE大语言模型。它在复杂逻辑推理、数学问题求解和代码生成方面表现出色。通过动态激活其专家网络它在保持计算效率的同时提供了卓越的性能。基于 R1DeepSeek-R1-W8A8 是该模型的完全量化版本。它对权重和激活均采用 8 位整数INT8量化这显著减少了模型的内存占用和计算需求使其能够在资源受限的环境中更高效地部署和应用。本文以 `DeepSeek-R1-W8A8` 版本为例,介绍 R1 系列模型的部署。"
#: ../../source/tutorials/models/DeepSeek-R1.md:8
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/DeepSeek-R1.md:10
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取模型支持的功能矩阵。"
#: ../../source/tutorials/models/DeepSeek-R1.md:12
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[功能指南](../../user_guide/feature_guide/index.md)以获取功能的配置方法。"
#: ../../source/tutorials/models/DeepSeek-R1.md:14
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/DeepSeek-R1.md:16
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/DeepSeek-R1.md:18
msgid ""
"`DeepSeek-R1-W8A8`(Quantized version): require 1 Atlas 800 A3 (64G × 16) "
"nodes or 2 Atlas 800 A2 (64G × 8) nodes. [Download model "
"weight](https://www.modelscope.cn/models/vllm-ascend/DeepSeek-R1-W8A8)"
msgstr ""
"`DeepSeek-R1-W8A8`(量化版本):需要 1 个 Atlas 800 A364G × 16节点或 2 个 Atlas 800 A264G × 8节点。[下载模型权重](https://www.modelscope.cn/models/vllm-ascend/DeepSeek-R1-W8A8)"
#: ../../source/tutorials/models/DeepSeek-R1.md:20
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes."
msgstr "建议将模型权重下载到多节点的共享目录中。"
#: ../../source/tutorials/models/DeepSeek-R1.md:22
msgid "Verify Multi-node Communication(Optional)"
msgstr "验证多节点通信(可选)"
#: ../../source/tutorials/models/DeepSeek-R1.md:24
msgid ""
"If you want to deploy multi-node environment, you need to verify multi-"
"node communication according to [verify multi-node communication "
"environment](../../installation.md#verify-multi-node-communication)."
msgstr "如果您想部署多节点环境,需要根据[验证多节点通信环境](../../installation.md#verify-multi-node-communication)来验证多节点通信。"
#: ../../source/tutorials/models/DeepSeek-R1.md:26
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/DeepSeek-R1.md:28
msgid "You can use our official docker image to run `DeepSeek-R1-W8A8` directly."
msgstr "您可以使用我们的官方 docker 镜像直接运行 `DeepSeek-R1-W8A8`。"
#: ../../source/tutorials/models/DeepSeek-R1.md:30
msgid ""
"Select an image based on your machine type and start the docker image on "
"your node, refer to [using docker](../../installation.md#set-up-using-"
"docker)."
msgstr "根据您的机器类型选择一个镜像并在您的节点上启动 docker 镜像,请参考[使用 docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/DeepSeek-R1.md:69
msgid ""
"If you want to deploy multi-node environment, you need to set up "
"environment on each node."
msgstr "如果您想部署多节点环境,需要在每个节点上设置环境。"
#: ../../source/tutorials/models/DeepSeek-R1.md:71
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/DeepSeek-R1.md:73
msgid "Service-oriented Deployment"
msgstr "面向服务的部署"
#: ../../source/tutorials/models/DeepSeek-R1.md:75
msgid ""
"`DeepSeek-R1-W8A8`: require 1 Atlas 800 A3 (64G × 16) nodes or 2 Atlas "
"800 A2 (64G × 8)."
msgstr "`DeepSeek-R1-W8A8`:需要 1 个 Atlas 800 A364G × 16节点或 2 个 Atlas 800 A264G × 8节点。"
#: ../../source/tutorials/models/DeepSeek-R1.md
msgid "DeepSeek-R1-W8A8 A3 series"
msgstr "DeepSeek-R1-W8A8 A3 系列"
#: ../../source/tutorials/models/DeepSeek-R1.md:120
msgid "**Notice:** The parameters are explained as follows:"
msgstr "**注意:** 参数解释如下:"
#: ../../source/tutorials/models/DeepSeek-R1.md:123
msgid ""
"Setting the environment variable `VLLM_ASCEND_BALANCE_SCHEDULING=1` "
"enables balance scheduling. This may help increase output throughput and "
"reduce TPOT in v1 scheduler. However, TTFT may degrade in some scenarios."
" Furthermore, enabling this feature is not recommended in scenarios where"
" PD is separated."
msgstr ""
"设置环境变量 `VLLM_ASCEND_BALANCE_SCHEDULING=1` 可启用均衡调度。这可能有助于在 v1 调度器中提高输出吞吐量并降低 TPOT。然而在某些场景下 TTFT 可能会下降。此外,在 PD 分离的场景中不建议启用此功能。"
#: ../../source/tutorials/models/DeepSeek-R1.md:124
msgid ""
"For single-node deployment, we recommend using `dp4tp4` instead of "
"`dp2tp8`."
msgstr "对于单节点部署,我们建议使用 `dp4tp4` 而不是 `dp2tp8`。"
#: ../../source/tutorials/models/DeepSeek-R1.md:125
msgid ""
"`--max-model-len` specifies the maximum context length - that is, the sum"
" of input and output tokens for a single request. For performance testing"
" with an input length of 3.5K and output length of 1.5K, a value of "
"`16384` is sufficient, however, for precision testing, please set it to "
"at least `35000`."
msgstr ""
"`--max-model-len` 指定最大上下文长度——即单个请求的输入和输出令牌总数。对于输入长度为 3.5K 和输出长度为 1.5K 的性能测试,`16384` 的值已足够,但对于精度测试,请将其设置为至少 `35000`。"
#: ../../source/tutorials/models/DeepSeek-R1.md:126
msgid ""
"`--no-enable-prefix-caching` indicates that prefix caching is disabled. "
"To enable it, remove this option."
msgstr "`--no-enable-prefix-caching` 表示前缀缓存被禁用。要启用它,请移除此选项。"
#: ../../source/tutorials/models/DeepSeek-R1.md:127
msgid ""
"If you use the w4a8 weight, more memory will be allocated to kvcache, and"
" you can try to increase system throughput to achieve greater throughput."
msgstr "如果您使用 w4a8 权重,将有更多内存分配给 kvcache您可以尝试增加系统吞吐量以实现更大的吞吐量。"
#: ../../source/tutorials/models/DeepSeek-R1.md
msgid "DeepSeek-R1-W8A8 A2 series"
msgstr "DeepSeek-R1-W8A8 A2 系列"
#: ../../source/tutorials/models/DeepSeek-R1.md:132
msgid "Run the following scripts on two nodes respectively."
msgstr "分别在两个节点上运行以下脚本。"
#: ../../source/tutorials/models/DeepSeek-R1.md:134
msgid "**Node 0**"
msgstr "**节点 0**"
#: ../../source/tutorials/models/DeepSeek-R1.md:179
msgid "**Node 1**"
msgstr "**节点 1**"
#: ../../source/tutorials/models/DeepSeek-R1.md:230
msgid "Prefill-Decode Disaggregation"
msgstr "Prefill-Decode 解耦"
#: ../../source/tutorials/models/DeepSeek-R1.md:232
msgid ""
"We recommend using DeepSeek-V3.1 for deployment: "
"[DeepSeek-V3.1](./DeepSeek-V3.1.md)."
msgstr "我们推荐使用 DeepSeek-V3.1 进行部署:[DeepSeek-V3.1](./DeepSeek-V3.1.md)。"
#: ../../source/tutorials/models/DeepSeek-R1.md:234
msgid "This solution has been tested and demonstrates excellent performance."
msgstr "此解决方案已经过测试,并展现出优异的性能。"
#: ../../source/tutorials/models/DeepSeek-R1.md:236
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/DeepSeek-R1.md:238
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "一旦您的服务器启动,您就可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/DeepSeek-R1.md:251
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/DeepSeek-R1.md:253
msgid "Here are two accuracy evaluation methods."
msgstr "这里有两种精度评估方法。"
#: ../../source/tutorials/models/DeepSeek-R1.md:255
#: ../../source/tutorials/models/DeepSeek-R1.md:286
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/models/DeepSeek-R1.md:257
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参考[使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/DeepSeek-R1.md:259
msgid ""
"After execution, you can get the result, here is the result of "
"`DeepSeek-R1-W8A8` in `vllm-ascend:0.11.0rc2` for reference only."
msgstr "执行后,您可以获得结果,以下是 `DeepSeek-R1-W8A8` 在 `vllm-ascend:0.11.0rc2` 中的结果,仅供参考。"
#: ../../source/tutorials/models/DeepSeek-R1.md:130
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/models/DeepSeek-R1.md:130
msgid "version"
msgstr "版本"
#: ../../source/tutorials/models/DeepSeek-R1.md:130
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/models/DeepSeek-R1.md:130
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/models/DeepSeek-R1.md:130
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/models/DeepSeek-R1.md:130
msgid "aime2024dataset"
msgstr "aime2024dataset"
#: ../../source/tutorials/models/DeepSeek-R1.md:130
msgid "-"
msgstr "-"
#: ../../source/tutorials/models/DeepSeek-R1.md:130
msgid "accuracy"
msgstr "准确率"
#: ../../source/tutorials/models/DeepSeek-R1.md:130
msgid "gen"
msgstr "gen"
#: ../../source/tutorials/models/DeepSeek-R1.md:130
msgid "80.00"
msgstr "80.00"
#: ../../source/tutorials/models/DeepSeek-R1.md:130
msgid "gpqadataset"
msgstr "gpqadataset"
#: ../../source/tutorials/models/DeepSeek-R1.md:130
msgid "72.22"
msgstr "72.22"
#: ../../source/tutorials/models/DeepSeek-R1.md:266
msgid "Using Language Model Evaluation Harness"
msgstr "使用 Language Model Evaluation Harness"
#: ../../source/tutorials/models/DeepSeek-R1.md:268
msgid ""
"As an example, take the `gsm8k` dataset as a test dataset, and run "
"accuracy evaluation of `DeepSeek-R1-W8A8` in online mode."
msgstr "以 `gsm8k` 数据集作为测试数据集为例,在在线模式下运行 `DeepSeek-R1-W8A8` 的精度评估。"
#: ../../source/tutorials/models/DeepSeek-R1.md:270
msgid ""
"Refer to [Using "
"lm_eval](../../developer_guide/evaluation/using_lm_eval.md) for `lm_eval`"
" installation."
msgstr "`lm_eval` 的安装请参考[使用 lm_eval](../../developer_guide/evaluation/using_lm_eval.md)。"
#: ../../source/tutorials/models/DeepSeek-R1.md:272
msgid "Run `lm_eval` to execute the accuracy evaluation."
msgstr "运行 `lm_eval` 以执行精度评估。"
#: ../../source/tutorials/models/DeepSeek-R1.md:282
msgid "After execution, you can get the result."
msgstr "执行后,您可以获得结果。"
#: ../../source/tutorials/models/DeepSeek-R1.md:284
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/DeepSeek-R1.md:288
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参考[使用 AISBench 进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/DeepSeek-R1.md:290
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM Benchmark"
#: ../../source/tutorials/models/DeepSeek-R1.md:292
msgid "Run performance evaluation of `DeepSeek-R1-W8A8` as an example."
msgstr "以运行 `DeepSeek-R1-W8A8` 的性能评估为例。"
#: ../../source/tutorials/models/DeepSeek-R1.md:294
msgid ""
"Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) "
"for more details."
msgstr "更多详情请参考 [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/models/DeepSeek-R1.md:296
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 有三个子命令:"
#: ../../source/tutorials/models/DeepSeek-R1.md:298
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批请求的延迟进行基准测试。"
#: ../../source/tutorials/models/DeepSeek-R1.md:299
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/models/DeepSeek-R1.md:300
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/models/DeepSeek-R1.md:302
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。运行代码如下。"
#: ../../source/tutorials/models/DeepSeek-R1.md:309
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您就可以获得性能评估结果。"

View File

@@ -0,0 +1,608 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:1
msgid "DeepSeek-V3/3.1"
msgstr "DeepSeek-V3/3.1"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:5
msgid ""
"DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-"
"thinking mode. Compared to the previous version, this upgrade brings "
"improvements in multiple aspects:"
msgstr ""
"DeepSeek-V3.1 是一个支持思考模式和非思考模式的混合模型。与前一版本相比,此"
"次升级在多个方面带来了改进:"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:7
msgid ""
"Hybrid thinking mode: One model supports both thinking mode and non-"
"thinking mode by changing the chat template."
msgstr ""
"混合思考模式:一个模型通过更改聊天模板,同时支持思考模式和非思考模式。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:9
msgid ""
"Smarter tool calling: Through post-training optimization, the model's "
"performance in tool usage and agent tasks has significantly improved."
msgstr ""
"更智能的工具调用:通过后训练优化,模型在工具使用和智能体任务方面的性能显著提"
"升。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:11
msgid ""
"Higher thinking efficiency: DeepSeek-V3.1-Think achieves comparable "
"answer quality to DeepSeek-R1-0528, while responding more quickly."
msgstr ""
"更高的思考效率DeepSeek-V3.1-Think 实现了与 DeepSeek-R1-0528 相当的答案质"
"量,同时响应速度更快。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:13
msgid "The `DeepSeek-V3.1` model is first supported in `vllm-ascend:v0.9.1rc3`."
msgstr "`DeepSeek-V3.1` 模型首次在 `vllm-ascend:v0.9.1rc3` 中得到支持。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:15
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, single-node and multi-node deployment, accuracy and "
"performance evaluation."
msgstr ""
"本文档将展示该模型的主要验证步骤,包括支持的特性、特性配置、环境准备、单节点"
"和多节点部署、精度和性能评估。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:17
msgid "Supported Features"
msgstr "支持的特性"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:19
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr ""
"请参考 [支持的特性](../../user_guide/support_matrix/supported_models.md) "
"以获取模型支持的特性矩阵。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:21
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr ""
"请参考 [特性指南](../../user_guide/feature_guide/index.md) 以获取特性的配"
"置。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:23
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:25
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:27
msgid ""
"`DeepSeek-V3.1`(BF16 version): [Download model "
"weight](https://www.modelscope.cn/models/deepseek-ai/DeepSeek-V3.1)."
msgstr ""
"`DeepSeek-V3.1`BF16 版本):[下载模型权重](https://www.modelscope.cn/"
"models/deepseek-ai/DeepSeek-V3.1)。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:28
msgid ""
"`DeepSeek-V3.1-w8a8-mtp-QuaRot`(Quantized version with mix mtp): "
"[Download model weight](https://www.modelscope.cn/models/Eco-"
"Tech/DeepSeek-V3.1-w8a8-mtp-QuaRot)."
msgstr ""
"`DeepSeek-V3.1-w8a8-mtp-QuaRot`(混合 MTP 量化版本):[下载模型权重]"
"(https://www.modelscope.cn/models/Eco-Tech/DeepSeek-V3.1-w8a8-mtp-"
"QuaRot)。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:29
msgid ""
"`DeepSeek-V3.1-Terminus-w4a8-mtp-QuaRot`(Quantized version with mix mtp):"
" [Download model weight](https://www.modelscope.cn/models/Eco-"
"Tech/DeepSeek-V3.1-Terminus-w4a8-mtp-QuaRot)."
msgstr ""
"`DeepSeek-V3.1-Terminus-w4a8-mtp-QuaRot`(混合 MTP 量化版本):[下载模型权"
"重](https://www.modelscope.cn/models/Eco-Tech/DeepSeek-V3.1-Terminus-w4a8-"
"mtp-QuaRot)。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:30
#, python-format
msgid ""
"`Quantization method`: "
"[msmodelslim](https://gitcode.com/Ascend/msit/blob/master/msmodelslim/example/DeepSeek/README.md#deepseek-v31-w8a8-%E6%B7%B7%E5%90%88%E9%87%8F%E5%8C%96-mtp-%E9%87%8F%E5%8C%96)."
" You can use this method to quantize the model."
msgstr ""
"`量化方法`"
"[msmodelslim](https://gitcode.com/Ascend/msit/blob/master/msmodelslim/example/DeepSeek/README.md#deepseek-v31-w8a8-%E6%B7%B7%E5%90%88%E9%87%8F%E5%8C%96-mtp-%E9%87%8F%E5%8C%96)。"
" 您可以使用此方法对模型进行量化。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:32
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`."
msgstr "建议将模型权重下载到多个节点的共享目录中,例如 `/root/.cache/`。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:34
msgid "Verify Multi-node Communication(Optional)"
msgstr "验证多节点通信(可选)"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:36
msgid ""
"If you want to deploy multi-node environment, you need to verify multi-"
"node communication according to [verify multi-node communication "
"environment](../../installation.md#verify-multi-node-communication)."
msgstr ""
"如果您想部署多节点环境,需要根据 [验证多节点通信环境](../../installation."
"md#verify-multi-node-communication) 验证多节点通信。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:38
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:40
msgid "You can use our official docker image to run `DeepSeek-V3.1` directly."
msgstr "您可以使用我们的官方 docker 镜像直接运行 `DeepSeek-V3.1`。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:42
msgid ""
"Select an image based on your machine type and start the docker image on "
"your node, refer to [using docker](../../installation.md#set-up-using-"
"docker)."
msgstr ""
"根据您的机器类型选择镜像并在节点上启动 docker 镜像,请参考 [使用 docker]"
"(../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:80
msgid ""
"If you want to deploy multi-node environment, you need to set up "
"environment on each node."
msgstr "如果您想部署多节点环境,需要在每个节点上设置环境。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:82
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:84
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:86
msgid ""
"Quantized model `DeepSeek-V3.1-w8a8-mtp-QuaRot` can be deployed on 1 "
"Atlas 800 A3 (64G × 16)."
msgstr ""
"量化模型 `DeepSeek-V3.1-w8a8-mtp-QuaRot` 可以部署在 1 台 Atlas 800 A3 "
"64G × 16上。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:88
msgid "Run the following script to execute online inference."
msgstr "运行以下脚本以执行在线推理。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:131
msgid "**Notice:** The parameters are explained as follows:"
msgstr "**注意:** 参数说明如下:"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:134
msgid ""
"Setting the environment variable `VLLM_ASCEND_BALANCE_SCHEDULING=1` "
"enables balance scheduling. This may help increase output throughput and "
"reduce TPOT in v1 scheduler. However, TTFT may degrade in some scenarios."
" Furthermore, enabling this feature is not recommended in scenarios where"
" PD is separated."
msgstr ""
"设置环境变量 `VLLM_ASCEND_BALANCE_SCHEDULING=1` 启用均衡调度。这可能有助于"
"在 v1 调度器中提高输出吞吐量并降低 TPOT。然而在某些场景下 TTFT 可能会下"
"降。此外,在 PD 分离的场景中不建议启用此功能。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:135
msgid ""
"For single-node deployment, we recommend using `dp4tp4` instead of "
"`dp2tp8`."
msgstr "对于单节点部署,我们建议使用 `dp4tp4` 而不是 `dp2tp8`。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:136
msgid ""
"`--max-model-len` specifies the maximum context length - that is, the sum"
" of input and output tokens for a single request. For performance testing"
" with an input length of 3.5K and output length of 1.5K, a value of "
"`16384` is sufficient, however, for precision testing, please set it at "
"least `35000`."
msgstr ""
"`--max-model-len` 指定最大上下文长度——即单个请求的输入和输出令牌之和。对于输"
"入长度为 3.5K 和输出长度为 1.5K 的性能测试,`16384` 的值就足够了,但是,对于"
"精度测试,请至少将其设置为 `35000`。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:137
msgid ""
"`--no-enable-prefix-caching` indicates that prefix caching is disabled. "
"To enable it, remove this option."
msgstr ""
"`--no-enable-prefix-caching` 表示前缀缓存被禁用。要启用它,请移除此选项。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:138
msgid ""
"If you use the w4a8 weight, more memory will be allocated to kvcache, and"
" you can try to increase system throughput to achieve greater throughput."
msgstr ""
"如果使用 w4a8 权重,将分配更多内存给 kvcache您可以尝试增加系统吞吐量以实现"
"更大的吞吐量。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:140
msgid "Multi-node Deployment"
msgstr "多节点部署"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:142
msgid ""
"`DeepSeek-V3.1-w8a8-mtp-QuaRot`: require at least 2 Atlas 800 A2 (64G × "
"8)."
msgstr ""
"`DeepSeek-V3.1-w8a8-mtp-QuaRot`:需要至少 2 台 Atlas 800 A264G × 8。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:144
msgid "Run the following scripts on two nodes respectively."
msgstr "分别在两个节点上运行以下脚本。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:146
msgid "**Node 0**"
msgstr "**节点 0**"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:198
msgid "**Node 1**"
msgstr "**节点 1**"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:252
msgid "Prefill-Decode Disaggregation"
msgstr "Prefill-Decode 解耦"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:254
msgid ""
"We recommend using Mooncake for deployment: "
"[Mooncake](../features/pd_disaggregation_mooncake_multi_node.md)."
msgstr ""
"我们建议使用 Mooncake 进行部署:[Mooncake](../features/"
"pd_disaggregation_mooncake_multi_node.md)。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:256
msgid ""
"Take Atlas 800 A3 (64G × 16) for example, we recommend to deploy 2P1D (4 "
"nodes) rather than 1P1D (2 nodes), because there is no enough NPU memory "
"to serve high concurrency in 1P1D case."
msgstr ""
"以 Atlas 800 A364G × 16为例我们建议部署 2P1D4 个节点)而不是 1P1D"
"2 个节点),因为在 1P1D 情况下没有足够的 NPU 内存来服务高并发。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:258
msgid ""
"`DeepSeek-V3.1-w8a8-mtp-QuaRot 2P1D Layerwise` require 4 Atlas 800 A3 "
"(64G × 16)."
msgstr ""
"`DeepSeek-V3.1-w8a8-mtp-QuaRot 2P1D Layerwise` 需要 4 台 Atlas 800 A3 "
"64G × 16。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:260
msgid ""
"To run the vllm-ascend `Prefill-Decode Disaggregation` service, you need "
"to deploy a `launch_dp_program.py` script and a `run_dp_template.sh` "
"script on each node and deploy a `proxy.sh` script on prefill master node"
" to forward requests."
msgstr ""
"要运行 vllm-ascend `Prefill-Decode 解耦`服务,您需要在每个节点上部署一个 "
"`launch_dp_program.py` 脚本和一个 `run_dp_template.sh` 脚本,并在 prefill "
"主节点上部署一个 `proxy.sh` 脚本来转发请求。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:262
msgid ""
"`launch_online_dp.py` to launch external dp vllm servers. "
"[launch\\_online\\_dp.py](https://github.com/vllm-project/vllm-"
"ascend/blob/main/examples/external_online_dp/launch_online_dp.py)"
msgstr ""
"`launch_online_dp.py` 用于启动外部 dp vllm 服务器。[launch\\_online\\_dp."
"py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/"
"external_online_dp/launch_online_dp.py)"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:265
msgid "Prefill Node 0 `run_dp_template.sh` script"
msgstr "Prefill 节点 0 `run_dp_template.sh` 脚本"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:342
msgid "Prefill Node 1 `run_dp_template.sh` script"
msgstr "Prefill 节点 1 `run_dp_template.sh` 脚本"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:419
msgid "Decode Node 0 `run_dp_template.sh` script"
msgstr "Decode 节点 0 `run_dp_template.sh` 脚本"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:495
msgid "Decode Node 1 `run_dp_template.sh` script"
msgstr "Decode 节点 1 `run_dp_template.sh` 脚本"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:571
msgid "**Notice:** The parameters are explained as follows:"
msgstr "**注意:** 参数说明如下:"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:574
msgid ""
"`VLLM_ASCEND_ENABLE_FLASHCOMM1=1`: enables the communication optimization"
" function on the prefill nodes."
msgstr "`VLLM_ASCEND_ENABLE_FLASHCOMM1=1`:在 prefill 节点上启用通信优化功能。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:575
msgid ""
"`VLLM_ASCEND_ENABLE_MLAPO=1`: enables the fusion operator, which can "
"significantly improve performance but consumes more NPU memory. In the "
"Prefill-Decode (PD) separation scenario, enable MLAPO only on decode "
"nodes."
msgstr ""
"`VLLM_ASCEND_ENABLE_MLAPO=1`:启用融合算子,这可以显著提高性能但会消耗更多 "
"NPU 内存。在 Prefill-Decode (PD) 分离场景中,仅在 decode 节点上启用 MLAPO。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:576
msgid ""
"`--async-scheduling`: enables the asynchronous scheduling function. When "
"Multi-Token Prediction (MTP) is enabled, asynchronous scheduling of "
"operator delivery can be implemented to overlap the operator delivery "
"latency."
msgstr ""
"`--async-scheduling`:启用异步调度功能。当启用多令牌预测 (MTP) 时,可以实现算"
"子交付的异步调度,以重叠算子交付延迟。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:577
msgid ""
"`cudagraph_capture_sizes`: The recommended value is `n x (mtp + 1)`. And "
"the min is `n = 1` and the max is `n = max-num-seqs`. For other values, "
"it is recommended to set them to the number of frequently occurring "
"requests on the Decode (D) node."
msgstr ""
"`cudagraph_capture_sizes`:推荐值为 `n x (mtp + 1)`。最小值为 `n = 1`,最大"
"值为 `n = max-num-seqs`。对于其他值,建议将其设置为 Decode (D) 节点上频繁出"
"现的请求数量。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:578
msgid ""
"`recompute_scheduler_enable: true`: enables the recomputation scheduler. "
"When the Key-Value Cache (KV Cache) of the decode node is insufficient, "
"requests will be sent to the prefill node to recompute the KV Cache. In "
"the PD separation scenario, it is recommended to enable this "
"configuration on both prefill and decode nodes simultaneously."
msgstr ""
"`recompute_scheduler_enable: true`:启用重计算调度器。当 decode 节点的键值缓"
"存 (KV Cache) 不足时,请求将被发送到 prefill 节点以重新计算 KV Cache。在 PD "
"分离场景中,建议同时在 prefill 和 decode 节点上启用此配置。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:579
msgid ""
"`multistream_overlap_shared_expert: true`: When the Tensor Parallelism "
"(TP) size is 1 or `enable_shared_expert_dp: true`, an additional stream "
"is enabled to overlap the computation process of shared experts for "
"improved efficiency."
msgstr ""
"`multistream_overlap_shared_expert: true`:当张量并行 (TP) 大小为 1 或 "
"`enable_shared_expert_dp: true` 时,启用额外的流来重叠共享专家的计算过程,以"
"提高效率。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:580
msgid ""
"`lmhead_tensor_parallel_size: 16`: When the Tensor Parallelism (TP) size "
"of the decode node is 1, this parameter allows the TP size of the LMHead "
"embedding layer to be greater than 1, which is used to reduce the "
"computational load of each card on the LMHead embedding layer."
msgstr ""
"`lmhead_tensor_parallel_size: 16`:当 decode 节点的张量并行 (TP) 大小为 1 "
"时,此参数允许 LMHead 嵌入层的 TP 大小大于 1用于减少每张卡在 LMHead 嵌入层"
"上的计算负载。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:582
msgid "run server for each node:"
msgstr "为每个节点运行服务器:"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:595
msgid "Run the `proxy.sh` script on the prefill master node"
msgstr "在 prefill 主节点上运行 `proxy.sh` 脚本"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:597
msgid ""
"Run a proxy server on the same node with the prefiller service instance. "
"You can get the proxy program in the repository's examples: "
"[load\\_balance\\_proxy\\_server\\_example.py](https://github.com/vllm-"
"project/vllm-"
"ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
msgstr "在与预填充服务实例相同的节点上运行一个代理服务器。您可以在仓库的示例中找到代理程序:[load\\_balance\\_proxy\\_server\\_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:653
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:655
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "服务器启动后,您可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:668
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:670
msgid "Here are two accuracy evaluation methods."
msgstr "以下是两种精度评估方法。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:672
#: ../../source/tutorials/models/DeepSeek-V3.1.md:689
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:674
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参考[使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:676
msgid ""
"After execution, you can get the result, here is the result of "
"`DeepSeek-V3.1-w8a8-mtp-QuaRot` in `vllm-ascend:0.11.0rc1` for reference "
"only."
msgstr "执行后,您可以获得结果。以下是 `vllm-ascend:0.11.0rc1` 中 `DeepSeek-V3.1-w8a8-mtp-QuaRot` 的结果,仅供参考。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "version"
msgstr "版本"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "note"
msgstr "备注"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "ceval"
msgstr "ceval"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "-"
msgstr "-"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "accuracy"
msgstr "准确率"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "gen"
msgstr "生成"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "90.94"
msgstr "90.94"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "1 Atlas 800 A3 (64G × 16)"
msgstr "1 Atlas 800 A3 (64G × 16)"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "gsm8k"
msgstr "gsm8k"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:44
msgid "96.28"
msgstr "96.28"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:683
msgid "Using Language Model Evaluation Harness"
msgstr "使用 Language Model Evaluation Harness"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:685
msgid "Not test yet."
msgstr "尚未测试。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:687
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:691
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参考[使用 AISBench 进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:693
msgid "The performance result is:"
msgstr "性能结果如下:"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:695
msgid "**Hardware**: A3-752T, 4 node"
msgstr "**硬件**A3-752T4 节点"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:697
msgid "**Deployment**: 2P1D, Prefill node: DP2+TP8, Decode Node: DP32+TP1"
msgstr "**部署方式**2P1D预填充节点DP2+TP8解码节点DP32+TP1"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:699
msgid "**Input/Output**: 3.5k/1.5k"
msgstr "**输入/输出**3.5k/1.5k"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:701
msgid ""
"**Performance**: TTFT = 6.16s, TPOT = 48.82ms, Average performance of "
"each card is 478 TPS (Token Per Second)."
msgstr "**性能**TTFT = 6.16sTPOT = 48.82ms,单卡平均性能为 478 TPS每秒令牌数。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:703
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM Benchmark"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:705
msgid ""
"Run performance evaluation of `DeepSeek-V3.1-w8a8-mtp-QuaRot` as an "
"example."
msgstr "以运行 `DeepSeek-V3.1-w8a8-mtp-QuaRot` 的性能评估为例。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:707
msgid ""
"Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) "
"for more details."
msgstr "更多详情请参考 [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:709
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 有三个子命令:"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:711
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批请求的延迟进行基准测试。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:712
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:713
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:715
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。按如下方式运行代码。"
#: ../../source/tutorials/models/DeepSeek-V3.1.md:721
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您将获得性能评估结果。"

View File

@@ -0,0 +1,396 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:1
msgid "DeepSeek-V3.2"
msgstr "DeepSeek-V3.2"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:5
msgid ""
"DeepSeek-V3.2 is a sparse attention model. The main architecture is "
"similar to DeepSeek-V3.1, but with a sparse attention mechanism, which is"
" designed to explore and validate optimizations for training and "
"inference efficiency in long-context scenarios."
msgstr ""
"DeepSeek-V3.2 是一个稀疏注意力模型。其主要架构与 DeepSeek-V3.1 类似,但引入了稀疏注意力机制,旨在探索和验证长上下文场景下训练和推理效率的优化方案。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:7
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, single-node and multi-node deployment, accuracy and "
"performance evaluation."
msgstr "本文档将展示该模型的主要验证步骤,包括支持的特性、特性配置、环境准备、单节点与多节点部署、精度和性能评估。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:9
msgid "Supported Features"
msgstr "支持的特性"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:11
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的特性](../../user_guide/support_matrix/supported_models.md)以获取模型支持的特性矩阵。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:13
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[特性指南](../../user_guide/feature_guide/index.md)以获取特性的配置方法。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:15
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:17
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:19
msgid ""
"`DeepSeek-V3.2-Exp-W8A8`(Quantized version): require 1 Atlas 800 A3 (64G "
"× 16) node or 2 Atlas 800 A2 (64G × 8) nodes. [Download model "
"weight](https://www.modelscope.cn/models/vllm-ascend/DeepSeek-V3.2-Exp-"
"W8A8)"
msgstr ""
"`DeepSeek-V3.2-Exp-W8A8`(量化版本):需要 1 个 Atlas 800 A364G × 16节点或 2 个 Atlas 800 A264G × 8节点。[下载模型权重](https://www.modelscope.cn/models/vllm-ascend/DeepSeek-V3.2-Exp-W8A8)"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:20
msgid ""
"`DeepSeek-V3.2-w8a8`(Quantized version): require 1 Atlas 800 A3 (64G × "
"16) node or 2 Atlas 800 A2 (64G × 8) nodes. [Download model "
"weight](https://www.modelscope.cn/models/vllm-ascend/DeepSeek-V3.2-W8A8/)"
msgstr ""
"`DeepSeek-V3.2-w8a8`(量化版本):需要 1 个 Atlas 800 A364G × 16节点或 2 个 Atlas 800 A264G × 8节点。[下载模型权重](https://www.modelscope.cn/models/vllm-ascend/DeepSeek-V3.2-W8A8/)"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:22
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`."
msgstr "建议将模型权重下载到多节点的共享目录中,例如 `/root/.cache/`。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:24
msgid "Verify Multi-node Communication(Optional)"
msgstr "验证多节点通信(可选)"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:26
msgid ""
"If you want to deploy multi-node environment, you need to verify multi-"
"node communication according to [verify multi-node communication "
"environment](../../installation.md#verify-multi-node-communication)."
msgstr "如果您想部署多节点环境,需要根据[验证多节点通信环境](../../installation.md#verify-multi-node-communication)来验证多节点通信。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:28
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:30
msgid "You can use our official docker image to run `DeepSeek-V3.2` directly."
msgstr "您可以使用我们的官方 docker 镜像直接运行 `DeepSeek-V3.2`。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md
msgid "A3 series"
msgstr "A3 系列"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:39
#: ../../source/tutorials/models/DeepSeek-V3.2.md:82
msgid "Start the docker image on your each node."
msgstr "在您的每个节点上启动 docker 镜像。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md
msgid "A2 series"
msgstr "A2 系列"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:115
msgid ""
"In addition, if you don't want to use the docker image as above, you can "
"also build all from source:"
msgstr "此外,如果您不想使用上述 docker 镜像,也可以从源码构建所有内容:"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:117
msgid ""
"Install `vllm-ascend` from source, refer to "
"[installation](../../installation.md)."
msgstr "从源码安装 `vllm-ascend`,请参考[安装指南](../../installation.md)。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:119
msgid ""
"If you want to deploy multi-node environment, you need to set up "
"environment on each node."
msgstr "如果您想部署多节点环境,需要在每个节点上设置环境。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:121
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:124
msgid ""
"In this tutorial, we suppose you downloaded the model weight to "
"`/root/.cache/`. Feel free to change it to your own path."
msgstr "在本教程中,我们假设您已将模型权重下载到 `/root/.cache/`。您可以随意更改为自己的路径。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:127
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:129
msgid ""
"Quantized model `DeepSeek-V3.2-w8a8` can be deployed on 1 Atlas 800 A3 "
"(64G × 16)."
msgstr "量化模型 `DeepSeek-V3.2-w8a8` 可以部署在 1 个 Atlas 800 A364G × 16节点上。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:131
msgid "Run the following script to execute online inference."
msgstr "运行以下脚本以执行在线推理。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:164
msgid ""
"In PD-disaggregated deployments, `layer_sharding` is supported only on "
"prefill/P nodes with `kv_role=\"kv_producer\"`. Do not enable it on "
"decode/D nodes or `kv_role=\"kv_both\"` nodes."
msgstr "在 PD 解耦部署中,`layer_sharding` 仅支持在具有 `kv_role=\"kv_producer\"` 的 prefill/P 节点上启用。不要在 decode/D 节点或 `kv_role=\"kv_both\"` 节点上启用它。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:166
msgid "Multi-node Deployment"
msgstr "多节点部署"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:168
msgid "`DeepSeek-V3.2-w8a8`: require at least 2 Atlas 800 A2 (64G × 8)."
msgstr "`DeepSeek-V3.2-w8a8`:需要至少 2 个 Atlas 800 A264G × 8节点。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:170
msgid "Run the following scripts on two nodes respectively."
msgstr "分别在两个节点上运行以下脚本。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:179
#: ../../source/tutorials/models/DeepSeek-V3.2.md:283
msgid "**Node0**"
msgstr "**节点0**"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:228
#: ../../source/tutorials/models/DeepSeek-V3.2.md:337
msgid "**Node1**"
msgstr "**节点1**"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:395
msgid "Prefill-Decode Disaggregation"
msgstr "Prefill-Decode 解耦"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:397
msgid ""
"We'd like to show the deployment guide of `DeepSeek-V3.2` on multi-node "
"environment with 1P1D for better performance."
msgstr "我们将展示 `DeepSeek-V3.2` 在多节点环境下采用 1P1D 部署的指南,以获得更好的性能。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:399
msgid "Before you start, please"
msgstr "在开始之前,请"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:401
msgid "prepare the script `launch_online_dp.py` on each node:"
msgstr "在每个节点上准备脚本 `launch_online_dp.py`"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:504
msgid "prepare the script `run_dp_template.sh` on each node."
msgstr "在每个节点上准备脚本 `run_dp_template.sh`。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:506
#: ../../source/tutorials/models/DeepSeek-V3.2.md:809
msgid "Prefill node 0"
msgstr "Prefill 节点 0"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:580
#: ../../source/tutorials/models/DeepSeek-V3.2.md:816
msgid "Prefill node 1"
msgstr "Prefill 节点 1"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:653
#: ../../source/tutorials/models/DeepSeek-V3.2.md:823
msgid "Decode node 0"
msgstr "Decode 节点 0"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:730
#: ../../source/tutorials/models/DeepSeek-V3.2.md:830
msgid "Decode node 1"
msgstr "Decode 节点 1"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:806
msgid ""
"Once the preparation is done, you can start the server with the following"
" command on each node: Refer to [Distributed DP Server With Large-Scale "
"Expert "
"Parallelism](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/large_scale_ep.html)"
" to get the detailed boot method."
msgstr "准备工作完成后,您可以在每个节点上使用以下命令启动服务器:请参考[分布式 DP 服务器与大规模专家并行](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/large_scale_ep.html)以获取详细的启动方法。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:837
msgid "Request Forwarding"
msgstr "请求转发"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:839
msgid ""
"To set up request forwarding, run the following script on any machine. "
"You can get the proxy program in the repository's examples: "
"[load_balance_proxy_layerwise_server_example.py](https://github.com/vllm-"
"project/vllm-"
"ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_layerwise_server_example.py)"
msgstr "要设置请求转发,请在任何机器上运行以下脚本。您可以在仓库的示例中找到代理程序:[load_balance_proxy_layerwise_server_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_layerwise_server_example.py)"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:868
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:870
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "服务器启动后,您可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:883
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:885
msgid "Here are two accuracy evaluation methods."
msgstr "这里有两种精度评估方法。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:887
#: ../../source/tutorials/models/DeepSeek-V3.2.md:913
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:889
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参考[使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:891
#: ../../source/tutorials/models/DeepSeek-V3.2.md:909
msgid "After execution, you can get the result."
msgstr "执行后,您可以获得结果。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:893
msgid "Using Language Model Evaluation Harness"
msgstr "使用 Language Model Evaluation Harness"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:895
msgid ""
"As an example, take the `gsm8k` dataset as a test dataset, and run "
"accuracy evaluation of `DeepSeek-V3.2-W8A8` in online mode."
msgstr "以 `gsm8k` 数据集作为测试数据集为例,运行 `DeepSeek-V3.2-W8A8` 的在线模式精度评估。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:897
msgid ""
"Refer to [Using "
"lm_eval](../../developer_guide/evaluation/using_lm_eval.md) for `lm_eval`"
" installation."
msgstr "`lm_eval` 的安装请参考[使用 lm_eval](../../developer_guide/evaluation/using_lm_eval.md)。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:899
msgid "Run `lm_eval` to execute the accuracy evaluation."
msgstr "运行 `lm_eval` 以执行精度评估。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:911
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:915
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参考[使用 AISBench 进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:917
msgid "The performance result is:"
msgstr "性能结果如下:"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:919
msgid "**Hardware**: A3-752T, 4 node"
msgstr "**硬件**A3-752T4 节点"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:921
msgid "**Deployment**: 1P1D, Prefill node: DP2+TP16, Decode Node: DP8+TP4"
msgstr "**部署**1P1DPrefill 节点DP2+TP16Decode 节点DP8+TP4"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:923
msgid "**Input/Output**: 64k/3k"
msgstr "**输入/输出**64k/3k"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:925
msgid "**Performance**: 533tps, TPOT 32ms"
msgstr "**性能**533tpsTPOT 32ms"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:927
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM Benchmark"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:929
msgid "Run performance evaluation of `DeepSeek-V3.2-W8A8` as an example."
msgstr "以运行 `DeepSeek-V3.2-W8A8` 的性能评估为例。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:931
msgid ""
"Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) "
"for more details."
msgstr "更多详情请参考 [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:933
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 有三个子命令:"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:935
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批请求的延迟进行基准测试。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:936
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:937
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:基准离线推理吞吐量。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:939
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例,按如下方式运行代码。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:946
msgid "Function Call"
msgstr "函数调用"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:948
msgid ""
"The function call feature is supported from v0.13.0rc1 on. Please use the"
" latest version."
msgstr "函数调用功能自 v0.13.0rc1 版本起支持。请使用最新版本。"
#: ../../source/tutorials/models/DeepSeek-V3.2.md:950
msgid ""
"Refer to [DeepSeek-V3.2 Usage "
"Guide](https://docs.vllm.ai/projects/recipes/en/latest/DeepSeek/DeepSeek-V3_2.html"
"#tool-calling-example) for details."
msgstr "详情请参阅 [DeepSeek-V3.2 使用指南](https://docs.vllm.ai/projects/recipes/en/latest/DeepSeek/DeepSeek-V3_2.html#tool-calling-example)。"

View File

@@ -0,0 +1,528 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/GLM4.x.md:1
msgid "GLM-4.5/4.6/4.7"
msgstr "GLM-4.5/4.6/4.7"
#: ../../source/tutorials/models/GLM4.x.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/GLM4.x.md:5
msgid ""
"GLM-4.x series models use a Mixture-of-Experts (MoE) architecture and are"
" foundational models specifically designed for agent applications."
msgstr "GLM-4.x 系列模型采用混合专家MoE架构是专为智能体应用设计的基础模型。"
#: ../../source/tutorials/models/GLM4.x.md:7
msgid "The `GLM-4.5` model is first supported in `vllm-ascend:v0.10.0rc1`."
msgstr "`GLM-4.5` 模型首次在 `vllm-ascend:v0.10.0rc1` 版本中得到支持。"
#: ../../source/tutorials/models/GLM4.x.md:9
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, single-node and multi-node deployment, accuracy and "
"performance evaluation."
msgstr "本文档将展示该模型的主要验证步骤,包括支持的功能、功能配置、环境准备、单节点与多节点部署、精度和性能评估。"
#: ../../source/tutorials/models/GLM4.x.md:11
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/GLM4.x.md:13
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取模型支持的功能矩阵。"
#: ../../source/tutorials/models/GLM4.x.md:15
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[功能指南](../../user_guide/feature_guide/index.md)以获取功能的配置信息。"
#: ../../source/tutorials/models/GLM4.x.md:17
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/GLM4.x.md:19
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/GLM4.x.md:21
msgid ""
"`GLM-4.5`(BF16 version): [Download model "
"weight](https://www.modelscope.cn/models/ZhipuAI/GLM-4.5)."
msgstr "`GLM-4.5`BF16 版本):[下载模型权重](https://www.modelscope.cn/models/ZhipuAI/GLM-4.5)。"
#: ../../source/tutorials/models/GLM4.x.md:22
msgid ""
"`GLM-4.6`(BF16 version): [Download model "
"weight](https://www.modelscope.cn/models/ZhipuAI/GLM-4.6)."
msgstr "`GLM-4.6`BF16 版本):[下载模型权重](https://www.modelscope.cn/models/ZhipuAI/GLM-4.6)。"
#: ../../source/tutorials/models/GLM4.x.md:23
msgid ""
"`GLM-4.7`(BF16 version): [Download model "
"weight](https://www.modelscope.cn/models/ZhipuAI/GLM-4.7)."
msgstr "`GLM-4.7`BF16 版本):[下载模型权重](https://www.modelscope.cn/models/ZhipuAI/GLM-4.7)。"
#: ../../source/tutorials/models/GLM4.x.md:24
msgid ""
"`GLM-4.5-w8a8-with-float-mtp`(Quantized version with mtp): [Download "
"model weight](https://modelers.cn/models/Modelers_Park/GLM-4.5-w8a8)."
msgstr "`GLM-4.5-w8a8-with-float-mtp`(带 mtp 的量化版本):[下载模型权重](https://modelers.cn/models/Modelers_Park/GLM-4.5-w8a8)。"
#: ../../source/tutorials/models/GLM4.x.md:25
msgid ""
"`GLM-4.6-w8a8`(Quantized version without mtp): [Download model "
"weight](https://modelers.cn/models/Modelers_Park/GLM-4.6-w8a8). Because "
"vllm do not support GLM4.6 mtp in October, so we do not provide mtp "
"version. And last month, it supported, you can use the following "
"quantization scheme to add mtp weights to Quantized weights."
msgstr "`GLM-4.6-w8a8`(不带 mtp 的量化版本):[下载模型权重](https://modelers.cn/models/Modelers_Park/GLM-4.6-w8a8)。由于 vllm 在十月份不支持 GLM4.6 的 mtp因此我们不提供 mtp 版本。上个月已支持,您可以使用以下量化方案将 mtp 权重添加到量化权重中。"
#: ../../source/tutorials/models/GLM4.x.md:26
msgid ""
"`GLM-4.7-w8a8-with-float-mtp`(Quantized version without mtp): [Download "
"model weight](https://modelscope.cn/models/Eco-"
"Tech/GLM-4.7-W8A8-floatmtp)."
msgstr "`GLM-4.7-w8a8-with-float-mtp`(不带 mtp 的量化版本):[下载模型权重](https://modelscope.cn/models/Eco-Tech/GLM-4.7-W8A8-floatmtp)。"
#: ../../source/tutorials/models/GLM4.x.md:27
msgid ""
"`Method of Quantify`: [quantization "
"scheme](https://blog.csdn.net/qq_37368095/article/details/156429653?spm=1011.2124.3001.6209)."
" You can use these methods to quantify the model."
msgstr "`量化方法`[量化方案](https://blog.csdn.net/qq_37368095/article/details/156429653?spm=1011.2124.3001.6209)。您可以使用这些方法对模型进行量化。"
#: ../../source/tutorials/models/GLM4.x.md:29
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`."
msgstr "建议将模型权重下载到多个节点的共享目录中,例如 `/root/.cache/`。"
#: ../../source/tutorials/models/GLM4.x.md:31
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/GLM4.x.md:33
msgid "You can use our official docker image to run `GLM-4.x` directly."
msgstr "您可以使用我们的官方 docker 镜像直接运行 `GLM-4.x`。"
#: ../../source/tutorials/models/GLM4.x.md
msgid "A3 series"
msgstr "A3 系列"
#: ../../source/tutorials/models/GLM4.x.md:42
#: ../../source/tutorials/models/GLM4.x.md:85
msgid "Start the docker image on your each node."
msgstr "在您的每个节点上启动 docker 镜像。"
#: ../../source/tutorials/models/GLM4.x.md
msgid "A2 series"
msgstr "A2 系列"
#: ../../source/tutorials/models/GLM4.x.md:118
msgid ""
"In addition, if you don't want to use the docker image as above, you can "
"also build all from source:"
msgstr "此外,如果您不想使用上述 docker 镜像,也可以从源码构建所有内容:"
#: ../../source/tutorials/models/GLM4.x.md:120
msgid ""
"Install `vllm-ascend` from source, refer to "
"[installation](../../installation.md)."
msgstr "从源码安装 `vllm-ascend`,请参考[安装指南](../../installation.md)。"
#: ../../source/tutorials/models/GLM4.x.md:122
msgid ""
"If you want to deploy multi-node environment, you need to set up "
"environment on each node."
msgstr "如果您想部署多节点环境,需要在每个节点上设置环境。"
#: ../../source/tutorials/models/GLM4.x.md:124
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/GLM4.x.md:126
msgid "**Notice:**"
msgstr "**注意:**"
#: ../../source/tutorials/models/GLM4.x.md:128
msgid ""
"We have optimized the FIA operator in CANN 8.5.1. Manual replacement of "
"the files related to the FIA operator is required. Please execute the FIA"
" operator replacement script: "
"[A2](../../../../tools/install_flash_infer_attention_score_ops_a2.sh) and"
" [A3](../../../../tools/install_flash_infer_attention_score_ops_a3.sh) "
"The optimization of the FIA operator will be enabled by default in CANN "
"9.x releases, and manual replacement will no longer be required. Please "
"stay tuned for updates to this document."
msgstr "我们已在 CANN 8.5.1 中优化了 FIA 算子。需要手动替换与 FIA 算子相关的文件。请执行 FIA 算子替换脚本:[A2](../../../../tools/install_flash_infer_attention_score_ops_a2.sh) 和 [A3](../../../../tools/install_flash_infer_attention_score_ops_a3.sh)。FIA 算子的优化将在 CANN 9.x 版本中默认启用,届时将不再需要手动替换。请关注本文档的更新。"
#: ../../source/tutorials/models/GLM4.x.md:132
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/models/GLM4.x.md:134
msgid "In low-latency scenarios, we recommend a single-machine deployment."
msgstr "在低延迟场景下,我们推荐单机部署。"
#: ../../source/tutorials/models/GLM4.x.md:135
msgid ""
"Quantized model `glm4.7_w8a8_with_float_mtp` can be deployed on 1 Atlas "
"800 A3 (64G × 16) or 1 Atlas 800 A2 (64G × 8)."
msgstr "量化模型 `glm4.7_w8a8_with_float_mtp` 可以部署在 1 台 Atlas 800 A364G × 16或 1 台 Atlas 800 A264G × 8上。"
#: ../../source/tutorials/models/GLM4.x.md:137
msgid "Run the following script to execute online inference."
msgstr "运行以下脚本以执行在线推理。"
#: ../../source/tutorials/models/GLM4.x.md:169
msgid "**Notice:** The parameters are explained as follows:"
msgstr "**注意:** 参数解释如下:"
#: ../../source/tutorials/models/GLM4.x.md:172
msgid ""
"`--async-scheduling` Asynchronous scheduling is a technique used to "
"optimize inference efficiency. It allows non-blocking task scheduling to "
"improve concurrency and throughput, especially when processing large-"
"scale models."
msgstr "`--async-scheduling` 异步调度是一种用于优化推理效率的技术。它允许非阻塞的任务调度,以提高并发性和吞吐量,特别是在处理大规模模型时。"
#: ../../source/tutorials/models/GLM4.x.md:173
msgid ""
"`fusion_ops_gmmswigluquant` The performance of the GmmSwigluQuant fusion "
"operator tends to degrade when the total number of NPUs is ≤ 16."
msgstr "`fusion_ops_gmmswigluquant` 当 NPU 总数 ≤ 16 时GmmSwigluQuant 融合算子的性能往往会下降。"
#: ../../source/tutorials/models/GLM4.x.md:175
msgid "Multi-node Deployment"
msgstr "多节点部署"
#: ../../source/tutorials/models/GLM4.x.md:177
msgid ""
"Although the former tutorial said \"Not recommended to deploy multi-node "
"on Atlas 800 A2 (64G × 8)\", but if you insist to deploy GLM-4.x model on"
" multi-node like 2 × Atlas 800 A2 (64G × 8), run the following scripts on"
" two nodes respectively."
msgstr "尽管之前的教程提到“不建议在 Atlas 800 A264G × 8上部署多节点”但如果您坚持要在类似 2 × Atlas 800 A264G × 8的多节点上部署 GLM-4.x 模型,请分别在两个节点上运行以下脚本。"
#: ../../source/tutorials/models/GLM4.x.md:179
msgid "**Node 0**"
msgstr "**节点 0**"
#: ../../source/tutorials/models/GLM4.x.md:230
msgid "**Node 1**"
msgstr "**节点 1**"
#: ../../source/tutorials/models/GLM4.x.md:283
msgid "Prefill-Decode Disaggregation"
msgstr "Prefill-Decode 解耦部署"
#: ../../source/tutorials/models/GLM4.x.md:285
msgid ""
"We'd like to show the deployment guide of `GLM4.7` on multi-node "
"environment with 2P1D for better performance."
msgstr "我们将展示 `GLM4.7` 在多节点环境2P1D下的部署指南以获得更好的性能。"
#: ../../source/tutorials/models/GLM4.x.md:287
msgid "Before you start, please"
msgstr "在开始之前,请"
#: ../../source/tutorials/models/GLM4.x.md:289
msgid "prepare the script `launch_online_dp.py` on each node:"
msgstr "在每个节点上准备脚本 `launch_online_dp.py`"
#: ../../source/tutorials/models/GLM4.x.md:392
msgid "prepare the script `run_dp_template.sh` on each node."
msgstr "在每个节点上准备脚本 `run_dp_template.sh`。"
#: ../../source/tutorials/models/GLM4.x.md:394
#: ../../source/tutorials/models/GLM4.x.md:669
msgid "Prefill node 0"
msgstr "Prefill 节点 0"
#: ../../source/tutorials/models/GLM4.x.md:460
#: ../../source/tutorials/models/GLM4.x.md:676
msgid "Prefill node 1"
msgstr "Prefill 节点 1"
#: ../../source/tutorials/models/GLM4.x.md:525
#: ../../source/tutorials/models/GLM4.x.md:683
msgid "Decode node 0"
msgstr "Decode 节点 0"
#: ../../source/tutorials/models/GLM4.x.md:596
#: ../../source/tutorials/models/GLM4.x.md:690
msgid "Decode node 1"
msgstr "Decode 节点 1"
#: ../../source/tutorials/models/GLM4.x.md:667
msgid ""
"Once the preparation is done, you can start the server with the following"
" command on each node:"
msgstr "准备工作完成后,您可以在每个节点上使用以下命令启动服务器:"
#: ../../source/tutorials/models/GLM4.x.md:697
msgid "Request Forwarding"
msgstr "请求转发"
#: ../../source/tutorials/models/GLM4.x.md:699
msgid ""
"To set up request forwarding, run the following script on any machine. "
"You can get the proxy program in the repository's examples: "
"[load_balance_proxy_server_example.py](https://github.com/vllm-project"
"/vllm-"
"ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
msgstr "要设置请求转发,请在任何机器上运行以下脚本。您可以在仓库的示例中找到代理程序:[load_balance_proxy_server_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
#: ../../source/tutorials/models/GLM4.x.md:728
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/GLM4.x.md:730
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "服务器启动后,您可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/GLM4.x.md:749
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/GLM4.x.md:751
msgid "Here are two accuracy evaluation methods."
msgstr "这里有两种精度评估方法。"
#: ../../source/tutorials/models/GLM4.x.md:753
#: ../../source/tutorials/models/GLM4.x.md:770
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/models/GLM4.x.md:755
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参考[使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/GLM4.x.md:757
msgid ""
"After execution, you can get the result, here is the result of `GLM4.7` "
"in `vllm-ascend:main` (after `vllm-ascend:0.14.0rc1`) for reference only."
msgstr "执行后,您可以获得结果,以下是 `GLM4.7` 在 `vllm-ascend:main``vllm-ascend:0.14.0rc1` 之后)中的结果,仅供参考。"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "version"
msgstr "版本"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "note"
msgstr "备注"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "GPQA"
msgstr "GPQA"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "-"
msgstr "-"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "accuracy"
msgstr "准确率"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "gen"
msgstr "生成"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "84.85"
msgstr "84.85"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "1 Atlas 800 A3 (64G × 16)"
msgstr "1 Atlas 800 A3 (64G × 16)"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "MATH500"
msgstr "MATH500"
#: ../../source/tutorials/models/GLM4.x.md:87
msgid "98.8"
msgstr "98.8"
#: ../../source/tutorials/models/GLM4.x.md:764
msgid "Using Language Model Evaluation Harness"
msgstr "使用语言模型评估工具"
#: ../../source/tutorials/models/GLM4.x.md:766
msgid "Not tested yet."
msgstr "尚未测试。"
#: ../../source/tutorials/models/GLM4.x.md:768
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/GLM4.x.md:772
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr ""
"详情请参考[使用AISBench进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/GLM4.x.md:774
msgid "Using vLLM Benchmark"
msgstr "使用vLLM基准测试"
#: ../../source/tutorials/models/GLM4.x.md:776
msgid "Run performance evaluation of `GLM-4.x` as an example."
msgstr "以运行 `GLM-4.x` 的性能评估为例。"
#: ../../source/tutorials/models/GLM4.x.md:778
msgid ""
"Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) "
"for more details."
msgstr ""
"更多详情请参考 [vllm基准测试](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/models/GLM4.x.md:780
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 包含三个子命令:"
#: ../../source/tutorials/models/GLM4.x.md:782
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:基准测试单批次请求的延迟。"
#: ../../source/tutorials/models/GLM4.x.md:783
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:基准测试在线服务吞吐量。"
#: ../../source/tutorials/models/GLM4.x.md:784
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:基准测试离线推理吞吐量。"
#: ../../source/tutorials/models/GLM4.x.md:786
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例,运行以下代码。"
#: ../../source/tutorials/models/GLM4.x.md:808
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您将获得性能评估结果。"
#: ../../source/tutorials/models/GLM4.x.md:810
msgid "Best Practices"
msgstr "最佳实践"
#: ../../source/tutorials/models/GLM4.x.md:812
msgid "In this chapter, we recommend best practices for three scenarios:"
msgstr "本章节,我们针对三种场景推荐最佳实践:"
#: ../../source/tutorials/models/GLM4.x.md:814
msgid ""
"Long-context: For long sequences with low concurrency (≤ 4): set `dp1 "
"tp16`; For long sequences with high concurrency (> 4): set `dp2 tp8`"
msgstr ""
"长上下文:对于低并发(≤ 4的长序列设置 `dp1 tp16`;对于高并发(> 4的长序列设置 `dp2 tp8`"
#: ../../source/tutorials/models/GLM4.x.md:815
msgid ""
"Low-latency: For short sequences with low latency: we recommend setting "
"`dp2 tp8`"
msgstr "低延迟:对于需要低延迟的短序列,我们推荐设置 `dp2 tp8`"
#: ../../source/tutorials/models/GLM4.x.md:816
msgid ""
"High-throughput: For short sequences with high throughput: we also "
"recommend setting `dp2 tp8`"
msgstr "高吞吐量:对于需要高吞吐量的短序列,我们也推荐设置 `dp2 tp8`"
#: ../../source/tutorials/models/GLM4.x.md:818
msgid ""
"**Notice:** `max-model-len` and `max-num-seqs` need to be set according "
"to the actual usage scenario. For other settings, please refer to the "
"**[Deployment](#deployment)** chapter."
msgstr ""
"**注意:** `max-model-len` 和 `max-num-seqs` 需要根据实际使用场景进行设置。其他设置请参考 **[部署](#deployment)** 章节。"
#: ../../source/tutorials/models/GLM4.x.md:821
msgid "FAQ"
msgstr "常见问题"
#: ../../source/tutorials/models/GLM4.x.md:823
msgid "**Q: Why is the TPOT performance poor in Long-context test?**"
msgstr "**问为什么在长上下文测试中TPOT性能不佳**"
#: ../../source/tutorials/models/GLM4.x.md:825
msgid ""
"A: Please ensure that the FIA operator replacement script has been "
"executed successfully to complete the replacement of FIA operators. Here "
"is the script: "
"[A2](../../../../tools/install_flash_infer_attention_score_ops_a2.sh) and"
" [A3](../../../../tools/install_flash_infer_attention_score_ops_a3.sh)"
msgstr ""
"答请确保已成功执行FIA算子替换脚本以完成FIA算子的替换。脚本如下"
"[A2](../../../../tools/install_flash_infer_attention_score_ops_a2.sh) 和 "
"[A3](../../../../tools/install_flash_infer_attention_score_ops_a3.sh)"
#: ../../source/tutorials/models/GLM4.x.md:827
msgid ""
"**Q: Startup fails with HCCL port conflicts (address already bound). What"
" should I do?**"
msgstr "**问启动失败提示HCCL端口冲突地址已被占用。我该怎么办**"
#: ../../source/tutorials/models/GLM4.x.md:829
msgid "A: Clean up old processes and restart: `pkill -f VLLM*`."
msgstr "答:清理旧进程并重启:`pkill -f VLLM*`。"
#: ../../source/tutorials/models/GLM4.x.md:831
msgid "**Q: How to handle OOM or unstable startup?**"
msgstr "**问如何处理OOM或启动不稳定的问题**"
#: ../../source/tutorials/models/GLM4.x.md:833
msgid ""
"A: Reduce `--max-num-seqs` and `--max-model-len` first. If needed, reduce"
" concurrency and load-testing pressure (e.g., `max-concurrency` / `num-"
"prompts`)."
msgstr ""
"答:首先减少 `--max-num-seqs` 和 `--max-model-len`。如有需要,降低并发度和负载测试压力(例如,`max-concurrency` / `num-prompts`)。"

View File

@@ -0,0 +1,475 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/GLM5.md:1
msgid "GLM-5"
msgstr "GLM-5"
#: ../../source/tutorials/models/GLM5.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/GLM5.md:5
msgid ""
"[GLM-5](https://huggingface.co/zai-org/GLM-5) use a Mixture-of-Experts "
"(MoE) architecture and targeting at complex systems engineering and long-"
"horizon agentic tasks."
msgstr ""
"[GLM-5](https://huggingface.co/zai-org/GLM-5) 采用混合专家 (Mixture-of-Experts, MoE) 架构,旨在处理复杂系统工程和长视野智能体任务。"
#: ../../source/tutorials/models/GLM5.md:7
msgid ""
"The `GLM-5` model is first supported in `vllm-ascend:v0.17.0rc1`. In "
"`vllm-ascend:v0.17.0rc1` and `vllm-ascend:v0.18.0rc1` , the version of "
"transformers need to be upgraded to 5.2.0."
msgstr ""
"`GLM-5` 模型首次在 `vllm-ascend:v0.17.0rc1` 版本中得到支持。在 `vllm-ascend:v0.17.0rc1` 和 `vllm-ascend:v0.18.0rc1` 版本中,需要将 transformers 的版本升级到 5.2.0。"
#: ../../source/tutorials/models/GLM5.md:9
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, single-node and multi-node deployment, accuracy and "
"performance evaluation."
msgstr ""
"本文档将展示该模型的主要验证步骤,包括支持的特性、特性配置、环境准备、单节点和多节点部署、精度和性能评估。"
#: ../../source/tutorials/models/GLM5.md:11
msgid "Supported Features"
msgstr "支持的特性"
#: ../../source/tutorials/models/GLM5.md:13
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr ""
"请参考[支持的特性](../../user_guide/support_matrix/supported_models.md)以获取模型支持的特性矩阵。"
#: ../../source/tutorials/models/GLM5.md:15
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr ""
"请参考[特性指南](../../user_guide/feature_guide/index.md)以获取特性的配置方法。"
#: ../../source/tutorials/models/GLM5.md:17
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/GLM5.md:19
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/GLM5.md:21
msgid ""
"`GLM-5`(BF16 version): [Download model "
"weight](https://www.modelscope.cn/models/ZhipuAI/GLM-5)."
msgstr ""
"`GLM-5` (BF16 版本): [下载模型权重](https://www.modelscope.cn/models/ZhipuAI/GLM-5)。"
#: ../../source/tutorials/models/GLM5.md:22
msgid ""
"`GLM-5-w4a8`: [Download model weight](https://modelscope.cn/models/Eco-"
"Tech/GLM-5-w4a8)."
msgstr ""
"`GLM-5-w4a8`: [下载模型权重](https://modelscope.cn/models/Eco-Tech/GLM-5-w4a8)。"
#: ../../source/tutorials/models/GLM5.md:23
msgid ""
"`GLM-5-w8a8`: [Download model weight](https://www.modelscope.cn/models"
"/Eco-Tech/GLM-5-w8a8)."
msgstr ""
"`GLM-5-w8a8`: [下载模型权重](https://www.modelscope.cn/models/Eco-Tech/GLM-5-w8a8)。"
#: ../../source/tutorials/models/GLM5.md:24
msgid ""
"You can use [msmodelslim](https://gitcode.com/Ascend/msmodelslim) to "
"quantify the model naively."
msgstr ""
"您可以使用 [msmodelslim](https://gitcode.com/Ascend/msmodelslim) 对模型进行简单的量化。"
#: ../../source/tutorials/models/GLM5.md:26
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`"
msgstr ""
"建议将模型权重下载到多个节点的共享目录中,例如 `/root/.cache/`"
#: ../../source/tutorials/models/GLM5.md:28
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/GLM5.md:30
msgid "You can use our official docker image to run GLM-5 directly."
msgstr "您可以使用我们的官方 docker 镜像直接运行 GLM-5。"
#: ../../source/tutorials/models/GLM5.md
msgid "A3 series"
msgstr "A3 系列"
#: ../../source/tutorials/models/GLM5.md:39
#: ../../source/tutorials/models/GLM5.md:86
msgid "Start the docker image on your each node."
msgstr "在您的每个节点上启动 docker 镜像。"
#: ../../source/tutorials/models/GLM5.md
msgid "A2 series"
msgstr "A2 系列"
#: ../../source/tutorials/models/GLM5.md:119
msgid ""
"In addition, if you don't want to use the docker image as above, you can "
"also build all from source:"
msgstr "此外,如果您不想使用上述的 docker 镜像,也可以从源码构建所有组件:"
#: ../../source/tutorials/models/GLM5.md:121
msgid ""
"Install `vllm-ascend` from source, refer to "
"[installation](https://docs.vllm.ai/projects/ascend/en/latest/installation.html)."
msgstr ""
"从源码安装 `vllm-ascend`,请参考[安装指南](https://docs.vllm.ai/projects/ascend/en/latest/installation.html)。"
#: ../../source/tutorials/models/GLM5.md:123
msgid ""
"If you want to deploy multi-node environment, you need to set up "
"environment on each node."
msgstr "如果您想部署多节点环境,需要在每个节点上设置环境。"
#: ../../source/tutorials/models/GLM5.md:125
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/GLM5.md:127
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/models/GLM5.md:136
msgid ""
"Quantized model `glm-5-w4a8` can be deployed on 1 Atlas 800 A3 (64G × 16)"
" ."
msgstr "量化模型 `glm-5-w4a8` 可以部署在 1 台 Atlas 800 A3 (64G × 16) 上。"
#: ../../source/tutorials/models/GLM5.md:138
#: ../../source/tutorials/models/GLM5.md:173
#: ../../source/tutorials/models/GLM5.md:213
msgid "Run the following script to execute online inference."
msgstr "运行以下脚本来执行在线推理。"
#: ../../source/tutorials/models/GLM5.md:171
msgid ""
"Quantized model `glm-5-w8a8` can be deployed on 1 Atlas 800 A3 (64G × 16)"
" ."
msgstr "量化模型 `glm-5-w8a8` 可以部署在 1 台 Atlas 800 A3 (64G × 16) 上。"
#: ../../source/tutorials/models/GLM5.md:211
msgid "Quantized model `glm-5-w4a8` can be deployed on 1 Atlas 800 A2 (64G × 8) ."
msgstr "量化模型 `glm-5-w4a8` 可以部署在 1 台 Atlas 800 A2 (64G × 8) 上。"
#: ../../source/tutorials/models/GLM5.md:248
msgid "**Notice:** The parameters are explained as follows:"
msgstr "**注意:** 参数解释如下:"
#: ../../source/tutorials/models/GLM5.md:251
msgid ""
"For single-node deployment, we recommend using `dp1tp16` and turn off "
"expert parallel in low-latency scenarios."
msgstr "对于单节点部署,在低延迟场景下,我们建议使用 `dp1tp16` 并关闭专家并行。"
#: ../../source/tutorials/models/GLM5.md:252
msgid ""
"`--async-scheduling` Asynchronous scheduling is a technique used to "
"optimize inference efficiency. It allows non-blocking task scheduling to "
"improve concurrency and throughput, especially when processing large-"
"scale models."
msgstr "`--async-scheduling` 异步调度是一种用于优化推理效率的技术。它允许非阻塞的任务调度,以提高并发性和吞吐量,尤其是在处理大规模模型时。"
#: ../../source/tutorials/models/GLM5.md:254
msgid "Multi-node Deployment"
msgstr "多节点部署"
#: ../../source/tutorials/models/GLM5.md:256
msgid ""
"If you want to deploy multi-node environment, you need to verify multi-"
"node communication according to [verify multi-node communication "
"environment](../../installation.md#verify-multi-node-communication)."
msgstr "如果您想部署多节点环境,需要根据[验证多节点通信环境](../../installation.md#verify-multi-node-communication)来验证多节点通信。"
#: ../../source/tutorials/models/GLM5.md:265
msgid "`glm-5-bf16`: require at least 2 Atlas 800 A3 (64G × 16)."
msgstr "`glm-5-bf16`: 需要至少 2 台 Atlas 800 A3 (64G × 16)。"
#: ../../source/tutorials/models/GLM5.md:267
#: ../../source/tutorials/models/GLM5.md:363
#: ../../source/tutorials/models/GLM5.md:528
msgid "Run the following scripts on two nodes respectively."
msgstr "分别在两个节点上运行以下脚本。"
#: ../../source/tutorials/models/GLM5.md:269
#: ../../source/tutorials/models/GLM5.md:365
#: ../../source/tutorials/models/GLM5.md:530
msgid "**node 0**"
msgstr "**节点 0**"
#: ../../source/tutorials/models/GLM5.md:313
#: ../../source/tutorials/models/GLM5.md:411
#: ../../source/tutorials/models/GLM5.md:580
msgid "**node 1**"
msgstr "**节点 1**"
#: ../../source/tutorials/models/GLM5.md:461
msgid ""
"For bf16 weight, use this script on each node to enable [Multi Token "
"Prediction "
"(MTP)](../../user_guide/feature_guide/Multi_Token_Prediction.md)."
msgstr "对于 bf16 权重,在每个节点上使用此脚本来启用[多令牌预测 (MTP)](../../user_guide/feature_guide/Multi_Token_Prediction.md)。"
#: ../../source/tutorials/models/GLM5.md:526
msgid "`glm-5-w8a8`: require 2 Atlas 800 A3 (64G × 16)."
msgstr "`glm-5-w8a8`: 需要 2 台 Atlas 800 A3 (64G × 16)。"
#: ../../source/tutorials/models/GLM5.md:634
msgid "Prefill-Decode Disaggregation"
msgstr "Prefill-Decode 解耦部署"
#: ../../source/tutorials/models/GLM5.md:636
msgid ""
"We'd like to show the deployment guide of `GLM-5` on multi-node "
"environment with 1P1D for better performance."
msgstr "我们将展示 `GLM-5` 在多节点环境下采用 1P1D 模式以获得更好性能的部署指南。"
#: ../../source/tutorials/models/GLM5.md:638
msgid "Before you start, please"
msgstr "在开始之前,请"
#: ../../source/tutorials/models/GLM5.md:640
msgid "prepare the script `launch_online_dp.py` on each node:"
msgstr "在每个节点上准备脚本 `launch_online_dp.py`"
#: ../../source/tutorials/models/GLM5.md:743
msgid "prepare the script `run_dp_template.sh` on each node."
msgstr "在每个节点上准备脚本 `run_dp_template.sh`。"
#: ../../source/tutorials/models/GLM5.md:745
msgid ""
"To support a 200k context window on the stage of prefill, the parameter "
"`\"layer_sharding\": [\"q_b_proj\"]` needs to be added to "
"`--additional_config` on each prefill node. In PD-disaggregated "
"deployment, `layer_sharding` is supported only on prefill/P nodes with "
"`kv_role=\"kv_producer\"`; do not enable it on decode/D nodes or "
"`kv_role=\"kv_both\"` nodes."
msgstr "为了在预填充阶段支持 200k 的上下文窗口,需要在每个预填充节点的 `--additional_config` 中添加参数 `\"layer_sharding\": [\"q_b_proj\"]`。在 PD 解耦部署中,`layer_sharding` 仅在 `kv_role=\"kv_producer\"` 的预填充/P 节点上受支持;不要在解码/D 节点或 `kv_role=\"kv_both\"` 的节点上启用它。"
#: ../../source/tutorials/models/GLM5.md:747
#: ../../source/tutorials/models/GLM5.md:1233
msgid "Prefill node 0"
msgstr "预填充节点 0"
#: ../../source/tutorials/models/GLM5.md:826
#: ../../source/tutorials/models/GLM5.md:1240
msgid "Prefill node 1"
msgstr "预填充节点 1"
#: ../../source/tutorials/models/GLM5.md:906
#: ../../source/tutorials/models/GLM5.md:1247
msgid "Decode node 0"
msgstr "解码节点 0"
#: ../../source/tutorials/models/GLM5.md:988
#: ../../source/tutorials/models/GLM5.md:1254
msgid "Decode node 1"
msgstr "解码节点 1"
#: ../../source/tutorials/models/GLM5.md:1069
#: ../../source/tutorials/models/GLM5.md:1261
msgid "Decode node 2"
msgstr "解码节点 2"
#: ../../source/tutorials/models/GLM5.md:1150
#: ../../source/tutorials/models/GLM5.md:1268
msgid "Decode node 3"
msgstr "解码节点 3"
#: ../../source/tutorials/models/GLM5.md:1231
msgid ""
"Once the preparation is done, you can start the server with the following"
" command on each node:"
msgstr "准备工作完成后,您可以在每个节点上使用以下命令启动服务器:"
#: ../../source/tutorials/models/GLM5.md:1275
msgid "Request Forwarding"
msgstr "请求转发"
#: ../../source/tutorials/models/GLM5.md:1277
msgid ""
"To set up request forwarding, run the following script on any machine. "
"You can get the proxy program in the repository's examples: "
"[load_balance_proxy_server_example.py](https://github.com/vllm-project"
"/vllm-"
"ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
msgstr "要设置请求转发,请在任何机器上运行以下脚本。您可以在仓库的示例中找到代理程序:[load_balance_proxy_server_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
#: ../../source/tutorials/models/GLM5.md:1318
msgid "**Notice:**"
msgstr "**注意:**"
#: ../../source/tutorials/models/GLM5.md:1320
msgid "Some configurations for optimization are shown below:"
msgstr "以下是一些用于优化的配置:"
#: ../../source/tutorials/models/GLM5.md:1322
msgid ""
"`VLLM_ASCEND_ENABLE_FLASHCOMM1`: Enable FlashComm optimization to reduce "
"communication and computation overhead on prefill node. With FlashComm "
"enabled, layer_sharding list cannot include o_proj as an element."
msgstr "`VLLM_ASCEND_ENABLE_FLASHCOMM1`: 启用 FlashComm 优化以减少预填充节点上的通信和计算开销。启用 FlashComm 后layer_sharding 列表不能包含 o_proj 作为元素。"
#: ../../source/tutorials/models/GLM5.md:1323
msgid ""
"`VLLM_ASCEND_ENABLE_FUSED_MC2`: Enable following fused operators: "
"dispatch_gmm_combine_decode and dispatch_ffn_combine operator."
msgstr "`VLLM_ASCEND_ENABLE_FUSED_MC2`: 启用以下融合算子dispatch_gmm_combine_decode 和 dispatch_ffn_combine 算子。"
#: ../../source/tutorials/models/GLM5.md:1324
msgid "`VLLM_ASCEND_ENABLE_MLAPO`: Enable fused operator MlaPreprocessOperation."
msgstr "`VLLM_ASCEND_ENABLE_MLAPO`: 启用融合算子 MlaPreprocessOperation。"
#: ../../source/tutorials/models/GLM5.md:1326
msgid ""
"Please refer to the following python file for further explanation and "
"restrictions of the environment variables above: "
"[envs.py](https://github.com/vllm-project/vllm-"
"ascend/blob/main/vllm_ascend/envs.py)"
msgstr "有关上述环境变量的进一步解释和限制,请参考以下 python 文件:[envs.py](https://github.com/vllm-project/vllm-ascend/blob/main/vllm_ascend/envs.py)"
#: ../../source/tutorials/models/GLM5.md:1328
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/GLM5.md:1330
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "服务器启动后,您可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/GLM5.md:1343
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/GLM5.md:1345
msgid "Here are two accuracy evaluation methods."
msgstr "以下是两种精度评估方法。"
#: ../../source/tutorials/models/GLM5.md:1347
#: ../../source/tutorials/models/GLM5.md:1359
msgid "Using AISBench"
msgstr "使用AISBench"
#: ../../source/tutorials/models/GLM5.md:1349
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参考[使用AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/GLM5.md:1351
msgid "After execution, you can get the result."
msgstr "执行后,您将获得结果。"
#: ../../source/tutorials/models/GLM5.md:1353
msgid "Using Language Model Evaluation Harness"
msgstr "使用Language Model Evaluation Harness"
#: ../../source/tutorials/models/GLM5.md:1355
msgid "Not tested yet."
msgstr "尚未测试。"
#: ../../source/tutorials/models/GLM5.md:1357
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/GLM5.md:1361
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参考[使用AISBench进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/GLM5.md:1363
msgid "Using vLLM Benchmark"
msgstr "使用vLLM基准测试"
#: ../../source/tutorials/models/GLM5.md:1365
msgid ""
"Refer to [vllm "
"benchmark](https://docs.vllm.ai/en/latest/contributing/benchmarks.html) "
"for more details."
msgstr "更多详情请参考[vllm基准测试](https://docs.vllm.ai/en/latest/contributing/benchmarks.html)。"
#: ../../source/tutorials/models/GLM5.md:1367
msgid "Best Practices"
msgstr "最佳实践"
#: ../../source/tutorials/models/GLM5.md:1369
msgid ""
"In this chapter, we recommend best practices in prefill-decode "
"disaggregation scenario with 1P1D architecture using 4 Atlas 800 A3 (64G "
"× 16):"
msgstr "本章节我们推荐在使用4台Atlas 800 A364G × 16的1P1D架构下预填充-解码分离场景的最佳实践:"
#: ../../source/tutorials/models/GLM5.md:1371
msgid ""
"Low-latency: We recommend setting `dp4 tp8` on prefill nodes and `dp4 "
"tp8` on decode nodes for low latency situation."
msgstr "低延迟场景:对于低延迟场景,我们建议在预填充节点上设置`dp4 tp8`,在解码节点上设置`dp4 tp8`。"
#: ../../source/tutorials/models/GLM5.md:1372
msgid ""
"High-throughput: `dp4 tp8` on prefill nodes and `dp8 tp4` on decode nodes"
" is recommended for high throughput situation."
msgstr "高吞吐场景:对于高吞吐场景,建议在预填充节点上设置`dp4 tp8`,在解码节点上设置`dp8 tp4`。"
#: ../../source/tutorials/models/GLM5.md:1374
msgid ""
"**Notice:** `max-model-len` and `max-num-seqs` need to be set according "
"to the actual usage scenario. For other settings, please refer to the "
"**[Deployment](#deployment)** chapter."
msgstr "**注意:** `max-model-len`和`max-num-seqs`需要根据实际使用场景进行设置。其他设置请参考**[部署](#deployment)**章节。"
#: ../../source/tutorials/models/GLM5.md:1377
msgid "FAQ"
msgstr "常见问题"
#: ../../source/tutorials/models/GLM5.md:1379
msgid ""
"**Q: How to solve ValueError: Tokenizer class TokenizersBackend does not "
"exist or is not currently imported?**"
msgstr "**问如何解决ValueError: Tokenizer class TokenizersBackend does not exist or is not currently imported?**"
#: ../../source/tutorials/models/GLM5.md:1381
msgid "A: Please update the version of transformers to 5.2.0"
msgstr "答请将transformers版本更新至5.2.0"
#: ../../source/tutorials/models/GLM5.md:1383
msgid "**Q: How to enable function calling for GLM-5?**"
msgstr "**问如何为GLM-5启用函数调用功能**"
#: ../../source/tutorials/models/GLM5.md:1385
msgid "A: Please add following configurations in vLLM startup command"
msgstr "答请在vLLM启动命令中添加以下配置"

View File

@@ -0,0 +1,134 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:1
msgid "Kimi-K2-Thinking"
msgstr "Kimi-K2-Thinking"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:5
msgid ""
"Kimi-K2-Thinking is a large-scale Mixture-of-Experts (MoE) model "
"developed by Moonshot AI. It features a hybrid thinking architecture that"
" excels in complex reasoning and problem-solving tasks."
msgstr "Kimi-K2-Thinking 是由 Moonshot AI 开发的大规模专家混合模型。它采用混合思维架构,在复杂推理和问题解决任务中表现出色。"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:7
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, environment preparation, single-node "
"deployment, and functional verification."
msgstr "本文档将展示该模型的主要验证步骤,包括支持的功能、环境准备、单节点部署和功能验证。"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:9
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:11
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取模型支持的功能矩阵。"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:13
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[功能指南](../../user_guide/feature_guide/index.md)以获取功能的配置信息。"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:15
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:17
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:19
msgid ""
"`Kimi-K2-Thinking`(bfloat16): require 1 Atlas 800 A3 (64G × 16) node. "
"[Download model "
"weight](https://huggingface.co/moonshotai/Kimi-K2-Thinking)."
msgstr "`Kimi-K2-Thinking`(bfloat16):需要 1 个 Atlas 800 A3 (64G × 16) 节点。[下载模型权重](https://huggingface.co/moonshotai/Kimi-K2-Thinking)。"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:21
msgid ""
"It is recommended to download the model weight to the shared directory, "
"such as `/mnt/sfs_turbo/.cache/`."
msgstr "建议将模型权重下载到共享目录,例如 `/mnt/sfs_turbo/.cache/`。"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:23
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:25
msgid "You can use our official docker image to run `Kimi-K2-Thinking` directly."
msgstr "您可以使用我们的官方 Docker 镜像直接运行 `Kimi-K2-Thinking`。"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:27
msgid ""
"Select an image based on your machine type and start the docker image on "
"your node, refer to [using docker](../../installation.md#set-up-using-"
"docker)."
msgstr "根据您的机器类型选择镜像并在节点上启动 Docker 镜像,请参考[使用 Docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:29
msgid "Run with Docker"
msgstr "使用 Docker 运行"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:72
msgid "Verify the Quantized Model"
msgstr "验证量化模型"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:74
msgid ""
"Please be advised to edit the value of "
"`\"quantization_config.config_groups.group_0.targets\"` from "
"`[\"Linear\"]` into `[\"MoE\"]` in `config.json` of original model "
"downloaded from [Hugging "
"Face](https://huggingface.co/moonshotai/Kimi-K2-Thinking)."
msgstr "请注意,请将从 [Hugging Face](https://huggingface.co/moonshotai/Kimi-K2-Thinking) 下载的原始模型的 `config.json` 文件中的 `\"quantization_config.config_groups.group_0.targets\"` 值从 `[\"Linear\"]` 修改为 `[\"MoE\"]`。"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:90
msgid "Your model files look like:"
msgstr "您的模型文件应类似如下结构:"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:109
msgid "Online Inference on Multi-NPU"
msgstr "多 NPU 在线推理"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:111
msgid "Run the following script to start the vLLM server on Multi-NPU:"
msgstr "运行以下脚本以在多 NPU 上启动 vLLM 服务器:"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:113
msgid ""
"For an Atlas 800 A3 (64G*16) node, tensor-parallel-size should be at "
"least 16."
msgstr "对于 Atlas 800 A3 (64G*16) 节点,张量并行大小应至少为 16。"
#: ../../source/tutorials/models/Kimi-K2-Thinking.md:136
msgid "Once your server is started, you can query the model with input prompts."
msgstr "服务器启动后,您可以使用输入提示词查询模型。"

View File

@@ -0,0 +1,582 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Kimi-K2.5.md:1
msgid "Kimi-K2.5"
msgstr "Kimi-K2.5"
#: ../../source/tutorials/models/Kimi-K2.5.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Kimi-K2.5.md:5
msgid ""
"Kimi K2.5 is an open-source, native multimodal agentic model built "
"through continual pretraining on approximately 15 trillion mixed visual "
"and text tokens atop Kimi-K2-Base. It seamlessly integrates vision and "
"language understanding with advanced agentic capabilities, instant and "
"thinking modes, as well as conversational and agentic paradigms."
msgstr ""
"Kimi K2.5 是一个开源的、原生的多模态智能体模型,通过在 Kimi-K2-Base 基础上持续预训练约 15 万亿视觉和文本混合令牌构建而成。它无缝集成了视觉与语言理解能力、先进的智能体能力、即时与思考模式,以及对话式和智能体范式。"
#: ../../source/tutorials/models/Kimi-K2.5.md:7
msgid "The `Kimi-K2.5` model is first supported in `vllm-ascend:v0.17.0rc1`."
msgstr "`Kimi-K2.5` 模型首次在 `vllm-ascend:v0.17.0rc1` 版本中得到支持。"
#: ../../source/tutorials/models/Kimi-K2.5.md:9
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, single-node and multi-node deployment, accuracy and "
"performance evaluation."
msgstr "本文档将展示该模型的主要验证步骤,包括支持的特性、特性配置、环境准备、单节点与多节点部署、精度和性能评估。"
#: ../../source/tutorials/models/Kimi-K2.5.md:11
msgid "Supported Features"
msgstr "支持的特性"
#: ../../source/tutorials/models/Kimi-K2.5.md:13
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考 [支持的特性](../../user_guide/support_matrix/supported_models.md) 获取模型支持的特性矩阵。"
#: ../../source/tutorials/models/Kimi-K2.5.md:15
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考 [特性指南](../../user_guide/feature_guide/index.md) 获取特性的配置信息。"
#: ../../source/tutorials/models/Kimi-K2.5.md:17
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Kimi-K2.5.md:19
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Kimi-K2.5.md:21
msgid ""
"`Kimi-K2.5-w4a8`(Quantized version for w4a8): [Download model "
"weight](https://modelscope.cn/models/Eco-Tech/Kimi-K2.5-W4A8)."
msgstr "`Kimi-K2.5-w4a8`w4a8量化版本[下载模型权重](https://modelscope.cn/models/Eco-Tech/Kimi-K2.5-W4A8)。"
#: ../../source/tutorials/models/Kimi-K2.5.md:22
msgid ""
"`kimi-k2.5-eagle3`(Eagle3 MTP draft model for accelerating inference of "
"Kimi-K2.5): [Download model "
"weight](https://huggingface.co/lightseekorg/kimi-k2.5-eagle3)"
msgstr "`kimi-k2.5-eagle3`(用于加速 Kimi-K2.5 推理的 Eagle3 MTP 草稿模型):[下载模型权重](https://huggingface.co/lightseekorg/kimi-k2.5-eagle3)"
#: ../../source/tutorials/models/Kimi-K2.5.md:24
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`."
msgstr "建议将模型权重下载到多节点的共享目录中,例如 `/root/.cache/`。"
#: ../../source/tutorials/models/Kimi-K2.5.md:26
msgid "Verify Multi-node Communication(Optional)"
msgstr "验证多节点通信(可选)"
#: ../../source/tutorials/models/Kimi-K2.5.md:28
msgid ""
"If you want to deploy multi-node environment, you need to verify multi-"
"node communication according to [verify multi-node communication "
"environment](../../installation.md#verify-multi-node-communication)."
msgstr "如果您想部署多节点环境,需要根据 [验证多节点通信环境](../../installation.md#verify-multi-node-communication) 验证多节点通信。"
#: ../../source/tutorials/models/Kimi-K2.5.md:30
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Kimi-K2.5.md:32
msgid "You can use our official docker image to run `Kimi-K2.5` directly."
msgstr "您可以使用我们的官方 docker 镜像直接运行 `Kimi-K2.5`。"
#: ../../source/tutorials/models/Kimi-K2.5.md:34
msgid ""
"Select an image based on your machine type and start the docker image on "
"your node, refer to [using docker](../../installation.md#set-up-using-"
"docker)."
msgstr "根据您的机器类型选择镜像,并在节点上启动 docker 镜像,请参考 [使用 docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/Kimi-K2.5.md
msgid "A3 series"
msgstr "A3 系列"
#: ../../source/tutorials/models/Kimi-K2.5.md:43
#: ../../source/tutorials/models/Kimi-K2.5.md:86
msgid "Start the docker image on your each node."
msgstr "在您的每个节点上启动 docker 镜像。"
#: ../../source/tutorials/models/Kimi-K2.5.md
msgid "A2 series"
msgstr "A2 系列"
#: ../../source/tutorials/models/Kimi-K2.5.md:119
msgid ""
"In addition, if you don't want to use the docker image as above, you can "
"also build all from source:"
msgstr "此外,如果您不想使用上述 docker 镜像,也可以从源码构建所有内容:"
#: ../../source/tutorials/models/Kimi-K2.5.md:121
msgid ""
"Install `vllm-ascend` from source, refer to "
"[installation](../../installation.md)."
msgstr "从源码安装 `vllm-ascend`,请参考 [安装](../../installation.md)。"
#: ../../source/tutorials/models/Kimi-K2.5.md:123
msgid ""
"If you want to deploy multi-node environment, you need to set up "
"environment on each node."
msgstr "如果您想部署多节点环境,需要在每个节点上设置环境。"
#: ../../source/tutorials/models/Kimi-K2.5.md:125
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Kimi-K2.5.md:127
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/models/Kimi-K2.5.md:129
msgid ""
"Quantized model `Kimi-K2.5-w4a8` can be deployed on 1 Atlas 800 A3 (64G ×"
" 16)."
msgstr "量化模型 `Kimi-K2.5-w4a8` 可以部署在 1 台 Atlas 800 A364G × 16上。"
#: ../../source/tutorials/models/Kimi-K2.5.md:131
msgid "Run the following script to execute online inference."
msgstr "运行以下脚本执行在线推理。"
#: ../../source/tutorials/models/Kimi-K2.5.md:176
#: ../../source/tutorials/models/Kimi-K2.5.md:645
msgid "**Notice:** The parameters are explained as follows:"
msgstr "**注意:** 参数解释如下:"
#: ../../source/tutorials/models/Kimi-K2.5.md:179
msgid ""
"Setting the environment variable `VLLM_ASCEND_BALANCE_SCHEDULING=1` "
"enables balance scheduling. This may help increase output throughput and "
"reduce TPOT in v1 scheduler. However, TTFT may degrade in some scenarios."
" Furthermore, enabling this feature is not recommended in scenarios where"
" PD is separated."
msgstr "设置环境变量 `VLLM_ASCEND_BALANCE_SCHEDULING=1` 启用均衡调度。这可能有助于提高 v1 调度器中的输出吞吐量并降低 TPOT。然而在某些场景下 TTFT 可能会下降。此外,在 PD 分离的场景中不建议启用此功能。"
#: ../../source/tutorials/models/Kimi-K2.5.md:180
msgid ""
"For single-node deployment, we recommend using `dp4tp4` instead of "
"`dp2tp8`."
msgstr "对于单节点部署,我们建议使用 `dp4tp4` 而不是 `dp2tp8`。"
#: ../../source/tutorials/models/Kimi-K2.5.md:181
msgid ""
"`--max-model-len` specifies the maximum context length - that is, the sum"
" of input and output tokens for a single request. For performance testing"
" with an input length of 3.5K and output length of 1.5K, a value of "
"`16384` is sufficient, however, for precision testing, please set it at "
"least `35000`."
msgstr "`--max-model-len` 指定最大上下文长度——即单个请求的输入和输出令牌总数。对于输入长度 3.5K 和输出长度 1.5K 的性能测试,`16384` 的值就足够了,但对于精度测试,请至少将其设置为 `35000`。"
#: ../../source/tutorials/models/Kimi-K2.5.md:182
msgid ""
"`--no-enable-prefix-caching` indicates that prefix caching is disabled. "
"To enable it, remove this option."
msgstr "`--no-enable-prefix-caching` 表示前缀缓存被禁用。要启用它,请移除此选项。"
#: ../../source/tutorials/models/Kimi-K2.5.md:183
msgid ""
"`--mm-encoder-tp-mode` indicates how to optimize multi-modal encoder "
"inference using tensor parallelism (TP). If you want to test the "
"multimodal inputs, we recommend using `data`."
msgstr "`--mm-encoder-tp-mode` 指示如何使用张量并行TP优化多模态编码器推理。如果您想测试多模态输入我们建议使用 `data`。"
#: ../../source/tutorials/models/Kimi-K2.5.md:184
msgid ""
"If you use the w4a8 weight, more memory will be allocated to kvcache, and"
" you can try to increase system throughput to achieve greater throughput."
msgstr "如果您使用 w4a8 权重,将有更多内存分配给 kvcache您可以尝试增加系统吞吐量以实现更高的吞吐量。"
#: ../../source/tutorials/models/Kimi-K2.5.md:186
msgid "Multi-node Deployment"
msgstr "多节点部署"
#: ../../source/tutorials/models/Kimi-K2.5.md:188
msgid "`Kimi-K2.5-w4a8`: require at least 2 Atlas 800 A2 (64G × 8)."
msgstr "`Kimi-K2.5-w4a8`:需要至少 2 台 Atlas 800 A264G × 8。"
#: ../../source/tutorials/models/Kimi-K2.5.md:190
msgid "Run the following scripts on two nodes respectively."
msgstr "分别在两个节点上运行以下脚本。"
#: ../../source/tutorials/models/Kimi-K2.5.md:192
msgid "**Node 0**"
msgstr "**节点 0**"
#: ../../source/tutorials/models/Kimi-K2.5.md:256
msgid "**Node 1**"
msgstr "**节点 1**"
#: ../../source/tutorials/models/Kimi-K2.5.md:322
msgid "Prefill-Decode Disaggregation"
msgstr "Prefill-Decode 分离"
#: ../../source/tutorials/models/Kimi-K2.5.md:324
msgid ""
"We recommend using Mooncake for deployment: "
"[Mooncake](../features/pd_disaggregation_mooncake_multi_node.md)."
msgstr "我们建议使用 Mooncake 进行部署:[Mooncake](../features/pd_disaggregation_mooncake_multi_node.md)。"
#: ../../source/tutorials/models/Kimi-K2.5.md:326
msgid ""
"Take Atlas 800 A3 (64G × 16) for example, we recommend to deploy 2P1D (4 "
"nodes) rather than 1P1D (2 nodes), because there is no enough NPU memory "
"to serve high concurrency in 1P1D case."
msgstr "以 Atlas 800 A364G × 16为例我们建议部署 2P1D4 个节点)而不是 1P1D2 个节点),因为在 1P1D 情况下没有足够的 NPU 内存来服务高并发。"
#: ../../source/tutorials/models/Kimi-K2.5.md:328
msgid "`Kimi-K2.5-w4a8 2P1D` require 4 Atlas 800 A3 (64G × 16)."
msgstr "`Kimi-K2.5-w4a8 2P1D` 需要 4 台 Atlas 800 A364G × 16。"
#: ../../source/tutorials/models/Kimi-K2.5.md:330
msgid ""
"To run the vllm-ascend `Prefill-Decode Disaggregation` service, you need "
"to deploy a `launch_dp_program.py` script and a `run_dp_template.sh` "
"script on each node and deploy a `proxy.sh` script on prefill master node"
" to forward requests."
msgstr "要运行 vllm-ascend `Prefill-Decode Disaggregation` 服务,您需要在每个节点上部署一个 `launch_dp_program.py` 脚本和一个 `run_dp_template.sh` 脚本,并在 prefill 主节点上部署一个 `proxy.sh` 脚本来转发请求。"
#: ../../source/tutorials/models/Kimi-K2.5.md:332
msgid ""
"`launch_online_dp.py` to launch external dp vllm servers. "
"[launch\\_online\\_dp.py](https://github.com/vllm-project/vllm-"
"ascend/blob/main/examples/external_online_dp/launch_online_dp.py)"
msgstr "`launch_online_dp.py` 用于启动外部 dp vllm 服务器。[launch\\_online\\_dp.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/external_online_dp/launch_online_dp.py)"
#: ../../source/tutorials/models/Kimi-K2.5.md:335
msgid "Prefill Node 0 `run_dp_template.sh` script"
msgstr "Prefill 节点 0 `run_dp_template.sh` 脚本"
#: ../../source/tutorials/models/Kimi-K2.5.md:413
msgid "Prefill Node 1 `run_dp_template.sh` script"
msgstr "Prefill 节点 1 `run_dp_template.sh` 脚本"
#: ../../source/tutorials/models/Kimi-K2.5.md:491
msgid "Decode Node 0 `run_dp_template.sh` script"
msgstr "Decode 节点 0 `run_dp_template.sh` 脚本"
#: ../../source/tutorials/models/Kimi-K2.5.md:568
msgid "Decode Node 1 `run_dp_template.sh` script"
msgstr "Decode 节点 1 `run_dp_template.sh` 脚本"
#: ../../source/tutorials/models/Kimi-K2.5.md:648
msgid ""
"`VLLM_ASCEND_ENABLE_FLASHCOMM1=1`: enables the communication optimization"
" function on the prefill nodes."
msgstr "`VLLM_ASCEND_ENABLE_FLASHCOMM1=1`:在 prefill 节点上启用通信优化功能。"
#: ../../source/tutorials/models/Kimi-K2.5.md:649
msgid ""
"`VLLM_ASCEND_ENABLE_MLAPO=1`: enables the fusion operator, which can "
"significantly improve performance but consumes more NPU memory. In the "
"Prefill-Decode (PD) separation scenario, enable MLAPO only on decode "
"nodes."
msgstr "`VLLM_ASCEND_ENABLE_MLAPO=1`:启用融合算子,这可以显著提高性能但会消耗更多 NPU 内存。在 Prefill-DecodePD分离场景中仅在 decode 节点上启用 MLAPO。"
#: ../../source/tutorials/models/Kimi-K2.5.md:650
msgid ""
"`--async-scheduling`: enables the asynchronous scheduling function. When "
"Multi-Token Prediction (MTP) is enabled, asynchronous scheduling of "
"operator delivery can be implemented to overlap the operator delivery "
"latency."
msgstr "`--async-scheduling`启用异步调度功能。当启用多令牌预测MTP可以实现算子交付的异步调度以重叠算子交付延迟。"
#: ../../source/tutorials/models/Kimi-K2.5.md:651
msgid ""
"`cudagraph_capture_sizes`: The recommended value is `n x (mtp + 1)`. And "
"the min is `n = 1` and the max is `n = max-num-seqs`. For other values, "
"it is recommended to set them to the number of frequently occurring "
"requests on the Decode (D) node."
msgstr "`cudagraph_capture_sizes`:推荐值为 `n x (mtp + 1)`。最小值为 `n = 1`,最大值为 `n = max-num-seqs`。对于其他值,建议将其设置为 DecodeD节点上频繁出现的请求数量。"
#: ../../source/tutorials/models/Kimi-K2.5.md:652
msgid ""
"`recompute_scheduler_enable: true`: enables the recomputation scheduler. "
"When the Key-Value Cache (KV Cache) of the decode node is insufficient, "
"requests will be sent to the prefill node to recompute the KV Cache. In "
"the PD separation scenario, it is recommended to enable this "
"configuration on both prefill and decode nodes simultaneously."
msgstr "`recompute_scheduler_enable: true`:启用重计算调度器。当 decode 节点的键值缓存KV Cache不足时请求将被发送到 prefill 节点以重新计算 KV Cache。在 PD 分离场景中,建议同时在 prefill 和 decode 节点上启用此配置。"
#: ../../source/tutorials/models/Kimi-K2.5.md:653
msgid ""
"`multistream_overlap_shared_expert: true`: When the Tensor Parallelism "
"(TP) size is 1 or `enable_shared_expert_dp: true`, an additional stream "
"is enabled to overlap the computation process of shared experts for "
"improved efficiency."
msgstr "`multistream_overlap_shared_expert: true`当张量并行TP大小为 1 或 `enable_shared_expert_dp: true` 时,启用额外的流来重叠共享专家的计算过程以提高效率。"
#: ../../source/tutorials/models/Kimi-K2.5.md:655
msgid "run server for each node:"
msgstr "为每个节点运行服务器:"
#: ../../source/tutorials/models/Kimi-K2.5.md:668
msgid "Run the `proxy.sh` script on the prefill master node"
msgstr "在 prefill 主节点上运行 `proxy.sh` 脚本"
#: ../../source/tutorials/models/Kimi-K2.5.md:670
msgid ""
"Run a proxy server on the same node with the prefiller service instance. "
"You can get the proxy program in the repository's examples: "
"[load\\_balance\\_proxy\\_server\\_example.py](https://github.com/vllm-"
"project/vllm-"
"ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
msgstr "在与 prefiller 服务实例相同的节点上运行一个代理服务器。您可以在仓库的示例中找到代理程序:[load\\_balance\\_proxy\\_server\\_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
#: ../../source/tutorials/models/Kimi-K2.5.md:726
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/Kimi-K2.5.md:728
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "一旦您的服务器启动,您就可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/Kimi-K2.5.md:749
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/Kimi-K2.5.md:751
msgid "Here are two accuracy evaluation methods."
msgstr "以下是两种精度评估方法。"
#: ../../source/tutorials/models/Kimi-K2.5.md:753
#: ../../source/tutorials/models/Kimi-K2.5.md:768
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/models/Kimi-K2.5.md:755
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参考 [使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/Kimi-K2.5.md:757
msgid ""
"After execution, you can get the result, here is the result of "
"`Kimi-K2.5-w4a8` in `vllm-ascend:v0.18.0rc1` for reference only."
msgstr "执行后,您将获得结果。以下为 `Kimi-K2.5-w4a8` 在 `vllm-ascend:v0.18.0rc1` 环境下的结果,仅供参考。"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "version"
msgstr "版本"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "note"
msgstr "备注"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "GSM8K"
msgstr "GSM8K"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "-"
msgstr "-"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "accuracy"
msgstr "准确率"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "gen"
msgstr "生成"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "96.07"
msgstr "96.07"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "1 Atlas 800 A3 (64G × 16)"
msgstr "1 Atlas 800 A3 (64G × 16)"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "AIME2025"
msgstr "AIME2025"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "90.00"
msgstr "90.00"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "GPQA"
msgstr "GPQA"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "84.85"
msgstr "84.85"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "TextVQA"
msgstr "TextVQA"
#: ../../source/tutorials/models/Kimi-K2.5.md:88
msgid "80.29"
msgstr "80.29"
#: ../../source/tutorials/models/Kimi-K2.5.md:766
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Kimi-K2.5.md:770
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参考 [使用 AISBench 进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/Kimi-K2.5.md:772
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM Benchmark"
#: ../../source/tutorials/models/Kimi-K2.5.md:774
msgid "Run performance evaluation of `Kimi-K2.5-w4a8` as an example."
msgstr "以运行 `Kimi-K2.5-w4a8` 的性能评估为例。"
#: ../../source/tutorials/models/Kimi-K2.5.md:776
msgid ""
"Refer to [vllm "
"benchmark](https://docs.vllm.ai/en/latest/contributing/benchmarks.html) "
"for more details."
msgstr "更多详情请参考 [vllm benchmark](https://docs.vllm.ai/en/latest/contributing/benchmarks.html)。"
#: ../../source/tutorials/models/Kimi-K2.5.md:778
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 包含三个子命令:"
#: ../../source/tutorials/models/Kimi-K2.5.md:780
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批请求的延迟进行基准测试。"
#: ../../source/tutorials/models/Kimi-K2.5.md:781
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/models/Kimi-K2.5.md:782
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/models/Kimi-K2.5.md:784
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。运行以下代码。"
#: ../../source/tutorials/models/Kimi-K2.5.md:791
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您将获得性能评估结果。"
#: ../../source/tutorials/models/Kimi-K2.5.md:793
msgid "Best Practices"
msgstr "最佳实践"
#: ../../source/tutorials/models/Kimi-K2.5.md:795
msgid "In this chapter, we recommend best practices for three scenarios:"
msgstr "本章节针对三种场景推荐最佳实践:"
#: ../../source/tutorials/models/Kimi-K2.5.md:797
msgid ""
"Long-context: For long sequences with low concurrency (≤ 4): set `dp1 "
"tp16`; For long sequences with high concurrency (> 4): set `dp2 tp8`"
msgstr "长上下文:对于低并发(≤ 4的长序列设置 `dp1 tp16`;对于高并发(> 4的长序列设置 `dp2 tp8`"
#: ../../source/tutorials/models/Kimi-K2.5.md:798
msgid ""
"Low-latency: For short sequences with low latency: we recommend setting "
"`dp2 tp8`"
msgstr "低延迟:对于需要低延迟的短序列:我们推荐设置 `dp2 tp8`"
#: ../../source/tutorials/models/Kimi-K2.5.md:799
msgid ""
"High-throughput: For short sequences with high throughput: we also "
"recommend setting `dp4 tp4`"
msgstr "高吞吐量:对于需要高吞吐量的短序列:我们也推荐设置 `dp4 tp4`"
#: ../../source/tutorials/models/Kimi-K2.5.md:801
msgid ""
"**Notice:** `max-model-len` and `max-num-seqs` need to be set according "
"to the actual usage scenario. For other settings, please refer to the "
"**[Deployment](#deployment)** chapter."
msgstr "**注意:** `max-model-len` 和 `max-num-seqs` 需要根据实际使用场景进行设置。其他设置请参考 **[部署](#deployment)** 章节。"
#: ../../source/tutorials/models/Kimi-K2.5.md:804
msgid "FAQ"
msgstr "常见问题"
#: ../../source/tutorials/models/Kimi-K2.5.md:806
msgid "**Q: Why is the TPOT performance poor in Long-context test?**"
msgstr "**问:为什么在长上下文测试中 TPOT 性能不佳?**"
#: ../../source/tutorials/models/Kimi-K2.5.md:808
msgid ""
"A: Please ensure that the FIA operator replacement script has been "
"executed successfully to complete the replacement of FIA operators. Here "
"is the script: "
"[A2](../../../../tools/install_flash_infer_attention_score_ops_a2.sh) and"
" [A3](../../../../tools/install_flash_infer_attention_score_ops_a3.sh)"
msgstr "答:请确保已成功执行 FIA 算子替换脚本以完成 FIA 算子的替换。脚本如下:[A2](../../../../tools/install_flash_infer_attention_score_ops_a2.sh) 和 [A3](../../../../tools/install_flash_infer_attention_score_ops_a3.sh)"
#: ../../source/tutorials/models/Kimi-K2.5.md:810
msgid ""
"**Q: Startup fails with HCCL port conflicts (address already bound). What"
" should I do?**"
msgstr "**问:启动失败,提示 HCCL 端口冲突(地址已被占用)。我该怎么办?**"
#: ../../source/tutorials/models/Kimi-K2.5.md:812
msgid "A: Clean up old processes and restart: `pkill -f VLLM*`."
msgstr "答:清理旧进程并重启:`pkill -f VLLM*`。"
#: ../../source/tutorials/models/Kimi-K2.5.md:814
msgid "**Q: How to handle OOM or unstable startup?**"
msgstr "**问:如何处理 OOM 或启动不稳定的问题?**"
#: ../../source/tutorials/models/Kimi-K2.5.md:816
msgid ""
"A: Reduce `--max-num-seqs` and `--max-model-len` first. If needed, reduce"
" concurrency and load-testing pressure (e.g., `max-concurrency` / `num-"
"prompts`)."
msgstr "答:首先减少 `--max-num-seqs` 和 `--max-model-len`。如有需要,降低并发度和压测压力(例如,`max-concurrency` / `num-prompts`)。"

View File

@@ -0,0 +1,574 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/MiniMax-M2.md:1
msgid "MiniMax-M2"
msgstr "MiniMax-M2"
#: ../../source/tutorials/models/MiniMax-M2.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/MiniMax-M2.md:5
msgid ""
"MiniMaxM2.5 is MiniMaxs flagship large language model, reinforced for "
"highvalue scenarios such as code generation, agentic tool "
"calling/search, and complex office workflows, with an emphasis on "
"reasoning efficiency and endtoend speed on challenging tasks."
msgstr ""
"MiniMaxM2.5 是 MiniMax 的旗舰大语言模型,针对代码生成、智能体工具调用/搜索以及复杂办公工作流等高价值场景进行了强化,重点在于推理效率和在挑战性任务上的端到端速度。"
#: ../../source/tutorials/models/MiniMax-M2.md:7
msgid ""
"MiniMax-M2.7 is MiniMax's first model deeply participating in its own "
"evolution. M2.7 is capable of building complex agent harnesses and "
"completing highly elaborate productivity tasks, leveraging Agent Teams, "
"complex Skills, and dynamic tool search."
msgstr ""
"MiniMax-M2.7 是 MiniMax 首个深度参与自身演进的模型。M2.7 能够构建复杂的智能体框架并完成高度精细的生产力任务,利用智能体团队、复杂技能和动态工具搜索。"
#: ../../source/tutorials/models/MiniMax-M2.md:9
msgid ""
"This document provides a unified deployment guide for `MiniMax-M2.5` and "
"`MiniMax-M2.7` on vLLM Ascend, covering both:"
msgstr "本文档提供了在 vLLM Ascend 上部署 `MiniMax-M2.5` 和 `MiniMax-M2.7` 的统一指南,涵盖以下两种部署方式:"
#: ../../source/tutorials/models/MiniMax-M2.md:11
msgid "**A3 single-node** deployment (Atlas 800 A3)"
msgstr "**A3 单节点**部署Atlas 800 A3"
#: ../../source/tutorials/models/MiniMax-M2.md:12
msgid "**A2 dual-node** deployment (2× Atlas 800I A2)"
msgstr "**A2 双节点**部署2× Atlas 800I A2"
#: ../../source/tutorials/models/MiniMax-M2.md:14
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/MiniMax-M2.md:16
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取模型支持的功能矩阵。"
#: ../../source/tutorials/models/MiniMax-M2.md:18
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[功能指南](../../user_guide/feature_guide/index.md)以获取功能的配置信息。"
#: ../../source/tutorials/models/MiniMax-M2.md:20
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/MiniMax-M2.md:22
msgid "Model Weights"
msgstr "模型权重"
#: ../../source/tutorials/models/MiniMax-M2.md:24
msgid ""
"`MiniMax-M2.5` (fp8 checkpoint): recommended to use **1× Atlas 800 A3** "
"or **2× Atlas 800I A2** nodes. Download the model weights from "
"[MiniMax/MiniMax-M2.5](https://modelscope.cn/models/MiniMax/MiniMax-M2.5)."
msgstr ""
"`MiniMax-M2.5`fp8 检查点):推荐使用 **1× Atlas 800 A3** 或 **2× Atlas 800I A2** 节点。从 "
"[MiniMax/MiniMax-M2.5](https://modelscope.cn/models/MiniMax/MiniMax-M2.5) 下载模型权重。"
#: ../../source/tutorials/models/MiniMax-M2.md:25
msgid ""
"`MiniMax-M2.5-w8a8-QuaRot` : Download the model weights from [Eco-"
"Tech/MiniMax-M2.5-w8a8-QuaRot](https://modelscope.cn/models/Eco-"
"Tech/MiniMax-M2.5-w8a8-QuaRot)."
msgstr ""
"`MiniMax-M2.5-w8a8-QuaRot`:从 [Eco-Tech/MiniMax-M2.5-w8a8-"
"QuaRot](https://modelscope.cn/models/Eco-Tech/MiniMax-M2.5-w8a8-QuaRot) 下载模型权重。"
#: ../../source/tutorials/models/MiniMax-M2.md:26
msgid ""
"`Eagle3` : Download the model weights from [vllm-ascend/MiniMax-M2.5"
"-eagel-model](https://modelscope.cn/models/vllm-ascend/MiniMax-M2.5"
"-eagel-model-0318)."
msgstr ""
"`Eagle3`:从 [vllm-ascend/MiniMax-M2.5-eagel-"
"model](https://modelscope.cn/models/vllm-ascend/MiniMax-M2.5-eagel-model-0318) 下载模型权重。"
#: ../../source/tutorials/models/MiniMax-M2.md:27
msgid ""
"`MiniMax-M2.7` (fp8 checkpoint): recommended to use **1× Atlas 800 A3** "
"or **2× Atlas 800I A2** nodes. Download the model weights from "
"[MiniMax/MiniMax-M2.7](https://modelscope.cn/models/MiniMax/MiniMax-M2.7)."
msgstr ""
"`MiniMax-M2.7`fp8 检查点):推荐使用 **1× Atlas 800 A3** 或 **2× Atlas 800I A2** 节点。从 "
"[MiniMax/MiniMax-M2.7](https://modelscope.cn/models/MiniMax/MiniMax-M2.7) 下载模型权重。"
#: ../../source/tutorials/models/MiniMax-M2.md:28
msgid ""
"`MiniMax-M2.7-w8a8-QuaRot` : Download the model weights from [Eco-"
"Tech/MiniMax-M2.7-w8a8-QuaRot](https://modelscope.cn/models/Eco-"
"Tech/MiniMax-M2.7-w8a8-QuaRot)."
msgstr ""
"`MiniMax-M2.7-w8a8-QuaRot`:从 [Eco-Tech/MiniMax-M2.7-w8a8-"
"QuaRot](https://modelscope.cn/models/Eco-Tech/MiniMax-M2.7-w8a8-QuaRot) 下载模型权重。"
#: ../../source/tutorials/models/MiniMax-M2.md:30
msgid ""
"It is recommended to download the model weights to a shared directory, "
"such as `/mnt/sfs_turbo/.cache/`. The current release automatically "
"detects the MiniMax-M2 fp8 checkpoint, disables fp8 quantization kernels "
"on NPU, and loads the weights by dequantizing to bf16. This behavior may "
"be removed once public bf16 weights are available."
msgstr ""
"建议将模型权重下载到共享目录,例如 `/mnt/sfs_turbo/.cache/`。当前版本会自动检测 MiniMax-M2 的 fp8 检查点,在 NPU 上禁用 fp8 "
"量化内核,并通过反量化为 bf16 来加载权重。一旦公开的 bf16 权重可用,此行为可能会被移除。"
#: ../../source/tutorials/models/MiniMax-M2.md:32
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/MiniMax-M2.md:34
msgid "You can use the official docker image to run `MiniMax-M2.5/M2.7` directly."
msgstr "您可以使用官方的 docker 镜像直接运行 `MiniMax-M2.5/M2.7`。"
#: ../../source/tutorials/models/MiniMax-M2.md:36
msgid ""
"Select an image based on your machine type and start the container on "
"your node. See [using docker](../../installation.md#set-up-using-docker)."
msgstr "根据您的机器类型选择镜像,并在您的节点上启动容器。请参阅[使用 docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/MiniMax-M2.md:38
msgid "Run with Docker"
msgstr "使用 Docker 运行"
#: ../../source/tutorials/models/MiniMax-M2.md:40
#: ../../source/tutorials/models/MiniMax-M2.md:129
#: ../../source/tutorials/models/MiniMax-M2.md:332
msgid "A3 (single node)"
msgstr "A3单节点"
#: ../../source/tutorials/models/MiniMax-M2.md:83
msgid "A2 (dual node, run on both nodes)"
msgstr "A2双节点在两个节点上运行"
#: ../../source/tutorials/models/MiniMax-M2.md:85
msgid "Create and run `minimax25-docker-run.sh` on **both** A2 nodes."
msgstr "在**两个** A2 节点上创建并运行 `minimax25-docker-run.sh`。"
#: ../../source/tutorials/models/MiniMax-M2.md:87
#: ../../source/tutorials/models/MiniMax-M2.md:133
msgid "Notes:"
msgstr "注意:"
#: ../../source/tutorials/models/MiniMax-M2.md:89
msgid ""
"The default configuration assumes an **Atlas 800I A2 8-NPU** node and "
"sets `ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7`. Update it based on your"
" hardware."
msgstr "默认配置假设为 **Atlas 800I A2 8-NPU** 节点,并设置 `ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7`。请根据您的硬件进行更新。"
#: ../../source/tutorials/models/MiniMax-M2.md:90
msgid ""
"Map your model weight directory into the container (the example maps it "
"to `/opt/data/verification/`)."
msgstr "将您的模型权重目录映射到容器中(示例中映射到 `/opt/data/verification/`)。"
#: ../../source/tutorials/models/MiniMax-M2.md:125
msgid "Online Inference on Multi-NPU"
msgstr "多 NPU 在线推理"
#: ../../source/tutorials/models/MiniMax-M2.md:127
msgid ""
"Below are recommended startup configurations for `MiniMax-M2.5`. Users "
"can simply change weights and model name to run this startup "
"configuration on `MiniMax-M2.7`. However it may not yet the best matchup "
"for `MiniMax-M2.7` if one is trying to reach the best performance."
msgstr "以下是 `MiniMax-M2.5` 的推荐启动配置。用户可以简单地更改权重和模型名称,即可在 `MiniMax-M2.7` 上运行此启动配置。但是,如果试图达到最佳性能,这可能还不是 `MiniMax-M2.7` 的最佳匹配。"
#: ../../source/tutorials/models/MiniMax-M2.md:131
msgid ""
"Below is a recommended startup configuration for short-context condition "
"like 3.5k/1.5k on `MiniMax-M2.5` to reach a good performance. If you wish"
" to run on long-context case, you may follow `Remarks` below to change "
"your config."
msgstr "以下是为 `MiniMax-M2.5` 在短上下文(如 3.5k/1.5k)条件下达到良好性能的推荐启动配置。如果您希望在长上下文情况下运行,可以按照下面的`备注`来更改配置。"
#: ../../source/tutorials/models/MiniMax-M2.md:135
msgid ""
"If you only care about short-context low latency, you can explicitly set "
"`--max-model-len 32768`. You may also set `tensor-parallel-size` to 16 "
"and set `data-parallel-size` to 1."
msgstr "如果您只关心短上下文的低延迟,可以显式设置 `--max-model-len 32768`。您也可以将 `tensor-parallel-size` 设置为 16并将 `data-parallel-size` 设置为 1。"
#: ../../source/tutorials/models/MiniMax-M2.md:136
msgid ""
"`export VLLM_ASCEND_BALANCE_SCHEDULING=1` is used to enhance scheduling "
"capacity between prefill and decode. This will work remarkably with a "
"lager `data-parallel-size`. This can increace performance when "
"cuncurrency gets closer to values equals to `data-parallel-size` times "
"`max-num-seqs`."
msgstr ""
"`export VLLM_ASCEND_BALANCE_SCHEDULING=1` 用于增强预填充和解码之间的调度能力。这在 `data-parallel-size` "
"较大时效果显著。当并发数接近 `data-parallel-size` 乘以 `max-num-seqs` 的值时,这可以提高性能。"
#: ../../source/tutorials/models/MiniMax-M2.md:137
msgid ""
"Running the current Eagle3 weights for `MiniMax-M2.7` yields no "
"performance improvement; it is recommended to remove the "
"`--speculative_config`."
msgstr "为 `MiniMax-M2.7` 运行当前的 Eagle3 权重不会带来性能提升;建议移除 `--speculative_config`。"
#: ../../source/tutorials/models/MiniMax-M2.md:174
msgid "Remarks:"
msgstr "备注:"
#: ../../source/tutorials/models/MiniMax-M2.md:176
msgid "`minimax_m2_append_think` keeps `<think>...</think>` inside `content`."
msgstr "`minimax_m2_append_think` 会将 `<think>...</think>` 保留在 `content` 内部。"
#: ../../source/tutorials/models/MiniMax-M2.md:177
msgid ""
"If you mainly rely on the reasoning semantics of `/v1/responses`, it is "
"recommended to use `--reasoning-parser minimax_m2` instead."
msgstr "如果您主要依赖 `/v1/responses` 的推理语义,建议改用 `--reasoning-parser minimax_m2`。"
#: ../../source/tutorials/models/MiniMax-M2.md:178
msgid ""
"To receive a better performance on long-context like 128k or 64k, we "
"recommend to do changes as shown below, and you can remove `export "
"VLLM_ASCEND_BALANCE_SCHEDULING=1`."
msgstr "为了在 128k 或 64k 等长上下文上获得更好的性能,我们建议进行如下更改,并且您可以移除 `export VLLM_ASCEND_BALANCE_SCHEDULING=1`。"
#: ../../source/tutorials/models/MiniMax-M2.md:193
msgid ""
"If you will to test with `curl` command, you can add following commands "
"addition to start up command above."
msgstr "如果您想使用 `curl` 命令进行测试,可以在上述启动命令的基础上添加以下命令。"
#: ../../source/tutorials/models/MiniMax-M2.md:201
msgid "A2 (dual node, tp=8 + dp=2)"
msgstr "A2双节点tp=8 + dp=2"
#: ../../source/tutorials/models/MiniMax-M2.md:203
msgid ""
"Since cross-node tensor parallelism (TP) can be unstable, the dual-node "
"guide uses a **tp=8 + dp=2** setup (8 NPUs per node, 16 NPUs total)."
msgstr "由于跨节点的张量并行TP可能不稳定双节点指南采用 **tp=8 + dp=2** 的设置(每个节点 8 个 NPU总共 16 个 NPU。"
#: ../../source/tutorials/models/MiniMax-M2.md:205
msgid "Node0 (primary) startup script"
msgstr "Node0主节点启动脚本"
#: ../../source/tutorials/models/MiniMax-M2.md:207
msgid ""
"Edit `minimax25_service_node0.sh` inside the node0 container, and replace"
" the placeholders with your actual values:"
msgstr "在 node0 容器内编辑 `minimax25_service_node0.sh`,并将占位符替换为您的实际值:"
#: ../../source/tutorials/models/MiniMax-M2.md:209
#, python-brace-format
msgid "`{PrimaryNodeIP}`: the primary node's IP address (public/cluster network)"
msgstr "`{PrimaryNodeIP}`:主节点的 IP 地址(公共/集群网络)"
#: ../../source/tutorials/models/MiniMax-M2.md:210
#, python-brace-format
msgid ""
"`{NIC}`: the NIC name for the public/cluster network (check via "
"`ifconfig`, e.g., `enp67s0f0np0`)"
msgstr "`{NIC}`:公共/集群网络的网卡名称(通过 `ifconfig` 检查,例如 `enp67s0f0np0`"
#: ../../source/tutorials/models/MiniMax-M2.md:211
msgid "`VLLM_TORCH_PROFILER_DIR`: optional, directory to store profiling outputs"
msgstr "`VLLM_TORCH_PROFILER_DIR`:可选,用于存储性能分析输出的目录"
#: ../../source/tutorials/models/MiniMax-M2.md:260
msgid "Node1 (secondary) startup script"
msgstr "Node1从节点启动脚本"
#: ../../source/tutorials/models/MiniMax-M2.md:262
msgid "Edit `minimax25_service_node1.sh` inside the node1 container:"
msgstr "在 node1 容器内编辑 `minimax25_service_node1.sh`"
#: ../../source/tutorials/models/MiniMax-M2.md:264
#, python-brace-format
msgid "`{SecondaryNodeIP}`: the secondary node's IP address"
msgstr "`{SecondaryNodeIP}`:从节点的 IP 地址"
#: ../../source/tutorials/models/MiniMax-M2.md:265
#, python-brace-format
msgid "`{PrimaryNodeIP}`: the primary node's IP address (same as node0)"
msgstr "`{PrimaryNodeIP}`:主节点的 IP 地址(与 node0 相同)"
#: ../../source/tutorials/models/MiniMax-M2.md:266
#, python-brace-format
msgid "`{NIC}`: same as above"
msgstr "`{NIC}`:同上"
#: ../../source/tutorials/models/MiniMax-M2.md:316
msgid "Startup order"
msgstr "启动顺序"
#: ../../source/tutorials/models/MiniMax-M2.md:318
msgid "Start the service on both nodes:"
msgstr "在两个节点上启动服务:"
#: ../../source/tutorials/models/MiniMax-M2.md:328
msgid "After node0 prints `service start` in logs, you can verify the service."
msgstr "在 node0 的日志中打印出 `service start` 后,您可以验证服务。"
#: ../../source/tutorials/models/MiniMax-M2.md:330
msgid "Verify the Service"
msgstr "验证服务"
#: ../../source/tutorials/models/MiniMax-M2.md:334
msgid "Test with an OpenAI-compatible client:"
msgstr "使用 OpenAI 兼容的客户端进行测试:"
#: ../../source/tutorials/models/MiniMax-M2.md:349
msgid "Or send a request using curl:"
msgstr "或者使用 curl 发送请求:"
#: ../../source/tutorials/models/MiniMax-M2.md:378
msgid "A2 (dual node)"
msgstr "A2双节点"
#: ../../source/tutorials/models/MiniMax-M2.md:380
#, python-brace-format
msgid ""
"Run the following from any machine that can reach the primary node "
"(replace `{PrimaryNodeIP}` with the real IP):"
msgstr "从任何可以访问主节点的机器上运行以下命令(将 `{PrimaryNodeIP}` 替换为真实 IP"
#: ../../source/tutorials/models/MiniMax-M2.md:396
msgid "Performance Reference (`MiniMax-M2.5`)"
msgstr "性能参考(`MiniMax-M2.5`"
#: ../../source/tutorials/models/MiniMax-M2.md:398
msgid "A3 (single node, tp=16, 4k/1k@bs16)"
msgstr "A3单节点tp=164k/1k@bs16"
#: ../../source/tutorials/models/MiniMax-M2.md:400
#: ../../source/tutorials/models/MiniMax-M2.md:446
msgid "Results"
msgstr "结果"
#: ../../source/tutorials/models/MiniMax-M2.md:402
msgid "**Baseline** (`3.5k/1k@bs=217`)"
msgstr "**基线**`3.5k/1k@bs=217`"
#: ../../source/tutorials/models/MiniMax-M2.md:382
#: ../../source/tutorials/models/MiniMax-M2.md:427
msgid "Metric"
msgstr "指标"
#: ../../source/tutorials/models/MiniMax-M2.md:382
#: ../../source/tutorials/models/MiniMax-M2.md:427
msgid "Result"
msgstr "结果"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "Success/Failure"
msgstr "成功/失败"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "`217/0`"
msgstr "`217/0`"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "Mean TTFT"
msgstr "平均TTFT"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "`10316.56 ms`"
msgstr "`10316.56 毫秒`"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "Mean TPOT"
msgstr "平均TPOT"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "`34.28 ms`"
msgstr "`34.28 毫秒`"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "Output tok/s"
msgstr "输出令牌/秒"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "`4803.81`"
msgstr "`4803.81`"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "Total tok/s"
msgstr "总令牌/秒"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "`16096.59`"
msgstr "`16096.59`"
#: ../../source/tutorials/models/MiniMax-M2.md:412
msgid "**Long-context reference** (`190k/1k@bs=4`)"
msgstr "**长上下文参考** (`190k/1k@bs=4`)"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "`37.12`"
msgstr "`37.12`"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "`2002.37 ms`"
msgstr "`2002.37 毫秒`"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "`105.54 ms`"
msgstr "`105.54 毫秒`"
#: ../../source/tutorials/models/MiniMax-M2.md:382
msgid "Mean ITL"
msgstr "平均ITL"
#: ../../source/tutorials/models/MiniMax-M2.md:421
msgid "A2 (dual node, 190k/1k, concurrency=4, 16 prompts)"
msgstr "A2 (双节点190k/1k并发数=416个提示词)"
#: ../../source/tutorials/models/MiniMax-M2.md:423
msgid "Benchmark method"
msgstr "基准测试方法"
#: ../../source/tutorials/models/MiniMax-M2.md:425
msgid "Use vLLM bench for the **190k/1k, concurrency=4, 16 prompts** scenario:"
msgstr "使用 vLLM bench 进行 **190k/1k并发数=416个提示词** 场景的测试:"
#: ../../source/tutorials/models/MiniMax-M2.md:448
msgid "**190k/1k, concurrency=4, 16 prompts**"
msgstr "**190k/1k并发数=416个提示词**"
#: ../../source/tutorials/models/MiniMax-M2.md:427
msgid "TTFT (avg)"
msgstr "TTFT (平均)"
#: ../../source/tutorials/models/MiniMax-M2.md:427
msgid "3305.25 ms"
msgstr "3305.25 毫秒"
#: ../../source/tutorials/models/MiniMax-M2.md:427
msgid "TPOT (avg)"
msgstr "TPOT (平均)"
#: ../../source/tutorials/models/MiniMax-M2.md:427
msgid "109.83 ms"
msgstr "109.83 毫秒"
#: ../../source/tutorials/models/MiniMax-M2.md:427
msgid "Output throughput"
msgstr "输出吞吐量"
#: ../../source/tutorials/models/MiniMax-M2.md:427
msgid "35.29 tok/s"
msgstr "35.29 令牌/秒"
#: ../../source/tutorials/models/MiniMax-M2.md:427
msgid "Prefix hit rate"
msgstr "前缀命中率"
#: ../../source/tutorials/models/MiniMax-M2.md:427
msgid "85%"
msgstr "85%"
#: ../../source/tutorials/models/MiniMax-M2.md:457
msgid "FAQ"
msgstr "常见问题"
#: ../../source/tutorials/models/MiniMax-M2.md:459
msgid "**Q: What should I do if the output is garbled in EP mode?**"
msgstr "**问:在 EP 模式下输出乱码怎么办?**"
#: ../../source/tutorials/models/MiniMax-M2.md:461
msgid ""
"A: It is recommended to keep `--enable-expert-parallel` and "
"`VLLM_ASCEND_ENABLE_FLASHCOMM1=1`."
msgstr "答:建议保持启用 `--enable-expert-parallel` 并设置 `VLLM_ASCEND_ENABLE_FLASHCOMM1=1`。"
#: ../../source/tutorials/models/MiniMax-M2.md:463
msgid ""
"**Q: Why is the `reasoning` field often empty after using "
"`minimax_m2_append_think`?**"
msgstr "**问:为什么使用 `minimax_m2_append_think` 后 `reasoning` 字段经常为空?**"
#: ../../source/tutorials/models/MiniMax-M2.md:465
msgid ""
"A: This is expected. The parser keeps `<think>...</think>` inside "
"`content`. If you mainly rely on the reasoning semantics of "
"`/v1/responses`, use `--reasoning-parser minimax_m2` instead."
msgstr "答:这是预期行为。解析器会将 `<think>...</think>` 保留在 `content` 字段内。如果您主要依赖 `/v1/responses` 的推理语义,请改用 `--reasoning-parser minimax_m2`。"
#: ../../source/tutorials/models/MiniMax-M2.md:467
msgid ""
"**Q: Startup fails with HCCL port conflicts (address already bound). What"
" should I do?**"
msgstr "**问:启动失败,提示 HCCL 端口冲突(地址已被占用)。该怎么办?**"
#: ../../source/tutorials/models/MiniMax-M2.md:469
msgid ""
"A: Clean up old processes and restart: `pkill -f \"vllm serve "
"/models/MiniMax-M2.5\"`."
msgstr "答:清理旧进程并重启:`pkill -f \"vllm serve /models/MiniMax-M2.5\"`。"
#: ../../source/tutorials/models/MiniMax-M2.md:471
msgid "**Q: How to handle OOM or unstable startup?**"
msgstr "**问:如何处理 OOM 或启动不稳定?**"
#: ../../source/tutorials/models/MiniMax-M2.md:473
msgid ""
"A: Reduce `--max-num-seqs` and `--max-num-batched-tokens` first. If "
"needed, reduce concurrency and load-testing pressure (e.g., `max-"
"concurrency` / `num-prompts`)."
msgstr "答:首先降低 `--max-num-seqs` 和 `--max-num-batched-tokens`。如有需要,降低并发数和负载测试压力(例如,`max-concurrency` / `num-prompts`)。"
#: ../../source/tutorials/models/MiniMax-M2.md:475
msgid "**Q: Why not use cross-node tp=16?**"
msgstr "**问:为什么不使用跨节点 tp=16**"
#: ../../source/tutorials/models/MiniMax-M2.md:477
msgid ""
"A: The referenced practice noted that cross-node TP may be unstable, so "
"`tp=8, dp=2` is recommended for dual-node deployment."
msgstr "答:参考实践指出跨节点 TP 可能不稳定,因此对于双节点部署,推荐使用 `tp=8, dp=2`。"
#: ../../source/tutorials/models/MiniMax-M2.md:479
msgid "**Q: How should I choose `--reasoning-parser`?**"
msgstr "**问:应该如何选择 `--reasoning-parser`**"
#: ../../source/tutorials/models/MiniMax-M2.md:481
msgid ""
"A: This guide uses `minimax_m2_append_think` so that `<think>...</think>`"
" is kept in `content`. If you mainly rely on the reasoning semantics of "
"`/v1/responses`, consider using `--reasoning-parser minimax_m2`."
msgstr "答:本指南使用 `minimax_m2_append_think`,以便将 `<think>...</think>` 保留在 `content` 中。如果您主要依赖 `/v1/responses` 的推理语义,请考虑使用 `--reasoning-parser minimax_m2`。"
#: ../../source/tutorials/models/MiniMax-M2.md:483
msgid "**Q: Which ports must be accessible?**"
msgstr "**问:哪些端口必须可访问?**"
#: ../../source/tutorials/models/MiniMax-M2.md:485
msgid ""
"A: At minimum, expose the serving port (e.g., `20004`) and the data-"
"parallel RPC port (e.g., `2347`), and ensure the two nodes can reach each"
" other over the network."
msgstr "答:至少需要暴露服务端口(例如 `20004`)和数据并行 RPC 端口(例如 `2347`),并确保两个节点可以通过网络互相访问。"

View File

@@ -0,0 +1,265 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/PaddleOCR-VL.md:1
msgid "PaddleOCR-VL"
msgstr "PaddleOCR-VL"
#: ../../source/tutorials/models/PaddleOCR-VL.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/PaddleOCR-VL.md:5
msgid ""
"PaddleOCR-VL is a SOTA and resource-efficient model tailored for document"
" parsing. Its core component is PaddleOCR-VL-0.9B, a compact yet powerful"
" vision-language model (VLM) that integrates a NaViT-style dynamic "
"resolution visual encoder with the ERNIE-4.5-0.3B language model to "
"enable accurate element recognition."
msgstr ""
"PaddleOCR-VL 是一款专为文档解析设计的 SOTA 且资源高效的模型。其核心组件是 PaddleOCR-VL-0.9B一个紧凑而强大的视觉语言模型VLM它集成了 NaViT 风格的动态分辨率视觉编码器和 ERNIE-4.5-0.3B 语言模型,以实现精确的元素识别。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:7
msgid ""
"This document provides a detailed workflow for the complete deployment "
"and verification of the model, including supported features, environment "
"preparation, single-node deployment, and functional verification. It is "
"designed to help users quickly complete model deployment and "
"verification."
msgstr ""
"本文档提供了完整的模型部署和验证的详细工作流程,包括支持的特性、环境准备、单节点部署和功能验证。旨在帮助用户快速完成模型部署和验证。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:9
msgid "Supported Features"
msgstr "支持的特性"
#: ../../source/tutorials/models/PaddleOCR-VL.md:11
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr ""
"请参考[支持的特性](../../user_guide/support_matrix/supported_models.md)以获取模型支持的特性矩阵。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:13
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[特性指南](../../user_guide/feature_guide/index.md)以获取特性的配置。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:15
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/PaddleOCR-VL.md:17
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/PaddleOCR-VL.md:19
msgid ""
"`PaddleOCR-VL-0.9B`: [PaddleOCR-"
"VL-0.9B](https://www.modelscope.cn/models/PaddlePaddle/PaddleOCR-VL)"
msgstr ""
"`PaddleOCR-VL-0.9B`: [PaddleOCR-VL-0.9B](https://www.modelscope.cn/models/PaddlePaddle/PaddleOCR-VL)"
#: ../../source/tutorials/models/PaddleOCR-VL.md:21
msgid ""
"It is recommended to download the model weights to a local directory "
"(e.g., `./PaddleOCR-VL`) for quick access during deployment."
msgstr "建议将模型权重下载到本地目录(例如 `./PaddleOCR-VL`),以便在部署期间快速访问。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:23
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/PaddleOCR-VL.md:25
msgid "You can use our official docker image to run `PaddleOCR-VL` directly."
msgstr "您可以使用我们的官方 docker 镜像直接运行 `PaddleOCR-VL`。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:27
msgid ""
"Select an image based on your machine type and start the docker image on "
"your node, refer to [using docker](../../installation.md#set-up-using-"
"docker)."
msgstr "根据您的机器类型选择镜像并在节点上启动 docker 镜像,请参考[使用 docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:51
msgid ""
"The 310P device is supported from version 0.15.0rc1. You need to select "
"the corresponding image for installation."
msgstr "310P 设备从版本 0.15.0rc1 开始支持。您需要选择对应的镜像进行安装。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:54
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/PaddleOCR-VL.md:56
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/models/PaddleOCR-VL.md:58
msgid "Single NPU (PaddleOCR-VL)"
msgstr "单 NPU (PaddleOCR-VL)"
#: ../../source/tutorials/models/PaddleOCR-VL.md:60
msgid ""
"PaddleOCR-VL supports single-node single-card deployment on the 910B4 and"
" 310P platform. Follow these steps to start the inference service:"
msgstr "PaddleOCR-VL 支持在 910B4 和 310P 平台上进行单节点单卡部署。请按照以下步骤启动推理服务:"
#: ../../source/tutorials/models/PaddleOCR-VL.md:62
msgid ""
"Prepare model weights: Ensure the downloaded model weights are stored in "
"the `PaddleOCR-VL` directory."
msgstr "准备模型权重:确保下载的模型权重存储在 `PaddleOCR-VL` 目录中。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:63
msgid "Create and execute the deployment script (save as `deploy.sh`):"
msgstr "创建并执行部署脚本(保存为 `deploy.sh`"
#: ../../source/tutorials/models/PaddleOCR-VL.md
msgid "910B4"
msgstr "910B4"
#: ../../source/tutorials/models/PaddleOCR-VL.md:72
msgid "Run the following script to start the vLLM server on single 910B4:"
msgstr "运行以下脚本在单张 910B4 上启动 vLLM 服务器:"
#: ../../source/tutorials/models/PaddleOCR-VL.md
msgid "310P"
msgstr "310P"
#: ../../source/tutorials/models/PaddleOCR-VL.md:97
msgid "Run the following script to start the vLLM server on single 310P:"
msgstr "运行以下脚本在单张 310P 上启动 vLLM 服务器:"
#: ../../source/tutorials/models/PaddleOCR-VL.md:116
msgid ""
"The `--max_model_len` option is added to prevent errors when generating "
"the attention operator mask on the 310P device."
msgstr "添加 `--max_model_len` 选项是为了防止在 310P 设备上生成注意力算子掩码时出错。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:121
msgid "Multiple NPU (PaddleOCR-VL)"
msgstr "多 NPU (PaddleOCR-VL)"
#: ../../source/tutorials/models/PaddleOCR-VL.md:123
msgid "Single-node deployment is recommended."
msgstr "推荐单节点部署。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:125
msgid "Prefill-Decode Disaggregation"
msgstr "Prefill-Decode 解耦"
#: ../../source/tutorials/models/PaddleOCR-VL.md:127
msgid "Not supported yet."
msgstr "暂不支持。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:129
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/PaddleOCR-VL.md:131
msgid "If your service start successfully, you can see the info shown below:"
msgstr "如果您的服务启动成功,您将看到如下信息:"
#: ../../source/tutorials/models/PaddleOCR-VL.md:139
msgid ""
"Once your server is started, you can use the OpenAI API client to make "
"queries."
msgstr "服务器启动后,您可以使用 OpenAI API 客户端进行查询。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:184
msgid ""
"If you query the server successfully, you can see the info shown below "
"(client):"
msgstr "如果您成功查询服务器,您将看到如下信息(客户端):"
#: ../../source/tutorials/models/PaddleOCR-VL.md:200
msgid "Offline Inference with vLLM and PP-DocLayoutV2"
msgstr "使用 vLLM 和 PP-DocLayoutV2 进行离线推理"
#: ../../source/tutorials/models/PaddleOCR-VL.md:202
msgid ""
"In the above example, we demonstrated how to use vLLM to infer the "
"PaddleOCR-VL-0.9B model. Typically, we also need to integrate the PP-"
"DocLayoutV2 model to fully unleash the capabilities of the PaddleOCR-VL "
"model, making it more consistent with the examples provided by the "
"official PaddlePaddle documentation."
msgstr "在上面的示例中,我们演示了如何使用 vLLM 推理 PaddleOCR-VL-0.9B 模型。通常,我们还需要集成 PP-DocLayoutV2 模型,以充分发挥 PaddleOCR-VL 模型的能力,使其更符合官方 PaddlePaddle 文档提供的示例。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:205
msgid ""
"Use separate virtual environments for VLLM and PP-DocLayoutV2 to prevent "
"dependency conflicts."
msgstr "为 VLLM 和 PP-DocLayoutV2 使用独立的虚拟环境,以防止依赖冲突。"
#: ../../source/tutorials/models/PaddleOCR-VL.md
msgid "PaddlePaddle"
msgstr "PaddlePaddle"
#: ../../source/tutorials/models/PaddleOCR-VL.md:215
msgid "The 910B4 device supports inference using the PaddlePaddle framework."
msgstr "910B4 设备支持使用 PaddlePaddle 框架进行推理。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:217
msgid "Pull the PaddlePaddle-compatible CANN image"
msgstr "拉取兼容 PaddlePaddle 的 CANN 镜像"
#: ../../source/tutorials/models/PaddleOCR-VL.md:223
msgid "Start the container using the following command:"
msgstr "使用以下命令启动容器:"
#: ../../source/tutorials/models/PaddleOCR-VL.md:235
msgid ""
"Install "
"[PaddlePaddle](https://www.paddlepaddle.org.cn/install/quick?docurl=undefined)"
" and [PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR)"
msgstr ""
"安装 [PaddlePaddle](https://www.paddlepaddle.org.cn/install/quick?docurl=undefined) 和 [PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR)"
#: ../../source/tutorials/models/PaddleOCR-VL.md:246
msgid "The OpenCV component may be missing:"
msgstr "可能缺少 OpenCV 组件:"
#: ../../source/tutorials/models/PaddleOCR-VL.md:253
msgid ""
"CANN-8.0.0 does not support some versions of NumPy and OpenCV. It is "
"recommended to install the specified versions."
msgstr "CANN-8.0.0 不支持某些版本的 NumPy 和 OpenCV。建议安装指定版本。"
#: ../../source/tutorials/models/PaddleOCR-VL.md
msgid "OM inference"
msgstr "OM 推理"
#: ../../source/tutorials/models/PaddleOCR-VL.md:264
msgid ""
"The 310P device supports only the OM model inference. For details about "
"the process, see the guide provided in "
"[ModelZoo](https://gitcode.com/Ascend/ModelZoo-"
"PyTorch/tree/master/ACL_PyTorch/built-in/ocr/PP-DocLayoutV2)."
msgstr "310P 设备仅支持 OM 模型推理。有关该过程的详细信息,请参阅 [ModelZoo](https://gitcode.com/Ascend/ModelZoo-PyTorch/tree/master/ACL_PyTorch/built-in/ocr/PP-DocLayoutV2) 中提供的指南。"
#: ../../source/tutorials/models/PaddleOCR-VL.md:268
msgid ""
"Using vLLM as the backend, combined with PP-DocLayoutV2 for offline "
"inference"
msgstr "使用 vLLM 作为后端,结合 PP-DocLayoutV2 进行离线推理"

View File

@@ -0,0 +1,360 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:1
msgid "Qwen-VL-Dense(Qwen2.5VL-3B/7B, Qwen3-VL-2B/4B/8B/32B)"
msgstr "Qwen-VL-Dense (Qwen2.5VL-3B/7B, Qwen3-VL-2B/4B/8B/32B)"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:5
msgid ""
"The Qwen-VL(Vision-Language)series from Alibaba Cloud comprises a family "
"of powerful Large Vision-Language Models (LVLMs) designed for "
"comprehensive multimodal understanding. They accept images, text, and "
"bounding boxes as input, and output text and detection boxes, enabling "
"advanced functions like image detection, multi-modal dialogue, and multi-"
"image reasoning."
msgstr ""
"阿里云的Qwen-VL视觉-语言系列是一组强大的大型视觉语言模型LVLM专为全面的多模态理解而设计。它们接受图像、文本和边界框作为输入并输出文本和检测框从而实现图像检测、多模态对话和多图像推理等高级功能。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:7
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, NPU deployment, accuracy and performance evaluation."
msgstr "本文档将展示该模型的主要验证步骤包括支持的功能、功能配置、环境准备、NPU部署、精度和性能评估。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:9
msgid ""
"This tutorial uses the vLLM-Ascend `v0.11.0rc3-a3` version for "
"demonstration, showcasing the `Qwen3-VL-8B-Instruct` model as an example "
"for single NPU deployment and the `Qwen2.5-VL-32B-Instruct` model as an "
"example for multi-NPU deployment."
msgstr "本教程使用 vLLM-Ascend `v0.11.0rc3-a3` 版本进行演示,以 `Qwen3-VL-8B-Instruct` 模型为例展示单NPU部署以 `Qwen2.5-VL-32B-Instruct` 模型为例展示多NPU部署。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:11
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:13
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取模型支持的功能矩阵。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:15
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[功能指南](../../user_guide/feature_guide/index.md)以获取功能的配置信息。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:17
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:19
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:21
msgid "require 1 Atlas 800I A2 (64G × 8) node or 1 Atlas 800 A3 (64G × 16) node:"
msgstr "需要 1 个 Atlas 800I A2 (64G × 8) 节点或 1 个 Atlas 800 A3 (64G × 16) 节点:"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:23
msgid ""
"`Qwen2.5-VL-3B-Instruct`: [Download model "
"weight](https://modelscope.cn/models/Qwen/Qwen2.5-VL-3B-Instruct)"
msgstr "`Qwen2.5-VL-3B-Instruct`: [下载模型权重](https://modelscope.cn/models/Qwen/Qwen2.5-VL-3B-Instruct)"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:24
msgid ""
"`Qwen2.5-VL-7B-Instruct`: [Download model "
"weight](https://modelscope.cn/models/Qwen/Qwen2.5-VL-7B-Instruct)"
msgstr "`Qwen2.5-VL-7B-Instruct`: [下载模型权重](https://modelscope.cn/models/Qwen/Qwen2.5-VL-7B-Instruct)"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:25
msgid ""
"`Qwen2.5-VL-32B-Instruct`:[Download model "
"weight](https://modelscope.cn/models/Qwen/Qwen2.5-VL-32B-Instruct)"
msgstr "`Qwen2.5-VL-32B-Instruct`:[下载模型权重](https://modelscope.cn/models/Qwen/Qwen2.5-VL-32B-Instruct)"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:26
msgid ""
"`Qwen2.5-VL-72B-Instruct`:[Download model "
"weight](https://modelscope.cn/models/Qwen/Qwen2.5-VL-72B-Instruct)"
msgstr "`Qwen2.5-VL-72B-Instruct`:[下载模型权重](https://modelscope.cn/models/Qwen/Qwen2.5-VL-72B-Instruct)"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:27
msgid ""
"`Qwen3-VL-2B-Instruct`: [Download model "
"weight](https://modelscope.cn/models/Qwen/Qwen3-VL-2B-Instruct)"
msgstr "`Qwen3-VL-2B-Instruct`: [下载模型权重](https://modelscope.cn/models/Qwen/Qwen3-VL-2B-Instruct)"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:28
msgid ""
"`Qwen3-VL-4B-Instruct`: [Download model "
"weight](https://modelscope.cn/models/Qwen/Qwen3-VL-4B-Instruct)"
msgstr "`Qwen3-VL-4B-Instruct`: [下载模型权重](https://modelscope.cn/models/Qwen/Qwen3-VL-4B-Instruct)"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:29
msgid ""
"`Qwen3-VL-8B-Instruct`: [Download model "
"weight](https://modelscope.cn/models/Qwen/Qwen3-VL-8B-Instruct)"
msgstr "`Qwen3-VL-8B-Instruct`: [下载模型权重](https://modelscope.cn/models/Qwen/Qwen3-VL-8B-Instruct)"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:30
msgid ""
"`Qwen3-VL-32B-Instruct`: [Download model "
"weight](https://modelscope.cn/models/Qwen/Qwen3-VL-32B-Instruct)"
msgstr "`Qwen3-VL-32B-Instruct`: [下载模型权重](https://modelscope.cn/models/Qwen/Qwen3-VL-32B-Instruct)"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:32
msgid ""
"A sample Qwen2.5-VL quantization script can be found in the modelslim "
"code repository. [Qwen2.5-VL Quantization Script "
"Example](https://gitcode.com/Ascend/msit/blob/master/msmodelslim/example/multimodal_vlm/Qwen2.5-VL/README.md)"
msgstr "可以在 modelslim 代码仓库中找到 Qwen2.5-VL 的量化脚本示例。[Qwen2.5-VL 量化脚本示例](https://gitcode.com/Ascend/msit/blob/master/msmodelslim/example/multimodal_vlm/Qwen2.5-VL/README.md)"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:34
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`"
msgstr "建议将模型权重下载到多个节点的共享目录中,例如 `/root/.cache/`"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:36
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen-VL-Dense.md
msgid "single-NPU"
msgstr "单NPU"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:45
#: ../../source/tutorials/models/Qwen-VL-Dense.md:73
msgid "Run docker container:"
msgstr "运行 Docker 容器:"
#: ../../source/tutorials/models/Qwen-VL-Dense.md
msgid "multi-NPU"
msgstr "多NPU"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:101
msgid "Setup environment variables:"
msgstr "设置环境变量:"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:112
msgid ""
"`max_split_size_mb` prevents the native allocator from splitting blocks "
"larger than this size (in MB). This can reduce fragmentation and may "
"allow some borderline workloads to complete without running out of "
"memory. You can find more details "
"[<u>here</u>](https://www.hiascend.com/document/detail/zh/CANNCommunityEdition/800alpha003/apiref/envref/envref_07_0061.html)."
msgstr ""
"`max_split_size_mb` 可防止原生分配器拆分大于此大小(以 MB 为单位)的内存块。这可以减少内存碎片,并可能使一些临界工作负载在内存耗尽前完成。您可以在"
"[<u>此处</u>](https://www.hiascend.com/document/detail/zh/CANNCommunityEdition/800alpha003/apiref/envref/envref_07_0061.html)找到更多详细信息。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:115
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:117
msgid "Offline Inference"
msgstr "离线推理"
#: ../../source/tutorials/models/Qwen-VL-Dense.md
msgid "Qwen3-VL-8B-Instruct"
msgstr "Qwen3-VL-8B-Instruct"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:126
msgid "Run the following script to execute offline inference on single-NPU:"
msgstr "运行以下脚本在单NPU上执行离线推理"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:191
#: ../../source/tutorials/models/Qwen-VL-Dense.md:287
msgid "If you run this script successfully, you can see the info shown below:"
msgstr "如果脚本运行成功,您将看到如下信息:"
#: ../../source/tutorials/models/Qwen-VL-Dense.md
msgid "Qwen2.5-VL-32B-Instruct"
msgstr "Qwen2.5-VL-32B-Instruct"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:221
msgid "Run the following script to execute offline inference on multi-NPU:"
msgstr "运行以下脚本在多NPU上执行离线推理"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:312
msgid "Online Serving"
msgstr "在线服务"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:321
msgid "Run docker container to start the vLLM server on single-NPU:"
msgstr "运行 Docker 容器以在单NPU上启动 vLLM 服务器:"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:332
msgid ""
"Add `--max_model_len` option to avoid ValueError that the Qwen3-VL-8B-"
"Instruct model's max seq len (256000) is larger than the maximum number "
"of tokens that can be stored in KV cache. This will differ with different"
" NPU series based on the HBM size. Please modify the value according to a"
" suitable value for your NPU series."
msgstr ""
"添加 `--max_model_len` 选项以避免 ValueError该错误提示 Qwen3-VL-8B-Instruct 模型的最大序列长度256000大于 KV 缓存可存储的最大令牌数。此值因不同 NPU 系列的 HBM 大小而异。请根据您 NPU 系列的合适值修改此值。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:335
#: ../../source/tutorials/models/Qwen-VL-Dense.md:422
msgid "If your service start successfully, you can see the info shown below:"
msgstr "如果服务启动成功,您将看到如下信息:"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:343
#: ../../source/tutorials/models/Qwen-VL-Dense.md:430
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "服务器启动后,您可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:360
#: ../../source/tutorials/models/Qwen-VL-Dense.md:447
msgid ""
"If you query the server successfully, you can see the info shown below "
"(client):"
msgstr "如果成功查询服务器,您将看到如下信息(客户端):"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:366
#: ../../source/tutorials/models/Qwen-VL-Dense.md:453
msgid "Logs of the vllm server:"
msgstr "vllm 服务器的日志:"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:381
msgid "Run docker container to start the vLLM server on multi-NPU:"
msgstr "运行 Docker 容器以在多NPU上启动 vLLM 服务器:"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:419
msgid ""
"Add `--max_model_len` option to avoid ValueError that the Qwen2.5-VL-32B-"
"Instruct model's max_model_len (128000) is larger than the maximum number"
" of tokens that can be stored in KV cache. This will differ with "
"different NPU series base on the HBM size. Please modify the value "
"according to a suitable value for your NPU series."
msgstr ""
"添加 `--max_model_len` 选项以避免 ValueError该错误提示 Qwen2.5-VL-32B-Instruct 模型的最大模型长度128000大于 KV 缓存可存储的最大令牌数。此值因不同 NPU 系列的 HBM 大小而异。请根据您 NPU 系列的合适值修改此值。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:468
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:470
msgid "Using Language Model Evaluation Harness"
msgstr "使用 Language Model Evaluation Harness"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:472
msgid ""
"The accuracy of some models is already within our CI monitoring scope, "
"including:"
msgstr "部分模型的精度已纳入我们的 CI 监控范围,包括:"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:474
msgid "`Qwen2.5-VL-7B-Instruct`"
msgstr "`Qwen2.5-VL-7B-Instruct`"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:475
msgid "`Qwen3-VL-8B-Instruct`"
msgstr "`Qwen3-VL-8B-Instruct`"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:484
msgid ""
"As an example, take the `mmmu_val` dataset as a test dataset, and run "
"accuracy evaluation of `Qwen3-VL-8B-Instruct` in offline mode."
msgstr "以 `mmmu_val` 数据集作为测试数据集为例,在离线模式下运行 `Qwen3-VL-8B-Instruct` 的精度评估。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:486
#: ../../source/tutorials/models/Qwen-VL-Dense.md:517
msgid ""
"Refer to [Using "
"lm_eval](../../developer_guide/evaluation/using_lm_eval.md) for more "
"details on `lm_eval` installation."
msgstr "有关 `lm_eval` 安装的更多详细信息,请参考[使用 lm_eval](../../developer_guide/evaluation/using_lm_eval.md)。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:492
#: ../../source/tutorials/models/Qwen-VL-Dense.md:523
msgid "Run `lm_eval` to execute the accuracy evaluation."
msgstr "运行 `lm_eval` 以执行精度评估。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:505
msgid ""
"After execution, you can get the result, here is the result of `Qwen3-VL-"
"8B-Instruct` in `vllm-ascend:0.11.0rc3` for reference only."
msgstr "执行后,您将获得结果。以下是 `vllm-ascend:0.11.0rc3` 中 `Qwen3-VL-8B-Instruct` 的结果,仅供参考。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:515
msgid ""
"As an example, take the `mmmu_val` dataset as a test dataset, and run "
"accuracy evaluation of `Qwen2.5-VL-32B-Instruct` in offline mode."
msgstr "以 `mmmu_val` 数据集作为测试数据集为例,在离线模式下运行 `Qwen2.5-VL-32B-Instruct` 的精度评估。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:535
msgid ""
"After execution, you can get the result, here is the result of `Qwen2.5"
"-VL-32B-Instruct` in `vllm-ascend:0.11.0rc3` for reference only."
msgstr "执行后,您将获得结果。以下是 `vllm-ascend:0.11.0rc3` 中 `Qwen2.5-VL-32B-Instruct` 的结果,仅供参考。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:543
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:545
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM Benchmark"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:547
msgid ""
"Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) "
"for more details."
msgstr "更多详细信息,请参考 [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:549
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 有三个子命令:"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:551
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`: 对单批请求的延迟进行基准测试。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:552
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`: 对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:553
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`: 对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:555
msgid ""
"The performance evaluation must be conducted in an online mode. Take the "
"`serve` as an example. Run the code as follows."
msgstr "性能评估必须在在线模式下进行。以 `serve` 为例。按如下方式运行代码。"
#: ../../source/tutorials/models/Qwen-VL-Dense.md:578
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您将获得性能评估结果。"

View File

@@ -0,0 +1,279 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen2.5-7B.md:1
msgid "Qwen2.5-7B"
msgstr "Qwen2.5-7B"
#: ../../source/tutorials/models/Qwen2.5-7B.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen2.5-7B.md:5
msgid ""
"Qwen2.5-7B-Instruct is the flagship instruction-tuned variant of Alibaba "
"Clouds Qwen 2.5 LLM series. It supports a maximum context window of "
"128K, enables generation of up to 8K tokens, and delivers enhanced "
"capabilities in multilingual processing, instruction following, "
"programming, mathematical computation, and structured data handling."
msgstr ""
"Qwen2.5-7B-Instruct 是阿里云 Qwen 2.5 大语言模型系列的旗舰指令调优变体。它支持最大 128K 的上下文窗口,能够生成最多 8K 个令牌,并在多语言处理、指令遵循、编程、数学计算和结构化数据处理方面提供增强的能力。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:7
msgid ""
"This document details the complete deployment and verification workflow "
"for the model, including supported features, environment preparation, "
"single-node deployment, functional verification, accuracy and performance"
" evaluation, and troubleshooting of common issues. It is designed to help"
" users quickly complete model deployment and validation."
msgstr ""
"本文档详细介绍了该模型的完整部署和验证工作流程,包括支持的功能、环境准备、单节点部署、功能验证、准确性和性能评估以及常见问题排查。旨在帮助用户快速完成模型部署和验证。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:9
msgid "The `Qwen2.5-7B-Instruct` model was supported since `vllm-ascend:v0.9.0`."
msgstr "`Qwen2.5-7B-Instruct` 模型自 `vllm-ascend:v0.9.0` 版本起获得支持。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:11
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/Qwen2.5-7B.md:13
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取模型支持的功能矩阵。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:15
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[功能指南](../../user_guide/feature_guide/index.md)以获取功能的配置信息。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:17
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen2.5-7B.md:19
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen2.5-7B.md:21
msgid ""
"`Qwen2.5-7B-Instruct`(BF16 version): require 1 Atlas 910B4 (32G × 1) "
"card. [Download model weight](https://modelscope.cn/models/Qwen/Qwen2.5"
"-7B-Instruct)"
msgstr ""
"`Qwen2.5-7B-Instruct`BF16 版本):需要 1 张 Atlas 910B432G × 1卡。[下载模型权重](https://modelscope.cn/models/Qwen/Qwen2.5-7B-Instruct)"
#: ../../source/tutorials/models/Qwen2.5-7B.md:23
msgid ""
"It is recommended to download the model weights to a local directory "
"(e.g., `./Qwen2.5-7B-Instruct/`) for quick access during deployment."
msgstr "建议将模型权重下载到本地目录(例如 `./Qwen2.5-7B-Instruct/`),以便在部署期间快速访问。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:25
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen2.5-7B.md:27
msgid ""
"You can use our official docker image and install extra operator for "
"supporting `Qwen2.5-7B-Instruct`."
msgstr "您可以使用我们的官方 docker 镜像,并安装额外的算子以支持 `Qwen2.5-7B-Instruct`。"
#: ../../source/tutorials/models/Qwen2.5-7B.md
msgid "A3 series"
msgstr "A3 系列"
#: ../../source/tutorials/models/Qwen2.5-7B.md:36
#: ../../source/tutorials/models/Qwen2.5-7B.md:64
msgid "Start the docker image on your each node."
msgstr "在您的每个节点上启动 docker 镜像。"
#: ../../source/tutorials/models/Qwen2.5-7B.md
msgid "A2 series"
msgstr "A2 系列"
#: ../../source/tutorials/models/Qwen2.5-7B.md:90
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen2.5-7B.md:92
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/models/Qwen2.5-7B.md:94
msgid ""
"Qwen2.5-7B-Instruct supports single-node single-card deployment on the "
"910B4 platform. Follow these steps to start the inference service:"
msgstr "Qwen2.5-7B-Instruct 支持在 910B4 平台上进行单节点单卡部署。请按照以下步骤启动推理服务:"
#: ../../source/tutorials/models/Qwen2.5-7B.md:96
msgid ""
"Prepare model weights: Ensure the downloaded model weights are stored in "
"the `./Qwen2.5-7B-Instruct/` directory."
msgstr "准备模型权重:确保下载的模型权重存储在 `./Qwen2.5-7B-Instruct/` 目录中。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:97
msgid "Create and execute the deployment script (save as `deploy.sh`):"
msgstr "创建并执行部署脚本(保存为 `deploy.sh`"
#: ../../source/tutorials/models/Qwen2.5-7B.md:112
msgid "Multi-node Deployment"
msgstr "多节点部署"
#: ../../source/tutorials/models/Qwen2.5-7B.md:114
msgid "Single-node deployment is recommended."
msgstr "推荐使用单节点部署。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:116
msgid "Prefill-Decode Disaggregation"
msgstr "预填充-解码分离"
#: ../../source/tutorials/models/Qwen2.5-7B.md:118
msgid "Not supported yet."
msgstr "暂不支持。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:120
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/Qwen2.5-7B.md:122
msgid "After starting the service, verify functionality using a `curl` request:"
msgstr "启动服务后,使用 `curl` 请求验证功能:"
#: ../../source/tutorials/models/Qwen2.5-7B.md:135
msgid ""
"A valid response (e.g., `\"Beijing is a vibrant and historic capital "
"city\"`) indicates successful deployment."
msgstr "有效的响应(例如 `\"Beijing is a vibrant and historic capital city\"`)表明部署成功。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:137
msgid "Accuracy Evaluation"
msgstr "准确性评估"
#: ../../source/tutorials/models/Qwen2.5-7B.md:139
#: ../../source/tutorials/models/Qwen2.5-7B.md:151
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/models/Qwen2.5-7B.md:141
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参考[使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:143
msgid ""
"Results and logs are saved to `benchmark/outputs/default/`. A sample "
"accuracy report is shown below:"
msgstr "结果和日志保存在 `benchmark/outputs/default/` 中。示例如下:"
#: ../../source/tutorials/models/Qwen2.5-7B.md:66
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/models/Qwen2.5-7B.md:66
msgid "version"
msgstr "版本"
#: ../../source/tutorials/models/Qwen2.5-7B.md:66
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/models/Qwen2.5-7B.md:66
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/models/Qwen2.5-7B.md:66
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/models/Qwen2.5-7B.md:66
msgid "gsm8k"
msgstr "gsm8k"
#: ../../source/tutorials/models/Qwen2.5-7B.md:66
msgid "-"
msgstr "-"
#: ../../source/tutorials/models/Qwen2.5-7B.md:66
msgid "accuracy"
msgstr "准确率"
#: ../../source/tutorials/models/Qwen2.5-7B.md:66
msgid "gen"
msgstr "生成"
#: ../../source/tutorials/models/Qwen2.5-7B.md:66
msgid "75.00"
msgstr "75.00"
#: ../../source/tutorials/models/Qwen2.5-7B.md:149
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen2.5-7B.md:153
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参考[使用 AISBench 进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:155
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM Benchmark"
#: ../../source/tutorials/models/Qwen2.5-7B.md:157
msgid "Run performance evaluation of `Qwen2.5-7B-Instruct` as an example."
msgstr "以运行 `Qwen2.5-7B-Instruct` 的性能评估为例。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:159
msgid ""
"Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) "
"for more details."
msgstr "更多详情请参考 [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:161
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 有三个子命令:"
#: ../../source/tutorials/models/Qwen2.5-7B.md:163
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批请求的延迟进行基准测试。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:164
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:165
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:167
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。按如下方式运行代码。"
#: ../../source/tutorials/models/Qwen2.5-7B.md:180
msgid "After several minutes, you can get the performance evaluation result."
msgstr "几分钟后,您将获得性能评估结果。"

View File

@@ -0,0 +1,301 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:1
msgid "Qwen2.5-Omni-7B"
msgstr "Qwen2.5-Omni-7B"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:5
msgid ""
"Qwen2.5-Omni is an end-to-end multimodal model designed to perceive "
"diverse modalities, including text, images, audio, and video, while "
"simultaneously generating text and natural speech responses in a "
"streaming manner."
msgstr "Qwen2.5-Omni 是一个端到端的多模态模型,旨在感知多种模态,包括文本、图像、音频和视频,同时以流式方式生成文本和自然语音响应。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:7
msgid ""
"The `Qwen2.5-Omni` model was supported since `vllm-ascend:v0.11.0rc0`. "
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, single-NPU and multi-NPU deployment, accuracy and "
"performance evaluation."
msgstr "`Qwen2.5-Omni` 模型自 `vllm-ascend:v0.11.0rc0` 版本起获得支持。本文档将展示该模型的主要验证步骤包括支持的特性、特性配置、环境准备、单NPU和多NPU部署、精度和性能评估。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:9
msgid "Supported Features"
msgstr "支持的特性"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:11
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的特性](../../user_guide/support_matrix/supported_models.md)以获取该模型支持的特性矩阵。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:13
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[特性指南](../../user_guide/feature_guide/index.md)以获取特性的配置方法。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:15
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:17
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:19
msgid ""
"`Qwen2.5-Omni-3B`(BF16): [Download model "
"weight](https://modelscope.cn/models/Qwen/Qwen2.5-Omni-3B)"
msgstr "`Qwen2.5-Omni-3B`(BF16): [下载模型权重](https://modelscope.cn/models/Qwen/Qwen2.5-Omni-3B)"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:20
msgid ""
"`Qwen2.5-Omni-7B`(BF16): [Download model "
"weight](https://modelscope.cn/models/Qwen/Qwen2.5-Omni-7B)"
msgstr "`Qwen2.5-Omni-7B`(BF16): [下载模型权重](https://modelscope.cn/models/Qwen/Qwen2.5-Omni-7B)"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:22
msgid "Following examples use the 7B version by default."
msgstr "以下示例默认使用 7B 版本。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:24
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:26
msgid "You can use our official docker image to run `Qwen2.5-Omni` directly."
msgstr "您可以使用我们的官方 docker 镜像直接运行 `Qwen2.5-Omni`。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:28
msgid ""
"Select an image based on your machine type and start the docker image on "
"your node, refer to [using docker](../../installation.md#set-up-using-"
"docker)."
msgstr "根据您的机器类型选择镜像并在节点上启动 docker 镜像,请参考[使用 docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:65
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:67
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:69
msgid "Single NPU (Qwen2.5-Omni-7B)"
msgstr "单 NPU (Qwen2.5-Omni-7B)"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:72
msgid ""
"The **environment variable** `LOCAL_MEDIA_PATH` which **allows** API "
"requests to read local images or videos from directories specified by the"
" server file system. Please note this is a security risk. Should only be "
"enabled in trusted environments."
msgstr "**环境变量** `LOCAL_MEDIA_PATH` **允许** API 请求从服务器文件系统指定的目录读取本地图像或视频。请注意,这存在安全风险。应仅在受信任的环境中启用。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:92
msgid ""
"Now vllm-ascend docker image should contain vllm[audio] build part, if "
"you encounter *audio not supported issue* by any chance, please re-build "
"vllm with [audio] flag."
msgstr "当前 vllm-ascend docker 镜像应包含 vllm[audio] 构建部分,如果您遇到*音频不支持的问题*,请使用 [audio] 标志重新构建 vllm。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:100
msgid ""
"`--allowed-local-media-path` is optional, only set it if you need infer "
"model with local media file."
msgstr "`--allowed-local-media-path` 是可选的,仅在需要使用本地媒体文件进行模型推理时设置。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:102
msgid ""
"`--gpu-memory-utilization` should not be set manually unless you know "
"what this parameter does."
msgstr "`--gpu-memory-utilization` 不应手动设置,除非您了解此参数的作用。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:104
msgid "Multiple NPU (Qwen2.5-Omni-7B)"
msgstr "多 NPU (Qwen2.5-Omni-7B)"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:123
msgid ""
"`--tensor_parallel_size` no need to set for this 7B model, but if you "
"really need tensor parallel, tp size can be one of `1/2/4`."
msgstr "对于此 7B 模型,无需设置 `--tensor_parallel_size`但如果确实需要张量并行tp 大小可以是 `1/2/4` 之一。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:125
msgid "Prefill-Decode Disaggregation"
msgstr "预填充-解码分离"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:127
msgid "Not supported yet."
msgstr "暂不支持。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:129
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:131
msgid "If your service **starts** successfully, you can see the info shown below:"
msgstr "如果您的服务**启动**成功,您可以看到如下所示的信息:"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:139
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "一旦您的服务器启动,您可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:167
msgid ""
"If you query the server successfully, you can see the info shown below "
"(client):"
msgstr "如果您成功查询服务器,您可以看到如下所示的信息(客户端):"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:173
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:175
msgid "Qwen2.5-Omni on vllm-ascend has been tested on AISBench."
msgstr "vllm-ascend 上的 Qwen2.5-Omni 已在 AISBench 上进行了测试。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:177
#: ../../source/tutorials/models/Qwen2.5-Omni.md:190
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:179
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参考[使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:181
msgid ""
"After execution, you can get the result, here is the result of `Qwen2.5"
"-Omni-7B` with `vllm-ascend:0.11.0rc0` for reference only."
msgstr "执行后,您可以获得结果,以下是 `Qwen2.5-Omni-7B` 在 `vllm-ascend:0.11.0rc0` 上的结果,仅供参考。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:91
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:91
msgid "platform"
msgstr "平台"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:91
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:91
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:91
msgid "vllm-api-stream-chat"
msgstr "vllm-api-stream-chat"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:91
msgid "textVQA"
msgstr "textVQA"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:91
msgid "A2"
msgstr "A2"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:91
msgid "accuracy"
msgstr "精度"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:91
msgid "gen_base64"
msgstr "gen_base64"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:91
msgid "83.47"
msgstr "83.47"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:91
msgid "A3"
msgstr "A3"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:91
msgid "84.04"
msgstr "84.04"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:188
msgid "Performance Evaluation"
msgstr "性能评估"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:192
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参考[使用 AISBench 进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:194
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM Benchmark"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:196
msgid "Run performance evaluation of `Qwen2.5-Omni-7B` as an example."
msgstr "以运行 `Qwen2.5-Omni-7B` 的性能评估为例。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:198
msgid ""
"Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) "
"for more details."
msgstr "更多详情请参考 [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:200
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 有三个子命令:"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:202
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批请求的延迟进行基准测试。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:203
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:204
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:206
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。按如下方式运行代码。"
#: ../../source/tutorials/models/Qwen2.5-Omni.md:212
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您就可以获得性能评估结果。"

View File

@@ -0,0 +1,665 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:1
msgid "Qwen3-235B-A22B"
msgstr "Qwen3-235B-A22B"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:5
msgid ""
"Qwen3 is the latest generation of large language models in Qwen series, "
"offering a comprehensive suite of dense and mixture-of-experts (MoE) "
"models. Built upon extensive training, Qwen3 delivers groundbreaking "
"advancements in reasoning, instruction-following, agent capabilities, and"
" multilingual support."
msgstr ""
"Qwen3 是 Qwen 系列最新一代的大语言模型提供了一套完整的稠密模型和专家混合模型。基于广泛的训练Qwen3 在推理、指令遵循、智能体能力和多语言支持方面实现了突破性进展。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:7
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, single-node and multi-node deployment, accuracy and "
"performance evaluation."
msgstr "本文档将展示该模型的主要验证步骤,包括支持的特性、特性配置、环境准备、单节点与多节点部署、精度和性能评估。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:9
msgid "The `Qwen3-235B-A22B` model is first supported in `vllm-ascend:v0.8.4rc2`."
msgstr "`Qwen3-235B-A22B` 模型首次在 `vllm-ascend:v0.8.4rc2` 版本中得到支持。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:11
msgid "Supported Features"
msgstr "支持的特性"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:13
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的特性](../../user_guide/support_matrix/supported_models.md)以获取该模型的支持特性矩阵。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:15
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[特性指南](../../user_guide/feature_guide/index.md)以获取特性的配置方法。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:17
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:19
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:21
msgid ""
"`Qwen3-235B-A22B`(BF16 version): require 1 Atlas 800 A3 (64G × 16) node, "
"1 Atlas 800 A2 (64G × 8) node or 2 Atlas 800 A2(32G × 8)nodes. [Download "
"model weight](https://www.modelscope.cn/models/Qwen/Qwen3-235B-A22B)"
msgstr ""
"`Qwen3-235B-A22B`(BF16 版本):需要 1 个 Atlas 800 A3 (64G × 16) 节点、1 个 Atlas 800 A2 (64G × 8) 节点或 2 个 Atlas 800 A2(32G × 8) 节点。[下载模型权重](https://www.modelscope.cn/models/Qwen/Qwen3-235B-A22B)"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:22
msgid ""
"`Qwen3-235B-A22B-w8a8`(Quantized version): require 1 Atlas 800 A3 (64G × "
"16) node or 1 Atlas 800 A2 (64G × 8) node or 2 Atlas 800 A2(32G × "
"8)nodes. [Download model weight](https://modelscope.cn/models/vllm-"
"ascend/Qwen3-235B-A22B-W8A8)"
msgstr ""
"`Qwen3-235B-A22B-w8a8`(量化版本):需要 1 个 Atlas 800 A3 (64G × 16) 节点、1 个 Atlas 800 A2 (64G × 8) 节点或 2 个 Atlas 800 A2(32G × 8) 节点。[下载模型权重](https://modelscope.cn/models/vllm-ascend/Qwen3-235B-A22B-W8A8)"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:24
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`."
msgstr "建议将模型权重下载到多节点的共享目录中,例如 `/root/.cache/`。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:26
msgid "Verify Multi-node Communication(Optional)"
msgstr "验证多节点通信(可选)"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:28
msgid ""
"If you want to deploy multi-node environment, you need to verify multi-"
"node communication according to [verify multi-node communication "
"environment](../../installation.md#verify-multi-node-communication)."
msgstr "如果您想部署多节点环境,需要根据[验证多节点通信环境](../../installation.md#verify-multi-node-communication)来验证多节点通信。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:30
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md
msgid "Use docker image"
msgstr "使用 Docker 镜像"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:36
msgid ""
"For example, using images `quay.io/ascend/vllm-ascend:v0.11.0rc2`(for "
"Atlas 800 A2) and `quay.io/ascend/vllm-ascend:v0.11.0rc2-a3`(for Atlas "
"800 A3)."
msgstr "例如,使用镜像 `quay.io/ascend/vllm-ascend:v0.11.0rc2`(适用于 Atlas 800 A2和 `quay.io/ascend/vllm-ascend:v0.11.0rc2-a3`(适用于 Atlas 800 A3。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:38
msgid ""
"Select an image based on your machine type and start the docker image on "
"your node, refer to [using docker](../../installation.md#set-up-using-"
"docker)."
msgstr "根据您的机器类型选择镜像并在节点上启动 Docker 容器,请参考[使用 Docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md
msgid "Build from source"
msgstr "从源码构建"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:78
msgid "You can build all from source."
msgstr "您可以从源码构建所有组件。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:80
msgid ""
"Install `vllm-ascend`, refer to [set up using "
"python](../../installation.md#set-up-using-python)."
msgstr "安装 `vllm-ascend`,请参考[使用 Python 设置](../../installation.md#set-up-using-python)。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:84
msgid ""
"If you want to deploy multi-node environment, you need to set up "
"environment on each node."
msgstr "如果您想部署多节点环境,需要在每个节点上设置环境。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:86
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:88
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:90
msgid ""
"`Qwen3-235B-A22B` and `Qwen3-235B-A22B-w8a8` can both be deployed on 1 "
"Atlas 800 A3(64G*16), 1 Atlas 800 A2(64G*8). Quantized version need to "
"start with parameter `--quantization ascend`."
msgstr "`Qwen3-235B-A22B` 和 `Qwen3-235B-A22B-w8a8` 都可以部署在 1 个 Atlas 800 A3(64G*16) 或 1 个 Atlas 800 A2(64G*8) 上。量化版本需要使用参数 `--quantization ascend` 启动。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:93
msgid "Run the following script to execute online 128k inference."
msgstr "运行以下脚本来执行在线 128k 推理。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:126
msgid "**Notice:**"
msgstr "**注意:**"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:128
msgid ""
"[Qwen3-235B-A22B](https://huggingface.co/Qwen/Qwen3-235B-A22B#processing-"
"long-texts) originally only supports 40960 "
"context(max_position_embeddings). If you want to use it and its related "
"quantization weights to run long seqs (such as 128k context), it is "
"required to use yarn rope-scaling technique."
msgstr ""
"[Qwen3-235B-A22B](https://huggingface.co/Qwen/Qwen3-235B-A22B#processing-long-texts) 原本仅支持 40960 上下文长度max_position_embeddings。如果您想使用它及其相关的量化权重来运行长序列例如 128k 上下文),需要使用 yarn rope-scaling 技术。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:129
#, python-brace-format
msgid ""
"For vLLM version same as or new than `v0.12.0`, use parameter: `--hf-"
"overrides '{\"rope_parameters\": "
"{\"rope_type\":\"yarn\",\"rope_theta\":1000000,\"factor\":4,\"original_max_position_embeddings\":32768}}'"
" \\`."
msgstr ""
"对于 `v0.12.0` 及以上版本的 vLLM使用参数`--hf-overrides '{\"rope_parameters\": "
"{\"rope_type\":\"yarn\",\"rope_theta\":1000000,\"factor\":4,\"original_max_position_embeddings\":32768}}' \\`。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:130
#, python-brace-format
msgid ""
"For vllm version below `v0.12.0`, use parameter: `--rope_scaling "
"'{\"rope_type\":\"yarn\",\"factor\":4,\"original_max_position_embeddings\":32768}'"
" \\`. If you are using weights like [Qwen3-235B-A22B-"
"Instruct-2507](https://huggingface.co/Qwen/Qwen3-235B-A22B-Instruct-2507)"
" which originally supports long contexts, there is no need to add this "
"parameter."
msgstr ""
"对于 `v0.12.0` 以下版本的 vLLM使用参数`--rope_scaling "
"'{\"rope_type\":\"yarn\",\"factor\":4,\"original_max_position_embeddings\":32768}' \\`。如果您使用的是像 [Qwen3-235B-A22B-Instruct-2507](https://huggingface.co/Qwen/Qwen3-235B-A22B-Instruct-2507) 这样原本就支持长上下文的权重,则无需添加此参数。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:133
msgid "The parameters are explained as follows:"
msgstr "参数解释如下:"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:135
msgid ""
"`--data-parallel-size` 1 and `--tensor-parallel-size` 8 are common "
"settings for data parallelism (DP) and tensor parallelism (TP) sizes."
msgstr "`--data-parallel-size` 1 和 `--tensor-parallel-size` 8 是数据并行DP和张量并行TP大小的常见设置。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:136
msgid ""
"`--max-model-len` represents the context length, which is the maximum "
"value of the input plus output for a single request."
msgstr "`--max-model-len` 表示上下文长度,即单个请求的输入加输出的最大值。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:137
msgid ""
"`--max-num-seqs` indicates the maximum number of requests that each DP "
"group is allowed to process. If the number of requests sent to the "
"service exceeds this limit, the excess requests will remain in a waiting "
"state and will not be scheduled. Note that the time spent in the waiting "
"state is also counted in metrics such as TTFT and TPOT. Therefore, when "
"testing performance, it is generally recommended that `--max-num-seqs` * "
"`--data-parallel-size` >= the actual total concurrency."
msgstr ""
"`--max-num-seqs` 表示每个 DP 组允许处理的最大请求数。如果发送到服务的请求数超过此限制,超出的请求将保持在等待状态,不会被调度。请注意,在等待状态所花费的时间也会计入 TTFT 和 TPOT 等指标。因此,在测试性能时,通常建议 `--max-num-seqs` * `--data-parallel-size` >= 实际总并发数。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:138
msgid ""
"`--max-num-batched-tokens` represents the maximum number of tokens that "
"the model can process in a single step. Currently, vLLM v1 scheduling "
"enables ChunkPrefill/SplitFuse by default, which means:"
msgstr "`--max-num-batched-tokens` 表示模型在单步中可以处理的最大 token 数。目前vLLM v1 调度默认启用 ChunkPrefill/SplitFuse这意味着"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:139
msgid ""
"(1) If the input length of a request is greater than `--max-num-batched-"
"tokens`, it will be divided into multiple rounds of computation according"
" to `--max-num-batched-tokens`;"
msgstr "(1) 如果一个请求的输入长度大于 `--max-num-batched-tokens`,它将根据 `--max-num-batched-tokens` 被分成多轮计算;"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:140
msgid ""
"(2) Decode requests are prioritized for scheduling, and prefill requests "
"are scheduled only if there is available capacity."
msgstr "(2) 解码请求优先被调度,只有在有可用容量时才会调度预填充请求。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:141
msgid ""
"Generally, if `--max-num-batched-tokens` is set to a larger value, the "
"overall latency will be lower, but the pressure on GPU memory (activation"
" value usage) will be greater."
msgstr "通常,如果将 `--max-num-batched-tokens` 设置为较大的值,整体延迟会更低,但 GPU 内存(激活值使用)的压力会更大。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:142
msgid ""
"`--gpu-memory-utilization` represents the proportion of HBM that vLLM "
"will use for actual inference. Its essential function is to calculate the"
" available kv_cache size. During the warm-up phase (referred to as "
"profile run in vLLM), vLLM records the peak GPU memory usage during an "
"inference process with an input size of `--max-num-batched-tokens`. The "
"available kv_cache size is then calculated as: `--gpu-memory-utilization`"
" * HBM size - peak GPU memory usage. Therefore, the larger the value of "
"`--gpu-memory-utilization`, the more kv_cache can be used. However, since"
" the GPU memory usage during the warm-up phase may differ from that "
"during actual inference (e.g., due to uneven EP load), setting `--gpu-"
"memory-utilization` too high may lead to OOM (Out of Memory) issues "
"during actual inference. The default value is `0.9`."
msgstr ""
"`--gpu-memory-utilization` 表示 vLLM 将用于实际推理的 HBM 比例。其核心功能是计算可用的 kv_cache 大小。在预热阶段(在 vLLM 中称为 profile runvLLM 会记录输入大小为 `--max-num-batched-tokens` 的推理过程中的峰值 GPU 内存使用量。然后,可用的 kv_cache 大小计算为:`--gpu-memory-utilization` * HBM 大小 - 峰值 GPU 内存使用量。因此,`--gpu-memory-utilization` 的值越大,可以使用的 kv_cache 就越多。然而,由于预热阶段的 GPU 内存使用量可能与实际推理期间不同(例如,由于 EP 负载不均),将 `--gpu-memory-utilization` 设置得过高可能会导致实际推理期间出现 OOM内存不足问题。默认值为 `0.9`。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:143
msgid ""
"`--enable-expert-parallel` indicates that EP is enabled. Note that vLLM "
"does not support a mixed approach of ETP and EP; that is, MoE can either "
"use pure EP or pure TP."
msgstr "`--enable-expert-parallel` 表示启用了 EP。请注意vLLM 不支持 ETP 和 EP 的混合方法也就是说MoE 可以使用纯 EP 或纯 TP。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:144
msgid ""
"`--no-enable-prefix-caching` indicates that prefix caching is disabled. "
"To enable it, remove this option."
msgstr "`--no-enable-prefix-caching` 表示前缀缓存被禁用。要启用它,请移除此选项。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:145
msgid ""
"`--quantization` \"ascend\" indicates that quantization is used. To "
"disable quantization, remove this option."
msgstr "`--quantization` \"ascend\" 表示使用了量化。要禁用量化,请移除此选项。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:146
msgid ""
"`--compilation-config` contains configurations related to the aclgraph "
"graph mode. The most significant configurations are \"cudagraph_mode\" "
"and \"cudagraph_capture_sizes\", which have the following meanings: "
"\"cudagraph_mode\": represents the specific graph mode. Currently, "
"\"PIECEWISE\" and \"FULL_DECODE_ONLY\" are supported. The graph mode is "
"mainly used to reduce the cost of operator dispatch. Currently, "
"\"FULL_DECODE_ONLY\" is recommended."
msgstr ""
"`--compilation-config` 包含与 aclgraph 图模式相关的配置。最重要的配置是 \"cudagraph_mode\" 和 \"cudagraph_capture_sizes\",其含义如下:\"cudagraph_mode\":表示特定的图模式。目前支持 \"PIECEWISE\" 和 \"FULL_DECODE_ONLY\"。图模式主要用于降低算子调度的开销。目前推荐使用 \"FULL_DECODE_ONLY\"。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:148
msgid ""
"\"cudagraph_capture_sizes\": represents different levels of graph modes. "
"The default value is [1, 2, 4, 8, 16, 24, 32, 40,..., `--max-num-seqs`]. "
"In the graph mode, the input for graphs at different levels is fixed, and"
" inputs between levels are automatically padded to the next level. "
"Currently, the default setting is recommended. Only in some scenarios is "
"it necessary to set this separately to achieve optimal performance."
msgstr ""
"\"cudagraph_capture_sizes\":表示不同级别的图模式。默认值为 [1, 2, 4, 8, 16, 24, 32, 40,..., `--max-num-seqs`]。在图模式下,不同级别图的输入是固定的,级别之间的输入会自动填充到下一个级别。目前推荐使用默认设置。只有在某些场景下,才需要单独设置此参数以达到最佳性能。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:149
msgid ""
"`export VLLM_ASCEND_ENABLE_FLASHCOMM1=1` indicates that Flashcomm1 "
"optimization is enabled. Currently, this optimization is only supported "
"for MoE in scenarios where tp_size > 1."
msgstr "`export VLLM_ASCEND_ENABLE_FLASHCOMM1=1` 表示启用了 Flashcomm1 优化。目前,此优化仅在 tp_size > 1 的场景下对 MoE 支持。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:151
msgid "Multi-node Deployment with MP (Recommended)"
msgstr "使用 MP 进行多节点部署(推荐)"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:153
msgid ""
"Assume you have Atlas 800 A3 (64G*16) nodes (or 2* A2), and want to "
"deploy the `Qwen3-VL-235B-A22B-Instruct` model across multiple nodes."
msgstr "假设您有 Atlas 800 A3 (64G*16) 节点(或 2* A2并希望跨多个节点部署 `Qwen3-VL-235B-A22B-Instruct` 模型。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:155
msgid "Node 0"
msgstr "节点 0"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:197
msgid "Node1"
msgstr "节点 1"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:243
msgid ""
"If the service starts successfully, the following information will be "
"displayed on node 0:"
msgstr "如果服务启动成功,节点 0 上将显示以下信息:"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:254
msgid "Multi-node Deployment with Ray"
msgstr "使用 Ray 进行多节点部署"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:256
msgid "refer to [Ray Distributed (Qwen/Qwen3-235B-A22B)](../features/ray.md)."
msgstr "请参考 [Ray 分布式 (Qwen/Qwen3-235B-A22B)](../features/ray.md)。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:258
msgid "Prefill-Decode Disaggregation"
msgstr "预填充-解码分离"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:260
msgid ""
"refer to [Prefill-Decode Disaggregation Mooncake Verification "
"(Qwen)](../features/pd_disaggregation_mooncake_multi_node.md)"
msgstr "请参阅 [Prefill-Decode 分离部署 Mooncake 验证 (Qwen)](../features/pd_disaggregation_mooncake_multi_node.md)"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:262
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:264
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "服务器启动后,您可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:277
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:279
msgid "Here are two accuracy evaluation methods."
msgstr "以下是两种精度评估方法。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:281
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:293
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:283
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参阅 [使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:285
msgid ""
"After execution, you can get the result, here is the result of `Qwen3"
"-235B-A22B-w8a8` in `vllm-ascend:0.11.0rc0` for reference only."
msgstr "执行后,您将获得结果。以下是 `vllm-ascend:0.11.0rc0` 中 `Qwen3-235B-A22B-w8a8` 的结果,仅供参考。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "version"
msgstr "版本"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "cevaldataset"
msgstr "cevaldataset"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "-"
msgstr "-"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "accuracy"
msgstr "准确率"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "gen"
msgstr "生成"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "91.16"
msgstr "91.16"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:291
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:295
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参阅 [使用 AISBench 进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:297
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM Benchmark"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:299
msgid "Run performance evaluation of `Qwen3-235B-A22B-w8a8` as an example."
msgstr "以运行 `Qwen3-235B-A22B-w8a8` 的性能评估为例。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:301
msgid ""
"Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) "
"for more details."
msgstr "更多详情请参阅 [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:303
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 包含三个子命令:"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:305
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批请求的延迟进行基准测试。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:306
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:307
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:309
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。按如下方式运行代码。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:316
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您将获得性能评估结果。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:318
msgid "Reproducing Performance Results"
msgstr "复现性能结果"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:320
msgid ""
"In this section, we provide simple scripts to re-produce our latest "
"performance. It is also recommended to read instructions above to "
"understand basic concepts or options in vLLM && vLLM-Ascend."
msgstr "本节提供简单的脚本来复现我们最新的性能结果。也建议阅读上方的说明,以了解 vLLM 和 vLLM-Ascend 中的基本概念或选项。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:322
msgid "Environment"
msgstr "环境"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:324
msgid "vLLM v0.13.0"
msgstr "vLLM v0.13.0"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:325
msgid "vLLM-Ascend v0.13.0rc1"
msgstr "vLLM-Ascend v0.13.0rc1"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:326
msgid "CANN 8.3.RC2"
msgstr "CANN 8.3.RC2"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:327
msgid "torch_npu 2.8.0"
msgstr "torch_npu 2.8.0"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:328
msgid "HDK/driver 25.3.RC1"
msgstr "HDK/驱动 25.3.RC1"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:329
msgid "triton_ascend 3.2.0"
msgstr "triton_ascend 3.2.0"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:331
msgid "Single Node A3 (64G*16)"
msgstr "单节点 A3 (64G*16)"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:333
msgid "Example server scripts:"
msgstr "服务器脚本示例:"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:368
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:597
msgid "Benchmark scripts:"
msgstr "基准测试脚本:"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:384
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:613
msgid "Reference test results:"
msgstr "参考测试结果:"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "num_requests"
msgstr "请求数量"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "concurrency"
msgstr "并发数"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "mean TTFT(ms)"
msgstr "平均 TTFT(毫秒)"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "mean TPOT(ms)"
msgstr "平均 TPOT(毫秒)"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "output token throughput (tok/s)"
msgstr "输出令牌吞吐量 (令牌/秒)"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "720"
msgstr "720"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "144"
msgstr "144"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "4717.45"
msgstr "4717.45"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "48.69"
msgstr "48.69"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "2761.72"
msgstr "2761.72"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:390
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:619
msgid "Note:"
msgstr "注意:"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:392
msgid ""
"Setting `export VLLM_ASCEND_ENABLE_FUSED_MC2=1` enables MoE fused "
"operators that reduce time consumption of MoE in both prefill and decode."
" This is an experimental feature which only supports W8A8 quantization on"
" Atlas A3 servers now. If you encounter any problems when using this "
"feature, you can disable it by setting `export "
"VLLM_ASCEND_ENABLE_FUSED_MC2=0` and update issues in vLLM-Ascend "
"community."
msgstr "设置 `export VLLM_ASCEND_ENABLE_FUSED_MC2=1` 可启用 MoE 融合算子,以减少预填充和解码阶段 MoE 的时间消耗。这是一个实验性功能,目前仅支持 Atlas A3 服务器上的 W8A8 量化。如果您在使用此功能时遇到任何问题,可以通过设置 `export VLLM_ASCEND_ENABLE_FUSED_MC2=0` 来禁用它,并在 vLLM-Ascend 社区更新问题。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:393
msgid ""
"Here we disable prefix cache because of random datasets. You can enable "
"prefix cache if requests have long common prefix."
msgstr "由于使用随机数据集,此处我们禁用了前缀缓存。如果请求具有较长的公共前缀,您可以启用前缀缓存。"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:395
msgid "Three Node A3 -- PD disaggregation"
msgstr "三节点 A3 -- PD 分离部署"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:397
msgid ""
"On three Atlas 800 A3(64G*16) server, we recommend to use one node as one"
" prefill instance and two nodes as one decode instance. Example server "
"scripts: Prefill Node 1"
msgstr "在三台 Atlas 800 A3(64G*16) 服务器上,我们建议使用一个节点作为一个预填充实例,两个节点作为一个解码实例。服务器脚本示例:预填充节点 1"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:462
msgid "Decode Node 1"
msgstr "解码节点 1"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:526
msgid "Decode Node 2"
msgstr "解码节点 2"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:591
msgid "PD proxy:"
msgstr "PD 代理:"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "2880"
msgstr "2880"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "576"
msgstr "576"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "3735.98"
msgstr "3735.98"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "52.07"
msgstr "52.07"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:76
msgid "8593.44"
msgstr "8593.44"
#: ../../source/tutorials/models/Qwen3-235B-A22B.md:621
msgid ""
"We recommend to set `export VLLM_ASCEND_ENABLE_FUSED_MC2=2` on this "
"scenario (typically EP32 for Qwen3-235B). This enables a different MoE "
"fusion operator."
msgstr "在此场景下(通常 Qwen3-235B 使用 EP32我们建议设置 `export VLLM_ASCEND_ENABLE_FUSED_MC2=2`。这将启用一个不同的 MoE 融合算子。"

View File

@@ -0,0 +1,67 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3-30B-A3B.md:1
msgid "Qwen3-30B-A3B"
msgstr "Qwen3-30B-A3B"
#: ../../source/tutorials/models/Qwen3-30B-A3B.md:3
msgid "Run vllm-ascend on Multi-NPU with Qwen3 MoE"
msgstr "在 Multi-NPU 上使用 Qwen3 MoE 运行 vllm-ascend"
#: ../../source/tutorials/models/Qwen3-30B-A3B.md:5
msgid "Run docker container:"
msgstr "运行 docker 容器:"
#: ../../source/tutorials/models/Qwen3-30B-A3B.md:34
msgid "Set up environment variables:"
msgstr "设置环境变量:"
#: ../../source/tutorials/models/Qwen3-30B-A3B.md:44
msgid "Online Inference on Multi-NPU"
msgstr "在 Multi-NPU 上进行在线推理"
#: ../../source/tutorials/models/Qwen3-30B-A3B.md:46
msgid "Run the following script to start the vLLM server on Multi-NPU:"
msgstr "运行以下脚本以在 Multi-NPU 上启动 vLLM 服务器:"
#: ../../source/tutorials/models/Qwen3-30B-A3B.md:48
msgid ""
"For an Atlas A2 with 64 GB of NPU card memory, tensor-parallel-size "
"should be at least 2, and for 32 GB of memory, tensor-parallel-size "
"should be at least 4."
msgstr "对于具有 64 GB NPU 卡内存的 Atlas A2tensor-parallel-size 应至少为 2对于 32 GB 内存tensor-parallel-size 应至少为 4。"
#: ../../source/tutorials/models/Qwen3-30B-A3B.md:54
msgid "Once your server is started, you can query the model with input prompts."
msgstr "服务器启动后,您可以使用输入提示词查询模型。"
#: ../../source/tutorials/models/Qwen3-30B-A3B.md:69
msgid "Offline Inference on Multi-NPU"
msgstr "在 Multi-NPU 上进行离线推理"
#: ../../source/tutorials/models/Qwen3-30B-A3B.md:71
msgid "Run the following script to execute offline inference on multi-NPU:"
msgstr "运行以下脚本以在 multi-NPU 上执行离线推理:"
#: ../../source/tutorials/models/Qwen3-30B-A3B.md:108
msgid "If you run this script successfully, you can see the info shown below:"
msgstr "如果成功运行此脚本,您将看到如下所示的信息:"

View File

@@ -0,0 +1,88 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:1
msgid "Qwen3-32B-W4A4"
msgstr "Qwen3-32B-W4A4"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:5
msgid ""
"W4A4 Flat Quantization is for better model compression and inference "
"efficiency on Ascend devices. And W4A4 is supported since `v0.11.0rc1`. "
"For modelslim, W4A4 is supported since `tag_MindStudio_8.2.RC1.B120_002`."
msgstr ""
"W4A4 扁平量化旨在提升模型在昇腾设备上的压缩率和推理效率。W4A4 自 `v0.11.0rc1` 版本起获得支持。对于 modelslimW4A4 自 `tag_MindStudio_8.2.RC1.B120_002` 版本起获得支持。"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:8
msgid "The following steps will show how to quantize Qwen3 32B to W4A4."
msgstr "以下步骤将展示如何将 Qwen3 32B 量化为 W4A4。"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:10
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:12
msgid "Run Docker Container"
msgstr "运行 Docker 容器"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:35
msgid "Install modelslim and Convert Model"
msgstr "安装 modelslim 并转换模型"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:38
msgid ""
"You can choose to convert the model yourself or use the quantized model "
"we uploaded, see <https://www.modelscope.cn/models/vllm-ascend/Qwen3-32B-"
"W4A4>"
msgstr ""
"您可以选择自行转换模型,或使用我们已上传的量化模型,详见 <https://www.modelscope.cn/models/vllm-ascend/Qwen3-32B-W4A4>"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:68
msgid "Verify the Quantized Model"
msgstr "验证量化模型"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:70
msgid "The converted model files look like:"
msgstr "转换后的模型文件结构如下:"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:95
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:97
msgid "Online Serving on Single NPU"
msgstr "单 NPU 在线服务"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:103
msgid "Once your server is started, you can query the model with input prompts."
msgstr "服务器启动后,您可以使用输入提示词查询模型。"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:118
msgid "Offline Inference on Single NPU"
msgstr "单 NPU 离线推理"
#: ../../source/tutorials/models/Qwen3-32B-W4A4.md:121
msgid "To enable quantization for ascend, quantization method must be \"ascend\"."
msgstr "要为昇腾启用量化,量化方法必须设置为 \"ascend\"。"

View File

@@ -0,0 +1,72 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3-8B-W4A8.md:1
msgid "Qwen3-8B-W4A8"
msgstr "Qwen3-8B-W4A8"
#: ../../source/tutorials/models/Qwen3-8B-W4A8.md:3
msgid "Run Docker Container"
msgstr "运行 Docker 容器"
#: ../../source/tutorials/models/Qwen3-8B-W4A8.md:6
msgid "w4a8 quantization feature is supported by v0.9.1rc2 and later."
msgstr "w4a8 量化特性由 v0.9.1rc2 及更高版本支持。"
#: ../../source/tutorials/models/Qwen3-8B-W4A8.md:30
msgid "Install modelslim and Convert Model"
msgstr "安装 modelslim 并转换模型"
#: ../../source/tutorials/models/Qwen3-8B-W4A8.md:33
msgid ""
"You can choose to convert the model yourself or use the quantized model "
"we uploaded, see <https://www.modelscope.cn/models/vllm-ascend/Qwen3-8B-"
"W4A8>"
msgstr ""
"您可以选择自行转换模型,或使用我们已上传的量化模型,请参阅 <https://www.modelscope.cn/models/vllm-ascend/Qwen3-8B-W4A8>"
#: ../../source/tutorials/models/Qwen3-8B-W4A8.md:73
msgid "Verify the Quantized Model"
msgstr "验证量化模型"
#: ../../source/tutorials/models/Qwen3-8B-W4A8.md:75
msgid "The converted model files look like:"
msgstr "转换后的模型文件结构如下:"
#: ../../source/tutorials/models/Qwen3-8B-W4A8.md:93
msgid ""
"Run the following script to start the vLLM server with the quantized "
"model:"
msgstr "运行以下脚本来启动使用量化模型的 vLLM 服务器:"
#: ../../source/tutorials/models/Qwen3-8B-W4A8.md:101
msgid "Once your server is started, you can query the model with input prompts."
msgstr "服务器启动后,您可以通过输入提示词来查询模型。"
#: ../../source/tutorials/models/Qwen3-8B-W4A8.md:116
msgid ""
"Run the following script to execute offline inference on single-NPU with "
"the quantized model:"
msgstr "运行以下脚本,使用量化模型在单 NPU 上进行离线推理:"
#: ../../source/tutorials/models/Qwen3-8B-W4A8.md:119
msgid "To enable quantization for ascend, quantization method must be \"ascend\"."
msgstr "要为 Ascend 启用量化,量化方法必须设置为 \"ascend\"。"

View File

@@ -0,0 +1,210 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:1
msgid "Qwen3-Coder-30B-A3B"
msgstr "Qwen3-Coder-30B-A3B"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:5
msgid ""
"The newly released Qwen3-Coder-30B-A3B employs a sparse MoE architecture "
"for efficient training and inference, delivering significant "
"optimizations in agentic coding, extended context support of up to 1M "
"tokens, and versatile function calling."
msgstr "新发布的 Qwen3-Coder-30B-A3B 采用稀疏 MoE 架构,以实现高效的训练和推理,在智能体编码、高达 1M token 的扩展上下文支持以及多功能调用方面带来了显著优化。"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:7
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, single-node deployment, accuracy and performance evaluation."
msgstr "本文档将展示该模型的主要验证步骤,包括支持的功能、功能配置、环境准备、单节点部署、精度和性能评估。"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:9
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:11
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取模型支持的功能矩阵。"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:13
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[功能指南](../../user_guide/feature_guide/index.md)以获取功能的配置信息。"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:15
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:17
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:19
msgid ""
"`Qwen3-Coder-30B-A3B-Instruct`(BF16 version): requires 1 Atlas 800 A3 "
"node (with 16x 64G NPUs) or 1 Atlas 800 A2 node (with 8x 64G/32G NPUs). "
"[Download model weight](https://modelscope.cn/models/Qwen/Qwen3-Coder-"
"30B-A3B-Instruct)"
msgstr "`Qwen3-Coder-30B-A3B-Instruct`BF16 版本):需要 1 个 Atlas 800 A3 节点(配备 16 个 64G NPU或 1 个 Atlas 800 A2 节点(配备 8 个 64G/32G NPU。[下载模型权重](https://modelscope.cn/models/Qwen/Qwen3-Coder-30B-A3B-Instruct)"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:21
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`"
msgstr "建议将模型权重下载到多节点的共享目录中,例如 `/root/.cache/`"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:23
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:25
msgid ""
"`Qwen3-Coder` is first supported in `vllm-ascend:v0.10.0rc1`, please run "
"this model using a later version."
msgstr "`Qwen3-Coder` 首次在 `vllm-ascend:v0.10.0rc1` 中得到支持,请使用此版本或更高版本来运行此模型。"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:27
msgid ""
"You can use our official docker image to run `Qwen3-Coder-30B-A3B-"
"Instruct` directly."
msgstr "您可以使用我们的官方 docker 镜像直接运行 `Qwen3-Coder-30B-A3B-Instruct`。"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:53
msgid ""
"In addition, if you don't want to use the docker image as above, you can "
"also build all from source:"
msgstr "此外,如果您不想使用上述 docker 镜像,也可以从源代码构建所有内容:"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:55
msgid ""
"Install `vllm-ascend` from source, refer to "
"[installation](../../installation.md)."
msgstr "从源代码安装 `vllm-ascend`,请参考[安装指南](../../installation.md)。"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:57
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:59
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:61
msgid "Run the following script to execute online inference."
msgstr "运行以下脚本来执行在线推理。"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:63
msgid ""
"For an Atlas A2 with 64 GB of NPU card memory, tensor-parallel-size "
"should be at least 2, and for 32 GB of memory, tensor-parallel-size "
"should be at least 4."
msgstr "对于配备 64 GB NPU 显存的 Atlas A2张量并行大小应至少为 2对于 32 GB 显存,张量并行大小应至少为 4。"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:72
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:74
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "服务器启动后,您可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:89
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:91
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:103
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:93
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参考[使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:95
msgid ""
"After execution, you can get the result, here is the result of `Qwen3"
"-Coder-30B-A3B-Instruct` in `vllm-ascend:0.11.0rc0` for reference only."
msgstr "执行后,您可以获得结果。以下是 `Qwen3-Coder-30B-A3B-Instruct` 在 `vllm-ascend:0.11.0rc0` 中的结果,仅供参考。"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:29
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:29
msgid "version"
msgstr "版本"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:29
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:29
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:29
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:29
msgid "openai_humaneval"
msgstr "openai_humaneval"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:29
msgid "f4a973"
msgstr "f4a973"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:29
msgid "humaneval_pass@1"
msgstr "humaneval_pass@1"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:29
msgid "gen"
msgstr "gen"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:29
msgid "94.51"
msgstr "94.51"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:101
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen3-Coder-30B-A3B.md:105
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参考[使用 AISBench 进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"

View File

@@ -0,0 +1,866 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3-Dense.md:1
msgid "Qwen3-Dense(Qwen3-0.6B/8B/32B)"
msgstr "Qwen3-Dense(Qwen3-0.6B/8B/32B)"
#: ../../source/tutorials/models/Qwen3-Dense.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3-Dense.md:5
msgid ""
"Qwen3 is the latest generation of large language models in Qwen series, "
"offering a comprehensive suite of dense and mixture-of-experts (MoE) "
"models. Built upon extensive training, Qwen3 delivers groundbreaking "
"advancements in reasoning, instruction-following, agent capabilities, and"
" multilingual support."
msgstr ""
"Qwen3 是 Qwen 系列最新一代的大语言模型,提供了一套完整的稠密模型和专家混合"
"(MoE) 模型。基于广泛的训练Qwen3 在推理、指令遵循、智能体能力和多语言支持方"
"面实现了突破性进展。"
#: ../../source/tutorials/models/Qwen3-Dense.md:7
msgid ""
"Welcome to the tutorial on optimizing Qwen Dense models in the vLLM-"
"Ascend environment. This guide will help you configure the most effective"
" settings for your use case, with practical examples that highlight key "
"optimization points. We will also explore how adjusting service "
"parameters can maximize throughput performance across various scenarios."
msgstr ""
"欢迎阅读在 vLLM-Ascend 环境中优化 Qwen 稠密模型的教程。本指南将帮助您为您的用"
"例配置最有效的设置,并通过实际示例突出关键优化点。我们还将探讨如何调整服务参"
"数以在各种场景下最大化吞吐性能。"
#: ../../source/tutorials/models/Qwen3-Dense.md:9
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, accuracy and performance evaluation."
msgstr ""
"本文档将展示模型的主要验证步骤,包括支持的特性、特性配置、环境准备、精度和性"
"能评估。"
#: ../../source/tutorials/models/Qwen3-Dense.md:11
msgid ""
"The Qwen3 Dense models are first supported in "
"[v0.8.4rc2](https://github.com/vllm-project/vllm-"
"ascend/blob/main/docs/source/user_guide/release_notes.md#v084rc2---"
"20250429). This example requires version **v0.11.0rc2**. Earlier versions"
" may lack certain features."
msgstr ""
"Qwen3 稠密模型首次在 "
"[v0.8.4rc2](https://github.com/vllm-project/vllm-"
"ascend/blob/main/docs/source/user_guide/release_notes.md#v084rc2---"
"20250429) 中得到支持。本示例需要版本 **v0.11.0rc2**。更早的版本可能缺少某些特"
"性。"
#: ../../source/tutorials/models/Qwen3-Dense.md:13
msgid "Supported Features"
msgstr "支持的特性"
#: ../../source/tutorials/models/Qwen3-Dense.md:15
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr ""
"请参考 [支持的特性](../../user_guide/support_matrix/supported_models."
"md) 以获取模型支持的特性矩阵。"
#: ../../source/tutorials/models/Qwen3-Dense.md:17
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr ""
"请参考 [特性指南](../../user_guide/feature_guide/index.md) 以获取特性的配置信"
"息。"
#: ../../source/tutorials/models/Qwen3-Dense.md:19
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen3-Dense.md:21
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen3-Dense.md:23
msgid ""
"`Qwen3-0.6B`(BF16 version): require 1 Atlas 800 A3 (64G × 2) card or 1 "
"Atlas 800I A2 (64G × 1) card. [Download model "
"weight](https://modelers.cn/models/Modelers_Park/Qwen3-0.6B)"
msgstr ""
"`Qwen3-0.6B`(BF16 版本): 需要 1 张 Atlas 800 A3 (64G × 2) 卡或 1 张 Atlas "
"800I A2 (64G × 1) 卡。[下载模型权重](https://modelers.cn/models/"
"Modelers_Park/Qwen3-0.6B)"
#: ../../source/tutorials/models/Qwen3-Dense.md:24
msgid ""
"`Qwen3-1.7B`(BF16 version): require 1 Atlas 800 A3 (64G × 2) card or 1 "
"Atlas 800I A2 (64G × 1) card. [Download model "
"weight](https://modelers.cn/models/Modelers_Park/Qwen3-1.7B)"
msgstr ""
"`Qwen3-1.7B`(BF16 版本): 需要 1 张 Atlas 800 A3 (64G × 2) 卡或 1 张 Atlas "
"800I A2 (64G × 1) 卡。[下载模型权重](https://modelers.cn/models/"
"Modelers_Park/Qwen3-1.7B)"
#: ../../source/tutorials/models/Qwen3-Dense.md:25
msgid ""
"`Qwen3-4B`(BF16 version): require 1 Atlas 800 A3 (64G × 2) card or 1 "
"Atlas 800I A2 (64G × 1) card. [Download model "
"weight](https://modelers.cn/models/Modelers_Park/Qwen3-4B)"
msgstr ""
"`Qwen3-4B`(BF16 版本): 需要 1 张 Atlas 800 A3 (64G × 2) 卡或 1 张 Atlas "
"800I A2 (64G × 1) 卡。[下载模型权重](https://modelers.cn/models/"
"Modelers_Park/Qwen3-4B)"
#: ../../source/tutorials/models/Qwen3-Dense.md:26
msgid ""
"`Qwen3-8B`(BF16 version): require 1 Atlas 800 A3 (64G × 2) card or 1 "
"Atlas 800I A2 (64G × 1) card. [Download model "
"weight](https://modelers.cn/models/Modelers_Park/Qwen3-8B)"
msgstr ""
"`Qwen3-8B`(BF16 版本): 需要 1 张 Atlas 800 A3 (64G × 2) 卡或 1 张 Atlas "
"800I A2 (64G × 1) 卡。[下载模型权重](https://modelers.cn/models/"
"Modelers_Park/Qwen3-8B)"
#: ../../source/tutorials/models/Qwen3-Dense.md:27
msgid ""
"`Qwen3-14B`(BF16 version): require 1 Atlas 800 A3 (64G × 2) card or 2 "
"Atlas 800I A2 (64G × 1) cards. [Download model "
"weight](https://modelers.cn/models/Modelers_Park/Qwen3-14B)"
msgstr ""
"`Qwen3-14B`(BF16 版本): 需要 1 张 Atlas 800 A3 (64G × 2) 卡或 2 张 Atlas "
"800I A2 (64G × 1) 卡。[下载模型权重](https://modelers.cn/models/"
"Modelers_Park/Qwen3-14B)"
#: ../../source/tutorials/models/Qwen3-Dense.md:28
msgid ""
"`Qwen3-32B`(BF16 version): require 2 Atlas 800 A3 (64G × 4) cards or 4 "
"Atlas 800I A2 (64G × 4) cards. [Download model "
"weight](https://modelers.cn/models/Modelers_Park/Qwen3-32B)"
msgstr ""
"`Qwen3-32B`(BF16 版本): 需要 2 张 Atlas 800 A3 (64G × 4) 卡或 4 张 Atlas "
"800I A2 (64G × 4) 卡。[下载模型权重](https://modelers.cn/models/"
"Modelers_Park/Qwen3-32B)"
#: ../../source/tutorials/models/Qwen3-Dense.md:29
msgid ""
"`Qwen3-32B-W8A8`(Quantized version): require 2 Atlas 800 A3 (64G × 4) "
"cards or 4 Atlas 800I A2 (64G × 4) cards. [Download model "
"weight](https://www.modelscope.cn/models/vllm-ascend/Qwen3-32B-W8A8)"
msgstr ""
"`Qwen3-32B-W8A8`(量化版本): 需要 2 张 Atlas 800 A3 (64G × 4) 卡或 4 张 "
"Atlas 800I A2 (64G × 4) 卡。[下载模型权重](https://www.modelscope.cn/"
"models/vllm-ascend/Qwen3-32B-W8A8)"
#: ../../source/tutorials/models/Qwen3-Dense.md:31
msgid ""
"These are the recommended numbers of cards, which can be adjusted "
"according to the actual situation."
msgstr "这些是推荐的卡数,可以根据实际情况进行调整。"
#: ../../source/tutorials/models/Qwen3-Dense.md:33
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`"
msgstr "建议将模型权重下载到多节点的共享目录,例如 `/root/.cache/`"
#: ../../source/tutorials/models/Qwen3-Dense.md:35
msgid "Verify Multi-node Communication(Optional)"
msgstr "验证多节点通信(可选)"
#: ../../source/tutorials/models/Qwen3-Dense.md:37
msgid ""
"If you want to deploy multi-node environment, you need to verify multi-"
"node communication according to [verify multi-node communication "
"environment](../../installation.md#verify-multi-node-communication)."
msgstr ""
"如果您想部署多节点环境,需要根据 [验证多节点通信环境](../../installation."
"md#verify-multi-node-communication) 来验证多节点通信。"
#: ../../source/tutorials/models/Qwen3-Dense.md:39
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen3-Dense.md:41
msgid ""
"You can use our official docker image for supporting Qwen3 Dense models. "
"Currently, we provide the all-in-one images.[Download "
"images](https://quay.io/repository/ascend/vllm-ascend?tab=tags)"
msgstr ""
"您可以使用我们的官方 docker 镜像来支持 Qwen3 稠密模型。目前,我们提供一体化镜"
"像。[下载镜像](https://quay.io/repository/ascend/vllm-ascend?tab=tags)"
#: ../../source/tutorials/models/Qwen3-Dense.md:44
msgid "Docker Pull (by tag)"
msgstr "Docker 拉取(通过标签)"
#: ../../source/tutorials/models/Qwen3-Dense.md:53
msgid "Docker run"
msgstr "Docker 运行"
#: ../../source/tutorials/models/Qwen3-Dense.md:90
msgid ""
"The default workdir is `/workspace`, vLLM and vLLM Ascend code are placed"
" in `/vllm-workspace` and installed in [development "
"mode](https://setuptools.pypa.io/en/latest/userguide/development_mode.html)"
" (`pip install -e`) to help developer immediately take place changes "
"without requiring a new installation."
msgstr ""
"默认工作目录是 `/workspace`vLLM 和 vLLM Ascend 代码放置在 `/vllm-"
"workspace` 中,并以 [开发模式](https://setuptools.pypa.io/en/latest/"
"userguide/development_mode.html) (`pip install -e`) 安装,以帮助开发者立即应用"
"更改而无需重新安装。"
#: ../../source/tutorials/models/Qwen3-Dense.md:92
msgid ""
"In the [Run docker container](./Qwen3-Dense.md#run-docker-container), "
"detailed explanations are provided through specific examples."
msgstr ""
"在 [运行 docker 容器](./Qwen3-Dense.md#run-docker-container) 中,通过具体示例"
"提供了详细说明。"
#: ../../source/tutorials/models/Qwen3-Dense.md:94
msgid ""
"In addition, if you don't want to use the docker image as above, you can "
"also build all from source:"
msgstr "此外,如果您不想使用上述 docker 镜像,也可以从源码构建所有内容:"
#: ../../source/tutorials/models/Qwen3-Dense.md:96
msgid ""
"Install `vllm-ascend` from source, refer to "
"[installation](../../installation.md)."
msgstr "从源码安装 `vllm-ascend`,请参考 [安装](../../installation.md)。"
#: ../../source/tutorials/models/Qwen3-Dense.md:98
msgid ""
"If you want to deploy multi-node environment, you need to set up "
"environment on each node."
msgstr "如果您想部署多节点环境,需要在每个节点上设置环境。"
#: ../../source/tutorials/models/Qwen3-Dense.md:100
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3-Dense.md:102
msgid ""
"In this section, we will demonstrate best practices for adjusting "
"hyperparameters in vLLM-Ascend to maximize inference throughput "
"performance. By tailoring service-level configurations to fit different "
"use cases, you can ensure that your system performs optimally across "
"various scenarios. We will guide you through how to fine-tune "
"hyperparameters based on observed phenomena, such as max_model_len, "
"max_num_batched_tokens, and cudagraph_capture_sizes, to achieve the best "
"performance."
msgstr ""
"在本节中,我们将演示在 vLLM-Ascend 中调整超参数以实现最大推理吞吐性能的最佳实"
"践。通过定制服务级配置以适应不同的用例,您可以确保您的系统在各种场景下都能达"
"到最佳性能。我们将指导您如何根据观察到的现象(例如 max_model_len、"
"max_num_batched_tokens 和 cudagraph_capture_sizes来微调超参数以获得最佳性"
"能。"
#: ../../source/tutorials/models/Qwen3-Dense.md:104
msgid "The specific example scenario is as follows:"
msgstr "具体示例如下:"
#: ../../source/tutorials/models/Qwen3-Dense.md:106
msgid "The machine environment is an Atlas 800 A3 (64G*16)"
msgstr "机器环境是 Atlas 800 A3 (64G*16)"
#: ../../source/tutorials/models/Qwen3-Dense.md:107
msgid "The LLM is Qwen3-32B-W8A8"
msgstr "LLM 是 Qwen3-32B-W8A8"
#: ../../source/tutorials/models/Qwen3-Dense.md:108
msgid "The data scenario is a fixed-length input of 3.5K and an output of 1.5K."
msgstr "数据场景是固定长度输入 3.5K 和输出 1.5K。"
#: ../../source/tutorials/models/Qwen3-Dense.md:109
msgid "The parallel configuration requirement is DP=1&TP=4"
msgstr "并行配置要求是 DP=1&TP=4"
#: ../../source/tutorials/models/Qwen3-Dense.md:110
msgid ""
"If the machine environment is an **Atlas 800I A2(64G*8)**, the deployment"
" approach stays identical."
msgstr "如果机器环境是 **Atlas 800I A2(64G*8)**,部署方法保持不变。"
#: ../../source/tutorials/models/Qwen3-Dense.md:112
msgid "Run docker container"
msgstr "运行 docker 容器"
#: ../../source/tutorials/models/Qwen3-Dense.md:116
#: ../../source/tutorials/models/Qwen3-Dense.md:192
#: ../../source/tutorials/models/Qwen3-Dense.md:222
#: ../../source/tutorials/models/Qwen3-Dense.md:303
msgid ""
"vllm-ascend/Qwen3-32B-W8A8 is the default model path, replace this with "
"your actual path."
msgstr "vllm-ascend/Qwen3-32B-W8A8 是默认模型路径,请替换为您的实际路径。"
#: ../../source/tutorials/models/Qwen3-Dense.md:117
msgid "v0.11.0rc2-a3 is image tag, replace this with your actual tag."
msgstr "v0.11.0rc2-a3 是镜像标签,请替换为您的实际标签。"
#: ../../source/tutorials/models/Qwen3-Dense.md:118
msgid "replace this with your actual port: '-p 8113:8113'."
msgstr "请替换为您的实际端口:'-p 8113:8113'。"
#: ../../source/tutorials/models/Qwen3-Dense.md:119
msgid "replace this with your actual card: '--device /dev/davinci0'."
msgstr "请替换为您的实际卡:'--device /dev/davinci0'。"
#: ../../source/tutorials/models/Qwen3-Dense.md:147
msgid "Online Inference on Multi-NPU"
msgstr "多 NPU 在线推理"
#: ../../source/tutorials/models/Qwen3-Dense.md:149
msgid "Run the following script to start the vLLM server on Multi-NPU."
msgstr "运行以下脚本以在多 NPU 上启动 vLLM 服务器。"
#: ../../source/tutorials/models/Qwen3-Dense.md:151
msgid ""
"This script is configured to achieve optimal performance under the above "
"specific example scenarios,with batchsize = 72 on two A3 cards."
msgstr "此脚本配置为在上述特定示例场景下实现最佳性能,在两块 A3 卡上 batchsize = 72。"
#: ../../source/tutorials/models/Qwen3-Dense.md:194
msgid ""
"If the model is not a quantized model, remove the `--quantization ascend`"
" parameter."
msgstr "如果模型不是量化模型,请移除 `--quantization ascend` 参数。"
#: ../../source/tutorials/models/Qwen3-Dense.md:196
#, python-brace-format
msgid ""
"**[Optional]** `--additional-config '{\"pa_shape_list\":[48,64,72,80]}'`:"
" `pa_shape_list` specifies the batch sizes where you want to switch to "
"the PA operator. This is a temporary tuning knob. Currently, the "
"attention operator dispatch defaults to the FIA operator. In some batch-"
"size (concurrency) settings, FIA may have suboptimal performance. By "
"setting `pa_shape_list`, when the runtime batch size matches one of the "
"listed values, vLLM-Ascend will replace FIA with the PA operator to "
"prevent performance degradation. In the future, FIA will be optimized for"
" these scenarios and this parameter will be removed."
msgstr ""
"**[可选]** `--additional-config '{\"pa_shape_list\":[48,64,72,80]}'`: "
"`pa_shape_list` 指定了您希望切换到 PA 算子的批次大小。这是一个临时的调优旋"
"钮。目前,注意力算子调度默认使用 FIA 算子。在某些批次大小并发设置下FIA "
"可能性能不佳。通过设置 `pa_shape_list`,当运行时批次大小与列出的值之一匹配时,"
"vLLM-Ascend 将用 PA 算子替换 FIA 算子以防止性能下降。未来FIA 将针对这些场景"
"进行优化,此参数将被移除。"
#: ../../source/tutorials/models/Qwen3-Dense.md:198
#, python-brace-format
msgid ""
"If the ultimate performance is desired, the cudagraph_capture_sizes "
"parameter can be enabled, reference: [key-optimization-"
"points](./Qwen3-Dense.md#key-optimization-points)、[optimization-"
"highlights](./Qwen3-Dense.md#optimization-highlights). Here is an example"
" of batchsize of 72: `--compilation-config '{\"cudagraph_mode\": "
"\"FULL_DECODE_ONLY\", "
"\"cudagraph_capture_sizes\":[1,8,24,48,60,64,72,76]}'`."
msgstr ""
"如果需要极致性能,可以启用 cudagraph_capture_sizes 参数,参考:[关键优化"
"点](./Qwen3-Dense.md#key-optimization-points)、[优化亮点](./Qwen3-"
"Dense.md#optimization-highlights)。以下是批次大小为 72 的示例:`--compilation-"
"config '{\"cudagraph_mode\": \"FULL_DECODE_ONLY\", "
"\"cudagraph_capture_sizes\":[1,8,24,48,60,64,72,76]}'`。"
#: ../../source/tutorials/models/Qwen3-Dense.md:201
msgid "Once your server is started, you can query the model with input prompts"
msgstr "服务器启动后,您可以使用输入提示词查询模型"
#: ../../source/tutorials/models/Qwen3-Dense.md:216
msgid "Offline Inference on Multi-NPU"
msgstr "多 NPU 离线推理"
#: ../../source/tutorials/models/Qwen3-Dense.md:218
msgid "Run the following script to execute offline inference on multi-NPU."
msgstr "运行以下脚本以在多 NPU 上执行离线推理。"
#: ../../source/tutorials/models/Qwen3-Dense.md:224
msgid ""
"If the model is not a quantized model,remove the "
"`quantization=\"ascend\"` parameter."
msgstr "如果模型不是量化模型,请移除 `quantization=\"ascend\"` 参数。"
#: ../../source/tutorials/models/Qwen3-Dense.md:265
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/Qwen3-Dense.md:267
msgid "Here is one accuracy evaluation methods."
msgstr "这里是一种精度评估方法。"
#: ../../source/tutorials/models/Qwen3-Dense.md:269
#: ../../source/tutorials/models/Qwen3-Dense.md:283
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/models/Qwen3-Dense.md:271
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参阅[使用AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/Qwen3-Dense.md:273
msgid ""
"After execution, you can get the result, here is the result of `Qwen3"
"-32B-W8A8` in `vllm-ascend:0.11.0rc2` for reference only."
msgstr "执行后,您将获得结果。此处展示的是 `Qwen3-32B-W8A8` 在 `vllm-ascend:0.11.0rc2` 环境下的结果,仅供参考。"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "version"
msgstr "版本"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "task name"
msgstr "任务名称"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "gsm8k"
msgstr "gsm8k"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "-"
msgstr "-"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "accuracy"
msgstr "准确率"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "gen"
msgstr "生成"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "gsm8k_gen_0_shot_noncot_chat_prompt"
msgstr "gsm8k_gen_0_shot_noncot_chat_prompt"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "96.44"
msgstr "96.44"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "math500"
msgstr "math500"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "math500_gen_0_shot_cot_chat_prompt"
msgstr "math500_gen_0_shot_cot_chat_prompt"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "97.60"
msgstr "97.60"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "aime"
msgstr "aime"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "aime2024_gen_0_shot_chat_prompt"
msgstr "aime2024_gen_0_shot_chat_prompt"
#: ../../source/tutorials/models/Qwen3-Dense.md:220
msgid "76.67"
msgstr "76.67"
#: ../../source/tutorials/models/Qwen3-Dense.md:281
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen3-Dense.md:285
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参阅[使用AISBench进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/Qwen3-Dense.md:287
msgid "Using vLLM Benchmark"
msgstr "使用vLLM基准测试"
#: ../../source/tutorials/models/Qwen3-Dense.md:289
msgid "Run performance evaluation of `Qwen3-32B-W8A8` as an example."
msgstr "以运行 `Qwen3-32B-W8A8` 的性能评估为例。"
#: ../../source/tutorials/models/Qwen3-Dense.md:291
msgid ""
"Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) "
"for more details."
msgstr "更多详情请参阅[vllm基准测试](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/models/Qwen3-Dense.md:293
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 包含三个子命令:"
#: ../../source/tutorials/models/Qwen3-Dense.md:295
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:基准测试单批次请求的延迟。"
#: ../../source/tutorials/models/Qwen3-Dense.md:296
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:基准测试在线服务吞吐量。"
#: ../../source/tutorials/models/Qwen3-Dense.md:297
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:基准测试离线推理吞吐量。"
#: ../../source/tutorials/models/Qwen3-Dense.md:299
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。按如下方式运行代码。"
#: ../../source/tutorials/models/Qwen3-Dense.md:310
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您将获得性能评估结果。"
#: ../../source/tutorials/models/Qwen3-Dense.md:312
msgid "Key Optimization Points"
msgstr "关键优化点"
#: ../../source/tutorials/models/Qwen3-Dense.md:314
msgid ""
"In this section, we will cover the key optimization points that can "
"significantly improve the performance of Qwen Dense models. These "
"techniques are designed to enhance throughput and efficiency across "
"various scenarios."
msgstr "本节将介绍能显著提升Qwen Dense模型性能的关键优化点。这些技术旨在提升各种场景下的吞吐量和效率。"
#: ../../source/tutorials/models/Qwen3-Dense.md:316
msgid "1. Rope Optimization"
msgstr "1. Rope优化"
#: ../../source/tutorials/models/Qwen3-Dense.md:318
msgid ""
"Rope optimization enhances the model's efficiency by modifying the "
"position encoding process. Specifically, it ensures that the "
"cos_sin_cache and the associated index selection operation are only "
"performed during the first layer of the forward pass. For subsequent "
"layers, the position encoding is directly reused, eliminating redundant "
"calculations and significantly speeding up inference in decode phase."
msgstr "Rope优化通过修改位置编码过程来提升模型效率。具体来说它确保 `cos_sin_cache` 及相关索引选择操作仅在正向传播的第一层执行。对于后续层,位置编码被直接复用,消除了冗余计算,并显著加快了解码阶段的推理速度。"
#: ../../source/tutorials/models/Qwen3-Dense.md:320
#: ../../source/tutorials/models/Qwen3-Dense.md:326
#: ../../source/tutorials/models/Qwen3-Dense.md:354
msgid ""
"This optimization is enabled by default and does not require any "
"additional environment variables to be set."
msgstr "此优化默认启用,无需设置任何额外的环境变量。"
#: ../../source/tutorials/models/Qwen3-Dense.md:322
msgid "2. AddRMSNormQuant Fusion"
msgstr "2. AddRMSNormQuant融合"
#: ../../source/tutorials/models/Qwen3-Dense.md:324
msgid ""
"AddRMSNormQuant fusion merges the Address-wise Multi-Scale Normalization "
"and Quantization operations, allowing for more efficient memory access "
"and computation, thereby enhancing throughput."
msgstr "AddRMSNormQuant融合将地址感知多尺度归一化与量化操作合并实现了更高效的内存访问和计算从而提升了吞吐量。"
#: ../../source/tutorials/models/Qwen3-Dense.md:328
msgid "3. FlashComm_v1"
msgstr "3. FlashComm_v1"
#: ../../source/tutorials/models/Qwen3-Dense.md:330
msgid ""
"FlashComm_v1 significantly improves performance in large-batch scenarios "
"by decomposing the traditional allreduce collective communication into "
"reduce-scatter and all-gather. This breakdown helps reduce the "
"computation of the RMSNorm token dimensions, leading to more efficient "
"processing. In quantization scenarios, FlashComm_v1 also reduces the "
"communication overhead by decreasing the bit-level data transfer, which "
"further minimizes the end-to-end latency during the prefill phase."
msgstr "FlashComm_v1通过将传统的allreduce集合通信分解为reduce-scatter和all-gather显著提升了大批量场景下的性能。这种分解有助于减少RMSNorm令牌维度的计算从而实现更高效的处理。在量化场景中FlashComm_v1还通过减少比特级数据传输来降低通信开销从而进一步最小化预填充阶段的端到端延迟。"
#: ../../source/tutorials/models/Qwen3-Dense.md:332
msgid ""
"It is important to note that the decomposition of the allreduce "
"communication into reduce-scatter and all-gather operations only provides"
" benefits in high-concurrency scenarios, where there is no significant "
"communication degradation. In other cases, this decomposition may result "
"in noticeable performance degradation. To mitigate this, the current "
"implementation uses a threshold-based approach, where FlashComm_v1 is "
"only enabled if the actual token count for each inference schedule "
"exceeds the threshold. This ensures that the feature is only activated in"
" scenarios where it improves performance, avoiding potential degradation "
"in lower-concurrency situations."
msgstr "需要注意的是将allreduce通信分解为reduce-scatter和all-gather操作仅在无显著通信降级的高并发场景下有益。在其他情况下这种分解可能导致明显的性能下降。为缓解此问题当前实现采用基于阈值的方法仅当每个推理调度的实际令牌数超过阈值时才启用FlashComm_v1。这确保了该功能仅在能提升性能的场景下激活避免了在低并发情况下可能出现的性能下降。"
#: ../../source/tutorials/models/Qwen3-Dense.md:334
msgid ""
"This optimization requires setting the environment variable "
"`VLLM_ASCEND_ENABLE_FLASHCOMM1 = 1` to be enabled."
msgstr "此优化需要设置环境变量 `VLLM_ASCEND_ENABLE_FLASHCOMM1 = 1` 来启用。"
#: ../../source/tutorials/models/Qwen3-Dense.md:336
msgid "4. Matmul and ReduceScatter Fusion"
msgstr "4. 矩阵乘法和ReduceScatter融合"
#: ../../source/tutorials/models/Qwen3-Dense.md:338
msgid ""
"Once FlashComm_v1 is enabled, an additional optimization can be applied. "
"This optimization fuses matrix multiplication and ReduceScatter "
"operations, along with tiling optimization. The Matmul computation is "
"treated as one pipeline, while the ReduceScatter and dequant operations "
"are handled in a separate pipeline. This approach significantly reduces "
"communication steps, improves computational efficiency, and allows for "
"better resource utilization, resulting in enhanced throughput, especially"
" in large-scale distributed environments."
msgstr "一旦启用FlashComm_v1可以应用额外的优化。此优化融合了矩阵乘法和ReduceScatter操作并包含分片优化。矩阵乘法计算被视为一个流水线而ReduceScatter和反量化操作则在另一个独立的流水线中处理。这种方法显著减少了通信步骤提高了计算效率并实现了更好的资源利用从而提升了吞吐量尤其在大规模分布式环境中效果显著。"
#: ../../source/tutorials/models/Qwen3-Dense.md:340
msgid ""
"This optimization is automatically enabled once FlashComm_v1 is "
"activated. However, due to an issue with performance degradation in "
"small-concurrency scenarios after this fusion, a threshold-based approach"
" is currently used to mitigate this problem. The optimization is only "
"applied when the token count exceeds the threshold, ensuring that it is "
"not enabled in cases where it could negatively impact performance."
msgstr "此优化在FlashComm_v1激活后会自动启用。然而由于融合后在小并发场景下存在性能下降的问题目前采用基于阈值的方法来缓解此问题。该优化仅在令牌数超过阈值时应用确保在可能对性能产生负面影响的情况下不被启用。"
#: ../../source/tutorials/models/Qwen3-Dense.md:342
msgid "5. Weight Prefetching"
msgstr "5. 权重预取"
#: ../../source/tutorials/models/Qwen3-Dense.md:344
msgid ""
"Weight prefetching optimizes memory usage by preloading weights into the "
"cache before they are needed, minimizing delays caused by memory access "
"during model execution."
msgstr "权重预取通过在需要之前将权重预加载到缓存中来优化内存使用,从而最小化模型执行期间因内存访问造成的延迟。"
#: ../../source/tutorials/models/Qwen3-Dense.md:346
msgid ""
"In dense model scenarios, the MLP's gate_up_proj and down_proj linear "
"layers often exhibit relatively high MTE utilization. To address this, we"
" create a separate pipeline specifically for weight prefetching, which "
"runs in parallel with the original vector computation pipeline, such as "
"RMSNorm and SiLU, before the MLP. This approach allows the weights to be "
"preloaded to L2 cache ahead of time, reducing MTE utilization during the "
"MLP computations and indirectly improving Cube computation efficiency by "
"minimizing resource contention and optimizing data flow."
msgstr "在稠密模型场景中MLP的gate_up_proj和down_proj线性层通常表现出相对较高的MTE利用率。为解决此问题我们创建了一个专门用于权重预取的独立流水线该流水线与MLP之前的原始向量计算流水线如RMSNorm和SiLU并行运行。这种方法允许权重提前预加载到L2缓存中从而降低MLP计算期间的MTE利用率并通过最小化资源争用和优化数据流间接提升Cube计算效率。"
#: ../../source/tutorials/models/Qwen3-Dense.md:348
#, python-brace-format
msgid ""
"Previously, the environment variables VLLM_ASCEND_ENABLE_PREFETCH_MLP "
"used to enable MLP weight prefetch and "
"VLLM_ASCEND_MLP_GATE_UP_PREFETCH_SIZE and "
"VLLM_ASCEND_MLP_DOWN_PREFETCH_SIZE used to set the weight prefetch size "
"for MLP gate_up_proj and down_proj were deprecated. Please use the "
"following configuration instead: \"weight_prefetch_config\": { "
"\"enabled\": true, \"prefetch_ratio\": { \"mlp\": { \"gate_up\": 1.0, "
"\"down\": 1.0}}}. See User Guide->Feature Guide->Weight Prefetch Guide "
"for details."
msgstr "之前用于启用MLP权重预取的环境变量 `VLLM_ASCEND_ENABLE_PREFETCH_MLP`以及用于设置MLP gate_up_proj和down_proj权重预取大小的 `VLLM_ASCEND_MLP_GATE_UP_PREFETCH_SIZE` 和 `VLLM_ASCEND_MLP_DOWN_PREFETCH_SIZE` 已被弃用。请改用以下配置:`\"weight_prefetch_config\": { \"enabled\": true, \"prefetch_ratio\": { \"mlp\": { \"gate_up\": 1.0, \"down\": 1.0}}}`。详情请参阅用户指南->功能指南->权重预取指南。"
#: ../../source/tutorials/models/Qwen3-Dense.md:350
msgid "6. Zerolike Elimination"
msgstr "6. Zerolike消除"
#: ../../source/tutorials/models/Qwen3-Dense.md:352
msgid ""
"This elimination removes unnecessary operations related to zero-like "
"tensors in Attention forward, improving the efficiency of matrix "
"operations and reducing memory usage."
msgstr "此消除操作移除了Attention前向传播中与类零张量相关的不必要操作提高了矩阵运算效率并减少了内存使用。"
#: ../../source/tutorials/models/Qwen3-Dense.md:356
msgid "7. FullGraph Optimization"
msgstr "7. 全图优化"
#: ../../source/tutorials/models/Qwen3-Dense.md:358
msgid ""
"ACLGraph offers several key optimizations to improve model execution "
"efficiency. By replaying the entire model execution graph at once, we "
"significantly reduce dispatch latency compared to multiple smaller "
"replays. This approach also stabilizes multi-device performance, as "
"capturing the model as a single static graph mitigates dispatch "
"fluctuations across devices. Additionally, consolidating graph captures "
"frees up streams, allowing for the capture of more graphs and optimizing "
"resource usage, ultimately leading to improved system efficiency and "
"reduced overhead."
msgstr "ACLGraph提供了多项关键优化以提升模型执行效率。通过一次性重放整个模型执行图与多次重放较小图相比我们显著降低了调度延迟。这种方法还能稳定多设备性能因为将模型捕获为单个静态图可以缓解跨设备的调度波动。此外整合图捕获可以释放流从而允许捕获更多图并优化资源使用最终提高系统效率并减少开销。"
#: ../../source/tutorials/models/Qwen3-Dense.md:360
#, python-brace-format
msgid ""
"The configuration compilation_config = { \"cudagraph_mode\": "
"\"FULL_DECODE_ONLY\"} is used when starting the service. This setup is "
"necessary to enable the aclgraph's full decode-only mode."
msgstr "启动服务时使用配置 `compilation_config = { \"cudagraph_mode\": \"FULL_DECODE_ONLY\"}`。此设置对于启用aclgraph的完全仅解码模式是必需的。"
#: ../../source/tutorials/models/Qwen3-Dense.md:362
msgid "8. Asynchronous Scheduling"
msgstr "8. 异步调度"
#: ../../source/tutorials/models/Qwen3-Dense.md:364
msgid ""
"Asynchronous scheduling is a technique used to optimize inference "
"efficiency. It allows non-blocking task scheduling to improve concurrency"
" and throughput, especially when processing large-scale models."
msgstr "异步调度是一种用于优化推理效率的技术。它允许非阻塞的任务调度,以提高并发性和吞吐量,尤其是在处理大规模模型时。"
#: ../../source/tutorials/models/Qwen3-Dense.md:366
msgid "This optimization is enabled by setting `--async-scheduling`."
msgstr "此优化通过设置 `--async-scheduling` 来启用。"
#: ../../source/tutorials/models/Qwen3-Dense.md:368
msgid "Optimization Highlights"
msgstr "优化亮点"
#: ../../source/tutorials/models/Qwen3-Dense.md:370
msgid ""
"Building on the specific example scenarios outlined earlier, this section"
" highlights the key tuning points that played a crucial role in achieving"
" optimal performance. By focusing on the most impactful adjustments to "
"hyperparameters and optimizations, well emphasize the strategies that "
"can be leveraged to maximize throughput, minimize latency, and ensure "
"efficient resource utilization in various environments. These insights "
"will help guide you in fine-tuning your own configurations for the best "
"possible results."
msgstr "基于前面概述的具体示例场景,本节重点介绍在实现最佳性能中起关键作用的关键调优点。通过关注对超参数和优化最具影响力的调整,我们将强调可用于最大化吞吐量、最小化延迟并确保在各种环境中高效利用资源的策略。这些见解将帮助指导您微调自己的配置,以获得最佳结果。"
#: ../../source/tutorials/models/Qwen3-Dense.md:372
msgid "1.Prefetch Buffer Size"
msgstr "1. 预取缓冲区大小"
#: ../../source/tutorials/models/Qwen3-Dense.md:374
msgid ""
"Setting the right prefetch buffer size is essential for optimizing weight"
" loading and the size of this prefetch buffer is directly related to the "
"time that can be hidden by vector computations. To achieve a near-perfect"
" overlap between the prefetch and computation streams, you can flexibly "
"adjust the buffer size by profiling and observing the degree of overlap "
"at different buffer sizes."
msgstr "设置正确的预取缓冲区大小对于优化权重加载至关重要,且此预取缓冲区的大小与向量计算可隐藏的时间直接相关。为了实现预取流与计算流近乎完美的重叠,您可以通过性能分析和观察不同缓冲区大小下的重叠程度来灵活调整缓冲区大小。"
#: ../../source/tutorials/models/Qwen3-Dense.md:376
msgid ""
"For example, in the real-world scenario mentioned above, I set the "
"prefetch buffer size for the gate_up_proj and down_proj in the MLP to "
"18MB. The reason for this is that, at this value, the vector computations"
" of RMSNorm and SiLU can effectively hide the prefetch stream, thereby "
"accelerating the Matmul computations of the two linear layers."
msgstr ""
"例如在上述实际场景中我将MLP中gate_up_proj和down_proj的预取缓冲区大小设置为18MB。"
"这样做的原因是在此数值下RMSNorm和SiLU的向量计算能够有效隐藏预取流从而加速两个线性层的Matmul计算。"
#: ../../source/tutorials/models/Qwen3-Dense.md:378
msgid "2.Max-num-batched-tokens"
msgstr "2.最大批处理令牌数"
#: ../../source/tutorials/models/Qwen3-Dense.md:380
msgid ""
"The max-num-batched-tokens parameter determines the maximum number of "
"tokens that can be processed in a single batch. Adjusting this value "
"helps to balance throughput and memory usage. Setting this value too "
"small can negatively impact end-to-end performance, as fewer tokens are "
"processed per batch, potentially leading to inefficiencies. Conversely, "
"setting it too large increases the risk of Out of Memory (OOM) errors due"
" to excessive memory consumption."
msgstr ""
"最大批处理令牌数参数决定了单批次可处理的令牌数量上限。调整此值有助于平衡吞吐量与内存使用。"
"若设置过小,每批次处理的令牌数较少,可能降低效率,从而对端到端性能产生负面影响。"
"反之若设置过大则会因内存消耗过高而增加内存溢出OOM错误的风险。"
#: ../../source/tutorials/models/Qwen3-Dense.md:382
msgid ""
"In the above real-world scenario, we not only conducted extensive testing"
" to determine the most cost-effective value, but also took into account "
"the accumulation of decode tokens when enabling chunked prefill. If the "
"value is set too small, a single request may被分块多次并且在推理的早期阶段一个批次可能只包含少量解码令牌。这可能导致端到端吞吐量达不到预期。"
msgstr ""
"在上述实际场景中,我们不仅通过大量测试确定了最具性价比的数值,还考虑了启用分块预填充时解码令牌的累积问题。"
"若该值设置过小,单个请求可能被多次分块处理,且在推理早期阶段,单个批次可能仅包含少量解码令牌,从而导致端到端吞吐量无法达到预期。"
#: ../../source/tutorials/models/Qwen3-Dense.md:384
msgid "3.Cudagraph_capture_sizes"
msgstr "3.CUDA图捕获尺寸"
#: ../../source/tutorials/models/Qwen3-Dense.md:386
msgid ""
"The cudagraph_capture_sizes parameter controls the granularity of graph "
"captures during the inference process. Adjusting this value determines "
"how much of the computation graph is captured at once, which can "
"significantly impact both performance and memory usage."
msgstr ""
"CUDA图捕获尺寸参数控制推理过程中图捕获的粒度。调整此值决定了单次捕获的计算图范围这对性能和内存使用均有显著影响。"
#: ../../source/tutorials/models/Qwen3-Dense.md:388
msgid ""
"If this list is not manually specified, it will be filled with a series "
"of evenly distributed values, which typically ensures good performance. "
"However, if you want to fine-tune it further, manually specifying the "
"values will yield better results. This is because if the batch size falls"
" between two sizes, the framework will automatically pad the token count "
"to the larger size. This often leads to actual performance deviating from"
" the expected or even degrading."
msgstr ""
"若未手动指定此列表,系统将自动填充一系列均匀分布的值,这通常能保证良好性能。"
"但若需进一步微调,手动指定数值将获得更佳效果。这是因为当批次大小介于两个尺寸之间时,框架会自动将令牌数填充至较大尺寸,这常导致实际性能偏离预期甚至下降。"
#: ../../source/tutorials/models/Qwen3-Dense.md:390
msgid ""
"Therefore, like the above real-world scenario, when adjusting the "
"benchmark request concurrency, we always ensure that the concurrency is "
"actually included in the cudagraph_capture_sizes list. This way, during "
"the decode phase, padding operations are essentially avoided, ensuring "
"the reliability of the experimental data."
msgstr ""
"因此如上述实际场景所示在调整基准测试请求并发度时我们始终确保并发度实际包含在CUDA图捕获尺寸列表中。"
"这样在解码阶段基本避免了填充操作,从而保证了实验数据的可靠性。"
#: ../../source/tutorials/models/Qwen3-Dense.md:392
msgid ""
"It's important to note that if you enable FlashComm_v1, the values in "
"this list must be integer multiples of the TP size. Any values that do "
"not meet this condition will be automatically filtered out. Therefore, I "
"recommend incrementally adding concurrency based on the TP size after "
"enabling FlashComm_v1."
msgstr ""
"需特别注意若启用FlashComm_v1此列表中的值必须是TP尺寸的整数倍。不满足此条件的任何值都将被自动过滤。"
"因此建议在启用FlashComm_v1后基于TP尺寸逐步增加并发度。"

View File

@@ -0,0 +1,269 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3-Next.md:1
msgid "Qwen3-Next"
msgstr "Qwen3-Next"
#: ../../source/tutorials/models/Qwen3-Next.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3-Next.md:5
msgid ""
"The Qwen3-Next model is a sparse MoE (Mixture of Experts) model with high"
" sparsity. Compared to the MoE architecture of Qwen3, it has introduced "
"key improvements in aspects such as the hybrid attention mechanism and "
"multi-token prediction mechanism, enhancing the training and inference "
"efficiency of the model under long contexts and large total parameter "
"scales."
msgstr ""
"Qwen3-Next 模型是一个具有高稀疏性的稀疏 MoE专家混合模型。与 Qwen3 的 MoE 架构相比,它在混合注意力机制和多令牌预测机制等方面引入了关键改进,提升了模型在长上下文和大总参数量规模下的训练和推理效率。"
#: ../../source/tutorials/models/Qwen3-Next.md:7
msgid ""
"This document will present the core verification steps of the model, "
"including supported features, environment preparation, as well as "
"accuracy and performance evaluation. Qwen3 Next is currently using Triton"
" Ascend, which is in the experimental phase. In subsequent versions, its "
"performance related to stability and accuracy may change, and performance"
" will be continuously optimized."
msgstr ""
"本文档将介绍该模型的核心验证步骤包括支持的功能、环境准备以及精度和性能评估。Qwen3 Next 目前使用处于实验阶段的 Triton Ascend。在后续版本中其与稳定性和精度相关的表现可能会发生变化性能将持续优化。"
#: ../../source/tutorials/models/Qwen3-Next.md:9
msgid "The `Qwen3-Next` model is first supported in `vllm-ascend:v0.10.2rc1`."
msgstr "`Qwen3-Next` 模型首次在 `vllm-ascend:v0.10.2rc1` 中得到支持。"
#: ../../source/tutorials/models/Qwen3-Next.md:11
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/Qwen3-Next.md:13
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取模型支持的功能矩阵。"
#: ../../source/tutorials/models/Qwen3-Next.md:15
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[功能指南](../../user_guide/feature_guide/index.md)以获取功能的配置信息。"
#: ../../source/tutorials/models/Qwen3-Next.md:17
msgid "Weight Preparation"
msgstr "权重准备"
#: ../../source/tutorials/models/Qwen3-Next.md:19
msgid ""
"Download Link for the `Qwen3-Next-80B-A3B-Instruct` Model Weights: "
"[Download model weight](https://modelscope.cn/models/Qwen/Qwen3-Next-80B-"
"A3B-Instruct)"
msgstr "`Qwen3-Next-80B-A3B-Instruct` 模型权重下载链接:[下载模型权重](https://modelscope.cn/models/Qwen/Qwen3-Next-80B-A3B-Instruct)"
#: ../../source/tutorials/models/Qwen3-Next.md:21
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3-Next.md:23
msgid ""
"If the machine environment is an Atlas 800I A3(64G*16), the deployment "
"approach stays identical."
msgstr "如果机器环境是 Atlas 800I A3(64G*16),部署方法保持不变。"
#: ../../source/tutorials/models/Qwen3-Next.md:25
msgid "Run docker container"
msgstr "运行 Docker 容器"
#: ../../source/tutorials/models/Qwen3-Next.md:54
msgid ""
"The Qwen3 Next is using [Triton Ascend](https://gitee.com/ascend/triton-"
"ascend) which is currently experimental. In future versions, there may be"
" behavioral changes related to stability, accuracy, and performance "
"improvement."
msgstr "Qwen3 Next 正在使用目前处于实验阶段的 [Triton Ascend](https://gitee.com/ascend/triton-ascend)。在未来的版本中,可能会有与稳定性、精度和性能改进相关的行为变化。"
#: ../../source/tutorials/models/Qwen3-Next.md:56
msgid "Inference"
msgstr "推理"
#: ../../source/tutorials/models/Qwen3-Next.md
msgid "Online Inference"
msgstr "在线推理"
#: ../../source/tutorials/models/Qwen3-Next.md:62
msgid "Run the following script to start the vLLM server on multi-NPU:"
msgstr "运行以下脚本在多 NPU 上启动 vLLM 服务器:"
#: ../../source/tutorials/models/Qwen3-Next.md:68
msgid "Once your server is started, you can query the model with input prompts."
msgstr "服务器启动后,您可以使用输入提示词查询模型。"
#: ../../source/tutorials/models/Qwen3-Next.md
msgid "Offline Inference"
msgstr "离线推理"
#: ../../source/tutorials/models/Qwen3-Next.md:87
msgid "Run the following script to execute offline inference on multi-NPU:"
msgstr "运行以下脚本在多 NPU 上执行离线推理:"
#: ../../source/tutorials/models/Qwen3-Next.md:125
msgid "If you run this script successfully, you can see the info shown below:"
msgstr "如果成功运行此脚本,您将看到如下信息:"
#: ../../source/tutorials/models/Qwen3-Next.md:133
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/Qwen3-Next.md:135
#: ../../source/tutorials/models/Qwen3-Next.md:147
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/models/Qwen3-Next.md:137
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参考[使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/Qwen3-Next.md:139
msgid ""
"After execution, you can get the result, here is the result of `Qwen3"
"-Next-80B-A3B-Instruct` in `vllm-ascend:0.13.0rc1` for reference only."
msgstr "执行后,您可以获得结果,以下是 `vllm-ascend:0.13.0rc1` 中 `Qwen3-Next-80B-A3B-Instruct` 的结果,仅供参考。"
#: ../../source/tutorials/models/Qwen3-Next.md:85
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/models/Qwen3-Next.md:85
msgid "version"
msgstr "版本"
#: ../../source/tutorials/models/Qwen3-Next.md:85
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/models/Qwen3-Next.md:85
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/models/Qwen3-Next.md:85
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/models/Qwen3-Next.md:85
msgid "gsm8k"
msgstr "gsm8k"
#: ../../source/tutorials/models/Qwen3-Next.md:85
msgid "-"
msgstr "-"
#: ../../source/tutorials/models/Qwen3-Next.md:85
msgid "accuracy"
msgstr "准确率"
#: ../../source/tutorials/models/Qwen3-Next.md:85
msgid "gen"
msgstr "生成"
#: ../../source/tutorials/models/Qwen3-Next.md:85
msgid "95.53"
msgstr "95.53"
#: ../../source/tutorials/models/Qwen3-Next.md:145
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen3-Next.md:149
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参考[使用 AISBench 进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/Qwen3-Next.md:151
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM Benchmark"
#: ../../source/tutorials/models/Qwen3-Next.md:153
msgid "Run performance evaluation of `Qwen3-Next` as an example."
msgstr "以运行 `Qwen3-Next` 的性能评估为例。"
#: ../../source/tutorials/models/Qwen3-Next.md:155
msgid ""
"Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) "
"for more details."
msgstr "更多详情请参考 [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/models/Qwen3-Next.md:157
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 有三个子命令:"
#: ../../source/tutorials/models/Qwen3-Next.md:159
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批请求的延迟进行基准测试。"
#: ../../source/tutorials/models/Qwen3-Next.md:160
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen3-Next.md:161
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen3-Next.md:163
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。按如下方式运行代码。"
#: ../../source/tutorials/models/Qwen3-Next.md:170
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您将获得性能评估结果。"
#: ../../source/tutorials/models/Qwen3-Next.md:172
msgid "The performance result is:"
msgstr "性能结果如下:"
#: ../../source/tutorials/models/Qwen3-Next.md:174
msgid "**Hardware**: A3-752T, 2 node"
msgstr "**硬件**A3-752T2 节点"
#: ../../source/tutorials/models/Qwen3-Next.md:176
msgid "**Deployment**: TP4 + Full Decode Only"
msgstr "**部署**TP4 + 仅全解码"
#: ../../source/tutorials/models/Qwen3-Next.md:178
msgid "**Input/Output**: 2k/2k"
msgstr "**输入/输出**2k/2k"
#: ../../source/tutorials/models/Qwen3-Next.md:180
msgid "**Concurrency**: 32"
msgstr "**并发数**32"
#: ../../source/tutorials/models/Qwen3-Next.md:182
msgid "**Performance**: 580tps, TPOT 54ms"
msgstr "**性能**580tpsTPOT 54ms"

View File

@@ -0,0 +1,230 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:1
msgid "Qwen3-Omni-30B-A3B-Thinking"
msgstr "Qwen3-Omni-30B-A3B-Thinking"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:5
msgid ""
"Qwen3-Omni is the natively end-to-end multilingual omni-modal foundation "
"models. It processes text, images, audio, and video, and delivers real-"
"time streaming responses in both text and natural speech. We introduce "
"several architectural upgrades to improve performance and efficiency. The"
" Thinking model of Qwen3-Omni-30B-A3B, containing the thinker component, "
"equipped with chain-of-thought reasoning, supporting audio, video, and "
"text input, with text output."
msgstr ""
"Qwen3-Omni 是原生端到端多语言全模态基础模型。它能处理文本、图像、音频和视频并以文本和自然语音形式提供实时流式响应。我们引入了多项架构升级以提升性能和效率。Qwen3-Omni-30B-A3B 的 Thinking 模型包含思考器组件,具备思维链推理能力,支持音频、视频和文本输入,输出为文本。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:7
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, single-node deployment, accuracy and performance evaluation."
msgstr "本文档将展示该模型的主要验证步骤,包括支持的功能、功能配置、环境准备、单节点部署、精度和性能评估。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:9
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:11
msgid ""
"Refer to [supported features](https://docs.vllm.ai/projects/ascend/zh-"
"cn/latest/user_guide/support_matrix/supported_models.html) to get the "
"model's supported feature matrix."
msgstr "请参考 [支持的功能](https://docs.vllm.ai/projects/ascend/zh-cn/latest/user_guide/support_matrix/supported_models.html) 以获取模型支持的功能矩阵。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:13
msgid ""
"Refer to [feature guide](https://docs.vllm.ai/projects/ascend/zh-"
"cn/latest/user_guide/feature_guide/index.html) to get the feature's "
"configuration."
msgstr "请参考 [功能指南](https://docs.vllm.ai/projects/ascend/zh-cn/latest/user_guide/feature_guide/index.html) 以获取功能的配置信息。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:15
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:17
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:19
msgid ""
"`Qwen3-Omni-30B-A3B-Thinking` requires 2 NPU Cards(64G × 2).[Download "
"model weight](https://modelscope.cn/models/Qwen/Qwen3-Omni-30B-A3B-"
"Thinking) It is recommended to download the model weight to the shared "
"directory of multiple nodes, such as `/root/.cache/`"
msgstr ""
"`Qwen3-Omni-30B-A3B-Thinking` 需要 2 张 NPU 卡 (64G × 2)。[下载模型权重](https://modelscope.cn/models/Qwen/Qwen3-Omni-30B-A3B-Thinking)。建议将模型权重下载到多节点的共享目录,例如 `/root/.cache/`。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:22
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md
msgid "Use docker image"
msgstr "使用 Docker 镜像"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:28
msgid ""
"You can use our official docker image to run Qwen3-Omni-30B-A3B-Thinking "
"directly"
msgstr "您可以使用我们的官方 Docker 镜像直接运行 Qwen3-Omni-30B-A3B-Thinking"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:30
msgid ""
"Select an image based on your machine type and start the docker image on "
"your node, refer to [using docker](../../installation.md#set-up-using-"
"docker)."
msgstr "根据您的机器类型选择镜像并在节点上启动 Docker 镜像,请参考 [使用 Docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md
msgid "Build from source"
msgstr "从源码构建"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:65
msgid "You can build all from source."
msgstr "您可以从源码构建所有组件。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:67
msgid ""
"Install `vllm-ascend`, refer to [set up using "
"python](../../installation.md#set-up-using-python)."
msgstr "安装 `vllm-ascend`,请参考 [使用 Python 设置](../../installation.md#set-up-using-python)。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:71
msgid "Please install system dependencies"
msgstr "请安装系统依赖"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:81
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:83
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:85
msgid "Offline Inference on Multi-NPU"
msgstr "多 NPU 离线推理"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:87
msgid "Run the following script to execute offline inference on multi-NPU:"
msgstr "运行以下脚本在多 NPU 上执行离线推理:"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:174
msgid "Online Inference on Multi-NPU"
msgstr "多 NPU 在线推理"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:176
msgid ""
"Run the following script to start the vLLM server on Multi-NPU: For an "
"Atlas A2 with 64 GB of NPU card memory, tensor-parallel-size should be at"
" least 1, and for 32 GB of memory, tensor-parallel-size should be at "
"least 2."
msgstr "运行以下脚本在多 NPU 上启动 vLLM 服务器:对于具有 64 GB NPU 卡内存的 Atlas A2tensor-parallel-size 应至少为 1对于 32 GB 内存tensor-parallel-size 应至少为 2。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:188
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:190
msgid "Once your server is started, you can query the model with input prompts."
msgstr "服务器启动后,您可以使用输入提示词查询模型。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:231
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:233
msgid "Here are accuracy evaluation methods."
msgstr "以下是精度评估方法。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:235
msgid "Using EvalScope"
msgstr "使用 EvalScope"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:237
msgid ""
"As an example, take the `gsm8k` `omnibench` `bbh` dataset as a test "
"dataset, and run accuracy evaluation of `Qwen3-Omni-30B-A3B-Thinking` in "
"online mode."
msgstr "以 `gsm8k`、`omnibench`、`bbh` 数据集作为测试数据集为例,在在线模式下运行 `Qwen3-Omni-30B-A3B-Thinking` 的精度评估。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:239
msgid ""
"Refer to Using "
"evalscope(<https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/evaluation/using_evalscope.html"
"#install-evalscope-using-pip>) for `evalscope`installation."
msgstr "关于 `evalscope` 的安装,请参考使用 evalscope (<https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/evaluation/using_evalscope.html#install-evalscope-using-pip>)。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:240
msgid "Run `evalscope` to execute the accuracy evaluation."
msgstr "运行 `evalscope` 以执行精度评估。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:255
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:296
msgid ""
"After execution, you can get the result, here is the result of `Qwen3"
"-Omni-30B-A3B-Thinking` in vllm-ascend:0.13.0rc1 for reference only."
msgstr "执行后,您可以获得结果。以下是 `Qwen3-Omni-30B-A3B-Thinking` 在 vllm-ascend:0.13.0rc1 中的结果,仅供参考。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:269
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:271
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM 基准测试"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:273
msgid ""
"Run performance evaluation of `Qwen3-Omni-30B-A3B-Thinking` as an "
"example. Refer to vllm benchmark for more details. Refer to [vllm "
"benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more "
"details."
msgstr "以运行 `Qwen3-Omni-30B-A3B-Thinking` 的性能评估为例。更多详情请参考 vllm 基准测试。更多详情请参考 [vllm 基准测试](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:277
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 有三个子命令:"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:279
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批次请求的延迟进行基准测试。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:280
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:281
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md:283
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。按如下方式运行代码。"

View File

@@ -0,0 +1,433 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:1
msgid "Qwen3-VL-235B-A22B-Instruct"
msgstr "Qwen3-VL-235B-A22B-Instruct"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:5
msgid ""
"The Qwen-VL(Vision-Language)series from Alibaba Cloud comprises a family "
"of powerful Large Vision-Language Models (LVLMs) designed for "
"comprehensive multimodal understanding. They accept images, text, and "
"bounding boxes as input, and output text and detection boxes, enabling "
"advanced functions like image detection, multi-modal dialogue, and multi-"
"image reasoning."
msgstr ""
"阿里云的Qwen-VL视觉-语言系列包含一系列强大的大型视觉语言模型LVLM专为全面的多模态理解而设计。它们接受图像、文本和边界框作为输入并输出文本和检测框从而实现图像检测、多模态对话和多图像推理等高级功能。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:7
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, NPU deployment, accuracy and performance evaluation."
msgstr "本文档将展示该模型的主要验证步骤包括支持的功能、功能配置、环境准备、NPU部署、精度和性能评估。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:9
msgid ""
"This tutorial uses the vLLM-Ascend `v0.11.0rc2` version for "
"demonstration, showcasing the `Qwen3-VL-235B-A22B-Instruct` model as an "
"example for multi-NPU deployment."
msgstr "本教程使用 vLLM-Ascend `v0.11.0rc2` 版本进行演示,以 `Qwen3-VL-235B-A22B-Instruct` 模型为例展示多NPU部署。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:11
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:13
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取模型支持的功能矩阵。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:15
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[功能指南](../../user_guide/feature_guide/index.md)以获取功能的配置信息。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:17
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:19
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:21
msgid ""
"`Qwen3-VL-235B-A22B-Instruct`(BF16 version): require 1 Atlas 800 A3 (64G "
"× 16) node2 Atlas 800 A264G × 8nodes. [Download model "
"weight](https://modelscope.cn/models/Qwen/Qwen3-VL-235B-A22B-Instruct/)"
msgstr ""
"`Qwen3-VL-235B-A22B-Instruct`BF16版本需要1个Atlas 800 A364G × 16节点2个Atlas 800 A264G × 8节点。[下载模型权重](https://modelscope.cn/models/Qwen/Qwen3-VL-235B-A22B-Instruct/)"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:23
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`"
msgstr "建议将模型权重下载到多个节点的共享目录中,例如 `/root/.cache/`"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:25
msgid "Verify Multi-node Communication(Optional)"
msgstr "验证多节点通信(可选)"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:27
msgid ""
"If you want to deploy multi-node environment, you need to verify multi-"
"node communication according to [verify multi-node communication "
"environment](../../installation.md#verify-multi-node-communication)."
msgstr "如果您想部署多节点环境,需要根据[验证多节点通信环境](../../installation.md#verify-multi-node-communication)来验证多节点通信。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:29
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md
msgid "Use docker image"
msgstr "使用Docker镜像"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:35
msgid ""
"For example, using images `quay.io/ascend/vllm-ascend:v0.11.0rc2`(for "
"Atlas 800 A2) and `quay.io/ascend/vllm-ascend:v0.11.0rc2-a3`(for Atlas "
"800 A3)."
msgstr "例如,使用镜像 `quay.io/ascend/vllm-ascend:v0.11.0rc2`适用于Atlas 800 A2和 `quay.io/ascend/vllm-ascend:v0.11.0rc2-a3`适用于Atlas 800 A3。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:37
msgid ""
"Select an image based on your machine type and start the docker image on "
"your node, refer to [using docker](../../installation.md#set-up-using-"
"docker)."
msgstr "根据您的机器类型选择镜像并在节点上启动Docker镜像请参考[使用Docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md
msgid "Build from source"
msgstr "从源码构建"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:78
msgid "You can build all from source."
msgstr "您可以从源码构建所有组件。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:80
msgid ""
"Install `vllm-ascend`, refer to [set up using "
"python](../../installation.md#set-up-using-python)."
msgstr "安装 `vllm-ascend`,请参考[使用Python设置](../../installation.md#set-up-using-python)。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:84
msgid ""
"If you want to deploy multi-node environment, you need to set up "
"environment on each node."
msgstr "如果您想部署多节点环境,需要在每个节点上设置环境。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:86
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:88
msgid "Multi-node Deployment with MP (Recommended)"
msgstr "使用MP进行多节点部署推荐"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:90
msgid ""
"Assume you have Atlas 800 A3 (64G*16) nodes (or 2* A2), and want to "
"deploy the `Qwen3-VL-235B-A22B-Instruct` model across multiple nodes."
msgstr "假设您拥有Atlas 800 A364G*16节点或2个A2节点并希望跨多个节点部署 `Qwen3-VL-235B-A22B-Instruct` 模型。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:92
msgid "Node 0"
msgstr "节点 0"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:135
msgid "Node1"
msgstr "节点 1"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:182
msgid "The parameters are explained as follows:"
msgstr "参数解释如下:"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:184
msgid ""
"`--max-model-len` represents the context length, which is the maximum "
"value of the input plus output for a single request."
msgstr "`--max-model-len` 表示上下文长度,即单个请求的输入加输出的最大值。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:185
msgid ""
"`--max-num-seqs` indicates the maximum number of requests that each DP "
"group is allowed to process. If the number of requests sent to the "
"service exceeds this limit, the excess requests will remain in a waiting "
"state and will not be scheduled. Note that the time spent in the waiting "
"state is also counted in metrics such as TTFT and TPOT. Therefore, when "
"testing performance, it is generally recommended that `--max-num-seqs` * "
"`--data-parallel-size` >= the actual total concurrency."
msgstr ""
"`--max-num-seqs` 表示每个DP组允许处理的最大请求数。如果发送到服务的请求数超过此限制超出的请求将保持在等待状态不会被调度。请注意等待状态所花费的时间也会计入TTFT和TPOT等指标。因此在测试性能时通常建议 `--max-num-seqs` * `--data-parallel-size` >= 实际总并发数。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:186
msgid ""
"`--max-num-batched-tokens` represents the maximum number of tokens that "
"the model can process in a single step. Currently, vLLM v1 scheduling "
"enables ChunkPrefill/SplitFuse by default, which means:"
msgstr "`--max-num-batched-tokens` 表示模型在单步中可以处理的最大token数。目前vLLM v1调度默认启用ChunkPrefill/SplitFuse这意味着"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:187
msgid ""
"(1) If the input length of a request is greater than `--max-num-batched-"
"tokens`, it will be divided into multiple rounds of computation according"
" to `--max-num-batched-tokens`;"
msgstr "1如果请求的输入长度大于 `--max-num-batched-tokens`,它将根据 `--max-num-batched-tokens` 被分成多轮计算;"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:188
msgid ""
"(2) Decode requests are prioritized for scheduling, and prefill requests "
"are scheduled only if there is available capacity."
msgstr "2解码请求优先被调度而预填充请求仅在有空闲容量时才会被调度。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:189
msgid ""
"Generally, if `--max-num-batched-tokens` is set to a larger value, the "
"overall latency will be lower, but the pressure on GPU memory (activation"
" value usage) will be greater."
msgstr "通常,如果将 `--max-num-batched-tokens` 设置为较大的值整体延迟会更低但GPU内存激活值使用的压力会更大。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:190
msgid ""
"`--gpu-memory-utilization` represents the proportion of HBM that vLLM "
"will use for actual inference. Its essential function is to calculate the"
" available kv_cache size. During the warm-up phase (referred to as "
"profile run in vLLM), vLLM records the peak GPU memory usage during an "
"inference process with an input size of `--max-num-batched-tokens`. The "
"available kv_cache size is then calculated as: `--gpu-memory-utilization`"
" * HBM size - peak GPU memory usage. Therefore, the larger the value of "
"`--gpu-memory-utilization`, the more kv_cache can be used. However, since"
" the GPU memory usage during the warm-up phase may differ from that "
"during actual inference (e.g., due to uneven EP load), setting `--gpu-"
"memory-utilization` too high may lead to OOM (Out of Memory) issues "
"during actual inference. The default value is `0.9`."
msgstr ""
"`--gpu-memory-utilization` 表示vLLM将用于实际推理的HBM比例。其主要功能是计算可用的kv_cache大小。在预热阶段在vLLM中称为profile runvLLM会记录输入大小为 `--max-num-batched-tokens` 的推理过程中的峰值GPU内存使用量。然后可用的kv_cache大小计算为`--gpu-memory-utilization` * HBM大小 - 峰值GPU内存使用量。因此`--gpu-memory-utilization` 的值越大可用的kv_cache就越多。然而由于预热阶段的GPU内存使用量可能与实际推理阶段不同例如由于EP负载不均衡将 `--gpu-memory-utilization` 设置得过高可能导致实际推理时出现OOM内存不足问题。默认值为 `0.9`。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:191
msgid ""
"`--enable-expert-parallel` indicates that EP is enabled. Note that vLLM "
"does not support a mixed approach of ETP and EP; that is, MoE can either "
"use pure EP or pure TP."
msgstr "`--enable-expert-parallel` 表示启用了EP。请注意vLLM不支持ETP和EP的混合方法也就是说MoE只能使用纯EP或纯TP。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:192
msgid ""
"`--no-enable-prefix-caching` indicates that prefix caching is disabled. "
"To enable it, remove this option."
msgstr "`--no-enable-prefix-caching` 表示前缀缓存被禁用。要启用它,请移除此选项。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:193
msgid ""
"`--quantization` \"ascend\" indicates that quantization is used. To "
"disable quantization, remove this option."
msgstr "`--quantization` \"ascend\" 表示使用了量化。要禁用量化,请移除此选项。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:194
msgid ""
"`--compilation-config` contains configurations related to the aclgraph "
"graph mode. The most significant configurations are \"cudagraph_mode\" "
"and \"cudagraph_capture_sizes\", which have the following meanings: "
"\"cudagraph_mode\": represents the specific graph mode. Currently, "
"\"PIECEWISE\" and \"FULL_DECODE_ONLY\" are supported. The graph mode is "
"mainly used to reduce the cost of operator dispatch. Currently, "
"\"FULL_DECODE_ONLY\" is recommended."
msgstr ""
"`--compilation-config` 包含与aclgraph图模式相关的配置。最重要的配置是 \"cudagraph_mode\" 和 \"cudagraph_capture_sizes\",其含义如下:\"cudagraph_mode\":表示特定的图模式。目前支持 \"PIECEWISE\" 和 \"FULL_DECODE_ONLY\"。图模式主要用于降低算子调度的开销。目前推荐使用 \"FULL_DECODE_ONLY\"。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:196
msgid ""
"\"cudagraph_capture_sizes\": represents different levels of graph modes. "
"The default value is [1, 2, 4, 8, 16, 24, 32, 40,..., `--max-num-seqs`]. "
"In the graph mode, the input for graphs at different levels is fixed, and"
" inputs between levels are automatically padded to the next level. "
"Currently, the default setting is recommended. Only in some scenarios is "
"it necessary to set this separately to achieve optimal performance."
msgstr ""
"\"cudagraph_capture_sizes\":表示不同级别的图模式。默认值为 [1, 2, 4, 8, 16, 24, 32, 40,..., `--max-num-seqs`]。在图模式下,不同级别图的输入是固定的,级别之间的输入会自动填充到下一个级别。目前推荐使用默认设置。仅在部分场景中需要单独设置此参数以达到最佳性能。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:197
msgid ""
"`export VLLM_ASCEND_ENABLE_FLASHCOMM1=1` indicates that Flashcomm1 "
"optimization is enabled. Currently, this optimization is only supported "
"for MoE in scenarios where tp_size > 1."
msgstr "`export VLLM_ASCEND_ENABLE_FLASHCOMM1=1` 表示启用了Flashcomm1优化。目前此优化仅在 tp_size > 1 的场景中支持MoE。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:199
msgid ""
"If the service starts successfully, the following information will be "
"displayed on node 0:"
msgstr "如果服务启动成功节点0上将显示以下信息"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:210
msgid "Multi-node Deployment with Ray"
msgstr "使用Ray进行多节点部署"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:212
msgid "refer to [Ray Distributed (Qwen/Qwen3-235B-A22B)](../features/ray.md)."
msgstr "请参考[Ray分布式Qwen/Qwen3-235B-A22B](../features/ray.md)。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:214
msgid "Prefill-Decode Disaggregation"
msgstr "预填充-解码解耦"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:216
msgid ""
"refer to [Prefill-Decode Disaggregation Mooncake "
"Verification](../features/pd_disaggregation_mooncake_multi_node.md)"
msgstr "请参考[预填充-解码解耦月饼验证](../features/pd_disaggregation_mooncake_multi_node.md)"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:218
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:220
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "一旦您的服务器启动,您可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:237
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:239
msgid "Here are two accuracy evaluation methods."
msgstr "这里有两种精度评估方法。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:241
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:253
msgid "Using AISBench"
msgstr "使用AISBench"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:243
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参考[使用AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:245
msgid ""
"After execution, you can get the result, here is the result of `Qwen3-VL-"
"235B-A22B-Instruct` in `vllm-ascend:0.11.0rc2` for reference only."
msgstr "执行后,您可以获得结果,以下是 `Qwen3-VL-235B-A22B-Instruct` 在 `vllm-ascend:0.11.0rc2` 中的结果,仅供参考。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:76
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:76
msgid "version"
msgstr "版本"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:76
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:76
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:76
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:76
msgid "aime2024"
msgstr "aime2024"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:76
msgid "-"
msgstr "-"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:76
msgid "accuracy"
msgstr "准确率"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:76
msgid "gen"
msgstr "生成"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:76
msgid "93"
msgstr "93"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:251
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:255
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr ""
"详情请参阅[使用 AISBench 进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:257
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM Benchmark"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:259
msgid "Run performance evaluation of `Qwen3-VL-235B-A22B-Instruct` as an example."
msgstr "以运行 `Qwen3-VL-235B-A22B-Instruct` 的性能评估为例。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:261
msgid ""
"Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) "
"for more details."
msgstr ""
"更多详情请参阅 [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/)。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:263
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 包含三个子命令:"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:265
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批次请求的延迟进行基准测试。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:266
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:267
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:269
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例,按如下方式运行代码。"
#: ../../source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md:276
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您将获得性能评估结果。"

View File

@@ -0,0 +1,199 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:1
msgid "Qwen3-VL-30B-A3B-Instruct"
msgstr "Qwen3-VL-30B-A3B-Instruct"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:5
msgid ""
"The Qwen-VL (Vision-Language) series from Alibaba Cloud comprises a "
"family of powerful Large Vision-Language Models (LVLMs) designed for "
"comprehensive multimodal understanding. They accept images, text, and "
"bounding boxes as input, and output text and detection boxes, enabling "
"advanced functions like image detection, multi-modal dialogue, and multi-"
"image reasoning."
msgstr ""
"阿里云的 Qwen-VL视觉-语言系列包含一系列强大的大型视觉语言模型LVLM专为全面的多模态理解而设计。它们接受图像、文本和边界框作为输入并输出文本和检测框从而实现图像检测、多模态对话和多图像推理等高级功能。"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:7
msgid ""
"This document will show the main verification steps of the `Qwen3-VL-30B-"
"A3B-Instruct`."
msgstr "本文档将展示 `Qwen3-VL-30B-A3B-Instruct` 的主要验证步骤。"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:9
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:11
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取模型支持的功能矩阵。"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:12
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[功能指南](../../user_guide/feature_guide/index.md)以获取功能的配置信息。"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:14
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:16
msgid "Prepare Model Weights"
msgstr "准备模型权重"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:18
msgid ""
"Running this model requires 1 Atlas 800I A2 (64G × 8) node or 1 Atlas 800"
" A3 (64G × 16) node."
msgstr "运行此模型需要 1 个 Atlas 800I A2 (64G × 8) 节点或 1 个 Atlas 800 A3 (64G × 16) 节点。"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:20
msgid ""
"Download model weight at [ModelScope "
"Website](https://modelscope.cn/models/Qwen/Qwen3-VL-30B-A3B-Instruct) or "
"download by below command:"
msgstr "从 [ModelScope 网站](https://modelscope.cn/models/Qwen/Qwen3-VL-30B-A3B-Instruct) 下载模型权重,或使用以下命令下载:"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:27
msgid ""
"It is recommended to download the model weights to the shared directory "
"of multiple nodes, such as `/root/.cache/`."
msgstr "建议将模型权重下载到多个节点的共享目录中,例如 `/root/.cache/`。"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:29
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:31
msgid "Run docker container:"
msgstr "运行 Docker 容器:"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:58
msgid "Setup environment variables:"
msgstr "设置环境变量:"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:69
msgid ""
"`max_split_size_mb` prevents the native allocator from splitting blocks "
"larger than this size (in MB). This can reduce fragmentation and may "
"allow some borderline workloads to complete without running out of "
"memory. You can find more details "
"[<u>here</u>](https://www.hiascend.com/document/detail/zh/CANNCommunityEdition/800alpha003/apiref/envref/envref_07_0061.html)."
msgstr ""
"`max_split_size_mb` 可防止原生分配器拆分大于此大小(以 MB 为单位)的内存块。这可以减少内存碎片,并可能使一些临界工作负载在内存耗尽前完成。您可以在[<u>此处</u>](https://www.hiascend.com/document/detail/zh/CANNCommunityEdition/800alpha003/apiref/envref/envref_07_0061.html)找到更多详细信息。"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:72
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:74
msgid "Online Serving"
msgstr "在线服务"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md
msgid "Image Inputs"
msgstr "图像输入"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:83
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:146
msgid ""
"Run the following command inside the container to start the vLLM server "
"on multi-NPU:"
msgstr "在容器内运行以下命令以在多 NPU 上启动 vLLM 服务器:"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:95
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:157
msgid ""
"vllm-ascend supports Expert Parallelism (EP) via `--enable-expert-"
"parallel`, which allows experts in MoE models to be deployed on separate "
"GPUs for better throughput."
msgstr "vllm-ascend 通过 `--enable-expert-parallel` 支持专家并行EP这允许将 MoE 模型中的专家部署在单独的 GPU 上以获得更好的吞吐量。"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:97
msgid ""
"It's highly recommended to specify `--limit-mm-per-prompt.video 0` if "
"your inference server will only process image inputs since enabling video"
" inputs consumes more memory reserved for long video embeddings."
msgstr "如果您的推理服务器仅处理图像输入,强烈建议指定 `--limit-mm-per-prompt.video 0`,因为启用视频输入会消耗更多为长视频嵌入保留的内存。"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:99
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:159
msgid ""
"You can set `--max-model-len` to preserve memory. By default the model's "
"context length is 262K, but `--max-model-len 128000` is good for most "
"scenarios."
msgstr "您可以设置 `--max-model-len` 以节省内存。默认情况下,模型的上下文长度为 262K但 `--max-model-len 128000` 适用于大多数场景。"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:102
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:164
msgid "If your service start successfully, you can see the info shown below:"
msgstr "如果您的服务启动成功,您可以看到如下所示的信息:"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:110
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:172
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "服务器启动后,您可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:128
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:190
msgid ""
"If you query the server successfully, you can see the info shown below "
"(client):"
msgstr "如果您成功查询服务器,您可以看到如下所示的信息(客户端):"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:134
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:196
msgid "Logs of the vllm server:"
msgstr "vllm 服务器的日志:"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md
msgid "Video Inputs"
msgstr "视频输入"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:161
msgid ""
"Set `--allowed-local-media-path /media` to use your local video that "
"located at `/media`, since directly download the video during serving can"
" be extremely slow due to network issues."
msgstr "设置 `--allowed-local-media-path /media` 以使用位于 `/media` 的本地视频,因为在服务期间直接下载视频可能因网络问题而极其缓慢。"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:205
msgid "Offline Inference"
msgstr "离线推理"
#: ../../source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md:207
msgid ""
"The usage of offline inference with `Qwen3-VL-30B-A3B-Instruct` is "
"totally the same as that of `Qwen3-VL-8B-Instruct`, find more details at "
"[Qwen3-VL-8B-"
"Instruct](https://docs.vllm.ai/projects/ascend/en/latest/tutorials/models"
"/Qwen-VL-Dense.html#offline-inference)."
msgstr "`Qwen3-VL-30B-A3B-Instruct` 的离线推理使用方法与 `Qwen3-VL-8B-Instruct` 完全相同,更多详细信息请参阅 [Qwen3-VL-8B-Instruct](https://docs.vllm.ai/projects/ascend/en/latest/tutorials/models/Qwen-VL-Dense.html#offline-inference)。"

View File

@@ -0,0 +1,172 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:1
msgid "Qwen3-VL-Embedding"
msgstr "Qwen3-VL-Embedding"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:5
msgid ""
"The Qwen3-VL-Embedding and Qwen3-VL-Reranker model series are the latest "
"additions to the Qwen family, built upon the recently open-sourced and "
"powerful Qwen3-VL foundation model. Specifically designed for multimodal "
"information retrieval and cross-modal understanding, this suite accepts "
"diverse inputs including text, images, screenshots, and videos, as well "
"as inputs containing a mixture of these modalities. This guide describes "
"how to run the model with vLLM Ascend."
msgstr ""
"Qwen3-VL-Embedding 和 Qwen3-VL-Reranker 模型系列是 Qwen 家族的最新成员,基于最近开源且强大的 Qwen3-VL 基础模型构建。该系列专为多模态信息检索和跨模态理解而设计,可接受包括文本、图像、截图和视频在内的多样化输入,以及包含这些模态混合的输入。本指南描述了如何使用 vLLM Ascend 运行该模型。"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:7
msgid "Supported Features"
msgstr "支持特性"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:9
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持特性](../../user_guide/support_matrix/supported_models.md)以获取模型的支持特性矩阵。"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:11
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:13
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:15
msgid ""
"`Qwen3-VL-Embedding-8B` [Download model "
"weight](https://www.modelscope.cn/models/Qwen/Qwen3-VL-Embedding-8B)"
msgstr ""
"`Qwen3-VL-Embedding-8B` [下载模型权重](https://www.modelscope.cn/models/Qwen/Qwen3-VL-Embedding-8B)"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:16
msgid ""
"`Qwen3-VL-Embedding-2B` [Download model "
"weight](https://www.modelscope.cn/models/Qwen/Qwen3-VL-Embedding-2B)"
msgstr ""
"`Qwen3-VL-Embedding-2B` [下载模型权重](https://www.modelscope.cn/models/Qwen/Qwen3-VL-Embedding-2B)"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:18
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`"
msgstr "建议将模型权重下载到多个节点的共享目录中,例如 `/root/.cache/`"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:20
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:22
msgid ""
"You can use our official docker image to run `Qwen3-VL-Embedding` series "
"models."
msgstr "您可以使用我们的官方 docker 镜像来运行 `Qwen3-VL-Embedding` 系列模型。"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:24
msgid ""
"Start the docker image on your node, refer to [using "
"docker](../../installation.md#set-up-using-docker)."
msgstr "在您的节点上启动 docker 镜像,请参考[使用 docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:26
msgid ""
"If you don't want to use the docker image as above, you can also build "
"all from source:"
msgstr "如果您不想使用上述 docker 镜像,也可以从源码构建所有内容:"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:28
msgid ""
"Install `vllm-ascend` from source, refer to "
"[installation](../../installation.md)."
msgstr "从源码安装 `vllm-ascend`,请参考[安装指南](../../installation.md)。"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:30
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:32
msgid ""
"Using the Qwen3-VL-Embedding-8B model as an example, first run the docker"
" container with the following command:"
msgstr "以 Qwen3-VL-Embedding-8B 模型为例,首先使用以下命令运行 docker 容器:"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:34
msgid "Online Inference"
msgstr "在线推理"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:40
msgid "Once your server is started, you can query the model with input prompts."
msgstr "服务器启动后,您可以使用输入提示词查询模型。"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:51
msgid "Offline Inference"
msgstr "离线推理"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:86
msgid "If you run this script successfully, you can see the info shown below:"
msgstr "如果成功运行此脚本,您将看到如下所示的信息:"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:96
msgid "For more examples, refer to the vLLM official examples:"
msgstr "更多示例,请参考 vLLM 官方示例:"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:98
msgid ""
"[Offline Vision Embedding Example](https://github.com/vllm-"
"project/vllm/blob/main/examples/pooling/embed/vision_embedding_offline.py)"
msgstr ""
"[离线视觉嵌入示例](https://github.com/vllm-project/vllm/blob/main/examples/pooling/embed/vision_embedding_offline.py)"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:99
msgid ""
"[Online Vision Embedding Example](https://github.com/vllm-"
"project/vllm/blob/main/examples/pooling/embed/vision_embedding_online.py)"
msgstr ""
"[在线视觉嵌入示例](https://github.com/vllm-project/vllm/blob/main/examples/pooling/embed/vision_embedding_online.py)"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:101
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:103
msgid ""
"Run performance of `Qwen3-VL-Embedding-8B` as an example. Refer to [vllm "
"benchmark](https://docs.vllm.ai/en/latest/benchmarking/cli/) for more "
"details."
msgstr "以 `Qwen3-VL-Embedding-8B` 的运行性能为例。更多详情请参考 [vllm 基准测试](https://docs.vllm.ai/en/latest/benchmarking/cli/)。"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:106
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。按如下方式运行代码。"
#: ../../source/tutorials/models/Qwen3-VL-Embedding.md:112
msgid ""
"After about several minutes, you can get the performance evaluation "
"result. With this tutorial, the performance result is:"
msgstr "大约几分钟后,您将获得性能评估结果。在本教程中,性能结果如下:"

View File

@@ -0,0 +1,190 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:1
msgid "Qwen3-VL-Reranker"
msgstr "Qwen3-VL-Reranker"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:5
msgid ""
"The Qwen3-VL-Embedding and Qwen3-VL-Reranker model series are the latest "
"additions to the Qwen family, built upon the recently open-sourced and "
"powerful Qwen3-VL foundation model. Specifically designed for multimodal "
"information retrieval and cross-modal understanding, this suite accepts "
"diverse inputs including text, images, screenshots, and videos, as well "
"as inputs containing a mixture of these modalities. This guide describes "
"how to run the model with vLLM Ascend."
msgstr ""
"Qwen3-VL-Embedding 和 Qwen3-VL-Reranker 模型系列是 Qwen 家族的最新成员,基于最近开源且功能强大的 Qwen3-VL 基础模型构建。该系列专为多模态信息检索和跨模态理解而设计,可接受包括文本、图像、截图和视频在内的多样化输入,以及包含这些模态混合的输入。本指南描述了如何使用 vLLM Ascend 运行该模型。"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:7
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:9
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取该模型的支持功能矩阵。"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:11
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:13
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:15
msgid ""
"`Qwen3-VL-Reranker-8B` [Download model "
"weight](https://www.modelscope.cn/models/Qwen/Qwen3-VL-Reranker-8B)"
msgstr "`Qwen3-VL-Reranker-8B` [下载模型权重](https://www.modelscope.cn/models/Qwen/Qwen3-VL-Reranker-8B)"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:16
msgid ""
"`Qwen3-VL-Reranker-2B` [Download model "
"weight](https://www.modelscope.cn/models/Qwen/Qwen3-VL-Reranker-2B)"
msgstr "`Qwen3-VL-Reranker-2B` [下载模型权重](https://www.modelscope.cn/models/Qwen/Qwen3-VL-Reranker-2B)"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:18
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`"
msgstr "建议将模型权重下载到多个节点的共享目录中,例如 `/root/.cache/`"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:20
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:22
msgid ""
"You can use our official docker image to run `Qwen3-VL-Reranker` series "
"models."
msgstr "您可以使用我们的官方 docker 镜像来运行 `Qwen3-VL-Reranker` 系列模型。"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:24
msgid ""
"Start the docker image on your node, refer to [using "
"docker](../../installation.md#set-up-using-docker)."
msgstr "在您的节点上启动 docker 镜像,请参考[使用 docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:26
msgid ""
"If you don't want to use the docker image as above, you can also build "
"all from source:"
msgstr "如果您不想使用上述 docker 镜像,也可以从源代码构建所有内容:"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:28
msgid ""
"Install `vllm-ascend` from source, refer to "
"[installation](../../installation.md)."
msgstr "从源代码安装 `vllm-ascend`,请参考[安装指南](../../installation.md)。"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:30
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:32
msgid "Using the Qwen3-VL-Reranker-8B model as an example:"
msgstr "以 Qwen3-VL-Reranker-8B 模型为例:"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:34
msgid "Chat Template"
msgstr "聊天模板"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:36
msgid ""
"The Qwen3-VL-Reranker model requires a specific chat template for proper "
"formatting. Create a file named `qwen3_vl_reranker.jinja` with the "
"following content:"
msgstr "Qwen3-VL-Reranker 模型需要一个特定的聊天模板以进行正确格式化。创建一个名为 `qwen3_vl_reranker.jinja` 的文件,内容如下:"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:64
msgid ""
"Save this file to a location of your choice (e.g., "
"`./qwen3_vl_reranker.jinja`)."
msgstr "将此文件保存到您选择的位置(例如,`./qwen3_vl_reranker.jinja`)。"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:66
msgid "Online Inference"
msgstr "在线推理"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:68
msgid "Start the server with the following command:"
msgstr "使用以下命令启动服务器:"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:78
msgid "Once your server is started, you can send request with follow examples."
msgstr "一旦您的服务器启动,您就可以按照以下示例发送请求。"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:118
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:203
msgid ""
"If you run this script successfully, you will see a list of scores "
"printed to the console, similar to this:"
msgstr "如果您成功运行此脚本,您将在控制台看到打印出的分数列表,类似于这样:"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:124
msgid "Offline Inference"
msgstr "离线推理"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:212
msgid "For more examples, refer to the vLLM official examples:"
msgstr "更多示例,请参考 vLLM 官方示例:"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:214
msgid ""
"[Offline Vision Embedding Example](https://github.com/vllm-"
"project/vllm/blob/main/examples/pooling/score/vision_reranker_offline.py)"
msgstr "[离线视觉重排示例](https://github.com/vllm-project/vllm/blob/main/examples/pooling/score/vision_reranker_offline.py)"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:215
msgid ""
"[Online Vision Embedding Example](https://github.com/vllm-"
"project/vllm/blob/main/examples/pooling/score/vision_rerank_api_online.py)"
msgstr "[在线视觉重排示例](https://github.com/vllm-project/vllm/blob/main/examples/pooling/score/vision_rerank_api_online.py)"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:217
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:219
msgid ""
"Run performance of `Qwen3-VL-Reranker-8B` as an example. Refer to [vllm "
"benchmark](https://docs.vllm.ai/en/latest/benchmarking/cli/) for more "
"details."
msgstr "以 `Qwen3-VL-Reranker-8B` 的运行性能为例。更多详情请参考 [vllm 基准测试](https://docs.vllm.ai/en/latest/benchmarking/cli/)。"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:222
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。按如下方式运行代码。"
#: ../../source/tutorials/models/Qwen3-VL-Reranker.md:228
msgid ""
"After about several minutes, you can get the performance evaluation "
"result. With this tutorial, the performance result is:"
msgstr "大约几分钟后,您将获得性能评估结果。在本教程中,性能结果如下:"

View File

@@ -0,0 +1,402 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3.5-27B.md:1
msgid "Qwen3.5-27B"
msgstr "Qwen3.5-27B"
#: ../../source/tutorials/models/Qwen3.5-27B.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3.5-27B.md:5
msgid ""
"Qwen3.5 represents a significant leap forward, integrating breakthroughs "
"in multimodal learning, architectural efficiency, reinforcement learning "
"scale, and global accessibility to empower developers and enterprises "
"with unprecedented capability and efficiency."
msgstr "Qwen3.5 代表了一次重大飞跃,它整合了多模态学习、架构效率、强化学习规模和全球可访问性方面的突破,为开发者和企业提供了前所未有的能力和效率。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:7
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, single-node and multi-node deployment, accuracy and "
"performance evaluation."
msgstr "本文档将展示该模型的主要验证步骤,包括支持的特性、特性配置、环境准备、单节点和多节点部署、精度和性能评估。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:9
msgid "The `Qwen3.5-27B` model is first supported in `vllm-ascend:v0.17.0rc1`."
msgstr "`Qwen3.5-27B` 模型首次在 `vllm-ascend:v0.17.0rc1` 版本中得到支持。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:11
msgid "Supported Features"
msgstr "支持的特性"
#: ../../source/tutorials/models/Qwen3.5-27B.md:13
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的特性](../../user_guide/support_matrix/supported_models.md)以获取模型支持的特性矩阵。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:15
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[特性指南](../../user_guide/feature_guide/index.md)以获取特性的配置信息。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:17
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen3.5-27B.md:19
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen3.5-27B.md:21
msgid ""
"`Qwen3.5-27B`(BF16 version): requires 1 Atlas 800 A3 (64G × 16) node or 1"
" Atlas 800 A2 (64G × 8) node. [Download model "
"weight](https://modelscope.cn/models/Qwen/Qwen3.5-27B)"
msgstr "`Qwen3.5-27B` (BF16 版本):需要 1 个 Atlas 800 A3 (64G × 16) 节点或 1 个 Atlas 800 A2 (64G × 8) 节点。[下载模型权重](https://modelscope.cn/models/Qwen/Qwen3.5-27B)"
#: ../../source/tutorials/models/Qwen3.5-27B.md:22
msgid ""
"`Qwen3.5-27B-w8a8`(Quantized version): requires 1 Atlas 800 A3 (64G × 16)"
" node or 1 Atlas 800 A2 (64G × 8) node. [Download model "
"weight](https://www.modelscope.cn/models/Eco-Tech/Qwen3.5-27B-w8a8-mtp)"
msgstr "`Qwen3.5-27B-w8a8` (量化版本):需要 1 个 Atlas 800 A3 (64G × 16) 节点或 1 个 Atlas 800 A2 (64G × 8) 节点。[下载模型权重](https://www.modelscope.cn/models/Eco-Tech/Qwen3.5-27B-w8a8-mtp)"
#: ../../source/tutorials/models/Qwen3.5-27B.md:24
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`."
msgstr "建议将模型权重下载到多个节点的共享目录中,例如 `/root/.cache/`。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:26
msgid "Verify Multi-node Communication(Optional)"
msgstr "验证多节点通信(可选)"
#: ../../source/tutorials/models/Qwen3.5-27B.md:28
msgid ""
"If you want to deploy multi-node environment, you need to verify multi-"
"node communication according to [verify multi-node communication "
"environment](../../installation.md#verify-multi-node-communication)."
msgstr "如果您想部署多节点环境,需要根据[验证多节点通信环境](../../installation.md#verify-multi-node-communication)来验证多节点通信。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:30
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen3.5-27B.md
msgid "Use docker image"
msgstr "使用 Docker 镜像"
#: ../../source/tutorials/models/Qwen3.5-27B.md:36
msgid ""
"For example, using images `quay.io/ascend/vllm-ascend:v0.17.0rc1`(for "
"Atlas 800 A2) and `quay.io/ascend/vllm-ascend:v0.17.0rc1-a3`(for Atlas "
"800 A3)."
msgstr "例如,使用镜像 `quay.io/ascend/vllm-ascend:v0.17.0rc1`(适用于 Atlas 800 A2和 `quay.io/ascend/vllm-ascend:v0.17.0rc1-a3`(适用于 Atlas 800 A3。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:38
msgid ""
"Select an image based on your machine type and start the docker image on "
"your node, refer to [using docker](../../installation.md#set-up-using-"
"docker)."
msgstr "根据您的机器类型选择镜像,并在您的节点上启动 Docker 镜像,请参考[使用 Docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/Qwen3.5-27B.md
msgid "Build from source"
msgstr "从源码构建"
#: ../../source/tutorials/models/Qwen3.5-27B.md:78
msgid "You can build all from source."
msgstr "您可以从源码构建所有组件。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:80
msgid ""
"Install `vllm-ascend`, refer to [set up using "
"python](../../installation.md#set-up-using-python)."
msgstr "安装 `vllm-ascend`,请参考[使用 Python 设置](../../installation.md#set-up-using-python)。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:84
msgid ""
"If you want to deploy multi-node environment, you need to set up "
"environment on each node."
msgstr "如果您想部署多节点环境,需要在每个节点上设置环境。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:86
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3.5-27B.md:88
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/models/Qwen3.5-27B.md:90
msgid ""
"`Qwen3.5-27B` and `Qwen3.5-27B-w8a8` can both be deployed on 1 Atlas 800 "
"A3(64G × 16), 1 Atlas 800 A2(64G × 8). Quantized version needs to start "
"with parameter --quantization ascend."
msgstr "`Qwen3.5-27B` 和 `Qwen3.5-27B-w8a8` 都可以部署在 1 个 Atlas 800 A3 (64G × 16) 或 1 个 Atlas 800 A2 (64G × 8) 节点上。量化版本需要使用参数 `--quantization ascend` 启动。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:92
msgid "Run the following script to execute online 128k inference."
msgstr "运行以下脚本来执行在线 128k 推理。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:125
msgid "**Notice:**"
msgstr "**注意:**"
#: ../../source/tutorials/models/Qwen3.5-27B.md:127
msgid "The parameters are explained as follows:"
msgstr "参数解释如下:"
#: ../../source/tutorials/models/Qwen3.5-27B.md:129
msgid ""
"`--data-parallel-size` 1 and `--tensor-parallel-size` 2 are common "
"settings for data parallelism (DP) and tensor parallelism (TP) sizes."
msgstr "`--data-parallel-size` 1 和 `--tensor-parallel-size` 2 是数据并行 (DP) 和张量并行 (TP) 大小的常见设置。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:130
msgid ""
"`--max-model-len` represents the context length, which is the maximum "
"value of the input plus output for a single request."
msgstr "`--max-model-len` 表示上下文长度,即单个请求的输入加输出的最大值。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:131
msgid ""
"`--max-num-seqs` indicates the maximum number of requests that each DP "
"group is allowed to process. If the number of requests sent to the "
"service exceeds this limit, the excess requests will remain in a waiting "
"state and will not be scheduled. Note that the time spent in the waiting "
"state is also counted in metrics such as TTFT and TPOT. Therefore, when "
"testing performance, it is generally recommended that `--max-num-seqs` * "
"`--data-parallel-size` >= the actual total concurrency."
msgstr "`--max-num-seqs` 表示每个 DP 组允许处理的最大请求数。如果发送到服务的请求数超过此限制,超出的请求将保持在等待状态,不会被调度。请注意,在等待状态所花费的时间也会计入 TTFT 和 TPOT 等指标。因此,在测试性能时,通常建议 `--max-num-seqs` * `--data-parallel-size` >= 实际总并发数。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:132
msgid ""
"`--max-num-batched-tokens` represents the maximum number of tokens that "
"the model can process in a single step. Currently, vLLM v1 scheduling "
"enables ChunkPrefill/SplitFuse by default, which means:"
msgstr "`--max-num-batched-tokens` 表示模型在单步中可以处理的最大 token 数。目前vLLM v1 调度默认启用 ChunkPrefill/SplitFuse这意味着"
#: ../../source/tutorials/models/Qwen3.5-27B.md:133
msgid ""
"(1) If the input length of a request is greater than `--max-num-batched-"
"tokens`, it will be divided into multiple rounds of computation according"
" to `--max-num-batched-tokens`;"
msgstr "(1) 如果一个请求的输入长度大于 `--max-num-batched-tokens`,它将根据 `--max-num-batched-tokens` 被分成多轮计算;"
#: ../../source/tutorials/models/Qwen3.5-27B.md:134
msgid ""
"(2) Decode requests are prioritized for scheduling, and prefill requests "
"are scheduled only if there is available capacity."
msgstr "(2) 解码请求优先调度,只有在有可用容量时才会调度预填充请求。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:135
msgid ""
"Generally, if `--max-num-batched-tokens` is set to a larger value, the "
"overall latency will be lower, but the pressure on GPU memory (activation"
" value usage) will be greater."
msgstr "通常,如果将 `--max-num-batched-tokens` 设置为较大的值,整体延迟会更低,但 GPU 内存(激活值使用)的压力会更大。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:136
msgid ""
"`--gpu-memory-utilization` represents the proportion of HBM that vLLM "
"will use for actual inference. Its essential function is to calculate the"
" available kv_cache size. During the warm-up phase (referred to as "
"profile run in vLLM), vLLM records the peak GPU memory usage during an "
"inference process with an input size of `--max-num-batched-tokens`. The "
"available kv_cache size is then calculated as: `--gpu-memory-utilization`"
" * HBM size - peak GPU memory usage. Therefore, the larger the value of "
"`--gpu-memory-utilization`, the more kv_cache can be used. However, since"
" the GPU memory usage during the warm-up phase may differ from that "
"during actual inference (e.g., due to uneven EP load), setting `--gpu-"
"memory-utilization` too high may lead to OOM (Out of Memory) issues "
"during actual inference. The default value is `0.9`."
msgstr "`--gpu-memory-utilization` 表示 vLLM 将用于实际推理的 HBM 比例。其核心功能是计算可用的 kv_cache 大小。在预热阶段(在 vLLM 中称为 profile runvLLM 会记录输入大小为 `--max-num-batched-tokens` 的推理过程中的峰值 GPU 内存使用量。然后,可用的 kv_cache 大小计算为:`--gpu-memory-utilization` * HBM 大小 - 峰值 GPU 内存使用量。因此,`--gpu-memory-utilization` 的值越大,可以使用的 kv_cache 就越多。然而,由于预热阶段的 GPU 内存使用量可能与实际推理期间不同(例如,由于 EP 负载不均衡),将 `--gpu-memory-utilization` 设置得过高可能会导致实际推理期间出现 OOM内存不足问题。默认值为 `0.9`。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:137
msgid ""
"`--no-enable-prefix-caching` indicates that prefix caching is disabled. "
"To enable it, for mamba-like models Qwen3.5, set `--enable-prefix-"
"caching` and `--mamba-cache-mode align`. Notice the current "
"implementation of hybrid kv cache might result in a very large block_size"
" when scheduling. For example, the block_size may be adjusted to 2048, "
"which means that any prefix shorter than 2048 will never be cached."
msgstr "`--no-enable-prefix-caching` 表示前缀缓存被禁用。要启用它,对于类似 Mamba 的模型 Qwen3.5,请设置 `--enable-prefix-caching` 和 `--mamba-cache-mode align`。请注意,当前混合 kv cache 的实现可能在调度时导致非常大的 block_size。例如block_size 可能被调整为 2048这意味着任何短于 2048 的前缀将永远不会被缓存。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:138
msgid ""
"`--quantization` \"ascend\" indicates that quantization is used. To "
"disable quantization, remove this option."
msgstr "`--quantization` \"ascend\" 表示使用量化。要禁用量化,请移除此选项。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:139
msgid ""
"`--compilation-config` contains configurations related to the aclgraph "
"graph mode. The most significant configurations are \"cudagraph_mode\" "
"and \"cudagraph_capture_sizes\", which have the following meanings: "
"\"cudagraph_mode\": represents the specific graph mode. Currently, "
"\"PIECEWISE\" and \"FULL_DECODE_ONLY\" are supported. The graph mode is "
"mainly used to reduce the cost of operator dispatch. Currently, "
"\"FULL_DECODE_ONLY\" is recommended."
msgstr "`--compilation-config` 包含与 aclgraph 图模式相关的配置。最重要的配置是 \"cudagraph_mode\" 和 \"cudagraph_capture_sizes\",其含义如下:\"cudagraph_mode\":表示特定的图模式。目前支持 \"PIECEWISE\" 和 \"FULL_DECODE_ONLY\"。图模式主要用于降低算子调度的开销。目前推荐使用 \"FULL_DECODE_ONLY\"。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:141
msgid ""
"\"cudagraph_capture_sizes\": represents different levels of graph modes. "
"The default value is [1, 2, 4, 8, 16, 24, 32, 40,..., `--max-num-seqs`]. "
"In the graph mode, the input for graphs at different levels is fixed, and"
" inputs between levels are automatically padded to the next level. "
"Currently, the default setting is recommended. Only in some scenarios is "
"it necessary to set this separately to achieve optimal performance."
msgstr "\"cudagraph_capture_sizes\":表示不同级别的图模式。默认值为 [1, 2, 4, 8, 16, 24, 32, 40,..., `--max-num-seqs`]。在图模式下,不同级别图的输入是固定的,级别之间的输入会自动填充到下一级别。目前推荐使用默认设置。只有在某些场景下,才需要单独设置此参数以达到最佳性能。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:143
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/Qwen3.5-27B.md:145
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "一旦您的服务器启动,您就可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/Qwen3.5-27B.md:158
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/Qwen3.5-27B.md:160
msgid "Here are two accuracy evaluation methods."
msgstr "以下是两种精度评估方法。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:162
#: ../../source/tutorials/models/Qwen3.5-27B.md:174
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/models/Qwen3.5-27B.md:164
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参考[使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:166
msgid ""
"After execution, you can get the result, here is the result of `Qwen3.5"
"-27B-w8a8` in `vllm-ascend:v0.17.0rc1` for reference only."
msgstr "执行后,您可以获得结果,以下是 `Qwen3.5-27B-w8a8` 在 `vllm-ascend:v0.17.0rc1` 中的结果,仅供参考。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:76
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/models/Qwen3.5-27B.md:76
msgid "version"
msgstr "版本"
#: ../../source/tutorials/models/Qwen3.5-27B.md:76
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/models/Qwen3.5-27B.md:76
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/models/Qwen3.5-27B.md:76
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/models/Qwen3.5-27B.md:76
msgid "gsm8k"
msgstr "gsm8k"
#: ../../source/tutorials/models/Qwen3.5-27B.md:76
msgid "-"
msgstr "-"
#: ../../source/tutorials/models/Qwen3.5-27B.md:76
msgid "accuracy"
msgstr "准确率"
#: ../../source/tutorials/models/Qwen3.5-27B.md:76
msgid "gen"
msgstr "生成"
#: ../../source/tutorials/models/Qwen3.5-27B.md:76
msgid "96.74"
msgstr "96.74"
#: ../../source/tutorials/models/Qwen3.5-27B.md:172
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen3.5-27B.md:176
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参阅[使用AISBench进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:178
msgid "Using vLLM Benchmark"
msgstr "使用vLLM基准测试"
#: ../../source/tutorials/models/Qwen3.5-27B.md:180
msgid "Run performance evaluation of `Qwen3.5-27B-w8a8` as an example."
msgstr "以运行 `Qwen3.5-27B-w8a8` 的性能评估为例。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:182
msgid ""
"Refer to [vllm "
"benchmark](https://docs.vllm.ai/en/latest/contributing/benchmarks.html) "
"for more details."
msgstr "更多详情请参阅[vllm基准测试](https://docs.vllm.ai/en/latest/contributing/benchmarks.html)。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:184
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 包含三个子命令:"
#: ../../source/tutorials/models/Qwen3.5-27B.md:186
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批次请求的延迟进行基准测试。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:187
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:188
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:190
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例,运行以下代码。"
#: ../../source/tutorials/models/Qwen3.5-27B.md:197
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您将获得性能评估结果。"

View File

@@ -0,0 +1,542 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:1
msgid "Qwen3.5-397B-A17B"
msgstr "Qwen3.5-397B-A17B"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:5
msgid ""
"Qwen3.5 represents a significant leap forward, integrating breakthroughs "
"in multimodal learning, architectural efficiency, reinforcement learning "
"scale, and global accessibility to empower developers and enterprises "
"with unprecedented capability and efficiency."
msgstr "Qwen3.5 代表了一次重大飞跃,它整合了多模态学习、架构效率、强化学习规模和全球可访问性方面的突破,为开发者和企业提供了前所未有的能力和效率。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:7
msgid ""
"This document will show the main verification steps of the model, "
"including supported features, feature configuration, environment "
"preparation, single-node and multi-node deployment, accuracy and "
"performance evaluation."
msgstr "本文档将展示该模型的主要验证步骤,包括支持的功能、功能配置、环境准备、单节点和多节点部署、精度和性能评估。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:9
msgid ""
"The `Qwen3.5-397B-A17B` model is first supported in `vllm-"
"ascend:v0.17.0rc1`."
msgstr "`Qwen3.5-397B-A17B` 模型首次在 `vllm-ascend:v0.17.0rc1` 版本中得到支持。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:11
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:13
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取模型支持的功能矩阵。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:15
msgid ""
"Refer to [feature guide](../../user_guide/feature_guide/index.md) to get "
"the feature's configuration."
msgstr "请参考[功能指南](../../user_guide/feature_guide/index.md)以获取功能的配置信息。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:17
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:19
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:21
msgid ""
"`Qwen3.5-397B-A17B`(BF16 version): require 2 Atlas 800 A3 (64G × 16) "
"nodes or 4 Atlas 800 A2 (64G × 8) nodes. [Download model "
"weight](https://www.modelscope.cn/models/Qwen/Qwen3.5-397B-A17B)"
msgstr "`Qwen3.5-397B-A17B` (BF16 版本):需要 2 个 Atlas 800 A3 (64G × 16) 节点或 4 个 Atlas 800 A2 (64G × 8) 节点。[下载模型权重](https://www.modelscope.cn/models/Qwen/Qwen3.5-397B-A17B)"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:22
msgid ""
"`Qwen3.5-397B-A17B-w8a8`(Quantized version): require 1 Atlas 800 A3 (64G "
"× 16) node or 2 Atlas 800 A2 (64G × 8) nodes. [Download model "
"weight](https://www.modelscope.cn/models/Eco-Tech/Qwen3.5-397B-A17B-"
"w8a8-mtp)"
msgstr "`Qwen3.5-397B-A17B-w8a8` (量化版本):需要 1 个 Atlas 800 A3 (64G × 16) 节点或 2 个 Atlas 800 A2 (64G × 8) 节点。[下载模型权重](https://www.modelscope.cn/models/Eco-Tech/Qwen3.5-397B-A17B-w8a8-mtp)"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:24
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`."
msgstr "建议将模型权重下载到多个节点的共享目录中,例如 `/root/.cache/`。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:26
msgid "Verify Multi-node Communication(Optional)"
msgstr "验证多节点通信(可选)"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:28
msgid ""
"If you want to deploy multi-node environment, you need to verify multi-"
"node communication according to [verify multi-node communication "
"environment](../../installation.md#verify-multi-node-communication)."
msgstr "如果您想部署多节点环境,需要根据[验证多节点通信环境](../../installation.md#verify-multi-node-communication)来验证多节点通信。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:30
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md
msgid "Use docker image"
msgstr "使用 Docker 镜像"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:36
msgid ""
"For example, using images `quay.io/ascend/vllm-ascend:v0.17.0rc1`(for "
"Atlas 800 A2) and `quay.io/ascend/vllm-ascend:v0.17.0rc1-a3`(for Atlas "
"800 A3)."
msgstr "例如,使用镜像 `quay.io/ascend/vllm-ascend:v0.17.0rc1`(适用于 Atlas 800 A2和 `quay.io/ascend/vllm-ascend:v0.17.0rc1-a3`(适用于 Atlas 800 A3。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:38
msgid ""
"Select an image based on your machine type and start the docker image on "
"your node, refer to [using docker](../../installation.md#set-up-using-"
"docker)."
msgstr "根据您的机器类型选择镜像并在节点上启动 Docker 镜像,请参考[使用 Docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md
msgid "Build from source"
msgstr "从源码构建"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:78
msgid "You can build all from source."
msgstr "您可以从源码构建所有组件。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:80
msgid ""
"Install `vllm-ascend`, refer to [set up using "
"python](../../installation.md#set-up-using-python)."
msgstr "安装 `vllm-ascend`,请参考[使用 Python 设置](../../installation.md#set-up-using-python)。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:84
msgid ""
"If you want to deploy multi-node environment, you need to set up "
"environment on each node."
msgstr "如果您想部署多节点环境,需要在每个节点上设置环境。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:86
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:88
msgid "Single-node Deployment"
msgstr "单节点部署"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:90
msgid ""
"`Qwen3.5-397B-A17B` can be deployed on 2 Atlas 800 A3(64G*16) or 4 Atlas "
"800 A2(64G*8). `Qwen3.5-397B-A17B-w8a8` can be deployed on 1 Atlas 800 "
"A3(64G*16) or 2 Atlas 800 A2(64G*8), need to start with parameter "
"`--quantization ascend`."
msgstr "`Qwen3.5-397B-A17B` 可以部署在 2 个 Atlas 800 A3(64G*16) 或 4 个 Atlas 800 A2(64G*8) 上。`Qwen3.5-397B-A17B-w8a8` 可以部署在 1 个 Atlas 800 A3(64G*16) 或 2 个 Atlas 800 A2(64G*8) 上,需要使用参数 `--quantization ascend` 启动。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:93
msgid ""
"Run the following script to execute online 128k inference On 1 Atlas 800 "
"A3(64G*16)."
msgstr "在 1 个 Atlas 800 A3(64G*16) 上运行以下脚本以执行在线 128k 推理。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:134
msgid "**Notice:**"
msgstr "**注意:**"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:136
msgid "The parameters are explained as follows:"
msgstr "参数解释如下:"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:138
msgid ""
"`--data-parallel-size` 1 and `--tensor-parallel-size` 16 are common "
"settings for data parallelism (DP) and tensor parallelism (TP) sizes."
msgstr "`--data-parallel-size` 1 和 `--tensor-parallel-size` 16 是数据并行 (DP) 和张量并行 (TP) 大小的常见设置。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:139
msgid ""
"`--max-model-len` represents the context length, which is the maximum "
"value of the input plus output for a single request."
msgstr "`--max-model-len` 表示上下文长度,即单个请求的输入加输出的最大值。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:140
msgid ""
"`--max-num-seqs` indicates the maximum number of requests that each DP "
"group is allowed to process. If the number of requests sent to the "
"service exceeds this limit, the excess requests will remain in a waiting "
"state and will not be scheduled. Note that the time spent in the waiting "
"state is also counted in metrics such as TTFT and TPOT. Therefore, when "
"testing performance, it is generally recommended that `--max-num-seqs` * "
"`--data-parallel-size` >= the actual total concurrency."
msgstr "`--max-num-seqs` 表示每个 DP 组允许处理的最大请求数。如果发送到服务的请求数超过此限制,多余的请求将保持在等待状态,不会被调度。请注意,在等待状态所花费的时间也会计入 TTFT 和 TPOT 等指标。因此,在测试性能时,通常建议 `--max-num-seqs` * `--data-parallel-size` >= 实际总并发数。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:141
msgid ""
"`--max-num-batched-tokens` represents the maximum number of tokens that "
"the model can process in a single step. Currently, vLLM v1 scheduling "
"enables ChunkPrefill/SplitFuse by default, which means:"
msgstr "`--max-num-batched-tokens` 表示模型单步可以处理的最大 token 数。目前vLLM v1 调度默认启用 ChunkPrefill/SplitFuse这意味着"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:142
msgid ""
"(1) If the input length of a request is greater than `--max-num-batched-"
"tokens`, it will be divided into multiple rounds of computation according"
" to `--max-num-batched-tokens`;"
msgstr "(1) 如果请求的输入长度大于 `--max-num-batched-tokens`,它将根据 `--max-num-batched-tokens` 被分成多轮计算;"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:143
msgid ""
"(2) Decode requests are prioritized for scheduling, and prefill requests "
"are scheduled only if there is available capacity."
msgstr "(2) 解码请求优先调度,只有在有可用容量时才调度预填充请求。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:144
msgid ""
"Generally, if `--max-num-batched-tokens` is set to a larger value, the "
"overall latency will be lower, but the pressure on GPU memory (activation"
" value usage) will be greater."
msgstr "通常,如果 `--max-num-batched-tokens` 设置得较大,整体延迟会更低,但 GPU 内存(激活值使用)的压力会更大。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:145
msgid ""
"`--gpu-memory-utilization` represents the proportion of HBM that vLLM "
"will use for actual inference. Its essential function is to calculate the"
" available kv_cache size. During the warm-up phase (referred to as "
"profile run in vLLM), vLLM records the peak GPU memory usage during an "
"inference process with an input size of `--max-num-batched-tokens`. The "
"available kv_cache size is then calculated as: `--gpu-memory-utilization`"
" * HBM size - peak GPU memory usage. Therefore, the larger the value of "
"`--gpu-memory-utilization`, the more kv_cache can be used. However, since"
" the GPU memory usage during the warm-up phase may differ from that "
"during actual inference (e.g., due to uneven EP load), setting `--gpu-"
"memory-utilization` too high may lead to OOM (Out of Memory) issues "
"during actual inference. The default value is `0.9`."
msgstr "`--gpu-memory-utilization` 表示 vLLM 将用于实际推理的 HBM 比例。其核心功能是计算可用的 kv_cache 大小。在预热阶段vLLM 中称为 profile runvLLM 会记录输入大小为 `--max-num-batched-tokens` 的推理过程中的峰值 GPU 内存使用量。然后,可用的 kv_cache 大小计算为:`--gpu-memory-utilization` * HBM 大小 - 峰值 GPU 内存使用量。因此,`--gpu-memory-utilization` 的值越大,可用的 kv_cache 就越多。然而,由于预热阶段的 GPU 内存使用量可能与实际推理时不同(例如,由于 EP 负载不均),将 `--gpu-memory-utilization` 设置得过高可能导致实际推理时出现 OOM内存不足问题。默认值为 `0.9`。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:146
msgid ""
"`--enable-expert-parallel` indicates that EP is enabled. Note that vLLM "
"does not support a mixed approach of ETP and EP; that is, MoE can either "
"use pure EP or pure TP."
msgstr "`--enable-expert-parallel` 表示启用了 EP。请注意vLLM 不支持 ETP 和 EP 的混合方法也就是说MoE 要么使用纯 EP要么使用纯 TP。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:147
msgid ""
"`--no-enable-prefix-caching` indicates that prefix caching is disabled. "
"To enable it, for mamba-like models Qwen3.5, set `--enable-prefix-"
"caching` and `--mamba-cache-mode align`. Notice the current "
"implementation of hybrid kv cache might result in a very large block_size"
" when scheduling. For example, the block_size may be adjusted to 2048, "
"which means that any prefix shorter than 2048 will never be cached."
msgstr "`--no-enable-prefix-caching` 表示前缀缓存被禁用。要启用它,对于类似 Mamba 的模型 Qwen3.5,请设置 `--enable-prefix-caching` 和 `--mamba-cache-mode align`。请注意,当前混合 kv cache 的实现可能在调度时导致非常大的 block_size。例如block_size 可能被调整为 2048这意味着任何短于 2048 的前缀将永远不会被缓存。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:148
msgid ""
"`--quantization` \"ascend\" indicates that quantization is used. To "
"disable quantization, remove this option."
msgstr "`--quantization` \"ascend\" 表示使用了量化。要禁用量化,请移除此选项。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:149
msgid ""
"`--compilation-config` contains configurations related to the aclgraph "
"graph mode. The most significant configurations are \"cudagraph_mode\" "
"and \"cudagraph_capture_sizes\", which have the following meanings: "
"\"cudagraph_mode\": represents the specific graph mode. Currently, "
"\"PIECEWISE\" and \"FULL_DECODE_ONLY\" are supported. The graph mode is "
"mainly used to reduce the cost of operator dispatch. Currently, "
"\"FULL_DECODE_ONLY\" is recommended."
msgstr "`--compilation-config` 包含与 aclgraph 图模式相关的配置。最重要的配置是 \"cudagraph_mode\" 和 \"cudagraph_capture_sizes\",其含义如下:\"cudagraph_mode\":表示特定的图模式。目前支持 \"PIECEWISE\" 和 \"FULL_DECODE_ONLY\"。图模式主要用于降低算子调度的开销。目前推荐使用 \"FULL_DECODE_ONLY\"。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:151
msgid ""
"\"cudagraph_capture_sizes\": represents different levels of graph modes. "
"The default value is [1, 2, 4, 8, 16, 24, 32, 40,..., `--max-num-seqs`]. "
"In the graph mode, the input for graphs at different levels is fixed, and"
" inputs between levels are automatically padded to the next level. "
"Currently, the default setting is recommended. Only in some scenarios is "
"it necessary to set this separately to achieve optimal performance."
msgstr "\"cudagraph_capture_sizes\":表示不同级别的图模式。默认值为 [1, 2, 4, 8, 16, 24, 32, 40,..., `--max-num-seqs`]。在图模式下,不同级别图的输入是固定的,级别之间的输入会自动填充到下一个级别。目前推荐使用默认设置。只有在某些场景下,才需要单独设置此参数以达到最佳性能。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:153
msgid "Multi-node Deployment with MP (Recommended)"
msgstr "使用 MP 的多节点部署(推荐)"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:155
msgid ""
"Assume you have 2 Atlas 800 A2 nodes, and want to deploy the `Qwen3.5"
"-397B-A17B` model across multiple nodes."
msgstr "假设您有 2 个 Atlas 800 A2 节点,并希望跨多个节点部署 `Qwen3.5-397B-A17B` 模型。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:157
msgid "Node 0"
msgstr "节点 0"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:203
msgid "Node1"
msgstr "节点 1"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:253
msgid ""
"If the service starts successfully, the following information will be "
"displayed on node 0:"
msgstr "如果服务启动成功,节点 0 上将显示以下信息:"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:264
msgid "Multi-node Deployment with Ray"
msgstr "使用 Ray 的多节点部署"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:266
msgid "refer to [Ray Distributed (Qwen/Qwen3-235B-A22B)](../features/ray.md)."
msgstr "请参考 [Ray 分布式 (Qwen/Qwen3-235B-A22B)](../features/ray.md)。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:268
msgid "Prefill-Decode Disaggregation"
msgstr "预填充-解码解耦"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:270
msgid ""
"We recommend using Mooncake for deployment: "
"[Mooncake](../features/pd_disaggregation_mooncake_multi_node.md)."
msgstr "我们推荐使用 Mooncake 进行部署:[Mooncake](../features/pd_disaggregation_mooncake_multi_node.md)。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:272
msgid ""
"Take Atlas 800 A3 (64G × 16) for example, we recommend to deploy 1P1D (3 "
"nodes) to run Qwen3.5-397B-A17B."
msgstr "以 Atlas 800 A3 (64G × 16) 为例,我们建议部署 1P1D3 个节点)来运行 Qwen3.5-397B-A17B。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:274
msgid "`Qwen3.5-397B-A17B-w8a8-mtp 1P1D` require 3 Atlas 800 A3 (64G × 16)."
msgstr "`Qwen3.5-397B-A17B-w8a8-mtp 1P1D` 需要 3 个 Atlas 800 A3 (64G × 16)。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:276
msgid ""
"To run the vllm-ascend `Prefill-Decode Disaggregation` service, you need "
"to deploy `run_p.sh` 、`run_d0.sh` and `run_d1.sh` script on each node and"
" deploy a `proxy.sh` script on prefill master node to forward requests."
msgstr "要运行 vllm-ascend `Prefill-Decode Disaggregation` 服务,您需要在每个节点上部署 `run_p.sh`、`run_d0.sh` 和 `run_d1.sh` 脚本,并在预填充主节点上部署一个 `proxy.sh` 脚本来转发请求。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:278
msgid "Prefill Node 0 `run_p.sh` script"
msgstr "预填充节点 0 `run_p.sh` 脚本"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:353
msgid "Decode Node 0 `run_d0.sh` script"
msgstr "解码节点 0 `run_d0.sh` 脚本"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:433
msgid "Decode Node 1 `run_d1.sh` script"
msgstr "解码节点 1 `run_d1.sh` 脚本"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:512
msgid "**Notice:** The parameters are explained as follows:"
msgstr "**注意:** 参数说明如下:"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:515
msgid ""
"`--async-scheduling`: enables the asynchronous scheduling function. When "
"Multi-Token Prediction (MTP) is enabled, asynchronous scheduling of "
"operator delivery can be implemented to overlap the operator delivery "
"latency."
msgstr ""
"`--async-scheduling`启用异步调度功能。当启用多令牌预测MTP可以实现算子交付的异步调度以重叠算子交付延迟。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:516
msgid ""
"`cudagraph_capture_sizes`: The recommended value is `n x (mtp + 1)`. And "
"the min is `n = 1` and the max is `n = max-num-seqs`. For other values, "
"it is recommended to set them to the number of frequently occurring "
"requests on the Decode (D) node."
msgstr ""
"`cudagraph_capture_sizes`:推荐值为 `n x (mtp + 1)`。最小值为 `n = 1`,最大值为 `n = max-num-seqs`。对于其他值建议设置为解码D节点上频繁出现的请求数量。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:517
msgid ""
"`recompute_scheduler_enable: true`: enables the recomputation scheduler. "
"When the Key-Value Cache (KV Cache) of the decode node is insufficient, "
"requests will be sent to the prefill node to recompute the KV Cache. In "
"the PD separation scenario, it is recommended to enable this "
"configuration on both prefill and decode nodes simultaneously."
msgstr ""
"`recompute_scheduler_enable: true`启用重计算调度器。当解码节点的键值缓存KV Cache不足时请求将被发送到预填充节点以重新计算 KV Cache。在 PD 分离场景下,建议同时在预填充节点和解码节点上启用此配置。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:518
msgid ""
"`no-enable-prefix-caching`: The prefix-cache feature is enabled by "
"default. You can use the `--no-enable-prefix-caching` parameter to "
"disable this feature. Notice: for Prefill-Decode disaggregation feature, "
"known issue on D node: [#7944](https://github.com/vllm-project/vllm-"
"ascend/issues/7944)"
msgstr ""
"`no-enable-prefix-caching`:前缀缓存功能默认启用。您可以使用 `--no-enable-prefix-caching` 参数禁用此功能。注意:对于预填充-解码分离功能D 节点上的已知问题:[#7944](https://github.com/vllm-project/vllm-ascend/issues/7944)"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:520
msgid "Run the `proxy.sh` script on the prefill master node"
msgstr "在预填充主节点上运行 `proxy.sh` 脚本"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:522
msgid ""
"Run a proxy server on the same node with the prefiller service instance. "
"You can get the proxy program in the repository's examples: "
"[load\\_balance\\_proxy\\_server\\_example.py](https://github.com/vllm-"
"project/vllm-"
"ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
msgstr ""
"在与预填充服务实例相同的节点上运行一个代理服务器。您可以在仓库的示例中找到代理程序:[load\\_balance\\_proxy\\_server\\_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:548
msgid "Functional Verification"
msgstr "功能验证"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:550
msgid "Once your server is started, you can query the model with input prompts:"
msgstr "服务器启动后,您可以使用输入提示词查询模型:"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:563
msgid "Accuracy Evaluation"
msgstr "精度评估"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:565
msgid "Here are two accuracy evaluation methods."
msgstr "以下是两种精度评估方法。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:567
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:579
msgid "Using AISBench"
msgstr "使用 AISBench"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:569
msgid ""
"Refer to [Using "
"AISBench](../../developer_guide/evaluation/using_ais_bench.md) for "
"details."
msgstr "详情请参阅[使用 AISBench](../../developer_guide/evaluation/using_ais_bench.md)。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:571
msgid ""
"After execution, you can get the result, here is the result of `Qwen3.5"
"-397B-A17B-w8a8` in `vllm-ascend:v0.17.0rc1` for reference only."
msgstr "执行后,您可以获得结果,以下是 `vllm-ascend:v0.17.0rc1` 中 `Qwen3.5-397B-A17B-w8a8` 的结果,仅供参考。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:76
msgid "dataset"
msgstr "数据集"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:76
msgid "version"
msgstr "版本"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:76
msgid "metric"
msgstr "指标"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:76
msgid "mode"
msgstr "模式"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:76
msgid "vllm-api-general-chat"
msgstr "vllm-api-general-chat"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:76
msgid "gsm8k"
msgstr "gsm8k"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:76
msgid "-"
msgstr "-"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:76
msgid "accuracy"
msgstr "准确率"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:76
msgid "gen"
msgstr "生成"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:76
msgid "96.74"
msgstr "96.74"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:577
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:581
msgid ""
"Refer to [Using AISBench for performance "
"evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-"
"performance-evaluation) for details."
msgstr "详情请参阅[使用 AISBench 进行性能评估](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation)。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:583
msgid "Using vLLM Benchmark"
msgstr "使用 vLLM Benchmark"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:585
msgid "Run performance evaluation of `Qwen3.5-397B-A17B-w8a8` as an example."
msgstr "以运行 `Qwen3.5-397B-A17B-w8a8` 的性能评估为例。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:587
msgid ""
"Refer to [vllm "
"benchmark](https://docs.vllm.ai/en/latest/contributing/benchmarks.html) "
"for more details."
msgstr "更多详情请参阅 [vllm benchmark](https://docs.vllm.ai/en/latest/contributing/benchmarks.html)。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:589
msgid "There are three `vllm bench` subcommands:"
msgstr "`vllm bench` 有三个子命令:"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:591
msgid "`latency`: Benchmark the latency of a single batch of requests."
msgstr "`latency`:对单批请求的延迟进行基准测试。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:592
msgid "`serve`: Benchmark the online serving throughput."
msgstr "`serve`:对在线服务吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:593
msgid "`throughput`: Benchmark offline inference throughput."
msgstr "`throughput`:对离线推理吞吐量进行基准测试。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:595
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。运行代码如下。"
#: ../../source/tutorials/models/Qwen3.5-397B-A17B.md:602
msgid ""
"After about several minutes, you can get the performance evaluation "
"result."
msgstr "大约几分钟后,您将获得性能评估结果。"

View File

@@ -0,0 +1,158 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3_embedding.md:1
msgid "Qwen3-Embedding"
msgstr "Qwen3-Embedding"
#: ../../source/tutorials/models/Qwen3_embedding.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3_embedding.md:5
msgid ""
"The Qwen3 Embedding model series is the latest proprietary model of the "
"Qwen family, specifically designed for text embedding and ranking tasks. "
"Building upon the dense foundational models of the Qwen3 series, it "
"provides a comprehensive range of text embeddings and reranking models in"
" various sizes (0.6B, 4B, and 8B). This guide describes how to run the "
"model with vLLM Ascend. Note that only 0.9.2rc1 and higher versions of "
"vLLM Ascend support the model."
msgstr ""
"Qwen3 Embedding 模型系列是 Qwen 家族最新的专有模型,专为文本嵌入和排序任务设计。它基于 Qwen3 系列的稠密基础模型提供了多种尺寸0.6B、4B 和 8B的全面文本嵌入和重排序模型。本指南描述了如何使用 vLLM Ascend 运行该模型。请注意,只有 vLLM Ascend 0.9.2rc1 及更高版本支持此模型。"
#: ../../source/tutorials/models/Qwen3_embedding.md:7
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/Qwen3_embedding.md:9
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取该模型的支持功能矩阵。"
#: ../../source/tutorials/models/Qwen3_embedding.md:11
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen3_embedding.md:13
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen3_embedding.md:15
msgid ""
"`Qwen3-Embedding-8B` [Download model "
"weight](https://www.modelscope.cn/models/Qwen/Qwen3-Embedding-8B)"
msgstr "`Qwen3-Embedding-8B` [下载模型权重](https://www.modelscope.cn/models/Qwen/Qwen3-Embedding-8B)"
#: ../../source/tutorials/models/Qwen3_embedding.md:16
msgid ""
"`Qwen3-Embedding-4B` [Download model "
"weight](https://www.modelscope.cn/models/Qwen/Qwen3-Embedding-4B)"
msgstr "`Qwen3-Embedding-4B` [下载模型权重](https://www.modelscope.cn/models/Qwen/Qwen3-Embedding-4B)"
#: ../../source/tutorials/models/Qwen3_embedding.md:17
msgid ""
"`Qwen3-Embedding-0.6B` [Download model "
"weight](https://www.modelscope.cn/models/Qwen/Qwen3-Embedding-0.6B)"
msgstr "`Qwen3-Embedding-0.6B` [下载模型权重](https://www.modelscope.cn/models/Qwen/Qwen3-Embedding-0.6B)"
#: ../../source/tutorials/models/Qwen3_embedding.md:19
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`"
msgstr "建议将模型权重下载到多个节点的共享目录中,例如 `/root/.cache/`"
#: ../../source/tutorials/models/Qwen3_embedding.md:21
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen3_embedding.md:23
msgid ""
"You can use our official docker image to run `Qwen3-Embedding` series "
"models."
msgstr "您可以使用我们的官方 docker 镜像来运行 `Qwen3-Embedding` 系列模型。"
#: ../../source/tutorials/models/Qwen3_embedding.md:25
msgid ""
"Start the docker image on your node, refer to [using "
"docker](../../installation.md#set-up-using-docker)."
msgstr "在您的节点上启动 docker 镜像,请参考[使用 docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/Qwen3_embedding.md:27
msgid ""
"if you don't want to use the docker image as above, you can also build "
"all from source:"
msgstr "如果您不想使用上述的 docker 镜像,也可以从源代码构建所有内容:"
#: ../../source/tutorials/models/Qwen3_embedding.md:29
msgid ""
"Install `vllm-ascend` from source, refer to "
"[installation](../../installation.md)."
msgstr "从源代码安装 `vllm-ascend`,请参考[安装指南](../../installation.md)。"
#: ../../source/tutorials/models/Qwen3_embedding.md:31
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3_embedding.md:33
msgid ""
"Using the Qwen3-Embedding-8B model as an example, first run the docker "
"container with the following command:"
msgstr "以 Qwen3-Embedding-8B 模型为例,首先使用以下命令运行 docker 容器:"
#: ../../source/tutorials/models/Qwen3_embedding.md:35
msgid "Online Inference"
msgstr "在线推理"
#: ../../source/tutorials/models/Qwen3_embedding.md:41
msgid "Once your server is started, you can query the model with input prompts."
msgstr "一旦您的服务器启动,您就可以使用输入提示词查询模型。"
#: ../../source/tutorials/models/Qwen3_embedding.md:52
msgid "Offline Inference"
msgstr "离线推理"
#: ../../source/tutorials/models/Qwen3_embedding.md:87
msgid "If you run this script successfully, you can see the info shown below:"
msgstr "如果您成功运行此脚本,您将看到如下所示的信息:"
#: ../../source/tutorials/models/Qwen3_embedding.md:96
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen3_embedding.md:98
msgid ""
"Run performance of `Qwen3-Reranker-8B` as an example. Refer to [vllm "
"benchmark](https://docs.vllm.ai/en/latest/contributing/) for more "
"details."
msgstr "以 `Qwen3-Reranker-8B` 的运行性能为例。更多详情请参考 [vllm 基准测试](https://docs.vllm.ai/en/latest/contributing/)。"
#: ../../source/tutorials/models/Qwen3_embedding.md:101
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。按如下方式运行代码。"
#: ../../source/tutorials/models/Qwen3_embedding.md:107
msgid ""
"After about several minutes, you can get the performance evaluation "
"result. With this tutorial, the performance result is:"
msgstr "大约几分钟后,您将获得性能评估结果。按照本教程,性能结果如下:"

View File

@@ -0,0 +1,165 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/Qwen3_reranker.md:1
msgid "Qwen3-Reranker"
msgstr "Qwen3-Reranker"
#: ../../source/tutorials/models/Qwen3_reranker.md:3
msgid "Introduction"
msgstr "简介"
#: ../../source/tutorials/models/Qwen3_reranker.md:5
msgid ""
"The Qwen3 Reranker model series is the latest proprietary model of the "
"Qwen family, specifically designed for text embedding and ranking tasks. "
"Building upon the dense foundational models of the Qwen3 series, it "
"provides a comprehensive range of text embeddings and reranking models in"
" various sizes (0.6B, 4B, and 8B). This guide describes how to run the "
"model with vLLM Ascend. Note that only 0.9.2rc1 and higher versions of "
"vLLM Ascend support the model."
msgstr ""
"Qwen3 Reranker 模型系列是 Qwen 家族最新的专有模型,专为文本嵌入和排序任务设计。它基于 Qwen3 系列的稠密基础模型提供了多种尺寸0.6B、4B 和 8B的全面文本嵌入和重排序模型。本指南描述了如何使用 vLLM Ascend 运行该模型。请注意,只有 vLLM Ascend 0.9.2rc1 及更高版本支持此模型。"
#: ../../source/tutorials/models/Qwen3_reranker.md:7
msgid "Supported Features"
msgstr "支持的功能"
#: ../../source/tutorials/models/Qwen3_reranker.md:9
msgid ""
"Refer to [supported "
"features](../../user_guide/support_matrix/supported_models.md) to get the"
" model's supported feature matrix."
msgstr "请参考[支持的功能](../../user_guide/support_matrix/supported_models.md)以获取该模型支持的功能矩阵。"
#: ../../source/tutorials/models/Qwen3_reranker.md:11
msgid "Environment Preparation"
msgstr "环境准备"
#: ../../source/tutorials/models/Qwen3_reranker.md:13
msgid "Model Weight"
msgstr "模型权重"
#: ../../source/tutorials/models/Qwen3_reranker.md:15
msgid ""
"`Qwen3-Reranker-8B` [Download model "
"weight](https://www.modelscope.cn/models/Qwen/Qwen3-Reranker-8B)"
msgstr "`Qwen3-Reranker-8B` [下载模型权重](https://www.modelscope.cn/models/Qwen/Qwen3-Reranker-8B)"
#: ../../source/tutorials/models/Qwen3_reranker.md:16
msgid ""
"`Qwen3-Reranker-4B` [Download model "
"weight](https://www.modelscope.cn/models/Qwen/Qwen3-Reranker-4B)"
msgstr "`Qwen3-Reranker-4B` [下载模型权重](https://www.modelscope.cn/models/Qwen/Qwen3-Reranker-4B)"
#: ../../source/tutorials/models/Qwen3_reranker.md:17
msgid ""
"`Qwen3-Reranker-0.6B` [Download model "
"weight](https://www.modelscope.cn/models/Qwen/Qwen3-Reranker-0.6B)"
msgstr "`Qwen3-Reranker-0.6B` [下载模型权重](https://www.modelscope.cn/models/Qwen/Qwen3-Reranker-0.6B)"
#: ../../source/tutorials/models/Qwen3_reranker.md:19
msgid ""
"It is recommended to download the model weight to the shared directory of"
" multiple nodes, such as `/root/.cache/`"
msgstr "建议将模型权重下载到多节点的共享目录中,例如 `/root/.cache/`"
#: ../../source/tutorials/models/Qwen3_reranker.md:21
msgid "Installation"
msgstr "安装"
#: ../../source/tutorials/models/Qwen3_reranker.md:23
msgid ""
"You can use our official docker image to run `Qwen3-Reranker` series "
"models."
msgstr "您可以使用我们的官方 docker 镜像来运行 `Qwen3-Reranker` 系列模型。"
#: ../../source/tutorials/models/Qwen3_reranker.md:25
msgid ""
"Start the docker image on your node, refer to [using "
"docker](../../installation.md#set-up-using-docker)."
msgstr "在您的节点上启动 docker 镜像,请参考[使用 docker](../../installation.md#set-up-using-docker)。"
#: ../../source/tutorials/models/Qwen3_reranker.md:27
msgid ""
"if you don't want to use the docker image as above, you can also build "
"all from source:"
msgstr "如果您不想使用上述 docker 镜像,也可以从源代码构建所有内容:"
#: ../../source/tutorials/models/Qwen3_reranker.md:29
msgid ""
"Install `vllm-ascend` from source, refer to "
"[installation](../../installation.md)."
msgstr "从源代码安装 `vllm-ascend`,请参考[安装](../../installation.md)。"
#: ../../source/tutorials/models/Qwen3_reranker.md:31
msgid "Deployment"
msgstr "部署"
#: ../../source/tutorials/models/Qwen3_reranker.md:33
msgid ""
"Using the Qwen3-Reranker-8B model as an example, first run the docker "
"container with the following command:"
msgstr "以 Qwen3-Reranker-8B 模型为例,首先使用以下命令运行 docker 容器:"
#: ../../source/tutorials/models/Qwen3_reranker.md:35
msgid "Online Inference"
msgstr "在线推理"
#: ../../source/tutorials/models/Qwen3_reranker.md:41
msgid "Once your server is started, you can send request with follow examples."
msgstr "服务器启动后,您可以按照以下示例发送请求。"
#: ../../source/tutorials/models/Qwen3_reranker.md:43
msgid "requests demo + formatting query & document"
msgstr "requests 演示 + 格式化查询和文档"
#: ../../source/tutorials/models/Qwen3_reranker.md:83
#: ../../source/tutorials/models/Qwen3_reranker.md:160
msgid ""
"If you run this script successfully, you will see a list of scores "
"printed to the console, similar to this:"
msgstr "如果成功运行此脚本,您将在控制台看到打印出的分数列表,类似于以下内容:"
#: ../../source/tutorials/models/Qwen3_reranker.md:89
msgid "Offline Inference"
msgstr "离线推理"
#: ../../source/tutorials/models/Qwen3_reranker.md:166
msgid "Performance"
msgstr "性能"
#: ../../source/tutorials/models/Qwen3_reranker.md:168
msgid ""
"Run performance of `Qwen3-Reranker-8B` as an example. Refer to [vllm "
"benchmark](https://docs.vllm.ai/en/latest/contributing/) for more "
"details."
msgstr "以 `Qwen3-Reranker-8B` 的运行性能为例。更多详情请参考 [vllm 基准测试](https://docs.vllm.ai/en/latest/contributing/)。"
#: ../../source/tutorials/models/Qwen3_reranker.md:171
msgid "Take the `serve` as an example. Run the code as follows."
msgstr "以 `serve` 为例。按如下方式运行代码。"
#: ../../source/tutorials/models/Qwen3_reranker.md:177
msgid ""
"After about several minutes, you can get the performance evaluation "
"result. With this tutorial, the performance result is:"
msgstr "大约几分钟后,您将获得性能评估结果。在本教程中,性能结果如下:"

View File

@@ -0,0 +1,29 @@
# SOME DESCRIPTIVE TITLE.
# Copyright (C) 2025, vllm-ascend team
# This file is distributed under the same license as the vllm-ascend
# package.
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
#
msgid ""
msgstr ""
"Project-Id-Version: vllm-ascend \n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language: zh_CN\n"
"Language-Team: zh_CN <LL@li.org>\n"
"Plural-Forms: nplurals=1; plural=0;\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: Babel 2.18.0\n"
#: ../../source/tutorials/models/index.md:1
#: ../../source/tutorials/models/index.md:5
msgid "Model Tutorials"
msgstr "模型教程"
#: ../../source/tutorials/models/index.md:3
msgid "This section provides tutorials for different models of vLLM Ascend."
msgstr "本节提供 vLLM Ascend 不同模型的使用教程。"