[v0.18.0][Doc] Translated Doc files 2026-04-14 (#8257)
## Auto-Translation Summary Translated **102** file(s): - <code>docs/source/locale/zh_CN/LC_MESSAGES/community/contributors.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/community/governance.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/community/user_stories/index.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/community/user_stories/llamafactory.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/community/versioning_policy.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/patch.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/contribution/index.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/contribution/testing.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/evaluation/using_evalscope.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/evaluation/using_lm_eval.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/evaluation/using_opencompass.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/performance_and_debug/msprobe_guide.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/performance_and_debug/performance_benchmark.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/performance_and_debug/service_profiling_guide.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/faqs.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/index.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/installation.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/quick_start.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/configuration/additional_config.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/graph_mode.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/lora.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/quantization.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/sleep_mode.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/structured_output.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/release_notes.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/support_matrix/index.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/support_matrix/supported_features.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/support_matrix/supported_models.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/ACL_Graph.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/KV_Cache_Pool_Guide.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/ModelRunner_prepare_inputs.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/add_custom_aclnn_op.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/context_parallel.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/cpu_binding.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/disaggregated_prefill.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/eplb_swift_balancer.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/npugraph_ex.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/Design_Documents/quantization.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/contribution/multi_node_test.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/evaluation/using_ais_bench.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/developer_guide/performance_and_debug/optimization_and_tuning.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/index.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/long_sequence_context_parallel_multi_node.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/long_sequence_context_parallel_single_node.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/pd_colocated_mooncake_multi_instance.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/pd_disaggregation_mooncake_multi_node.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/pd_disaggregation_mooncake_single_node.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/ray.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/features/suffix_speculative_decoding.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/hardwares/310p.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/hardwares/index.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/DeepSeek-R1.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/DeepSeek-V3.1.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/DeepSeek-V3.2.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/GLM4.x.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/GLM5.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Kimi-K2-Thinking.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Kimi-K2.5.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/MiniMax-M2.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/PaddleOCR-VL.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen-VL-Dense.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen2.5-7B.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen2.5-Omni.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-235B-A22B.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-30B-A3B.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-32B-W4A4.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-8B-W4A8.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-Coder-30B-A3B.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-Dense.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-Next.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-VL-235B-A22B-Instruct.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-VL-30B-A3B-Instruct.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-VL-Embedding.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3-VL-Reranker.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3.5-27B.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3.5-397B-A17B.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3_embedding.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/Qwen3_reranker.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/tutorials/models/index.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/deployment_guide/index.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/deployment_guide/using_volcano_kthena.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/Fine_grained_TP.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/Multi_Token_Prediction.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/batch_invariance.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/context_parallel.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/cpu_binding.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/dynamic_batch.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/epd_disaggregation.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/eplb_swift_balancer.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/external_dp.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/kv_pool.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/large_scale_ep.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/layer_sharding.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/lmcache_ascend_deployment.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/netloader.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/npugraph_ex.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/rfork.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/sequence_parallelism.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/speculative_decoding.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/ucm_deployment.po</code> - <code>docs/source/locale/zh_CN/LC_MESSAGES/user_guide/feature_guide/weight_prefetch.po</code> --- [Workflow run](https://github.com/vllm-project/vllm-ascend/actions/runs/24390263284) Signed-off-by: vllm-ascend-ci <vllm-ascend-ci@users.noreply.github.com> Co-authored-by: vllm-ascend-ci <vllm-ascend-ci@users.noreply.github.com>
This commit is contained in:
@@ -0,0 +1,283 @@
|
||||
# SOME DESCRIPTIVE TITLE.
|
||||
# Copyright (C) 2025, vllm-ascend team
|
||||
# This file is distributed under the same license as the vllm-ascend
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend \n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:1
|
||||
msgid "ACL Graph"
|
||||
msgstr "ACL 图"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:3
|
||||
msgid "Why do we need ACL Graph?"
|
||||
msgstr "为什么需要 ACL 图?"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:5
|
||||
msgid ""
|
||||
"In LLM inference, each token requires nearly a thousand operator "
|
||||
"executions. When host launching operators are slower than device, it will"
|
||||
" cause host bound. In severe cases, the device will be idle for more than"
|
||||
" half of the time. To solve this problem, we use graph in LLM inference."
|
||||
msgstr ""
|
||||
"在 LLM 推理中,每个 token 需要执行近千次算子。当主机(host)启动算子的速度慢于设备(device)时,会导致主机瓶颈(host bound)。在严重情况下,设备超过一半的时间将处于空闲状态。为了解决这个问题,我们在 LLM 推理中使用图(graph)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:26
|
||||
msgid "How to use ACL Graph?"
|
||||
msgstr "如何使用 ACL 图?"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:28
|
||||
msgid ""
|
||||
"ACL Graph is enabled by default in V1 Engine, you just need to check that"
|
||||
" `enforce_eager` is not set to `True`. More details see: [Graph Mode "
|
||||
"Guide](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/graph_mode.html)"
|
||||
msgstr ""
|
||||
"ACL 图在 V1 引擎中默认启用,您只需确认 `enforce_eager` 未设置为 `True`。更多详情请参阅:[图模式指南](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/graph_mode.html)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:30
|
||||
msgid "How it works?"
|
||||
msgstr "工作原理"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:32
|
||||
msgid ""
|
||||
"In short, graph mode works in two steps: **capture and replay**. When the"
|
||||
" engine starts, we capture all of the ops in the model forward and save "
|
||||
"it as a graph. When a request comes in, we just replay the graph on the "
|
||||
"device and wait for the result."
|
||||
msgstr ""
|
||||
"简而言之,图模式分两步工作:**捕获(capture)和重放(replay)**。当引擎启动时,我们捕获模型前向传播中的所有算子并将其保存为一个图。当请求到达时,我们只需在设备上重放该图并等待结果。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:34
|
||||
msgid "But in reality, graph mode is not that simple."
|
||||
msgstr "但实际上,图模式并非如此简单。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:36
|
||||
msgid "Padding and Bucketing"
|
||||
msgstr "填充与分桶"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:38
|
||||
msgid ""
|
||||
"Due to the fact that a graph can only replay the ops captured before, "
|
||||
"without doing tiling and checking graph input, we need to ensure the "
|
||||
"consistency of the graph input. However, we know that the model input's "
|
||||
"shape depends on the request scheduled by the Scheduler, so we can't "
|
||||
"ensure consistency."
|
||||
msgstr ""
|
||||
"由于图只能重放之前捕获的算子,而不会进行分片(tiling)或检查图输入,因此我们需要确保图输入的一致性。然而,我们知道模型输入的形状取决于调度器(Scheduler)安排的请求,因此无法保证一致性。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:40
|
||||
msgid ""
|
||||
"Obviously, we can solve this problem by capturing the biggest shape and "
|
||||
"padding all of the model inputs to it. But this will bring a lot of "
|
||||
"redundant computing and make performance worse. So we can capture "
|
||||
"multiple graphs with different shapes, and pad the model input to the "
|
||||
"nearest graph, which will greatly reduce redundant computing. But when "
|
||||
"`max_num_batched_tokens` is very large, the number of graphs that need to"
|
||||
" be captured will also become very large. We know that when the input "
|
||||
"tensor's shape is large, the computing time will be very long, and graph "
|
||||
"mode is not necessary in this case. So all of the things we need to do "
|
||||
"are:"
|
||||
msgstr ""
|
||||
"显然,我们可以通过捕获最大形状并将所有模型输入填充到该形状来解决此问题。但这会带来大量冗余计算并使性能变差。因此,我们可以捕获多个不同形状的图,并将模型输入填充到最接近的图,这将大大减少冗余计算。但当 `max_num_batched_tokens` 非常大时,需要捕获的图数量也会变得非常大。我们知道,当输入张量的形状很大时,计算时间会很长,在这种情况下图模式并非必要。因此,我们需要做的所有事情是:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:42
|
||||
msgid "Set a threshold;"
|
||||
msgstr "设置一个阈值;"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:43
|
||||
msgid ""
|
||||
"When `num_scheduled_tokens` is bigger than the threshold, use "
|
||||
"`eager_mode`;"
|
||||
msgstr "当 `num_scheduled_tokens` 大于阈值时,使用 `eager_mode`;"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:44
|
||||
msgid "Capture multiple graphs within a range below the threshold;"
|
||||
msgstr "在低于阈值的范围内捕获多个图;"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:59
|
||||
msgid "Piecewise and Full graph"
|
||||
msgstr "分段图与完整图"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:61
|
||||
msgid ""
|
||||
"Due to the increasing complexity of the attention layer in current LLMs, "
|
||||
"we can't ensure all types of attention can run in graph. In MLA, "
|
||||
"prefill_tokens and decode_tokens have different calculation methods, so "
|
||||
"when a batch has both prefills and decodes in MLA, graph mode is "
|
||||
"difficult to handle this situation."
|
||||
msgstr ""
|
||||
"由于当前 LLM 中注意力层的复杂性不断增加,我们无法确保所有类型的注意力都能在图模式下运行。在 MLA 中,prefill_tokens 和 decode_tokens 有不同的计算方法,因此当 MLA 中的一个批次同时包含预填充和解码时,图模式难以处理这种情况。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:63
|
||||
msgid ""
|
||||
"vLLM solves this problem with piecewise graph mode. We use eager mode to "
|
||||
"launch attention's ops, and use graph to deal with others. But this also "
|
||||
"brings some problems: The cost of launching ops has become large again. "
|
||||
"Although much smaller than eager mode, it will also lead to host bound "
|
||||
"when the CPU is poor or `num_tokens` is small."
|
||||
msgstr ""
|
||||
"vLLM 通过分段图模式解决了这个问题。我们使用 eager 模式来启动注意力算子,并使用图来处理其他算子。但这也会带来一些问题:启动算子的开销再次变大。虽然比 eager 模式小得多,但当 CPU 性能较差或 `num_tokens` 较小时,仍会导致主机瓶颈。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:65
|
||||
msgid "Altogether, we need to support both piecewise and full graph mode."
|
||||
msgstr "总之,我们需要同时支持分段图和完整图模式。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:67
|
||||
msgid ""
|
||||
"When attention can run in graph, we tend to choose full graph mode to "
|
||||
"achieve optimal performance;"
|
||||
msgstr "当注意力可以在图中运行时,我们倾向于选择完整图模式以获得最佳性能;"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:68
|
||||
msgid "When full graph does not work, use piecewise graph as a substitute;"
|
||||
msgstr "当完整图无法工作时,使用分段图作为替代;"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:69
|
||||
msgid ""
|
||||
"When piecewise graph's performance is not good and full graph mode is "
|
||||
"blocked, separate prefills and decodes, and use full graph mode in "
|
||||
"**decode_only** situations. Because when a batch includes prefill "
|
||||
"requests, usually `num_tokens` will be quite big and not cause host "
|
||||
"bound."
|
||||
msgstr ""
|
||||
"当分段图性能不佳且完整图模式受阻时,将预填充和解码分离,并在 **decode_only** 情况下使用完整图模式。因为当一个批次包含预填充请求时,通常 `num_tokens` 会相当大,不会导致主机瓶颈。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:71
|
||||
msgid ""
|
||||
"Currently, due to stream resource constraint, we can only support a few "
|
||||
"buckets in piecewise graph mode now, which will cause redundant computing"
|
||||
" and may lead to performance degradation compared with eager mode."
|
||||
msgstr ""
|
||||
"目前,由于流资源限制,我们现在只能在分段图模式下支持少数几个桶(buckets),这会导致冗余计算,并且与 eager 模式相比可能导致性能下降。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:73
|
||||
msgid "How is it implemented?"
|
||||
msgstr "如何实现?"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:75
|
||||
msgid ""
|
||||
"vLLM has already implemented most of the modules in graph mode. You can "
|
||||
"see more details at: [CUDA "
|
||||
"Graphs](https://docs.vllm.ai/en/latest/design/cuda_graphs.html)"
|
||||
msgstr ""
|
||||
"vLLM 已经在图模式下实现了大部分模块。您可以在以下链接查看更多详情:[CUDA 图](https://docs.vllm.ai/en/latest/design/cuda_graphs.html)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:77
|
||||
msgid ""
|
||||
"When in graph mode, vLLM will call "
|
||||
"`current_platform.get_static_graph_wrapper_cls` to get the current "
|
||||
"device's graph model wrapper, so what we need to do is implement the "
|
||||
"graph mode wrapper on Ascend: `ACLGraphWrapper`."
|
||||
msgstr ""
|
||||
"在图模式下,vLLM 会调用 `current_platform.get_static_graph_wrapper_cls` 来获取当前设备的图模型包装器,因此我们需要做的是在 Ascend 上实现图模式包装器:`ACLGraphWrapper`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:79
|
||||
msgid ""
|
||||
"vLLM has added `support_torch_compile` decorator to all models. This "
|
||||
"decorator will replace the `__init__` and `forward` interface of the "
|
||||
"model class. When `forward` is called, the code inside the "
|
||||
"`ACLGraphWrapper` will be executed, and it will do capture or replay as "
|
||||
"mentioned above."
|
||||
msgstr ""
|
||||
"vLLM 已为所有模型添加了 `support_torch_compile` 装饰器。此装饰器将替换模型类的 `__init__` 和 `forward` 接口。当调用 `forward` 时,`ACLGraphWrapper` 内部的代码将被执行,并执行如上所述的捕获或重放操作。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:81
|
||||
msgid ""
|
||||
"When using piecewise graph, we just need to follow the above-mentioned "
|
||||
"process. But when in full graph, due to the complexity of the attention, "
|
||||
"sometimes we need to update attention op's params before execution. So we"
|
||||
" implement `update_attn_params` and `update_mla_attn_params` functions "
|
||||
"for full graph mode. During forward, memory will be reused between "
|
||||
"different ops, so we can't update attention op's params before forward. "
|
||||
"In ACL Graph, we use `torch.npu.graph_task_update_begin` and "
|
||||
"`torch.npu.graph_task_update_end` to do it, and use "
|
||||
"`torch.npu.ExternalEvent` to ensure order between param updates and op "
|
||||
"executions."
|
||||
msgstr ""
|
||||
"使用分段图时,我们只需遵循上述流程。但在完整图模式下,由于注意力的复杂性,有时我们需要在执行前更新注意力算子的参数。因此,我们为完整图模式实现了 `update_attn_params` 和 `update_mla_attn_params` 函数。在前向传播期间,内存会在不同算子之间重用,因此我们无法在前向传播之前更新注意力算子的参数。在 ACL 图中,我们使用 `torch.npu.graph_task_update_begin` 和 `torch.npu.graph_task_update_end` 来实现这一点,并使用 `torch.npu.ExternalEvent` 来确保参数更新与算子执行之间的顺序。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:83
|
||||
msgid "DFX"
|
||||
msgstr "DFX"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:85
|
||||
msgid "Stream resource constraint"
|
||||
msgstr "流资源限制"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:87
|
||||
msgid ""
|
||||
"Currently, we can only capture 1800 graphs at most, due to the limitation"
|
||||
" of ACL graph that a graph requires at least a separate stream. This "
|
||||
"number is bounded by the number of streams, which is 2048; we save 248 "
|
||||
"streams as a buffer. Besides, there are many variables that can affect "
|
||||
"the number of buckets:"
|
||||
msgstr ""
|
||||
"目前,由于 ACL 图的限制(一个图至少需要一个独立的流),我们最多只能捕获 1800 个图。这个数字受限于流的数量,即 2048;我们保留 248 个流作为缓冲区。此外,还有许多变量会影响桶的数量:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:89
|
||||
msgid ""
|
||||
"Piecewise graph divides the model into `num_hidden_layers + 1` sub "
|
||||
"modules, based on the attention layer. Every sub module is a single graph"
|
||||
" which needs to cost a stream, so the number of buckets in piecewise "
|
||||
"graph mode is very tight compared with full graph mode."
|
||||
msgstr ""
|
||||
"分段图根据注意力层将模型划分为 `num_hidden_layers + 1` 个子模块。每个子模块都是一个单独的图,需要消耗一个流,因此与完整图模式相比,分段图模式下的桶数量非常紧张。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:91
|
||||
msgid ""
|
||||
"The number of streams required for a graph is related to the number of "
|
||||
"comm domains. Each comm domain will increase one stream consumed by a "
|
||||
"graph."
|
||||
msgstr "一个图所需的流数量与通信域(comm domain)的数量有关。每个通信域都会增加一个图消耗的流。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:93
|
||||
msgid ""
|
||||
"When multi-stream is explicitly called in a sub module, it will consume "
|
||||
"an additional stream."
|
||||
msgstr "当在子模块中显式调用多流(multi-stream)时,它将消耗一个额外的流。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:95
|
||||
msgid ""
|
||||
"There are some other rules about ACL Graph and stream. Currently, we use "
|
||||
"func `update_aclgraph_sizes` to calculate the maximum number of buckets "
|
||||
"and update `graph_batch_sizes` to ensure stream resource is sufficient."
|
||||
msgstr ""
|
||||
"关于 ACL 图和流还有一些其他规则。目前,我们使用函数 `update_aclgraph_sizes` 来计算最大桶数并更新 `graph_batch_sizes`,以确保流资源充足。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:97
|
||||
msgid "We will expand the stream resource limitation in the future."
|
||||
msgstr "我们将在未来扩展流资源限制。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:99
|
||||
msgid "Limitations"
|
||||
msgstr "限制"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:101
|
||||
msgid "`FULL` and `FULL_AND_PIECEWISE` are not supported now;"
|
||||
msgstr "目前不支持 `FULL` 和 `FULL_AND_PIECEWISE`;"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:102
|
||||
msgid ""
|
||||
"When use ACL Graph and MTP and `num_speculative_tokens > 1`, as vLLM "
|
||||
"don't support this case in v0.11.0, we need to set "
|
||||
"`cudagraph_capture_sizes` explicitly."
|
||||
msgstr ""
|
||||
"当使用 ACL 图和 MTP 且 `num_speculative_tokens > 1` 时,由于 vLLM 在 v0.11.0 中不支持此情况,我们需要显式设置 `cudagraph_capture_sizes`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ACL_Graph.md:103
|
||||
msgid "`use_inductor` is not supported now;"
|
||||
msgstr "目前不支持 `use_inductor`;"
|
||||
@@ -0,0 +1,309 @@
|
||||
# SOME DESCRIPTIVE TITLE.
|
||||
# Copyright (C) 2025, vllm-ascend team
|
||||
# This file is distributed under the same license as the vllm-ascend
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend \n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:1
|
||||
msgid "KV Cache Pool"
|
||||
msgstr "KV 缓存池"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:3
|
||||
msgid "Why KV Cache Pool?"
|
||||
msgstr "为什么需要 KV 缓存池?"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:5
|
||||
msgid ""
|
||||
"Prefix caching is an important feature in LLM inference that can reduce "
|
||||
"prefill computation time drastically."
|
||||
msgstr "前缀缓存是大语言模型推理中的一项重要特性,可以显著减少预填充计算时间。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:7
|
||||
msgid ""
|
||||
"However, the performance gain from prefix caching is highly dependent on "
|
||||
"the cache hit rate, while the cache hit rate can be limited if one only "
|
||||
"uses HBM for KV cache storage."
|
||||
msgstr "然而,前缀缓存带来的性能提升高度依赖于缓存命中率,而如果仅使用 HBM 存储 KV 缓存,缓存命中率会受到限制。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:9
|
||||
msgid ""
|
||||
"Hence, KV Cache Pool is proposed to utilize various types of storage "
|
||||
"including HBM, DRAM, and SSD, making a pool for KV Cache storage while "
|
||||
"making the prefix of requests visible across all nodes, increasing the "
|
||||
"cache hit rate for all requests."
|
||||
msgstr "因此,我们提出了 KV 缓存池,旨在利用包括 HBM、DRAM 和 SSD 在内的多种存储类型,构建一个 KV 缓存存储池,同时使请求的前缀在所有节点间可见,从而提高所有请求的缓存命中率。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:11
|
||||
msgid ""
|
||||
"vLLM Ascend currently supports [MooncakeStore](https://github.com"
|
||||
"/kvcache-ai/Mooncake), one of the most recognized KV Cache storage "
|
||||
"engines."
|
||||
msgstr "vLLM Ascend 目前支持 [MooncakeStore](https://github.com/kvcache-ai/Mooncake),这是最受认可的 KV 缓存存储引擎之一。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:13
|
||||
msgid ""
|
||||
"While one can utilize Mooncake Store in vLLM V1 engine by setting it as a"
|
||||
" remote backend of LMCache with GPU (see "
|
||||
"[Tutorial](https://github.com/LMCache/LMCache/blob/dev/examples/kv_cache_reuse/remote_backends/mooncakestore/README.md)),"
|
||||
" we find it would be better to integrate a connector that directly "
|
||||
"supports Mooncake Store and can utilize the data transfer strategy that "
|
||||
"best fits Huawei NPU hardware."
|
||||
msgstr "虽然可以通过将 Mooncake Store 设置为 GPU 上 LMCache 的远程后端来在 vLLM V1 引擎中使用它(参见[教程](https://github.com/LMCache/LMCache/blob/dev/examples/kv_cache_reuse/remote_backends/mooncakestore/README.md)),但我们认为集成一个直接支持 Mooncake Store 并能利用最适合华为 NPU 硬件的数据传输策略的连接器会更好。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:15
|
||||
msgid ""
|
||||
"Hence, we propose to integrate Mooncake Store with a brand new "
|
||||
"**MooncakeStoreConnectorV1**, which is indeed largely inspired by "
|
||||
"**LMCacheConnectorV1** (see the `How is MooncakeStoreConnectorV1 "
|
||||
"Implemented?` section)."
|
||||
msgstr "因此,我们提议将 Mooncake Store 与全新的 **MooncakeStoreConnectorV1** 集成,该连接器的设计在很大程度上受到了 **LMCacheConnectorV1** 的启发(参见 `MooncakeStoreConnectorV1 是如何实现的?` 部分)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:17
|
||||
msgid "Usage"
|
||||
msgstr "使用方法"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:19
|
||||
msgid ""
|
||||
"vLLM Ascend currently supports Mooncake Store for KV Cache Pool. To "
|
||||
"enable Mooncake Store, one needs to configure `kv-transfer-config` and "
|
||||
"choose `MooncakeStoreConnector` as the KV Connector."
|
||||
msgstr "vLLM Ascend 目前支持使用 Mooncake Store 作为 KV 缓存池。要启用 Mooncake Store,需要配置 `kv-transfer-config` 并选择 `MooncakeStoreConnector` 作为 KV 连接器。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:21
|
||||
msgid ""
|
||||
"For step-by-step deployment and configuration, please refer to the [KV "
|
||||
"Pool User "
|
||||
"Guide](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/kv_pool.html)."
|
||||
msgstr "关于逐步部署和配置,请参考 [KV 池用户指南](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/kv_pool.html)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:23
|
||||
msgid "How it works?"
|
||||
msgstr "工作原理"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:25
|
||||
msgid ""
|
||||
"The KV Cache Pool integrates multiple memory tiers (HBM, DRAM, SSD, etc.)"
|
||||
" through a connector-based architecture."
|
||||
msgstr "KV 缓存池通过基于连接器的架构,整合了多个内存层级(HBM、DRAM、SSD 等)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:27
|
||||
msgid ""
|
||||
"Each connector implements a unified interface for storing, retrieving, "
|
||||
"and transferring KV blocks between tiers, depending on access frequency "
|
||||
"and hardware bandwidth."
|
||||
msgstr "每个连接器实现了一个统一的接口,用于根据访问频率和硬件带宽在不同层级之间存储、检索和传输 KV 块。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:29
|
||||
msgid ""
|
||||
"When combined with vLLM’s Prefix Caching mechanism, the pool enables "
|
||||
"efficient caching both locally (in HBM) and globally (via Mooncake), "
|
||||
"ensuring that frequently used prefixes remain hot while less frequently "
|
||||
"accessed KV data can spill over to lower-cost memory."
|
||||
msgstr "当与 vLLM 的前缀缓存机制结合时,该池能够实现本地(HBM 中)和全局(通过 Mooncake)的高效缓存,确保常用前缀保持热状态,而访问频率较低的 KV 数据则可以溢出到成本更低的内存中。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:31
|
||||
msgid "1. Combining KV Cache Pool with HBM Prefix Caching"
|
||||
msgstr "1. 将 KV 缓存池与 HBM 前缀缓存结合"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:33
|
||||
msgid ""
|
||||
"Prefix Caching with HBM is already supported by the vLLM V1 Engine. By "
|
||||
"introducing KV Connector V1, users can seamlessly combine HBM-based "
|
||||
"Prefix Caching with Mooncake-backed KV Pool."
|
||||
msgstr "vLLM V1 引擎已支持基于 HBM 的前缀缓存。通过引入 KV Connector V1,用户可以无缝地将基于 HBM 的前缀缓存与 Mooncake 支持的 KV 池结合起来。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:36
|
||||
msgid ""
|
||||
"The user can enable both features simply by enabling Prefix Caching, "
|
||||
"which is enabled by default in vLLM V1 unless the "
|
||||
"`--no_enable_prefix_caching` flag is set, and setting up the KV Connector"
|
||||
" for KV Pool (e.g., the MooncakeStoreConnector)."
|
||||
msgstr "用户只需启用前缀缓存(在 vLLM V1 中默认启用,除非设置了 `--no_enable_prefix_caching` 标志)并为 KV 池设置 KV 连接器(例如 MooncakeStoreConnector),即可同时启用这两个功能。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:38
|
||||
msgid "**Workflow**:"
|
||||
msgstr "**工作流程**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:40
|
||||
msgid "The engine first checks for prefix hits in the HBM cache."
|
||||
msgstr "引擎首先检查 HBM 缓存中的前缀命中情况。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:42
|
||||
msgid ""
|
||||
"After getting the number of hit tokens on HBM, it queries the KV Pool via"
|
||||
" the connector. If there are additional hits in the KV Pool, we get the "
|
||||
"**additional blocks only** from the KV Pool, and get the rest of the "
|
||||
"blocks directly from HBM to minimize the data transfer latency."
|
||||
msgstr "获取 HBM 上的命中令牌数量后,引擎通过连接器查询 KV 池。如果在 KV 池中有额外的命中,我们**仅从 KV 池获取额外的块**,其余块则直接从 HBM 获取,以最小化数据传输延迟。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:44
|
||||
msgid ""
|
||||
"After the KV Caches in the KV Pool are loaded into HBM, the remaining "
|
||||
"process is the same as Prefix Caching in HBM."
|
||||
msgstr "将 KV 池中的 KV 缓存加载到 HBM 后,剩余过程与 HBM 中的前缀缓存相同。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:46
|
||||
msgid "2. Combining KV Cache Pool with Mooncake PD Disaggregation"
|
||||
msgstr "2. 将 KV 缓存池与 Mooncake PD 解耦结合"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:48
|
||||
msgid ""
|
||||
"When used together with Mooncake PD (Prefill-Decode) Disaggregation, the "
|
||||
"KV Cache Pool can further decouple prefill and decode stages across "
|
||||
"devices or nodes."
|
||||
msgstr "当与 Mooncake PD(预填充-解码)解耦功能结合使用时,KV 缓存池可以进一步在设备或节点间解耦预填充和解码阶段。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:50
|
||||
msgid ""
|
||||
"Currently, we only perform put and get operations of KV Pool for "
|
||||
"**Prefill Nodes**, and Decode Nodes get their KV Cache from Mooncake P2P "
|
||||
"KV Connector, i.e., MooncakeConnector."
|
||||
msgstr "目前,我们仅对**预填充节点**执行 KV 池的 put 和 get 操作,解码节点则通过 Mooncake P2P KV 连接器(即 MooncakeConnector)获取其 KV 缓存。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:52
|
||||
msgid ""
|
||||
"The key benefit of doing this is that we can keep the gain in performance"
|
||||
" by computing less with Prefix Caching from HBM and KV Pool for Prefill "
|
||||
"Nodes, while not sacrificing the data transfer efficiency between Prefill"
|
||||
" and Decode nodes with P2P KV Connector that transfers KV Caches between "
|
||||
"NPU devices directly."
|
||||
msgstr "这样做的主要好处是,我们可以通过为预填充节点使用来自 HBM 和 KV 池的前缀缓存来减少计算量,从而保持性能增益,同时又不牺牲预填充节点与解码节点之间的数据传输效率,因为 P2P KV 连接器直接在 NPU 设备间传输 KV 缓存。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:54
|
||||
msgid ""
|
||||
"To enable this feature, we need to set up both Mooncake Connector and "
|
||||
"Mooncake Store Connector with a Multi Connector, which is a KV Connector "
|
||||
"class provided by vLLM that can call multiple KV Connectors in a specific"
|
||||
" order."
|
||||
msgstr "要启用此功能,我们需要使用 Multi Connector 来设置 Mooncake Connector 和 Mooncake Store Connector。Multi Connector 是 vLLM 提供的一个 KV 连接器类,可以按特定顺序调用多个 KV 连接器。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:56
|
||||
msgid ""
|
||||
"For details, please also refer to the Mooncake Connector Store Deployment"
|
||||
" Guide."
|
||||
msgstr "详情请参阅 Mooncake Connector Store 部署指南。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:58
|
||||
msgid "How is MooncakeStoreConnectorV1 Implemented?"
|
||||
msgstr "MooncakeStoreConnectorV1 是如何实现的?"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:60
|
||||
msgid ""
|
||||
"**MooncakeStoreConnectorV1** inherits the KV Connector V1 class in vLLM "
|
||||
"V1: through implementing the required methods defined in the KV connector"
|
||||
" V1 base class, one can integrate a third-party KV cache transfer/storage"
|
||||
" backend into the vLLM framework."
|
||||
msgstr "**MooncakeStoreConnectorV1** 继承自 vLLM V1 中的 KV Connector V1 类:通过实现 KV 连接器 V1 基类中定义的必要方法,可以将第三方 KV 缓存传输/存储后端集成到 vLLM 框架中。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:62
|
||||
msgid ""
|
||||
"MooncakeStoreConnectorV1 is also largely inspired by LMCacheConnectorV1 "
|
||||
"in terms of the `Lookup Engine`/`Lookup Client` design for looking up KV "
|
||||
"cache keys, and the `ChunkedTokenDatabase` class for processing tokens "
|
||||
"into prefix-aware hashes as well as other hashing related designs. On top"
|
||||
" of this, we have also added our own design including `KVTransferThread` "
|
||||
"that allows async `get` and `put` of KV caches with multi-threading, and "
|
||||
"NPU-related data transfer optimization such as removing the `LocalBuffer`"
|
||||
" in LMCache to remove redundant data transfer."
|
||||
msgstr "MooncakeStoreConnectorV1 也在很大程度上借鉴了 LMCacheConnectorV1,包括用于查找 KV 缓存键的 `Lookup Engine`/`Lookup Client` 设计,以及用于将令牌处理为前缀感知哈希的 `ChunkedTokenDatabase` 类和其他哈希相关设计。在此基础上,我们还添加了自己的设计,包括允许通过多线程异步 `get` 和 `put` KV 缓存的 `KVTransferThread`,以及与 NPU 相关的数据传输优化,例如移除 LMCache 中的 `LocalBuffer` 以消除冗余数据传输。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:64
|
||||
msgid ""
|
||||
"The KV Connector methods that need to be implemented can be categorized "
|
||||
"into scheduler-side methods that are called in V1 scheduler and worker-"
|
||||
"side methods that are called in V1 worker, namely:"
|
||||
msgstr "需要实现的 KV 连接器方法可以分为在 V1 调度器中调用的调度器端方法和在 V1 工作器中调用的工作器端方法,即:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:66
|
||||
msgid "KV Connector Scheduler-Side Methods"
|
||||
msgstr "KV 连接器调度器端方法"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:68
|
||||
msgid ""
|
||||
"`get_num_new_matched_tokens`: Get prefix cache hit in number of tokens "
|
||||
"through looking up into the KV pool. `update_states_after_alloc`: "
|
||||
"Update KVConnector state after temporary buffer alloc. "
|
||||
"`build_connector_meta`: Attach the connector metadata to the request "
|
||||
"object. `request_finished`: Once a request is finished, determine "
|
||||
"whether request blocks should be freed now or will be sent asynchronously"
|
||||
" and freed later."
|
||||
msgstr ""
|
||||
"`get_num_new_matched_tokens`:通过查询 KV 池,获取以令牌数表示的前缀缓存命中数。\n"
|
||||
"`update_states_after_alloc`:临时缓冲区分配后更新 KVConnector 状态。\n"
|
||||
"`build_connector_meta`:将连接器元数据附加到请求对象。\n"
|
||||
"`request_finished`:请求完成后,确定请求块是应立即释放,还是将异步发送并稍后释放。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:73
|
||||
msgid "Connector Worker-Side Methods"
|
||||
msgstr "连接器工作器端方法"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:75
|
||||
msgid ""
|
||||
"`register_kv_caches`: Register KV cache buffers needed for KV cache "
|
||||
"transfer. `start_load_kv`: Perform KV cache load operation that transfers"
|
||||
" KV cache from storage to device. `wait_for_layer_load`: Optional; Wait "
|
||||
"for layer load in layerwise + async KV load scenario. `save_kv_layer`: "
|
||||
"Optional; Do layerwise KV cache put into KV Pool. `wait_for_save`: Wait "
|
||||
"for KV Save to finish if async KV cache save/put. `get_finished`: Get "
|
||||
"request that finished KV transfer, `done_sending` if `put` finished, "
|
||||
"`done_receiving` if `get` finished."
|
||||
msgstr ""
|
||||
"`register_kv_caches`:注册 KV 缓存传输所需的 KV 缓存缓冲区。\n"
|
||||
"`start_load_kv`:执行 KV 缓存加载操作,将 KV 缓存从存储传输到设备。\n"
|
||||
"`wait_for_layer_load`:可选;在分层 + 异步 KV 加载场景中等待层加载。\n"
|
||||
"`save_kv_layer`:可选;执行分层 KV 缓存放入 KV 池的操作。\n"
|
||||
"`wait_for_save`:如果异步保存/放入 KV 缓存,则等待 KV 保存完成。\n"
|
||||
"`get_finished`:获取已完成 KV 传输的请求,如果 `put` 完成则为 `done_sending`,如果 `get` 完成则为 `done_receiving`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:82
|
||||
msgid "DFX"
|
||||
msgstr "DFX(可诊断性、可维护性、可服务性)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:84
|
||||
msgid ""
|
||||
"When looking up a key in KV Pool, if we cannot find the key, there is no "
|
||||
"Cache Hit for this specific block; we return no hit for this block and do"
|
||||
" not look up further blocks for the current request."
|
||||
msgstr "在 KV 池中查找键时,如果找不到该键,则此特定块没有缓存命中;我们返回此块未命中,并且不再为当前请求查找后续块。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:85
|
||||
msgid ""
|
||||
"Similarly, when we are trying to put a block into KV Pool and it fails, "
|
||||
"we do not put further blocks (subject to change)."
|
||||
msgstr "类似地,当我们尝试将一个块放入 KV 池但失败时,我们不会放入后续块(可能更改)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:87
|
||||
msgid "Limitations"
|
||||
msgstr "限制"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:89
|
||||
msgid ""
|
||||
"Currently, Mooncake Store for vLLM-Ascend only supports DRAM as the "
|
||||
"storage for KV Cache pool."
|
||||
msgstr "目前,vLLM-Ascend 的 Mooncake Store 仅支持 DRAM 作为 KV 缓存池的存储。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/KV_Cache_Pool_Guide.md:91
|
||||
msgid ""
|
||||
"For now, if we successfully looked up a key and found it exists, but "
|
||||
"failed to get it when calling KV Pool's get function, we just output a "
|
||||
"log indicating the get operation failed and keep going; hence, the "
|
||||
"accuracy of that specific request may be affected. We will handle this "
|
||||
"situation by falling back the request and re-compute everything assuming "
|
||||
"there's no prefix cache hit (or even better, revert only one block and "
|
||||
"keep using the Prefix Caches before that)."
|
||||
msgstr "目前,如果我们成功查找到一个键并发现它存在,但在调用 KV 池的 get 函数时失败,我们仅输出一条日志表明 get 操作失败并继续执行;因此,该特定请求的准确性可能会受到影响。我们将通过回退请求并假设没有前缀缓存命中来重新计算所有内容(或者更好的是,仅回退一个块并继续使用该块之前的前缀缓存)来处理这种情况。"
|
||||
@@ -0,0 +1,629 @@
|
||||
# SOME DESCRIPTIVE TITLE.
|
||||
# Copyright (C) 2025, vllm-ascend team
|
||||
# This file is distributed under the same license as the vllm-ascend
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend \n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:1
|
||||
msgid "Prepare inputs for model forwarding"
|
||||
msgstr "为模型前向传播准备输入"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:3
|
||||
msgid "Purpose"
|
||||
msgstr "目的"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:5
|
||||
msgid "Information required to perform model forward pass:"
|
||||
msgstr "执行模型前向传播所需的信息:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:7
|
||||
msgid "the inputs"
|
||||
msgstr "输入"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:8
|
||||
msgid "the corresponding attention metadata of the inputs"
|
||||
msgstr "输入对应的注意力元数据"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:10
|
||||
msgid "The following diagram shows what we should prepare for model inference."
|
||||
msgstr "下图展示了我们需要为模型推理准备的内容。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:20
|
||||
msgid ""
|
||||
"Therefore, as long as we have these two pieces of information mentioned "
|
||||
"above, we can perform the model's forward propagation."
|
||||
msgstr "因此,只要我们拥有上述两方面的信息,就可以执行模型的前向传播。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:22
|
||||
msgid ""
|
||||
"This document will explain **how we obtain the inputs and their "
|
||||
"corresponding attention metadata**."
|
||||
msgstr "本文将解释**我们如何获取输入及其对应的注意力元数据**。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:24
|
||||
msgid "Overview"
|
||||
msgstr "概述"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:26
|
||||
msgid "1. Obtain inputs"
|
||||
msgstr "1. 获取输入"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:28
|
||||
msgid "The workflow of obtaining inputs:"
|
||||
msgstr "获取输入的工作流程:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:30
|
||||
msgid ""
|
||||
"Get `token positions`: relative position of each token within its request"
|
||||
" sequence."
|
||||
msgstr "获取 `token positions`:每个 token 在其请求序列中的相对位置。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:32
|
||||
msgid "Get `token indices`: index of each scheduled token in the token table."
|
||||
msgstr "获取 `token indices`:每个已调度 token 在 token 表中的索引。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:34
|
||||
msgid ""
|
||||
"Get `Token IDs`: using token indices to retrieve the Token IDs from "
|
||||
"**token id table**."
|
||||
msgstr "获取 `Token IDs`:使用 token indices 从 **token id table** 中检索 Token IDs。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:36
|
||||
msgid ""
|
||||
"At last, these `Token IDs` are required to be fed into a model, and "
|
||||
"`positions` should also be sent into the model to create `Rope` (Rotary "
|
||||
"positional embedding). Both of them are the inputs of the model."
|
||||
msgstr "最后,这些 `Token IDs` 需要输入到模型中,`positions` 也需要送入模型以创建 `Rope`(旋转位置编码)。两者共同构成模型的输入。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:38
|
||||
msgid ""
|
||||
"**Note**: The `Token IDs` are the inputs of a model, so we also call them"
|
||||
" `Inputs IDs`."
|
||||
msgstr "**注意**:`Token IDs` 是模型的输入,因此我们也称它们为 `Inputs IDs`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:40
|
||||
msgid "2. Build inputs attention metadata"
|
||||
msgstr "2. 构建输入注意力元数据"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:42
|
||||
msgid "A model requires these attention metadata during the forward pass:"
|
||||
msgstr "模型在前向传播过程中需要以下注意力元数据:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:44
|
||||
msgid ""
|
||||
"`query start location`: start and end location of each request "
|
||||
"corresponding to the scheduled tokens."
|
||||
msgstr "`query start location`:每个请求对应的已调度 token 的起始和结束位置。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:45
|
||||
msgid ""
|
||||
"`sequence length`: length of each request including both computed tokens "
|
||||
"and newly scheduled tokens."
|
||||
msgstr "`sequence length`:每个请求的长度,包括已计算 token 和新调度的 token。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:46
|
||||
msgid "`number of computed tokens`: number of computed tokens for each request."
|
||||
msgstr "`number of computed tokens`:每个请求已计算 token 的数量。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:47
|
||||
msgid "`number of requests`: number of requests in this batch."
|
||||
msgstr "`number of requests`:本批次中的请求数量。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:48
|
||||
msgid "`number of tokens`: total number of scheduled tokens in this batch."
|
||||
msgstr "`number of tokens`:本批次中已调度 token 的总数。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:49
|
||||
msgid ""
|
||||
"**`block table`**: translates the logical address (within its sequence) "
|
||||
"of each block to its global physical address in the device's memory."
|
||||
msgstr "**`block table`**:将每个块在其序列内的逻辑地址转换为其在设备内存中的全局物理地址。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:50
|
||||
msgid ""
|
||||
"`max query len`: the longest scheduled tokens length in this request "
|
||||
"batch."
|
||||
msgstr "`max query len`:本请求批次中最长的已调度 token 长度。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:51
|
||||
msgid ""
|
||||
"`slot mapping`: indices of each token that input token will be stored "
|
||||
"into."
|
||||
msgstr "`slot mapping`:输入 token 将被存储到的每个 token 的索引。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:52
|
||||
msgid ""
|
||||
"`attention mask`: mask matrix applied to attention scores before softmax "
|
||||
"to control which tokens can attend to each other (usually a causal "
|
||||
"attention)."
|
||||
msgstr "`attention mask`:在 softmax 之前应用于注意力分数的掩码矩阵,用于控制哪些 token 可以相互关注(通常是因果注意力)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:54
|
||||
msgid "Before start"
|
||||
msgstr "开始之前"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:56
|
||||
msgid "There are mainly three types of variables."
|
||||
msgstr "主要有三种类型的变量。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:58
|
||||
msgid ""
|
||||
"token level: represents one attribute corresponding to each scheduled "
|
||||
"token, so the length of this variable is the number of scheduled tokens."
|
||||
msgstr "token 级别:代表每个已调度 token 对应的一个属性,因此该变量的长度等于已调度 token 的数量。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:59
|
||||
msgid ""
|
||||
"request level: represents one attribute of each scheduled request, whose "
|
||||
"length usually is the number of scheduled requests. (`query start "
|
||||
"location` is a special case, which has one more element.)"
|
||||
msgstr "请求级别:代表每个已调度请求的一个属性,其长度通常等于已调度请求的数量。(`query start location` 是一个特例,它多一个元素。)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:60
|
||||
msgid "system level:"
|
||||
msgstr "系统级别:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:61
|
||||
msgid ""
|
||||
"**Token IDs table**: stores the token IDs (i.e. the inputs of a model) of"
|
||||
" each request. The shape of this table is `(max num request, max model "
|
||||
"len)`. Here, `max num request` is the maximum count of concurrent "
|
||||
"requests allowed in a forward batch and `max model len` is the maximum "
|
||||
"token count that can be handled at one request sequence in this model."
|
||||
msgstr "**Token IDs table**:存储每个请求的 token IDs(即模型的输入)。此表的形状为 `(max num request, max model len)`。其中,`max num request` 是前向批次中允许的最大并发请求数,`max model len` 是该模型中单个请求序列可以处理的最大 token 数量。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:62
|
||||
msgid ""
|
||||
"**Block table**: translates the logical address (within its sequence) of "
|
||||
"each block to its global physical address in the device's memory. The "
|
||||
"shape of this table is `(max num request, max model len / block size)`"
|
||||
msgstr "**Block table**:将每个块在其序列内的逻辑地址转换为其在设备内存中的全局物理地址。此表的形状为 `(max num request, max model len / block size)`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:64
|
||||
msgid ""
|
||||
"**Note**: Both of these two tables come from the `_update_states` method "
|
||||
"before **preparing inputs**. You can take a look if you need more "
|
||||
"inspiration."
|
||||
msgstr "**注意**:这两个表都来自 **准备输入** 之前的 `_update_states` 方法。如果需要更多启发,可以查看一下。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:66
|
||||
msgid "Tips"
|
||||
msgstr "提示"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:68
|
||||
msgid ""
|
||||
"Simply put, a `token ID` is an **integer** (usually `int32`), which "
|
||||
"represents a token. Example of `Token ID`:"
|
||||
msgstr "简而言之,一个 `token ID` 是一个**整数**(通常是 `int32`),它代表一个 token。`Token ID` 示例:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:88
|
||||
msgid "Go through details"
|
||||
msgstr "深入细节"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:90
|
||||
msgid "Assumptions:"
|
||||
msgstr "假设:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:92
|
||||
msgid "maximum number of tokens that can be scheduled at once: 10"
|
||||
msgstr "一次可调度的最大 token 数:10"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:93
|
||||
msgid "`block size`: 2"
|
||||
msgstr "`block size`:2"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:94
|
||||
msgid ""
|
||||
"Totally schedule 3 requests. Their prompt lengths are 3, 2, and 8 "
|
||||
"respectively."
|
||||
msgstr "总共调度 3 个请求。它们的提示长度分别为 3、2 和 8。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:95
|
||||
msgid ""
|
||||
"`max model length`: 12 (the maximum token count that can be handled at "
|
||||
"one request sequence in a model)."
|
||||
msgstr "`max model length`:12(模型中单个请求序列可以处理的最大 token 数量)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:97
|
||||
msgid ""
|
||||
"These assumptions are configured at the beginning when starting vLLM. "
|
||||
"They are not fixed, so you can manually set them."
|
||||
msgstr "这些假设是在启动 vLLM 时配置的。它们不是固定的,因此可以手动设置。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:99
|
||||
msgid "Step 1: All requests in the prefill phase"
|
||||
msgstr "步骤 1:所有请求均处于预填充阶段"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:101
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:213
|
||||
msgid "Obtain inputs"
|
||||
msgstr "获取输入"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:103
|
||||
#, python-brace-format
|
||||
msgid ""
|
||||
"As the maximum number of tokens that can be scheduled is 10, the "
|
||||
"scheduled tokens of each request can be represented as `{'0': 3, '1': 2, "
|
||||
"'2': 5}`. Note that `request_2` uses chunked prefill, leaving 3 prompt "
|
||||
"tokens unscheduled."
|
||||
msgstr "由于一次可调度的最大 token 数为 10,每个请求的已调度 token 可以表示为 `{'0': 3, '1': 2, '2': 5}`。注意 `request_2` 使用了分块预填充,留下了 3 个提示 token 未调度。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:105
|
||||
msgid "1. Get token positions"
|
||||
msgstr "1. 获取 token positions"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:107
|
||||
msgid ""
|
||||
"First, determine which request each token belongs to: tokens 0–2 are "
|
||||
"assigned to **request_0**, tokens 3–4 to **request_1**, and tokens 5–9 to"
|
||||
" **request_2**. To represent this mapping, we use `request indices`, for "
|
||||
"example, `request indices`: `[0, 0, 0, 1, 1, 2, 2, 2, 2, 2]`."
|
||||
msgstr "首先,确定每个 token 属于哪个请求:token 0–2 分配给 **request_0**,token 3–4 分配给 **request_1**,token 5–9 分配给 **request_2**。为了表示这种映射,我们使用 `request indices`,例如,`request indices`:`[0, 0, 0, 1, 1, 2, 2, 2, 2, 2]`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:109
|
||||
msgid ""
|
||||
"For each request, use **the number of computed tokens** + **the relative "
|
||||
"position of current scheduled tokens** (`request_0: [0 + 0, 0 + 1, 0 + "
|
||||
"2]`, `request_1: [0 + 0, 0 + 1]`, `request_2: [0 + 0, 0 + 1,..., 0 + 4]`)"
|
||||
" and then concatenate them together (`[0, 1, 2, 0, 1, 0, 1, 2, 3, 4]`)."
|
||||
msgstr "对于每个请求,使用 **已计算 token 的数量** + **当前调度 token 的相对位置**(`request_0: [0 + 0, 0 + 1, 0 + 2]`,`request_1: [0 + 0, 0 + 1]`,`request_2: [0 + 0, 0 + 1,..., 0 + 4]`),然后将它们连接在一起(`[0, 1, 2, 0, 1, 0, 1, 2, 3, 4]`)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:111
|
||||
msgid ""
|
||||
"Note: there is a more efficient way (using `request indices`) to create "
|
||||
"positions in actual code."
|
||||
msgstr "注意:在实际代码中,有一种更高效的方法(使用 `request indices`)来创建 positions。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:113
|
||||
msgid ""
|
||||
"Finally, `token positions` can be obtained as `[0, 1, 2, 0, 1, 0, 1, 2, "
|
||||
"3, 4]`. This variable is **token level**."
|
||||
msgstr "最后,`token positions` 可以获取为 `[0, 1, 2, 0, 1, 0, 1, 2, 3, 4]`。此变量是 **token 级别** 的。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:115
|
||||
msgid "2. Get token indices"
|
||||
msgstr "2. 获取 token indices"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:117
|
||||
msgid ""
|
||||
"The shape of the current **Token IDs table** is `(max num request, max "
|
||||
"model len)`."
|
||||
msgstr "当前 **Token IDs table** 的形状为 `(max num request, max model len)`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:119
|
||||
msgid ""
|
||||
"Why are these `T_3_5`, `T_3_6`, `T_3_7` in this table without being "
|
||||
"scheduled?"
|
||||
msgstr "为什么表中的 `T_3_5`、`T_3_6`、`T_3_7` 没有被调度?"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:121
|
||||
msgid ""
|
||||
"We fill all Token IDs in one request sequence to this table at once, but "
|
||||
"we only retrieve the tokens we scheduled this time. Then we retrieve the "
|
||||
"remaining Token IDs next time."
|
||||
msgstr "我们将一个请求序列中的所有 Token IDs 一次性填充到此表中,但我们只检索本次调度的 token。然后下次再检索剩余的 Token IDs。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:133
|
||||
msgid "Note that `T_x_x` is an `int32`."
|
||||
msgstr "注意 `T_x_x` 是一个 `int32`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:135
|
||||
msgid ""
|
||||
"Let's say `M = max model len`. Then we can use `token positions` together"
|
||||
" with `request indices` of each token to construct `token indices`."
|
||||
msgstr "假设 `M = max model len`。那么我们可以使用 `token positions` 以及每个 token 的 `request indices` 来构造 `token indices`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:137
|
||||
msgid ""
|
||||
"So `token indices` = `[0 + 0 * M, 1 + 0 * M, 2 + 0 * M, 0 + 1 * M, 1 + 1 "
|
||||
"* M, 0 + 2 * M, 1 + 2 * M, 2 + 2 * M, 3 + 2 * M, 4 + 2 * M]` = `[0, 1, 2,"
|
||||
" 12, 13, 24, 25, 26, 27, 28]`"
|
||||
msgstr "所以 `token indices` = `[0 + 0 * M, 1 + 0 * M, 2 + 0 * M, 0 + 1 * M, 1 + 1 * M, 0 + 2 * M, 1 + 2 * M, 2 + 2 * M, 3 + 2 * M, 4 + 2 * M]` = `[0, 1, 2, 12, 13, 24, 25, 26, 27, 28]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:139
|
||||
msgid "3. Retrieve the Token IDs"
|
||||
msgstr "3. 检索 Token IDs"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:141
|
||||
msgid ""
|
||||
"We use `token indices` to select out the corresponding `Input IDs` from "
|
||||
"the token table. The pseudocode is as follows:"
|
||||
msgstr "我们使用 `token indices` 从 token 表中选择出对应的 `Input IDs`。伪代码如下:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:147
|
||||
msgid "As mentioned before, we refer to these `Token IDs` as `Input IDs`."
|
||||
msgstr "如前所述,我们将这些 `Token IDs` 称为 `Input IDs`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:149
|
||||
msgid ""
|
||||
"`Input IDs` = `[T_0_0, T_0_1, T_0_2, T_1_0, T_1_1, T_2_0, T_2_1, T_3_2, "
|
||||
"T_3_3, T_3_4]`"
|
||||
msgstr "`Input IDs` = `[T_0_0, T_0_1, T_0_2, T_1_0, T_1_1, T_2_0, T_2_1, T_3_2, T_3_3, T_3_4]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:151
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:237
|
||||
msgid "Build inputs attention metadata"
|
||||
msgstr "构建输入注意力元数据"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:153
|
||||
msgid ""
|
||||
"In the current **Block Table**, we use the first block (i.e. block_0) to "
|
||||
"mark the unused block. The shape of the block is `(max num request, max "
|
||||
"model len / block size)`, where `max model len / block size = 12 / 2 = "
|
||||
"6`."
|
||||
msgstr ""
|
||||
"在当前的**块表**中,我们使用第一个块(即 block_0)来标记未使用的块。块的形状为 `(最大请求数, 最大模型长度 / 块大小)`,其中 `最大模型长度 / 块大小 = 12 / 2 = 6`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:165
|
||||
msgid "The KV cache block in the device memory is like:"
|
||||
msgstr "设备内存中的 KV 缓存块如下所示:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:171
|
||||
msgid ""
|
||||
"Let's say `K = max model len / block size = 6`, and we can get token "
|
||||
"`device block number`."
|
||||
msgstr "假设 `K = 最大模型长度 / 块大小 = 6`,我们可以得到令牌的`设备块编号`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:173
|
||||
msgid "The workflow of achieving slot mapping:"
|
||||
msgstr "实现槽映射的工作流程:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:175
|
||||
msgid "Get `block table indices` using `K`, `positions` and `request indices`."
|
||||
msgstr "使用 `K`、`positions` 和 `request indices` 获取`块表索引`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:177
|
||||
msgid ""
|
||||
"Purpose: For each token, it could be used to select `device block number`"
|
||||
" from `block table`."
|
||||
msgstr "目的:对于每个令牌,它可用于从`块表`中选择`设备块编号`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:179
|
||||
msgid "Get `device block number` using `block table indices`."
|
||||
msgstr "使用`块表索引`获取`设备块编号`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:181
|
||||
msgid ""
|
||||
"Purpose: `device block number` indicates which device block each token "
|
||||
"belongs to."
|
||||
msgstr "目的:`设备块编号`指示每个令牌属于哪个设备块。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:183
|
||||
msgid "Get `block offsets` using `positions` and `block size`."
|
||||
msgstr "使用 `positions` 和 `block size` 获取`块内偏移`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:185
|
||||
msgid ""
|
||||
"Purpose: `block offsets` indicates the offsets of each token within a "
|
||||
"block."
|
||||
msgstr "目的:`块内偏移`指示每个令牌在块内的偏移量。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:187
|
||||
msgid "construct `slot mapping` using `device block number` and `block offsets`."
|
||||
msgstr "使用`设备块编号`和`块内偏移`构建`槽映射`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:189
|
||||
msgid "Purpose: we can use `slot mapping` to store Token IDs into token slots."
|
||||
msgstr "目的:我们可以使用`槽映射`将令牌 ID 存储到令牌槽中。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:191
|
||||
msgid "Details:"
|
||||
msgstr "详细信息:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:193
|
||||
msgid ""
|
||||
"(**Token level**) Use a simple formula to calculate `block table "
|
||||
"indices`: `request indices * K + positions / block size`. So it equals "
|
||||
"`[0 * 6 + 0 / 2, 0 * 6 + 1 / 2, 0 * 6 + 2 / 2, 1 * 6 + 0 / 2, 1 * 6 + 1 /"
|
||||
" 2, 2 * 6 + 0 / 2, 2 * 6 + 1 / 2, 2 * 6 + 2 / 2, 2 * 6 + 3 / 2, 2 * 6 + 4"
|
||||
" / 2] = [0, 0, 1, 6, 6, 12, 12, 13, 13, 14]`. This could be used to "
|
||||
"select `device block number` from `block table`."
|
||||
msgstr ""
|
||||
"(**令牌级别**) 使用一个简单的公式计算`块表索引`:`request indices * K + positions / block size`。因此它等于 `[0 * 6 + 0 / 2, 0 * 6 + 1 / 2, 0 * 6 + 2 / 2, 1 * 6 + 0 / 2, 1 * 6 + 1 / 2, 2 * 6 + 0 / 2, 2 * 6 + 1 / 2, 2 * 6 + 2 / 2, 2 * 6 + 3 / 2, 2 * 6 + 4 / 2] = [0, 0, 1, 6, 6, 12, 12, 13, 13, 14]`。这可用于从`块表`中选择`设备块编号`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:194
|
||||
msgid ""
|
||||
"(**Token level**) Use `block table indices` to select out `device block "
|
||||
"number` for each scheduled token. The pseudocode is `block_numbers = "
|
||||
"block_table[block_table_indices]`. So `device block number=[1, 1, 2, 3, "
|
||||
"3, 4, 4, 5, 5, 6]`"
|
||||
msgstr ""
|
||||
"(**令牌级别**) 使用`块表索引`为每个已调度的令牌选择出`设备块编号`。伪代码为 `block_numbers = block_table[block_table_indices]`。因此 `设备块编号=[1, 1, 2, 3, 3, 4, 4, 5, 5, 6]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:195
|
||||
msgid ""
|
||||
"(**Token level**) `block offsets` could be computed by `block offsets = "
|
||||
"positions % block size = [0, 1, 0, 0, 1, 0, 1, 0, 1, 0]`."
|
||||
msgstr ""
|
||||
"(**令牌级别**) `块内偏移`可以通过 `block offsets = positions % block size = [0, 1, 0, 0, 1, 0, 1, 0, 1, 0]` 计算得出。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:196
|
||||
msgid ""
|
||||
"Finally, use `block offsets` and `device block number` to create `slot "
|
||||
"mapping`: `device block number * block size + block_offsets = [2, 3, 4, "
|
||||
"6, 7, 8, 9, 10, 11, 12]`"
|
||||
msgstr ""
|
||||
"最后,使用`块内偏移`和`设备块编号`创建`槽映射`:`设备块编号 * 块大小 + 块内偏移 = [2, 3, 4, 6, 7, 8, 9, 10, 11, 12]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:198
|
||||
msgid "(**Request level**) As we know the scheduled token count is `[3, 2, 5]`:"
|
||||
msgstr "(**请求级别**) 已知已调度的令牌数量为 `[3, 2, 5]`:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:200
|
||||
msgid ""
|
||||
"(**Request level**) Use prefix sum to calculate `query start location`: "
|
||||
"`[0, 3, 5, 10]`."
|
||||
msgstr "(**请求级别**) 使用前缀和计算`查询起始位置`:`[0, 3, 5, 10]`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:201
|
||||
msgid ""
|
||||
"(**Request level**) All tokens in step 1 are in the prefill stage, and "
|
||||
"the computed tokens count is 0; then `sequence length` = `[3, 2, 5]`."
|
||||
msgstr "(**请求级别**) 步骤 1 中的所有令牌都处于预填充阶段,已计算的令牌数量为 0;因此 `序列长度` = `[3, 2, 5]`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:202
|
||||
msgid ""
|
||||
"(**Request level**) As mentioned above, `number of computed tokens` are "
|
||||
"all 0s: `[0, 0, 0]`."
|
||||
msgstr "(**请求级别**) 如上所述,`已计算令牌数`均为 0:`[0, 0, 0]`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:203
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:272
|
||||
msgid "`number of requests`: `3`"
|
||||
msgstr "`请求数量`:`3`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:204
|
||||
msgid "(**Request level**) `number of tokens`: `[3, 2, 5]`"
|
||||
msgstr "(**请求级别**) `令牌数量`:`[3, 2, 5]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:205
|
||||
msgid "`max query len`: `5`"
|
||||
msgstr "`最大查询长度`:`5`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:206
|
||||
msgid "(**Token level**) `slot mapping`: `[2, 3, 4, 6, 7, 8, 9, 10, 11, 12]`"
|
||||
msgstr "(**令牌级别**) `槽映射`:`[2, 3, 4, 6, 7, 8, 9, 10, 11, 12]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:207
|
||||
msgid ""
|
||||
"`attention mask`: For all requests that initiate a prefill process, we "
|
||||
"simply create only one mask matrix for reuse across different requests. "
|
||||
"The shape of this mask matrix is `5 * 5`:"
|
||||
msgstr "`注意力掩码`:对于所有发起预填充过程的请求,我们仅创建一个掩码矩阵,以便在不同请求间复用。该掩码矩阵的形状为 `5 * 5`:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:209
|
||||
msgid "Step 2: Chunked prefill"
|
||||
msgstr "步骤 2:分块预填充"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:211
|
||||
msgid ""
|
||||
"In Step 2, we no longer provide explanations or perform calculations; "
|
||||
"instead, we directly present the final result."
|
||||
msgstr "在步骤 2 中,我们不再提供解释或进行计算;而是直接呈现最终结果。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:215
|
||||
#, python-brace-format
|
||||
msgid "Scheduled token of each request: `{'0': 1, '1': 1, '2': 3}`"
|
||||
msgstr "每个请求的已调度令牌:`{'0': 1, '1': 1, '2': 3}`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:217
|
||||
msgid "`request indices`: `[0, 1, 2, 2, 2]`"
|
||||
msgstr "`请求索引`:`[0, 1, 2, 2, 2]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:218
|
||||
msgid "`token positions`: `[3, 2, 5, 6, 7]`"
|
||||
msgstr "`令牌位置`:`[3, 2, 5, 6, 7]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:220
|
||||
msgid "Current **Token IDs table**:"
|
||||
msgstr "当前**令牌 ID 表**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:232
|
||||
msgid ""
|
||||
"**Note**: **T_0_3**, **T_1_2** are new Token IDs of **request_0** and "
|
||||
"**request_1** respectively. They are sampled from the output of the "
|
||||
"model."
|
||||
msgstr "**注意**:**T_0_3**、**T_1_2** 分别是 **request_0** 和 **request_1** 的新令牌 ID。它们是从模型输出中采样得到的。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:234
|
||||
msgid "`token indices`: `[3, 14, 29, 30, 31]`"
|
||||
msgstr "`令牌索引`:`[3, 14, 29, 30, 31]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:235
|
||||
msgid "`Input IDs`: `[T_0_3, T_1_2, T_3_5, T_3_6, T_3_7]`"
|
||||
msgstr "`输入 ID`:`[T_0_3, T_1_2, T_3_5, T_3_6, T_3_7]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:239
|
||||
msgid ""
|
||||
"We allocate the blocks `7` and `8` to `request_1` and `request_2` "
|
||||
"respectively, as they need more space in device to store KV cache "
|
||||
"following token generation or chunked prefill."
|
||||
msgstr "我们将块 `7` 和 `8` 分别分配给 `request_1` 和 `request_2`,因为它们在令牌生成或分块预填充后需要更多设备空间来存储 KV 缓存。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:241
|
||||
msgid "Current **Block Table**:"
|
||||
msgstr "当前**块表**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:253
|
||||
msgid "KV cache block in the device memory:"
|
||||
msgstr "设备内存中的 KV 缓存块:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:259
|
||||
msgid "(**Token level**) `block table indices`: `[1, 7, 14, 15, 15]`"
|
||||
msgstr "(**令牌级别**) `块表索引`:`[1, 7, 14, 15, 15]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:260
|
||||
msgid "(**Token level**) `device block number`: `[2, 7, 6, 8, 8]`"
|
||||
msgstr "(**令牌级别**) `设备块编号`:`[2, 7, 6, 8, 8]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:261
|
||||
msgid "(**Token level**) `block offsets`: `[1, 0, 1, 0, 1]`"
|
||||
msgstr "(**令牌级别**) `块内偏移`:`[1, 0, 1, 0, 1]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:262
|
||||
msgid "(**Token level**) `slot mapping`: `[5, 14, 13, 16, 17]`"
|
||||
msgstr "(**令牌级别**) `槽映射`:`[5, 14, 13, 16, 17]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:264
|
||||
msgid "Scheduled token count: `[1, 1, 3]`"
|
||||
msgstr "已调度令牌数量:`[1, 1, 3]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:266
|
||||
msgid "`query start location`: `[0, 1, 2, 5]`"
|
||||
msgstr "`查询起始位置`:`[0, 1, 2, 5]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:268
|
||||
msgid "`sequence length`: `[4, 3, 8]`"
|
||||
msgstr "`序列长度`:`[4, 3, 8]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:270
|
||||
msgid "`number of computed tokens`: `[3, 2, 5]`"
|
||||
msgstr "`已计算令牌数`:`[3, 2, 5]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:274
|
||||
msgid "`max query len`: `3`"
|
||||
msgstr "`最大查询长度`:`3`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:276
|
||||
msgid "`slot mapping`: `[5, 14, 13, 16, 17]`"
|
||||
msgstr "`槽映射`:`[5, 14, 13, 16, 17]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:278
|
||||
msgid "`attention mask`: `5 * 8`"
|
||||
msgstr "`注意力掩码`:`5 * 8`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:280
|
||||
msgid "Each token has a `1 * 8` vector, and there are 5 scheduled tokens."
|
||||
msgstr "每个令牌有一个 `1 * 8` 的向量,共有 5 个已调度的令牌。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:282
|
||||
msgid "At last"
|
||||
msgstr "最后"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:284
|
||||
msgid ""
|
||||
"If you understand step 1 and step 2, you will know all the following "
|
||||
"steps."
|
||||
msgstr "如果您理解了步骤 1 和步骤 2,您就会知道所有后续步骤。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/ModelRunner_prepare_inputs.md:286
|
||||
msgid ""
|
||||
"Hope this document helps you better understand how vLLM prepares inputs "
|
||||
"for model forwarding. If you have any good ideas, you are welcome to "
|
||||
"contribute to us."
|
||||
msgstr "希望本文档能帮助您更好地理解 vLLM 如何为模型前向传播准备输入。如果您有任何好的想法,欢迎向我们贡献。"
|
||||
@@ -0,0 +1,84 @@
|
||||
# SOME DESCRIPTIVE TITLE.
|
||||
# Copyright (C) 2025, vllm-ascend team
|
||||
# This file is distributed under the same license as the vllm-ascend
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend \n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/add_custom_aclnn_op.md:1
|
||||
msgid "Adding a custom aclnn operation"
|
||||
msgstr "添加自定义 aclnn 算子"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/add_custom_aclnn_op.md:3
|
||||
msgid ""
|
||||
"This document describes how to add a custom aclnn operation to vllm-"
|
||||
"ascend."
|
||||
msgstr "本文档描述了如何向 vllm-ascend 添加自定义 aclnn 算子。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/add_custom_aclnn_op.md:5
|
||||
msgid "How custom aclnn operation works in vllm-ascend?"
|
||||
msgstr "自定义 aclnn 算子在 vllm-ascend 中如何工作?"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/add_custom_aclnn_op.md:7
|
||||
msgid ""
|
||||
"Custom aclnn operations are built and installed into "
|
||||
"`vllm_ascend/cann_ops_custom` directory during the build process of vllm-"
|
||||
"ascend. Then the aclnn operators are bound to `torch.ops._C_ascend` "
|
||||
"module, enabling users to invoke them in vllm-ascend python code."
|
||||
msgstr "自定义 aclnn 算子在 vllm-ascend 的构建过程中被编译并安装到 `vllm_ascend/cann_ops_custom` 目录。然后,这些 aclnn 算子被绑定到 `torch.ops._C_ascend` 模块,使用户能够在 vllm-ascend 的 Python 代码中调用它们。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/add_custom_aclnn_op.md:9
|
||||
msgid "To enable custom operations, use the following code:"
|
||||
msgstr "要启用自定义算子,请使用以下代码:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/add_custom_aclnn_op.md:17
|
||||
msgid "How to add a custom aclnn operation?"
|
||||
msgstr "如何添加自定义 aclnn 算子?"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/add_custom_aclnn_op.md:19
|
||||
msgid "Create a new operation folder under `csrc` directory."
|
||||
msgstr "在 `csrc` 目录下创建一个新的算子文件夹。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/add_custom_aclnn_op.md:20
|
||||
msgid ""
|
||||
"Create `op_host` and `op_kernel` directories for host and kernel source "
|
||||
"code."
|
||||
msgstr "为宿主端和内核源代码创建 `op_host` 和 `op_kernel` 目录。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/add_custom_aclnn_op.md:21
|
||||
msgid ""
|
||||
"Add build options in `csrc/build_aclnn.sh` for supported SOC. Note that "
|
||||
"multiple ops should be separated with `;`, i.e. `CUSTOM_OPS=op1;op2;op3`."
|
||||
msgstr "在 `csrc/build_aclnn.sh` 中为支持的 SOC 添加构建选项。注意多个算子应用 `;` 分隔,例如 `CUSTOM_OPS=op1;op2;op3`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/add_custom_aclnn_op.md:22
|
||||
msgid ""
|
||||
"Bind aclnn operators to torch.ops._C_ascend module in "
|
||||
"`csrc/torch_binding.cpp`."
|
||||
msgstr "在 `csrc/torch_binding.cpp` 中将 aclnn 算子绑定到 torch.ops._C_ascend 模块。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/add_custom_aclnn_op.md:23
|
||||
msgid ""
|
||||
"Write a meta implementation in `csrc/torch_binding_meta.cpp` for the op "
|
||||
"to be captured into the aclgraph."
|
||||
msgstr "在 `csrc/torch_binding_meta.cpp` 中为算子编写一个元实现,以便其能被捕获到 aclgraph 中。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/add_custom_aclnn_op.md:25
|
||||
msgid ""
|
||||
"After a successful build of vllm-ascend, the custom aclnn operation can "
|
||||
"be invoked in python code."
|
||||
msgstr "成功构建 vllm-ascend 后,即可在 Python 代码中调用自定义的 aclnn 算子。"
|
||||
@@ -0,0 +1,391 @@
|
||||
# SOME DESCRIPTIVE TITLE.
|
||||
# Copyright (C) 2025, vllm-ascend team
|
||||
# This file is distributed under the same license as the vllm-ascend
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend \n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:1
|
||||
msgid "Context Parallel (CP)"
|
||||
msgstr "上下文并行 (CP)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:3
|
||||
msgid ""
|
||||
"TL;DR PCP accelerates prefill via sequence splitting. DCP eliminates KV "
|
||||
"cache redundancy."
|
||||
msgstr "TL;DR PCP 通过序列分割加速预填充。DCP 消除 KV 缓存冗余。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:5
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:5
|
||||
msgid "ContextParallel"
|
||||
msgstr "ContextParallel"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:7
|
||||
msgid ""
|
||||
"For the main discussions during the development process, please refer to "
|
||||
"the [RFC](https://github.com/vllm-project/vllm/issues/25749) and the "
|
||||
"relevant links referenced by or referencing this RFC."
|
||||
msgstr "关于开发过程中的主要讨论,请参阅 [RFC](https://github.com/vllm-project/vllm/issues/25749) 以及该 RFC 引用或被引用的相关链接。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:9
|
||||
msgid "What is CP?"
|
||||
msgstr "什么是 CP?"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:11
|
||||
msgid ""
|
||||
"**Context Parallel (CP)** is a strategy for parallelizing computation "
|
||||
"along the sequence dimension across multiple devices."
|
||||
msgstr "**上下文并行 (CP)** 是一种沿序列维度在多个设备间并行计算的策略。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:13
|
||||
msgid ""
|
||||
"**Prefill Context Parallel (PCP)** expands the world size of devices and "
|
||||
"uses dedicated communication domains. Its primary goal is to partition "
|
||||
"the sequence dimension during the prefill phase, enabling different "
|
||||
"devices to compute distinct chunks of the sequence simultaneously. The KV"
|
||||
" cache is sharded along the sequence dimension across devices. This "
|
||||
"approach impacts the computational logic of both the Prefill and Decode "
|
||||
"stages to varying degrees."
|
||||
msgstr "**预填充上下文并行 (PCP)** 扩展了设备的世界大小并使用专用的通信域。其主要目标是在预填充阶段对序列维度进行分区,使不同设备能同时计算序列的不同分块。KV 缓存沿序列维度跨设备分片。此方法在不同程度上影响了预填充和解码阶段的计算逻辑。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:18
|
||||
msgid ""
|
||||
"**Decode Context Parallel (DCP)** reuses the communication domain of "
|
||||
"Tensor Parallelism (TP) and does not require additional devices. Its main"
|
||||
" objective is to eliminate duplicated storage of the KV cache by sharding"
|
||||
" it along the sequence dimension across devices within the TP domain that"
|
||||
" would otherwise hold redundant copies. DCP primarily influences the "
|
||||
"Decode logic, as well as the logic for chunked prefill and cached "
|
||||
"prefill."
|
||||
msgstr "**解码上下文并行 (DCP)** 复用张量并行 (TP) 的通信域,且不需要额外的设备。其主要目标是通过在 TP 域内沿序列维度对 KV 缓存进行分片,消除原本会存储冗余副本的设备间的重复存储。DCP 主要影响解码逻辑,以及分块预填充和缓存预填充的逻辑。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:22
|
||||
msgid "How to Use CP?"
|
||||
msgstr "如何使用 CP?"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:24
|
||||
msgid ""
|
||||
"Please refer to the [context parallel user "
|
||||
"guide](../../user_guide/feature_guide/context_parallel.md) for detailed "
|
||||
"information."
|
||||
msgstr "详细信息请参阅 [上下文并行用户指南](../../user_guide/feature_guide/context_parallel.md)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:26
|
||||
msgid "How It Works?"
|
||||
msgstr "工作原理"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:28
|
||||
msgid "Device Distribution"
|
||||
msgstr "设备分布"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:30
|
||||
msgid ""
|
||||
"We introduce new communication domains for PCP and reuse TP for DCP, and "
|
||||
"this is the new layout of devices for PCP2, DCP2, and TP4. "
|
||||
""
|
||||
msgstr "我们为 PCP 引入了新的通信域,并为 DCP 复用了 TP 的通信域,这是 PCP2、DCP2 和 TP4 的新设备布局。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:30
|
||||
msgid "device_world"
|
||||
msgstr "device_world"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:33
|
||||
msgid "Block Table"
|
||||
msgstr "块表"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:35
|
||||
msgid ""
|
||||
"CP performs sequence sharding on the KV cache storage. To facilitate "
|
||||
"efficient storage and access, tokens are stored in an interleaved manner "
|
||||
"across devices, with the interleaving granularity determined by "
|
||||
"`cp_kv_cache_interleave_size`, whose default value is "
|
||||
"`cp_kv_cache_interleave_size=1`, a.k.a. 'token interleave'."
|
||||
msgstr "CP 对 KV 缓存存储执行序列分片。为了便于高效存储和访问,令牌以交错方式跨设备存储,交错粒度由 `cp_kv_cache_interleave_size` 决定,其默认值为 `cp_kv_cache_interleave_size=1`,也称为“令牌交错”。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:37
|
||||
msgid ""
|
||||
"Given that PCP and DCP behave similarly for KV cache sharding, we refer "
|
||||
"to them collectively as CP. Specifically, `cp_size = pcp_size * "
|
||||
"dcp_size`, and `cp_rank = pcp_rank * dcp_size + dcp_rank`."
|
||||
msgstr "鉴于 PCP 和 DCP 在 KV 缓存分片方面的行为相似,我们将它们统称为 CP。具体来说,`cp_size = pcp_size * dcp_size`,且 `cp_rank = pcp_rank * dcp_size + dcp_rank`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:39
|
||||
msgid ""
|
||||
"As illustrated, a virtual block is defined in the block table, where "
|
||||
"blocks within the same CP device group form a virtual block. The virtual "
|
||||
"block size is `virtual_block_size = block_size * cp_size`."
|
||||
msgstr "如图所示,块表中定义了一个虚拟块,同一 CP 设备组内的块构成一个虚拟块。虚拟块大小为 `virtual_block_size = block_size * cp_size`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:41
|
||||
#, python-format
|
||||
msgid ""
|
||||
"For any token `x`, referencing the following figure, its (virtual) block "
|
||||
"index is `x // virtual_block_size`, and the offset within the virtual "
|
||||
"block is `offset_within_virtual_block = x % virtual_block_size`. The "
|
||||
"local block index is `local_block_index = offset_within_virtual_block // "
|
||||
"cp_kv_cache_interleave_size`, and the device number is `target_rank = "
|
||||
"local_block_index % cp_size`. The offset within the local block is "
|
||||
"`(local_block_index // cp_size) * cp_kv_cache_interleave_size + "
|
||||
"offset_within_virtual_block % cp_kv_cache_interleave_size`."
|
||||
msgstr "对于任意令牌 `x`,参考下图,其(虚拟)块索引为 `x // virtual_block_size`,在虚拟块内的偏移量为 `offset_within_virtual_block = x % virtual_block_size`。本地块索引为 `local_block_index = offset_within_virtual_block // cp_kv_cache_interleave_size`,设备号为 `target_rank = local_block_index % cp_size`。在本地块内的偏移量为 `(local_block_index // cp_size) * cp_kv_cache_interleave_size + offset_within_virtual_block % cp_kv_cache_interleave_size`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:45
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:45
|
||||
msgid "BlockTable"
|
||||
msgstr "BlockTable"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:47
|
||||
msgid ""
|
||||
"Based on the logic above, the `slot_mapping` calculation process is "
|
||||
"adjusted, and the `slot_mapping` values on each device are modified to "
|
||||
"ensure the KV cache is sharded along the sequence dimension and stored "
|
||||
"across different devices as expected."
|
||||
msgstr "基于上述逻辑,调整了 `slot_mapping` 的计算过程,并修改了每个设备上的 `slot_mapping` 值,以确保 KV 缓存沿序列维度分片并按预期存储在不同设备上。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:49
|
||||
#, python-format
|
||||
msgid ""
|
||||
"The current implementation requires that `block_size % "
|
||||
"cp_kv_cache_interleave_size == 0`."
|
||||
msgstr "当前实现要求 `block_size % cp_kv_cache_interleave_size == 0`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:51
|
||||
msgid "Decode Context Parallel (DCP)"
|
||||
msgstr "解码上下文并行 (DCP)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:53
|
||||
msgid ""
|
||||
"As mentioned above, the primary function of DCP is to shard the KV cache "
|
||||
"along the sequence dimension for storage. Its impact lies in the logic of"
|
||||
" the decode and chunked prefill phases."
|
||||
msgstr "如上所述,DCP 的主要功能是沿序列维度对 KV 缓存进行分片存储。其影响在于解码和分块预填充阶段的逻辑。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:55
|
||||
msgid ""
|
||||
"**Prefill Phase:** As illustrated, during the Chunked Prefill "
|
||||
"computation, two distinct logic implementations are employed for MLA and "
|
||||
"GQA backends."
|
||||
msgstr "**预填充阶段:** 如图所示,在分块预填充计算期间,MLA 和 GQA 后端采用了两种不同的逻辑实现。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:58
|
||||
msgid ""
|
||||
"In the **MLA backend**, a Context KV Cache `all_gather` operation is "
|
||||
"performed to aggregate the full KV values. These are then used for "
|
||||
"attention computation with the Q values of the current chunk. Note that "
|
||||
"in multi-request scenarios, the directly gathered KV results are "
|
||||
"interleaved across requests. The `reorg_kvcache` function is used to "
|
||||
"reorganize the KV cache, ensuring that the KV cache of the same request "
|
||||
"is stored contiguously."
|
||||
msgstr "在 **MLA 后端** 中,执行上下文 KV 缓存 `all_gather` 操作以聚合完整的 KV 值。然后这些值与当前分块的 Q 值一起用于注意力计算。请注意,在多请求场景中,直接收集的 KV 结果在请求间是交错的。使用 `reorg_kvcache` 函数来重新组织 KV 缓存,确保同一请求的 KV 缓存被连续存储。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:63
|
||||
msgid ""
|
||||
"In the **GQA backend**, an `all_gather` is performed along the head "
|
||||
"dimension for Q. This is because DCP overlaps with the TP communication "
|
||||
"domain, and the Q heads within a DCP group differ. However, they need to "
|
||||
"exchange results with the locally computed KV cache for online Softmax "
|
||||
"updates. To ensure correctness during result updates, the Q values are "
|
||||
"synchronized across the DCP group via head-dimension `all_gather`. During"
|
||||
" the result update process, `cp_lse_ag_out_rs` is invoked to aggregate "
|
||||
"`attn_output` and `attn_lse`, update the results, and perform a reduce-"
|
||||
"scatter operation on the outputs. Alternatively, we can use an all-to-all"
|
||||
" communication to exchange the output and LSE results, followed by direct"
|
||||
" local updates. This approach aligns with the logic adapted for PCP "
|
||||
"compatibility."
|
||||
msgstr "在 **GQA 后端** 中,沿头维度对 Q 执行 `all_gather`。这是因为 DCP 与 TP 通信域重叠,且 DCP 组内的 Q 头不同。然而,它们需要与本地计算的 KV 缓存交换结果以进行在线 Softmax 更新。为确保结果更新过程中的正确性,Q 值通过头维度的 `all_gather` 在 DCP 组内同步。在结果更新过程中,调用 `cp_lse_ag_out_rs` 来聚合 `attn_output` 和 `attn_lse`,更新结果,并对输出执行 reduce-scatter 操作。或者,我们可以使用 all-to-all 通信来交换输出和 LSE 结果,然后直接进行本地更新。这种方法与为 PCP 兼容性而调整的逻辑一致。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:70
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:70
|
||||
msgid "DCP-Prefill"
|
||||
msgstr "DCP-Prefill"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:72
|
||||
msgid ""
|
||||
"**Decode Phase:** The logic during the decode phase is consistent with "
|
||||
"that of GQA's chunked prefill: an all-gather operation is first performed"
|
||||
" along the Q head dimension to ensure consistency within the DCP group. "
|
||||
"After computing the results with the local KV cache, the results are "
|
||||
"updated via the `cp_lse_ag_out_rs` function."
|
||||
msgstr "**解码阶段:** 解码阶段的逻辑与 GQA 的分块预填充一致:首先沿 Q 头维度执行 all-gather 操作以确保 DCP 组内的一致性。使用本地 KV 缓存计算结果后,通过 `cp_lse_ag_out_rs` 函数更新结果。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:76
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:76
|
||||
msgid "DCP-Decode"
|
||||
msgstr "DCP-Decode"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:78
|
||||
msgid "Prefill Context Parallel (PCP)"
|
||||
msgstr "预填充上下文并行 (PCP)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:80
|
||||
msgid "**Tokens Partition in Head-Tail Style**"
|
||||
msgstr "**头尾式令牌分区**"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:82
|
||||
msgid ""
|
||||
"PCP requires splitting the input sequence and ensuring balanced "
|
||||
"computational load across devices during the prefill phase. We employ a "
|
||||
"head-tail style for splitting and concatenation: specifically, the "
|
||||
"sequence is first padded to a length of `2*pcp_size`, then divided into "
|
||||
"`2*pcp_size` equal parts. The first part is merged with the last part, "
|
||||
"the second part with the second last part, and so on, thereby assigning "
|
||||
"computationally balanced chunks to each device. Additionally, since "
|
||||
"allgather aggregation of KV or Q results in interleaved chunks from "
|
||||
"different requests, we compute `pcp_allgather_restore_idx` to quickly "
|
||||
"restore the original order."
|
||||
msgstr "PCP 需要在预填充阶段分割输入序列并确保跨设备的计算负载均衡。我们采用头尾式进行分割和连接:具体来说,首先将序列填充到长度为 `2*pcp_size`,然后分成 `2*pcp_size` 个相等的部分。第一部分与最后一部分合并,第二部分与倒数第二部分合并,依此类推,从而为每个设备分配计算上均衡的分块。此外,由于 KV 或 Q 的 allgather 聚合会导致来自不同请求的交错分块,我们计算 `pcp_allgather_restore_idx` 以快速恢复原始顺序。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:87
|
||||
msgid "These logics are implemented in the function `_update_tokens_for_pcp`."
|
||||
msgstr "这些逻辑在函数 `_update_tokens_for_pcp` 中实现。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:89
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:89
|
||||
msgid "PCP-Partition"
|
||||
msgstr "PCP-Partition"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:91
|
||||
msgid "**Prefill Phase:**"
|
||||
msgstr "**预填充阶段:**"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:93
|
||||
msgid ""
|
||||
"During the Prefill phase (excluding chunked prefill), we employ an all-"
|
||||
"gather KV approach to address the issue of incomplete sequences on "
|
||||
"individual GPUs. It is important to note that we only aggregate the KV "
|
||||
"values for the current layer at a time, and these are discarded "
|
||||
"immediately after use, avoiding excessive peak memory usage. This method "
|
||||
"can also be directly applied to KV cache storage (since the KV cache "
|
||||
"partitioning method differs from PCP sequence partitioning, it is "
|
||||
"inevitable that each GPU requires a complete copy of the KV values). All "
|
||||
"attention backends maintain consistency in this logic."
|
||||
msgstr "在预填充阶段(不包括分块预填充),我们采用 all-gather KV 的方法来解决单个 GPU 上序列不完整的问题。需要注意的是,我们一次只聚合当前层的 KV 值,并且在使用后立即丢弃,以避免过高的峰值内存使用。此方法也可直接应用于 KV 缓存存储(由于 KV 缓存的分区方法与 PCP 序列分区不同,每个 GPU 都需要一份完整的 KV 值副本是不可避免的)。所有注意力后端在此逻辑上保持一致。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:98
|
||||
msgid ""
|
||||
"Note: While a Ring Attention approach could also facilitate information "
|
||||
"exchange with lower peak memory and enable computation-communication "
|
||||
"overlap, we prioritized the all-gather KV implementation after evaluating"
|
||||
" that the development complexity was high and the benefits of overlap "
|
||||
"were limited."
|
||||
msgstr "注意:虽然环形注意力方法也能以更低的峰值内存促进信息交换并实现计算-通信重叠,但在评估了开发复杂度高且重叠收益有限后,我们优先实现了 all-gather KV 方案。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:100
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:100
|
||||
msgid "PCP-Prefill"
|
||||
msgstr "PCP-Prefill"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:102
|
||||
msgid "**Decode Phase:**"
|
||||
msgstr "**解码阶段:**"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:104
|
||||
msgid ""
|
||||
"During the decode phase, we only need to add an allgather within the PCP "
|
||||
"group after the DCP all-to-all communication exchanges the output and "
|
||||
"LSE, before proceeding with the output update."
|
||||
msgstr "在解码阶段,我们只需要在 DCP all-to-all 通信交换输出和 LSE 之后,于 PCP 组内添加一个 allgather,然后再进行输出更新。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:106
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:106
|
||||
msgid "PCP-Decode"
|
||||
msgstr "PCP-Decode"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:108
|
||||
msgid "**Chunked Prefill:**"
|
||||
msgstr "**分块预填充:**"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:110
|
||||
msgid ""
|
||||
"Currently, there are three viable approaches for Chunked Prefill "
|
||||
"compatibility: **AllGatherQ**, **AllGatherKV**, and **Ring-Attn**. Since "
|
||||
"PCP performs sequence sharding on both the query sequence and the KV "
|
||||
"cache, we need to ensure that one side has complete information or employ"
|
||||
" a method like Ring-Attn to perform computations sequentially. The "
|
||||
"advantages and disadvantages of Ring-Attn will not be elaborated here."
|
||||
msgstr "目前,有三种可行的分块预填充兼容性方法:**AllGatherQ**、**AllGatherKV** 和 **Ring-Attn**。由于 PCP 对查询序列和 KV 缓存都执行序列分片,我们需要确保其中一方拥有完整信息,或者采用类似 Ring-Attn 的方法顺序执行计算。Ring-Attn 的优缺点在此不赘述。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:114
|
||||
msgid ""
|
||||
"We have implemented the **AllGatherQ** approach in the GQA attention "
|
||||
"backend and the **AllGatherKV** approach in the MLA attention backend. "
|
||||
"The workflow after **AllGatherQ** is identical to the decode phase, while"
|
||||
" the workflow after **AllGatherKV** is the same as the standard prefill "
|
||||
"phase. For details, please refer to the diagram below; specific steps "
|
||||
"will not be repeated."
|
||||
msgstr "我们已在 GQA 注意力后端实现了 **AllGatherQ** 方法,并在 MLA 注意力后端实现了 **AllGatherKV** 方法。**AllGatherQ** 之后的工作流与解码阶段相同,而 **AllGatherKV** 之后的工作流与标准预填充阶段相同。详情请参考下图;具体步骤不再赘述。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:118
|
||||
msgid ""
|
||||
"One important note: **AllGatherKV** may lead to significant peak memory "
|
||||
"usage when the context length becomes excessively long. To mitigate this,"
|
||||
" we adopt a segmented processing strategy. By predefining the maximum "
|
||||
"amount of KV cache processed per round, we sequentially complete the "
|
||||
"attention computation and online softmax updates for each segment."
|
||||
msgstr ""
|
||||
"一个重要注意事项:当上下文长度变得过长时,**AllGatherKV** 可能导致显著的峰值内存使用。为了缓解这个问题,我们采用了分段处理策略。通过预定义每轮处理的 KV 缓存最大量,我们依次完成每个分段的注意力计算和在线 softmax 更新。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:122
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:122
|
||||
msgid "PCP-ChunkedPrefill"
|
||||
msgstr "PCP-ChunkedPrefill"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:124
|
||||
msgid "Related Files"
|
||||
msgstr "相关文件"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:126
|
||||
msgid "slot_mapping computation: `vllm_ascend/worker/block_table.py`"
|
||||
msgstr "slot_mapping 计算:`vllm_ascend/worker/block_table.py`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:127
|
||||
msgid ""
|
||||
"sequences splitting and metadata prepare: "
|
||||
"`vllm_ascend/worker/model_runner_v1.py`"
|
||||
msgstr "序列拆分与元数据准备:`vllm_ascend/worker/model_runner_v1.py`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:128
|
||||
msgid "GQA backend: `vllm_ascend/attention/attention_cp.py`"
|
||||
msgstr "GQA 后端:`vllm_ascend/attention/attention_cp.py`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/context_parallel.md:129
|
||||
msgid "MLA backend: `vllm_ascend/attention/mla_cp.py`"
|
||||
msgstr "MLA 后端:`vllm_ascend/attention/mla_cp.py`"
|
||||
@@ -0,0 +1,814 @@
|
||||
# SOME DESCRIPTIVE TITLE.
|
||||
# Copyright (C) 2025, vllm-ascend team
|
||||
# This file is distributed under the same license as the vllm-ascend
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend \n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:1
|
||||
msgid "CPU Binding"
|
||||
msgstr "CPU 绑定"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:3
|
||||
msgid "Overview"
|
||||
msgstr "概述"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:5
|
||||
msgid ""
|
||||
"CPU binding pins vLLM Ascend worker processes and key threads to specific"
|
||||
" CPU cores to reduce CPU–NPU cross‑NUMA traffic and stabilize latency "
|
||||
"under multi‑process workloads. It is designed for ARM servers running "
|
||||
"Ascend NPUs and is automatically executed during worker initialization "
|
||||
"when enabled."
|
||||
msgstr ""
|
||||
"CPU 绑定将 vLLM Ascend 工作进程和关键线程固定到特定的 CPU 核心,以减少 CPU-"
|
||||
"NPU 跨 NUMA 流量,并在多进程工作负载下稳定延迟。它专为运行 Ascend NPU 的 ARM "
|
||||
"服务器设计,启用后会在工作进程初始化期间自动执行。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:7
|
||||
msgid "Background"
|
||||
msgstr "背景"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:9
|
||||
msgid ""
|
||||
"On multi‑socket ARM systems, the OS scheduler may place vLLM threads on "
|
||||
"CPUs far from the local NPU, causing NUMA cross‑traffic and jitter. CPU "
|
||||
"binding enforces a deterministic CPU placement strategy and optionally "
|
||||
"binds NPU IRQs to the same CPU pool. This is distinct from other "
|
||||
"performance features (e.g., graph mode or dynamic batch) because it is "
|
||||
"purely a host‑side affinity policy and does not change model execution "
|
||||
"logic."
|
||||
msgstr ""
|
||||
"在多插槽 ARM 系统上,操作系统调度器可能会将 vLLM 线程放置在远离本地 NPU 的 "
|
||||
"CPU 上,从而导致 NUMA 跨域流量和延迟抖动。CPU 绑定强制执行一种确定性的 CPU "
|
||||
"放置策略,并可选地将 NPU IRQ 绑定到同一个 CPU 池。这与其他性能特性(如图模式"
|
||||
"或动态批处理)不同,因为它纯粹是主机端的亲和性策略,不改变模型执行逻辑。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:11
|
||||
msgid "Design & How it works"
|
||||
msgstr "设计与工作原理"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:13
|
||||
msgid "Key concepts"
|
||||
msgstr "关键概念"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:15
|
||||
msgid ""
|
||||
"**Allowed CPU list**: The cpuset from /proc/self/status "
|
||||
"(Cpus_allowed_list). All allocations are constrained to this list."
|
||||
msgstr ""
|
||||
"**允许的 CPU 列表**:来自 /proc/self/status (Cpus_allowed_list) 的 cpuset。"
|
||||
"所有分配都受限于此列表。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:16
|
||||
msgid ""
|
||||
"**Running NPU list**: Logical NPU IDs extracted from npu‑smi process "
|
||||
"listing, optionally filtered by ASCEND_RT_VISIBLE_DEVICES."
|
||||
msgstr ""
|
||||
"**运行中的 NPU 列表**:从 npu-smi 进程列表中提取的逻辑 NPU ID,可选地由 "
|
||||
"ASCEND_RT_VISIBLE_DEVICES 过滤。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:17
|
||||
msgid ""
|
||||
"**CPU pool per NPU**: The CPU list assigned to each logical NPU ID based "
|
||||
"on the binding mode."
|
||||
msgstr ""
|
||||
"**每个 NPU 的 CPU 池**:根据绑定模式分配给每个逻辑 NPU ID 的 CPU 列表。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:18
|
||||
msgid "**Binding modes & Device behavior**:"
|
||||
msgstr "**绑定模式与设备行为**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "Device type"
|
||||
msgstr "设备类型"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "Default mode"
|
||||
msgstr "默认模式"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "Description"
|
||||
msgstr "描述"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "A3 (No Affinity)"
|
||||
msgstr "A3 (无亲和性)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`global_slice`"
|
||||
msgstr "`global_slice`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid ""
|
||||
"Splits the allowed CPU list evenly based on the **total number of global "
|
||||
"logical NPUs**, ensuring each NPU is assigned a contiguous segment of CPU"
|
||||
" cores. This prevents CPU core overlap across multiple process groups."
|
||||
msgstr ""
|
||||
"根据**全局逻辑 NPU 总数**均匀分割允许的 CPU 列表,确保每个 NPU 被分配一个连"
|
||||
"续的 CPU 核心段。这可以防止多个进程组之间的 CPU 核心重叠。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "A2 / 310P / Others"
|
||||
msgstr "A2 / 310P / 其他"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`topo_affinity`"
|
||||
msgstr "`topo_affinity`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid ""
|
||||
"Allocates CPUs based on NPU topology affinity (`npu‑smi info -t topo`). "
|
||||
"If multiple NPUs are assigned to a single NUMA node (which may cause "
|
||||
"bandwidth contention), the CPU allocation extends to adjacent NUMA nodes."
|
||||
msgstr ""
|
||||
"基于 NPU 拓扑亲和性 (`npu-smi info -t topo`) 分配 CPU。如果多个 NPU 被分配"
|
||||
"到单个 NUMA 节点(可能导致带宽争用),则 CPU 分配会扩展到相邻的 NUMA 节点。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:25
|
||||
msgid "**Default**: enabled (enable_cpu_binding = true)."
|
||||
msgstr "**默认**:启用 (enable_cpu_binding = true)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:26
|
||||
msgid "**Fallback**: If NPU topo affinity is unavailable, global_slice is used."
|
||||
msgstr "**回退**:如果 NPU 拓扑亲和性不可用,则使用 global_slice。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:27
|
||||
msgid ""
|
||||
"**Failure handling**: Any exception in binding is logged as a warning and"
|
||||
" **binding is skipped for that rank**."
|
||||
msgstr ""
|
||||
"**故障处理**:绑定过程中的任何异常都会记录为警告,并且**跳过该等级的绑定**。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:29
|
||||
msgid "Execution flow (simplified)"
|
||||
msgstr "执行流程(简化版)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:31
|
||||
msgid ""
|
||||
"**Feature entry**: worker initialization calls `bind_cpus(local_rank)` "
|
||||
"when `enable_cpu_binding` is true."
|
||||
msgstr ""
|
||||
"**功能入口**:当 `enable_cpu_binding` 为 true 时,工作进程初始化会调用 "
|
||||
"`bind_cpus(local_rank)`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:32
|
||||
msgid ""
|
||||
"**CPU architecture gate**: If the CPU is not ARM, binding is skipped with"
|
||||
" a log."
|
||||
msgstr "**CPU 架构门控**:如果 CPU 不是 ARM,则记录日志并跳过绑定。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:33
|
||||
msgid "**Collect device info**:"
|
||||
msgstr "**收集设备信息**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:34
|
||||
msgid "Map logical NPU IDs from `npu‑smi info -m`."
|
||||
msgstr "从 `npu-smi info -m` 映射逻辑 NPU ID。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:35
|
||||
msgid "Detect running NPU IDs from npu‑smi info process table."
|
||||
msgstr "从 npu-smi info 进程表中检测运行中的 NPU ID。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:36
|
||||
msgid "Read cpuset from /proc/self/status."
|
||||
msgstr "从 /proc/self/status 读取 cpuset。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:37
|
||||
msgid "Read topo affinity from `npu‑smi info -t topo`."
|
||||
msgstr "从 `npu-smi info -t topo` 读取拓扑亲和性。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:38
|
||||
msgid "**Build CPU pools**:"
|
||||
msgstr "**构建 CPU 池**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:39
|
||||
msgid "Use **global_slice** for A3 devices; **topo_affinity** for A2 and 310P."
|
||||
msgstr "对 A3 设备使用 **global_slice**;对 A2 和 310P 使用 **topo_affinity**。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:40
|
||||
msgid "If topo affinity is missing, fall back to global_slice."
|
||||
msgstr "如果缺少拓扑亲和性,则回退到 global_slice。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:41
|
||||
msgid "Ensure each NPU has at least 5 CPUs."
|
||||
msgstr "确保每个 NPU 至少有 5 个 CPU。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:42
|
||||
msgid "**Allocate per‑role CPUs**:"
|
||||
msgstr "**分配按角色划分的 CPU**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:43
|
||||
msgid "Reserve the first two CPUs for IRQ binding."
|
||||
msgstr "保留前两个 CPU 用于 IRQ 绑定。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:44
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:62
|
||||
msgid "`main`: pool[2:-2]"
|
||||
msgstr "`main`: pool[2:-2]"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:45
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:63
|
||||
msgid "`acl`: pool[-2]"
|
||||
msgstr "`acl`: pool[-2]"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:46
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:64
|
||||
msgid "`release`: pool[-1]"
|
||||
msgstr "`release`: pool[-1]"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:47
|
||||
msgid "**Bind threads**:"
|
||||
msgstr "**绑定线程**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:48
|
||||
msgid "Main process is pinned to `main` CPUs."
|
||||
msgstr "主进程被固定到 `main` CPU。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:49
|
||||
msgid "ACL threads (named with acl_thread) are pinned to `acl` CPU."
|
||||
msgstr "ACL 线程(以 acl_thread 命名)被固定到 `acl` CPU。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:50
|
||||
msgid "Release threads (named with release_thread) are pinned to `release` CPU."
|
||||
msgstr "释放线程(以 release_thread 命名)被固定到 `release` CPU。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:51
|
||||
msgid "**Bind NPU IRQs (optional)**:"
|
||||
msgstr "**绑定 NPU IRQ(可选)**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:52
|
||||
msgid ""
|
||||
"If /proc/irq is writable, bind SQ/CQ IRQs to the first two CPUs in the "
|
||||
"pool."
|
||||
msgstr "如果 /proc/irq 可写,则将 SQ/CQ IRQ 绑定到池中的前两个 CPU。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:53
|
||||
msgid "irqbalance may be stopped to prevent overrides."
|
||||
msgstr "可能会停止 irqbalance 以防止覆盖。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:54
|
||||
msgid "**Memory binding (optional)**:"
|
||||
msgstr "**内存绑定(可选)**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:55
|
||||
msgid ""
|
||||
"If migratepages is available, memory for ACL threads is migrated to the "
|
||||
"NPU’s NUMA node."
|
||||
msgstr "如果 migratepages 可用,则将 ACL 线程的内存迁移到 NPU 的 NUMA 节点。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:57
|
||||
msgid "Allocation plan examples"
|
||||
msgstr "分配方案示例"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:59
|
||||
msgid ""
|
||||
"The allocation plan is derived directly from the CPU pool per NPU and "
|
||||
"then split into roles:"
|
||||
msgstr "分配方案直接来源于每个 NPU 的 CPU 池,然后按角色划分:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:61
|
||||
msgid "IRQ CPUs: pool[0], pool[1]"
|
||||
msgstr "IRQ CPU: pool[0], pool[1]"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:66
|
||||
msgid "Below are concrete examples that reflect the actual code paths."
|
||||
msgstr "以下是反映实际代码路径的具体示例。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:68
|
||||
msgid "Example 1: A3 inference server with 640 CPUs and 16 NPUs"
|
||||
msgstr "示例 1:具有 640 个 CPU 和 16 个 NPU 的 A3 推理服务器"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:70
|
||||
msgid "allowed_cpus = [0..639] (640 CPUs)"
|
||||
msgstr "allowed_cpus = [0..639] (640 个 CPU)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:71
|
||||
msgid "NUMA nodes = 0..7 (8 NUMA nodes, symmetric layout)"
|
||||
msgstr "NUMA 节点 = 0..7 (8 个 NUMA 节点,对称布局)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:72
|
||||
msgid "total_npus = 16"
|
||||
msgstr "total_npus = 16"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:73
|
||||
msgid "running_npu_list = [0..15]"
|
||||
msgstr "running_npu_list = [0..15]"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:74
|
||||
msgid "base = 640 // 16 = 40, extra = 0"
|
||||
msgstr "base = 640 // 16 = 40, extra = 0"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:75
|
||||
msgid "Each NPU gets a 40‑CPU pool."
|
||||
msgstr "每个 NPU 获得一个 40 个 CPU 的池。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "NPU ID"
|
||||
msgstr "NPU ID"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "Assigned CPU Cores (global_slice)"
|
||||
msgstr "分配的 CPU 核心 (global_slice)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "Role Division (IRQ/Main/ACL/Release)"
|
||||
msgstr "角色划分 (IRQ/Main/ACL/Release)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "0"
|
||||
msgstr "0"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "0-39"
|
||||
msgstr "0-39"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`IRQ`: 0-1, `Main`: 2-37, `ACL`: 38, `Release`: 39"
|
||||
msgstr "`IRQ`: 0-1, `Main`: 2-37, `ACL`: 38, `Release`: 39"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "1"
|
||||
msgstr "1"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "40-79"
|
||||
msgstr "40-79"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`IRQ`: 40-41, `Main`: 42-77, `ACL`: 78, `Release`: 79"
|
||||
msgstr "`IRQ`: 40-41, `Main`: 42-77, `ACL`: 78, `Release`: 79"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "..."
|
||||
msgstr "..."
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "15"
|
||||
msgstr "15"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "600-639"
|
||||
msgstr "600-639"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`IRQ`: 600-601, `Main`: 602-637, `ACL`: 638, `Release`: 639"
|
||||
msgstr "`IRQ`: 600-601, `Main`: 602-637, `ACL`: 638, `Release`: 639"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:84
|
||||
msgid ""
|
||||
"This layout remains deterministic even when multiple processes share the "
|
||||
"same cpuset, because slicing is based on the global logical NPU ID."
|
||||
msgstr ""
|
||||
"即使多个进程共享同一个 cpuset,此布局也保持确定性,因为切片是基于全局逻辑 "
|
||||
"NPU ID 的。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:86
|
||||
msgid "Example 2: A3 global_slice, even split"
|
||||
msgstr "示例 2:A3 global_slice,均匀分割"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:88
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:109
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:142
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:161
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:182
|
||||
msgid "**Inputs**:"
|
||||
msgstr "**输入**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:91
|
||||
msgid ""
|
||||
"NUMA nodes = 0..1 (2 NUMA nodes, symmetric layout; NUMA0 = 0..11, NUMA1 ="
|
||||
" 12..23)"
|
||||
msgstr "NUMA 节点 = 0..1 (2个NUMA节点,对称布局;NUMA0 = 0..11, NUMA1 = 12..23)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:92
|
||||
msgid "total_npus = 4 (from npu-smi info -m)"
|
||||
msgstr "total_npus = 4 (来自 npu-smi info -m)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:93
|
||||
msgid "running_npu_list = [0, 1, 2, 3]"
|
||||
msgstr "running_npu_list = [0, 1, 2, 3]"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:95
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:116
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:149
|
||||
msgid "**Global slice**:"
|
||||
msgstr "**全局切片**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:97
|
||||
msgid "base = 24 // 4 = 6, extra = 0"
|
||||
msgstr "base = 24 // 4 = 6, extra = 0"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:98
|
||||
msgid "Each NPU gets a 6‑CPU pool."
|
||||
msgstr "每个NPU获得一个包含6个CPU的池。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "0-5"
|
||||
msgstr "0-5"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`IRQ`: 0-1, `Main`: 2-3, `ACL`: 4, `Release`: 5"
|
||||
msgstr "`IRQ`: 0-1, `Main`: 2-3, `ACL`: 4, `Release`: 5"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "6-11"
|
||||
msgstr "6-11"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`IRQ`: 6-7, `Main`: 8-9, `ACL`: 10, `Release`: 11"
|
||||
msgstr "`IRQ`: 6-7, `Main`: 8-9, `ACL`: 10, `Release`: 11"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "2"
|
||||
msgstr "2"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "12-17"
|
||||
msgstr "12-17"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`IRQ`: 12-13, `Main`: 14-15, `ACL`: 16, `Release`: 17"
|
||||
msgstr "`IRQ`: 12-13, `Main`: 14-15, `ACL`: 16, `Release`: 17"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "3"
|
||||
msgstr "3"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "18-23"
|
||||
msgstr "18-23"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`IRQ`: 18-19, `Main`: 20-21, `ACL`: 22, `Release`: 23"
|
||||
msgstr "`IRQ`: 18-19, `Main`: 20-21, `ACL`: 22, `Release`: 23"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:107
|
||||
msgid "Example 3: A3 global_slice, remainder distribution"
|
||||
msgstr "示例 3: A3 global_slice,余数分配"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:111
|
||||
msgid "allowed_cpus = [0..16] (17 CPUs)"
|
||||
msgstr "allowed_cpus = [0..16] (17个CPU)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:112
|
||||
msgid ""
|
||||
"NUMA nodes = 0..1 (2 NUMA nodes, symmetric layout; NUMA0 = 0..7, NUMA1 = "
|
||||
"8..16)"
|
||||
msgstr "NUMA 节点 = 0..1 (2个NUMA节点,对称布局;NUMA0 = 0..7, NUMA1 = 8..16)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:113
|
||||
msgid "total_npus = 3"
|
||||
msgstr "total_npus = 3"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:114
|
||||
msgid "running_npu_list = [0, 1, 2]"
|
||||
msgstr "running_npu_list = [0, 1, 2]"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:118
|
||||
msgid "base = 17 // 3 = 5, extra = 2"
|
||||
msgstr "base = 17 // 3 = 5, extra = 2"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:119
|
||||
msgid "NPU0 pool size = 6 (base+1)"
|
||||
msgstr "NPU0 池大小 = 6 (base+1)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:120
|
||||
msgid "NPU1 pool size = 6 (base+1)"
|
||||
msgstr "NPU1 池大小 = 6 (base+1)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:121
|
||||
msgid "NPU2 pool size = 5 (base)"
|
||||
msgstr "NPU2 池大小 = 5 (base)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "12-16"
|
||||
msgstr "12-16"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`IRQ`: 12-13, `Main`: 14, `ACL`: 15, `Release`: 16"
|
||||
msgstr "`IRQ`: 12-13, `Main`: 14, `ACL`: 15, `Release`: 16"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:129
|
||||
msgid ""
|
||||
"Note: When a pool size is exactly 5, `main` has a single CPU (pool[2]). "
|
||||
"If any pool is <5, binding raises an error."
|
||||
msgstr "注意:当池大小恰好为5时,`main` 只有一个CPU (pool[2])。如果任何池小于5,绑定将引发错误。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:131
|
||||
msgid "**NUMA analysis**:"
|
||||
msgstr "**NUMA 分析**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:133
|
||||
msgid ""
|
||||
"With the symmetric NUMA layout above (NUMA0 = 0..7, NUMA1 = 8..16), NPU0 "
|
||||
"stays within NUMA0, NPU2 stays within NUMA1, but NPU1 spans both NUMA0 "
|
||||
"(6,7) and NUMA1 (8..11). This is a direct consequence of global slicing "
|
||||
"over the ordered cpuset; the remainder distribution does not enforce NUMA"
|
||||
" boundaries."
|
||||
msgstr "在上述对称NUMA布局中 (NUMA0 = 0..7, NUMA1 = 8..16),NPU0保持在NUMA0内,NPU2保持在NUMA1内,但NPU1跨越了NUMA0 (6,7) 和 NUMA1 (8..11)。这是对有序cpuset进行全局切片的直接结果;余数分配不强制NUMA边界。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:134
|
||||
msgid ""
|
||||
"If the cpuset numbering is interleaved across NUMA nodes (non‑symmetric "
|
||||
"layout), cross‑NUMA pools can happen even earlier. This is why symmetric "
|
||||
"NUMA layout is recommended for best locality."
|
||||
msgstr "如果cpuset编号在NUMA节点间交错(非对称布局),跨NUMA池可能更早发生。这就是为什么推荐对称NUMA布局以获得最佳局部性。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:136
|
||||
msgid "Known limitations and future improvements"
|
||||
msgstr "已知限制与未来改进"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:138
|
||||
msgid ""
|
||||
"With the current `global_slice` strategy, some CPU/NPU layouts cannot "
|
||||
"avoid cross‑NUMA pools. A future enhancement should incorporate NUMA node"
|
||||
" boundaries into the slicing logic so that pools remain within a single "
|
||||
"NUMA node whenever possible."
|
||||
msgstr "使用当前的 `global_slice` 策略,某些CPU/NPU布局无法避免跨NUMA池。未来的增强应将NUMA节点边界纳入切片逻辑,以便池尽可能保持在单个NUMA节点内。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:140
|
||||
msgid "Example 4: global_slice with visible subset of NPUs"
|
||||
msgstr "示例 4: 使用NPU可见子集的 global_slice"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:144
|
||||
msgid "total_npus = 8 (from npu-smi info -m)"
|
||||
msgstr "total_npus = 8 (来自 npu-smi info -m)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:145
|
||||
msgid "running_npu_list = [2, 3] (filtered by ASCEND_RT_VISIBLE_DEVICES)"
|
||||
msgstr "running_npu_list = [2, 3] (由 ASCEND_RT_VISIBLE_DEVICES 过滤)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:146
|
||||
msgid "allowed_cpus = [0..39] (40 CPUs)"
|
||||
msgstr "allowed_cpus = [0..39] (40个CPU)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:147
|
||||
msgid ""
|
||||
"NUMA nodes = 0..3 (4 NUMA nodes, symmetric layout; 0..9, 10..19, 20..29, "
|
||||
"30..39)"
|
||||
msgstr "NUMA 节点 = 0..3 (4个NUMA节点,对称布局;0..9, 10..19, 20..29, 30..39)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:151
|
||||
msgid "base = 40 // 8 = 5, extra = 0"
|
||||
msgstr "base = 40 // 8 = 5, extra = 0"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:152
|
||||
msgid ""
|
||||
"Only the visible logical NPUs get pools, but slicing uses the global NPU "
|
||||
"ID so different processes do not overlap."
|
||||
msgstr "只有可见的逻辑NPU获得池,但切片使用全局NPU ID,因此不同进程不会重叠。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "10-14"
|
||||
msgstr "10-14"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`IRQ`: 10-11, `Main`: 12, `ACL`: 13, `Release`: 14"
|
||||
msgstr "`IRQ`: 10-11, `Main`: 12, `ACL`: 13, `Release`: 14"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "15-19"
|
||||
msgstr "15-19"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`IRQ`: 15-16, `Main`: 17, `ACL`: 18, `Release`: 19"
|
||||
msgstr "`IRQ`: 15-16, `Main`: 17, `ACL`: 18, `Release`: 19"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:159
|
||||
msgid "Example 5: A2/310P topo_affinity with NUMA extension"
|
||||
msgstr "示例 5: 具有NUMA扩展的 A2/310P topo_affinity"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:163
|
||||
#, python-brace-format
|
||||
msgid "npu_affinity = {0: [0..7], 1: [0..7]} (from `npu-smi info -t topo`)"
|
||||
msgstr "npu_affinity = {0: [0..7], 1: [0..7]} (来自 `npu-smi info -t topo`)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:164
|
||||
msgid "allowed_cpus = [0..15] (16 CPUs)"
|
||||
msgstr "allowed_cpus = [0..15] (16个CPU)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:165
|
||||
msgid "NUMA nodes = 0..1 (2 NUMA nodes; NUMA0 = 0..7, NUMA1 = 8..15)"
|
||||
msgstr "NUMA 节点 = 0..1 (2个NUMA节点;NUMA0 = 0..7, NUMA1 = 8..15)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:167
|
||||
msgid "**NUMA extension**:"
|
||||
msgstr "**NUMA 扩展**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:169
|
||||
msgid ""
|
||||
"Both NPUs are on NUMA0, so each pool extends to the nearest NUMA node to "
|
||||
"reduce contention."
|
||||
msgstr "两个NPU都在NUMA0上,因此每个池扩展到最近的NUMA节点以减少争用。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:170
|
||||
msgid "NPU0 extends to NUMA1 -> [0..15]"
|
||||
msgstr "NPU0 扩展到 NUMA1 -> [0..15]"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:171
|
||||
msgid "NPU1 extends to NUMA1 -> [0..15]"
|
||||
msgstr "NPU1 扩展到 NUMA1 -> [0..15]"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:173
|
||||
msgid ""
|
||||
"Because both pools are identical, the allocator applies average "
|
||||
"distribution across NPUs to avoid overlap. With a pool [0..15] and 2 "
|
||||
"NPUs, the final pools become:"
|
||||
msgstr "由于两个池相同,分配器应用跨NPU的平均分配以避免重叠。对于池 [0..15] 和 2个NPU,最终池变为:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "Assigned CPU Cores (topo_affinity)"
|
||||
msgstr "分配的CPU核心 (topo_affinity)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "0-7"
|
||||
msgstr "0-7"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`IRQ`: 0-1, `Main`: 2-5, `ACL`: 6, `Release`: 7"
|
||||
msgstr "`IRQ`: 0-1, `Main`: 2-5, `ACL`: 6, `Release`: 7"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "8-15"
|
||||
msgstr "8-15"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "`IRQ`: 8-9, `Main`: 10-13, `ACL`: 14, `Release`: 15"
|
||||
msgstr "`IRQ`: 8-9, `Main`: 10-13, `ACL`: 14, `Release`: 15"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:180
|
||||
msgid "Example 6: Minimum CPUs per NPU"
|
||||
msgstr "示例 6: 每个NPU的最小CPU数"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:184
|
||||
msgid "total_npus = 2"
|
||||
msgstr "total_npus = 2"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:185
|
||||
msgid "allowed_cpus = [0..7] (8 CPUs)"
|
||||
msgstr "allowed_cpus = [0..7] (8个CPU)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:186
|
||||
msgid ""
|
||||
"NUMA nodes = 0..1 (2 NUMA nodes, symmetric layout; NUMA0 = 0..3, NUMA1 = "
|
||||
"4..7)"
|
||||
msgstr "NUMA 节点 = 0..1 (2个NUMA节点,对称布局;NUMA0 = 0..3, NUMA1 = 4..7)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:188
|
||||
msgid "**Result**:"
|
||||
msgstr "**结果**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:190
|
||||
msgid ""
|
||||
"base = 4, which is < 5, so binding fails with: \"Insufficient CPUs for "
|
||||
"binding with IRQ/ACL/REL reservations...\""
|
||||
msgstr "base = 4,小于5,因此绑定失败,错误信息为:\"用于IRQ/ACL/REL预留绑定的CPU不足...\""
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "Assigned CPU Cores"
|
||||
msgstr "分配的CPU核心"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "N/A"
|
||||
msgstr "不适用"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md
|
||||
msgid "Binding error (insufficient CPUs per NPU)"
|
||||
msgstr "绑定错误(每个NPU的CPU不足)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:197
|
||||
msgid ""
|
||||
"To resolve, either reduce total_npus or enlarge the cpuset so that each "
|
||||
"NPU has at least 5 CPUs."
|
||||
msgstr "要解决此问题,要么减少 total_npus,要么扩大 cpuset,使每个NPU至少有5个CPU。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:199
|
||||
msgid "Logging and verification"
|
||||
msgstr "日志记录与验证"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:201
|
||||
msgid "Logs show the selected binding mode and the allocation plan, for example:"
|
||||
msgstr "日志显示选定的绑定模式和分配计划,例如:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:202
|
||||
msgid "`[cpu_bind_mode] mode=global_slice rank=0 visible_npus=[...]`"
|
||||
msgstr "`[cpu_bind_mode] mode=global_slice rank=0 visible_npus=[...]`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:203
|
||||
msgid "`The CPU allocation plan is as follows: ...`"
|
||||
msgstr "`CPU分配计划如下:...`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:204
|
||||
msgid "You can verify affinity via taskset or `/proc/<pid>/status` after startup."
|
||||
msgstr "启动后,您可以通过 taskset 或 `/proc/<pid>/status` 验证亲和性。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:206
|
||||
msgid "Limitations & Notes"
|
||||
msgstr "限制与注意事项"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:208
|
||||
msgid "**ARM‑only**: Binding is skipped on non‑ARM CPUs."
|
||||
msgstr "**仅限ARM**:在非ARM CPU上跳过绑定。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:209
|
||||
msgid ""
|
||||
"**Minimum CPU requirement**: Each logical NPU requires at least 5 CPUs. "
|
||||
"If the cpuset is smaller, binding fails with an error."
|
||||
msgstr "**最小CPU要求**:每个逻辑NPU至少需要5个CPU。如果cpuset更小,绑定将失败并报错。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:210
|
||||
msgid ""
|
||||
"**NUMA symmetry assumption**: For best locality, the current strategies "
|
||||
"assume the cpuset is evenly distributed across NUMA nodes and CPU "
|
||||
"numbering aligns with NUMA layout; otherwise NUMA locality may be "
|
||||
"suboptimal."
|
||||
msgstr "**NUMA对称性假设**:为获得最佳局部性,当前策略假设cpuset在NUMA节点间均匀分布,且CPU编号与NUMA布局对齐;否则NUMA局部性可能不理想。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:211
|
||||
msgid ""
|
||||
"Example (symmetric layout): 2 NUMA nodes, 64 CPUs total. NUMA0 = CPUs "
|
||||
"0–31, NUMA1 = CPUs 32–63, and the cpuset is 0–63. With 4 logical NPUs, "
|
||||
"global slicing yields 16 CPUs per NPU (0–15, 16–31, 32–47, 48–63), so "
|
||||
"each NPU’s pool stays within a single NUMA node."
|
||||
msgstr "示例(对称布局):2个NUMA节点,总共64个CPU。NUMA0 = CPU 0–31,NUMA1 = CPU 32–63,cpuset为0–63。对于4个逻辑NPU,全局切片每个NPU产生16个CPU (0–15, 16–31, 32–47, 48–63),因此每个NPU的池保持在单个NUMA节点内。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:212
|
||||
msgid "**Runtime dependencies**:"
|
||||
msgstr "**运行时依赖**:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:213
|
||||
msgid "Requires npu‑smi and lscpu commands."
|
||||
msgstr "需要 npu‑smi 和 lscpu 命令。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:214
|
||||
msgid "IRQ binding requires write access to /proc/irq."
|
||||
msgstr "IRQ绑定需要对 /proc/irq 的写访问权限。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:215
|
||||
msgid "Memory binding requires migratepages; otherwise it is skipped."
|
||||
msgstr "内存绑定需要 migratepages;否则将被跳过。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:216
|
||||
msgid ""
|
||||
"**IRQ side effects**: irqbalance may be stopped to avoid overriding "
|
||||
"bindings."
|
||||
msgstr "**IRQ副作用**:可能会停止 irqbalance 以避免覆盖绑定。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:217
|
||||
msgid ""
|
||||
"**Per‑process behavior**: Only the current rank’s NPU is used for IRQ "
|
||||
"binding to avoid cross‑process overwrite."
|
||||
msgstr "**每进程行为**:仅使用当前 rank 的 NPU 进行 IRQ 绑定,以避免跨进程覆盖。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:219
|
||||
msgid "Debug logging"
|
||||
msgstr "调试日志"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:221
|
||||
msgid ""
|
||||
"Use the standard vLLM logging configuration to enable debug logs. The "
|
||||
"binding process emits debug messages (e.g., `[cpu_global_slice] ...`) "
|
||||
"when debug level is enabled."
|
||||
msgstr "使用标准的 vLLM 日志配置来启用调试日志。当启用调试级别时,绑定过程会发出调试消息(例如 `[cpu_global_slice] ...`)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:223
|
||||
msgid "References"
|
||||
msgstr "参考"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:225
|
||||
msgid ""
|
||||
"CPU binding implementation: vllm_ascend/cpu_binding.py (`DeviceInfo`, "
|
||||
"`CpuAlloc`, `bind_cpus`)"
|
||||
msgstr "CPU 绑定实现:vllm_ascend/cpu_binding.py (`DeviceInfo`, `CpuAlloc`, `bind_cpus`)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:226
|
||||
msgid ""
|
||||
"Worker integration: vllm_ascend/worker/worker.py "
|
||||
"(`NPUWorker._init_device`)"
|
||||
msgstr "Worker 集成:vllm_ascend/worker/worker.py (`NPUWorker._init_device`)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:227
|
||||
msgid ""
|
||||
"Additional config option: "
|
||||
"docs/source/user_guide/configuration/additional_config.md "
|
||||
"(`enable_cpu_binding`)"
|
||||
msgstr "附加配置选项:docs/source/user_guide/configuration/additional_config.md (`enable_cpu_binding`)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/cpu_binding.md:228
|
||||
msgid "Tests: tests/ut/device_allocator/test_cpu_binding.py"
|
||||
msgstr "测试:tests/ut/device_allocator/test_cpu_binding.py"
|
||||
@@ -0,0 +1,360 @@
|
||||
# SOME DESCRIPTIVE TITLE.
|
||||
# Copyright (C) 2025, vllm-ascend team
|
||||
# This file is distributed under the same license as the vllm-ascend
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend \n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:1
|
||||
msgid "Disaggregated-prefill"
|
||||
msgstr "解耦式预填充"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:3
|
||||
msgid "Why disaggregated-prefill?"
|
||||
msgstr "为何需要解耦式预填充?"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:5
|
||||
msgid ""
|
||||
"This feature addresses the need to optimize the **Time Per Output Token "
|
||||
"(TPOT)** and **Time To First Token (TTFT)** in large-scale inference "
|
||||
"tasks. The motivation is two-fold:"
|
||||
msgstr ""
|
||||
"此功能旨在优化大规模推理任务中的**单输出令牌时间 (TPOT)** 和**首令牌时间 "
|
||||
"(TTFT)**。其动机主要有两方面:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:7
|
||||
msgid ""
|
||||
"**Adjusting Parallel Strategy and Instance Count for P and D Nodes** "
|
||||
"Using the disaggregated-prefill strategy, this feature allows the system "
|
||||
"to flexibly adjust the parallelization strategy (e.g., data parallelism "
|
||||
"(dp), tensor parallelism (tp), and expert parallelism (ep)) and the "
|
||||
"instance count for both P (Prefiller) and D (Decoder) nodes. This leads "
|
||||
"to better system performance tuning, particularly for **TTFT** and "
|
||||
"**TPOT**."
|
||||
msgstr ""
|
||||
"**调整 P 节点和 D 节点的并行策略与实例数量** 采用解耦式预填充策略,此功能允许系统灵活调整 P(预填充器)节点和 D(解码器)节点的并行化策略(例如数据并行 (dp)、张量并行 (tp) 和专家并行 (ep))以及实例数量。这有助于实现更好的系统性能调优,特别是针对 **TTFT** 和 **TPOT**。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:10
|
||||
msgid ""
|
||||
"**Optimizing TPOT** Without the disaggregated-prefill strategy, prefill "
|
||||
"tasks are inserted during decoding, which results in inefficiencies and "
|
||||
"delays. Disaggregated-prefill solves this by allowing for better control "
|
||||
"over the system’s **TPOT**. By managing chunked prefill tasks "
|
||||
"effectively, the system avoids the challenge of determining the optimal "
|
||||
"chunk size and provides more reliable control over the time taken for "
|
||||
"generating output tokens."
|
||||
msgstr ""
|
||||
"**优化 TPOT** 在没有解耦式预填充策略的情况下,预填充任务会在解码过程中插入,导致效率低下和延迟。解耦式预填充通过允许更好地控制系统 **TPOT** 来解决此问题。通过有效管理分块的预填充任务,系统避免了确定最佳分块大小的挑战,并对生成输出令牌所需时间提供了更可靠的控制。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:15
|
||||
msgid "Usage"
|
||||
msgstr "使用方法"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:17
|
||||
msgid ""
|
||||
"vLLM Ascend currently supports two types of connectors for handling KV "
|
||||
"cache management:"
|
||||
msgstr "vLLM Ascend 目前支持两种用于处理 KV 缓存管理的连接器:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:19
|
||||
msgid "**MooncakeConnector**: D nodes pull KV cache from P nodes."
|
||||
msgstr "**MooncakeConnector**:D 节点从 P 节点拉取 KV 缓存。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:20
|
||||
msgid ""
|
||||
"**MooncakeLayerwiseConnector**: P nodes push KV cache to D nodes in a "
|
||||
"layered manner."
|
||||
msgstr "**MooncakeLayerwiseConnector**:P 节点以分层方式将 KV 缓存推送到 D 节点。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:22
|
||||
msgid ""
|
||||
"For step-by-step deployment and configuration, refer to the following "
|
||||
"guide: "
|
||||
"[https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/pd_disaggregation_mooncake_multi_node.html](https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/pd_disaggregation_mooncake_multi_node.html)"
|
||||
msgstr ""
|
||||
"有关分步部署和配置,请参考以下指南: "
|
||||
"[https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/pd_disaggregation_mooncake_multi_node.html](https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/pd_disaggregation_mooncake_multi_node.html)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:27
|
||||
msgid "How It Works"
|
||||
msgstr "工作原理"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:29
|
||||
msgid "1. Design Approach"
|
||||
msgstr "1. 设计思路"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:31
|
||||
msgid ""
|
||||
"Under the disaggregated-prefill, a global proxy receives external "
|
||||
"requests, forwarding prefill to P nodes and decode to D nodes; the KV "
|
||||
"cache (key–value cache) is exchanged between P and D nodes via peer-to-"
|
||||
"peer (P2P) communication."
|
||||
msgstr ""
|
||||
"在解耦式预填充架构下,一个全局代理接收外部请求,将预填充请求转发给 P 节点,将解码请求转发给 D 节点;KV 缓存(键值缓存)通过点对点 (P2P) 通信在 P 节点和 D 节点之间交换。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:33
|
||||
msgid "2. Implementation Design"
|
||||
msgstr "2. 实现设计"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:35
|
||||
msgid ""
|
||||
"Our design diagram is shown below, illustrating the pull and push schemes"
|
||||
" respectively.  "
|
||||
""
|
||||
msgstr ""
|
||||
"我们的设计图如下所示,分别展示了拉取和推送方案。 "
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:35
|
||||
msgid "alt text"
|
||||
msgstr "替代文本"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:39
|
||||
msgid "Mooncake Connector"
|
||||
msgstr "Mooncake 连接器"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:41
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:49
|
||||
msgid "The request is sent to the Proxy’s `_handle_completions` endpoint."
|
||||
msgstr "请求被发送到代理的 `_handle_completions` 端点。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:42
|
||||
msgid ""
|
||||
"The Proxy calls `select_prefiller` to choose a P node and forwards the "
|
||||
"request, configuring `kv_transfer_params` with `do_remote_decode=True`, "
|
||||
"`max_completion_tokens=1`, and `min_tokens=1`."
|
||||
msgstr ""
|
||||
"代理调用 `select_prefiller` 选择一个 P 节点并转发请求,配置 `kv_transfer_params` 为 `do_remote_decode=True`、`max_completion_tokens=1` 和 `min_tokens=1`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:43
|
||||
msgid ""
|
||||
"After the P node’s scheduler finishes prefill, `update_from_output` "
|
||||
"invokes the schedule connector’s `request_finished` to defer KV cache "
|
||||
"release, constructs `kv_transfer_params` with `do_remote_prefill=True`, "
|
||||
"and returns to the Proxy."
|
||||
msgstr ""
|
||||
"P 节点的调度器完成预填充后,`update_from_output` 调用调度连接器的 `request_finished` 以延迟释放 KV 缓存,构建 `kv_transfer_params` 为 `do_remote_prefill=True`,并返回给代理。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:44
|
||||
msgid ""
|
||||
"The Proxy calls `select_decoder` to choose a D node and forwards the "
|
||||
"request."
|
||||
msgstr "代理调用 `select_decoder` 选择一个 D 节点并转发请求。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:45
|
||||
msgid ""
|
||||
"On the D node, the scheduler marks the request as "
|
||||
"`RequestStatus.WAITING_FOR_REMOTE_KVS`, pre-allocates KV cache, calls "
|
||||
"`kv_connector_no_forward` to pull the remote KV cache, then notifies the "
|
||||
"P node to release KV cache and proceeds with decoding to return the "
|
||||
"result."
|
||||
msgstr ""
|
||||
"在 D 节点上,调度器将请求标记为 `RequestStatus.WAITING_FOR_REMOTE_KVS`,预分配 KV 缓存,调用 `kv_connector_no_forward` 拉取远程 KV 缓存,然后通知 P 节点释放 KV 缓存并继续解码以返回结果。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:47
|
||||
msgid "Mooncake Layerwise Connector"
|
||||
msgstr "Mooncake 分层连接器"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:50
|
||||
msgid ""
|
||||
"The Proxy calls `select_decoder` to choose a D node and forwards the "
|
||||
"request, configuring `kv_transfer_params` with `do_remote_prefill=True` "
|
||||
"and setting the `metaserver` endpoint."
|
||||
msgstr ""
|
||||
"代理调用 `select_decoder` 选择一个 D 节点并转发请求,配置 `kv_transfer_params` 为 `do_remote_prefill=True` 并设置 `metaserver` 端点。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:51
|
||||
msgid ""
|
||||
"On the D node, the scheduler uses `kv_transfer_params` to mark the "
|
||||
"request as `RequestStatus.WAITING_FOR_REMOTE_KVS`, pre-allocates KV "
|
||||
"cache, then calls `kv_connector_no_forward` to send a request to the "
|
||||
"metaserver and waits for the KV cache transfer to complete."
|
||||
msgstr ""
|
||||
"在 D 节点上,调度器使用 `kv_transfer_params` 将请求标记为 `RequestStatus.WAITING_FOR_REMOTE_KVS`,预分配 KV 缓存,然后调用 `kv_connector_no_forward` 向元服务器发送请求并等待 KV 缓存传输完成。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:52
|
||||
msgid ""
|
||||
"The Proxy’s `metaserver` endpoint receives the request, calls "
|
||||
"`select_prefiller` to choose a P node, and forwards it with "
|
||||
"`kv_transfer_params` set to `do_remote_decode=True`, "
|
||||
"`max_completion_tokens=1`, and `min_tokens=1`."
|
||||
msgstr ""
|
||||
"代理的 `metaserver` 端点接收请求,调用 `select_prefiller` 选择一个 P 节点,并转发请求,设置 `kv_transfer_params` 为 `do_remote_decode=True`、`max_completion_tokens=1` 和 `min_tokens=1`。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:53
|
||||
msgid ""
|
||||
"During processing, the P node’s scheduler pushes KV cache layer-wise; "
|
||||
"once all layers pushing is complete, it releases the request and notifies"
|
||||
" the D node to begin decoding."
|
||||
msgstr "在处理过程中,P 节点的调度器逐层推送 KV 缓存;所有层推送完成后,它释放请求并通知 D 节点开始解码。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:54
|
||||
msgid "The D node performs decoding and returns the result."
|
||||
msgstr "D 节点执行解码并返回结果。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:56
|
||||
msgid "3. Interface Design"
|
||||
msgstr "3. 接口设计"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:58
|
||||
msgid ""
|
||||
"Taking MooncakeConnector as an example, the system is organized into "
|
||||
"three primary classes:"
|
||||
msgstr "以 MooncakeConnector 为例,系统被组织成三个主要类:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:60
|
||||
msgid "**MooncakeConnector**: Base class that provides core interfaces."
|
||||
msgstr "**MooncakeConnector**:提供核心接口的基类。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:61
|
||||
msgid ""
|
||||
"**MooncakeConnectorScheduler**: Interface for scheduling the connectors "
|
||||
"within the engine core, responsible for managing KV cache transfer "
|
||||
"requirements and completion."
|
||||
msgstr "**MooncakeConnectorScheduler**:用于在引擎核心内调度连接器的接口,负责管理 KV 缓存传输需求和完成情况。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:62
|
||||
msgid ""
|
||||
"**MooncakeConnectorWorker**: Interface for managing KV cache registration"
|
||||
" and transfer in worker processes."
|
||||
msgstr "**MooncakeConnectorWorker**:用于在工作进程中管理 KV 缓存注册和传输的接口。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:64
|
||||
msgid "4. Specifications Design"
|
||||
msgstr "4. 规格设计"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:66
|
||||
msgid ""
|
||||
"This feature is flexible and supports various configurations, including "
|
||||
"setups with MLA and GQA models. It is compatible with A2 and A3 hardware "
|
||||
"configurations and facilitates scenarios involving both equal and unequal"
|
||||
" TP setups across multiple P and D nodes."
|
||||
msgstr ""
|
||||
"此功能灵活,支持多种配置,包括使用 MLA 和 GQA 模型的设置。它与 A2 和 A3 硬件配置兼容,并支持跨多个 P 节点和 D 节点的相等和不相等 TP 设置场景。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md
|
||||
msgid "Feature"
|
||||
msgstr "功能"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md
|
||||
msgid "Status"
|
||||
msgstr "状态"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md
|
||||
msgid "A2"
|
||||
msgstr "A2"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md
|
||||
msgid "🟢 Functional"
|
||||
msgstr "🟢 功能正常"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md
|
||||
msgid "A3"
|
||||
msgstr "A3"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md
|
||||
msgid "equal TP configuration"
|
||||
msgstr "相等 TP 配置"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md
|
||||
msgid "unequal TP configuration"
|
||||
msgstr "不相等 TP 配置"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md
|
||||
msgid "MLA"
|
||||
msgstr "MLA"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md
|
||||
msgid "GQA"
|
||||
msgstr "GQA"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:77
|
||||
msgid "🟢 Functional: Fully operational, with ongoing optimizations."
|
||||
msgstr "🟢 功能正常:完全可运行,正在进行优化。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:78
|
||||
msgid "🔵 Experimental: Experimental support, interfaces and functions may change."
|
||||
msgstr "🔵 实验性:实验性支持,接口和功能可能发生变化。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:79
|
||||
msgid "🚧 WIP: Under active development, will be supported soon."
|
||||
msgstr "🚧 开发中:正在积极开发,即将支持。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:80
|
||||
msgid ""
|
||||
"🟡 Planned: Scheduled for future implementation (some may have open "
|
||||
"PRs/RFCs)."
|
||||
msgstr "🟡 计划中:计划在未来实现(部分可能已有开放的 PR/RFC)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:81
|
||||
msgid "🔴 NO plan/Deprecated: No plan or deprecated by vLLM."
|
||||
msgstr "🔴 无计划/已弃用:无计划或已被 vLLM 弃用。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:85
|
||||
msgid "DFX Analysis"
|
||||
msgstr "DFX 分析"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:87
|
||||
msgid "1. Config Parameter Validation"
|
||||
msgstr "1. 配置参数验证"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:89
|
||||
msgid ""
|
||||
"Validate KV transfer config by checking whether the kv_connector type is "
|
||||
"supported and whether kv_connector_module_path exists and is loadable. On"
|
||||
" transfer failures, emit clear error logs for diagnostics."
|
||||
msgstr ""
|
||||
"通过检查 kv_connector 类型是否受支持以及 kv_connector_module_path 是否存在且可加载来验证 KV 传输配置。传输失败时,发出清晰的错误日志以供诊断。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:91
|
||||
msgid "2. Port Conflict Detection"
|
||||
msgstr "2. 端口冲突检测"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:93
|
||||
msgid ""
|
||||
"Before startup, perform a port-usage check on configured ports (e.g., "
|
||||
"rpc_port, metrics_port, http_port/metaserver) by attempting to bind. If a"
|
||||
" port is already in use, fail fast and log an error."
|
||||
msgstr "启动前,通过尝试绑定来对配置的端口(例如 rpc_port、metrics_port、http_port/metaserver)进行端口使用情况检查。如果端口已被占用,快速失败并记录错误。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:95
|
||||
msgid "3. PD Ratio Validation"
|
||||
msgstr "3. PD 比例验证"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:97
|
||||
msgid ""
|
||||
"Under non-symmetric PD scenarios, validate the P-to-D tp ratio against "
|
||||
"expected and scheduling constraints to ensure correct and reliable "
|
||||
"operation."
|
||||
msgstr "在非对称 PD 场景下,根据预期和调度约束验证 P 到 D 的 tp 比例,以确保正确可靠的操作。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:101
|
||||
msgid "Limitations"
|
||||
msgstr "限制"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:103
|
||||
msgid ""
|
||||
"Heterogeneous P and D nodes are not supported—for example, running P "
|
||||
"nodes on A2 and D nodes on A3."
|
||||
msgstr "不支持异构的 P 节点和 D 节点——例如,在 A2 上运行 P 节点,在 A3 上运行 D 节点。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/disaggregated_prefill.md:105
|
||||
msgid ""
|
||||
"In non-symmetric TP configurations, only cases where the P nodes have a "
|
||||
"higher TP degree than the D nodes and the P TP count is an integer "
|
||||
"multiple of the D TP count are supported (i.e., P_tp > D_tp and P_tp % "
|
||||
"D_tp = 0)."
|
||||
msgstr "在非对称 TP 配置中,仅支持 P 节点的 TP 度数高于 D 节点且 P 节点的 TP 数量是 D 节点 TP 数量的整数倍的情况(即 P_tp > D_tp 且 P_tp % D_tp = 0)。"
|
||||
@@ -0,0 +1,467 @@
|
||||
# SOME DESCRIPTIVE TITLE.
|
||||
# Copyright (C) 2025, vllm-ascend team
|
||||
# This file is distributed under the same license as the vllm-ascend
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend \n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:1
|
||||
msgid "Expert Parallelism Load Balancer (EPLB)"
|
||||
msgstr "专家并行负载均衡器 (EPLB)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:3
|
||||
msgid "Why We Need EPLB?"
|
||||
msgstr "为什么需要 EPLB?"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:5
|
||||
msgid ""
|
||||
"When using Expert Parallelism (EP), different experts are assigned to "
|
||||
"different NPUs. Given that the load of various experts may vary depending"
|
||||
" on the current workload, it is crucial to maintain balanced loads across"
|
||||
" different NPUs. We adopt a redundant experts strategy by duplicating "
|
||||
"heavily-loaded experts. Then, we heuristically pack these duplicated "
|
||||
"experts onto NPUs to ensure load balancing across them. Moreover, thanks "
|
||||
"to the group-limited expert routing used in MoE models, we also attempt "
|
||||
"to place experts of the same group on the same node to reduce inter-node "
|
||||
"data traffic, whenever possible."
|
||||
msgstr ""
|
||||
"在使用专家并行 (EP) 时,不同的专家被分配到不同的 NPU 上。鉴于不同专家的负载可能因当前工作负载而异,保持不同 NPU 之间的负载均衡至关重要。我们采用冗余专家策略,通过复制高负载的专家来实现。然后,我们启发式地将这些复制的专家打包到 NPU 上,以确保它们之间的负载均衡。此外,得益于 MoE 模型中使用的组限制专家路由,我们也尽可能将同一组的专家放置在同一节点上,以减少节点间的数据流量。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:7
|
||||
msgid ""
|
||||
"To facilitate reproduction and deployment, vLLM Ascend supports the "
|
||||
"deployed EP load balancing algorithm in `vllm_ascend/eplb/core/policy`. "
|
||||
"The algorithm computes a balanced expert replication and placement plan "
|
||||
"based on the estimated expert loads. Note that the exact method for "
|
||||
"predicting expert loads is outside the scope of this repository. A common"
|
||||
" method is to use a moving average of historical statistics."
|
||||
msgstr ""
|
||||
"为了方便复现和部署,vLLM Ascend 在 `vllm_ascend/eplb/core/policy` 中支持已部署的 EP 负载均衡算法。该算法根据估计的专家负载计算一个均衡的专家复制和放置计划。请注意,预测专家负载的具体方法不在本仓库的讨论范围内。一种常见的方法是使用历史统计数据的移动平均值。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:9
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:9
|
||||
msgid "eplb"
|
||||
msgstr "eplb"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:11
|
||||
msgid "How to Use EPLB?"
|
||||
msgstr "如何使用 EPLB?"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:13
|
||||
msgid ""
|
||||
"Please refer to the EPLB section of the user guide for detailed "
|
||||
"information: [How to Use "
|
||||
"EPLB](../../user_guide/feature_guide/eplb_swift_balancer.md)"
|
||||
msgstr ""
|
||||
"请参阅用户指南中的 EPLB 部分以获取详细信息:[如何使用 "
|
||||
"EPLB](../../user_guide/feature_guide/eplb_swift_balancer.md)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:15
|
||||
msgid "How It Works?"
|
||||
msgstr "工作原理"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:17
|
||||
msgid "**EPLB Module Architecture**"
|
||||
msgstr "**EPLB 模块架构**"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:40
|
||||
msgid ""
|
||||
"**1. Adaptor Module** *Handles registration and adaptation for "
|
||||
"different MoE model types*"
|
||||
msgstr "**1. 适配器模块** *处理不同 MoE 模型类型的注册和适配*"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:43
|
||||
msgid ""
|
||||
"`abstract_adaptor.py` Abstract base class defining unified registration"
|
||||
" interfaces for EPLB adapters"
|
||||
msgstr "`abstract_adaptor.py` 定义 EPLB 适配器统一注册接口的抽象基类"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:45
|
||||
msgid ""
|
||||
"`vllm_adaptor.py` Implementation supporting Qwen3-MoE and DeepSeek "
|
||||
"models, standardizing parameter handling for policy algorithms"
|
||||
msgstr "`vllm_adaptor.py` 支持 Qwen3-MoE 和 DeepSeek 模型的实现,标准化策略算法的参数处理"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:48
|
||||
msgid ""
|
||||
"**2. Core Module** *Implements core algorithms, updates, and "
|
||||
"asynchronous processing*"
|
||||
msgstr "**2. 核心模块** *实现核心算法、更新和异步处理*"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:51
|
||||
msgid ""
|
||||
"**Policy Submodule** *Load balancing algorithms with factory pattern "
|
||||
"instantiation*"
|
||||
msgstr "**策略子模块** *采用工厂模式实例化的负载均衡算法*"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:53
|
||||
msgid ""
|
||||
"`policy_abstract.py` Abstract class for load balancing strategy "
|
||||
"interfaces"
|
||||
msgstr "`policy_abstract.py` 负载均衡策略接口的抽象类"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:55
|
||||
msgid ""
|
||||
"`policy_default_eplb.py` Default implementation of open-source EPLB "
|
||||
"paper algorithm"
|
||||
msgstr "`policy_default_eplb.py` 开源 EPLB 论文算法的默认实现"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:57
|
||||
msgid ""
|
||||
"`policy_swift_balancer.py` Enhanced version optimizing expert swaps for"
|
||||
" low-bandwidth devices (e.g., A2)"
|
||||
msgstr "`policy_swift_balancer.py` 针对低带宽设备(例如 A2)优化专家交换的增强版本"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:59
|
||||
msgid ""
|
||||
"`policy_flashlb.py` Threshold-based adjustment reducing operational "
|
||||
"costs through layer-wise fluctuation detection"
|
||||
msgstr "`policy_flashlb.py` 基于阈值的调整,通过逐层波动检测降低操作成本"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:61
|
||||
msgid ""
|
||||
"`policy_factory.py` Strategy factory for automatic algorithm "
|
||||
"instantiation"
|
||||
msgstr "`policy_factory.py` 用于自动算法实例化的策略工厂"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:64
|
||||
msgid ""
|
||||
"`eplb_device_transfer_loader.py` Manages expert table/weight "
|
||||
"transmission and updates"
|
||||
msgstr "`eplb_device_transfer_loader.py` 管理专家表/权重的传输和更新"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:66
|
||||
msgid "`eplb_utils.py` Utilities for expert table initialization and mapping"
|
||||
msgstr "`eplb_utils.py` 用于专家表初始化和映射的实用工具"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:68
|
||||
msgid ""
|
||||
"`eplb_worker.py` Asynchronous algorithm orchestration and result "
|
||||
"processing"
|
||||
msgstr "`eplb_worker.py` 异步算法编排和结果处理"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:71
|
||||
msgid "**3. System Components**"
|
||||
msgstr "**3. 系统组件**"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:73
|
||||
msgid ""
|
||||
"`eplb_updator.py` Central coordinator for load balancing during "
|
||||
"inference workflows"
|
||||
msgstr "`eplb_updator.py` 推理工作流中负载均衡的中心协调器"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:75
|
||||
msgid "`utils.py` General utilities for EPLB interface registration"
|
||||
msgstr "`utils.py` EPLB 接口注册的通用实用工具"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:78
|
||||
msgid "*Key Optimizations:*"
|
||||
msgstr "*关键优化点:*"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:80
|
||||
msgid "Maintained original structure while improving technical clarity"
|
||||
msgstr "保持原始结构的同时提高了技术清晰度"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:81
|
||||
msgid "Standardized terminology"
|
||||
msgstr "标准化术语"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:82
|
||||
msgid "Enhanced algorithm differentiation through concise descriptors"
|
||||
msgstr "通过简洁的描述符增强了算法区分度"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:83
|
||||
msgid "Improved scoping through hierarchical presentation"
|
||||
msgstr "通过分层展示改进了范围界定"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:84
|
||||
msgid "Preserved file/class relationships while optimizing readability"
|
||||
msgstr "在优化可读性的同时保留了文件/类关系"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:86
|
||||
msgid "Default Algorithm"
|
||||
msgstr "默认算法"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:88
|
||||
msgid "Hierarchical Load Balancing"
|
||||
msgstr "分层负载均衡"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:90
|
||||
msgid ""
|
||||
"When the number of server nodes evenly divides the number of expert "
|
||||
"groups, we use the hierarchical load balancing policy to leverage group-"
|
||||
"limited expert routing. We first pack the expert groups onto nodes "
|
||||
"evenly, ensuring balanced loads across different nodes. Then, we "
|
||||
"replicate the experts within each node. Finally, we pack the replicated "
|
||||
"experts onto individual NPUs to ensure load balancing across them. The "
|
||||
"hierarchical load balancing policy can be used in the prefilling stage "
|
||||
"with a smaller expert-parallel size."
|
||||
msgstr ""
|
||||
"当服务器节点数量能整除专家组数量时,我们使用分层负载均衡策略来利用组限制专家路由。我们首先将专家组均匀地打包到节点上,确保不同节点间的负载均衡。然后,我们在每个节点内复制专家。最后,我们将复制的专家打包到各个 NPU 上,以确保它们之间的负载均衡。分层负载均衡策略可以在预填充阶段使用,此时专家并行规模较小。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:92
|
||||
msgid "Global Load Balancing"
|
||||
msgstr "全局负载均衡"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:94
|
||||
msgid ""
|
||||
"In other cases, we use the global load balancing policy, which replicates"
|
||||
" experts globally regardless of expert groups, and packs the replicated "
|
||||
"experts onto individual NPUs. This policy can be adopted in the decoding "
|
||||
"stage with a larger expert-parallel size."
|
||||
msgstr ""
|
||||
"在其他情况下,我们使用全局负载均衡策略,该策略不考虑专家组,而是在全局范围内复制专家,并将复制的专家打包到各个 NPU 上。此策略可以在解码阶段采用,此时专家并行规模较大。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:96
|
||||
msgid "Add a New EPLB Policy"
|
||||
msgstr "添加新的 EPLB 策略"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:98
|
||||
msgid ""
|
||||
"If you want to add a new eplb policy to vllm_ascend, you must follow "
|
||||
"these steps:"
|
||||
msgstr "如果你想向 vllm_ascend 添加一个新的 eplb 策略,必须遵循以下步骤:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:100
|
||||
msgid ""
|
||||
"Inherit the `EplbPolicy` abstract class of `policy_abstract.py` and "
|
||||
"override the `rebalance_experts` interface, ensuring consistent input "
|
||||
"parameters `current_expert_table`, `expert_workload` and return types "
|
||||
"`newplacement`. For example:"
|
||||
msgstr ""
|
||||
"继承 `policy_abstract.py` 中的 `EplbPolicy` 抽象类,并重写 `rebalance_experts` 接口,确保输入参数 "
|
||||
"`current_expert_table`、`expert_workload` 和返回类型 `newplacement` 保持一致。例如:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:126
|
||||
msgid ""
|
||||
"To add a new EPLB algorithm, include the policy type and its "
|
||||
"corresponding implementation class in the `PolicyFactory` of "
|
||||
"`policy_factory.py`."
|
||||
msgstr "要添加新的 EPLB 算法,请在 `policy_factory.py` 的 `PolicyFactory` 中包含策略类型及其对应的实现类。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:128
|
||||
msgid "Add a New MoE Model"
|
||||
msgstr "添加新的 MoE 模型"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:130
|
||||
msgid "**Implementation Guide for Model Integration**"
|
||||
msgstr "**模型集成实施指南**"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:132
|
||||
msgid "**Adapter File Modification**"
|
||||
msgstr "**适配器文件修改**"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:133
|
||||
msgid "Inherit or modify `vllm_ascend/eplb/adaptor/vllm_adaptor.py`"
|
||||
msgstr "继承或修改 `vllm_ascend/eplb/adaptor/vllm_adaptor.py`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:134
|
||||
msgid "Add processing logic for key parameters:"
|
||||
msgstr "为关键参数添加处理逻辑:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:135
|
||||
msgid "`num_dense_layers`"
|
||||
msgstr "`num_dense_layers`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:136
|
||||
msgid "`global_expert_num`"
|
||||
msgstr "`global_expert_num`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:137
|
||||
msgid "`num_roe_layers`"
|
||||
msgstr "`num_roe_layers`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:138
|
||||
msgid "Ensure parameter synchronization in the `model_register` function."
|
||||
msgstr "确保在 `model_register` 函数中进行参数同步。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:140
|
||||
msgid "For example:"
|
||||
msgstr "例如:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:142
|
||||
msgid "Modify `__init__` of `vllm_adaptor.py` to add a new moe model eplb params:"
|
||||
msgstr "修改 `vllm_adaptor.py` 的 `__init__` 以添加新 MoE 模型的 eplb 参数:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:150
|
||||
msgid ""
|
||||
"Modify `model_register` of `vllm_adaptor.py` to register eplb params for "
|
||||
"new moe model:"
|
||||
msgstr "修改 `vllm_adaptor.py` 的 `model_register` 以注册新 MoE 模型的 eplb 参数:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:157
|
||||
msgid "**MoE Feature Integration**"
|
||||
msgstr "**MoE 功能集成**"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:158
|
||||
msgid "Extend `vllm_ascend/eplb/utils.py` with MoE-specific methods"
|
||||
msgstr "使用 MoE 特定方法扩展 `vllm_ascend/eplb/utils.py`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:159
|
||||
msgid "Implement required functionality for expert routing or weight management"
|
||||
msgstr "实现专家路由或权重管理所需的功能"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:161
|
||||
msgid "**Registration Logic Update**"
|
||||
msgstr "**注册逻辑更新**"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:162
|
||||
msgid "Add patch logic within the `model_register` function"
|
||||
msgstr "在 `model_register` 函数内添加补丁逻辑"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:163
|
||||
msgid "Maintain backward compatibility with existing model types"
|
||||
msgstr "保持与现有模型类型的向后兼容性"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:165
|
||||
msgid "**Validation & Testing**"
|
||||
msgstr "**验证与测试**"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:166
|
||||
msgid "Verify parameter consistency across layers"
|
||||
msgstr "验证跨层的参数一致性"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:167
|
||||
msgid "Test cross-device communication for expert tables"
|
||||
msgstr "测试专家表的跨设备通信"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:168
|
||||
msgid "Benchmark against baseline implementations (e.g., Qwen3-MoE)"
|
||||
msgstr "与基线实现(例如 Qwen3-MoE)进行基准测试"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:170
|
||||
msgid "*Key Implementation Notes:*"
|
||||
msgstr "*关键实施说明:*"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:172
|
||||
msgid "Preserve existing interface contracts in abstract classes"
|
||||
msgstr "在抽象类中保留现有的接口契约"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:173
|
||||
msgid "Use decorators for non-intrusive patch integration"
|
||||
msgstr "使用装饰器进行非侵入式补丁集成"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:174
|
||||
msgid "Leverage `eplb_utils.py` for shared expert mapping operations"
|
||||
msgstr "利用 `eplb_utils.py` 进行共享的专家映射操作"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:176
|
||||
msgid "DFX"
|
||||
msgstr "DFX"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:178
|
||||
msgid "Parameter Validation"
|
||||
msgstr "参数验证"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:180
|
||||
msgid "Integer Parameters"
|
||||
msgstr "整数参数"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:182
|
||||
msgid ""
|
||||
"All integer input parameters must explicitly specify their maximum and "
|
||||
"minimum values and be subject to valid value validation. For example, "
|
||||
"`expert_heat_collection_interval` must be greater than 0:"
|
||||
msgstr ""
|
||||
"所有整型输入参数必须明确指定其最大值和最小值,并接受有效值验证。例如,`expert_heat_collection_interval` 必须大于0:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:197
|
||||
msgid "File Path"
|
||||
msgstr "文件路径"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:199
|
||||
msgid ""
|
||||
"The file path for EPLB must be checked for legality, such as whether the "
|
||||
"file path is valid and whether it has appropriate read and write "
|
||||
"permissions. For example:"
|
||||
msgstr "必须检查 EPLB 文件路径的合法性,例如文件路径是否有效以及是否具有适当的读写权限。例如:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:225
|
||||
msgid "Function Specifications"
|
||||
msgstr "功能规范"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:227
|
||||
msgid "Initialization Function"
|
||||
msgstr "初始化函数"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:229
|
||||
msgid ""
|
||||
"All EPLB parameters must be initialized by default during initialization,"
|
||||
" with specified parameter types and default values for proper handling."
|
||||
msgstr "所有 EPLB 参数在初始化期间必须默认初始化,并指定参数类型和默认值以便正确处理。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:231
|
||||
msgid "General Functions"
|
||||
msgstr "通用函数"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:233
|
||||
msgid ""
|
||||
"All method arguments must specify parameter types and default values, and"
|
||||
" functions must include default return value handling for default "
|
||||
"arguments. It is recommended to use `try-except` blocks to handle the "
|
||||
"function body, specifying the type of exception captured and the failure "
|
||||
"handling (e.g., logging exceptions or returning a failure status)."
|
||||
msgstr ""
|
||||
"所有方法参数必须指定参数类型和默认值,并且函数必须包含针对默认参数的默认返回值处理。建议使用 `try-except` 块来处理函数体,指定捕获的异常类型和失败处理(例如,记录异常或返回失败状态)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:235
|
||||
msgid "Consistency"
|
||||
msgstr "一致性"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:237
|
||||
msgid "Expert Map"
|
||||
msgstr "专家映射"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:239
|
||||
msgid ""
|
||||
"The expert map must be globally unique during initialization and update. "
|
||||
"In a multi-node scenario during initialization, distributed communication"
|
||||
" should be used to verify the consistency of expert maps across each "
|
||||
"rank. If they are inconsistent, the user should be notified which ranks "
|
||||
"have inconsistent maps. During the update process, if only a few layers "
|
||||
"or the expert table of a certain rank has been changed, the updated "
|
||||
"expert table must be synchronized with the EPLB's context to ensure "
|
||||
"global consistency."
|
||||
msgstr ""
|
||||
"专家映射在初始化和更新期间必须是全局唯一的。在初始化期间的多节点场景中,应使用分布式通信来验证每个 rank 上专家映射的一致性。如果不一致,应通知用户哪些 rank 的映射不一致。在更新过程中,如果只有少数层或某个 rank 的专家表被更改,则必须将更新后的专家表与 EPLB 的上下文同步,以确保全局一致性。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:242
|
||||
msgid "Expert Weight"
|
||||
msgstr "专家权重"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:244
|
||||
msgid ""
|
||||
"When updating expert weights, ensure that the memory allocated for the "
|
||||
"expert weights has been released, or that the expert (referring to the "
|
||||
"old version) is no longer in use."
|
||||
msgstr "更新专家权重时,确保为专家权重分配的内存已被释放,或者专家(指旧版本)不再被使用。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:246
|
||||
msgid "Limitations"
|
||||
msgstr "限制"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/eplb_swift_balancer.md:248
|
||||
msgid ""
|
||||
"Before using EPLB, start the script and add `export "
|
||||
"DYNAMIC_EPLB=\"true\"`. Before performing load data collection (or "
|
||||
"performance data collection), start the script and add `export "
|
||||
"EXPERT_MAP_RECORD=\"true\"`."
|
||||
msgstr ""
|
||||
"在使用 EPLB 之前,启动脚本并添加 `export DYNAMIC_EPLB=\"true\"`。在执行负载数据收集(或性能数据收集)之前,启动脚本并添加 `export EXPERT_MAP_RECORD=\"true\"`。"
|
||||
@@ -0,0 +1,220 @@
|
||||
# SOME DESCRIPTIVE TITLE.
|
||||
# Copyright (C) 2025, vllm-ascend team
|
||||
# This file is distributed under the same license as the vllm-ascend
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend \n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:1
|
||||
msgid "Npugraph_ex"
|
||||
msgstr "Npugraph_ex"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:3
|
||||
msgid "How Does It Work?"
|
||||
msgstr "工作原理"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:5
|
||||
msgid ""
|
||||
"This is an optimization based on Fx graphs, which can be considered an "
|
||||
"acceleration solution for the aclgraph mode."
|
||||
msgstr "这是一种基于 Fx 图的优化,可视为 aclgraph 模式的一种加速方案。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:7
|
||||
msgid "You can get its code [code](https://gitcode.com/Ascend/torchair)"
|
||||
msgstr "您可以在 [code](https://gitcode.com/Ascend/torchair) 获取其代码"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:9
|
||||
msgid "Default Fx Graph Optimization"
|
||||
msgstr "默认 Fx 图优化"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:11
|
||||
msgid "Fx Graph pass"
|
||||
msgstr "Fx 图处理过程"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:13
|
||||
msgid ""
|
||||
"For the intermediate nodes of the model, replace the non-in-place "
|
||||
"operators contained in the nodes with in-place operators to reduce memory"
|
||||
" movement during computation and improve performance."
|
||||
msgstr "对于模型的中间节点,将其包含的非原位运算符替换为原位运算符,以减少计算过程中的内存移动,提升性能。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:14
|
||||
msgid ""
|
||||
"For the original input parameters of the model, if they include in-place "
|
||||
"operators, Dynamo's Functionalize process will replace the in-place "
|
||||
"operators with a form of non-in-place operators + copy operators. "
|
||||
"npugraph_ex will reverse this process, restoring the in-place operators "
|
||||
"and reducing memory movement."
|
||||
msgstr "对于模型的原始输入参数,如果包含原位运算符,Dynamo 的 Functionalize 过程会将其替换为非原位运算符 + 复制运算符的形式。npugraph_ex 将逆转此过程,恢复原位运算符,减少内存移动。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:16
|
||||
msgid "Fx fusion pass"
|
||||
msgstr "Fx 融合处理过程"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:18
|
||||
msgid ""
|
||||
"npugraph_ex now provides three default operator fusion passes, and more "
|
||||
"will be added in the future."
|
||||
msgstr "npugraph_ex 目前提供三种默认的算子融合处理过程,未来将添加更多。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:20
|
||||
msgid ""
|
||||
"Operator combinations that meet the replacement rules can be replaced "
|
||||
"with the corresponding fused operators."
|
||||
msgstr "符合替换规则的算子组合可以被替换为相应的融合算子。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:22
|
||||
msgid ""
|
||||
"You can get the default [fusion pass "
|
||||
"list](https://www.hiascend.com/document/detail/zh/Pytorch/730/modthirdparty/torchairuseguide/torchair_00017.html)"
|
||||
msgstr "您可以查看默认的[融合处理过程列表](https://www.hiascend.com/document/detail/zh/Pytorch/730/modthirdparty/torchairuseguide/torchair_00017.html)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:24
|
||||
msgid "Custom fusion pass"
|
||||
msgstr "自定义融合处理过程"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:26
|
||||
msgid ""
|
||||
"Users can register a custom graph fusion pass in TorchAir to modify "
|
||||
"PyTorch FX graphs. The registration relies on the register_replacement "
|
||||
"API."
|
||||
msgstr "用户可以在 TorchAir 中注册自定义的图融合处理过程,以修改 PyTorch FX 图。注册依赖于 register_replacement API。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:28
|
||||
msgid "Below is the declaration of this API and a demo of its usage."
|
||||
msgstr "以下是该 API 的声明及其使用示例。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid "Parameter Name"
|
||||
msgstr "参数名称"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid "Input/Output"
|
||||
msgstr "输入/输出"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid "Explanation"
|
||||
msgstr "说明"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid "Is necessary"
|
||||
msgstr "是否必需"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid "search_fn"
|
||||
msgstr "search_fn"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid "Input"
|
||||
msgstr "输入"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid ""
|
||||
"This function is the operator combination or calculation logic that you "
|
||||
"want to recognize in the FX graph, such as the operator combination that "
|
||||
"needs to be fused"
|
||||
msgstr "此函数是您希望在 FX 图中识别的算子组合或计算逻辑,例如需要融合的算子组合"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid "Yes"
|
||||
msgstr "是"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid "replace_fn"
|
||||
msgstr "replace_fn"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid ""
|
||||
"When the combination corresponding to search_fn is found in the target "
|
||||
"graph, this function's computation logic will replace the original "
|
||||
"subgraph to achieve operator fusion or optimization."
|
||||
msgstr "当在目标图中找到与 search_fn 对应的组合时,此函数的计算逻辑将替换原子图,以实现算子融合或优化。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid "example_inputs"
|
||||
msgstr "example_inputs"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid ""
|
||||
"Example input tensors used to track search_fn and replace_fn. The shape "
|
||||
"and dtype of the input should match the actual scenario."
|
||||
msgstr "用于追踪 search_fn 和 replace_fn 的示例输入张量。输入的形状和数据类型应与实际场景匹配。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid "trace_fn"
|
||||
msgstr "trace_fn"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid ""
|
||||
"By default, only the forward computation graph is tracked, which is "
|
||||
"suitable for optimization during the inference phase; if training "
|
||||
"scenarios need to be supported, a function that supports backward "
|
||||
"tracking can be provided."
|
||||
msgstr "默认情况下,仅追踪前向计算图,这适用于推理阶段的优化;如果需要支持训练场景,可以提供支持反向追踪的函数。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid "No"
|
||||
msgstr "否"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid "extra_check"
|
||||
msgstr "extra_check"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid ""
|
||||
"Find the extra verification function after operator fusion. The "
|
||||
"function's input parameter must be a Match object from "
|
||||
"torch._inductor.pattern_matcher, and it is used for further custom checks"
|
||||
" on the matching result, such as checking whether the fused operators are"
|
||||
" on the same stream, checking the device type, checking the input shapes,"
|
||||
" and so on."
|
||||
msgstr "算子融合后的额外验证函数。该函数的输入参数必须是来自 torch._inductor.pattern_matcher 的 Match 对象,用于对匹配结果进行进一步的自定义检查,例如检查融合后的算子是否在同一流上、检查设备类型、检查输入形状等。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid "search_fn_pattern"
|
||||
msgstr "search_fn_pattern"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md
|
||||
msgid ""
|
||||
"A custom pattern object is generally unnecessary to provide. Its "
|
||||
"definition follows the rules of the native PyTorch MultiOutputPattern "
|
||||
"object. After passing this parameter, search_fn will no longer be used to"
|
||||
" match operator combinations; instead, this parameter will be used "
|
||||
"directly as the matching rule."
|
||||
msgstr "通常无需提供自定义模式对象。其定义遵循原生 PyTorch MultiOutputPattern 对象的规则。传入此参数后,将不再使用 search_fn 来匹配算子组合,而是直接使用此参数作为匹配规则。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:43
|
||||
msgid "Usage Example"
|
||||
msgstr "使用示例"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:97
|
||||
msgid ""
|
||||
"The default fusion pass in npugraph_ex is also implemented based on this "
|
||||
"API. You can see more examples of using this API in the vllm-ascend and "
|
||||
"npugraph_ex code repositories."
|
||||
msgstr "npugraph_ex 中的默认融合处理过程也是基于此 API 实现的。您可以在 vllm-ascend 和 npugraph_ex 代码仓库中查看更多使用此 API 的示例。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:99
|
||||
msgid "DFX"
|
||||
msgstr "DFX"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/npugraph_ex.md:101
|
||||
msgid ""
|
||||
"By reusing the TORCH_COMPILE_DEBUG environment variable from the PyTorch "
|
||||
"community, when TORCH_COMPILE_DEBUG=1 is set, it will output the FX "
|
||||
"graphs throughout the entire process."
|
||||
msgstr "通过复用 PyTorch 社区的 TORCH_COMPILE_DEBUG 环境变量,当设置 TORCH_COMPILE_DEBUG=1 时,将输出整个过程中的 FX 图。"
|
||||
@@ -4,245 +4,218 @@
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2025.
|
||||
#
|
||||
#, fuzzy
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2025-07-18 09:01+0800\n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"Generated-By: Babel 2.17.0\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:1
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:1
|
||||
msgid "Patch in vLLM Ascend"
|
||||
msgstr "在 vLLM Ascend 中的补丁"
|
||||
msgstr "vLLM Ascend 中的补丁"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:3
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:3
|
||||
msgid ""
|
||||
"vLLM Ascend is a platform plugin for vLLM. Due to the release cycle of vLLM "
|
||||
"and vLLM Ascend is different, and the hardware limitation in some case, we "
|
||||
"need to patch some code in vLLM to make it compatible with vLLM Ascend."
|
||||
"vLLM Ascend is a platform plugin for vLLM. Due to the different release "
|
||||
"cycle of vLLM and vLLM Ascend and their hardware limitations, we need to "
|
||||
"patch some code in vLLM to make it compatible with vLLM Ascend."
|
||||
msgstr ""
|
||||
"vLLM Ascend 是 vLLM 的一个平台插件。由于 vLLM 和 vLLM Ascend "
|
||||
"的发布周期不同,并且在某些情况下存在硬件限制,我们需要对 vLLM 进行一些代码补丁,以使其能够兼容 vLLM Ascend。"
|
||||
"的发布周期不同且存在硬件限制,我们需要对 vLLM 中的部分代码打补丁,以使其兼容 vLLM Ascend。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:5
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:5
|
||||
msgid ""
|
||||
"In vLLM Ascend code, we provide a patch module `vllm_ascend/patch` to "
|
||||
"address the change for vLLM."
|
||||
msgstr "在 vLLM Ascend 代码中,我们提供了一个补丁模块 `vllm_ascend/patch` 用于应对 vLLM 的变更。"
|
||||
"adapt to changes in vLLM."
|
||||
msgstr "在 vLLM Ascend 代码中,我们提供了一个补丁模块 `vllm_ascend/patch` 来适配 vLLM 的变更。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:7
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:7
|
||||
msgid "Principle"
|
||||
msgstr "原理"
|
||||
msgstr "原则"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:9
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:9
|
||||
msgid ""
|
||||
"We should keep in mind that Patch is not the best way to make vLLM Ascend "
|
||||
"compatible. It's just a temporary solution. The best way is to contribute "
|
||||
"the change to vLLM to make it compatible with vLLM Ascend originally. In "
|
||||
"vLLM Ascend, we have the basic principle for Patch strategy:"
|
||||
"We should keep in mind that Patch is not the best way to make vLLM Ascend"
|
||||
" compatible. It's just a temporary solution. The best way is to "
|
||||
"contribute the change to vLLM to make it compatible with vLLM Ascend "
|
||||
"initially. In vLLM Ascend, we have the basic principle for Patch "
|
||||
"strategy:"
|
||||
msgstr ""
|
||||
"我们需要记住,Patch 不是让 vLLM 兼容 Ascend 的最佳方式,这只是一个临时的解决方案。最好的方法是将修改贡献到 vLLM 项目中,从而让"
|
||||
" vLLM 原生支持 Ascend。对于 vLLM Ascend,我们对 Patch 策略有一个基本原则:"
|
||||
"我们需要牢记,补丁并非实现 vLLM Ascend 兼容性的最佳方式,它只是一个临时解决方案。最佳方式是将修改贡献给 vLLM,使其原生兼容 "
|
||||
"vLLM Ascend。在 vLLM Ascend 中,我们遵循以下补丁策略基本原则:"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:11
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:11
|
||||
msgid "Less is more. Please do not patch unless it's the only way currently."
|
||||
msgstr "少即是多。请不要打补丁,除非这是目前唯一的方法。"
|
||||
msgstr "少即是多。除非是当前唯一的方法,否则请不要打补丁。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:12
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:12
|
||||
msgid ""
|
||||
"Once a patch is added, it's required to describe the future plan for "
|
||||
"removing the patch."
|
||||
msgstr "一旦补丁被添加,必须说明将来移除该补丁的计划。"
|
||||
msgstr "一旦添加补丁,必须描述未来移除该补丁的计划。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:13
|
||||
msgid "Anytime, clean the patch code is welcome."
|
||||
msgstr "任何时候,欢迎清理补丁代码。"
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:13
|
||||
msgid "Anytime, cleaning the patch code is welcome."
|
||||
msgstr "随时欢迎清理补丁代码。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:15
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:15
|
||||
msgid "How it works"
|
||||
msgstr "工作原理"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:17
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:17
|
||||
msgid "In `vllm_ascend/patch`, you can see the code structure as follows:"
|
||||
msgstr "在 `vllm_ascend/patch` 目录中,你可以看到如下代码结构:"
|
||||
msgstr "在 `vllm_ascend/patch` 中,你可以看到如下代码结构:"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:33
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:29
|
||||
msgid ""
|
||||
"**platform**: The patch code in this directory is for patching the code in "
|
||||
"vLLM main process. It's called by "
|
||||
"`vllm_ascend/platform::NPUPlatform::pre_register_and_update` very early when"
|
||||
" vLLM is initialized."
|
||||
"**platform**: The patch code in this directory is for patching the code "
|
||||
"in vLLM main process. It's called by "
|
||||
"`vllm_ascend/platform::NPUPlatform::pre_register_and_update` very early "
|
||||
"when vLLM is initialized."
|
||||
msgstr ""
|
||||
"**platform**:此目录下的补丁代码用于修补 vLLM 主进程中的代码。当 vLLM 初始化时,会在很早的阶段由 "
|
||||
"**platform**:此目录中的补丁代码用于修补 vLLM 主进程中的代码。它在 vLLM 初始化早期由 "
|
||||
"`vllm_ascend/platform::NPUPlatform::pre_register_and_update` 调用。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:34
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:30
|
||||
msgid ""
|
||||
"For online mode, vLLM process calls the platform patch here "
|
||||
"`vllm/vllm/engine/arg_utils.py::AsyncEngineArgs.add_cli_args` when parsing "
|
||||
"the cli args."
|
||||
"For online mode, vLLM process calls the platform patch in "
|
||||
"`vllm/vllm/engine/arg_utils.py::AsyncEngineArgs.add_cli_args` when "
|
||||
"parsing the cli args."
|
||||
msgstr ""
|
||||
"对于在线模式,vLLM 进程在解析命令行参数时,会在 "
|
||||
"`vllm/vllm/engine/arg_utils.py::AsyncEngineArgs.add_cli_args` 这里调用平台补丁。"
|
||||
"`vllm/vllm/engine/arg_utils.py::AsyncEngineArgs.add_cli_args` 处调用平台补丁。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:35
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:31
|
||||
msgid ""
|
||||
"For offline mode, vLLM process calls the platform patch here "
|
||||
"For offline mode, vLLM process calls the platform patch in "
|
||||
"`vllm/vllm/engine/arg_utils.py::EngineArgs.create_engine_config` when "
|
||||
"parsing the input parameters."
|
||||
msgstr ""
|
||||
"对于离线模式,vLLM 进程在解析输入参数时,会在此处调用平台补丁 "
|
||||
"`vllm/vllm/engine/arg_utils.py::EngineArgs.create_engine_config`。"
|
||||
"对于离线模式,vLLM 进程在解析输入参数时,会在 "
|
||||
"`vllm/vllm/engine/arg_utils.py::EngineArgs.create_engine_config` 处调用平台补丁。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:36
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:32
|
||||
msgid ""
|
||||
"**worker**: The patch code in this directory is for patching the code in "
|
||||
"vLLM worker process. It's called by "
|
||||
"`vllm_ascend/worker/worker::NPUWorker::__init__` when the vLLM worker "
|
||||
"process is initialized."
|
||||
msgstr ""
|
||||
"**worker**:此目录中的补丁代码用于修补 vLLM worker 进程中的代码。在初始化 vLLM worker 进程时,会被 "
|
||||
"**worker**:此目录中的补丁代码用于修补 vLLM worker 进程中的代码。它在 vLLM worker 进程初始化时由 "
|
||||
"`vllm_ascend/worker/worker::NPUWorker::__init__` 调用。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:37
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:33
|
||||
msgid ""
|
||||
"For both online and offline mode, vLLM engine core process calls the worker "
|
||||
"patch here `vllm/vllm/worker/worker_base.py::WorkerWrapperBase.init_worker` "
|
||||
"when initializing the worker process."
|
||||
"For both online and offline mode, vLLM engine core process calls the "
|
||||
"worker patch in "
|
||||
"`vllm/vllm/worker/worker_base.py::WorkerWrapperBase.init_worker` when "
|
||||
"initializing the worker process."
|
||||
msgstr ""
|
||||
"无论是在线还是离线模式,vLLM 引擎核心进程在初始化 worker 进程时,都会在这里调用 worker "
|
||||
"补丁:`vllm/vllm/worker/worker_base.py::WorkerWrapperBase.init_worker`。"
|
||||
"对于在线和离线模式,vLLM 引擎核心进程在初始化 worker 进程时,会在 "
|
||||
"`vllm/vllm/worker/worker_base.py::WorkerWrapperBase.init_worker` 处调用 worker 补丁。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:39
|
||||
msgid ""
|
||||
"In both **platform** and **worker** folder, there are several patch modules."
|
||||
" They are used for patching different version of vLLM."
|
||||
msgstr "在 **platform** 和 **worker** 文件夹中都有一些补丁模块。它们用于修补不同版本的 vLLM。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:41
|
||||
msgid ""
|
||||
"`patch_0_9_2`: This module is used for patching vLLM 0.9.2. The version is "
|
||||
"always the nearest version of vLLM. Once vLLM is released, we will drop this"
|
||||
" patch module and bump to a new version. For example, `patch_0_9_2` is used "
|
||||
"for patching vLLM 0.9.2."
|
||||
msgstr ""
|
||||
"`patch_0_9_2`:此模块用于修补 vLLM 0.9.2。该版本始终对应于 vLLM 的最近版本。一旦 vLLM "
|
||||
"发布新版本,我们将移除此补丁模块并升级到新版本。例如,`patch_0_9_2` 就是用于修补 vLLM 0.9.2 的。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:42
|
||||
msgid ""
|
||||
"`patch_main`: This module is used for patching the code in vLLM main branch."
|
||||
msgstr "`patch_main`:该模块用于修补 vLLM 主分支代码。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:43
|
||||
msgid ""
|
||||
"`patch_common`: This module is used for patching both vLLM 0.9.2 and vLLM "
|
||||
"main branch."
|
||||
msgstr "`patch_common`:此模块用于同时修补 vLLM 0.9.2 版本和 vLLM 主分支。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:45
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:35
|
||||
msgid "How to write a patch"
|
||||
msgstr "如何撰写补丁"
|
||||
msgstr "如何编写补丁"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:47
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:37
|
||||
msgid ""
|
||||
"Before writing a patch, following the principle above, we should patch the "
|
||||
"least code. If it's necessary, we can patch the code in either **platform** "
|
||||
"and **worker** folder. Here is an example to patch `distributed` module in "
|
||||
"vLLM."
|
||||
"Before writing a patch, following the principle above, we should patch "
|
||||
"the least code. If it's necessary, we can patch the code in either "
|
||||
"**platform** or **worker** folder. Here is an example to patch "
|
||||
"`distributed` module in vLLM."
|
||||
msgstr ""
|
||||
"在编写补丁之前,遵循上述原则,我们应尽量修改最少的代码。如果有必要,我们可以修改 **platform** 和 **worker** "
|
||||
"文件夹中的代码。下面是一个在 vLLM 中修改 `distributed` 模块的示例。"
|
||||
"在编写补丁前,遵循上述原则,我们应尽可能少地修改代码。如果确有必要,我们可以在 **platform** 或 **worker** "
|
||||
"文件夹中打补丁。以下是一个修补 vLLM 中 `distributed` 模块的示例。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:49
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:39
|
||||
msgid ""
|
||||
"Decide which version of vLLM we should patch. For example, after analysis, "
|
||||
"here we want to patch both 0.9.2 and main of vLLM."
|
||||
msgstr "决定我们应该修补哪个版本的 vLLM。例如,经过分析后,这里我们想要同时修补 vLLM 的 0.9.2 版和主分支(main)。"
|
||||
"Decide which version of vLLM we should patch. For example, after "
|
||||
"analysis, here we want to patch both `0.10.0` and `main` of vLLM."
|
||||
msgstr "确定我们需要修补哪个版本的 vLLM。例如,经过分析,这里我们想要同时修补 vLLM 的 `0.10.0` 版本和 `main` 分支。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:50
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:40
|
||||
msgid ""
|
||||
"Decide which process we should patch. For example, here `distributed` "
|
||||
"belongs to the vLLM main process, so we should patch `platform`."
|
||||
msgstr "决定我们应该修补哪个进程。例如,这里 `distributed` 属于 vLLM 主进程,所以我们应该修补 `platform`。"
|
||||
msgstr "确定我们需要修补哪个进程。例如,这里的 `distributed` 属于 vLLM 主进程,因此我们应该修补 `platform`。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:51
|
||||
#, python-brace-format
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:41
|
||||
msgid ""
|
||||
"Create the patch file in the right folder. The file should be named as "
|
||||
"`patch_{module_name}.py`. The example here is "
|
||||
"`vllm_ascend/patch/platform/patch_common/patch_distributed.py`."
|
||||
"`vllm_ascend/patch/platform/patch_distributed.py`."
|
||||
msgstr ""
|
||||
"在正确的文件夹中创建补丁文件。文件应命名为 `patch_{module_name}.py`。此处的示例是 "
|
||||
"`vllm_ascend/patch/platform/patch_common/patch_distributed.py`。"
|
||||
"`vllm_ascend/patch/platform/patch_distributed.py`。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:52
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:42
|
||||
msgid "Write your patch code in the new file. Here is an example:"
|
||||
msgstr "在新文件中编写你的补丁代码。以下是一个示例:"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:62
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:54
|
||||
msgid ""
|
||||
"Import the patch file in `__init__.py`. In this example, add `import "
|
||||
"vllm_ascend.patch.platform.patch_common.patch_distributed` into "
|
||||
"`vllm_ascend/patch/platform/patch_common/__init__.py`."
|
||||
"vllm_ascend.patch.platform.patch_distributed` into "
|
||||
"`vllm_ascend/patch/platform/__init__.py`."
|
||||
msgstr ""
|
||||
"在 `__init__.py` 中导入补丁文件。在这个示例中,将 `import "
|
||||
"vllm_ascend.patch.platform.patch_common.patch_distributed` 添加到 "
|
||||
"`vllm_ascend/patch/platform/patch_common/__init__.py` 中。"
|
||||
"在 `__init__.py` 中导入补丁文件。在此示例中,将 `import "
|
||||
"vllm_ascend.patch.platform.patch_distributed` 添加到 `vllm_ascend/patch/platform/__init__.py` 中。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:63
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:55
|
||||
msgid ""
|
||||
"Add the description of the patch in `vllm_ascend/patch/__init__.py`. The "
|
||||
"description format is as follows:"
|
||||
msgstr "在 `vllm_ascend/patch/__init__.py` 中添加补丁的描述。描述格式如下:"
|
||||
msgstr "在 `vllm_ascend/patch/__init__.py` 中添加补丁描述。描述格式如下:"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:77
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:71
|
||||
msgid ""
|
||||
"Add the Unit Test and E2E Test. Any newly added code in vLLM Ascend should "
|
||||
"contain the Unit Test and E2E Test as well. You can find more details in "
|
||||
"[test guide](../contribution/testing.md)"
|
||||
"Add the Unit Test and E2E Test. Any newly added code in vLLM Ascend "
|
||||
"should contain the Unit Test and E2E Test as well. You can find more "
|
||||
"details in [test guide](../contribution/testing.md)"
|
||||
msgstr ""
|
||||
"添加单元测试和端到端(E2E)测试。在 vLLM Ascend 中新增的任何代码也应包含单元测试和端到端测试。更多详情请参见 "
|
||||
"[测试指南](../contribution/testing.md)。"
|
||||
"添加单元测试和端到端测试。vLLM Ascend 中任何新增的代码都应包含单元测试和端到端测试。更多详情请参阅 [测试指南]"
|
||||
"(../contribution/testing.md)。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:80
|
||||
msgid "Limitation"
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:73
|
||||
msgid "Limitations"
|
||||
msgstr "限制"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:81
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:75
|
||||
msgid ""
|
||||
"In V1 Engine, vLLM starts three kinds of process: Main process, EngineCore "
|
||||
"process and Worker process. Now vLLM Ascend only support patch the code in "
|
||||
"Main process and Worker process by default. If you want to patch the code "
|
||||
"runs in EngineCore process, you should patch EngineCore process entirely "
|
||||
"during setup, the entry code is here `vllm.v1.engine.core`. Please override "
|
||||
"`EngineCoreProc` and `DPEngineCoreProc` entirely."
|
||||
"In V1 Engine, vLLM starts three kinds of processes: Main process, "
|
||||
"EngineCore process and Worker process. Now vLLM Ascend can only patch the"
|
||||
" code in Main process and Worker process by default. If you want to patch"
|
||||
" the code running in EngineCore process, you should patch EngineCore "
|
||||
"process entirely during setup. Find the entire code in "
|
||||
"`vllm.v1.engine.core`. Please override `EngineCoreProc` and "
|
||||
"`DPEngineCoreProc` entirely."
|
||||
msgstr ""
|
||||
"在 V1 引擎中,vLLM 会启动三种类型的进程:主进程、EngineCore 进程和 Worker 进程。现在 vLLM Ascend "
|
||||
"默认只支持在主进程和 Worker 进程中打补丁代码。如果你想要在 EngineCore 进程中打补丁,你需要在设置阶段对 EngineCore "
|
||||
"进程整体打补丁,入口代码在 `vllm.v1.engine.core`。请完全重写 `EngineCoreProc` 和 "
|
||||
"`DPEngineCoreProc`。"
|
||||
"在 V1 引擎中,vLLM 启动三种进程:主进程、EngineCore 进程和 Worker 进程。目前 vLLM Ascend "
|
||||
"默认只能修补主进程和 Worker 进程中的代码。如果你想修补 EngineCore 进程中运行的代码,你需要在设置阶段完全修补 EngineCore "
|
||||
"进程。相关完整代码位于 `vllm.v1.engine.core`。请完全重写 `EngineCoreProc` 和 `DPEngineCoreProc`。"
|
||||
|
||||
#: ../../developer_guide/Design_Documents/patch.md:82
|
||||
#: ../../source/developer_guide/Design_Documents/patch.md:76
|
||||
msgid ""
|
||||
"If you are running an edited vLLM code, the version of the vLLM may be "
|
||||
"changed automatically. For example, if you runs an edited vLLM based on "
|
||||
"v0.9.n, the version of vLLM may be change to v0.9.nxxx, in this case, the "
|
||||
"patch for v0.9.n in vLLM Ascend would not work as expect, because that vLLM "
|
||||
"Ascend can't distinguish the version of vLLM you're using. In this case, you"
|
||||
" can set the environment variable `VLLM_VERSION` to specify the version of "
|
||||
"vLLM you're using, then the patch for v0.9.2 should work."
|
||||
"If you are running edited vLLM code, the version of vLLM may be changed "
|
||||
"automatically. For example, if you run the edited vLLM based on v0.9.n, "
|
||||
"the version of vLLM may be changed to v0.9.nxxx. In this case, the patch "
|
||||
"for v0.9.n in vLLM Ascend would not work as expected, because vLLM Ascend"
|
||||
" can't distinguish the version of the vLLM you're using. In this case, "
|
||||
"you can set the environment variable `VLLM_VERSION` to specify the "
|
||||
"version of the vLLM you're using, and then the patch for v0.10.0 should "
|
||||
"work."
|
||||
msgstr ""
|
||||
"如果你运行的是经过编辑的 vLLM 代码,vLLM 的版本可能会被自动更改。例如,如果你基于 v0.9.n 运行了编辑后的 vLLM,vLLM "
|
||||
"的版本可能会变为 v0.9.nxxx,在这种情况下,vLLM Ascend 的 v0.9.n 补丁将无法正常工作,因为 vLLM Ascend "
|
||||
"无法区分你所使用的 vLLM 版本。这时,你可以设置环境变量 `VLLM_VERSION` 来指定你所使用的 vLLM 版本,这样对 v0.9.2 "
|
||||
"的补丁就应该可以正常工作。"
|
||||
"如果你运行的是经过编辑的 vLLM 代码,vLLM 的版本可能会自动更改。例如,如果你基于 v0.9.n 运行编辑后的 vLLM,vLLM "
|
||||
"的版本可能会变为 v0.9.nxxx。在这种情况下,vLLM Ascend 中针对 v0.9.n 的补丁将无法按预期工作,因为 vLLM Ascend "
|
||||
"无法区分你正在使用的 vLLM 版本。此时,你可以设置环境变量 `VLLM_VERSION` 来指定你使用的 vLLM 版本,这样针对 v0.10.0 "
|
||||
"的补丁就应该能正常工作了。"
|
||||
@@ -0,0 +1,359 @@
|
||||
# SOME DESCRIPTIVE TITLE.
|
||||
# Copyright (C) 2025, vllm-ascend team
|
||||
# This file is distributed under the same license as the vllm-ascend
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend \n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:1
|
||||
msgid "Quantization Adaptation Guide"
|
||||
msgstr "量化适配指南"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:3
|
||||
msgid ""
|
||||
"This document provides guidance for adapting quantization algorithms and "
|
||||
"models related to **ModelSlim**."
|
||||
msgstr "本文档为适配与 **ModelSlim** 相关的量化算法和模型提供指导。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:5
|
||||
msgid "Quantization Feature Introduction"
|
||||
msgstr "量化特性介绍"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:7
|
||||
msgid "Quantization Inference Process"
|
||||
msgstr "量化推理流程"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:9
|
||||
msgid ""
|
||||
"The current process for registering and obtaining quantization methods in"
|
||||
" vLLM Ascend is as follows:"
|
||||
msgstr "当前 vLLM Ascend 中注册和获取量化方法的流程如下:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:11
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:11
|
||||
msgid "get_quant_method"
|
||||
msgstr "get_quant_method"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:13
|
||||
msgid ""
|
||||
"vLLM Ascend registers a custom Ascend quantization method. By configuring"
|
||||
" the `--quantization ascend` parameter (or `quantization=\"ascend\"` for "
|
||||
"offline), the quantization feature is enabled. When constructing the "
|
||||
"`quant_config`, the registered `AscendModelSlimConfig` is initialized and"
|
||||
" `get_quant_method` is called to obtain the quantization method "
|
||||
"corresponding to each weight part, stored in the `quant_method` "
|
||||
"attribute."
|
||||
msgstr "vLLM Ascend 注册了一个自定义的 Ascend 量化方法。通过配置 `--quantization ascend` 参数(或离线时使用 `quantization=\"ascend\"`),即可启用量化功能。在构建 `quant_config` 时,会初始化已注册的 `AscendModelSlimConfig`,并调用 `get_quant_method` 来获取每个权重部分对应的量化方法,存储在 `quant_method` 属性中。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:15
|
||||
msgid ""
|
||||
"Currently supported quantization methods include `AscendLinearMethod`, "
|
||||
"`AscendFusedMoEMethod`, `AscendEmbeddingMethod`, and their corresponding "
|
||||
"non-quantized methods:"
|
||||
msgstr "当前支持的量化方法包括 `AscendLinearMethod`、`AscendFusedMoEMethod`、`AscendEmbeddingMethod` 及其对应的非量化方法:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:17
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:17
|
||||
msgid "quant_methods_overview"
|
||||
msgstr "quant_methods_overview"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:19
|
||||
msgid ""
|
||||
"The quantization method base class defined by vLLM and the overall call "
|
||||
"flow of quantization methods are as follows:"
|
||||
msgstr "vLLM 定义的量化方法基类及量化方法的整体调用流程如下:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:21
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:21
|
||||
msgid "quant_method_call_flow"
|
||||
msgstr "quant_method_call_flow"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:23
|
||||
msgid ""
|
||||
"The `embedding` method is generally not implemented for quantization, "
|
||||
"focusing only on the other three methods."
|
||||
msgstr "`embedding` 方法通常不实现量化,仅关注其他三种方法。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:25
|
||||
msgid ""
|
||||
"The `create_weights` method is used for weight initialization; the "
|
||||
"`process_weights_after_loading` method is used for weight post-"
|
||||
"processing, such as transposition, format conversion, data type "
|
||||
"conversion, etc.; the `apply` method is used to perform activation "
|
||||
"quantization and quantized matrix multiplication calculations during the "
|
||||
"forward process."
|
||||
msgstr "`create_weights` 方法用于权重初始化;`process_weights_after_loading` 方法用于权重后处理,例如转置、格式转换、数据类型转换等;`apply` 方法用于在前向传播过程中执行激活量化和量化矩阵乘法计算。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:27
|
||||
msgid ""
|
||||
"We need to implement the `create_weights`, "
|
||||
"`process_weights_after_loading`, and `apply` methods for different "
|
||||
"**layers** (**attention**, **mlp**, **moe**)."
|
||||
msgstr "我们需要为不同的**层**(**attention**、**mlp**、**moe**)实现 `create_weights`、`process_weights_after_loading` 和 `apply` 方法。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:29
|
||||
msgid ""
|
||||
"**Supplement**: When loading the model, the quantized model's description"
|
||||
" file **quant_model_description.json** needs to be read. This file "
|
||||
"describes the quantization configuration and parameters for each part of "
|
||||
"the model weights, for example:"
|
||||
msgstr "**补充说明**:加载模型时,需要读取量化模型的描述文件 **quant_model_description.json**。该文件描述了模型各部分权重的量化配置和参数,例如:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:49
|
||||
msgid ""
|
||||
"Based on the above content, we present a brief description of the "
|
||||
"adaptation process for quantization algorithms and quantized models."
|
||||
msgstr "基于以上内容,我们对量化算法和量化模型的适配过程进行简要描述。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:51
|
||||
msgid "Quantization Algorithm Adaptation"
|
||||
msgstr "量化算法适配"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:53
|
||||
msgid ""
|
||||
"**Step 1: Algorithm Design**. Define the algorithm ID (e.g., "
|
||||
"`W4A8_DYNAMIC`), determine supported layers (linear, moe, attention), and"
|
||||
" design the quantization scheme (static/dynamic, "
|
||||
"pertensor/perchannel/pergroup)."
|
||||
msgstr "**步骤 1:算法设计**。定义算法 ID(例如 `W4A8_DYNAMIC`),确定支持的层(linear、moe、attention),并设计量化方案(静态/动态、pertensor/perchannel/pergroup)。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:54
|
||||
msgid ""
|
||||
"**Step 2: Registration**. Use the `@register_scheme` decorator in "
|
||||
"`vllm_ascend/quantization/methods/registry.py` to register your "
|
||||
"quantization scheme class."
|
||||
msgstr "**步骤 2:注册**。在 `vllm_ascend/quantization/methods/registry.py` 中使用 `@register_scheme` 装饰器注册您的量化方案类。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:68
|
||||
msgid ""
|
||||
"**Step 3: Implementation**. Create an algorithm implementation file, such"
|
||||
" as `vllm_ascend/quantization/methods/w4a8.py`, and implement the method "
|
||||
"class and logic."
|
||||
msgstr "**步骤 3:实现**。创建一个算法实现文件,例如 `vllm_ascend/quantization/methods/w4a8.py`,并实现方法类和逻辑。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:69
|
||||
msgid ""
|
||||
"**Step 4: Testing**. Use your algorithm to generate quantization "
|
||||
"configurations and verify correctness and performance on target models "
|
||||
"and hardware."
|
||||
msgstr "**步骤 4:测试**。使用您的算法生成量化配置,并在目标模型和硬件上验证正确性和性能。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:71
|
||||
msgid "Quantized Model Adaptation"
|
||||
msgstr "量化模型适配"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:73
|
||||
msgid ""
|
||||
"Adapting a new quantized model requires ensuring the following three "
|
||||
"points:"
|
||||
msgstr "适配一个新的量化模型需要确保以下三点:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:75
|
||||
msgid "The original model has been successfully adapted in `vLLM Ascend`."
|
||||
msgstr "原始模型已在 `vLLM Ascend` 中成功适配。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:76
|
||||
msgid ""
|
||||
"**Fused Module Mapping**: Add the model's `model_type` to "
|
||||
"`packed_modules_model_mapping` in "
|
||||
"`vllm_ascend/quantization/modelslim_config.py` (e.g., `qkv_proj`, "
|
||||
"`gate_up_proj`, `experts`) to ensure sharding consistency and correct "
|
||||
"loading."
|
||||
msgstr "**融合模块映射**:将模型的 `model_type` 添加到 `vllm_ascend/quantization/modelslim_config.py` 中的 `packed_modules_model_mapping`(例如 `qkv_proj`、`gate_up_proj`、`experts`),以确保分片一致性和正确加载。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:96
|
||||
msgid ""
|
||||
"All quantization algorithms used by the quantized model have been "
|
||||
"integrated into the `quantization` module."
|
||||
msgstr "量化模型使用的所有量化算法都已集成到 `quantization` 模块中。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:98
|
||||
msgid "Currently Supported Quantization Algorithms"
|
||||
msgstr "当前支持的量化算法"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:100
|
||||
msgid ""
|
||||
"vLLM Ascend supports multiple quantization algorithms. The following "
|
||||
"table provides an overview of each quantization algorithm based on the "
|
||||
"implementation in the `vllm_ascend.quantization` module:"
|
||||
msgstr "vLLM Ascend 支持多种量化算法。下表基于 `vllm_ascend.quantization` 模块中的实现,概述了每种量化算法:"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Algorithm"
|
||||
msgstr "算法"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Weight"
|
||||
msgstr "权重"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Activation"
|
||||
msgstr "激活"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Weight Granularity"
|
||||
msgstr "权重粒度"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Activation Granularity"
|
||||
msgstr "激活粒度"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Type"
|
||||
msgstr "类型"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Description"
|
||||
msgstr "描述"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "`W4A16`"
|
||||
msgstr "`W4A16`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "INT4"
|
||||
msgstr "INT4"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "FP16/BF16"
|
||||
msgstr "FP16/BF16"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Per-Group"
|
||||
msgstr "Per-Group"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Per-Tensor"
|
||||
msgstr "Per-Tensor"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Static"
|
||||
msgstr "静态"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid ""
|
||||
"4-bit weight quantization with 16-bit activation precision, specifically "
|
||||
"designed for MoE model expert layers, supporting int32 format weight "
|
||||
"packing"
|
||||
msgstr "4位权重量化,16位激活精度,专为 MoE 模型专家层设计,支持 int32 格式权重打包"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "`W8A16`"
|
||||
msgstr "`W8A16`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "INT8"
|
||||
msgstr "INT8"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Per-Channel"
|
||||
msgstr "Per-Channel"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid ""
|
||||
"8-bit weight quantization with 16-bit activation precision, balancing "
|
||||
"accuracy and performance, suitable for linear layers"
|
||||
msgstr "8位权重量化,16位激活精度,平衡精度与性能,适用于线性层"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "`W8A8`"
|
||||
msgstr "`W8A8`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid ""
|
||||
"Static activation quantization, suitable for scenarios requiring high "
|
||||
"precision"
|
||||
msgstr "静态激活量化,适用于需要高精度的场景"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "`W8A8_DYNAMIC`"
|
||||
msgstr "`W8A8_DYNAMIC`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Per-Token"
|
||||
msgstr "Per-Token"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Dynamic"
|
||||
msgstr "动态"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Dynamic activation quantization with per-token scaling factor calculation"
|
||||
msgstr "动态激活量化,按 token 计算缩放因子"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "`W4A8_DYNAMIC`"
|
||||
msgstr "`W4A8_DYNAMIC`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid ""
|
||||
"Supports both direct per-channel quantization to 4-bit and two-step "
|
||||
"quantization (per-channel to 8-bit then per-group to 4-bit)"
|
||||
msgstr "支持直接按通道量化到4位,以及两步量化(先按通道量化到8位,再按组量化到4位)"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "`W4A4_FLATQUANT_DYNAMIC`"
|
||||
msgstr "`W4A4_FLATQUANT_DYNAMIC`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid ""
|
||||
"Uses FlatQuant for activation distribution smoothing before 4-bit dynamic"
|
||||
" quantization, with additional matrix multiplications for precision "
|
||||
"preservation"
|
||||
msgstr "在4位动态量化前使用 FlatQuant 平滑激活分布,并通过额外的矩阵乘法来保持精度"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "`W8A8_MIX`"
|
||||
msgstr "`W8A8_MIX`"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Per-Tensor/Token"
|
||||
msgstr "Per-Tensor/Token"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid "Mixed"
|
||||
msgstr "混合"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md
|
||||
msgid ""
|
||||
"PD Colocation Scenario uses dynamic quantization for both P node and D "
|
||||
"node; PD Disaggregation Scenario uses dynamic quantization for P node and"
|
||||
" static for D node"
|
||||
msgstr "PD 共部署场景下,P节点和D节点均使用动态量化;PD 分离部署场景下,P节点使用动态量化,D节点使用静态量化"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:112
|
||||
msgid ""
|
||||
"**Static vs Dynamic:** Static quantization uses pre-computed scaling "
|
||||
"factors with better performance, while dynamic quantization computes "
|
||||
"scaling factors on-the-fly for each token/activation tensor with higher "
|
||||
"precision."
|
||||
msgstr "**静态与动态:** 静态量化使用预计算的缩放因子,性能更好;而动态量化则为每个 token/激活张量实时计算缩放因子,精度更高。"
|
||||
|
||||
#: ../../source/developer_guide/Design_Documents/quantization.md:114
|
||||
msgid ""
|
||||
"**Granularity:** Refers to the scope of scaling factor computation (e.g.,"
|
||||
" per-tensor, per-channel, per-group)."
|
||||
msgstr "**粒度:** 指缩放因子计算的范围(例如,per-tensor、per-channel、per-group)。"
|
||||
@@ -4,184 +4,178 @@
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2025.
|
||||
#
|
||||
#, fuzzy
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2025-07-18 09:01+0800\n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"Generated-By: Babel 2.17.0\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:107
|
||||
#: ../../source/developer_guide/contribution/index.md:108
|
||||
msgid "Index"
|
||||
msgstr "索引"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:1
|
||||
#: ../../source/developer_guide/contribution/index.md:1
|
||||
msgid "Contributing"
|
||||
msgstr "贡献"
|
||||
msgstr "贡献指南"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:3
|
||||
msgid "Building and testing"
|
||||
#: ../../source/developer_guide/contribution/index.md:3
|
||||
msgid "Building and Testing"
|
||||
msgstr "构建与测试"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:4
|
||||
#: ../../source/developer_guide/contribution/index.md:5
|
||||
msgid ""
|
||||
"It's recommended to set up a local development environment to build and test"
|
||||
" before you submit a PR."
|
||||
msgstr "建议先搭建本地开发环境来进行构建和测试,再提交 PR。"
|
||||
"It's recommended to set up a local development environment to build vllm-"
|
||||
"ascend and run tests before you submit a PR."
|
||||
msgstr "建议在提交 PR 之前,先搭建本地开发环境来构建 vllm-ascend 并运行测试。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:7
|
||||
msgid "Setup development environment"
|
||||
msgstr "搭建开发环境"
|
||||
#: ../../source/developer_guide/contribution/index.md:8
|
||||
msgid "Set up a development environment"
|
||||
msgstr "设置开发环境"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:9
|
||||
#: ../../source/developer_guide/contribution/index.md:10
|
||||
msgid ""
|
||||
"Theoretically, the vllm-ascend build is only supported on Linux because "
|
||||
"`vllm-ascend` dependency `torch_npu` only supports Linux."
|
||||
msgstr ""
|
||||
"理论上,vllm-ascend 构建仅支持 Linux,因为 `vllm-ascend` 的依赖项 `torch_npu` 只支持 Linux。"
|
||||
msgstr "理论上,vllm-ascend 的构建仅支持 Linux,因为其依赖项 `torch_npu` 仅支持 Linux。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:12
|
||||
#: ../../source/developer_guide/contribution/index.md:13
|
||||
msgid ""
|
||||
"But you can still set up dev env on Linux/Windows/macOS for linting and "
|
||||
"basic test as following commands:"
|
||||
msgstr "但你仍然可以在 Linux/Windows/macOS 上按照以下命令设置开发环境,用于代码规约检查和基本测试:"
|
||||
"But you can still set up a development environment on Linux/Windows/macOS"
|
||||
" for linting and running basic tests."
|
||||
msgstr "但你仍然可以在 Linux/Windows/macOS 上设置开发环境,用于代码规范检查和运行基本测试。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:15
|
||||
#: ../../source/developer_guide/contribution/index.md:16
|
||||
msgid "Run lint locally"
|
||||
msgstr "在本地运行 lint"
|
||||
msgstr "本地运行代码检查"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:33
|
||||
#: ../../source/developer_guide/contribution/index.md:35
|
||||
msgid "Run CI locally"
|
||||
msgstr "本地运行CI"
|
||||
msgstr "本地运行 CI"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:35
|
||||
msgid "After complete \"Run lint\" setup, you can run CI locally:"
|
||||
msgstr "在完成“运行 lint”设置后,你可以在本地运行 CI:"
|
||||
#: ../../source/developer_guide/contribution/index.md:37
|
||||
msgid "After completing \"Run lint\" setup, you can run CI locally:"
|
||||
msgstr "完成“运行代码检查”设置后,你可以在本地运行 CI:"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:61
|
||||
#: ../../source/developer_guide/contribution/index.md:63
|
||||
msgid "Submit the commit"
|
||||
msgstr "提交该提交"
|
||||
msgstr "提交更改"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:68
|
||||
msgid ""
|
||||
"🎉 Congratulations! You have completed the development environment setup."
|
||||
msgstr "🎉 恭喜!你已经完成了开发环境的搭建。"
|
||||
#: ../../source/developer_guide/contribution/index.md:70
|
||||
msgid "🎉 Congratulations! You have completed the development environment setup."
|
||||
msgstr "🎉 恭喜!您已完成开发环境的设置。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:70
|
||||
msgid "Test locally"
|
||||
#: ../../source/developer_guide/contribution/index.md:72
|
||||
msgid "Testing locally"
|
||||
msgstr "本地测试"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:72
|
||||
#: ../../source/developer_guide/contribution/index.md:74
|
||||
msgid ""
|
||||
"You can refer to [Testing](./testing.md) doc to help you setup testing "
|
||||
"environment and running tests locally."
|
||||
msgstr "你可以参考 [测试](./testing.md) 文档,帮助你搭建测试环境并在本地运行测试。"
|
||||
"You can refer to [Testing](./testing.md) to set up a testing environment"
|
||||
" and running tests locally."
|
||||
msgstr "你可以参考 [测试](./testing.md) 文档来设置测试环境并在本地运行测试。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:74
|
||||
#: ../../source/developer_guide/contribution/index.md:76
|
||||
msgid "DCO and Signed-off-by"
|
||||
msgstr "DCO 和签名确认"
|
||||
msgstr "DCO 与签署确认"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:76
|
||||
#: ../../source/developer_guide/contribution/index.md:78
|
||||
msgid ""
|
||||
"When contributing changes to this project, you must agree to the DCO. "
|
||||
"Commits must include a `Signed-off-by:` header which certifies agreement "
|
||||
"with the terms of the DCO."
|
||||
msgstr "当为本项目贡献更改时,您必须同意 DCO。提交必须包含 `Signed-off-by:` 头部,以证明您同意 DCO 的条款。"
|
||||
msgstr "向本项目贡献更改时,您必须同意 DCO。提交必须包含 `Signed-off-by:` 标头,以证明您同意 DCO 的条款。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:78
|
||||
#: ../../source/developer_guide/contribution/index.md:80
|
||||
msgid "Using `-s` with `git commit` will automatically add this header."
|
||||
msgstr "在使用 `git commit` 时加上 `-s` 参数会自动添加这个头部信息。"
|
||||
msgstr "在 `git commit` 命令中使用 `-s` 参数会自动添加此标头。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:80
|
||||
#: ../../source/developer_guide/contribution/index.md:82
|
||||
msgid "PR Title and Classification"
|
||||
msgstr "PR 标题与分类"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:82
|
||||
#: ../../source/developer_guide/contribution/index.md:84
|
||||
msgid ""
|
||||
"Only specific types of PRs will be reviewed. The PR title is prefixed "
|
||||
"appropriately to indicate the type of change. Please use one of the "
|
||||
"following:"
|
||||
msgstr "只有特定类型的 PR 会被审核。PR 标题应使用合适的前缀以指明更改类型。请使用以下之一:"
|
||||
msgstr "只有特定类型的 PR 会被审核。PR 标题应使用适当的前缀来指明更改类型。请使用以下前缀之一:"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:84
|
||||
#: ../../source/developer_guide/contribution/index.md:86
|
||||
msgid "`[Attention]` for new features or optimization in attention."
|
||||
msgstr "`[Attention]` 用于注意力机制中新特性或优化。"
|
||||
msgstr "`[Attention]` 用于注意力机制的新功能或优化。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:85
|
||||
#: ../../source/developer_guide/contribution/index.md:87
|
||||
msgid "`[Communicator]` for new features or optimization in communicators."
|
||||
msgstr "`[Communicator]` 适用于通信器中的新特性或优化。"
|
||||
msgstr "`[Communicator]` 用于通信器的新功能或优化。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:86
|
||||
#: ../../source/developer_guide/contribution/index.md:88
|
||||
msgid "`[ModelRunner]` for new features or optimization in model runner."
|
||||
msgstr "`[ModelRunner]` 用于模型运行器中的新功能或优化。"
|
||||
msgstr "`[ModelRunner]` 用于模型运行器的新功能或优化。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:87
|
||||
#: ../../source/developer_guide/contribution/index.md:89
|
||||
msgid "`[Platform]` for new features or optimization in platform."
|
||||
msgstr "`[Platform]` 用于平台中新功能或优化。"
|
||||
msgstr "`[Platform]` 用于平台的新功能或优化。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:88
|
||||
#: ../../source/developer_guide/contribution/index.md:90
|
||||
msgid "`[Worker]` for new features or optimization in worker."
|
||||
msgstr "`[Worker]` 用于 worker 的新功能或优化。"
|
||||
msgstr "`[Worker]` 用于工作器的新功能或优化。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:89
|
||||
#: ../../source/developer_guide/contribution/index.md:91
|
||||
msgid ""
|
||||
"`[Core]` for new features or optimization in the core vllm-ascend logic "
|
||||
"(such as platform, attention, communicators, model runner)"
|
||||
msgstr "`[Core]` 用于核心 vllm-ascend 逻辑中的新特性或优化(例如平台、注意力机制、通信器、模型运行器)。"
|
||||
msgstr "`[Core]` 用于核心 vllm-ascend 逻辑中的新功能或优化(例如平台、注意力机制、通信器、模型运行器)。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:90
|
||||
msgid "`[Kernel]` changes affecting compute kernels and ops."
|
||||
msgstr "`[Kernel]` 影响计算内核和操作的更改。"
|
||||
#: ../../source/developer_guide/contribution/index.md:92
|
||||
msgid "`[Kernel]` for changes affecting compute kernels and ops."
|
||||
msgstr "`[Kernel]` 用于影响计算内核和操作的更改。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:91
|
||||
#: ../../source/developer_guide/contribution/index.md:93
|
||||
msgid "`[Bugfix]` for bug fixes."
|
||||
msgstr "`[Bugfix]` 用于表示错误修复。"
|
||||
msgstr "`[Bugfix]` 用于错误修复。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:92
|
||||
#: ../../source/developer_guide/contribution/index.md:94
|
||||
msgid "`[Doc]` for documentation fixes and improvements."
|
||||
msgstr "`[Doc]` 用于文档修复和改进。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:93
|
||||
#: ../../source/developer_guide/contribution/index.md:95
|
||||
msgid "`[Test]` for tests (such as unit tests)."
|
||||
msgstr "`[Test]` 用于测试(如单元测试)。"
|
||||
msgstr "`[Test]` 用于测试(例如单元测试)。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:94
|
||||
#: ../../source/developer_guide/contribution/index.md:96
|
||||
msgid "`[CI]` for build or continuous integration improvements."
|
||||
msgstr "`[CI]` 用于构建或持续集成的改进。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:95
|
||||
#: ../../source/developer_guide/contribution/index.md:97
|
||||
msgid ""
|
||||
"`[Misc]` for PRs that do not fit the above categories. Please use this "
|
||||
"sparingly."
|
||||
msgstr "对于不属于上述类别的 PR,请使用 `[Misc]`。请谨慎使用此标签。"
|
||||
msgstr "`[Misc]` 用于不属于上述类别的 PR。请谨慎使用此标签。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:98
|
||||
#: ../../source/developer_guide/contribution/index.md:100
|
||||
msgid ""
|
||||
"If the PR spans more than one category, please include all relevant "
|
||||
"prefixes."
|
||||
msgstr "如果拉取请求(PR)涵盖多个类别,请包含所有相关的前缀。"
|
||||
msgstr "如果 PR 涉及多个类别,请包含所有相关的前缀。"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:101
|
||||
#: ../../source/developer_guide/contribution/index.md:103
|
||||
msgid "Others"
|
||||
msgstr "其他"
|
||||
|
||||
#: ../../developer_guide/contribution/index.md:103
|
||||
#: ../../source/developer_guide/contribution/index.md:105
|
||||
msgid ""
|
||||
"You may find more information about contributing to vLLM Ascend backend "
|
||||
"plugin on "
|
||||
"[<u>docs.vllm.ai</u>](https://docs.vllm.ai/en/latest/contributing/overview.html)."
|
||||
" If you find any problem when contributing, you can feel free to submit a PR"
|
||||
" to improve the doc to help other developers."
|
||||
msgstr ""
|
||||
"你可以在 "
|
||||
"[<u>docs.vllm.ai</u>](https://docs.vllm.ai/en/latest/contributing/overview.html)"
|
||||
" 上找到有关为 vLLM Ascend 后端插件做贡献的更多信息。如果你在贡献过程中遇到任何问题,欢迎随时提交 PR 来改进文档,以帮助其他开发者。"
|
||||
"[<u>docs.vllm.ai</u>](https://docs.vllm.ai/en/latest/contributing). If "
|
||||
"you encounter any problems while contributing, feel free to submit a PR "
|
||||
"to improve the documentation to help other developers."
|
||||
msgstr "你可以在 [<u>docs.vllm.ai</u>](https://docs.vllm.ai/en/latest/contributing) 上找到有关为 vLLM Ascend 后端插件做贡献的更多信息。如果在贡献过程中遇到任何问题,欢迎随时提交 PR 来改进文档,以帮助其他开发者。"
|
||||
@@ -0,0 +1,222 @@
|
||||
# SOME DESCRIPTIVE TITLE.
|
||||
# Copyright (C) 2025, vllm-ascend team
|
||||
# This file is distributed under the same license as the vllm-ascend
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend \n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:1
|
||||
msgid "Multi Node Test"
|
||||
msgstr "多节点测试"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:3
|
||||
msgid ""
|
||||
"Multi-Node CI is designed to test distributed scenarios of very large "
|
||||
"models, eg: disaggregated_prefill multi DP across multi nodes and so on."
|
||||
msgstr "多节点CI旨在测试超大规模模型的分布式场景,例如:跨多节点的解耦预填充(disaggregated_prefill)、多数据并行(multi DP)等。"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:5
|
||||
msgid "How it works"
|
||||
msgstr "工作原理"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:7
|
||||
msgid ""
|
||||
"The following picture shows the basic deployment view of the multi-node "
|
||||
"CI mechanism. It shows how the GitHub action interacts with "
|
||||
"[lws](https://lws.sigs.k8s.io/docs/overview/) (a kind of kubernetes crd "
|
||||
"resource)."
|
||||
msgstr "下图展示了多节点CI机制的基本部署视图。它说明了GitHub Action如何与[lws](https://lws.sigs.k8s.io/docs/overview/)(一种Kubernetes CRD资源)进行交互。"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:9
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:9
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:13
|
||||
msgid "alt text"
|
||||
msgstr "替代文本"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:11
|
||||
msgid ""
|
||||
"From the workflow perspective, we can see how the final test script is "
|
||||
"executed, The key point is that these two [lws.yaml and "
|
||||
"run.sh](https://github.com/vllm-project/vllm-"
|
||||
"ascend/tree/main/tests/e2e/nightly/multi_node/scripts), The former "
|
||||
"defines how our k8s cluster is pulled up, and the latter defines the "
|
||||
"entry script when the pod is started, Each node executes different logic "
|
||||
"according to the "
|
||||
"[LWS_WORKER_INDEX](https://lws.sigs.k8s.io/docs/reference/labels-"
|
||||
"annotations-and-environment-variables/) environment variable, so that "
|
||||
"multiple nodes can form a distributed cluster to perform tasks."
|
||||
msgstr "从工作流的角度,我们可以看到最终的测试脚本是如何执行的。关键在于这两个文件:[lws.yaml和run.sh](https://github.com/vllm-project/vllm-ascend/tree/main/tests/e2e/nightly/multi_node/scripts)。前者定义了我们的k8s集群如何被拉起,后者定义了Pod启动时的入口脚本。每个节点根据[LWS_WORKER_INDEX](https://lws.sigs.k8s.io/docs/reference/labels-annotations-and-environment-variables/)环境变量执行不同的逻辑,从而使多个节点能够组成一个分布式集群来执行任务。"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:13
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:15
|
||||
msgid "How to contribute"
|
||||
msgstr "如何贡献"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:17
|
||||
msgid "Upload custom weights"
|
||||
msgstr "上传自定义权重"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:19
|
||||
msgid ""
|
||||
"If you need customized weights, for example, you quantized a w8a8 weight "
|
||||
"for DeepSeek-V3 and you want your weight to run on CI, uploading weights "
|
||||
"to ModelScope's [vllm-ascend](https://www.modelscope.cn/organization"
|
||||
"/vllm-ascend) organization is welcome. If you do not have permission to "
|
||||
"upload, please contact @Potabk"
|
||||
msgstr "如果您需要自定义权重,例如,您为DeepSeek-V3量化了一个w8a8权重,并希望您的权重能在CI上运行,欢迎将权重上传至ModelScope的[vllm-ascend](https://www.modelscope.cn/organization/vllm-ascend)组织。如果您没有上传权限,请联系@Potabk。"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:21
|
||||
msgid "Add config yaml"
|
||||
msgstr "添加配置YAML"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:23
|
||||
msgid ""
|
||||
"As the entrypoint script [run.sh](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/0bf3f21a987aede366ec4629ad0ffec8e32fe90d/tests/e2e/nightly/multi_node/scripts/run.sh#L106)"
|
||||
" shows, a k8s pod startup means traversing all *.yaml files in the "
|
||||
"[directory](https://github.com/vllm-project/vllm-"
|
||||
"ascend/tree/main/tests/e2e/nightly/multi_node/config/), reading and "
|
||||
"executing according to different configurations, so what we need to do is"
|
||||
" just add \"yamls\" like [DeepSeek-V3.yaml](https://github.com/vllm-"
|
||||
"project/vllm-"
|
||||
"ascend/blob/main/tests/e2e/nightly/multi_node/config/DeepSeek-V3.yaml)."
|
||||
msgstr "如入口脚本[run.sh](https://github.com/vllm-project/vllm-ascend/blob/0bf3f21a987aede366ec4629ad0ffec8e32fe90d/tests/e2e/nightly/multi_node/scripts/run.sh#L106)所示,一个k8s Pod的启动意味着遍历[目录](https://github.com/vllm-project/vllm-ascend/tree/main/tests/e2e/nightly/multi_node/config/)中的所有*.yaml文件,并根据不同的配置读取和执行。因此,我们需要做的就是添加类似[DeepSeek-V3.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/tests/e2e/nightly/multi_node/config/DeepSeek-V3.yaml)的\"yaml\"文件。"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:25
|
||||
msgid ""
|
||||
"Suppose you have **2 nodes** running a 1P1D setup (1 Prefillers + 1 "
|
||||
"Decoder):"
|
||||
msgstr "假设您有**2个节点**运行1P1D设置(1个预填充器 + 1个解码器):"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:27
|
||||
msgid "you may add a config file looks like:"
|
||||
msgstr "您可以添加一个类似这样的配置文件:"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:73
|
||||
msgid "Add the case to nightly workflow"
|
||||
msgstr "将用例添加到夜间工作流"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:75
|
||||
msgid ""
|
||||
"Currently, the multi-node test workflow is defined in the "
|
||||
"[nightly_test_a3.yaml](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/main/.github/workflows/schedule_nightly_test_a3.yaml)"
|
||||
msgstr "目前,多节点测试工作流定义在[nightly_test_a3.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/schedule_nightly_test_a3.yaml)中。"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:110
|
||||
msgid ""
|
||||
"The matrix above defines all the parameters required to add a multi-"
|
||||
"machine use case. The parameters worth noting (if you are adding a new "
|
||||
"use case) are `size` and the path to the yaml configuration file. The "
|
||||
"former defines the number of nodes required for your use case, and the "
|
||||
"latter defines the path to the configuration file you have completed in "
|
||||
"step 2."
|
||||
msgstr "上面的矩阵定义了添加一个多机用例所需的所有参数。值得注意的参数(如果您正在添加一个新用例)是`size`和yaml配置文件的路径。前者定义了您的用例所需的节点数量,后者定义了您在步骤2中完成的配置文件的路径。"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:112
|
||||
msgid "Run Multi-Node tests locally"
|
||||
msgstr "本地运行多节点测试"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:114
|
||||
msgid "1. Use kubernetes"
|
||||
msgstr "1. 使用Kubernetes"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:116
|
||||
msgid ""
|
||||
"This section assumes that you already have a "
|
||||
"[Kubernetes](https://kubernetes.io/docs/setup/) NPU cluster environment "
|
||||
"locally. Then you can easily start our test with one click."
|
||||
msgstr "本节假设您本地已经有一个[Kubernetes](https://kubernetes.io/docs/setup/) NPU集群环境。然后您可以轻松地一键启动我们的测试。"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:118
|
||||
msgid "Step 1. Install LWS CRD resources"
|
||||
msgstr "步骤 1. 安装LWS CRD资源"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:120
|
||||
msgid ""
|
||||
"See <https://lws.sigs.k8s.io/docs/installation/> Which can be used as a "
|
||||
"reference"
|
||||
msgstr "参考<https://lws.sigs.k8s.io/docs/installation/>"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:122
|
||||
msgid "Step 2. Deploy the following yaml file `lws.yaml` as what you want"
|
||||
msgstr "步骤 2. 按需部署以下yaml文件`lws.yaml`"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:258
|
||||
msgid "Verify the status of the pods:"
|
||||
msgstr "验证Pod的状态:"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:264
|
||||
msgid "Should get an output similar to this:"
|
||||
msgstr "应该会得到类似这样的输出:"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:272
|
||||
msgid "Verify that the distributed inference works:"
|
||||
msgstr "验证分布式推理是否正常工作:"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:278
|
||||
msgid "Should get something similar to this:"
|
||||
msgstr "应该会得到类似这样的结果:"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:312
|
||||
msgid "2. Test without kubernetes"
|
||||
msgstr "2. 不使用Kubernetes进行测试"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:314
|
||||
msgid ""
|
||||
"Since our script is Kubernetes-friendly, we need to actively pass in some"
|
||||
" cluster information if you don't have a Kubernetes environment."
|
||||
msgstr "由于我们的脚本对Kubernetes友好,如果您没有Kubernetes环境,则需要主动传入一些集群信息。"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:316
|
||||
msgid "Step 1. Add cluster_hosts to config yamls"
|
||||
msgstr "步骤 1. 向配置YAML文件添加cluster_hosts"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:318
|
||||
msgid ""
|
||||
"Modify on every cluster host, commands just like "
|
||||
"[DeepSeek-V3.yaml](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/e760aae1df7814073a4180172385505c1ec0fd83/tests/e2e/nightly/multi_node/config/DeepSeek-V3.yaml#L25)"
|
||||
" after the configure item `num_nodes` , for example: `cluster_hosts: "
|
||||
"[\"xxx.xxx.xxx.188\", \"xxx.xxx.xxx.212\"]`"
|
||||
msgstr "在每个集群主机上进行修改,就像[DeepSeek-V3.yaml](https://github.com/vllm-project/vllm-ascend/blob/e760aae1df7814073a4180172385505c1ec0fd83/tests/e2e/nightly/multi_node/config/DeepSeek-V3.yaml#L25)那样,在配置项`num_nodes`之后添加,例如:`cluster_hosts: [\"xxx.xxx.xxx.188\", \"xxx.xxx.xxx.212\"]`"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:321
|
||||
msgid "Step 2. Install develop environment"
|
||||
msgstr "步骤 2. 安装开发环境"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:322
|
||||
msgid "Install vllm-ascend develop packages on every cluster host"
|
||||
msgstr "在每个集群主机上安装vllm-ascend开发包"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:329
|
||||
msgid "Install AISBench on the first host(leader node) in cluster_hosts"
|
||||
msgstr "在cluster_hosts中的第一个主机(主节点)上安装AISBench"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:341
|
||||
msgid "Step 3. Running test locally"
|
||||
msgstr "步骤 3. 本地运行测试"
|
||||
|
||||
#: ../../source/developer_guide/contribution/multi_node_test.md:343
|
||||
msgid "Run the script on **each node separately**"
|
||||
msgstr "在**每个节点上分别**运行脚本"
|
||||
@@ -4,234 +4,289 @@
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2025.
|
||||
#
|
||||
#, fuzzy
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2025-07-18 09:01+0800\n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"Generated-By: Babel 2.17.0\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:1
|
||||
#: ../../source/developer_guide/contribution/testing.md:1
|
||||
msgid "Testing"
|
||||
msgstr "测试"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:3
|
||||
#: ../../source/developer_guide/contribution/testing.md:3
|
||||
msgid ""
|
||||
"This secition explains how to write e2e tests and unit tests to verify the "
|
||||
"implementation of your feature."
|
||||
msgstr "本节介绍如何编写端到端测试和单元测试,以验证你的功能实现。"
|
||||
"This document explains how to write E2E tests and unit tests to verify "
|
||||
"the implementation of your feature."
|
||||
msgstr "本文档介绍如何编写端到端测试和单元测试,以验证您实现的功能。"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:5
|
||||
msgid "Setup test environment"
|
||||
#: ../../source/developer_guide/contribution/testing.md:5
|
||||
msgid "Set up a test environment"
|
||||
msgstr "设置测试环境"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:7
|
||||
#: ../../source/developer_guide/contribution/testing.md:7
|
||||
msgid ""
|
||||
"The fastest way to setup test environment is to use the main branch "
|
||||
"The fastest way to set up a test environment is to use the main branch's "
|
||||
"container image:"
|
||||
msgstr "搭建测试环境最快的方法是使用 main 分支的容器镜像:"
|
||||
msgstr "设置测试环境最快的方法是使用 main 分支的容器镜像:"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md
|
||||
#: ../../source/developer_guide/contribution/testing.md
|
||||
msgid "Local (CPU)"
|
||||
msgstr "本地(CPU)"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:18
|
||||
msgid "You can run the unit tests on CPU with the following steps:"
|
||||
msgstr "你可以按照以下步骤在 CPU 上运行单元测试:"
|
||||
#: ../../source/developer_guide/contribution/testing.md:18
|
||||
msgid "You can run the unit tests on CPUs with the following steps:"
|
||||
msgstr "您可以按照以下步骤在 CPU 上运行单元测试:"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md
|
||||
#: ../../source/developer_guide/contribution/testing.md
|
||||
msgid "Single card"
|
||||
msgstr "单张卡片"
|
||||
msgstr "单卡"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:85
|
||||
#: ../../developer_guide/contribution/testing.md:123
|
||||
msgid ""
|
||||
"After starting the container, you should install the required packages:"
|
||||
msgstr "启动容器后,你应该安装所需的软件包:"
|
||||
#: ../../source/developer_guide/contribution/testing.md:96
|
||||
#: ../../source/developer_guide/contribution/testing.md:135
|
||||
msgid "After starting the container, you should install the required packages:"
|
||||
msgstr "启动容器后,您应该安装所需的软件包:"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md
|
||||
#: ../../source/developer_guide/contribution/testing.md
|
||||
msgid "Multi cards"
|
||||
msgstr "多卡"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:137
|
||||
#: ../../source/developer_guide/contribution/testing.md:149
|
||||
msgid "Running tests"
|
||||
msgstr "运行测试"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:139
|
||||
msgid "Unit test"
|
||||
#: ../../source/developer_guide/contribution/testing.md:151
|
||||
msgid "Unit tests"
|
||||
msgstr "单元测试"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:141
|
||||
#: ../../source/developer_guide/contribution/testing.md:153
|
||||
msgid "There are several principles to follow when writing unit tests:"
|
||||
msgstr "编写单元测试时需要遵循几个原则:"
|
||||
msgstr "编写单元测试时需要遵循以下几个原则:"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:143
|
||||
#: ../../source/developer_guide/contribution/testing.md:155
|
||||
msgid ""
|
||||
"The test file path should be consistent with source file and start with "
|
||||
"`test_` prefix, such as: `vllm_ascend/worker/worker.py` --> "
|
||||
"The test file path should be consistent with the source file and start "
|
||||
"with the `test_` prefix, such as: `vllm_ascend/worker/worker.py` --> "
|
||||
"`tests/ut/worker/test_worker.py`"
|
||||
msgstr ""
|
||||
"测试文件的路径应与源文件保持一致,并以 `test_` 前缀开头,例如:`vllm_ascend/worker/worker.py` --> "
|
||||
"测试文件路径应与源文件保持一致,并以 `test_` 前缀开头,例如:`vllm_ascend/worker/worker.py` --> "
|
||||
"`tests/ut/worker/test_worker.py`"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:144
|
||||
#: ../../source/developer_guide/contribution/testing.md:156
|
||||
msgid ""
|
||||
"The vLLM Ascend test are using unittest framework, see "
|
||||
"[here](https://docs.python.org/3/library/unittest.html#module-unittest) to "
|
||||
"understand how to write unit tests."
|
||||
"The vLLM Ascend test uses unittest framework. See [the Python unittest "
|
||||
"documentation](https://docs.python.org/3/library/unittest.html#module-"
|
||||
"unittest) to understand how to write unit tests."
|
||||
msgstr ""
|
||||
"vLLM Ascend 测试使用 unittest "
|
||||
"框架,参见[这里](https://docs.python.org/3/library/unittest.html#module-"
|
||||
"unittest)了解如何编写单元测试。"
|
||||
"vLLM Ascend 测试使用 unittest 框架。请参阅 [Python unittest "
|
||||
"文档](https://docs.python.org/3/library/unittest.html#module-unittest) 以了解如何编写单元测试。"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:145
|
||||
#: ../../source/developer_guide/contribution/testing.md:157
|
||||
msgid ""
|
||||
"All unit tests can be run on CPU, so you must mock the device-related "
|
||||
"function to host."
|
||||
msgstr "所有单元测试都可以在 CPU 上运行,因此你必须将与设备相关的函数模拟为 host。"
|
||||
"All unit tests can be run on CPUs, so you must mock the device-related "
|
||||
"functions on the host."
|
||||
msgstr "所有单元测试都可以在 CPU 上运行,因此您必须在主机上模拟与设备相关的函数。"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:146
|
||||
#: ../../source/developer_guide/contribution/testing.md:158
|
||||
msgid ""
|
||||
"Example: [tests/ut/test_ascend_config.py](https://github.com/vllm-"
|
||||
"project/vllm-ascend/blob/main/tests/ut/test_ascend_config.py)."
|
||||
"Example: [tests/ut/test_ascend_config.py](https://github.com/vllm-project"
|
||||
"/vllm-ascend/blob/main/tests/ut/test_ascend_config.py)."
|
||||
msgstr ""
|
||||
"示例:[tests/ut/test_ascend_config.py](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/main/tests/ut/test_ascend_config.py)。"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:147
|
||||
#: ../../source/developer_guide/contribution/testing.md:159
|
||||
msgid "You can run the unit tests using `pytest`:"
|
||||
msgstr "你可以使用 `pytest` 运行单元测试:"
|
||||
msgstr "您可以使用 `pytest` 运行单元测试:"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md
|
||||
msgid "Multi cards test"
|
||||
msgstr "多卡测试"
|
||||
#: ../../source/developer_guide/contribution/testing.md
|
||||
msgid "Single-card"
|
||||
msgstr "单卡"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:192
|
||||
#: ../../source/developer_guide/contribution/testing.md
|
||||
msgid "Multi-card"
|
||||
msgstr "多卡"
|
||||
|
||||
#: ../../source/developer_guide/contribution/testing.md:206
|
||||
msgid "E2E test"
|
||||
msgstr "端到端测试"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:194
|
||||
#: ../../source/developer_guide/contribution/testing.md:208
|
||||
msgid ""
|
||||
"Although vllm-ascend CI provide [e2e test](https://github.com/vllm-"
|
||||
"project/vllm-ascend/blob/main/.github/workflows/vllm_ascend_test.yaml) on "
|
||||
"Ascend CI, you can run it locally."
|
||||
"Although vllm-ascend CI provides E2E tests on Ascend CI (for example, "
|
||||
"[schedule_nightly_test_a2.yaml](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/main/.github/workflows/schedule_nightly_test_a2.yaml), "
|
||||
"[schedule_nightly_test_a3.yaml](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/main/.github/workflows/schedule_nightly_test_a3.yaml), "
|
||||
"[pr_test_full.yaml](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/main/.github/workflows/pr_test_full.yaml)), you can run them "
|
||||
"locally."
|
||||
msgstr ""
|
||||
"虽然 vllm-ascend CI 在 Ascend CI 上提供了 [端到端测试](https://github.com/vllm-"
|
||||
"project/vllm-"
|
||||
"ascend/blob/main/.github/workflows/vllm_ascend_test.yaml),你也可以在本地运行它。"
|
||||
"虽然 vllm-ascend CI 在 Ascend CI 上提供了端到端测试(例如,[schedule_nightly_test_a2.yaml](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/main/.github/workflows/schedule_nightly_test_a2.yaml)、[schedule_nightly_test_a3.yaml](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/main/.github/workflows/schedule_nightly_test_a3.yaml)、[pr_test_full.yaml](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/main/.github/workflows/pr_test_full.yaml)),但您也可以在本地运行它们。"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:204
|
||||
msgid "You can't run e2e test on CPU."
|
||||
msgstr "你无法在 CPU 上运行 e2e 测试。"
|
||||
#: ../../source/developer_guide/contribution/testing.md:218
|
||||
msgid "You can't run the E2E test on CPUs."
|
||||
msgstr "您无法在 CPU 上运行端到端测试。"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:240
|
||||
#: ../../source/developer_guide/contribution/testing.md:257
|
||||
msgid ""
|
||||
"This will reproduce e2e test: "
|
||||
"This will reproduce the E2E test. See "
|
||||
"[vllm_ascend_test.yaml](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/main/.github/workflows/vllm_ascend_test.yaml)."
|
||||
msgstr ""
|
||||
"这将复现端到端测试:[vllm_ascend_test.yaml](https://github.com/vllm-project/vllm-"
|
||||
"这将复现端到端测试。请参阅 [vllm_ascend_test.yaml](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/main/.github/workflows/vllm_ascend_test.yaml)。"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:242
|
||||
msgid "E2E test example:"
|
||||
msgstr "E2E 测试示例:"
|
||||
#: ../../source/developer_guide/contribution/testing.md:259
|
||||
msgid ""
|
||||
"For running nightly multi-node test cases locally, refer to the `Running "
|
||||
"Locally` section in [Multi Node Test](./multi_node_test.md)."
|
||||
msgstr "要在本地运行夜间多节点测试用例,请参阅 [多节点测试](./multi_node_test.md) 中的 `本地运行` 部分。"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:244
|
||||
#: ../../source/developer_guide/contribution/testing.md:261
|
||||
msgid "E2E test example"
|
||||
msgstr "端到端测试示例"
|
||||
|
||||
#: ../../source/developer_guide/contribution/testing.md:263
|
||||
msgid ""
|
||||
"Offline test example: "
|
||||
"[`tests/e2e/singlecard/test_offline_inference.py`](https://github.com/vllm-"
|
||||
"project/vllm-"
|
||||
"[`tests/e2e/singlecard/test_offline_inference.py`](https://github.com"
|
||||
"/vllm-project/vllm-"
|
||||
"ascend/blob/main/tests/e2e/singlecard/test_offline_inference.py)"
|
||||
msgstr ""
|
||||
"离线测试示例:[`tests/e2e/singlecard/test_offline_inference.py`](https://github.com/vllm-"
|
||||
"project/vllm-"
|
||||
"离线测试示例:[`tests/e2e/singlecard/test_offline_inference.py`](https://github.com"
|
||||
"/vllm-project/vllm-"
|
||||
"ascend/blob/main/tests/e2e/singlecard/test_offline_inference.py)"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:245
|
||||
#: ../../source/developer_guide/contribution/testing.md:264
|
||||
msgid ""
|
||||
"Online test examples: "
|
||||
"[`tests/e2e/singlecard/test_prompt_embedding.py`](https://github.com/vllm-"
|
||||
"project/vllm-ascend/blob/main/tests/e2e/singlecard/test_prompt_embedding.py)"
|
||||
"[`tests/e2e/singlecard/test_prompt_embedding.py`](https://github.com"
|
||||
"/vllm-project/vllm-"
|
||||
"ascend/blob/main/tests/e2e/singlecard/test_prompt_embedding.py)"
|
||||
msgstr ""
|
||||
"在线测试示例:[`tests/e2e/singlecard/test_prompt_embedding.py`](https://github.com/vllm-"
|
||||
"project/vllm-ascend/blob/main/tests/e2e/singlecard/test_prompt_embedding.py)"
|
||||
"在线测试示例:[`tests/e2e/singlecard/test_prompt_embedding.py`](https://github.com"
|
||||
"/vllm-project/vllm-"
|
||||
"ascend/blob/main/tests/e2e/singlecard/test_prompt_embedding.py)"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:246
|
||||
#: ../../source/developer_guide/contribution/testing.md:265
|
||||
msgid ""
|
||||
"Correctness test example: "
|
||||
"[`tests/e2e/singlecard/test_aclgraph_accuracy.py`](https://github.com/vllm-"
|
||||
"project/vllm-ascend/blob/main/tests/e2e/singlecard/test_aclgraph_accuracy.py)"
|
||||
"[`tests/e2e/singlecard/test_aclgraph_accuracy.py`](https://github.com"
|
||||
"/vllm-project/vllm-"
|
||||
"ascend/blob/main/tests/e2e/singlecard/test_aclgraph_accuracy.py)"
|
||||
msgstr ""
|
||||
"正确性测试示例:[`tests/e2e/singlecard/test_aclgraph_accuracy.py`](https://github.com/vllm-"
|
||||
"project/vllm-ascend/blob/main/tests/e2e/singlecard/test_aclgraph_accuracy.py)"
|
||||
"正确性测试示例:[`tests/e2e/singlecard/test_aclgraph_accuracy.py`](https://github.com"
|
||||
"/vllm-project/vllm-"
|
||||
"ascend/blob/main/tests/e2e/singlecard/test_aclgraph_accuracy.py)"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:247
|
||||
#: ../../source/developer_guide/contribution/testing.md:267
|
||||
msgid ""
|
||||
"Reduced Layer model test example: [test_torchair_graph_mode.py - "
|
||||
"DeepSeek-V3-Pruning](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/20767a043cccb3764214930d4695e53941de87ec/tests/e2e/multicard/test_torchair_graph_mode.py#L48)"
|
||||
msgstr ""
|
||||
"简化层模型测试示例:[test_torchair_graph_mode.py - "
|
||||
"DeepSeek-V3-Pruning](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/20767a043cccb3764214930d4695e53941de87ec/tests/e2e/multicard/test_torchair_graph_mode.py#L48)"
|
||||
"The CI resource is limited, and you might need to reduce the number of "
|
||||
"layers of a model. Below is an example of how to generate a reduced layer"
|
||||
" model:"
|
||||
msgstr "CI 资源有限,您可能需要减少模型的层数。以下是如何生成缩减层数模型的示例:"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:249
|
||||
#: ../../source/developer_guide/contribution/testing.md:268
|
||||
msgid ""
|
||||
"The CI resource is limited, you might need to reduce layer number of the "
|
||||
"model, below is an example of how to generate a reduced layer model:"
|
||||
msgstr "CI 资源有限,您可能需要减少模型的层数,下面是一个生成减少层数模型的示例:"
|
||||
"Fork the original model repo in modelscope. All the files in the repo "
|
||||
"except for weights are required."
|
||||
msgstr "在 ModelScope 中 Fork 原始模型仓库。需要仓库中除权重文件外的所有文件。"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:250
|
||||
msgid ""
|
||||
"Fork the original model repo in modelscope, we need all the files in the "
|
||||
"repo except for weights."
|
||||
msgstr "在 modelscope 中 fork 原始模型仓库,我们需要仓库中的所有文件,除了权重文件。"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:251
|
||||
#: ../../source/developer_guide/contribution/testing.md:269
|
||||
#, python-brace-format
|
||||
msgid ""
|
||||
"Set `num_hidden_layers` to the expected number of layers, e.g., "
|
||||
"`{\"num_hidden_layers\": 2,}`"
|
||||
msgstr "将 `num_hidden_layers` 设置为期望的层数,例如 `{\"num_hidden_layers\": 2,}`"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:252
|
||||
#: ../../source/developer_guide/contribution/testing.md:270
|
||||
msgid ""
|
||||
"Copy the following python script as `generate_random_weight.py`. Set the "
|
||||
"relevant parameters `MODEL_LOCAL_PATH`, `DIST_DTYPE` and `DIST_MODEL_PATH` "
|
||||
"as needed:"
|
||||
"relevant parameters `MODEL_LOCAL_PATH`, `DIST_DTYPE` and "
|
||||
"`DIST_MODEL_PATH` as needed:"
|
||||
msgstr ""
|
||||
"将以下 Python 脚本复制为 `generate_random_weight.py`。根据需要设置相关参数 "
|
||||
"`MODEL_LOCAL_PATH`、`DIST_DTYPE` 和 `DIST_MODEL_PATH`:"
|
||||
"将以下 Python 脚本复制为 `generate_random_weight.py`。根据需要设置相关参数 `MODEL_LOCAL_PATH`、`DIST_DTYPE` 和 `DIST_MODEL_PATH`:"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:270
|
||||
#: ../../source/developer_guide/contribution/testing.md:288
|
||||
msgid "View CI log summary in GitHub Actions"
|
||||
msgstr "在 GitHub Actions 中查看 CI 日志摘要"
|
||||
|
||||
#: ../../source/developer_guide/contribution/testing.md:290
|
||||
msgid ""
|
||||
"After a CI job finishes, you can open the corresponding GitHub Actions "
|
||||
"job page and check the `Summary` tab to view the generated CI log "
|
||||
"summary."
|
||||
msgstr "CI 作业完成后,您可以打开相应的 GitHub Actions 作业页面,并查看 `Summary` 选项卡以查看生成的 CI 日志摘要。"
|
||||
|
||||
#: ../../source/developer_guide/contribution/testing.md:293
|
||||
msgid ""
|
||||
msgstr ""
|
||||
|
||||
#: ../../source/developer_guide/contribution/testing.md:293
|
||||
msgid "GitHub Actions CI log summary"
|
||||
msgstr "GitHub Actions CI 日志摘要"
|
||||
|
||||
#: ../../source/developer_guide/contribution/testing.md:295
|
||||
msgid ""
|
||||
"The summary is intended to help developers triage failures more quickly. "
|
||||
"It may include:"
|
||||
msgstr "该摘要旨在帮助开发者更快地排查故障。它可能包括:"
|
||||
|
||||
#: ../../source/developer_guide/contribution/testing.md:297
|
||||
msgid "failed test files"
|
||||
msgstr "失败的测试文件"
|
||||
|
||||
#: ../../source/developer_guide/contribution/testing.md:298
|
||||
msgid "failed test cases"
|
||||
msgstr "失败的测试用例"
|
||||
|
||||
#: ../../source/developer_guide/contribution/testing.md:299
|
||||
msgid "distinct root-cause errors"
|
||||
msgstr "不同的根本原因错误"
|
||||
|
||||
#: ../../source/developer_guide/contribution/testing.md:300
|
||||
msgid "short error context extracted from the job log"
|
||||
msgstr "从作业日志中提取的简短错误上下文"
|
||||
|
||||
#: ../../source/developer_guide/contribution/testing.md:302
|
||||
msgid ""
|
||||
"This summary is generated from the job log by "
|
||||
"`/.github/workflows/scripts/ci_log_summary_v2.py` for unit-test and e2e "
|
||||
"workflows."
|
||||
msgstr "该摘要是由 `/.github/workflows/scripts/ci_log_summary_v2.py` 从作业日志中为单元测试和端到端测试工作流生成的。"
|
||||
|
||||
#: ../../source/developer_guide/contribution/testing.md:305
|
||||
msgid "Run doctest"
|
||||
msgstr "运行 doctest"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:272
|
||||
#: ../../source/developer_guide/contribution/testing.md:307
|
||||
msgid ""
|
||||
"vllm-ascend provides a `vllm-ascend/tests/e2e/run_doctests.sh` command to "
|
||||
"run all doctests in the doc files. The doctest is a good way to make sure "
|
||||
"the docs are up to date and the examples are executable, you can run it "
|
||||
"vllm-ascend provides a `vllm-ascend/tests/e2e/run_doctests.sh` command to"
|
||||
" run all doctests in the doc files. The doctest is a good way to make "
|
||||
"sure docs stay current and examples remain executable, which can be run "
|
||||
"locally as follows:"
|
||||
msgstr ""
|
||||
"vllm-ascend 提供了一个 `vllm-ascend/tests/e2e/run_doctests.sh` 命令,用于运行文档文件中的所有 "
|
||||
"doctest。doctest 是确保文档保持最新且示例可执行的好方法,你可以按照以下方式在本地运行它:"
|
||||
"vllm-ascend 提供了一个 `vllm-ascend/tests/e2e/run_doctests.sh` 命令来运行文档文件中的所有 doctest。doctest "
|
||||
"是确保文档保持最新且示例保持可执行性的好方法,可以按如下方式在本地运行:"
|
||||
|
||||
#: ../../developer_guide/contribution/testing.md:280
|
||||
#: ../../source/developer_guide/contribution/testing.md:315
|
||||
msgid ""
|
||||
"This will reproduce the same environment as the CI: "
|
||||
"This will reproduce the same environment as the CI. See "
|
||||
"[vllm_ascend_doctest.yaml](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/main/.github/workflows/vllm_ascend_doctest.yaml)."
|
||||
msgstr ""
|
||||
"这将复现与 CI 相同的环境:[vllm_ascend_doctest.yaml](https://github.com/vllm-"
|
||||
"project/vllm-ascend/blob/main/.github/workflows/vllm_ascend_doctest.yaml)。"
|
||||
"这将复现与 CI 相同的环境。请参阅 [vllm_ascend_doctest.yaml](https://github.com/vllm-project/vllm-"
|
||||
"ascend/blob/main/.github/workflows/vllm_ascend_doctest.yaml)。"
|
||||
@@ -0,0 +1,238 @@
|
||||
# SOME DESCRIPTIVE TITLE.
|
||||
# Copyright (C) 2025, vllm-ascend team
|
||||
# This file is distributed under the same license as the vllm-ascend
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend \n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:1
|
||||
msgid "Using AISBench"
|
||||
msgstr "使用 AISBench"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:3
|
||||
msgid ""
|
||||
"This document guides you to conduct accuracy testing using "
|
||||
"[AISBench](https://gitee.com/aisbench/benchmark/tree/master). AISBench "
|
||||
"provides accuracy and performance evaluation for many datasets."
|
||||
msgstr "本文档指导您如何使用 [AISBench](https://gitee.com/aisbench/benchmark/tree/master) 进行精度测试。AISBench 为许多数据集提供了精度和性能评估。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:5
|
||||
msgid "Online Server"
|
||||
msgstr "在线服务器"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:7
|
||||
msgid "1. Start the vLLM server"
|
||||
msgstr "1. 启动 vLLM 服务器"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:9
|
||||
msgid "You can run docker container to start the vLLM server on a single NPU:"
|
||||
msgstr "您可以运行 docker 容器在单个 NPU 上启动 vLLM 服务器:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:37
|
||||
msgid "Run the vLLM server in the docker."
|
||||
msgstr "在 docker 中运行 vLLM 服务器。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:45
|
||||
msgid ""
|
||||
"`--max_model_len` should be greater than `35000`, this will be suitable "
|
||||
"for most datasets. Otherwise the accuracy evaluation may be affected."
|
||||
msgstr "`--max_model_len` 应大于 `35000`,这适用于大多数数据集。否则可能会影响精度评估。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:48
|
||||
msgid "The vLLM server is started successfully, if you see logs as below:"
|
||||
msgstr "如果看到如下日志,则 vLLM 服务器启动成功:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:56
|
||||
msgid "2. Run different datasets using AISBench"
|
||||
msgstr "2. 使用 AISBench 运行不同数据集"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:58
|
||||
msgid "Install AISBench"
|
||||
msgstr "安装 AISBench"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:60
|
||||
msgid ""
|
||||
"Refer to [AISBench](https://gitee.com/aisbench/benchmark/tree/master) for"
|
||||
" details. Install AISBench from source."
|
||||
msgstr "详情请参考 [AISBench](https://gitee.com/aisbench/benchmark/tree/master)。从源码安装 AISBench。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:69
|
||||
msgid "Install extra AISBench dependencies."
|
||||
msgstr "安装额外的 AISBench 依赖项。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:76
|
||||
msgid "Run `ais_bench -h` to check the installation."
|
||||
msgstr "运行 `ais_bench -h` 以检查安装。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:78
|
||||
msgid "Download Dataset"
|
||||
msgstr "下载数据集"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:80
|
||||
msgid "You can choose one or multiple datasets to execute accuracy evaluation."
|
||||
msgstr "您可以选择一个或多个数据集来执行精度评估。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:82
|
||||
msgid "`C-Eval` dataset."
|
||||
msgstr "`C-Eval` 数据集。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:84
|
||||
msgid ""
|
||||
"Take `C-Eval` dataset as an example. You can refer to "
|
||||
"[Datasets](https://gitee.com/aisbench/benchmark/tree/master/ais_bench/benchmark/configs/datasets)"
|
||||
" for more datasets. Each dataset has a `README.md` with detailed download"
|
||||
" and installation instructions."
|
||||
msgstr "以 `C-Eval` 数据集为例。更多数据集请参考 [Datasets](https://gitee.com/aisbench/benchmark/tree/master/ais_bench/benchmark/configs/datasets)。每个数据集都有一个 `README.md` 文件,包含详细的下载和安装说明。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:86
|
||||
msgid "Download dataset and install it to specific path."
|
||||
msgstr "下载数据集并安装到指定路径。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:98
|
||||
msgid "`MMLU` dataset."
|
||||
msgstr "`MMLU` 数据集。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:107
|
||||
msgid "`GPQA` dataset."
|
||||
msgstr "`GPQA` 数据集。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:116
|
||||
msgid "`MATH` dataset."
|
||||
msgstr "`MATH` 数据集。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:125
|
||||
msgid "`LiveCodeBench` dataset."
|
||||
msgstr "`LiveCodeBench` 数据集。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:133
|
||||
msgid "`AIME 2024` dataset."
|
||||
msgstr "`AIME 2024` 数据集。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:144
|
||||
msgid "`GSM8K` dataset."
|
||||
msgstr "`GSM8K` 数据集。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:153
|
||||
msgid "Configuration"
|
||||
msgstr "配置"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:155
|
||||
msgid ""
|
||||
"Update the file "
|
||||
"`benchmark/ais_bench/benchmark/configs/models/vllm_api/vllm_api_general_chat.py`."
|
||||
" There are several arguments that you should update according to your "
|
||||
"environment."
|
||||
msgstr "更新文件 `benchmark/ais_bench/benchmark/configs/models/vllm_api/vllm_api_general_chat.py`。有几个参数需要根据您的环境进行更新。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:158
|
||||
msgid ""
|
||||
"`attr`: Identifier for the inference backend type, fixed as `service` "
|
||||
"(serving-based inference) or `local` (local model)."
|
||||
msgstr "`attr`:推理后端类型的标识符,固定为 `service`(基于服务的推理)或 `local`(本地模型)。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:159
|
||||
msgid "`type`: Used to select different backend API types."
|
||||
msgstr "`type`:用于选择不同的后端 API 类型。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:160
|
||||
msgid ""
|
||||
"`abbr`: Unique identifier for a local task, used to distinguish between "
|
||||
"multiple tasks."
|
||||
msgstr "`abbr`:本地任务的唯一标识符,用于区分多个任务。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:161
|
||||
msgid "`path`: Update to your model weight path."
|
||||
msgstr "`path`:更新为您的模型权重路径。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:162
|
||||
msgid "`model`: Update to your model name in vLLM."
|
||||
msgstr "`model`:更新为您的 vLLM 中的模型名称。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:163
|
||||
msgid "`host_ip` and `host_port`: Update to your vLLM server ip and port."
|
||||
msgstr "`host_ip` 和 `host_port`:更新为您的 vLLM 服务器的 IP 和端口。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:164
|
||||
msgid ""
|
||||
"`max_out_len`: Note `max_out_len` + LLM input length should be less than "
|
||||
"`max-model-len`(config in your vllm server), `32768` will be suitable for"
|
||||
" most datasets."
|
||||
msgstr "`max_out_len`:注意 `max_out_len` + LLM 输入长度应小于 `max-model-len`(在您的 vllm 服务器中配置),`32768` 适用于大多数数据集。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:165
|
||||
msgid "`batch_size`: Update according to your dataset."
|
||||
msgstr "`batch_size`:根据您的数据集进行更新。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:166
|
||||
msgid "`temperature`: Update inference argument."
|
||||
msgstr "`temperature`:更新推理参数。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:199
|
||||
msgid "Execute Accuracy Evaluation"
|
||||
msgstr "执行精度评估"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:201
|
||||
msgid "Run the following code to execute different accuracy evaluation."
|
||||
msgstr "运行以下代码以执行不同的精度评估。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:224
|
||||
msgid ""
|
||||
"After each dataset execution, you can get the result from saved files "
|
||||
"such as `outputs/default/20250628_151326`, there is an example as "
|
||||
"follows:"
|
||||
msgstr "每个数据集执行后,您可以从保存的文件(例如 `outputs/default/20250628_151326`)中获取结果,示例如下:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:249
|
||||
msgid "Execute Performance Evaluation"
|
||||
msgstr "执行性能评估"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:251
|
||||
msgid "Text-only benchmarks:"
|
||||
msgstr "纯文本基准测试:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:273
|
||||
msgid "Multi-modal benchmarks (text + images):"
|
||||
msgstr "多模态基准测试(文本 + 图像):"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:280
|
||||
msgid ""
|
||||
"After execution, you can get the result from saved files, there is an "
|
||||
"example as follows:"
|
||||
msgstr "执行后,您可以从保存的文件中获取结果,示例如下:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:300
|
||||
msgid "3. Troubleshooting"
|
||||
msgstr "3. 故障排除"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:302
|
||||
msgid "Invalid Image Path Error"
|
||||
msgstr "无效图像路径错误"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:304
|
||||
msgid "If you download the TextVQA dataset following the AISBench documentation:"
|
||||
msgstr "如果您按照 AISBench 文档下载 TextVQA 数据集:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:316
|
||||
msgid "you may encounter the following error:"
|
||||
msgstr "您可能会遇到以下错误:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_ais_bench.md:322
|
||||
msgid ""
|
||||
"You need to manually replace the dataset image paths with absolute paths,"
|
||||
" changing `/path/to/benchmark/ais_bench/datasets/textvqa/train_images/` "
|
||||
"to the actual absolute directory where the images are stored:"
|
||||
msgstr "您需要手动将数据集图像路径替换为绝对路径,将 `/path/to/benchmark/ais_bench/datasets/textvqa/train_images/` 更改为图像存储的实际绝对目录:"
|
||||
@@ -4,109 +4,111 @@
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2025.
|
||||
#
|
||||
#, fuzzy
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2025-07-18 09:01+0800\n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"Generated-By: Babel 2.17.0\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:1
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:1
|
||||
msgid "Using EvalScope"
|
||||
msgstr "使用 EvalScope"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:3
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:3
|
||||
msgid ""
|
||||
"This document will guide you have model inference stress testing and "
|
||||
"accuracy testing using [EvalScope](https://github.com/modelscope/evalscope)."
|
||||
"This document will guide you through model inference stress testing and "
|
||||
"accuracy testing using "
|
||||
"[EvalScope](https://github.com/modelscope/evalscope)."
|
||||
msgstr ""
|
||||
"本文档将指导您如何使用 [EvalScope](https://github.com/modelscope/evalscope) "
|
||||
"进行模型推理压力测试和精度测试。"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:5
|
||||
msgid "1. Online serving"
|
||||
msgstr "1. 在线服务"
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:5
|
||||
msgid "1. Online server"
|
||||
msgstr "1. 在线服务器"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:7
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:7
|
||||
msgid "You can run docker container to start the vLLM server on a single NPU:"
|
||||
msgstr "你可以运行 docker 容器,在单个 NPU 上启动 vLLM 服务器:"
|
||||
msgstr "你可以运行 Docker 容器,在单个 NPU 上启动 vLLM 服务器:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:34
|
||||
msgid "If your service start successfully, you can see the info shown below:"
|
||||
msgstr "如果你的服务启动成功,你会看到如下所示的信息:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:42
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:35
|
||||
msgid ""
|
||||
"Once your server is started, you can query the model with input prompts in "
|
||||
"new terminal:"
|
||||
msgstr "一旦你的服务器启动后,你可以在新的终端中用输入提示词查询模型:"
|
||||
"If the vLLM server is started successfully, you can see information shown"
|
||||
" below:"
|
||||
msgstr "如果 vLLM 服务器启动成功,你将看到如下所示的信息:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:55
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:43
|
||||
msgid ""
|
||||
"Once your server is started, you can query the model with input prompts "
|
||||
"in a new terminal:"
|
||||
msgstr "服务器启动后,你可以在新的终端中使用输入提示词查询模型:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:56
|
||||
msgid "2. Install EvalScope using pip"
|
||||
msgstr "2. 使用 pip 安装 EvalScope"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:57
|
||||
msgid "You can install EvalScope by using:"
|
||||
msgstr "你可以使用以下方式安装 EvalScope:"
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:58
|
||||
msgid "You can install EvalScope as follows:"
|
||||
msgstr "你可以通过以下方式安装 EvalScope:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:65
|
||||
msgid "3. Run gsm8k accuracy test using EvalScope"
|
||||
msgstr "3. 使用 EvalScope 运行 gsm8k 准确率测试"
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:66
|
||||
msgid "3. Run GSM8K using EvalScope for accuracy testing"
|
||||
msgstr "3. 使用 EvalScope 运行 GSM8K 进行精度测试"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:67
|
||||
msgid "You can `evalscope eval` run gsm8k accuracy test:"
|
||||
msgstr "你可以使用 `evalscope eval` 运行 gsm8k 准确率测试:"
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:68
|
||||
msgid "You can use `evalscope eval` to run GSM8K for accuracy testing:"
|
||||
msgstr "你可以使用 `evalscope eval` 运行 GSM8K 进行精度测试:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:78
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:114
|
||||
msgid "After 1-2 mins, the output is as shown below:"
|
||||
msgstr "1-2 分钟后,输出如下所示:"
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:80
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:117
|
||||
msgid "After 1 to 2 minutes, the output is shown below:"
|
||||
msgstr "1 到 2 分钟后,输出结果如下所示:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:88
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:90
|
||||
msgid ""
|
||||
"See more detail in: [EvalScope doc - Model API Service "
|
||||
"Evaluation](https://evalscope.readthedocs.io/en/latest/get_started/basic_usage.html#model-"
|
||||
"api-service-evaluation)."
|
||||
"See more details in [EvalScope doc - Model API Service "
|
||||
"Evaluation](https://evalscope.readthedocs.io/en/latest/get_started/basic_usage.html"
|
||||
"#model-api-service-evaluation)."
|
||||
msgstr ""
|
||||
"更多详情请见:[EvalScope 文档 - 模型 API "
|
||||
"服务评测](https://evalscope.readthedocs.io/en/latest/get_started/basic_usage.html#model-"
|
||||
"api-service-evaluation)。"
|
||||
"更多详情请参阅 [EvalScope 文档 - 模型 API "
|
||||
"服务评估](https://evalscope.readthedocs.io/en/latest/get_started/basic_usage.html"
|
||||
"#model-api-service-evaluation)。"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:90
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:92
|
||||
msgid "4. Run model inference stress testing using EvalScope"
|
||||
msgstr "4. 使用 EvalScope 运行模型推理压力测试"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:92
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:94
|
||||
msgid "Install EvalScope[perf] using pip"
|
||||
msgstr "使用 pip 安装 EvalScope[perf]"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:98
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:100
|
||||
msgid "Basic usage"
|
||||
msgstr "基本用法"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:100
|
||||
msgid "You can use `evalscope perf` run perf test:"
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:102
|
||||
msgid "You can use `evalscope perf` to run perf testing:"
|
||||
msgstr "你可以使用 `evalscope perf` 运行性能测试:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:112
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:115
|
||||
msgid "Output results"
|
||||
msgstr "输出结果"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_evalscope.md:173
|
||||
#: ../../source/developer_guide/evaluation/using_evalscope.md:176
|
||||
msgid ""
|
||||
"See more detail in: [EvalScope doc - Model Inference Stress "
|
||||
"Testing](https://evalscope.readthedocs.io/en/latest/user_guides/stress_test/quick_start.html#basic-"
|
||||
"usage)."
|
||||
"See more detail in [EvalScope doc - Model Inference Stress "
|
||||
"Testing](https://evalscope.readthedocs.io/en/latest/user_guides/stress_test/quick_start.html"
|
||||
"#basic-usage)."
|
||||
msgstr ""
|
||||
"更多详情见:[EvalScope 文档 - "
|
||||
"模型推理压力测试](https://evalscope.readthedocs.io/en/latest/user_guides/stress_test/quick_start.html#basic-"
|
||||
"usage)。"
|
||||
"更多详情请参阅 [EvalScope 文档 - "
|
||||
"模型推理压力测试](https://evalscope.readthedocs.io/en/latest/user_guides/stress_test/quick_start.html"
|
||||
"#basic-usage)。"
|
||||
@@ -4,62 +4,115 @@
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2025.
|
||||
#
|
||||
#, fuzzy
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2025-07-18 09:01+0800\n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"Generated-By: Babel 2.17.0\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_lm_eval.md:1
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:1
|
||||
msgid "Using lm-eval"
|
||||
msgstr "使用 lm-eval"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_lm_eval.md:2
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:3
|
||||
msgid "This document guides you to conduct accuracy testing using [lm-eval][1]."
|
||||
msgstr "本文档指导您如何使用 [lm-eval][1] 进行准确率测试。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:5
|
||||
msgid "Online Server"
|
||||
msgstr "在线服务器"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:7
|
||||
msgid "1. Start the vLLM server"
|
||||
msgstr "1. 启动 vLLM 服务器"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:9
|
||||
msgid "You can run docker container to start the vLLM server on a single NPU:"
|
||||
msgstr "您可以在单个 NPU 上运行 Docker 容器来启动 vLLM 服务器:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:38
|
||||
msgid "The vLLM server is started successfully, if you see logs as below:"
|
||||
msgstr "如果您看到如下日志,则表示 vLLM 服务器已成功启动:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:46
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:175
|
||||
msgid "2. Run GSM8K using lm-eval for accuracy testing"
|
||||
msgstr "2. 使用 lm-eval 运行 GSM8K 进行准确率测试"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:48
|
||||
msgid "You can query the result with input prompts:"
|
||||
msgstr "您可以使用输入提示词查询结果:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:75
|
||||
msgid "The output format matches the following:"
|
||||
msgstr "输出格式符合以下形式:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:105
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:177
|
||||
msgid "Install lm-eval in the container:"
|
||||
msgstr "在容器中安装 lm-eval:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:114
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:186
|
||||
msgid ""
|
||||
"This document will guide you have a accuracy testing using [lm-"
|
||||
"eval](https://github.com/EleutherAI/lm-evaluation-harness)."
|
||||
msgstr ""
|
||||
"本文将指导你如何使用 [lm-eval](https://github.com/EleutherAI/lm-evaluation-harness) "
|
||||
"进行准确率测试。"
|
||||
"The Docker container is launched with `VLLM_USE_MODELSCOPE=True`, which "
|
||||
"may cause lm-eval to download datasets from ModelScope instead of "
|
||||
"HuggingFace. Setting `USE_MODELSCOPE_HUB=0` disables this behavior so "
|
||||
"that lm-eval can fetch datasets from HuggingFace correctly."
|
||||
msgstr "Docker 容器以 `VLLM_USE_MODELSCOPE=True` 启动,这可能导致 lm-eval 从 ModelScope 而非 HuggingFace 下载数据集。设置 `USE_MODELSCOPE_HUB=0` 可禁用此行为,使 lm-eval 能够正确从 HuggingFace 获取数据集。"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_lm_eval.md:4
|
||||
msgid "1. Run docker container"
|
||||
msgstr "1. 运行 docker 容器"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_lm_eval.md:6
|
||||
msgid "You can run docker container on a single NPU:"
|
||||
msgstr "你可以在单个NPU上运行docker容器:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_lm_eval.md:33
|
||||
msgid "2. Run ceval accuracy test using lm-eval"
|
||||
msgstr "2. 使用 lm-eval 运行 ceval 准确性测试"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_lm_eval.md:34
|
||||
msgid "Install lm-eval in the container."
|
||||
msgstr "在容器中安装 lm-eval。"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_lm_eval.md:39
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:120
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:192
|
||||
msgid "Run the following command:"
|
||||
msgstr "运行以下命令:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_lm_eval.md:50
|
||||
msgid "After 1-2 mins, the output is as shown below:"
|
||||
msgstr "1-2 分钟后,输出如下所示:"
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:131
|
||||
msgid "After 30 minutes, the output is as shown below:"
|
||||
msgstr "30 分钟后,输出如下所示:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_lm_eval.md:62
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:143
|
||||
msgid "Offline Server"
|
||||
msgstr "离线服务器"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:145
|
||||
msgid "1. Run docker container"
|
||||
msgstr "1. 运行 docker 容器"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:147
|
||||
msgid "You can run docker container on a single NPU:"
|
||||
msgstr "您可以在单个 NPU 上运行 docker 容器:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:203
|
||||
msgid "After 1 to 2 minutes, the output is shown below:"
|
||||
msgstr "1 到 2 分钟后,输出如下所示:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:215
|
||||
msgid "Use Offline Datasets"
|
||||
msgstr "使用离线数据集"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:217
|
||||
msgid ""
|
||||
"You can see more usage on [Lm-eval Docs](https://github.com/EleutherAI/lm-"
|
||||
"evaluation-harness/blob/main/docs/README.md)."
|
||||
msgstr ""
|
||||
"你可以在 [Lm-eval 文档](https://github.com/EleutherAI/lm-evaluation-"
|
||||
"harness/blob/main/docs/README.md) 上查看更多用法。"
|
||||
"Take GSM8K (single dataset) and MMLU (multi-subject dataset) as examples,"
|
||||
" and you can see more from [using-local-datasets][2]."
|
||||
msgstr "以 GSM8K(单数据集)和 MMLU(多学科数据集)为例,您可以在 [using-local-datasets][2] 中查看更多信息。"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:231
|
||||
msgid "Set [gsm8k.yaml][3] as follows:"
|
||||
msgstr "按如下方式设置 [gsm8k.yaml][3]:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:294
|
||||
msgid "Set [_default_template_yaml][4] as follows:"
|
||||
msgstr "按如下方式设置 [_default_template_yaml][4]:"
|
||||
|
||||
#: ../../source/developer_guide/evaluation/using_lm_eval.md:317
|
||||
msgid "You can see more usage on [Lm-eval Docs][5]."
|
||||
msgstr "您可以在 [Lm-eval 文档][5] 中查看更多用法。"
|
||||
@@ -4,80 +4,79 @@
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2025.
|
||||
#
|
||||
#, fuzzy
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2025-07-18 09:01+0800\n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"Generated-By: Babel 2.17.0\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_opencompass.md:1
|
||||
#: ../../source/developer_guide/evaluation/using_opencompass.md:1
|
||||
msgid "Using OpenCompass"
|
||||
msgstr "使用 OpenCompass"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_opencompass.md:2
|
||||
#: ../../source/developer_guide/evaluation/using_opencompass.md:3
|
||||
msgid ""
|
||||
"This document will guide you have a accuracy testing using "
|
||||
"This document guides you to conduct accuracy testing using "
|
||||
"[OpenCompass](https://github.com/open-compass/opencompass)."
|
||||
msgstr ""
|
||||
"本文档将指导你如何使用 [OpenCompass](https://github.com/open-compass/opencompass) "
|
||||
"进行准确率测试。"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_opencompass.md:4
|
||||
msgid "1. Online Serving"
|
||||
#: ../../source/developer_guide/evaluation/using_opencompass.md:5
|
||||
msgid "1. Online Server"
|
||||
msgstr "1. 在线服务"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_opencompass.md:6
|
||||
msgid "You can run docker container to start the vLLM server on a single NPU:"
|
||||
msgstr "你可以运行 docker 容器,在单个 NPU 上启动 vLLM 服务器:"
|
||||
#: ../../source/developer_guide/evaluation/using_opencompass.md:7
|
||||
msgid "You can run a docker container to start the vLLM server on a single NPU:"
|
||||
msgstr "你可以运行一个 Docker 容器,在单个 NPU 上启动 vLLM 服务器:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_opencompass.md:32
|
||||
msgid "If your service start successfully, you can see the info shown below:"
|
||||
msgstr "如果你的服务启动成功,你会看到如下所示的信息:"
|
||||
#: ../../source/developer_guide/evaluation/using_opencompass.md:35
|
||||
msgid "The vLLM server is started successfully, if you see information as below:"
|
||||
msgstr "如果看到如下信息,则表明 vLLM 服务器已成功启动:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_opencompass.md:39
|
||||
#: ../../source/developer_guide/evaluation/using_opencompass.md:43
|
||||
msgid ""
|
||||
"Once your server is started, you can query the model with input prompts in "
|
||||
"new terminal:"
|
||||
msgstr "一旦你的服务器启动后,你可以在新的终端中用输入提示词查询模型:"
|
||||
"Once your server is started, you can query the model with input prompts "
|
||||
"in a new terminal."
|
||||
msgstr "服务器启动后,你可以在新的终端中使用输入提示词来查询模型。"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_opencompass.md:51
|
||||
msgid "2. Run ceval accuracy test using OpenCompass"
|
||||
msgstr "2. 使用 OpenCompass 运行 ceval 准确率测试"
|
||||
#: ../../source/developer_guide/evaluation/using_opencompass.md:56
|
||||
msgid "2. Run C-Eval using OpenCompass for accuracy testing"
|
||||
msgstr "2. 使用 OpenCompass 运行 C-Eval 进行准确率测试"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_opencompass.md:52
|
||||
#: ../../source/developer_guide/evaluation/using_opencompass.md:58
|
||||
msgid ""
|
||||
"Install OpenCompass and configure the environment variables in the "
|
||||
"container."
|
||||
msgstr "在容器中安装 OpenCompass 并配置环境变量。"
|
||||
"container:"
|
||||
msgstr "在容器中安装 OpenCompass 并配置环境变量:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_opencompass.md:64
|
||||
#: ../../source/developer_guide/evaluation/using_opencompass.md:70
|
||||
msgid ""
|
||||
"Add `opencompass/configs/eval_vllm_ascend_demo.py` with the following "
|
||||
"content:"
|
||||
msgstr "添加 `opencompass/configs/eval_vllm_ascend_demo.py`,内容如下:"
|
||||
"Add the following content to "
|
||||
"`opencompass/configs/eval_vllm_ascend_demo.py`:"
|
||||
msgstr "将以下内容添加到 `opencompass/configs/eval_vllm_ascend_demo.py` 文件中:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_opencompass.md:104
|
||||
#: ../../source/developer_guide/evaluation/using_opencompass.md:110
|
||||
msgid "Run the following command:"
|
||||
msgstr "运行以下命令:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_opencompass.md:110
|
||||
msgid "After 1-2 mins, the output is as shown below:"
|
||||
msgstr "1-2 分钟后,输出如下所示:"
|
||||
#: ../../source/developer_guide/evaluation/using_opencompass.md:116
|
||||
msgid "After 1 to 2 minutes, the output is shown below:"
|
||||
msgstr "1 到 2 分钟后,输出结果如下所示:"
|
||||
|
||||
#: ../../developer_guide/evaluation/using_opencompass.md:120
|
||||
#: ../../source/developer_guide/evaluation/using_opencompass.md:126
|
||||
msgid ""
|
||||
"You can see more usage on [OpenCompass "
|
||||
"Docs](https://opencompass.readthedocs.io/en/latest/index.html)."
|
||||
msgstr ""
|
||||
"你可以在 [OpenCompass "
|
||||
"文档](https://opencompass.readthedocs.io/en/latest/index.html) 查看更多用法。"
|
||||
"文档](https://opencompass.readthedocs.io/en/latest/index.html) 中查看更多用法。"
|
||||
@@ -6,183 +6,187 @@
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Report-Msgid-Bugs-To: EMAIL@ADDRESS\n"
|
||||
"POT-Creation-Date: 2025-11-21 10:19+0800\n"
|
||||
"PO-Revision-Date: 2025-11-21 10:31\n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: 2025-11-21 10:31+0000\n"
|
||||
"Last-Translator: Codex <codex@example.com>\n"
|
||||
"Language-Team: Chinese (Simplified) <LL@li.org>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: Chinese (Simplified) <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"Generated-By: Babel 2.17.0\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../developer_guide/performance_and_performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:1
|
||||
msgid "MSProbe Debugging Guide"
|
||||
msgstr "MSProbe 调试指南"
|
||||
|
||||
#: ../../developer_guide/performance_and_performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:3
|
||||
msgid ""
|
||||
"During inference or training runs we often encounter accuracy anomalies such"
|
||||
" as outputs drifting away from the expectation, unstable numerical behavior "
|
||||
"(NaN/Inf), or predictions that no longer match the labels. To pinpoint the "
|
||||
"root cause we have to monitor and capture intermediate data produced while "
|
||||
"the model executes—feature maps, weights, activations, and layer outputs. By"
|
||||
" capturing key tensors at specific stages, logging I/O pairs for the core "
|
||||
"layers, and retaining contextual metadata (prompts, tensor dtypes, hardware "
|
||||
"configuration, etc.), we can systematically trace where the accuracy "
|
||||
"degradation or numerical error started. This guide describes the end-to-end "
|
||||
"workflow for diagnosing accuracy issues for AI models (with a focus on vllm-"
|
||||
"ascend services): preparation, data capture, and analysis & verification."
|
||||
"During inference or training runs we often encounter accuracy anomalies "
|
||||
"such as outputs drifting away from the expectation, unstable numerical "
|
||||
"behavior (NaN/Inf), or predictions that no longer match the labels. To "
|
||||
"pinpoint the root cause we have to monitor and capture intermediate data "
|
||||
"produced while the model executes—feature maps, weights, activations, and"
|
||||
" layer outputs. By capturing key tensors at specific stages, logging I/O "
|
||||
"pairs for the core layers, and retaining contextual metadata (prompts, "
|
||||
"tensor dtypes, hardware configuration, etc.), we can systematically trace"
|
||||
" where the accuracy degradation or numerical error started. This guide "
|
||||
"describes the end-to-end workflow for diagnosing accuracy issues for AI "
|
||||
"models (with a focus on vllm-ascend services): preparation, data capture,"
|
||||
" and analysis & verification."
|
||||
msgstr ""
|
||||
"在推理或训练过程中,我们经常会遇到输出偏离预期、出现 NaN/Inf "
|
||||
"等数值不稳定现象,或者模型预测与标签不一致等精度异常。要定位根因,就必须监控并采集模型执行过程中的中间数据——例如特征图、权重、激活值及各层输出。通过在关键阶段捕获核心张量、记录核心层的输入输出对,并保留提示词、张量"
|
||||
" dtype、硬件配置等上下文元数据,我们可以系统追踪精度退化或数值错误的源头。本指南聚焦 vllm-ascend 服务,介绍 AI "
|
||||
"模型精度问题排查的完整流程:准备、数据采集以及分析与验证。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:5
|
||||
msgid "0. Background Concepts"
|
||||
msgstr "0. 前置概念"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:7
|
||||
msgid "`msprobe` supports three accuracy levels:"
|
||||
msgstr "`msprobe` 支持三种精度级别:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:9
|
||||
msgid ""
|
||||
"**L0**: dumps tensors at the module level and generates `construct.json` so "
|
||||
"that visualization tools can rebuild the network structure. A model or "
|
||||
"submodule handle must be passed in."
|
||||
msgstr "**L0**:在`nn.Module`级别保存`tensor`,并生成 `construct.json` 以便可视化工具还原网络结构,需要传入模型或子模块句柄。"
|
||||
"**L0**: dumps tensors at the module level and generates `construct.json` "
|
||||
"so that visualization tools can rebuild the network structure. A model or"
|
||||
" submodule handle must be passed in."
|
||||
msgstr ""
|
||||
"**L0**:在`nn.Module`级别保存`tensor`,并生成 `construct.json` "
|
||||
"以便可视化工具还原网络结构,需要传入模型或子模块句柄。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:10
|
||||
msgid ""
|
||||
"**L1**: collects operator-level statistics only, which is suitable for "
|
||||
"lightweight troubleshooting."
|
||||
msgstr "**L1**:仅采集算子级统计信息,适合轻量排查。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:11
|
||||
msgid ""
|
||||
"**mix**: captures both structural information and operator statistics, which"
|
||||
" is useful when you need both graph reconstruction and numerical "
|
||||
"**mix**: captures both structural information and operator statistics, "
|
||||
"which is useful when you need both graph reconstruction and numerical "
|
||||
"comparisons."
|
||||
msgstr "**mix**:同时获取结构信息与算子统计,适用于既要构图又要进行数值对比的场景。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:13
|
||||
msgid "1. Prerequisites"
|
||||
msgstr "1. 前提条件"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:15
|
||||
msgid "1.1 Install `msprobe`"
|
||||
msgstr "1.1 安装 `msprobe`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:17
|
||||
msgid "Install msprobe with pip:"
|
||||
msgstr "使用 pip 安装 msprobe:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:23
|
||||
msgid "1.2 Visualization dependencies (optional)"
|
||||
msgstr "1.2 可视化依赖(可选)"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:25
|
||||
msgid ""
|
||||
"Install additional dependencies if you need to visualize the captured data."
|
||||
"Install additional dependencies if you need to visualize the captured "
|
||||
"data."
|
||||
msgstr "如需对采集的数据进行可视化,请安装以下依赖。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:27
|
||||
msgid "Install `tb_graph_ascend`:"
|
||||
msgstr "安装 `tb_graph_ascend`:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:33
|
||||
msgid "2. Collecting Data with `msprobe`"
|
||||
msgstr "2. 使用 `msprobe` 采集数据"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:35
|
||||
msgid ""
|
||||
"We generally follow a coarse-to-fine strategy when capturing data. First "
|
||||
"identify the token where the issue shows up, and then decide which range "
|
||||
"needs to be sampled around that token. The typical workflow is described "
|
||||
"below."
|
||||
"We generally follow a coarse-to-fine strategy when capturing data. First,"
|
||||
" identify the token where the issue shows up, and then decide which range"
|
||||
" needs to be sampled around that token. The typical workflow is described"
|
||||
" below."
|
||||
msgstr "采集通常遵循由粗到细的策略:先确定问题出现的 token,再围绕该 token 决定采样范围,常规流程如下。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:37
|
||||
msgid "2.1 Prepare the dump configuration file"
|
||||
msgstr "2.1 准备 dump 配置文件"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:39
|
||||
msgid ""
|
||||
"Create a `config.json` that can be parsed by `PrecisionDebugger` and place "
|
||||
"it in an accessible path. Common fields are:"
|
||||
"Create a `config.json` that can be parsed by `PrecisionDebugger` and "
|
||||
"place it in an accessible path. Common fields are:"
|
||||
msgstr "创建可被 `PrecisionDebugger` 解析的 `config.json` 并放置在可访问路径,常见字段如下:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "Field"
|
||||
msgstr "字段"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "Description"
|
||||
msgstr "说明"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "Required"
|
||||
msgstr "必填"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "`task`"
|
||||
msgstr "`task`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid ""
|
||||
"Type of dump task. Common PyTorch values include `\"statistics\"` and "
|
||||
"`\"tensor\"`. A statistics task collects tensor statistics (mean, variance, "
|
||||
"max, min, etc.) while a tensor task captures arbitrary tensors."
|
||||
"`\"tensor\"`. A statistics task collects tensor statistics (mean, "
|
||||
"variance, max, min, etc.) while a tensor task captures arbitrary tensors."
|
||||
msgstr ""
|
||||
"dump 任务类型。PyTorch 常见取值包括 `\"statistics\"` 和 `\"tensor\"`:statistics "
|
||||
"任务采集张量统计量(均值、方差、最大值、最小值等),tensor 任务可采集任意张量。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "Yes"
|
||||
msgstr "是"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "`dump_path`"
|
||||
msgstr "`dump_path`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid ""
|
||||
"Directory where dump results are stored. When omitted, `msprobe` uses its "
|
||||
"default path."
|
||||
"Directory where dump results are stored. When omitted, `msprobe` uses its"
|
||||
" default path."
|
||||
msgstr "dump 结果保存目录,未配置时使用 `msprobe` 默认路径。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "No"
|
||||
msgstr "否"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "`rank`"
|
||||
msgstr "`rank`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid ""
|
||||
"Ranks to sample. An empty list collects every rank. For single-card tasks "
|
||||
"you must set this field to `[]`."
|
||||
"Ranks to sample. An empty list collects every rank. For single-card "
|
||||
"tasks, you must set this field to `[]`."
|
||||
msgstr "指定需要采集的设备 rank,空列表表示全部 rank;单卡任务必须配置为 `[]`。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "`step`"
|
||||
msgstr "`step`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "Token iteration(s) to sample. An empty list means every iteration."
|
||||
msgstr "指定采集的 token 轮次,空列表表示全部迭代。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "`level`"
|
||||
msgstr "`level`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid ""
|
||||
"Dump level string (`\"L0\"`, `\"L1\"`, or `\"mix\"`). `L0` targets "
|
||||
"`nn.Module`, `L1` targets `torch.api`, and `mix` collects both."
|
||||
@@ -190,372 +194,354 @@ msgstr ""
|
||||
"dump 级别字符串(`\"L0\"`、`\"L1\"`、`\"mix\"`),L0 面向 `nn.Module`,L1 面向 "
|
||||
"`torch.api`,mix 同时采集两者。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "`async_dump`"
|
||||
msgstr "`async_dump`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid ""
|
||||
"Whether to enable asynchronous dump (supported for PyTorch "
|
||||
"`statistics`/`tensor` tasks). Defaults to `false`."
|
||||
msgstr "是否启用异步 dump(PyTorch `statistics`/`tensor` 任务可用),默认 `false`。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "`scope`"
|
||||
msgstr "`scope`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "Module range to sample. An empty list collects every module."
|
||||
msgstr "指定需要采集的模块范围,空列表表示全部模块。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "`list`"
|
||||
msgstr "`list`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "Operator range to sample. An empty list collects every operator."
|
||||
msgstr "指定需要采集的算子范围,空列表表示全部算子。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid ""
|
||||
"To restrict the operators that are captured, configure the `list` block:"
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:52
|
||||
msgid "To restrict the operators that are captured, configure the `list` block:"
|
||||
msgstr "如需进一步限定算子范围,请配置 `list`:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:54
|
||||
msgid ""
|
||||
"`scope` (list[str]): In PyTorch pynative scenarios this field restricts the "
|
||||
"dump range. Provide two module or API names that follow the tool's naming "
|
||||
"convention to lock a range; only data between the two names will be dumped. "
|
||||
"Examples:"
|
||||
"`scope` (list[str]): In PyTorch PyNative scenarios this field restricts "
|
||||
"the dump range. Provide two module or API names that follow the tool's "
|
||||
"naming convention to lock a range; only data between the two names will "
|
||||
"be dumped. Examples:"
|
||||
msgstr ""
|
||||
"`scope`(list[str]):在 PyTorch 动态图场景下用于限定 dump 区间。按照工具命名格式提供两个模块或 API 名称,只会 "
|
||||
"dump 这一区间内的数据。示例:"
|
||||
"`scope`(list[str]):在 PyTorch 动态图场景下用于限定 dump 区间。按照工具命名格式提供两个模块或 API 名称,只会"
|
||||
" dump 这一区间内的数据。示例:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:62
|
||||
msgid ""
|
||||
"The `level` setting determines what can be provided—modules when `level=L0`,"
|
||||
" APIs when `level=L1`, and either modules or APIs when `level=mix`."
|
||||
msgstr ""
|
||||
"`level` 的取值决定可配置内容:`level=L0` 填模块名,`level=L1` 填 API 名,`level=mix` 则二者皆可。"
|
||||
"The `level` setting determines what can be provided—modules when "
|
||||
"`level=L0`, APIs when `level=L1`, and either modules or APIs when "
|
||||
"`level=mix`."
|
||||
msgstr "`level` 的取值决定可配置内容:`level=L0` 填模块名,`level=L1` 填 API 名,`level=mix` 则二者皆可。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:64
|
||||
msgid "`list` (list[str]): Custom operator list. Options include:"
|
||||
msgstr "`list`(list[str]):用于自定义采集的算子范围,常见方式包括:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:65
|
||||
msgid ""
|
||||
"Supply the full names of specific APIs in PyTorch pynative scenarios to only"
|
||||
" dump those APIs. Example: `\"list\": [\"Tensor.permute.1.forward\", "
|
||||
"\"Tensor.transpose.2.forward\", \"Torch.relu.3.backward\"]`."
|
||||
"Supply the full names of specific APIs in PyTorch pynative scenarios to "
|
||||
"only dump those APIs. Example: `\"list\": [\"Tensor.permute.1.forward\", "
|
||||
"\"Tensor.transpose.2.forward\", \"Torch.relu.3.forward\"]`."
|
||||
msgstr ""
|
||||
"在 PyTorch 动态图场景中配置 API 全称,仅 dump 这些 API,例如 `\"list\": "
|
||||
"[\"Tensor.permute.1.forward\", \"Tensor.transpose.2.forward\", "
|
||||
"\"Torch.relu.3.backward\"]`。"
|
||||
"\"Torch.relu.3.forward\"]`。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:66
|
||||
msgid ""
|
||||
"When `level=mix`, you can provide module names so that the dump expands to "
|
||||
"everything produced while the module is running. Example: `\"list\": "
|
||||
"When `level=mix`, you can provide module names so that the dump expands "
|
||||
"to everything produced while the module is running. Example: `\"list\": "
|
||||
"[\"Module.module.language_model.encoder.layers.0.mlp.ParallelMlp.forward.0\"]`."
|
||||
msgstr ""
|
||||
"当 `level=mix` 时可以填写模块名称,工具会在该模块执行期间展开并 dump 所有数据,例如 `\"list\": "
|
||||
"[\"Module.module.language_model.encoder.layers.0.mlp.ParallelMlp.forward.0\"]`。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:67
|
||||
msgid ""
|
||||
"Provide a substring such as `\"list\": [\"relu\"]` to dump every API whose "
|
||||
"name contains the substring. When `level=mix`, modules whose names contain "
|
||||
"the substring are also expanded."
|
||||
"Provide a substring such as `\"list\": [\"relu\"]` to dump every API "
|
||||
"whose name contains the substring. When `level=mix`, modules whose names "
|
||||
"contain the substring are also expanded."
|
||||
msgstr ""
|
||||
"也可以仅提供子串(如 `\"list\": [\"relu\"]`),会 dump 名称包含该字符串的 API,且 `level=mix` "
|
||||
"时会展开名称包含该字符串的模块。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:69
|
||||
msgid "Example configuration:"
|
||||
msgstr "示例配置:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "2. Enable `msprobe` in vllm-ascend"
|
||||
msgstr "2. 在 vllm-ascend 中启用 `msprobe`"
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:90
|
||||
msgid "3. Enable `msprobe` in vllm-ascend"
|
||||
msgstr "3. 在 vllm-ascend 中启用 `msprobe`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:92
|
||||
msgid ""
|
||||
"Start vLLM in eager mode by adding `--enforce-eager` (static-graph scenarios"
|
||||
" are not supported yet) and pass the config path through `--additional-"
|
||||
"config`:"
|
||||
"Start vLLM in eager mode by adding `--enforce-eager` (static-graph "
|
||||
"scenarios are not supported yet) and pass the config path through "
|
||||
"`--additional-config`:"
|
||||
msgstr ""
|
||||
"通过添加 `--enforce-eager` 以 eager 模式启动 vLLM(静态图暂不支持),并通过 `--additional-config` "
|
||||
"传入配置路径:"
|
||||
"通过添加 `--enforce-eager` 以 eager 模式启动 vLLM(静态图暂不支持),并通过 `--additional-"
|
||||
"config` 传入配置路径:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "3. Send requests and collect dumps"
|
||||
msgstr "3. 发送请求并采集 dump"
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:103
|
||||
msgid "4. Send requests and collect dumps"
|
||||
msgstr "4. 发送请求并采集 dump"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:105
|
||||
msgid "Send inference requests as usual, for example:"
|
||||
msgstr "按常规方式发送推理请求,例如:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:118
|
||||
msgid ""
|
||||
"Each request drives the sequence `msprobe: start -> forward/backward -> stop"
|
||||
" -> step`. The runner invokes `step()` on every code path, so you always get"
|
||||
" a complete dataset even if inference returns early."
|
||||
"Each request drives the sequence `msprobe: start -> forward -> stop -> "
|
||||
"step`. The runner invokes `step()` on every code path, so you always get "
|
||||
"a complete dataset even if inference returns early."
|
||||
msgstr ""
|
||||
"每个请求都会执行 `msprobe: start -> forward/backward -> stop -> step`,Runner "
|
||||
"每个请求都会执行 `msprobe: start -> forward -> stop -> step`,Runner "
|
||||
"在所有路径都会调用 `step()`,即使推理提前结束也能拿到完整数据。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:120
|
||||
msgid "Dump files are written into `dump_path`. They usually contain:"
|
||||
msgstr "dump 文件写入 `dump_path`,通常包含:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:121
|
||||
msgid "Tensor files grouped by operator/module."
|
||||
msgstr "按算子或模块划分的张量文件。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:122
|
||||
msgid ""
|
||||
"`dump.json`, which records metadata such as dtype, shape, min/max, and "
|
||||
"`requires_grad`."
|
||||
msgstr "描述 dtype、shape、最小/最大值以及 `requires_grad` 等信息的 `dump.json`。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:123
|
||||
msgid ""
|
||||
"`construct.json`, which is generated when `level` is `L0` or `mix` (required"
|
||||
" for visualization)."
|
||||
"`construct.json`, which is generated when `level` is `L0` or `mix` "
|
||||
"(required for visualization)."
|
||||
msgstr "当级别为 `L0` 或 `mix` 时生成的 `construct.json`(可视化必需)。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:125
|
||||
msgid "Example directory layout:"
|
||||
msgstr "目录结构示例:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#, python-brace-format
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:156
|
||||
msgid ""
|
||||
"`rank`: Device ID. Each card writes its data to the corresponding `rank{ID}`"
|
||||
" directory. In non-distributed scenarios the directory is simply named "
|
||||
"`rank`."
|
||||
"`rank`: Device ID. Each card writes its data to the corresponding "
|
||||
"`rank{ID}` directory. In non-distributed scenarios the directory is "
|
||||
"simply named `rank`."
|
||||
msgstr "`rank`:设备 ID。每张卡写入对应的 `rank{ID}` 目录,非分布式场景目录名称为 `rank`。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:157
|
||||
msgid "`dump_tensor_data`: Tensor payloads that were collected."
|
||||
msgstr "`dump_tensor_data`:采集到的张量数据。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:158
|
||||
msgid ""
|
||||
"`dump.json`: Statistics for the forward/backward data of each API or module,"
|
||||
" including names, dtype, shape, max, min, mean, L2 norm (square root of the "
|
||||
"L2 variance), and CRC-32 when `summary_mode=\"md5\"`. See [dump.json file "
|
||||
"description](#dumpjson-file-description) for details."
|
||||
"`dump.json`: Statistics for the forward data of each API or module, "
|
||||
"including names, dtype, shape, max, min, mean, L2 norm (square root of "
|
||||
"the L2 variance), and CRC-32 when `summary_mode=\"md5\"`. See [dump.json "
|
||||
"file description](#dumpjson-file-description) for details."
|
||||
msgstr ""
|
||||
"`dump.json`:保存各 API 或模块前/反向数据统计,包含名称、dtype、shape、max、min、mean、L2 "
|
||||
"norm(平方根)以及在 `summary_mode=\"md5\"` 下的 CRC-32。详见 [dump.json file "
|
||||
"description](#dumpjson-file-description)。"
|
||||
"`dump.json`:各 API 或模块前向数据的统计信息,包括名称、dtype、shape、最大值、最小值、平均值、L2 范数(L2 方差的平方根),以及在 `summary_mode=\"md5\"` 时的 CRC-32 值。详见 [dump.json 文件说明](#dumpjson-file-description)。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:159
|
||||
msgid ""
|
||||
"`dump_error_info.log`: Present only when the dump tool encountered an error "
|
||||
"and records the failure log."
|
||||
msgstr "`dump_error_info.log`:仅在 dump 工具报错时生成,记录错误日志。"
|
||||
"`dump_error_info.log`: Present only when the dump tool encountered an "
|
||||
"error and records the failure log."
|
||||
msgstr "`dump_error_info.log`:仅在 dump 工具遇到错误时生成,记录失败日志。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:160
|
||||
msgid "`stack.json`: Call stacks for APIs/modules."
|
||||
msgstr "`stack.json`:API/Module 的调用栈信息。"
|
||||
msgstr "`stack.json`:API/模块的调用栈信息。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:161
|
||||
msgid ""
|
||||
"`construct.json`: Hierarchical structure description. Empty when `level=L1`."
|
||||
msgstr "`construct.json`:分层结构描述,`level=L1` 时为空。"
|
||||
"`construct.json`: Hierarchical structure description. Empty when "
|
||||
"`level=L1`."
|
||||
msgstr "`construct.json`:分层结构描述,当 `level=L1` 时为空。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "4. Analyze the results"
|
||||
msgstr "4. 分析结果"
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:163
|
||||
msgid "5. Analyze the results"
|
||||
msgstr "5. 分析结果"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "4.1 Prerequisites"
|
||||
msgstr "4.1 前置条件"
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:165
|
||||
msgid "5.1 Prerequisites"
|
||||
msgstr "5.1 前置条件"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:167
|
||||
msgid ""
|
||||
"You typically need two dump datasets: one from the \"problem side\" (the run"
|
||||
" that exposes the accuracy or numerical error) and another from the "
|
||||
"You typically need two dump datasets: one from the \"problem side\" (the "
|
||||
"run that exposes the accuracy or numerical error) and another from the "
|
||||
"\"benchmark side\" (a good baseline). These datasets do not have to be "
|
||||
"identical—they can come from different branches, framework versions, or even"
|
||||
" alternative implementations (operator substitutions, different graph-"
|
||||
"optimization switches, etc.). As long as they use the same or similar "
|
||||
"inputs, hardware topology, and sampling points (step/token), `msprobe` can "
|
||||
"compare them and locate the divergent nodes. If you cannot find a perfectly "
|
||||
"clean benchmark, start by capturing the problem-side data, craft the "
|
||||
"smallest reproducible case by hand, and perform a self-comparison. Below we "
|
||||
"assume the problem dump is `problem_dump` and the benchmark dump is "
|
||||
"`bench_dump`."
|
||||
"identical—they can come from different branches, framework versions, or "
|
||||
"even alternative implementations (operator substitutions, different "
|
||||
"graph-optimization switches, etc.). As long as they use the same or "
|
||||
"similar inputs, hardware topology, and sampling points (step/token), "
|
||||
"`msprobe` can compare them and locate the divergent nodes. If you cannot "
|
||||
"find a perfectly clean benchmark, start by capturing the problem-side "
|
||||
"data, craft the smallest reproducible case by hand, and perform a self-"
|
||||
"comparison. Below we assume the problem dump is `problem_dump` and the "
|
||||
"benchmark dump is `bench_dump`."
|
||||
msgstr ""
|
||||
"通常需要准备两份 dump "
|
||||
"数据:一份来自出现精度或数值异常的“问题侧”,另一份来自表现正常的“标杆侧”。两份数据无需完全一致,可以来自不同分支、不同框架版本,甚至不同实现(算子替换、图优化开关差异等)。只要输入、硬件拓扑和采样点(step/token)保持一致或相近,msprobe"
|
||||
" 就能对比并定位差异节点。若无法找到足够干净的标杆,可先采集问题侧数据,手动构造最小复现用例并进行自对比。下文默认问题侧目录为 "
|
||||
"`problem_dump`,标杆侧为 `bench_dump`。"
|
||||
"通常需要两份 dump 数据集:一份来自“问题侧”(暴露精度或数值错误的运行),另一份来自“标杆侧”(良好的基线)。这些数据集不必完全相同——它们可以来自不同的分支、框架版本,甚至是替代实现(算子替换、不同的图优化开关等)。只要它们使用相同或相似的输入、硬件拓扑和采样点(step/token),`msprobe` 就可以比较它们并定位差异节点。如果找不到完全干净的标杆,可以先捕获问题侧数据,手动构建最小的可复现案例,并进行自比较。下文假设问题侧 dump 为 `problem_dump`,标杆侧 dump 为 `bench_dump`。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "4.2 Visualization"
|
||||
msgstr "4.2 可视化"
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:169
|
||||
msgid "5.2 Visualization"
|
||||
msgstr "5.2 可视化"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:171
|
||||
msgid ""
|
||||
"Use `msprobe graph_visualize` to generate results that can be opened inside "
|
||||
"`tb_graph_ascend`."
|
||||
msgstr "使用 `msprobe graph_visualize` 生成结果,并在 `tb_graph_ascend` 中查看。"
|
||||
"Use `msprobe -f pytorch graph` to generate results that can be opened "
|
||||
"inside `tb_graph_ascend`."
|
||||
msgstr "使用 `msprobe -f pytorch graph` 生成结果,可在 `tb_graph_ascend` 中打开。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:173
|
||||
msgid ""
|
||||
"Ensure the dump contains `construct.json` (i.e., `level = L0` or `level = "
|
||||
"mix`)."
|
||||
msgstr "确保 dump 中包含 `construct.json`(即 `level=L0` 或 `level=mix`)。"
|
||||
"Ensure the dump contains `construct.json` (i.e., `level = L0` or `level ="
|
||||
" mix`)."
|
||||
msgstr "确保 dump 包含 `construct.json`(即 `level = L0` 或 `level = mix`)。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:174
|
||||
msgid ""
|
||||
"Prepare a comparison file such as `compare.json`. Its format and generation "
|
||||
"flow are described in section 3.1.3 of `msprobe_visualization.md`. Example "
|
||||
"(minimal runnable snippet):"
|
||||
msgstr ""
|
||||
"准备 `compare.json` 等对比文件,其格式与生成方式见 `msprobe_visualization.md` 3.1.3 节。示例:"
|
||||
"Prepare a comparison file such as `compare.json`. Its format and "
|
||||
"generation flow are described in section 3.1.3 of "
|
||||
"`msprobe_visualization.md`. Example (minimal runnable snippet):"
|
||||
msgstr "准备一个比较文件,例如 `compare.json`。其格式和生成流程在 `msprobe_visualization.md` 的 3.1.3 节中描述。示例(最小可运行片段):"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:184
|
||||
msgid ""
|
||||
"Replace the paths with your dump directories before invoking `msprobe "
|
||||
"graph_visualize`. **If you only need to build a single graph**, omit "
|
||||
"Replace the paths with your dump directories before invoking `msprobe -f "
|
||||
"pytorch graph`. **If you only need to build a single graph**, omit "
|
||||
"`bench_path` to visualize one dump. Multi-rank scenarios (single rank, "
|
||||
"multi-rank, or multi-step multi-rank) are also supported. `npu_path` or "
|
||||
"`bench_path` must contain folders named `rank+number`, and every rank folder"
|
||||
" must contain a non-empty `construct.json` together with `dump.json` and "
|
||||
"`stack.json`. If any `construct.json` is empty, verify that the dump level "
|
||||
"includes `L0` or `mix`. When comparing graphs, both `npu_path` and "
|
||||
"`bench_path` must contain the same set of rank folders so they can be paired"
|
||||
" one-to-one."
|
||||
"`bench_path` must contain folders named `rank+number`, and every rank "
|
||||
"folder must contain a non-empty `construct.json` together with "
|
||||
"`dump.json` and `stack.json`. If any `construct.json` is empty, verify "
|
||||
"that the dump level includes `L0` or `mix`. When comparing graphs, both "
|
||||
"`npu_path` and `bench_path` must contain the same set of rank folders so "
|
||||
"they can be paired one-to-one."
|
||||
msgstr ""
|
||||
"在执行 `msprobe graph_visualize` 前,将路径替换为实际 dump 目录。**若只需构建单图**,可省略 "
|
||||
"`bench_path`。单 rank、多 rank 以及多 step 多 rank 场景均受支持:`npu_path` 或 `bench_path` "
|
||||
"下必须只有名为 `rank+数字` 的文件夹,并且每个 rank 目录都包含非空的 `construct.json`、`dump.json` 与 "
|
||||
"`stack.json`。若某个 `construct.json` 为空,请确认 dump 级别包含 L0 或 mix。做图比较时,两侧的 rank "
|
||||
"目录数量和名称必须一一对应。"
|
||||
"在调用 `msprobe -f pytorch graph` 之前,将路径替换为你的 dump 目录。**如果只需要构建单个图**,省略 `bench_path` 以可视化一个 dump。多 rank 场景(单 rank、多 rank 或多 step 多 rank)也受支持。`npu_path` 或 `bench_path` 必须包含名为 `rank+数字` 的文件夹,并且每个 rank 文件夹必须包含一个非空的 `construct.json` 以及 `dump.json` 和 `stack.json`。如果任何 `construct.json` 为空,请验证 dump 级别是否包含 `L0` 或 `mix`。比较图时,`npu_path` 和 `bench_path` 必须包含相同的 rank 文件夹集合,以便它们可以一一配对。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:209
|
||||
msgid "Run:"
|
||||
msgstr "执行:"
|
||||
msgstr "运行:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:217
|
||||
msgid ""
|
||||
"After the comparison finishes, a `*.vis.db` file is created under "
|
||||
"`graph_output`."
|
||||
msgstr "对比完成后会在 `graph_output` 下生成 `*.vis.db` 文件。"
|
||||
msgstr "比较完成后,会在 `graph_output` 下创建一个 `*.vis.db` 文件。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:219
|
||||
#, python-brace-format
|
||||
msgid "Graph build: `build_{timestamp}.vis.db`"
|
||||
msgstr "图构建:`build_{timestamp}.vis.db`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:220
|
||||
#, python-brace-format
|
||||
msgid "Graph comparison: `compare_{timestamp}.vis.db`"
|
||||
msgstr "图对比:`compare_{timestamp}.vis.db`"
|
||||
msgstr "图比较:`compare_{timestamp}.vis.db`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:222
|
||||
msgid ""
|
||||
"Launch `tensorboard` and load the output directory to inspect structural "
|
||||
"differences, numerical comparisons, overflow detection results, cross-device"
|
||||
" communication nodes, and filters/search. Pass the directory containing the "
|
||||
"`.vis.db` files to `--logdir`:"
|
||||
"differences, numerical comparisons, overflow detection results, cross-"
|
||||
"device communication nodes, and filters/search. Pass the directory "
|
||||
"containing the `.vis.db` files to `--logdir`:"
|
||||
msgstr ""
|
||||
"启动 `tensorboard` 并加载输出目录,可查看结构差异、精度对比、溢出检测、跨卡通信节点以及多级目录搜索/筛选。将包含 `.vis.db` "
|
||||
"的目录传给 `--logdir`:"
|
||||
"启动 `tensorboard` 并加载输出目录,以检查结构差异、数值比较、溢出检测结果、跨设备通信节点以及过滤器/搜索。将包含 `.vis.db` 文件的目录传递给 `--logdir`:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:228
|
||||
msgid ""
|
||||
"Inspect the visualization. The UI usually displays the overall model "
|
||||
"structure with operators, parameters, and tensor I/O. Click any node to "
|
||||
"expand its children."
|
||||
msgstr "在可视化界面中可查看模型整体结构(算子、参数、张量 I/O),点击节点可展开其子结构。"
|
||||
msgstr "检查可视化界面。UI 通常显示包含算子、参数和张量 I/O 的整体模型结构。点击任何节点以展开其子节点。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:229
|
||||
msgid ""
|
||||
"**Difference visualization**: Comparison results highlight divergent nodes "
|
||||
"with different colors (the larger the difference, the redder the node). "
|
||||
"Click a node to view its detailed information including tensor "
|
||||
"inputs/outputs, parameters, and operator type. Analyze the data difference "
|
||||
"and the surrounding connections to pinpoint the exact divergence."
|
||||
"**Difference visualization**: Comparison results highlight divergent "
|
||||
"nodes with different colors (the larger the difference, the redder the "
|
||||
"node). Click a node to view its detailed information including tensor "
|
||||
"inputs/outputs, parameters, and operator type. Analyze the data "
|
||||
"difference and the surrounding connections to pinpoint the exact "
|
||||
"divergence."
|
||||
msgstr ""
|
||||
"**差异可视化**:对比结果会使用不同颜色突出显示差异节点(差异越大颜色越红)。点击节点可查看输入输出张量、参数以及算子类型,据此结合上下游关系定位具体差异点。"
|
||||
"**差异可视化**:比较结果用不同颜色突出显示差异节点(差异越大,节点越红)。点击节点可查看其详细信息,包括张量输入/输出、参数和算子类型。分析数据差异和周围连接,以精确定位确切的差异点。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:230
|
||||
msgid "**Helper features**:"
|
||||
msgstr "**辅助功能**:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:231
|
||||
msgid ""
|
||||
"Switch rank/step: Quickly check difference nodes on different ranks and "
|
||||
"steps."
|
||||
msgstr "切换 rank/step:快速查看不同 rank 和 step 下的差异节点。"
|
||||
msgstr "切换 rank/step:快速检查不同 rank 和 step 上的差异节点。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:232
|
||||
msgid "Search/filter: Use the search box to filter nodes by operator name, etc."
|
||||
msgstr "搜索/过滤:使用搜索框按算子名称等过滤节点。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:233
|
||||
msgid ""
|
||||
"Search/filter: Use the search box to filter nodes by operator name, etc."
|
||||
msgstr "搜索/筛选:可根据算子名称等快速过滤节点。"
|
||||
"Manual mapping: Automatic mapping cannot cover every case, so the tool "
|
||||
"lets you manually map nodes between the problem and benchmark graphs "
|
||||
"before generating comparison results."
|
||||
msgstr "手动映射:自动映射无法覆盖所有情况,因此该工具允许你在生成比较结果之前,手动映射问题图和标杆图之间的节点。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid ""
|
||||
"Manual mapping: Automatic mapping cannot cover every case, so the tool lets "
|
||||
"you manually map nodes between the problem and benchmark graphs before "
|
||||
"generating comparison results."
|
||||
msgstr "手动映射:当自动映射无法覆盖所有情况时,可手动匹配问题侧与标杆侧节点后再生成对比结果。"
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:235
|
||||
msgid "6. Troubleshooting"
|
||||
msgstr "6. 故障排除"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid "5. Troubleshooting"
|
||||
msgstr "5. 故障排查"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:237
|
||||
msgid ""
|
||||
"`RuntimeError: Please enforce eager mode`: Restart vLLM and add the "
|
||||
"`--enforce-eager` flag."
|
||||
msgstr ""
|
||||
"`RuntimeError: Please enforce eager mode`:重启 vLLM 并加上 `--enforce-eager` 参数。"
|
||||
msgstr "`RuntimeError: Please enforce eager mode`:重启 vLLM 并添加 `--enforce-eager` 标志。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:238
|
||||
msgid ""
|
||||
"No dump files: Confirm that the JSON path is correct and every node has "
|
||||
"write permission. In distributed scenarios set `keep_all_ranks` so that "
|
||||
"every rank writes its own dump."
|
||||
msgstr ""
|
||||
"缺少 dump 文件:检查 JSON 路径是否正确、各节点是否具有写权限;分布式场景可启用 `keep_all_ranks` 让每个 rank "
|
||||
"单独写入。"
|
||||
msgstr "没有 dump 文件:确认 JSON 路径正确且每个节点都有写权限。在分布式场景中,设置 `keep_all_ranks` 以便每个 rank 写入自己的 dump。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:239
|
||||
msgid ""
|
||||
"Dumps are too large: Start with a `statistics` task to locate abnormal "
|
||||
"tensors, then narrow the scope with `scope`/`list`/`tensor_list`, `filters`,"
|
||||
" `token_range`, etc."
|
||||
msgstr ""
|
||||
"dump 体积过大:建议先运行 `statistics` 任务定位异常张量,再通过 "
|
||||
"`scope`/`list`/`tensor_list`、`filters`、`token_range` 等方式缩小范围。"
|
||||
"tensors, then narrow the scope with `scope`/`list`/`tensor_list`, "
|
||||
"`filters`, `token_range`, etc."
|
||||
msgstr "Dump 文件过大:从 `statistics` 任务开始,定位异常张量,然后使用 `scope`/`list`/`tensor_list`、`filters`、`token_range` 等缩小范围。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:243
|
||||
msgid "Appendix"
|
||||
msgstr "附录"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:245
|
||||
msgid "dump.json file description"
|
||||
msgstr "dump.json 文件说明"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:247
|
||||
msgid "L0 level"
|
||||
msgstr "L0 级别"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:249
|
||||
msgid ""
|
||||
"An L0 `dump.json` contains forward/backward I/O for modules together with "
|
||||
"parameters and parameter gradients. Using PyTorch's `Conv2d` as an example, "
|
||||
"the network code looks like:"
|
||||
msgstr ""
|
||||
"L0 级别的 `dump.json` 包含模块的前/反向输入输出以及参数与参数梯度。以下以 PyTorch 的 `Conv2d` 为例,网络代码如下:"
|
||||
"An L0 `dump.json` contains forward I/O for modules together with "
|
||||
"parameters. Using PyTorch's `Conv2d` as an example, the network code "
|
||||
"looks like:"
|
||||
msgstr "L0 级别的 `dump.json` 包含模块的前向 I/O 以及参数。以 PyTorch 的 `Conv2d` 为例,网络代码如下:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:251
|
||||
msgid ""
|
||||
"`output = self.conv2(input) # self.conv2 = torch.nn.Conv2d(64, 128, 5, "
|
||||
"padding=2, bias=True)`"
|
||||
@@ -563,36 +549,19 @@ msgstr ""
|
||||
"`output = self.conv2(input) # self.conv2 = torch.nn.Conv2d(64, 128, 5, "
|
||||
"padding=2, bias=True)`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:253
|
||||
msgid "`dump.json` contains the following entries:"
|
||||
msgstr "`dump.json` 包含以下条目:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:255
|
||||
msgid ""
|
||||
"`Module.conv2.Conv2d.forward.0`: Forward data of the module. `input_args` "
|
||||
"represents positional inputs, `input_kwargs` represents keyword inputs, "
|
||||
"`Module.conv2.Conv2d.forward.0`: Forward data of the module. `input_args`"
|
||||
" represents positional inputs, `input_kwargs` represents keyword inputs, "
|
||||
"`output` stores forward outputs, and `parameters` stores weights/biases."
|
||||
msgstr ""
|
||||
"`Module.conv2.Conv2d.forward.0`:模块的前向数据,`input_args` 为位置参数,`input_kwargs` "
|
||||
"为关键字参数,`output` 存放前向输出,`parameters` 存放权重和偏置。"
|
||||
"`Module.conv2.Conv2d.forward.0`:模块的前向数据。`input_args` 表示位置输入,`input_kwargs` 表示关键字输入,`output` 存储前向输出,`parameters` 存储权重/偏置。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid ""
|
||||
"`Module.conv2.Conv2d.parameters_grad`: Parameter gradients (weight and "
|
||||
"bias)."
|
||||
msgstr "`Module.conv2.Conv2d.parameters_grad`:模块参数的梯度(weight 与 bias)。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid ""
|
||||
"`Module.conv2.Conv2d.backward.0`: Backward data of the module. `input` "
|
||||
"represents gradients that flow into the module (gradients of the forward "
|
||||
"outputs) and `output` represents gradients that flow out (gradients of the "
|
||||
"module inputs)."
|
||||
msgstr ""
|
||||
"`Module.conv2.Conv2d.backward.0`:模块的反向数据,`input` 表示流入模块的梯度(对应前向输出),`output` "
|
||||
"表示流出的梯度(对应模块输入)。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:257
|
||||
#, python-brace-format
|
||||
msgid ""
|
||||
"**Note**: When the `model` parameter passed to the dump API is "
|
||||
@@ -600,47 +569,32 @@ msgid ""
|
||||
"include the index inside the list (`{Module}.{index}.*`). Example: "
|
||||
"`Module.0.conv1.Conv2d.forward.0`."
|
||||
msgstr ""
|
||||
"**说明**:当 dump API 的 `model` 参数为 `List[torch.nn.Module]` 或 "
|
||||
"`Tuple[torch.nn.Module]` 时,模块级名称会包含其在列表中的索引(`{Module}.{index}.*`),例如 "
|
||||
"`Module.0.conv1.Conv2d.forward.0`。"
|
||||
"**注意**:当传递给 dump API 的 `model` 参数是 `List[torch.nn.Module]` 或 `Tuple[torch.nn.Module]` 时,模块级名称包含列表内的索引(`{Module}.{index}.*`)。例如:`Module.0.conv1.Conv2d.forward.0`。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:341
|
||||
msgid "L1 level"
|
||||
msgstr "L1 级别"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:343
|
||||
msgid ""
|
||||
"An L1 `dump.json` records forward/backward I/O for APIs. Using PyTorch's "
|
||||
"`relu` function as an example (`output = torch.nn.functional.relu(input)`), "
|
||||
"the file contains:"
|
||||
msgstr ""
|
||||
"L1 级别的 `dump.json` 记录 API 的前/反向输入输出。以下以 PyTorch 的 `relu` 函数(`output = "
|
||||
"torch.nn.functional.relu(input)`)为例:"
|
||||
"An L1 `dump.json` records forward I/O for APIs. Using PyTorch's `relu` "
|
||||
"function as an example (`output = torch.nn.functional.relu(input)`), the "
|
||||
"file contains:"
|
||||
msgstr "L1 级别的 `dump.json` 记录 API 的前向 I/O。以 PyTorch 的 `relu` 函数为例(`output = torch.nn.functional.relu(input)`),该文件包含:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:345
|
||||
msgid ""
|
||||
"`Functional.relu.0.forward`: Forward data of the API. `input_args` are "
|
||||
"positional inputs, `input_kwargs` are keyword inputs, and `output` stores "
|
||||
"the forward outputs."
|
||||
msgstr ""
|
||||
"`Functional.relu.0.forward`:API 的前向数据,`input_args` 为位置输入,`input_kwargs` "
|
||||
"为关键字输入,`output` 存放前向输出。"
|
||||
"positional inputs, `input_kwargs` are keyword inputs, and `output` stores"
|
||||
" the forward outputs."
|
||||
msgstr "`Functional.relu.0.forward`:API 的前向数据。`input_args` 是位置输入,`input_kwargs` 是关键字输入,`output` 存储前向输出。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
msgid ""
|
||||
"`Functional.relu.0.backward`: Backward data of the API. `input` represents "
|
||||
"the gradients of the forward outputs, and `output` represents the gradients "
|
||||
"that flow back to the forward inputs."
|
||||
msgstr ""
|
||||
"`Functional.relu.0.backward`:API 的反向数据,`input` 表示前向输出的梯度,`output` "
|
||||
"表示回传到前向输入的梯度。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:398
|
||||
msgid "mix level"
|
||||
msgstr "mix 级别"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/msprobe_guide.md
|
||||
#: ../../source/developer_guide/performance_and_debug/msprobe_guide.md:400
|
||||
msgid ""
|
||||
"A `mix` dump.json contains both L0 and L1 level data; the file format is the"
|
||||
" same as the examples above."
|
||||
msgstr "`mix` 级别的 dump.json 同时包含 L0 与 L1 数据,文件格式与上述示例相同。"
|
||||
"A `mix` dump.json contains both L0 and L1 level data; the file format is "
|
||||
"the same as the examples above."
|
||||
msgstr "`mix` 级别的 dump.json 包含 L0 和 L1 级别的数据;文件格式与上述示例相同。"
|
||||
|
||||
@@ -0,0 +1,348 @@
|
||||
# SOME DESCRIPTIVE TITLE.
|
||||
# Copyright (C) 2025, vllm-ascend team
|
||||
# This file is distributed under the same license as the vllm-ascend
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2026.
|
||||
#
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend \n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:1
|
||||
msgid "Optimization and Tuning"
|
||||
msgstr "优化与调优"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:3
|
||||
msgid ""
|
||||
"This guide aims to help users improve vLLM-Ascend performance at the "
|
||||
"system level. It includes OS configuration, library optimization, "
|
||||
"deployment guide, and so on. Any feedback is welcome."
|
||||
msgstr "本指南旨在帮助用户在系统层面提升 vLLM-Ascend 的性能。内容包括操作系统配置、库优化、部署指南等。欢迎提供任何反馈。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:5
|
||||
msgid "Preparation"
|
||||
msgstr "准备工作"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:7
|
||||
msgid "Run the container:"
|
||||
msgstr "运行容器:"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:31
|
||||
msgid "Configure your environment:"
|
||||
msgstr "配置您的环境:"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:49
|
||||
msgid "Install vllm and vllm-ascend:"
|
||||
msgstr "安装 vllm 和 vllm-ascend:"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:61
|
||||
msgid ""
|
||||
"Please follow the [Installation "
|
||||
"Guide](https://docs.vllm.ai/projects/ascend/en/latest/installation.html) "
|
||||
"to make sure vLLM and vllm-ascend are installed correctly."
|
||||
msgstr "请遵循[安装指南](https://docs.vllm.ai/projects/ascend/en/latest/installation.html)以确保 vLLM 和 vllm-ascend 正确安装。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:64
|
||||
msgid ""
|
||||
"Make sure your vLLM and vllm-ascend are installed after your Python "
|
||||
"configuration is completed, because these packages will build binary "
|
||||
"files using python in current environment. If you install vLLM and vllm-"
|
||||
"ascend before completing section 1.1, the binary files will not use the "
|
||||
"optimized python."
|
||||
msgstr "请确保在完成 Python 配置后再安装 vLLM 和 vllm-ascend,因为这些软件包将使用当前环境中的 python 构建二进制文件。如果您在完成第 1.1 节之前就安装了 vLLM 和 vllm-ascend,则二进制文件将不会使用优化后的 python。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:67
|
||||
msgid "Optimizations"
|
||||
msgstr "优化措施"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:69
|
||||
msgid "1. Compilation Optimization"
|
||||
msgstr "1. 编译优化"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:71
|
||||
msgid "1.1. Install optimized `python`"
|
||||
msgstr "1.1. 安装优化版 `python`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:73
|
||||
msgid ""
|
||||
"Python supports **LTO** and **PGO** optimization starting from version "
|
||||
"`3.6` and above, which can be enabled at compile time. And we have "
|
||||
"offered optimized `python` packages directly to users for the sake of "
|
||||
"convenience. You can also reproduce the `python` build following this "
|
||||
"[tutorial](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0063.html)"
|
||||
" according to your specific scenarios."
|
||||
msgstr "Python 从 `3.6` 及以上版本开始支持 **LTO** 和 **PGO** 优化,可以在编译时启用。为了方便用户,我们直接提供了优化版的 `python` 软件包。您也可以根据具体场景,按照此[教程](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0063.html)自行构建 `python`。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:101
|
||||
msgid "2. OS Optimization"
|
||||
msgstr "2. 操作系统优化"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:103
|
||||
msgid "2.1. jemalloc"
|
||||
msgstr "2.1. jemalloc"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:105
|
||||
msgid ""
|
||||
"**jemalloc** is a memory allocator that improves performance for multi-"
|
||||
"threaded scenarios and can reduce memory fragmentation. jemalloc uses a "
|
||||
"local thread memory manager to allocate variables, which can avoid lock "
|
||||
"competition between threads and can hugely optimize performance."
|
||||
msgstr "**jemalloc** 是一个内存分配器,可提升多线程场景下的性能并减少内存碎片。jemalloc 使用本地线程内存管理器来分配变量,这可以避免线程间的锁竞争,从而大幅优化性能。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:117
|
||||
msgid "2.2. Tcmalloc"
|
||||
msgstr "2.2. Tcmalloc"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:119
|
||||
msgid ""
|
||||
"**TCMalloc (Thread Caching Malloc)** is a universal memory allocator that"
|
||||
" improves overall performance while ensuring low latency by introducing a"
|
||||
" multi-level cache structure, reducing mutex contention and optimizing "
|
||||
"large object processing flow. Find more "
|
||||
"[details](https://www.hiascend.com/document/detail/zh/Pytorch/700/ptmoddevg/trainingmigrguide/performance_tuning_0068.html)."
|
||||
msgstr "**TCMalloc (Thread Caching Malloc)** 是一个通用内存分配器,通过引入多级缓存结构、减少互斥锁竞争以及优化大对象处理流程,在确保低延迟的同时提升整体性能。更多[详情](https://www.hiascend.com/document/detail/zh/Pytorch/700/ptmoddevg/trainingmigrguide/performance_tuning_0068.html)。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:140
|
||||
msgid "3. `torch_npu` Optimization"
|
||||
msgstr "3. `torch_npu` 优化"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:142
|
||||
msgid ""
|
||||
"Some performance tuning features in `torch_npu` are controlled by "
|
||||
"environment variables. Some features and their related environment "
|
||||
"variables are shown below."
|
||||
msgstr "`torch_npu` 中的一些性能调优功能由环境变量控制。部分功能及其相关环境变量如下所示。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:144
|
||||
msgid "Memory optimization:"
|
||||
msgstr "内存优化:"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:155
|
||||
msgid "Scheduling optimization:"
|
||||
msgstr "调度优化:"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:166
|
||||
msgid "4. CANN Optimization"
|
||||
msgstr "4. CANN 优化"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:168
|
||||
msgid "4.1. HCCL Optimization"
|
||||
msgstr "4.1. HCCL 优化"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:170
|
||||
msgid ""
|
||||
"There are some performance tuning features in HCCL, which are controlled "
|
||||
"by environment variables."
|
||||
msgstr "HCCL 中有一些性能调优功能,由环境变量控制。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:172
|
||||
msgid ""
|
||||
"You can configure HCCL to use \"AIV\" mode to optimize performance by "
|
||||
"setting the environment variable shown below. In \"AIV\" mode, the "
|
||||
"communication is scheduled by AI vector core directly with RoCE, instead "
|
||||
"of being scheduled by AI CPU."
|
||||
msgstr "您可以通过设置如下所示的环境变量,将 HCCL 配置为使用 \"AIV\" 模式以优化性能。在 \"AIV\" 模式下,通信由 AI 向量核通过 RoCE 直接调度,而非由 AI CPU 调度。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:179
|
||||
msgid ""
|
||||
"Plus, there are more features for performance optimization in specific "
|
||||
"scenarios, which are shown below."
|
||||
msgstr "此外,针对特定场景还有更多性能优化功能,如下所示。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:181
|
||||
msgid ""
|
||||
"`HCCL_INTRA_ROCE_ENABLE`: Use RDMA link instead of SDMA link between two "
|
||||
"8Ps as the mesh interconnect link. Find more "
|
||||
"[details](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0044.html)."
|
||||
msgstr "`HCCL_INTRA_ROCE_ENABLE`:在两个 8P 之间使用 RDMA 链路而非 SDMA 链路作为网状互连链路。更多[详情](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0044.html)。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:182
|
||||
msgid ""
|
||||
"`HCCL_RDMA_TC`: Use this var to configure traffic class of RDMA NIC. Find"
|
||||
" more "
|
||||
"[details](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0045.html)."
|
||||
msgstr "`HCCL_RDMA_TC`:使用此变量配置 RDMA 网卡的流量类别。更多[详情](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0045.html)。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:183
|
||||
msgid ""
|
||||
"`HCCL_RDMA_SL`: Use this var to configure service level of RDMA NIC. Find"
|
||||
" more "
|
||||
"[details](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0046.html)."
|
||||
msgstr "`HCCL_RDMA_SL`:使用此变量配置 RDMA 网卡的服务级别。更多[详情](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0046.html)。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:184
|
||||
msgid ""
|
||||
"`HCCL_BUFFSIZE`: Use this var to control the cache size for sharing data "
|
||||
"between two NPUs. Find more "
|
||||
"[details](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0047.html)."
|
||||
msgstr "`HCCL_BUFFSIZE`:使用此变量控制两个 NPU 之间共享数据的缓存大小。更多[详情](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0047.html)。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:186
|
||||
msgid "5. OS Optimization"
|
||||
msgstr "5. 操作系统优化"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:188
|
||||
msgid ""
|
||||
"This section describes operating system–level optimizations applied on "
|
||||
"the host machine (bare metal or Kubernetes node) to improve performance "
|
||||
"stability, latency, and throughput for inference workloads."
|
||||
msgstr "本节描述了在主机(裸机或 Kubernetes 节点)上应用的操作系统级优化,旨在提升推理工作负载的性能稳定性、延迟和吞吐量。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:191
|
||||
msgid ""
|
||||
"These settings must be applied on the host OS and with root privileges. "
|
||||
"Not inside containers."
|
||||
msgstr "这些设置必须在主机操作系统上以 root 权限应用,而不是在容器内部。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:194
|
||||
msgid "5.1"
|
||||
msgstr "5.1"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:196
|
||||
msgid "Set CPU Frequency Governor to `performance`"
|
||||
msgstr "将 CPU 频率调节器设置为 `performance`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:202
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:219
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:239
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:261
|
||||
msgid "Purpose"
|
||||
msgstr "目的"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:204
|
||||
msgid "Forces all CPU cores to run under the `performance` governor"
|
||||
msgstr "强制所有 CPU 核心在 `performance` 调节器下运行"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:205
|
||||
msgid "Disables dynamic frequency scaling (e.g., `ondemand`, `powersave`)"
|
||||
msgstr "禁用动态频率调节(例如 `ondemand`、`powersave`)"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:207
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:223
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:243
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:265
|
||||
msgid "Benefits"
|
||||
msgstr "优势"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:209
|
||||
msgid "Keeps CPU cores at maximum frequency"
|
||||
msgstr "使 CPU 核心保持最高频率"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:210
|
||||
msgid "Reduces latency jitter"
|
||||
msgstr "减少延迟抖动"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:211
|
||||
msgid "Improves predictability for inference workloads"
|
||||
msgstr "提高推理工作负载的可预测性"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:213
|
||||
msgid "5.2 Disable Swap Usage"
|
||||
msgstr "5.2 禁用交换空间使用"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:221
|
||||
msgid "Minimizes the kernel’s tendency to swap memory pages to disk"
|
||||
msgstr "最小化内核将内存页交换到磁盘的倾向"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:225
|
||||
msgid "Prevents severe latency spikes caused by swapping"
|
||||
msgstr "防止因交换导致的严重延迟峰值"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:226
|
||||
msgid "Improves stability for large in-memory models"
|
||||
msgstr "提高大型内存模型的稳定性"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:228
|
||||
msgid "Notes"
|
||||
msgstr "备注"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:230
|
||||
msgid "For inference workloads, swap can introduce second-level latency"
|
||||
msgstr "对于推理工作负载,交换可能导致秒级延迟"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:231
|
||||
msgid "Recommended values are `0` or `1`"
|
||||
msgstr "推荐值为 `0` 或 `1`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:233
|
||||
msgid "5.3 Disable Automatic NUMA Balancing"
|
||||
msgstr "5.3 禁用自动 NUMA 平衡"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:241
|
||||
msgid "Disables the kernel’s automatic NUMA page migration mechanism"
|
||||
msgstr "禁用内核的自动 NUMA 页面迁移机制"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:245
|
||||
msgid "Prevents background memory page migrations"
|
||||
msgstr "防止后台内存页迁移"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:246
|
||||
msgid "Reduces unpredictable memory access latency"
|
||||
msgstr "减少不可预测的内存访问延迟"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:247
|
||||
msgid "Improves performance stability on NUMA systems"
|
||||
msgstr "提高 NUMA 系统上的性能稳定性"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:249
|
||||
msgid "Recommended For"
|
||||
msgstr "推荐用于"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:251
|
||||
msgid "Multi-socket servers"
|
||||
msgstr "多插槽服务器"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:252
|
||||
msgid "Ascend / NPU deployments with explicit NUMA binding"
|
||||
msgstr "具有显式 NUMA 绑定的 Ascend / NPU 部署"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:253
|
||||
msgid "Systems with manually managed CPU and memory affinity"
|
||||
msgstr "手动管理 CPU 和内存亲和性的系统"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:255
|
||||
msgid "5.4 Increase Scheduler Migration Cost"
|
||||
msgstr "5.4 增加调度器迁移成本"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:263
|
||||
msgid "Increases the cost for the scheduler to migrate tasks between CPU cores"
|
||||
msgstr "增加调度器在 CPU 核心间迁移任务的成本"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:267
|
||||
msgid "Reduces frequent thread migration"
|
||||
msgstr "减少频繁的线程迁移"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:268
|
||||
msgid "Improves CPU cache locality"
|
||||
msgstr "提高 CPU 缓存局部性"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:269
|
||||
msgid "Lowers latency jitter for inference workloads"
|
||||
msgstr "降低推理工作负载的延迟抖动"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:271
|
||||
msgid "Parameter Details"
|
||||
msgstr "参数详情"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:273
|
||||
msgid "Unit: nanoseconds (ns)"
|
||||
msgstr "单位:纳秒 (ns)"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:274
|
||||
msgid "Typical recommended range: 50000–100000"
|
||||
msgstr "典型推荐范围:50000–100000"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/optimization_and_tuning.md:275
|
||||
msgid "Higher values encourage threads to stay on the same CPU core"
|
||||
msgstr "更高的值鼓励线程保持在同一个 CPU 核心上"
|
||||
@@ -4,85 +4,338 @@
|
||||
# package.
|
||||
# FIRST AUTHOR <EMAIL@ADDRESS>, 2025.
|
||||
#
|
||||
#, fuzzy
|
||||
msgid ""
|
||||
msgstr ""
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Project-Id-Version: vllm-ascend\n"
|
||||
"Report-Msgid-Bugs-To: \n"
|
||||
"POT-Creation-Date: 2025-07-18 09:01+0800\n"
|
||||
"POT-Creation-Date: 2026-04-14 09:08+0000\n"
|
||||
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
|
||||
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Language: zh_CN\n"
|
||||
"Language-Team: zh_CN <LL@li.org>\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"MIME-Version: 1.0\n"
|
||||
"Content-Type: text/plain; charset=utf-8\n"
|
||||
"Content-Transfer-Encoding: 8bit\n"
|
||||
"Plural-Forms: nplurals=1; plural=0;\n"
|
||||
"Generated-By: Babel 2.17.0\n"
|
||||
"Generated-By: Babel 2.18.0\n"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/performance_benchmark.md:1
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:1
|
||||
msgid "Performance Benchmark"
|
||||
msgstr "性能基准"
|
||||
msgstr "性能基准测试"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/performance_benchmark.md:2
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:3
|
||||
msgid ""
|
||||
"This document details the benchmark methodology for vllm-ascend, aimed at "
|
||||
"evaluating the performance under a variety of workloads. To maintain "
|
||||
"This document details the benchmark methodology for vllm-ascend, aimed at"
|
||||
" evaluating the performance under a variety of workloads. To maintain "
|
||||
"alignment with vLLM, we use the [benchmark](https://github.com/vllm-"
|
||||
"project/vllm/tree/main/benchmarks) script provided by the vllm project."
|
||||
msgstr ""
|
||||
"本文档详细说明了 vllm-ascend 的基准测试方法,旨在评估其在多种工作负载下的性能。为了与 vLLM 保持一致,我们使用 vllm 项目提供的 "
|
||||
"[benchmark](https://github.com/vllm-project/vllm/tree/main/benchmarks) 脚本。"
|
||||
"本文档详细说明了 vllm-ascend 的基准测试方法,旨在评估其在多种工作负载下的性能。为了与 vLLM 保持一致,我们使用 vllm "
|
||||
"项目提供的 [benchmark](https://github.com/vllm-"
|
||||
"project/vllm/tree/main/benchmarks) 脚本。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/performance_benchmark.md:4
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:5
|
||||
msgid ""
|
||||
"**Benchmark Coverage**: We measure offline e2e latency and throughput, and "
|
||||
"fixed-QPS online serving benchmarks, for more details see [vllm-ascend "
|
||||
"benchmark scripts](https://github.com/vllm-project/vllm-"
|
||||
"**Benchmark Coverage**: We measure offline E2E latency and throughput, "
|
||||
"and fixed-QPS online serving benchmarks. For more details, see [vllm-"
|
||||
"ascend benchmark scripts](https://github.com/vllm-project/vllm-"
|
||||
"ascend/tree/main/benchmarks)."
|
||||
msgstr ""
|
||||
"**基准测试覆盖范围**:我们测量离线端到端延迟和吞吐量,以及固定 QPS 的在线服务基准测试。更多详情请参见 [vllm-ascend "
|
||||
"基准测试脚本](https://github.com/vllm-project/vllm-ascend/tree/main/benchmarks)。"
|
||||
"基准测试脚本](https://github.com/vllm-project/vllm-"
|
||||
"ascend/tree/main/benchmarks)。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/performance_benchmark.md:6
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:7
|
||||
msgid "**Legend Description**:"
|
||||
msgstr "**图例说明**:"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:9
|
||||
msgid "✅ = Supported"
|
||||
msgstr "✅ = 已支持"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:10
|
||||
msgid "🟡 = Partial / Work in progress"
|
||||
msgstr "🟡 = 部分支持 / 开发中"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:11
|
||||
msgid "🚧 = Under development"
|
||||
msgstr "🚧 = 开发中"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:13
|
||||
msgid "1. Run docker container"
|
||||
msgstr "1. 运行 docker 容器"
|
||||
msgstr "1. 运行 Docker 容器"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/performance_benchmark.md:31
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:39
|
||||
msgid "2. Install dependencies"
|
||||
msgstr "2. 安装依赖项"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/performance_benchmark.md:38
|
||||
msgid "3. (Optional)Prepare model weights"
|
||||
msgstr "3.(可选)准备模型权重"
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:47
|
||||
msgid "3. Run basic benchmarks"
|
||||
msgstr "3. 运行基础基准测试"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/performance_benchmark.md:39
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:49
|
||||
msgid ""
|
||||
"For faster running speed, we recommend downloading the model in advance:"
|
||||
msgstr "为了更快的运行速度,建议提前下载模型:"
|
||||
"This section introduces how to perform performance testing using the "
|
||||
"benchmark suite built into VLLM."
|
||||
msgstr "本节介绍如何使用 VLLM 内置的基准测试套件进行性能测试。"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/performance_benchmark.md:44
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:51
|
||||
msgid "3.1 Dataset"
|
||||
msgstr "3.1 数据集"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:53
|
||||
msgid ""
|
||||
"You can also replace all model paths in the [json](https://github.com/vllm-"
|
||||
"project/vllm-ascend/tree/main/benchmarks/tests) files with your local paths:"
|
||||
"VLLM supports a variety of [datasets](https://github.com/vllm-"
|
||||
"project/vllm/blob/main/vllm/benchmarks/datasets.py)."
|
||||
msgstr "VLLM 支持多种[数据集](https://github.com/vllm-project/vllm/blob/main/vllm/benchmarks/datasets.py)。"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "Dataset"
|
||||
msgstr "数据集"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "Online"
|
||||
msgstr "在线"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "Offline"
|
||||
msgstr "离线"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "Data Path"
|
||||
msgstr "数据路径"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "ShareGPT"
|
||||
msgstr "ShareGPT"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "✅"
|
||||
msgstr "✅"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid ""
|
||||
"`wget "
|
||||
"https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.json`"
|
||||
msgstr ""
|
||||
"你也可以将 [json](https://github.com/vllm-project/vllm-"
|
||||
"ascend/tree/main/benchmarks/tests) 文件中的所有模型路径替换为你的本地路径:"
|
||||
"`wget "
|
||||
"https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.json`"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/performance_benchmark.md:60
|
||||
msgid "4. Run benchmark script"
|
||||
msgstr "4. 运行基准测试脚本"
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "ShareGPT4V (Image)"
|
||||
msgstr "ShareGPT4V (图像)"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/performance_benchmark.md:61
|
||||
msgid "Run benchmark script:"
|
||||
msgstr "运行基准测试脚本:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/performance_benchmark.md:66
|
||||
msgid "After about 10 mins, the output is as shown below:"
|
||||
msgstr "大约 10 分钟后,输出如下所示:"
|
||||
|
||||
#: ../../developer_guide/performance_and_debug/performance_benchmark.md:176
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid ""
|
||||
"The result json files are generated into the path `benchmark/results` These "
|
||||
"files contain detailed benchmarking results for further analysis."
|
||||
msgstr "结果 json 文件会生成到路径 `benchmark/results`。这些文件包含了用于进一步分析的详细基准测试结果。"
|
||||
"`wget https://huggingface.co/datasets/Lin-"
|
||||
"Chen/ShareGPT4V/resolve/main/sharegpt4v_instruct_gpt4-vision_cap100k.json`<br>Note"
|
||||
" that the images need to be downloaded separately. For example, to "
|
||||
"download COCO's 2017 Train images:<br>`wget "
|
||||
"http://images.cocodataset.org/zips/train2017.zip`"
|
||||
msgstr ""
|
||||
"`wget https://huggingface.co/datasets/Lin-"
|
||||
"Chen/ShareGPT4V/resolve/main/sharegpt4v_instruct_gpt4-vision_cap100k.json`<br>请注意,图像需要单独下载。例如,要下载"
|
||||
" COCO 2017 训练集图像:<br>`wget "
|
||||
"http://images.cocodataset.org/zips/train2017.zip`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "ShareGPT4Video (Video)"
|
||||
msgstr "ShareGPT4Video (视频)"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "`git clone https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video`"
|
||||
msgstr "`git clone https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "BurstGPT"
|
||||
msgstr "BurstGPT"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid ""
|
||||
"`wget "
|
||||
"https://github.com/HPMLL/BurstGPT/releases/download/v1.1/BurstGPT_without_fails_2.csv`"
|
||||
msgstr ""
|
||||
"`wget "
|
||||
"https://github.com/HPMLL/BurstGPT/releases/download/v1.1/BurstGPT_without_fails_2.csv`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "Sonnet (deprecated)"
|
||||
msgstr "Sonnet (已弃用)"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "Local file: `benchmarks/sonnet.txt`"
|
||||
msgstr "本地文件:`benchmarks/sonnet.txt`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "Random"
|
||||
msgstr "随机"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "`synthetic`"
|
||||
msgstr "`synthetic`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "RandomMultiModal (Image/Video)"
|
||||
msgstr "RandomMultiModal (图像/视频)"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "🟡"
|
||||
msgstr "🟡"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "🚧"
|
||||
msgstr "🚧"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "RandomForReranking"
|
||||
msgstr "RandomForReranking"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "Prefix Repetition"
|
||||
msgstr "前缀重复"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "HuggingFace-VisionArena"
|
||||
msgstr "HuggingFace-VisionArena"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "`lmarena-ai/VisionArena-Chat`"
|
||||
msgstr "`lmarena-ai/VisionArena-Chat`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "HuggingFace-MMVU"
|
||||
msgstr "HuggingFace-MMVU"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "`yale-nlp/MMVU`"
|
||||
msgstr "`yale-nlp/MMVU`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "HuggingFace-InstructCoder"
|
||||
msgstr "HuggingFace-InstructCoder"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "`likaixin/InstructCoder`"
|
||||
msgstr "`likaixin/InstructCoder`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "HuggingFace-AIMO"
|
||||
msgstr "HuggingFace-AIMO"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid ""
|
||||
"`AI-MO/aimo-validation-aime`, `AI-MO/NuminaMath-1.5`, `AI-MO/NuminaMath-"
|
||||
"CoT`"
|
||||
msgstr "`AI-MO/aimo-validation-aime`, `AI-MO/NuminaMath-1.5`, `AI-MO/NuminaMath-CoT`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "HuggingFace-Other"
|
||||
msgstr "HuggingFace-其他"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "`lmms-lab/LLaVA-OneVision-Data`, `Aeala/ShareGPT_Vicuna_unfiltered`"
|
||||
msgstr "`lmms-lab/LLaVA-OneVision-Data`, `Aeala/ShareGPT_Vicuna_unfiltered`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "HuggingFace-MTBench"
|
||||
msgstr "HuggingFace-MTBench"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "`philschmid/mt-bench`"
|
||||
msgstr "`philschmid/mt-bench`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "HuggingFace-Blazedit"
|
||||
msgstr "HuggingFace-Blazedit"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "`vdaita/edit_5k_char`, `vdaita/edit_10k_char`"
|
||||
msgstr "`vdaita/edit_5k_char`, `vdaita/edit_10k_char`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "Spec Bench"
|
||||
msgstr "Spec Bench"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid ""
|
||||
"`wget https://raw.githubusercontent.com/hemingkx/Spec-"
|
||||
"Bench/refs/heads/main/data/spec_bench/question.jsonl`"
|
||||
msgstr "`wget https://raw.githubusercontent.com/hemingkx/Spec-Bench/refs/heads/main/data/spec_bench/question.jsonl`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "Custom"
|
||||
msgstr "自定义"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:15
|
||||
msgid "Local file: `data.jsonl`"
|
||||
msgstr "本地文件:`data.jsonl`"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:83
|
||||
msgid ""
|
||||
"The datasets mentioned above are all links to datasets on huggingface. "
|
||||
"The dataset's `dataset-name` should be set to `hf`. For local `dataset-"
|
||||
"path`, please set `hf-name` to its Hugging Face ID like"
|
||||
msgstr "上述提到的数据集均为 Hugging Face 上数据集的链接。数据集的 `dataset-name` 应设置为 `hf`。对于本地的 `dataset-path`,请将 `hf-name` 设置为其 Hugging Face ID,例如:"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:93
|
||||
msgid "3.2 Run basic benchmark"
|
||||
msgstr "3.2 运行基础基准测试"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:95
|
||||
msgid "3.2.1 Online serving"
|
||||
msgstr "3.2.1 在线服务"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:97
|
||||
msgid "First start serving your model:"
|
||||
msgstr "首先启动模型服务:"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:103
|
||||
msgid "Then run the benchmarking script:"
|
||||
msgstr "然后运行基准测试脚本:"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:118
|
||||
msgid "If successful, you will see the following output:"
|
||||
msgstr "如果成功,您将看到以下输出:"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:147
|
||||
msgid "3.2.2 Offline Throughput Benchmark"
|
||||
msgstr "3.2.2 离线吞吐量基准测试"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:158
|
||||
msgid "If successful, you will see the following output"
|
||||
msgstr "如果成功,您将看到以下输出"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:167
|
||||
msgid "3.2.4 Multi-Modal Benchmark"
|
||||
msgstr "3.2.4 多模态基准测试"
|
||||
|
||||
#: ../../source/developer_guide/performance_and_debug/performance_benchmark.md:216
|
||||
msgid "3.2.5 Embedding Benchmark"
|
||||
msgstr "3.2.5 嵌入基准测试"
|
||||
|
||||
#~ msgid "3. (Optional)Prepare model weights"
|
||||
#~ msgstr "3.(可选)准备模型权重"
|
||||
|
||||
#~ msgid ""
|
||||
#~ "For faster running speed, we recommend"
|
||||
#~ " downloading the model in advance:"
|
||||
#~ msgstr "为了获得更快的运行速度,我们建议提前下载模型:"
|
||||
|
||||
#~ msgid ""
|
||||
#~ "You can also replace all model "
|
||||
#~ "paths in the [json](https://github.com/vllm-"
|
||||
#~ "project/vllm-ascend/tree/main/benchmarks/tests) files "
|
||||
#~ "with your local paths:"
|
||||
#~ msgstr ""
|
||||
#~ "您也可以将 [json](https://github.com/vllm-project/vllm-"
|
||||
#~ "ascend/tree/main/benchmarks/tests) 文件中的所有模型路径替换为您的本地路径:"
|
||||
|
||||
#~ msgid "After about 10 mins, the output is as shown below:"
|
||||
#~ msgstr "大约 10 分钟后,输出如下所示:"
|
||||
|
||||
#~ msgid ""
|
||||
#~ "The result json files are generated "
|
||||
#~ "into the path `benchmark/results` These "
|
||||
#~ "files contain detailed benchmarking results"
|
||||
#~ " for further analysis."
|
||||
#~ msgstr "结果 JSON 文件将生成到路径 `benchmark/results`。这些文件包含详细的基准测试结果,可用于进一步分析。"
|
||||
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user