xc-llm-ascend/tutorials at 195eac665b2d42b8287c59128490586c6931d54c - xc-llm-ascend - Gitea: Git with a cup of tea

EngineX/xc-llm-ascend

Files

History

InSec a5cb8e40f5 [doc]Modify quantization tutorials (#5026 )

### What this PR does / why we need it?
Modify quantization tutorials to correct a few mistakes:
Qwen3-32B-W4A4.md and Qwen3-8B-W4A8.md
Qwen3-8B-W4A8: need to set one idle npu card.
Qwen3-32B-W4A4: need to set two idle npu cards for the flatquant
training and modify the calib_file path which does not match the
ModeSlim version.
### Does this PR introduce _any_ user-facing change?
N/A
### How was this patch tested?

- vLLM version: v0.12.0
- vLLM main:
ad32e3e19c

Signed-off-by: IncSec <1790766300@qq.com>

2025-12-15 20:12:06 +08:00

..

310p.md

[Doc] Update tutorial index (#4920 )

2025-12-11 20:53:13 +08:00

DeepSeek-R1.md

[Doc] Update tutorial index (#4920 )

2025-12-11 20:53:13 +08:00

DeepSeek-V3.1.md

[Doc] Update tutorial index (#4920 )

2025-12-11 20:53:13 +08:00

DeepSeek-V3.2.md

add release note for 0.12.0 (#4995 )

2025-12-13 22:09:59 +08:00

index.md

[doc][main] Correct more doc mistakes (#4958 )

2025-12-13 18:36:58 +08:00

Kimi-K2-Thinking.md

[Doc] Update tutorial index (#4920 )

2025-12-11 20:53:13 +08:00

pd_disaggregation_mooncake_multi_node.md

[Doc] Update tutorial index (#4920 )

2025-12-11 20:53:13 +08:00

pd_disaggregation_mooncake_single_node.md

update qwen2.5vl readme (#4938 )

2025-12-12 15:40:07 +08:00

Qwen2.5-7B.md

[Doc] Update tutorial index (#4920 )

2025-12-11 20:53:13 +08:00

Qwen2.5-Omni.md

[Doc] Add single NPU tutorial for Qwen2.5-Omni-7B (#4446 )

2025-11-29 11:57:29 +08:00

Qwen3_embedding.md

[Misc] Update pooling example (#5002 )

2025-12-15 08:36:19 +08:00

Qwen3-8B-W4A8.md

[doc]Modify quantization tutorials (#5026 )

2025-12-15 20:12:06 +08:00

Qwen3-30B-A3B.md

[Doc] Update tutorial index (#4920 )

2025-12-11 20:53:13 +08:00

Qwen3-32B-W4A4.md

[doc]Modify quantization tutorials (#5026 )

2025-12-15 20:12:06 +08:00

Qwen3-235B-A22B.md

[Doc] Update tutorial index (#4920 )

2025-12-11 20:53:13 +08:00

Qwen3-Coder-30B-A3B.md

[Doc] Add tutorial for Qwen3-Coder-30B-A3B (#4391 )

2025-12-02 16:03:37 +08:00

Qwen3-Dense.md

[Doc] Update tutorial index (#4920 )

2025-12-11 20:53:13 +08:00

Qwen3-Next.md

Add Qwen3-Next tutorials (#4607 )

2025-12-15 11:48:22 +08:00

Qwen-VL-Dense.md

[doc][main] Correct mistakes in doc (#4945 )

2025-12-12 19:17:10 +08:00

ray.md

[doc][main] Correct more doc mistakes (#4958 )

2025-12-13 18:36:58 +08:00