xc-llm-ascend

Files

sdmyzlp e72f94e38f Support multistream of MLA vector operations (#1135 )

### What this PR does / why we need it?
Move all vector operations to a secondary stream, with the expected
overlaping being:
```
              | q_rmsnorm |                  | kv_norm_rope_cache |       | q_rope |
| matmul W_DQ | matmul W_DKV | index | index |    matmul W_UQ     | split | matmul W_KV_T |
```

Currently, the `IndexByTensor` operators introduced by computation of
`cos` and `sin` can't be offloaded to the secondary stream due to a
known bug of graph fusion optimization pass. So we instead keep it in
the main stream, only requires it be computed before `matmul W_UQ` to
avoid hindering later overlapping. The problem may be solved by later
optimization (#993), which hoists the computation of `cos` and `sin` up
to the first layer.

### Does this PR introduce _any_ user-facing change?
Controlled by `torchair_graph_config.enable_multistream_mla`, defaulted
to False.

### How was this patch tested?
Tested on 1x16 910 node, with tailored 2 layer DSKv2.

Signed-off-by: sdmyzlp <lrwei2@petalmail.com>

2025-06-12 21:42:09 +08:00

additional_config.md

Support multistream of MLA vector operations (#1135 )

2025-06-12 21:42:09 +08:00

env_vars.md

[Doc] Add environment variables doc (#519 )

2025-04-15 16:09:36 +08:00

graph_mode.md

[Doc] Fix the config parameter name "enable" in graph_mode.md. (#1159 )

2025-06-11 11:03:37 +08:00

release_notes.md

[Bugfix] add compilation/__init__.py to fix import error (#1152 )