xc-llm-ascend

Files

Mengqing Cao b64ee7d346 [Dist] Set device as rank (#202 )

### What this PR does / why we need it?
The rank returned by `torch.distributed.get_rank(device_group)` is the
local rank, but rank (or rank in process group (PG)) is expected.
Thus we change to use `torch.npu.current_device()` to set device

```python
    # difference between `local_rank` and `rank_in_group`:
    # if we have a group of size 4 across two nodes:
    # Process | Node | Rank | Local Rank | Rank in Group
    #   0     |   0  |  0   |     0      |       0
    #   1     |   0  |  1   |     1      |       1
    #   2     |   1  |  2   |     0      |       2
    #   3     |   1  |  3   |     1      |       3
```

Tested by @wwfu109 with
`vllm/tests/distributed/test_customops::test_multi_process_tensor_parallel_pipeline_parallel`

Signed-off-by: MengqingCao <cmq0113@163.com>

2025-03-03 09:23:13 +08:00

ops

[CI] Fix unsolved bugs caused by pta api change. (#190 )

2025-02-27 19:52:28 +08:00

quantization

[Core] Cherry pick from 0.7.1 to keep the main code newest (#127 )

2025-02-21 17:07:37 +08:00

__init__.py

[Core] Init vllm-ascend (#3 )

2025-02-05 10:53:12 +08:00

attention.py

[CI] Upgrade to newest pta.(MLA and FusedMoE) (#189 )