xc-llm-ascend

Files

csoulnd 97dbcaf919 [BugFix][310P][v0.18.0] Use CPU generator cache for sampling (#8624 )

### What this PR does / why we need it?
This PR introduces a caching mechanism for CPU-based `torch.Generator`
objects in the `_random_sample_310p` function to optimize sampling
performance. It includes unit tests for cache persistence and state
recovery. Feedback highlights a critical bug where keying the cache by
batch index instead of generator ID can break RNG reproducibility during
request re-scheduling, and notes a potential memory leak in the global
cache.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Tested via new unit tests in `tests/ut/_310p/sample/test_sampler_310.py`
verifying cache logic and error handling.

---------

Signed-off-by: csoulnd <daidaicurry@foxmail.com>

2026-04-24 09:34:14 +08:00

e2e

[Doc][releases/v0.18.0] fix documentation error or non-standard description (#8626 )

2026-04-23 18:55:44 +08:00

[BugFix][310P][v0.18.0] Use CPU generator cache for sampling (#8624 )

2026-04-24 09:34:14 +08:00

__init__.py

[SpecDecode] Add spec decode support (#500 )

2025-04-17 20:16:32 +08:00