[v0.11.0][Fix] Prevent memory leak in MLA decode graph (#3743) (#3774)

### What this PR does / why we need it? The cache for MLA decode graph parameters was holding strong references to tensors, preventing them from being garbage collected and leading to increased memory usage. This change wraps the cached tensors in weak references, allowing them to be deallocated when no longer in use and reducing overall memory pressure. ### Does this PR introduce _any_ user-facing change? None. ### How was this patch tested? None. Signed-off-by: Yizhou Liu <liu_yizhou@outlook.com>
2025-10-27 16:00:20 +08:00
parent 825fdfb197
commit 43276fd822
4 changed files with 26 additions and 16 deletions
--- a/vllm_ascend/attention/attention_v1.py
+++ b/vllm_ascend/attention/attention_v1.py
@@ -443,7 +443,8 @@ class AscendAttentionBackendImpl(AttentionImpl):
                            block_table=attn_metadata.block_tables,
                            context_lens=attn_metadata.seq_lens,
                            out=output)
-                        update_graph_params_workspaces(num_tokens, workspace)
+                        update_graph_params_workspaces(
+                            num_tokens, weak_ref_tensors(workspace))

                # Handle graph capturing mode
                stream = torch_npu.npu.current_stream()
@@ -459,7 +460,7 @@ class AscendAttentionBackendImpl(AttentionImpl):
                    self.num_kv_heads,
                    self.num_heads,
                    self.scale,
-                    weak_ref_tensors(attn_metadata.block_tables),
+                    attn_metadata.block_tables,
                    attn_metadata.seq_lens,
                    weak_ref_tensors(output),
                ))