[Refactor] Fix AttentionMaskBuilder singleton and remove redundant pcp_prefill_mask (#4870)
## What this PR does / why we need it? This PR fixes the `AttentionMaskBuilder` singleton initialization issue introduced in PR #4779 and removes the unused `pcp_prefill_mask` field. ### Background After PR #4779 made `AttentionMaskBuilder` a singleton with `@singleton` decorator, the class constructor now requires a `device` parameter. However, two initialization sites were still using the old parameterless constructor, causing failures. ### Changes 1. **Fix singleton initialization** - Fixed `AttentionMaskBuilder()` → `AttentionMaskBuilder(self.device)` in `AscendMLAMetadataBuilder.__init__()` - Fixed `AttentionMaskBuilder()` → `AttentionMaskBuilder(self.device)` in `AscendAttentionMetadataBuilder.__init__()` 2. **Remove unused field** - Removed `pcp_prefill_mask` field from `AscendPrefillContextParallelMetadata` (never used in codebase) - Updated related test assertions ### Related - Issue #5463 - PR #4779 (Unify all mask generation methods) - PR #5389 (Make AttentionMaskBuilder singleton) ## Does this PR introduce _any_ user-facing change? No. This is an internal refactoring. ## How was this patch tested? - ✅ Local testing: No linter errors - ✅ Unit tests for attention modules verified - ⏳ CI pipeline Signed-off-by: lico67373 <918688502@qq.com> Co-authored-by: weijinqian0 <1184188277@qq.com>
This commit is contained in:
@@ -291,8 +291,6 @@ class TestMtpProposer:
|
||||
|
||||
mock_runner = MagicMock()
|
||||
mock_runner.actual_seq_lengths_q = MagicMock()
|
||||
mock_runner.attn_mask = MagicMock()
|
||||
mock_runner.spec_attn_mask = MagicMock()
|
||||
mock_runner.attn_state = MagicMock()
|
||||
mock_runner.graph_pad_size = 0
|
||||
mock_runner.decode_token_per_req = MagicMock()
|
||||
@@ -334,5 +332,3 @@ class TestMtpProposer:
|
||||
assert spec_common_attn_metadata.num_actual_tokens == total_num_tokens
|
||||
assert spec_common_attn_metadata.max_query_len == 8
|
||||
assert spec_common_attn_metadata.actual_seq_lengths_q == proposer.runner.actual_seq_lengths_q
|
||||
assert spec_common_attn_metadata.attn_mask == proposer.runner.attn_mask
|
||||
assert spec_common_attn_metadata.spec_attn_mask == proposer.runner.spec_attn_mask
|
||||
|
||||
Reference in New Issue
Block a user