fix(CRITICAL): CCCL overflow guard — clamp before cumsum + max-num-seqs=2
Three fixes derived from CCCL source code patterns: 1. CCCL accumulator_t pattern (dispatch_segmented_scan.cuh): - Clamp g to [-5, 2] BEFORE cumsum (was: no pre-clamp, post-clamp ±80) - Tighten post-cumsum clamp to ±20 (was ±80) - Clamp A_log to [-8, 4] before exp() (was: unclamped) - Clamp softplus output to max=10 (was: unclamped) - Clamp g before exp_() in decode path (was: NO clamp at all) 2. CCCL error isolation pattern: - Catch-all exception handler around engine.generate() - max-num-seqs 1→2 to prevent t2_n_2 crash cascade 3. Reduce _DNN_CHUNK 4096→2048 (fewer cumsum steps = less overflow) Root cause: Sub508/509 scored 0 because t2_n_2 killed engine process. NaN (99.98-100% per GatedDeltaNet layer) from unclamped cumsum→exp overflow.
This commit is contained in:
@@ -175,6 +175,6 @@ if [ -n "$VLLM2" ]; then
|
||||
cp ./chat_utils.py "$VLLM2/entrypoints/chat_utils.py" 2>/dev/null || true
|
||||
fi
|
||||
|
||||
echo "[patch_ops] DONE — CoreX dispatch + serving layer + engine patches deployed"
|
||||
echo "[patch_ops] Deployed: qwen3_5.py(CoreX dispatch), paged_attn.py, mamba_cache.py, sequence.py, scheduler.py, xformers patches, tool/reasoning parsers, serving layer"
|
||||
echo "[patch_ops] NOT deployed (base image native): model_runner.py, _custom_ops.py, sampler.py, logits_processor.py, arg_utils.py"
|
||||
echo "[patch_ops] DONE — all patches deployed"
|
||||
echo "[patch_ops] Deployed: qwen3_5.py, paged_attn.py, mamba_cache.py, sequence.py, scheduler.py, xformers patches, tool/reasoning parsers, serving layer"
|
||||
echo "[patch_ops] NOT deployed (using base image native): model_runner.py, _custom_ops.py, sampler.py, logits_processor.py, arg_utils.py"
|
||||
|
||||
Reference in New Issue
Block a user