fix(CRITICAL): CCCL overflow guard — clamp before cumsum + max-num-seqs=2
Three fixes derived from CCCL source code patterns: 1. CCCL accumulator_t pattern (dispatch_segmented_scan.cuh): - Clamp g to [-5, 2] BEFORE cumsum (was: no pre-clamp, post-clamp ±80) - Tighten post-cumsum clamp to ±20 (was ±80) - Clamp A_log to [-8, 4] before exp() (was: unclamped) - Clamp softplus output to max=10 (was: unclamped) - Clamp g before exp_() in decode path (was: NO clamp at all) 2. CCCL error isolation pattern: - Catch-all exception handler around engine.generate() - max-num-seqs 1→2 to prevent t2_n_2 crash cascade 3. Reduce _DNN_CHUNK 4096→2048 (fewer cumsum steps = less overflow) Root cause: Sub508/509 scored 0 because t2_n_2 killed engine process. NaN (99.98-100% per GatedDeltaNet layer) from unclamped cumsum→exp overflow.
This commit is contained in:
@@ -364,6 +364,12 @@ class OpenAIServingChat(OpenAIServing):
|
||||
except ValueError as e:
|
||||
# TODO: Use a vllm-specific Validation Error
|
||||
return self.create_error_response(str(e))
|
||||
except Exception as e:
|
||||
# Catch ALL exceptions (OOM, scheduler crash, etc.) to prevent
|
||||
# a single request from killing the entire engine process.
|
||||
logger.exception("Engine error (non-fatal, returning 500): %s", e)
|
||||
return self.create_error_response(
|
||||
f"Internal engine error: {type(e).__name__}: {e}")
|
||||
|
||||
if raw_request:
|
||||
result_generator = iterate_with_cancellation(
|
||||
|
||||
Reference in New Issue
Block a user