840fe923cceefd15527bd24bebba4a0a579aad51
Docker log reveals: 'NaN in prefill GatedDeltaNet layer 0 (frac=0.9998)' Every DeltaNet (linear attention) layer produces 99.98% NaN values. nan_to_num replaces them with zeros, destroying model output quality. This is the root cause of d10_thinking_disable_ctk gibberish output. Root cause: g.cumsum(dim=-1) accumulates unbounded gate logits. When fed to exp(), large values overflow to Inf, which propagates as NaN through subsequent matmul and forward_sub operations. Fix: Clamp cumulative gate logits to [-20, 20] before any exp(). Range keeps exp in [~2e-9, ~5e8] — safe for float32 accumulation. Inspired by CCCL dispatch_reduce_deterministic.cuh: numerical stability requires bounded intermediate values (RFA pattern). Also in this log: - FusedMoE: 'vllm_moe_topk_softmax' not in ixformer → PyTorch fallback (expected, cannot fix without BI-V100 kernel rebuild) - OOM at end of sub168: 31.72 GiB GPU with 30.86 GiB allocated CCCL input: dispatch_reduce_deterministic.cuh RFA pattern, tuning_batch_memcpy.cuh (small=128t×4buf, large=256t×32B)
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%