fix(critical): match Sub168 config exactly + disable risky numerical patch
1. computility-run.yaml: restore Sub168's proven params: - max-model-len=256000 (not 100000) - gpu-memory-utilization=0.95 (not 0.90) - max-num-seqs=2 (not 1) - max-num-batched-tokens=4096 (restored) - enable-chunked-prefill (restored) These params worked for Sub168. Now that pip install is removed, they should work for us too. 2. patch_ops.sh: disable patch_numerical_stability.py If corex_gdn loads (which it should without pip install breaking deps), Python GatedDeltaNet fallback never runs, so numerical patches are unnecessary. Running regex replacements on qwen3_5.py risks breaking corex import conditions.
This commit is contained in:
@@ -8,14 +8,16 @@ command:
|
|||||||
- --served-model-name
|
- --served-model-name
|
||||||
- llm
|
- llm
|
||||||
- --max-model-len
|
- --max-model-len
|
||||||
- '100000'
|
- '256000'
|
||||||
- --gpu-memory-utilization
|
- --gpu-memory-utilization
|
||||||
- '0.90'
|
- '0.95'
|
||||||
- --trust-remote-code
|
- --trust-remote-code
|
||||||
- -tp
|
- -tp
|
||||||
- '4'
|
- '4'
|
||||||
- --max-num-seqs
|
- --max-num-seqs
|
||||||
- '1'
|
- '2'
|
||||||
|
- --max-num-batched-tokens
|
||||||
|
- '4096'
|
||||||
- --disable-log-requests
|
- --disable-log-requests
|
||||||
- --disable-frontend-multiprocessing
|
- --disable-frontend-multiprocessing
|
||||||
- --enforce-eager
|
- --enforce-eager
|
||||||
@@ -25,6 +27,7 @@ command:
|
|||||||
- --reasoning-parser
|
- --reasoning-parser
|
||||||
- qwen3
|
- qwen3
|
||||||
- --enable-prefix-caching
|
- --enable-prefix-caching
|
||||||
|
- --enable-chunked-prefill
|
||||||
- --dtype
|
- --dtype
|
||||||
- half
|
- half
|
||||||
env:
|
env:
|
||||||
|
|||||||
@@ -107,19 +107,15 @@ done
|
|||||||
echo "[patch_ops] reasoning parser + serving files installed"
|
echo "[patch_ops] reasoning parser + serving files installed"
|
||||||
|
|
||||||
# ============================================================
|
# ============================================================
|
||||||
# 4. CCCL Agent-pattern: numerical stability patch for qwen3_5.py
|
# 4. Numerical stability patch — DISABLED
|
||||||
# Sub509 docker logs: 99.98% NaN in every GatedDeltaNet layer.
|
# If corex_gdn loads (which it should without pip install),
|
||||||
# Base image has NaN detection + nan_to_num(nan=0.0), but that
|
# the Python _torch_chunk_gated_delta_rule is NEVER called.
|
||||||
# means DeltaNet layers output all-zeros → model "brain dead"
|
# Patching qwen3_5.py risks breaking corex import conditions.
|
||||||
# → can't produce <tool_call> XML → d03 FAIL.
|
# Only enable this if docker logs still show NaN after corex fix.
|
||||||
#
|
|
||||||
# Strategy (CCCL optionally_static): detect what guards exist,
|
|
||||||
# inject ONLY what's missing. Preserve corex kernel paths.
|
|
||||||
# Agent flow: Init → Detect → Patch → Verify.
|
|
||||||
# ============================================================
|
# ============================================================
|
||||||
python3 ./patch_numerical_stability.py 2>&1 || \
|
# python3 ./patch_numerical_stability.py 2>&1 || \
|
||||||
echo "[patch_ops] WARNING: numerical stability patch failed (non-fatal)"
|
# echo "[patch_ops] WARNING: numerical stability patch failed (non-fatal)"
|
||||||
echo "[patch_ops] numerical stability patch complete"
|
echo "[patch_ops] numerical stability patch SKIPPED (corex_gdn handles DeltaNet)"
|
||||||
|
|
||||||
# ============================================================
|
# ============================================================
|
||||||
# 5. DO NOT full-replace these files — base image has optimized versions.
|
# 5. DO NOT full-replace these files — base image has optimized versions.
|
||||||
|
|||||||
Reference in New Issue
Block a user