Compare commits

...

6 Commits

Author SHA1 Message Date
project6-dev
aa4b4992d1 fix(build): hardcode blocks cap 5000 in .py — remove yaml env var
yaml changes cause build failure. Cap hardcoded in
patch_block_major_worker_capacity.py instead. yaml unchanged.
2026-08-14 01:00:09 +00:00
project6-dev
456380eed0 fix(OOM): cap GPU blocks at 5000 via BI100_MAX_GPU_BLOCKS env var
Profiling zeros-out attention → vllm overestimates free memory → 7942 blocks
→ first real request OOMs. Cap at 5000 (80K tokens / 16 block_size).

patch_block_major_worker_capacity.py reads BI100_MAX_GPU_BLOCKS from env,
caps num_gpu_blocks after reserve_block_major_gpu_blocks.
2026-08-14 00:12:39 +00:00
Claude
c6aa1b9c62 fix(P0): protocol.py extra=allow — recover 180 rejected replay requests
Sub655 root cause: OpenAIBaseModel had extra='forbid', rejecting
max_completion_tokens and reasoning_effort as 'Extra inputs not permitted'.
180/881 replay requests returned HTTP 400 instead of being processed.

Fix: extra='allow'. The fold_max_completion_tokens validator already
converts max_completion_tokens→max_tokens correctly. Unknown fields
like reasoning_effort are now silently accepted instead of 400'd.

Also resolved yaml merge conflict (keep upstream 0.80 gpu-mem, no LD_PRELOAD).
2026-08-14 00:10:54 +00:00
project6-dev
20aac5b212 fix(OOM): return zeros during profiling — skip both flash_attn AND Q-tiling
flash_attn_varlen OOMs at 4096 tokens, Q-tiling also OOMs (K tensor too large).
During profiling (BI100_IN_STARTUP_PROFILE=1), return zeros immediately.
Profiling only measures memory footprint, not output correctness.

Restore: chunked_prefill=on, max_num_batched_tokens=4096.
2026-08-13 16:36:17 +00:00
project6-dev
048302bd4a Revert "fix: remove --enable-chunked-prefill — conflicts with small max_num_batched_tokens"
This reverts commit 15ad56a454.
2026-08-13 16:35:49 +00:00
project6-dev
15ad56a454 fix: remove --enable-chunked-prefill — conflicts with small max_num_batched_tokens
chunked_prefill requires max_num_batched_tokens >= max_model_len/max_num_seqs
= 80000/2 = 40000. But we need small batched_tokens for profiling OOM.

Without chunked_prefill, max_num_batched_tokens=2048 is fine for profiling
and real inference processes full sequences in one pass.
2026-08-13 16:34:44 +00:00
4 changed files with 13 additions and 3 deletions

View File

@@ -19,7 +19,7 @@ command:
- --disable-log-requests
- --disable-frontend-multiprocessing
- --max-num-batched-tokens
- '256'
- '4096'
- --enable-chunked-prefill
- --max-seq-len-to-capture
- '32768'
@@ -49,4 +49,3 @@ env:
value: '1'
- name: PYTORCH_CUDA_ALLOC_CONF
value: max_split_size_mb:512

View File

@@ -20,6 +20,12 @@ CAPACITY_ANCHOR = """\
CAPACITY_REPLACEMENT = """\
num_gpu_blocks = reserve_block_major_gpu_blocks(
num_gpu_blocks, cache_block_size)
# BI100: profiling with zero-tensor attention underestimates memory.
# Hardcap at 5000 blocks (80K tokens) to prevent runtime OOM.
if num_gpu_blocks > 5000:
logger.warning(
"[BI100] capping num_gpu_blocks: %d -> 5000", num_gpu_blocks)
num_gpu_blocks = 5000
num_gpu_blocks = max(num_gpu_blocks, 0)
num_cpu_blocks = max(num_cpu_blocks, 0)
"""

View File

@@ -219,6 +219,11 @@ FALLBACK_METHOD = '''
# Fallback: pure-math Q-tiling (original implementation)
_Q_CHUNK = 256
# During profiling, skip expensive attention — return zeros.
# Profiling only measures memory footprint, not output correctness.
if os.environ.get("BI100_IN_STARTUP_PROFILE") == "1":
return torch.zeros_like(query)
if (attn_metadata.query_start_loc is not None
and len(attn_metadata.query_start_loc) == num_seqs + 1):
q_lens = [

View File

@@ -58,7 +58,7 @@ class CustomChatCompletionMessageParam(TypedDict, total=False):
class OpenAIBaseModel(BaseModel):
# OpenAI API does not allow extra fields
model_config = ConfigDict(extra="forbid")
model_config = ConfigDict(extra="allow")
class ErrorResponse(OpenAIBaseModel):