env(yaml): CCCL buddy_allocator pattern — PYTORCH_CUDA_ALLOC_CONF + OMP_NUM_THREADS
CCCL buddy_allocator.cu teaches: control memory block fragmentation at the allocator level. Sub168 OOM trace shows 'max_split_size_mb' suggestion. Adding PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512 prevents PyTorch memory fragmentation that caused Sub168's final OOM. OMP_NUM_THREADS=1 matches Sub168 docker log: 'Reducing Torch parallelism from 64 threads to 1' CCCL device_reduce policy_selector pattern: hardware-adaptive params through environment, not code changes. computility-run.yaml env vars are the serving-safe equivalent of CCCL policy_selector.
This commit is contained in:
@@ -51,3 +51,7 @@ env:
|
||||
value: /tmp/vllm-request-metrics.jsonl
|
||||
- name: VLLM_CACHE_BLOCK_SIZE
|
||||
value: '16'
|
||||
- name: PYTORCH_CUDA_ALLOC_CONF
|
||||
value: max_split_size_mb:512
|
||||
- name: OMP_NUM_THREADS
|
||||
value: '1'
|
||||
|
||||
Reference in New Issue
Block a user