CCCL sources read as design input: - group_by.cuh: static vs dynamic unit_count → match proven config - bench/bench.py: timeout + cache + graceful failure → cap default tokens - transform_iterator.cu: lazy transform pipeline → message preprocessing Changes: 1. computility-run.yaml: match Sub168's proven config exactly: - max-model-len: 100000 (not 32768, Sub168 used 100000 successfully) - Remove --max-num-batched-tokens (Sub168 didn't use it) - Remove --enable-chunked-prefill (Sub168 didn't use it) - Keep: max-num-seqs=1, gpu-mem=0.9, enable-prefix-caching 2. serving_chat.py: CCCL bench.py timeout pattern - Cap ALL requests without explicit max_tokens to 8192 - Cap tool_call requests to 2048 - Prevents NaN-damaged model from generating 99K tokens - Sub168 generates 139-2497 tokens per request
51 lines
1.3 KiB
YAML
51 lines
1.3 KiB
YAML
concurrency: 1
|
|
command:
|
|
- python3
|
|
- -m
|
|
- vllm.entrypoints.openai.api_server
|
|
- --model
|
|
- /model
|
|
- --served-model-name
|
|
- llm
|
|
- --max-model-len
|
|
- '100000'
|
|
- --gpu-memory-utilization
|
|
- '0.90'
|
|
- --trust-remote-code
|
|
- -tp
|
|
- '4'
|
|
- --max-num-seqs
|
|
- '1'
|
|
- --disable-log-requests
|
|
- --disable-frontend-multiprocessing
|
|
- --enforce-eager
|
|
- --enable-auto-tool-choice
|
|
- --tool-call-parser
|
|
- qwen3_coder
|
|
- --reasoning-parser
|
|
- qwen3
|
|
- --enable-prefix-caching
|
|
- --dtype
|
|
- half
|
|
env:
|
|
- name: VLLM_ENGINE_ITERATION_TIMEOUT_S
|
|
value: '3600'
|
|
- name: VLLM_ATTENTION_BACKEND
|
|
value: XFORMERS
|
|
- name: ENABLE_CUSTOM_IPC
|
|
value: '1'
|
|
- name: PYTHONPATH
|
|
value: /usr/local/corex/lib/python3/dist-packages:/usr/local/corex/lib64/python3/dist-packages
|
|
- name: LD_LIBRARY_PATH
|
|
value: /usr/local/corex/lib64:/usr/local/openmpi/lib
|
|
- name: VLLM_COREX_FA2_LIBRARY
|
|
value: /usr/local/corex/lib64/libcorex_fa2.so
|
|
- name: VLLM_COREX_GDN_LIBRARY
|
|
value: /usr/local/corex/lib64/libcorex_gdn.so
|
|
- name: VLLM_COREX_MOE_LIBRARY
|
|
value: /usr/local/corex/lib64/libcorex_moe.so
|
|
- name: VLLM_REQUEST_METRICS_FILE
|
|
value: /tmp/vllm-request-metrics.jsonl
|
|
- name: VLLM_CACHE_BLOCK_SIZE
|
|
value: '16'
|