fix: max_num_seqs=2 for n=2 support + remove protocol n clamp
Sub168 (competitor) passes t2_n_2 with n=2 at 1.50s even with max_num_seqs likely >1. Our max_num_seqs=1 made n=2 crash. Changes: - computility-run.yaml: max-num-seqs 1→2 (200GB total VRAM sufficient) - protocol.py: remove n>1 clamp, let serving_chat scheduler guard handle it - serving_chat.py retains try/except guard for get_scheduler_config Risk: if 2 concurrent seqs OOM, service crashes. But concurrency=1 means only 1 request at a time, so n=2 just generates 2 answers sequentially. CCCL input: tuning_topk.cuh (bits_per_pass=11 for float32, threads=512), tuning_transform.cuh (cc_to_min_bytes_in_flight: B200=64KB, A100=16KB, BI-V100 should use 48-64KB based on per-SM BW=56GB/s)
This commit is contained in:
@@ -15,7 +15,7 @@ command:
|
||||
- -tp
|
||||
- '4'
|
||||
- --max-num-seqs
|
||||
- '1'
|
||||
- '2'
|
||||
- --disable-log-requests
|
||||
- --disable-frontend-multiprocessing
|
||||
- --enforce-eager
|
||||
|
||||
Reference in New Issue
Block a user