f1a408c75f0ca7b21c53e631367351e20a1cdc86
t3_max_tokens_none test sends request without max_tokens. With max_model_len=262144, default_max_tokens was ~260K tokens. Model generates indefinitely at 12 TPS, container gets OOM killed after ~6min. Cap at 8192 prevents memory exhaustion while still allowing long outputs.
Description
No description provided
Languages
Python
88.5%
Cuda
9.2%
Shell
2.3%