Support penalty in overlap mode; return logprob with chunked prefill; improve benchmark scripts (#3988)
Co-authored-by: SangBin Cho <rkooo567@gmail.com> Co-authored-by: dhou-xai <dhou@x.ai> Co-authored-by: Hanming Lu <hanming_lu@berkeley.edu>
This commit is contained in:
@@ -96,7 +96,6 @@ Please consult the documentation below to learn more about the parameters you ma
|
||||
* `schedule_policy`: The scheduling policy to control the processing order of waiting prefill requests in a single engine.
|
||||
* `schedule_conservativeness`: Can be used to decrease/increase the conservativeness of the server when taking new requests. Highly conservative behavior leads to starvation, but low conservativeness leads to slowed-down performance.
|
||||
* `cpu_offload_gb`: Reserve this amount of RAM in GB for offloading of model parameters to the CPU.
|
||||
* `prefill_only_one_req`: When this flag is turned on, the engine prefills only one request at a time.
|
||||
|
||||
## Other runtime options
|
||||
|
||||
|
||||
Reference in New Issue
Block a user