Deprecate --disable-flashinfer and introduce --attention-backend (#1380)

2024-09-10 17:11:16 -07:00
parent 3a6e8b6d78
commit 46094e0c1b
13 changed files with 99 additions and 61 deletions
--- a/docs/en/install.md
+++ b/docs/en/install.md
@@ -92,5 +92,5 @@ sky status --endpoint 30000 sglang
 </details>

 ### Common Notes
- [FlashInfer](https://github.com/flashinfer-ai/flashinfer) is currently one of the dependencies that must be installed for SGLang. It only supports sm75 and above. If you encounter any FlashInfer-related issues on sm75+ devices (e.g., T4, A10, A100, L4, L40S, H100), consider using Triton's kernel by `--disable-flashinfer --disable-flashinfer-sampling` and raise an issue.
+- [FlashInfer](https://github.com/flashinfer-ai/flashinfer) is the default attention kernel backend. It only supports sm75 and above. If you encounter any FlashInfer-related issues on sm75+ devices (e.g., T4, A10, A100, L4, L40S, H100), please disable it by adding `--disable-flashinfer --disable-flashinfer-sampling` and open an issue on GitHub.
 - If you only need to use the OpenAI backend, you can avoid installing other dependencies by using `pip install "sglang[openai]"`.