sglang/sgl-kernel/README.md

# SGL Kernel

[Kernel Library](https://github.com/sgl-project/sglang/tree/main/sgl-kernel) for SGLang

[![PyPI](https://img.shields.io/pypi/v/sgl-kernel)](https://pypi.org/project/sgl-kernel)

## Installation

For CUDA 11.8:

```bash
pip3 install sgl-kernel -i https://docs.sglang.ai/whl/cu118
```

For CUDA 12.1 or CUDA 12.4:

```bash
pip3 install sgl-kernel
```

# Developer Guide

## Development Environment Setup

Use Docker to set up the development environment. See [Docker setup guide](https://github.com/sgl-project/sglang/blob/main/docs/developer/development_guide_using_docker.md#setup-docker-container).

Create and enter development container:
```bash
docker run -itd --shm-size 32g --gpus all -v $HOME/.cache:/root/.cache --ipc=host --name sglang_zhyncs lmsysorg/sglang:dev /bin/zsh
docker exec -it sglang_zhyncs /bin/zsh
```

## Project Structure

### Dependencies

Third-party libraries:

- [CUTLASS](https://github.com/NVIDIA/cutlass)
- [FlashInfer](https://github.com/flashinfer-ai/flashinfer)
- [DeepGEMM](https://github.com/deepseek-ai/DeepGEMM)
- [FlashAttention](https://github.com/Dao-AILab/flash-attention)

### Kernel Development

Steps to add a new kernel:

1. Implement the kernel in [csrc](https://github.com/sgl-project/sglang/tree/main/sgl-kernel/csrc)
2. Expose the interface in [include/sgl_kernel_ops.h](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/include/sgl_kernel_ops.h)
3. Create torch extension in [csrc/torch_extension.cc](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/csrc/torch_extension.cc)
4. Update [CMakeLists.txt](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/CMakeLists.txt) to include new CUDA source
5. Expose Python interface in [python](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/python/sgl_kernel)

### Development Tips

1. When implementing kernels in [csrc](https://github.com/sgl-project/sglang/tree/main/sgl-kernel/csrc), only define pure CUDA files and C++ interfaces. If you need to use `Torch::tensor`, use `<torch/all.h>` instead of `<torch/extension.h>`. Using `<torch/extension.h>` will cause compilation errors when using SABI.

2. When creating torch extensions, add the function definition with `m.def`, and device binding with `m.impl`:
- Using torch.compile need `m.def` with schema, it helps auto capture the custom kernel. Reference: [How to add FakeTensor](https://docs.google.com/document/d/1_W62p8WJOQQUzPsJYa7s701JXt0qf2OfLub2sbkHOaU/edit?tab=t.0#heading=h.ptttacy8y1u9)

- How to write schema: [Schema reference](https://github.com/pytorch/pytorch/blob/main/aten/src/ATen/native/README.md#func)

   ```cpp
   // We need def with schema here for torch.compile
   m.def(
    "bmm_fp8(Tensor A, Tensor B, Tensor! D, Tensor A_scale, Tensor B_scale, Tensor workspace_buffer, int "
    "cublas_handle, int cuda_stream) -> ()");
   m.impl("bmm_fp8", torch::kCUDA, &bmm_fp8);
   ```

3. When exposing Python interfaces, avoid using kwargs in C++ interface kernels.

    **Avoid this:**

    ```cpp
    torch.ops.sgl_kernel.apply_rope_pos_ids_cos_sin_cache.default(
        q=query.view(query.shape[0], -1, head_size),
        k=key.view(key.shape[0], -1, head_size),
        q_rope=query.view(query.shape[0], -1, head_size),
        k_rope=key.view(key.shape[0], -1, head_size),
        cos_sin_cache=cos_sin_cache,
        pos_ids=positions.long(),
        interleave=(not is_neox),
        cuda_stream=get_cuda_stream(),
    )
    ```

    **Use this instead:**

    ```cpp
    torch.ops.sgl_kernel.apply_rope_pos_ids_cos_sin_cache.default(
        query.view(query.shape[0], -1, head_size),
        key.view(key.shape[0], -1, head_size),
        query.view(query.shape[0], -1, head_size),
        key.view(key.shape[0], -1, head_size),
        cos_sin_cache,
        positions.long(),
        (not is_neox),
        get_cuda_stream(),
    )
    ```

### Integrating Third-Party Libraries with Data Type Conversion

When integrating new third-party libraries like flash-attention, you may encounter data type compatibility issues between the C++ interface and PyTorch bindings. For example, the third-party code might use `float` or `int` types, while PyTorch requires `double` and `int64_t`.

> The reason we need `double` and `int64_t` in torch binding is that TORCH_LIBRARY handles the `Python-to-C++` conversion process. Python's `float` data type actually corresponds to `double` in C++, while Python's `int` corresponds to `int64_t` in C++.

To address this issue, we provide the `make_pytorch_shim` function in [sgl_kernel_torch_shim](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/include/sgl_kernel_torch_shim.h) that handles data type conversions automatically.

When you need to support new data type conversions, you can easily add conversion functions like this:

```cpp
// Map `int` -> `int64_t`
template <>
struct pytorch_library_compatible_type<int> {
  using type = int64_t;
  static int convert_from_type(int64_t arg) {
    TORCH_CHECK(arg <= std::numeric_limits<int>::max(), "int64_t value is too large to be converted  to int");
    TORCH_CHECK(arg >= std::numeric_limits<int>::min(), "int64_t value is too small to be converted to int");
    return arg;
  }
};
```

To use this with your library functions, simply wrap them with make_pytorch_shim:

```cpp
/*
 * From flash-attention
 */
 m.impl("fwd", torch::kCUDA, make_pytorch_shim(&mha_fwd));
```

### Build & Install

Development build:

```bash
make build
```

Note:

The `sgl-kernel` is rapidly evolving. If you experience a compilation failure, try using `make rebuild`.

#### Build with [ccache](https://github.com/ccache/ccache)
```bash
# or `yum install -y ccache`.
apt-get install -y ccache
# Building with ccache is enabled when ccache is installed and CCACHE_DIR is set.
export CCACHE_DIR=/path/to/your/ccache/dir
export CCACHE_BACKEND=""
export CCACHE_KEEP_LOCAL_STORAGE="TRUE"
unset CCACHE_READONLY
python -m uv build --wheel -Cbuild-dir=build --color=always .
```

##### Configuring CMake Build Options
Cmake options can be configuring by adding `-Ccmake.define.<option>=<value>` to the `uv build` flags.
For example, to enable building FP4 kernels, use:
```bash
python -m uv build --wheel -Cbuild-dir=build -Ccmake.define.SGL_KERNEL_ENABLE_FP4=1 --color=always .
```
See CMakeLists.txt for more options.

### Testing & Benchmarking

1. Add pytest tests in [tests/](https://github.com/sgl-project/sglang/tree/main/sgl-kernel/tests), if you need to skip some test, please use `@pytest.mark.skipif`

```python
@pytest.mark.skipif(
    skip_condition, reason="Nvfp4 Requires compute capability of 10 or above."
)
```

2. Add benchmarks using [triton benchmark](https://triton-lang.org/main/python-api/generated/triton.testing.Benchmark.html) in [benchmark/](https://github.com/sgl-project/sglang/tree/main/sgl-kernel/benchmark)
3. Run test suite


### Release new version

Update version in [pyproject.toml](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/pyproject.toml) and [version.py](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/python/sgl_kernel/version.py)
Add new contributors so they can trigger CI automatically (#2269) Co-authored-by: Qun Yang <qun.yang@intel.com> Co-authored-by: zhengy001 <zhengy.gator@gmail.com> Co-authored-by: HandH1998 <1335248067@qq.com> Co-authored-by: xiaobo <xiaob.chen@outlook.com> 2024-11-29 16:37:52 -08:00			`# SGL Kernel`
minor: add sgl-kernel dir (#2261) 2024-11-30 02:27:35 +08:00
update installation doc for sgl-kernel (#3129) 2025-01-26 00:00:13 +08:00			`[Kernel Library](https://github.com/sgl-project/sglang/tree/main/sgl-kernel) for SGLang`
feat: support sgl-kernel pypi (#2302) 2024-12-01 20:11:21 +08:00
			`[![PyPI](https://img.shields.io/pypi/v/sgl-kernel)](https://pypi.org/project/sgl-kernel)`
update installation doc for sgl-kernel (#3129) 2025-01-26 00:00:13 +08:00
			`## Installation`

			`For CUDA 11.8:`

			```bash
			`pip3 install sgl-kernel -i https://docs.sglang.ai/whl/cu118`
			```

			`For CUDA 12.1 or CUDA 12.4:`

			```bash
			`pip3 install sgl-kernel`
			```
Misc clean up; Remove the support of jump forward (#4032) 2025-03-03 07:02:14 -08:00
			`# Developer Guide`

			`## Development Environment Setup`

			`Use Docker to set up the development environment. See [Docker setup guide](https://github.com/sgl-project/sglang/blob/main/docs/developer/development_guide_using_docker.md#setup-docker-container).`

			`Create and enter development container:`
			```bash
			`docker run -itd --shm-size 32g --gpus all -v $HOME/.cache:/root/.cache --ipc=host --name sglang_zhyncs lmsysorg/sglang:dev /bin/zsh`
			`docker exec -it sglang_zhyncs /bin/zsh`
			```

			`## Project Structure`

			`### Dependencies`

			`Third-party libraries:`

			`- [CUTLASS](https://github.com/NVIDIA/cutlass)`
			`- [FlashInfer](https://github.com/flashinfer-ai/flashinfer)`
[Docs] Update DeepGEMM at README.md (#4886) 2025-03-30 00:53:39 +08:00			`- [DeepGEMM](https://github.com/deepseek-ai/DeepGEMM)`
cleanup sgl-kernel (#4933) 2025-03-30 14:12:30 -07:00			`- [FlashAttention](https://github.com/Dao-AILab/flash-attention)`
Misc clean up; Remove the support of jump forward (#4032) 2025-03-03 07:02:14 -08:00
			`### Kernel Development`

			`Steps to add a new kernel:`

Rename files in sgl kernel to avoid nested folder structure (#4213) Co-authored-by: zhyncs <me@zhyncs.com> 2025-03-08 22:54:51 -08:00			`1. Implement the kernel in [csrc](https://github.com/sgl-project/sglang/tree/main/sgl-kernel/csrc)`
			`2. Expose the interface in [include/sgl_kernel_ops.h](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/include/sgl_kernel_ops.h)`
			`3. Create torch extension in [csrc/torch_extension.cc](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/csrc/torch_extension.cc)`
remove setup for sgl-kernel (#4899) 2025-03-29 12:47:38 -07:00			`4. Update [CMakeLists.txt](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/CMakeLists.txt) to include new CUDA source`
Rename files in sgl kernel to avoid nested folder structure (#4213) Co-authored-by: zhyncs <me@zhyncs.com> 2025-03-08 22:54:51 -08:00			`5. Expose Python interface in [python](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/python/sgl_kernel)`
Misc clean up; Remove the support of jump forward (#4032) 2025-03-03 07:02:14 -08:00
[Misc] Clean m.def and add Development Tips (#4890) 2025-03-30 14:06:18 +08:00			`### Development Tips`

			1. When implementing kernels in [csrc](https://github.com/sgl-project/sglang/tree/main/sgl-kernel/csrc), only define pure CUDA files and C++ interfaces. If you need to use `Torch::tensor`, use `<torch/all.h>` instead of `<torch/extension.h>`. Using `<torch/extension.h>` will cause compilation errors when using SABI.

[Fix] revert clean m.def for cudagraph (#4944) 2025-03-31 17:08:55 +08:00			2. When creating torch extensions, add the function definition with `m.def`, and device binding with `m.impl`:
			- Using torch.compile need `m.def` with schema, it helps auto capture the custom kernel. Reference: [How to add FakeTensor](https://docs.google.com/document/d/1_W62p8WJOQQUzPsJYa7s701JXt0qf2OfLub2sbkHOaU/edit?tab=t.0#heading=h.ptttacy8y1u9)

			`- How to write schema: [Schema reference](https://github.com/pytorch/pytorch/blob/main/aten/src/ATen/native/README.md#func)`

[Misc] Clean m.def and add Development Tips (#4890) 2025-03-30 14:06:18 +08:00			```cpp
[Fix] revert clean m.def for cudagraph (#4944) 2025-03-31 17:08:55 +08:00			`// We need def with schema here for torch.compile`
			`m.def(`
			`"bmm_fp8(Tensor A, Tensor B, Tensor! D, Tensor A_scale, Tensor B_scale, Tensor workspace_buffer, int "`
			`"cublas_handle, int cuda_stream) -> ()");`
			`m.impl("bmm_fp8", torch::kCUDA, &bmm_fp8);`
[Misc] Clean m.def and add Development Tips (#4890) 2025-03-30 14:06:18 +08:00			```

			`3. When exposing Python interfaces, avoid using kwargs in C++ interface kernels.`

			`Avoid this:`

			```cpp
			`torch.ops.sgl_kernel.apply_rope_pos_ids_cos_sin_cache.default(`
			`q=query.view(query.shape[0], -1, head_size),`
			`k=key.view(key.shape[0], -1, head_size),`
			`q_rope=query.view(query.shape[0], -1, head_size),`
			`k_rope=key.view(key.shape[0], -1, head_size),`
			`cos_sin_cache=cos_sin_cache,`
			`pos_ids=positions.long(),`
			`interleave=(not is_neox),`
			`cuda_stream=get_cuda_stream(),`
			`)`
			```

			`Use this instead:`

			```cpp
			`torch.ops.sgl_kernel.apply_rope_pos_ids_cos_sin_cache.default(`
			`query.view(query.shape[0], -1, head_size),`
			`key.view(key.shape[0], -1, head_size),`
			`query.view(query.shape[0], -1, head_size),`
			`key.view(key.shape[0], -1, head_size),`
			`cos_sin_cache,`
			`positions.long(),`
			`(not is_neox),`
			`get_cuda_stream(),`
			`)`
			```

[feat] add fa3 in sgl-kernel (#4902) Co-authored-by: Sleepcoo <Sleepcoo@gmail.com> 2025-03-31 03:57:10 +08:00			`### Integrating Third-Party Libraries with Data Type Conversion`

			When integrating new third-party libraries like flash-attention, you may encounter data type compatibility issues between the C++ interface and PyTorch bindings. For example, the third-party code might use `float` or `int` types, while PyTorch requires `double` and `int64_t`.

[Fix] revert clean m.def for cudagraph (#4944) 2025-03-31 17:08:55 +08:00			> The reason we need `double` and `int64_t` in torch binding is that TORCH_LIBRARY handles the `Python-to-C++` conversion process. Python's `float` data type actually corresponds to `double` in C++, while Python's `int` corresponds to `int64_t` in C++.

[feat] add fa3 in sgl-kernel (#4902) Co-authored-by: Sleepcoo <Sleepcoo@gmail.com> 2025-03-31 03:57:10 +08:00			To address this issue, we provide the `make_pytorch_shim` function in [sgl_kernel_torch_shim](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/include/sgl_kernel_torch_shim.h) that handles data type conversions automatically.

			`When you need to support new data type conversions, you can easily add conversion functions like this:`

			```cpp
			// Map `int` -> `int64_t`
			`template <>`
			`struct pytorch_library_compatible_type<int> {`
			`using type = int64_t;`
			`static int convert_from_type(int64_t arg) {`
			`TORCH_CHECK(arg <= std::numeric_limits<int>::max(), "int64_t value is too large to be converted to int");`
			`TORCH_CHECK(arg >= std::numeric_limits<int>::min(), "int64_t value is too small to be converted to int");`
			`return arg;`
			`}`
			`};`
			```

			`To use this with your library functions, simply wrap them with make_pytorch_shim:`

			```cpp
			`/*`
			`* From flash-attention`
			`*/`
[Fix] revert clean m.def for cudagraph (#4944) 2025-03-31 17:08:55 +08:00			`m.impl("fwd", torch::kCUDA, make_pytorch_shim(&mha_fwd));`
[feat] add fa3 in sgl-kernel (#4902) Co-authored-by: Sleepcoo <Sleepcoo@gmail.com> 2025-03-31 03:57:10 +08:00			```

Misc clean up; Remove the support of jump forward (#4032) 2025-03-03 07:02:14 -08:00			`### Build & Install`

			`Development build:`

			```bash
			`make build`
			```

			`Note:`

			The `sgl-kernel` is rapidly evolving. If you experience a compilation failure, try using `make rebuild`.

[Build] Support build sgl-kernel with ccache (#5020) 2025-04-03 15:22:37 +08:00			`#### Build with [ccache](https://github.com/ccache/ccache)`
			```bash
			# or `yum install -y ccache`.
			`apt-get install -y ccache`
			`# Building with ccache is enabled when ccache is installed and CCACHE_DIR is set.`
			`export CCACHE_DIR=/path/to/your/ccache/dir`
			`export CCACHE_BACKEND=""`
			`export CCACHE_KEEP_LOCAL_STORAGE="TRUE"`
			`unset CCACHE_READONLY`
			`python -m uv build --wheel -Cbuild-dir=build --color=always .`
			```

FP4 weight loading and inference (2/2) (#3972) 2025-04-08 17:26:21 -07:00			`##### Configuring CMake Build Options`
			Cmake options can be configuring by adding `-Ccmake.define.<option>=<value>` to the `uv build` flags.
			`For example, to enable building FP4 kernels, use:`
			```bash
			`python -m uv build --wheel -Cbuild-dir=build -Ccmake.define.SGL_KERNEL_ENABLE_FP4=1 --color=always .`
			```
			`See CMakeLists.txt for more options.`

Misc clean up; Remove the support of jump forward (#4032) 2025-03-03 07:02:14 -08:00			`### Testing & Benchmarking`

[Misc] Use pytest.mark.skipif in sgl-kernel test (#5137) 2025-04-08 12:35:14 +08:00			1. Add pytest tests in [tests/](https://github.com/sgl-project/sglang/tree/main/sgl-kernel/tests), if you need to skip some test, please use `@pytest.mark.skipif`

			```python
			`@pytest.mark.skipif(`
			`skip_condition, reason="Nvfp4 Requires compute capability of 10 or above."`
			`)`
			```

Misc clean up; Remove the support of jump forward (#4032) 2025-03-03 07:02:14 -08:00			`2. Add benchmarks using [triton benchmark](https://triton-lang.org/main/python-api/generated/triton.testing.Benchmark.html) in [benchmark/](https://github.com/sgl-project/sglang/tree/main/sgl-kernel/benchmark)`
			`3. Run test suite`

[Misc] Use pytest.mark.skipif in sgl-kernel test (#5137) 2025-04-08 12:35:14 +08:00

Misc clean up; Remove the support of jump forward (#4032) 2025-03-03 07:02:14 -08:00			`### Release new version`

Rename files in sgl kernel to avoid nested folder structure (#4213) Co-authored-by: zhyncs <me@zhyncs.com> 2025-03-08 22:54:51 -08:00			`Update version in [pyproject.toml](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/pyproject.toml) and [version.py](https://github.com/sgl-project/sglang/blob/main/sgl-kernel/python/sgl_kernel/version.py)`