EngineX-Hygon/sglang

Fork 0

Files

History

Yi Zhang 5ec5eaf760 fix allreduce test (#4909 )

2025-03-29 23:16:53 -07:00

3rdparty

support cmake for sgl-kernel (#4706 )

2025-03-27 01:42:28 -07:00

benchmark

Add deepseek style fused moe group gate selection kernel (#4530 )

2025-03-29 11:51:45 -07:00

csrc

[Misc] Clean m.def and add Development Tips (#4890 )

2025-03-29 23:06:18 -07:00

include

Add deepseek style fused moe group gate selection kernel (#4530 )

2025-03-29 11:51:45 -07:00

python/sgl_kernel

[Misc] Clean m.def and add Development Tips (#4890 )

2025-03-29 23:06:18 -07:00

tests

fix allreduce test (#4909 )

2025-03-29 23:16:53 -07:00

.clang-format

New clang format for sgl kernel (#4194 )

2025-03-07 20:21:08 -08:00

build.sh

fix sgl-kernel cu118 build (#4872 )

2025-03-28 17:23:51 -07:00

CMakeLists.txt

Add deepseek style fused moe group gate selection kernel (#4530 )

2025-03-29 11:51:45 -07:00

LICENSE

feat: support sgl-kernel pypi (#2302 )

2024-12-01 20:11:21 +08:00

Makefile

bump sgl-kernel 0.0.5.post4 (#4768 )

2025-03-28 14:40:53 -07:00

pyproject_rocm.toml

bump sgl-kernel 0.0.5.post4 (#4768 )

2025-03-28 14:40:53 -07:00

pyproject.toml

bump sgl-kernel 0.0.5.post4 (#4768 )

2025-03-28 14:40:53 -07:00

README.md

[Misc] Clean m.def and add Development Tips (#4890 )

2025-03-29 23:06:18 -07:00

rename_wheels.sh

support cmake for sgl-kernel (#4706 )

2025-03-27 01:42:28 -07:00

setup_rocm.py

support cmake for sgl-kernel (#4706 )

2025-03-27 01:42:28 -07:00

THIRDPARTYNOTICES.txt

add THIRDPARTYNOTICES for DeepGEMM (#4272 )

2025-03-10 11:10:57 -07:00

README.md

SGL Kernel

Kernel Library for SGLang

Installation

For CUDA 11.8:

pip3 install sgl-kernel -i https://docs.sglang.ai/whl/cu118

For CUDA 12.1 or CUDA 12.4:

pip3 install sgl-kernel

Developer Guide

Development Environment Setup

Use Docker to set up the development environment. See Docker setup guide.

Create and enter development container:

docker run -itd --shm-size 32g --gpus all -v $HOME/.cache:/root/.cache --ipc=host --name sglang_zhyncs lmsysorg/sglang:dev /bin/zsh
docker exec -it sglang_zhyncs /bin/zsh

Project Structure

Dependencies

Third-party libraries:

Kernel Development

Steps to add a new kernel:

Implement the kernel in csrc
Expose the interface in include/sgl_kernel_ops.h
Create torch extension in csrc/torch_extension.cc
Update CMakeLists.txt to include new CUDA source
Expose Python interface in python

Development Tips

When implementing kernels in csrc, only define pure CUDA files and C++ interfaces. If you need to use Torch::tensor, use <torch/all.h> instead of <torch/extension.h>. Using <torch/extension.h> will cause compilation errors when using SABI.
When creating torch extensions, simply add the function definition with m.def:
```
m.def("register_graph_buffers", register_graph_buffers);
```

When exposing Python interfaces, avoid using kwargs in C++ interface kernels.

Avoid this:

torch.ops.sgl_kernel.apply_rope_pos_ids_cos_sin_cache.default(
    q=query.view(query.shape[0], -1, head_size),
    k=key.view(key.shape[0], -1, head_size),
    q_rope=query.view(query.shape[0], -1, head_size),
    k_rope=key.view(key.shape[0], -1, head_size),
    cos_sin_cache=cos_sin_cache,
    pos_ids=positions.long(),
    interleave=(not is_neox),
    cuda_stream=get_cuda_stream(),
)

Use this instead:

torch.ops.sgl_kernel.apply_rope_pos_ids_cos_sin_cache.default(
    query.view(query.shape[0], -1, head_size),
    key.view(key.shape[0], -1, head_size),
    query.view(query.shape[0], -1, head_size),
    key.view(key.shape[0], -1, head_size),
    cos_sin_cache,
    positions.long(),
    (not is_neox),
    get_cuda_stream(),
)

Build & Install

Development build:

make build

Note:

The sgl-kernel is rapidly evolving. If you experience a compilation failure, try using make rebuild.

Testing & Benchmarking

Add pytest tests in tests/
Add benchmarks using triton benchmark in benchmark/
Run test suite

Release new version

Update version in pyproject.toml and version.py