51 lines
1.9 KiB
Markdown
51 lines
1.9 KiB
Markdown
# GPT-OSS 20B GGUF with MoE CPU Offloading Support
|
|
|
|
This is a specialized conversion of OpenAI's GPT-OSS 20B model that has been tested and verified to work with MoE (Mixture of Experts) CPU offloading capabilities.
|
|
|
|
## Key Features
|
|
|
|
- **Full F16 precision** (13.8GB file size)
|
|
- **MoE Architecture**: 32 experts, 4 active per token
|
|
- **CPU Offloading Support**: Expert tensors can be offloaded to CPU for massive VRAM savings
|
|
- **Verified Working**: Tested with shimmy feat/moe-cpu-offload branch
|
|
- **Memory Efficiency**: Reduces GPU memory usage from ~15GB to ~2MB when using CPU offloading
|
|
|
|
## Model Specifications
|
|
|
|
- **Architecture**: gpt-oss
|
|
- **Parameters**: 20.91B total, 3.6B active
|
|
- **Context Length**: 131,072 tokens
|
|
- **Quantization**: F16 base + MXFP4 expert weights
|
|
- **Expert Count**: 32 experts, 4 used per token
|
|
- **Sliding Window**: 128 tokens
|
|
|
|
## Usage with MoE CPU Offloading
|
|
|
|
This model requires a compatible inference engine that supports MoE CPU offloading. Tested with:
|
|
- shimmy inference server (feat/moe-cpu-offload branch)
|
|
- llama-cpp-rs with MoE CPU offloading support
|
|
|
|
Example usage:
|
|
```bash
|
|
# Enable CPU offloading for all expert tensors
|
|
./shimmy serve --cpu-moe --model gpt-oss-20b-f16.gguf
|
|
```
|
|
|
|
## Performance Characteristics
|
|
|
|
- **VRAM Usage**: ~2 MiB (with CPU offloading) vs ~15GB (standard)
|
|
- **Generation Speed**: Maintained quality with reasonable speed
|
|
- **Memory Savings**: 99.9% VRAM reduction for large MoE models
|
|
|
|
## Original Model
|
|
|
|
Based on OpenAI's official GPT-OSS 20B model:
|
|
- Source: https://huggingface.co/openai/gpt-oss-20b
|
|
- License: Apache 2.0
|
|
- Paper: https://arxiv.org/abs/2508.10925
|
|
|
|
## Conversion Details
|
|
|
|
Converted from the original HuggingFace format specifically to support MoE CPU offloading features in GGUF format while preserving the full model precision and MoE architecture.
|
|
|
|
**Note**: This is an experimental build demonstrating MoE CPU offloading capabilities. Standard GGUF versions may not support this feature. |