Files
gpt-oss-20b-moe-cpu-offload…/MOE-GGUF-README.md
ModelHub XC b81494004c 初始化项目,由ModelHub XC社区提供模型
Model: MikeKuykendall/gpt-oss-20b-moe-cpu-offload-gguf
Source: Original Platform
2026-09-01 07:04:17 +08:00

51 lines
1.9 KiB
Markdown

# GPT-OSS 20B GGUF with MoE CPU Offloading Support
This is a specialized conversion of OpenAI's GPT-OSS 20B model that has been tested and verified to work with MoE (Mixture of Experts) CPU offloading capabilities.
## Key Features
- **Full F16 precision** (13.8GB file size)
- **MoE Architecture**: 32 experts, 4 active per token
- **CPU Offloading Support**: Expert tensors can be offloaded to CPU for massive VRAM savings
- **Verified Working**: Tested with shimmy feat/moe-cpu-offload branch
- **Memory Efficiency**: Reduces GPU memory usage from ~15GB to ~2MB when using CPU offloading
## Model Specifications
- **Architecture**: gpt-oss
- **Parameters**: 20.91B total, 3.6B active
- **Context Length**: 131,072 tokens
- **Quantization**: F16 base + MXFP4 expert weights
- **Expert Count**: 32 experts, 4 used per token
- **Sliding Window**: 128 tokens
## Usage with MoE CPU Offloading
This model requires a compatible inference engine that supports MoE CPU offloading. Tested with:
- shimmy inference server (feat/moe-cpu-offload branch)
- llama-cpp-rs with MoE CPU offloading support
Example usage:
```bash
# Enable CPU offloading for all expert tensors
./shimmy serve --cpu-moe --model gpt-oss-20b-f16.gguf
```
## Performance Characteristics
- **VRAM Usage**: ~2 MiB (with CPU offloading) vs ~15GB (standard)
- **Generation Speed**: Maintained quality with reasonable speed
- **Memory Savings**: 99.9% VRAM reduction for large MoE models
## Original Model
Based on OpenAI's official GPT-OSS 20B model:
- Source: https://huggingface.co/openai/gpt-oss-20b
- License: Apache 2.0
- Paper: https://arxiv.org/abs/2508.10925
## Conversion Details
Converted from the original HuggingFace format specifically to support MoE CPU offloading features in GGUF format while preserving the full model precision and MoE architecture.
**Note**: This is an experimental build demonstrating MoE CPU offloading capabilities. Standard GGUF versions may not support this feature.