1.9 KiB
1.9 KiB
GPT-OSS 20B GGUF with MoE CPU Offloading Support
This is a specialized conversion of OpenAI's GPT-OSS 20B model that has been tested and verified to work with MoE (Mixture of Experts) CPU offloading capabilities.
Key Features
- Full F16 precision (13.8GB file size)
- MoE Architecture: 32 experts, 4 active per token
- CPU Offloading Support: Expert tensors can be offloaded to CPU for massive VRAM savings
- Verified Working: Tested with shimmy feat/moe-cpu-offload branch
- Memory Efficiency: Reduces GPU memory usage from ~15GB to ~2MB when using CPU offloading
Model Specifications
- Architecture: gpt-oss
- Parameters: 20.91B total, 3.6B active
- Context Length: 131,072 tokens
- Quantization: F16 base + MXFP4 expert weights
- Expert Count: 32 experts, 4 used per token
- Sliding Window: 128 tokens
Usage with MoE CPU Offloading
This model requires a compatible inference engine that supports MoE CPU offloading. Tested with:
- shimmy inference server (feat/moe-cpu-offload branch)
- llama-cpp-rs with MoE CPU offloading support
Example usage:
# Enable CPU offloading for all expert tensors
./shimmy serve --cpu-moe --model gpt-oss-20b-f16.gguf
Performance Characteristics
- VRAM Usage: ~2 MiB (with CPU offloading) vs ~15GB (standard)
- Generation Speed: Maintained quality with reasonable speed
- Memory Savings: 99.9% VRAM reduction for large MoE models
Original Model
Based on OpenAI's official GPT-OSS 20B model:
- Source: https://huggingface.co/openai/gpt-oss-20b
- License: Apache 2.0
- Paper: https://arxiv.org/abs/2508.10925
Conversion Details
Converted from the original HuggingFace format specifically to support MoE CPU offloading features in GGUF format while preserving the full model precision and MoE architecture.
Note: This is an experimental build demonstrating MoE CPU offloading capabilities. Standard GGUF versions may not support this feature.