# GPT-OSS 20B GGUF with MoE CPU Offloading Support This is a specialized conversion of OpenAI's GPT-OSS 20B model that has been tested and verified to work with MoE (Mixture of Experts) CPU offloading capabilities. ## Key Features - **Full F16 precision** (13.8GB file size) - **MoE Architecture**: 32 experts, 4 active per token - **CPU Offloading Support**: Expert tensors can be offloaded to CPU for massive VRAM savings - **Verified Working**: Tested with shimmy feat/moe-cpu-offload branch - **Memory Efficiency**: Reduces GPU memory usage from ~15GB to ~2MB when using CPU offloading ## Model Specifications - **Architecture**: gpt-oss - **Parameters**: 20.91B total, 3.6B active - **Context Length**: 131,072 tokens - **Quantization**: F16 base + MXFP4 expert weights - **Expert Count**: 32 experts, 4 used per token - **Sliding Window**: 128 tokens ## Usage with MoE CPU Offloading This model requires a compatible inference engine that supports MoE CPU offloading. Tested with: - shimmy inference server (feat/moe-cpu-offload branch) - llama-cpp-rs with MoE CPU offloading support Example usage: ```bash # Enable CPU offloading for all expert tensors ./shimmy serve --cpu-moe --model gpt-oss-20b-f16.gguf ``` ## Performance Characteristics - **VRAM Usage**: ~2 MiB (with CPU offloading) vs ~15GB (standard) - **Generation Speed**: Maintained quality with reasonable speed - **Memory Savings**: 99.9% VRAM reduction for large MoE models ## Original Model Based on OpenAI's official GPT-OSS 20B model: - Source: https://huggingface.co/openai/gpt-oss-20b - License: Apache 2.0 - Paper: https://arxiv.org/abs/2508.10925 ## Conversion Details Converted from the original HuggingFace format specifically to support MoE CPU offloading features in GGUF format while preserving the full model precision and MoE architecture. **Note**: This is an experimental build demonstrating MoE CPU offloading capabilities. Standard GGUF versions may not support this feature.