Files
gpt-oss-20b-moe-cpu-offload…/MOE-GGUF-README.md
ModelHub XC b81494004c 初始化项目,由ModelHub XC社区提供模型
Model: MikeKuykendall/gpt-oss-20b-moe-cpu-offload-gguf
Source: Original Platform
2026-09-01 07:04:17 +08:00

1.9 KiB

GPT-OSS 20B GGUF with MoE CPU Offloading Support

This is a specialized conversion of OpenAI's GPT-OSS 20B model that has been tested and verified to work with MoE (Mixture of Experts) CPU offloading capabilities.

Key Features

  • Full F16 precision (13.8GB file size)
  • MoE Architecture: 32 experts, 4 active per token
  • CPU Offloading Support: Expert tensors can be offloaded to CPU for massive VRAM savings
  • Verified Working: Tested with shimmy feat/moe-cpu-offload branch
  • Memory Efficiency: Reduces GPU memory usage from ~15GB to ~2MB when using CPU offloading

Model Specifications

  • Architecture: gpt-oss
  • Parameters: 20.91B total, 3.6B active
  • Context Length: 131,072 tokens
  • Quantization: F16 base + MXFP4 expert weights
  • Expert Count: 32 experts, 4 used per token
  • Sliding Window: 128 tokens

Usage with MoE CPU Offloading

This model requires a compatible inference engine that supports MoE CPU offloading. Tested with:

  • shimmy inference server (feat/moe-cpu-offload branch)
  • llama-cpp-rs with MoE CPU offloading support

Example usage:

# Enable CPU offloading for all expert tensors
./shimmy serve --cpu-moe --model gpt-oss-20b-f16.gguf

Performance Characteristics

  • VRAM Usage: ~2 MiB (with CPU offloading) vs ~15GB (standard)
  • Generation Speed: Maintained quality with reasonable speed
  • Memory Savings: 99.9% VRAM reduction for large MoE models

Original Model

Based on OpenAI's official GPT-OSS 20B model:

Conversion Details

Converted from the original HuggingFace format specifically to support MoE CPU offloading features in GGUF format while preserving the full model precision and MoE architecture.

Note: This is an experimental build demonstrating MoE CPU offloading capabilities. Standard GGUF versions may not support this feature.