初始化项目,由ModelHub XC社区提供模型
Model: MikeKuykendall/gpt-oss-20b-moe-cpu-offload-gguf Source: Original Platform
This commit is contained in:
51
MOE-GGUF-README.md
Normal file
51
MOE-GGUF-README.md
Normal file
@@ -0,0 +1,51 @@
|
||||
# GPT-OSS 20B GGUF with MoE CPU Offloading Support
|
||||
|
||||
This is a specialized conversion of OpenAI's GPT-OSS 20B model that has been tested and verified to work with MoE (Mixture of Experts) CPU offloading capabilities.
|
||||
|
||||
## Key Features
|
||||
|
||||
- **Full F16 precision** (13.8GB file size)
|
||||
- **MoE Architecture**: 32 experts, 4 active per token
|
||||
- **CPU Offloading Support**: Expert tensors can be offloaded to CPU for massive VRAM savings
|
||||
- **Verified Working**: Tested with shimmy feat/moe-cpu-offload branch
|
||||
- **Memory Efficiency**: Reduces GPU memory usage from ~15GB to ~2MB when using CPU offloading
|
||||
|
||||
## Model Specifications
|
||||
|
||||
- **Architecture**: gpt-oss
|
||||
- **Parameters**: 20.91B total, 3.6B active
|
||||
- **Context Length**: 131,072 tokens
|
||||
- **Quantization**: F16 base + MXFP4 expert weights
|
||||
- **Expert Count**: 32 experts, 4 used per token
|
||||
- **Sliding Window**: 128 tokens
|
||||
|
||||
## Usage with MoE CPU Offloading
|
||||
|
||||
This model requires a compatible inference engine that supports MoE CPU offloading. Tested with:
|
||||
- shimmy inference server (feat/moe-cpu-offload branch)
|
||||
- llama-cpp-rs with MoE CPU offloading support
|
||||
|
||||
Example usage:
|
||||
```bash
|
||||
# Enable CPU offloading for all expert tensors
|
||||
./shimmy serve --cpu-moe --model gpt-oss-20b-f16.gguf
|
||||
```
|
||||
|
||||
## Performance Characteristics
|
||||
|
||||
- **VRAM Usage**: ~2 MiB (with CPU offloading) vs ~15GB (standard)
|
||||
- **Generation Speed**: Maintained quality with reasonable speed
|
||||
- **Memory Savings**: 99.9% VRAM reduction for large MoE models
|
||||
|
||||
## Original Model
|
||||
|
||||
Based on OpenAI's official GPT-OSS 20B model:
|
||||
- Source: https://huggingface.co/openai/gpt-oss-20b
|
||||
- License: Apache 2.0
|
||||
- Paper: https://arxiv.org/abs/2508.10925
|
||||
|
||||
## Conversion Details
|
||||
|
||||
Converted from the original HuggingFace format specifically to support MoE CPU offloading features in GGUF format while preserving the full model precision and MoE architecture.
|
||||
|
||||
**Note**: This is an experimental build demonstrating MoE CPU offloading capabilities. Standard GGUF versions may not support this feature.
|
||||
Reference in New Issue
Block a user