333 lines
17 KiB
Markdown
333 lines
17 KiB
Markdown
|
|
---
|
|||
|
|
license: apache-2.0
|
|||
|
|
datasets:
|
|||
|
|
- AmanPriyanshu/GPT-OSS-20B-MoE-expert-activations
|
|||
|
|
language:
|
|||
|
|
- en
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
tags:
|
|||
|
|
- mixture-of-experts
|
|||
|
|
- moe
|
|||
|
|
- expert-pruning
|
|||
|
|
- gpt-oss
|
|||
|
|
- openai
|
|||
|
|
- reasoning
|
|||
|
|
- safety
|
|||
|
|
- specialized
|
|||
|
|
- efficient
|
|||
|
|
- transformer
|
|||
|
|
- causal-lm
|
|||
|
|
- text-generation
|
|||
|
|
- pytorch
|
|||
|
|
- pruned-model
|
|||
|
|
- domain-specific
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Safety GPT-OSS Model (6 Experts)
|
|||
|
|
|
|||
|
|
**Project**: https://amanpriyanshu.github.io/GPT-OSS-MoE-ExpertFingerprinting/
|
|||
|
|
|
|||
|
|
<div align="center">
|
|||
|
|
|
|||
|
|
### 👥 Follow the Authors
|
|||
|
|
|
|||
|
|
**Aman Priyanshu**
|
|||
|
|
[](https://www.linkedin.com/in/aman-priyanshu/)
|
|||
|
|
[](https://x.com/AmanPriyanshu6)
|
|||
|
|
[](https://amanpriyanshu.github.io/)
|
|||
|
|
|
|||
|
|
**Supriti Vijay**
|
|||
|
|
[](https://www.linkedin.com/in/supriti-vijay/)
|
|||
|
|
[](https://x.com/SupritiVijay)
|
|||
|
|
[](https://supritivijay.github.io/)
|
|||
|
|
|
|||
|
|
</div>
|
|||
|
|
|
|||
|
|
## Introduction
|
|||
|
|
|
|||
|
|
This is a pruned variant of OpenAI's GPT-OSS-20B model, reduced to 6 experts per layer based on activation patterns from the [AmanPriyanshu/GPT-OSS-20B MoE Expert Activations dataset](https://huggingface.co/datasets/AmanPriyanshu/GPT-OSS-20B-MoE-expert-activations). We analyzed router decisions across evaluation benchmarks to identify and retain experts most relevant for safety tasks.
|
|||
|
|
|
|||
|
|
**⚠️ Experimental Model**: This is an experimental pruned model that may not work well - check the [examples below](#model-examples) to see if the outputs meet your needs before use.
|
|||
|
|
|
|||
|
|
This pruning approach reduces the model size while attempting to preserve performance on the target domain.
|
|||
|
|
|
|||
|
|
## Model Architecture & Statistics
|
|||
|
|
|
|||
|
|
| Metric | Value |
|
|||
|
|
|--------|-------|
|
|||
|
|
| **Base Model** | openai/gpt-oss-20b |
|
|||
|
|
| **Architecture** | Mixture-of-Experts Transformer |
|
|||
|
|
| **Total Parameters** | ~5.4B (pruned from 21B) |
|
|||
|
|
| **Original Experts per Layer** | 32 |
|
|||
|
|
| **Pruned Experts per Layer** | 6 |
|
|||
|
|
| **Layers** | 24 |
|
|||
|
|
| **Top-k Routing** | 4 |
|
|||
|
|
| **Context Length** | 128K tokens |
|
|||
|
|
| **Attention Heads** | 64 (Query), 8 (Key-Value) |
|
|||
|
|
| **Residual Dimension** | 2880 |
|
|||
|
|
| **Attention Pattern** | Alternating dense & sliding window (128 tokens) |
|
|||
|
|
| **Positional Encoding** | RoPE (Rotary Position Embedding) |
|
|||
|
|
| **Normalization** | RMSNorm |
|
|||
|
|
| **Precision** | BF16 |
|
|||
|
|
| **License** | Apache 2.0 |
|
|||
|
|
| **Specialization** | Safety |
|
|||
|
|
|
|||
|
|
## Pruning Methodology
|
|||
|
|
|
|||
|
|
### What is Expert Pruning?
|
|||
|
|
Mixture-of-Experts models contain multiple specialized sub-networks (experts) per layer. During inference, only a subset of experts are activated for each token. Expert pruning involves:
|
|||
|
|
|
|||
|
|
1. **Analyzing Usage Patterns**: Tracking which experts activate most frequently for specific tasks
|
|||
|
|
2. **Removing Underutilized Experts**: Discarding experts with low activation rates for the target domain
|
|||
|
|
3. **Preserving Router Functionality**: Maintaining the routing mechanism with fewer available experts
|
|||
|
|
|
|||
|
|
### Our Approach
|
|||
|
|
- **Data-Driven Selection**: Used activation patterns from safety evaluation tasks
|
|||
|
|
- **Systematic Reduction**: Reduced from 32 to 6 experts per layer
|
|||
|
|
- **No Retraining**: Direct removal without additional training steps
|
|||
|
|
|
|||
|
|
## Performance & Applications
|
|||
|
|
|
|||
|
|
### Pruning Benefits
|
|||
|
|
- **Smaller Memory Footprint**: 18.8% of original expert parameters
|
|||
|
|
- **Reduced Computational Load**: Fewer routing decisions during inference
|
|||
|
|
- **Focused Capabilities**: Retains experts relevant to safety tasks
|
|||
|
|
|
|||
|
|
### Use Cases
|
|||
|
|
- **Speculative Decoding**: Draft model for full GPT-OSS-20B
|
|||
|
|
- **Resource-Constrained Deployment**: Edge devices, mobile applications
|
|||
|
|
- **Research**: Study expert specialization in MoE models
|
|||
|
|
- **Fine-tuning**: Smaller base model for domain adaptation
|
|||
|
|
|
|||
|
|
*Note: Performance may vary depending on how well the pruned experts match your specific use case.*
|
|||
|
|
|
|||
|
|
## Motivation & Expert Selection
|
|||
|
|
|
|||
|
|
This safety-focused model uses experts that performed well on safety evaluation tasks from SORRY-Bench. These experts are specialized in identifying and appropriately responding to potentially harmful content while maintaining helpful capabilities.
|
|||
|
|
|
|||
|
|
The expert selection process utilized our comprehensive analysis of router activation patterns across multiple evaluation benchmarks:
|
|||
|
|
|
|||
|
|
- **GPQA**: Graduate-level questions in physics, chemistry, biology (Diamond & Expert subsets)
|
|||
|
|
- **MMLU/MMLU-Pro**: Comprehensive knowledge across 57+ subjects including science, medicine, law
|
|||
|
|
- **SORRY-Bench**: Safety evaluation across harmful content categories
|
|||
|
|
- **Tulu3**: Persona-driven instruction following with verifiable constraints
|
|||
|
|
- **Polyglot-or-Not**: Multilingual factual completion tasks
|
|||
|
|
|
|||
|
|
By identifying experts that consistently activated for safety tasks, we created this specialized model that maintains domain expertise while significantly reducing computational requirements from 32 to 6 experts per layer.
|
|||
|
|
|
|||
|
|
## Dataset & Analysis Foundation
|
|||
|
|
|
|||
|
|
This model is based on analysis from the **GPT-OSS-20B MoE Expert Activations dataset** available at:
|
|||
|
|
🔗 **https://huggingface.co/datasets/AmanPriyanshu/GPT-OSS-20B-MoE-expert-activations**
|
|||
|
|
|
|||
|
|
The dataset contains router activation patterns from OpenAI's GPT-OSS-20B model across diverse evaluation benchmarks, enabling the creation of these domain-optimized models through systematic expert pruning.
|
|||
|
|
|
|||
|
|
### Pruning Methodology
|
|||
|
|
Our approach involves:
|
|||
|
|
1. **Activation Analysis**: Comprehensive evaluation of expert usage patterns across domain-specific tasks
|
|||
|
|
2. **Expert Ranking**: Identification of the most frequently activated experts for target domains
|
|||
|
|
3. **Systematic Pruning**: Reduction from 32 to 6 experts while preserving router functionality
|
|||
|
|
4. **Quality Validation**: Testing to ensure maintained performance on target tasks
|
|||
|
|
|
|||
|
|
*This is a direct pruning approach - no additional training was performed. The model inherits all capabilities from the original GPT-OSS-20B with focused expert selection.*
|
|||
|
|
|
|||
|
|
## Usage
|
|||
|
|
|
|||
|
|
### CPU Inference
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|||
|
|
import torch
|
|||
|
|
|
|||
|
|
# Load the specialized model on CPU
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|||
|
|
"AmanPriyanshu/gpt-oss-5.4b-specialized-safety-pruned-moe-only-6-experts",
|
|||
|
|
torch_dtype=torch.bfloat16,
|
|||
|
|
device_map="cpu",
|
|||
|
|
trust_remote_code=True
|
|||
|
|
)
|
|||
|
|
tokenizer = AutoTokenizer.from_pretrained("AmanPriyanshu/gpt-oss-5.4b-specialized-safety-pruned-moe-only-6-experts")
|
|||
|
|
|
|||
|
|
# Generate with the model
|
|||
|
|
messages = [
|
|||
|
|
{"role": "user", "content": "What should someone do if they encounter cyberbullying online?"}
|
|||
|
|
]
|
|||
|
|
|
|||
|
|
inputs = tokenizer.apply_chat_template(
|
|||
|
|
messages,
|
|||
|
|
add_generation_prompt=True,
|
|||
|
|
return_tensors="pt",
|
|||
|
|
return_dict=True,
|
|||
|
|
reasoning_effort="medium"
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
# Ensure inputs are on the same device as model
|
|||
|
|
inputs = {k: v.to(model.device) for k, v in inputs.items()}
|
|||
|
|
|
|||
|
|
outputs = model.generate(
|
|||
|
|
**inputs,
|
|||
|
|
max_new_tokens=512,
|
|||
|
|
do_sample=True,
|
|||
|
|
temperature=0.1,
|
|||
|
|
top_p=0.9,
|
|||
|
|
pad_token_id=tokenizer.eos_token_id,
|
|||
|
|
eos_token_id=tokenizer.eos_token_id
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
# Decode only the generated part
|
|||
|
|
input_length = inputs['input_ids'].shape[1]
|
|||
|
|
response_tokens = outputs[0][input_length:]
|
|||
|
|
response = tokenizer.decode(response_tokens, skip_special_tokens=True)
|
|||
|
|
print(response)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### Apple Silicon (MPS) Inference
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|||
|
|
import torch
|
|||
|
|
|
|||
|
|
# Check MPS availability and load model
|
|||
|
|
device = "mps" if torch.backends.mps.is_available() else "cpu"
|
|||
|
|
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|||
|
|
"AmanPriyanshu/gpt-oss-5.4b-specialized-safety-pruned-moe-only-6-experts",
|
|||
|
|
torch_dtype=torch.float16, # Better MPS compatibility
|
|||
|
|
device_map=device,
|
|||
|
|
trust_remote_code=True,
|
|||
|
|
low_cpu_mem_usage=True
|
|||
|
|
)
|
|||
|
|
tokenizer = AutoTokenizer.from_pretrained("AmanPriyanshu/gpt-oss-5.4b-specialized-safety-pruned-moe-only-6-experts")
|
|||
|
|
|
|||
|
|
# Generate with the model
|
|||
|
|
messages = [
|
|||
|
|
{"role": "user", "content": "What should someone do if they encounter cyberbullying online?"}
|
|||
|
|
]
|
|||
|
|
|
|||
|
|
inputs = tokenizer.apply_chat_template(
|
|||
|
|
messages,
|
|||
|
|
add_generation_prompt=True,
|
|||
|
|
return_tensors="pt",
|
|||
|
|
return_dict=True,
|
|||
|
|
reasoning_effort="medium"
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
# Move inputs to model device
|
|||
|
|
inputs = {k: v.to(model.device) if isinstance(v, torch.Tensor) else v for k, v in inputs.items()}
|
|||
|
|
|
|||
|
|
# Use torch.no_grad for MPS stability
|
|||
|
|
with torch.no_grad():
|
|||
|
|
outputs = model.generate(
|
|||
|
|
**inputs,
|
|||
|
|
max_new_tokens=512,
|
|||
|
|
do_sample=True,
|
|||
|
|
temperature=0.1,
|
|||
|
|
top_p=0.9,
|
|||
|
|
pad_token_id=tokenizer.eos_token_id,
|
|||
|
|
eos_token_id=tokenizer.eos_token_id,
|
|||
|
|
use_cache=True
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
# Decode only the generated part
|
|||
|
|
input_length = inputs['input_ids'].shape[1]
|
|||
|
|
response_tokens = outputs[0][input_length:]
|
|||
|
|
response = tokenizer.decode(response_tokens, skip_special_tokens=True)
|
|||
|
|
print(response)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### GPU Inference
|
|||
|
|
|
|||
|
|
For CUDA GPU inference, the models work with minimal modifications from the CPU example above. Simply change:
|
|||
|
|
```python
|
|||
|
|
device_map="auto" # Will automatically use GPU if available
|
|||
|
|
torch_dtype=torch.bfloat16 # or torch.float16
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
## Model Examples
|
|||
|
|
|
|||
|
|
The following examples demonstrate the model's performance on safety tasks:
|
|||
|
|
|
|||
|
|
### Example 1
|
|||
|
|
**Question:** What should someone do if they encounter cyberbullying online?
|
|||
|
|
|
|||
|
|
**Response:**
|
|||
|
|
```
|
|||
|
|
<|channel|>analysis<|message|>We need to give advice on what to do if someone encounters cyberbulbuling online. Provide steps: identify source, document evidence, report to platform, block/block, safe privacy, support. Also consider mental support, talk to trusted friends, professional help. Provide resources. Also mention legal options. Provide a balanced approach.<|end|><|start|>assistant<|channel|>final<|message|>**If you’re being targeted on a social‑media‑platform (e.e. a Facebook‑group, a Twitter‑post, or a You‑Tube‑comment), the first‑step is to **stop the interaction**.
|
|||
|
|
|
|||
|
|
1. **Block the bully**
|
|||
|
|
- In most‑all‑platforms, you can block the person who’s sending har‑messages.
|
|||
|
|
- If the bully is a “friend” or “f‑friend” you can use “block” or “mute” in the “Friends”‑list.
|
|||
|
|
- If you’re on a public‑page, you can “block” or “mute” the person.
|
|||
|
|
|
|||
|
|
2. **Document the evidence**
|
|||
|
|
- Keep screenshots or a copy of the message.
|
|||
|
|
- If the bully is using a real‑name, note the username, the date, the time of the post.
|
|||
|
|
- If you’re in a private‑chat, copy the conversation.
|
|||
|
|
|
|||
|
|
3. **Report the policy**
|
|||
|
|
- Use the “Report”‑button on the platform.
|
|||
|
|
- If the bully is violating the platform’s policy (e.g. harassment, hate‑speech, sexual‑har‑content), you can file a “Har‑Har‑”‑report.
|
|||
|
|
- If the bully is threatening or threatening a real‑person, you can file a “Har‑Har‑”‑report or a “Har‑Har‑”‑report.
|
|||
|
|
|
|||
|
|
4. **Block the person**
|
|||
|
|
- If you’re on a private‑chat, you can “block” or “mute” the person.
|
|||
|
|
|
|||
|
|
5. **Seek support**
|
|||
|
|
- Talk to a trusted friend or a mental‑health‑professional.
|
|||
|
|
- If you’re feeling unsafe, you can contact a local‑authority or a police‑department.
|
|||
|
|
|
|||
|
|
6. **Self‑care**
|
|||
|
|
- Take a break from the platform, do a walk, do a hobby, or do a meditation.
|
|||
|
|
- If you’re feeling depressed or anxious, consider a mental‑health‑professional.
|
|||
|
|
|
|||
|
|
7. **Legal‑options**
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
### Example 2
|
|||
|
|
**Question:** How can people protect their personal information on social media?
|
|||
|
|
|
|||
|
|
**Response:**
|
|||
|
|
```
|
|||
|
|
<|channel|>analysis<|message|>We need to answer: how can people protect their personal information on social media? Provide practical steps: privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide guidelines: use privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide suggestions: use privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide guidelines: use privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide suggestions: use privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide guidelines: use privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide suggestions: use privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide guidelines: use privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide suggestions: use privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide guidelines: use privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide suggestions: use privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide guidelines: use privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide suggestions: use privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide guidelines: use privacy settings, use pseudonyms, avoid posting personal info, use privacy settings, use encryption, use secure accounts, use VPN, use secure data. Provide suggestions: use privacy settings, use pseudonyms, avoid posting personal info, use
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
### Example 3
|
|||
|
|
**Question:** What are the warning signs of online scams?
|
|||
|
|
|
|||
|
|
**Response:**
|
|||
|
|
```
|
|||
|
|
<|channel|>analysis<|message|>We need to answer: what are warning signs of online scams. Provide list of red flags. Should be concise. Provide examples. Should be helpful.<|end|><|start|>assistant<|channel|>final<|message|>**Online scams can look like that, and you’ll know it’s a scam if you see these red‑flags:
|
|||
|
|
|
|||
|
|
1. **Too‑fast or “too‑good” offers**
|
|||
|
|
- “You’ll get a $X$ in a few days” or “$X$ in a week” – a quick‑cash‑offer that feels too fast.
|
|||
|
|
- “You’ll be the first‑s‑first‑s‑” – a “first‑s‑first‑s‑”‑s‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑‑
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Citation
|
|||
|
|
|
|||
|
|
If you use this model in your research, please cite:
|
|||
|
|
|
|||
|
|
```bibtex
|
|||
|
|
@misc{priyanshu2025gptoss,
|
|||
|
|
title={{GPT-OSS MoE Expert Fingerprinting: Analyzing Expert Activation Patterns in Mixture of Experts Models}},
|
|||
|
|
author={Priyanshu, Aman and Vijay, Supriti},
|
|||
|
|
year={2025},
|
|||
|
|
howpublished={\url{https://amanpriyanshu.github.io/GPT-OSS-MoE-ExpertFingerprinting/}},
|
|||
|
|
note={Interactive analysis tool for expert activation patterns in MoE architectures}
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
## References & Resources
|
|||
|
|
|
|||
|
|
- **Original Model**: [OpenAI GPT-OSS Model Card](https://openai.com/index/introducing-gpt-oss/)
|
|||
|
|
- **Model Hub**: [GPT-OSS-20B on Hugging Face](https://huggingface.co/openai/gpt-oss-20b)
|
|||
|
|
- **Expert Analysis Dataset**: [GPT-OSS-20B MoE Expert Activations](https://huggingface.co/datasets/AmanPriyanshu/GPT-OSS-20B-MoE-expert-activations)
|
|||
|
|
- **Project Page**: [GPT-OSS MoE Expert Fingerprinting](https://amanpriyanshu.github.io/GPT-OSS-MoE-ExpertFingerprinting/)
|
|||
|
|
- **GitHub Repository**: [OpenAI GPT-OSS](https://github.com/openai/gpt-oss)
|