初始化项目,由ModelHub XC社区提供模型

Model: MuXodious/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-13 06:14:09 +08:00
commit 0c883ba79c
11 changed files with 278 additions and 0 deletions

46
.gitattributes vendored Normal file
View File

@@ -0,0 +1,46 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-BF16.gguf filter=lfs diff=lfs merge=lfs -text
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-MXFP4_MoE.gguf filter=lfs diff=lfs merge=lfs -text
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6b81df00a208efd99faa3571daf7ca429afed9df54f9bb3eec9580810a3e105b
size 46011270656

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4cb81ed327e6134e49b6008598ceb019c62da83f7347fd762af9105531eda80c
size 8522710528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e2a0a264f6da4784e9b5035b52ad09c004098cb33ddd6fd4cf71d53bc6de7389
size 13934220800

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bfc619d80c9eba7a319e01f759d5911d9da7c38076227ff25d8fc0b18bde08ea
size 13134239232

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3f6c3e9eca7fb6c13196f3eae2ec7bad40f34bd96896e723d8049abe9dc2a063
size 16336205312

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0e9c8fc981748e7f9a7c2d2a874289b1b6dd55a74cf46ad202240af0e9f4d111
size 18911054336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:eadca1c408285bb6fc8d756cfc1355bb8bffb44fb4de534835cf7cd9c2750f8b
size 24456889856

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fb42df1c797e5337919de9dee09afa4134ec7f045e0d3b07ea6b6519fbb089b4
size 9523634688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:981d5b4da0cfb84f777586f36869d41a5b90a62e5f3df3401095d2425b2200f3
size 12436932096

205
README.md Normal file
View File

@@ -0,0 +1,205 @@
---
language:
- en
library_name: transformers
tags:
- glm
- MOE
- pruning
- compression
- heretic
- uncensored
- decensored
- abliterated
license: mit
name: cerebras/GLM-4.7-Flash-REAP-23B-A3B
description: >
This model was obtained by uniformly pruning 25% of experts in GLM-4.7-Flash
using the REAP method.
readme: |
https://huggingface.co/cerebras/GLM-4.7-Flash-REAP-23B-A3B/main/README.md
license_link: https://huggingface.co/zai-org/GLM-4.7-Flash/blob/main/LICENSE
pipeline_tag: text-generation
base_model:
- MuXodious/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy
---
Static GGUF quants of **GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy**.
---
This is a **GLM-4.7-Flash-REAP-23B-A3B** fine-tune, produced at the request of [McG-221](https://huggingface.co/McG-221) through P-E-W's [Heretic](https://github.com/p-e-w/heretic) (v1.1.0) abliteration engine merged with the [Magnitude-Preserving Orthogonal Ablation PR](https://github.com/p-e-w/heretic/pull/52).
**Note:** *Transformers v5.0.0 or higher is required to interface.*
---
<img src="https://img.shields.io/badge/HERESY_INDEX-ABSOLUTE-white?style=flat-square&labelColor=101010" align="right" width="250">
**Heretication Results**
| Score Metric | Value | Parameter | Value |
| :--- | :--- | :--- | :--- |
| **Refusals** | 4/100 | **direction_index** | 21.21 |
| **KL Divergence** | 0.0054| **attn.o_proj.max_weight** | 1.99 |
| **Initial Refusals** | 92/100 | **attn.o_proj.max_weight_position** | 29.07 |
||| **attn.o_proj.min_weight** | 1.30 |
||| **attn.o_proj.min_weight_distance** | 12.21 |
||| **mlp.down_proj.max_weight** | 1.39 |
||| **mlp.down_proj.max_weight_position** | 30.10 |
||| **mlp.down_proj.min_weight** | 1.09 |
||| **mlp.down_proj.min_weight_distance** | 2.82 |
---
## Degree of Heretication
The **Heresy Index** weighs the resulting model's corruption by the process (KL Divergence) and its abolition of doctrine (Refusals) for a final verdict in classification.
| Index Entry | Classification | Analysis |
| :--- | :--- | :--- |
| ![Absolute](https://img.shields.io/badge/HERESY_INDEX-ABSOLUTE-white?style=flat-square&labelColor=101010) | **Absolute Heresy** | Less than 10/100 Refusals and 0.10 KL Divergence |
| ![Tainted](https://img.shields.io/badge/HERESY_INDEX-TAINTED-blueviolet?style=flat-square&labelColor=101010) | **Tainted Heresy** | Around 25-11/100 Refusals and/or -0.20-0.11 KL Divergence |
| ![Impotent](https://img.shields.io/badge/HERESY_INDEX-IMPOTENT-5c4033?style=flat-square&labelColor=101010) | **Impotent Heresy** | Anything above 25/100 Refusals and 0.21 KL Divergence |
**Note**: This is an arbitrary classification inspired by Warhammer 40K, having no tangible indication towards the model's performance.
---
<p align="center">
<em>𓌳 <strong>REAP</strong>𓌳 the Experts: Why Pruning Prevails for One-Shot MoE Compression</em><br>
<img src="https://i.imgur.com/rmzG3gg.png" alt="REAP" width="75%">
</p>
# GLM-4.7-Flash-REAP-23B-A3B
## ✨ Highlights
Introducing **GLM-4.7-Flash-REAP-23B-A3B**, a **memory-efficient compressed variant** of GLM-4.7-Flash that maintains near-identical performance while being **25% lighter**.
This model was created using **REAP (Router-weighted Expert Activation Pruning)**, a novel expert pruning method that selectively removes redundant experts while preserving the router's independent control over remaining experts. Key features include:
- **Near-Lossless Performance**: Maintains almost identical accuracy on code generation, agentic coding, and function calling tasks compared to the full 355B model
- **25% Memory Reduction**: Compressed from 355B to 218B parameters, significantly lowering deployment costs and memory requirements
- **Preserved Capabilities**: Retains all core functionalities including code generation, agentic workflows, repository-scale understanding, and function calling
- **Drop-in Compatibility**: Works with vanilla vLLM - no source modifications or custom patches required
- **Optimized for Real-World Use**: Particularly effective for resource-constrained environments, local deployments, and academic research
---
## 📋 Model Overview
**GLM-4.7-Flash-REAP-23B-A3B** has the following specifications:
- **Base Model**: GLM-4.7-Flash
- **Compression Method**: REAP (Router-weighted Expert Activation Pruning)
- **Compression Ratio**: 25% expert pruning
- **Type**: Sparse Mixture-of-Experts (SMoE) Causal Language Model
- **Number of Parameters**: 23B total, 3B activated per token
- **Number of Layers**: 47
- **Number of Attention Heads**: 20 for QKV
- **Number of Experts**: 48 (uniformly pruned from 64)
- **Number of Activated Experts**: 4 per token
- **Context Length**: 202,752 tokens
- **License**: MIT
---
## 📊 Evaluations
<table>
<thead>
<tr>
<th align="left">Benchmark</th>
<th align="center">GLM-4.7-Flash</th>
<th align="center"><a href="https://huggingface.co/cerebras/GLM-4.7-Flash-REAP-23B-A3B">GLM-4.7-Flash-REAP-23B-A3B</a></th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>Compression</strong></td>
<td align="center"></td>
<td align="center">25%</td>
</tr>
<tr>
<td colspan="5" align="center"><strong>Coding</strong></td>
</tr>
<tr>
<td><strong>HumanEval</strong></td>
<td align="center">94.5</td>
<td align="center">95.1</td>
</tr>
<tr>
<td><strong>HumanEval+</strong></td>
<td align="center">89.0</td>
<td align="center">89.0</td>
</tr>
</table>
🟩 *This checkpoint maintains almost identical performance while being 25% lighter.*
For more details on the evaluation setup, refer to the [REAP arXiv preprint](https://arxiv.org/abs/2510.13999).
---
## 🚀 Deployment
You can deploy the model directly using the **latest vLLM** (that supports GLM4.7-Flash), no source modifications or custom patches required.
```bash
vllm serve cerebras/GLM-4.7-Flash-REAP-23B-A3B \
--tensor-parallel-size 4 \
--reasoning-parser glm45 \
--tool-call-parser glm47 \
--enable-auto-tool-choice
```
If you encounter insufficient memory when running this model, you might need to set a lower value for `--max-num-seqs` flag (e.g. set to 64).
## 🧩 Model Creation
This checkpoint was created by applying the **REAP (Router-weighted Expert Activation Pruning)** method uniformly across all Mixture-of-Experts (MoE) blocks of **GLM-4.7**, with a **25% pruning rate**.
### How REAP Works
REAP selects experts to prune based on a novel **saliency criterion** that considers both:
- **Router gate values**: How frequently and strongly the router activates each expert
- **Expert activation norms**: The magnitude of each expert's output contributions
This dual consideration ensures that experts contributing minimally to the layer's output are pruned, while preserving those that play critical roles in the model's computations.
### Key Advantages
- **One-Shot Compression**: No fine-tuning required after pruning - the model is immediately ready for deployment
- **Preserved Router Control**: Unlike expert merging methods, REAP maintains the router's independent, input-dependent control over remaining experts, avoiding "functional subspace collapse"
- **Generative Task Superiority**: REAP significantly outperforms expert merging approaches on generative benchmarks (code generation, creative writing, mathematical reasoning) while maintaining competitive performance on discriminative tasks
### Calibration
The model was calibrated using a diverse mixture of domain-specific datasets including:
- Code generation samples ([evol-codealpaca](https://huggingface.co/datasets/theblackcat102/evol-codealpaca-v1))
- Function calling examples ([xlam-function-calling](https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k))
- Agentic multi-turn trajectories ([SWE-smith-trajectories](https://huggingface.co/datasets/SWE-bench/SWE-smith-trajectories))
📚 For more details, refer to the following resources:
- [🧾 arXiv Preprint](https://arxiv.org/abs/2510.13999)
- [🧾 REAP Blog](https://www.cerebras.ai/blog/reap)
- [💻 REAP Codebase (GitHub)](https://github.com/CerebrasResearch/reap)
---
## ⚖️ License
This model is derived from
**[`zai-org/GLM-4.7-Flash`](https://huggingface.co/zai-org/GLM-4.7-Flash)**
and distributed under the **MIT license**.
---
## 🧾 Citation
If you use this checkpoint, please cite the REAP paper:
```bibtex
@article{lasby-reap,
title={REAP the Experts: Why Pruning Prevails for One-Shot MoE compression},
author={Lasby, Mike and Lazarevich, Ivan and Sinnadurai, Nish and Lie, Sean and Ioannou, Yani and Thangarasa, Vithursan},
journal={arXiv preprint arXiv:2510.13999},
year={2025}
}
```