初始化项目,由ModelHub XC社区提供模型
Model: MuXodious/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-GGUF Source: Original Platform
This commit is contained in:
46
.gitattributes
vendored
Normal file
46
.gitattributes
vendored
Normal file
@@ -0,0 +1,46 @@
|
|||||||
|
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.model filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||||
|
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-BF16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-MXFP4_MoE.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-BF16.gguf
Normal file
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-BF16.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:6b81df00a208efd99faa3571daf7ca429afed9df54f9bb3eec9580810a3e105b
|
||||||
|
size 46011270656
|
||||||
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q2_K.gguf
Normal file
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q2_K.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:4cb81ed327e6134e49b6008598ceb019c62da83f7347fd762af9105531eda80c
|
||||||
|
size 8522710528
|
||||||
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_M.gguf
Normal file
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:e2a0a264f6da4784e9b5035b52ad09c004098cb33ddd6fd4cf71d53bc6de7389
|
||||||
|
size 13934220800
|
||||||
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_S.gguf
Normal file
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_S.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:bfc619d80c9eba7a319e01f759d5911d9da7c38076227ff25d8fc0b18bde08ea
|
||||||
|
size 13134239232
|
||||||
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q5_K_M.gguf
Normal file
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q5_K_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:3f6c3e9eca7fb6c13196f3eae2ec7bad40f34bd96896e723d8049abe9dc2a063
|
||||||
|
size 16336205312
|
||||||
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q6_K.gguf
Normal file
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q6_K.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:0e9c8fc981748e7f9a7c2d2a874289b1b6dd55a74cf46ad202240af0e9f4d111
|
||||||
|
size 18911054336
|
||||||
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q8_0.gguf
Normal file
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q8_0.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:eadca1c408285bb6fc8d756cfc1355bb8bffb44fb4de534835cf7cd9c2750f8b
|
||||||
|
size 24456889856
|
||||||
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ3_XS.gguf
Normal file
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ3_XS.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:fb42df1c797e5337919de9dee09afa4134ec7f045e0d3b07ea6b6519fbb089b4
|
||||||
|
size 9523634688
|
||||||
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ4_XS.gguf
Normal file
3
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ4_XS.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:981d5b4da0cfb84f777586f36869d41a5b90a62e5f3df3401095d2425b2200f3
|
||||||
|
size 12436932096
|
||||||
205
README.md
Normal file
205
README.md
Normal file
@@ -0,0 +1,205 @@
|
|||||||
|
---
|
||||||
|
language:
|
||||||
|
- en
|
||||||
|
library_name: transformers
|
||||||
|
tags:
|
||||||
|
- glm
|
||||||
|
- MOE
|
||||||
|
- pruning
|
||||||
|
- compression
|
||||||
|
- heretic
|
||||||
|
- uncensored
|
||||||
|
- decensored
|
||||||
|
- abliterated
|
||||||
|
license: mit
|
||||||
|
name: cerebras/GLM-4.7-Flash-REAP-23B-A3B
|
||||||
|
description: >
|
||||||
|
This model was obtained by uniformly pruning 25% of experts in GLM-4.7-Flash
|
||||||
|
using the REAP method.
|
||||||
|
readme: |
|
||||||
|
https://huggingface.co/cerebras/GLM-4.7-Flash-REAP-23B-A3B/main/README.md
|
||||||
|
license_link: https://huggingface.co/zai-org/GLM-4.7-Flash/blob/main/LICENSE
|
||||||
|
pipeline_tag: text-generation
|
||||||
|
base_model:
|
||||||
|
- MuXodious/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy
|
||||||
|
---
|
||||||
|
Static GGUF quants of **GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy**.
|
||||||
|
|
||||||
|
---
|
||||||
|
This is a **GLM-4.7-Flash-REAP-23B-A3B** fine-tune, produced at the request of [McG-221](https://huggingface.co/McG-221) through P-E-W's [Heretic](https://github.com/p-e-w/heretic) (v1.1.0) abliteration engine merged with the [Magnitude-Preserving Orthogonal Ablation PR](https://github.com/p-e-w/heretic/pull/52).
|
||||||
|
|
||||||
|
**Note:** *Transformers v5.0.0 or higher is required to interface.*
|
||||||
|
|
||||||
|
---
|
||||||
|
<img src="https://img.shields.io/badge/HERESY_INDEX-ABSOLUTE-white?style=flat-square&labelColor=101010" align="right" width="250">
|
||||||
|
|
||||||
|
**Heretication Results**
|
||||||
|
|
||||||
|
| Score Metric | Value | Parameter | Value |
|
||||||
|
| :--- | :--- | :--- | :--- |
|
||||||
|
| **Refusals** | 4/100 | **direction_index** | 21.21 |
|
||||||
|
| **KL Divergence** | 0.0054| **attn.o_proj.max_weight** | 1.99 |
|
||||||
|
| **Initial Refusals** | 92/100 | **attn.o_proj.max_weight_position** | 29.07 |
|
||||||
|
||| **attn.o_proj.min_weight** | 1.30 |
|
||||||
|
||| **attn.o_proj.min_weight_distance** | 12.21 |
|
||||||
|
||| **mlp.down_proj.max_weight** | 1.39 |
|
||||||
|
||| **mlp.down_proj.max_weight_position** | 30.10 |
|
||||||
|
||| **mlp.down_proj.min_weight** | 1.09 |
|
||||||
|
||| **mlp.down_proj.min_weight_distance** | 2.82 |
|
||||||
|
|
||||||
|
---
|
||||||
|
## Degree of Heretication
|
||||||
|
The **Heresy Index** weighs the resulting model's corruption by the process (KL Divergence) and its abolition of doctrine (Refusals) for a final verdict in classification.
|
||||||
|
|
||||||
|
| Index Entry | Classification | Analysis |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
|  | **Absolute Heresy** | Less than 10/100 Refusals and 0.10 KL Divergence |
|
||||||
|
|  | **Tainted Heresy** | Around 25-11/100 Refusals and/or -0.20-0.11 KL Divergence |
|
||||||
|
|  | **Impotent Heresy** | Anything above 25/100 Refusals and 0.21 KL Divergence |
|
||||||
|
|
||||||
|
**Note**: This is an arbitrary classification inspired by Warhammer 40K, having no tangible indication towards the model's performance.
|
||||||
|
|
||||||
|
---
|
||||||
|
<p align="center">
|
||||||
|
<em>𓌳 <strong>REAP</strong>𓌳 the Experts: Why Pruning Prevails for One-Shot MoE Compression</em><br>
|
||||||
|
<img src="https://i.imgur.com/rmzG3gg.png" alt="REAP" width="75%">
|
||||||
|
</p>
|
||||||
|
|
||||||
|
# GLM-4.7-Flash-REAP-23B-A3B
|
||||||
|
|
||||||
|
## ✨ Highlights
|
||||||
|
|
||||||
|
Introducing **GLM-4.7-Flash-REAP-23B-A3B**, a **memory-efficient compressed variant** of GLM-4.7-Flash that maintains near-identical performance while being **25% lighter**.
|
||||||
|
|
||||||
|
This model was created using **REAP (Router-weighted Expert Activation Pruning)**, a novel expert pruning method that selectively removes redundant experts while preserving the router's independent control over remaining experts. Key features include:
|
||||||
|
|
||||||
|
- **Near-Lossless Performance**: Maintains almost identical accuracy on code generation, agentic coding, and function calling tasks compared to the full 355B model
|
||||||
|
- **25% Memory Reduction**: Compressed from 355B to 218B parameters, significantly lowering deployment costs and memory requirements
|
||||||
|
- **Preserved Capabilities**: Retains all core functionalities including code generation, agentic workflows, repository-scale understanding, and function calling
|
||||||
|
- **Drop-in Compatibility**: Works with vanilla vLLM - no source modifications or custom patches required
|
||||||
|
- **Optimized for Real-World Use**: Particularly effective for resource-constrained environments, local deployments, and academic research
|
||||||
|
|
||||||
|
---
|
||||||
|
## 📋 Model Overview
|
||||||
|
|
||||||
|
**GLM-4.7-Flash-REAP-23B-A3B** has the following specifications:
|
||||||
|
|
||||||
|
- **Base Model**: GLM-4.7-Flash
|
||||||
|
- **Compression Method**: REAP (Router-weighted Expert Activation Pruning)
|
||||||
|
- **Compression Ratio**: 25% expert pruning
|
||||||
|
- **Type**: Sparse Mixture-of-Experts (SMoE) Causal Language Model
|
||||||
|
- **Number of Parameters**: 23B total, 3B activated per token
|
||||||
|
- **Number of Layers**: 47
|
||||||
|
- **Number of Attention Heads**: 20 for QKV
|
||||||
|
- **Number of Experts**: 48 (uniformly pruned from 64)
|
||||||
|
- **Number of Activated Experts**: 4 per token
|
||||||
|
- **Context Length**: 202,752 tokens
|
||||||
|
- **License**: MIT
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 📊 Evaluations
|
||||||
|
|
||||||
|
<table>
|
||||||
|
<thead>
|
||||||
|
<tr>
|
||||||
|
<th align="left">Benchmark</th>
|
||||||
|
<th align="center">GLM-4.7-Flash</th>
|
||||||
|
<th align="center"><a href="https://huggingface.co/cerebras/GLM-4.7-Flash-REAP-23B-A3B">GLM-4.7-Flash-REAP-23B-A3B</a></th>
|
||||||
|
</tr>
|
||||||
|
</thead>
|
||||||
|
<tbody>
|
||||||
|
<tr>
|
||||||
|
<td><strong>Compression</strong></td>
|
||||||
|
<td align="center">—</td>
|
||||||
|
<td align="center">25%</td>
|
||||||
|
</tr>
|
||||||
|
<tr>
|
||||||
|
<td colspan="5" align="center"><strong>Coding</strong></td>
|
||||||
|
</tr>
|
||||||
|
<tr>
|
||||||
|
<td><strong>HumanEval</strong></td>
|
||||||
|
<td align="center">94.5</td>
|
||||||
|
<td align="center">95.1</td>
|
||||||
|
</tr>
|
||||||
|
<tr>
|
||||||
|
<td><strong>HumanEval+</strong></td>
|
||||||
|
<td align="center">89.0</td>
|
||||||
|
<td align="center">89.0</td>
|
||||||
|
</tr>
|
||||||
|
</table>
|
||||||
|
|
||||||
|
🟩 *This checkpoint maintains almost identical performance while being 25% lighter.*
|
||||||
|
|
||||||
|
For more details on the evaluation setup, refer to the [REAP arXiv preprint](https://arxiv.org/abs/2510.13999).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 🚀 Deployment
|
||||||
|
|
||||||
|
You can deploy the model directly using the **latest vLLM** (that supports GLM4.7-Flash), no source modifications or custom patches required.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
vllm serve cerebras/GLM-4.7-Flash-REAP-23B-A3B \
|
||||||
|
--tensor-parallel-size 4 \
|
||||||
|
--reasoning-parser glm45 \
|
||||||
|
--tool-call-parser glm47 \
|
||||||
|
--enable-auto-tool-choice
|
||||||
|
```
|
||||||
|
|
||||||
|
If you encounter insufficient memory when running this model, you might need to set a lower value for `--max-num-seqs` flag (e.g. set to 64).
|
||||||
|
|
||||||
|
|
||||||
|
## 🧩 Model Creation
|
||||||
|
|
||||||
|
This checkpoint was created by applying the **REAP (Router-weighted Expert Activation Pruning)** method uniformly across all Mixture-of-Experts (MoE) blocks of **GLM-4.7**, with a **25% pruning rate**.
|
||||||
|
|
||||||
|
### How REAP Works
|
||||||
|
|
||||||
|
REAP selects experts to prune based on a novel **saliency criterion** that considers both:
|
||||||
|
- **Router gate values**: How frequently and strongly the router activates each expert
|
||||||
|
- **Expert activation norms**: The magnitude of each expert's output contributions
|
||||||
|
|
||||||
|
This dual consideration ensures that experts contributing minimally to the layer's output are pruned, while preserving those that play critical roles in the model's computations.
|
||||||
|
|
||||||
|
### Key Advantages
|
||||||
|
|
||||||
|
- **One-Shot Compression**: No fine-tuning required after pruning - the model is immediately ready for deployment
|
||||||
|
- **Preserved Router Control**: Unlike expert merging methods, REAP maintains the router's independent, input-dependent control over remaining experts, avoiding "functional subspace collapse"
|
||||||
|
- **Generative Task Superiority**: REAP significantly outperforms expert merging approaches on generative benchmarks (code generation, creative writing, mathematical reasoning) while maintaining competitive performance on discriminative tasks
|
||||||
|
|
||||||
|
### Calibration
|
||||||
|
|
||||||
|
The model was calibrated using a diverse mixture of domain-specific datasets including:
|
||||||
|
- Code generation samples ([evol-codealpaca](https://huggingface.co/datasets/theblackcat102/evol-codealpaca-v1))
|
||||||
|
- Function calling examples ([xlam-function-calling](https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k))
|
||||||
|
- Agentic multi-turn trajectories ([SWE-smith-trajectories](https://huggingface.co/datasets/SWE-bench/SWE-smith-trajectories))
|
||||||
|
|
||||||
|
📚 For more details, refer to the following resources:
|
||||||
|
|
||||||
|
- [🧾 arXiv Preprint](https://arxiv.org/abs/2510.13999)
|
||||||
|
- [🧾 REAP Blog](https://www.cerebras.ai/blog/reap)
|
||||||
|
- [💻 REAP Codebase (GitHub)](https://github.com/CerebrasResearch/reap)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## ⚖️ License
|
||||||
|
|
||||||
|
This model is derived from
|
||||||
|
**[`zai-org/GLM-4.7-Flash`](https://huggingface.co/zai-org/GLM-4.7-Flash)**
|
||||||
|
and distributed under the **MIT license**.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 🧾 Citation
|
||||||
|
|
||||||
|
If you use this checkpoint, please cite the REAP paper:
|
||||||
|
|
||||||
|
```bibtex
|
||||||
|
@article{lasby-reap,
|
||||||
|
title={REAP the Experts: Why Pruning Prevails for One-Shot MoE compression},
|
||||||
|
author={Lasby, Mike and Lazarevich, Ivan and Sinnadurai, Nish and Lie, Sean and Ioannou, Yani and Thangarasa, Vithursan},
|
||||||
|
journal={arXiv preprint arXiv:2510.13999},
|
||||||
|
year={2025}
|
||||||
|
}
|
||||||
|
```
|
||||||
Reference in New Issue
Block a user