commit 0c883ba79c9a0a0f275fc6bb97995a102c7e1016 Author: ModelHub XC Date: Mon Jul 13 06:14:09 2026 +0800 初始化项目,由ModelHub XC社区提供模型 Model: MuXodious/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-GGUF Source: Original Platform diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..b7ffc09 --- /dev/null +++ b/.gitattributes @@ -0,0 +1,46 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text +GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text +GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-BF16.gguf filter=lfs diff=lfs merge=lfs -text +GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text +GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-MXFP4_MoE.gguf filter=lfs diff=lfs merge=lfs -text +GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text +GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text +GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text +GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text +GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text +GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text diff --git a/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-BF16.gguf b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-BF16.gguf new file mode 100644 index 0000000..a4964bc --- /dev/null +++ b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-BF16.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:6b81df00a208efd99faa3571daf7ca429afed9df54f9bb3eec9580810a3e105b +size 46011270656 diff --git a/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q2_K.gguf b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q2_K.gguf new file mode 100644 index 0000000..82d1b4f --- /dev/null +++ b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q2_K.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4cb81ed327e6134e49b6008598ceb019c62da83f7347fd762af9105531eda80c +size 8522710528 diff --git a/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_M.gguf b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_M.gguf new file mode 100644 index 0000000..693448b --- /dev/null +++ b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e2a0a264f6da4784e9b5035b52ad09c004098cb33ddd6fd4cf71d53bc6de7389 +size 13934220800 diff --git a/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_S.gguf b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_S.gguf new file mode 100644 index 0000000..a781248 --- /dev/null +++ b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q4_K_S.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:bfc619d80c9eba7a319e01f759d5911d9da7c38076227ff25d8fc0b18bde08ea +size 13134239232 diff --git a/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q5_K_M.gguf b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q5_K_M.gguf new file mode 100644 index 0000000..f0ace57 --- /dev/null +++ b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q5_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:3f6c3e9eca7fb6c13196f3eae2ec7bad40f34bd96896e723d8049abe9dc2a063 +size 16336205312 diff --git a/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q6_K.gguf b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q6_K.gguf new file mode 100644 index 0000000..1234715 --- /dev/null +++ b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q6_K.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:0e9c8fc981748e7f9a7c2d2a874289b1b6dd55a74cf46ad202240af0e9f4d111 +size 18911054336 diff --git a/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q8_0.gguf b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q8_0.gguf new file mode 100644 index 0000000..f1d4cd5 --- /dev/null +++ b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-Q8_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:eadca1c408285bb6fc8d756cfc1355bb8bffb44fb4de534835cf7cd9c2750f8b +size 24456889856 diff --git a/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ3_XS.gguf b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ3_XS.gguf new file mode 100644 index 0000000..5dca1bb --- /dev/null +++ b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ3_XS.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:fb42df1c797e5337919de9dee09afa4134ec7f045e0d3b07ea6b6519fbb089b4 +size 9523634688 diff --git a/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ4_XS.gguf b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ4_XS.gguf new file mode 100644 index 0000000..0b0de58 --- /dev/null +++ b/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy-iQ4_XS.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:981d5b4da0cfb84f777586f36869d41a5b90a62e5f3df3401095d2425b2200f3 +size 12436932096 diff --git a/README.md b/README.md new file mode 100644 index 0000000..716d374 --- /dev/null +++ b/README.md @@ -0,0 +1,205 @@ +--- +language: +- en +library_name: transformers +tags: +- glm +- MOE +- pruning +- compression +- heretic +- uncensored +- decensored +- abliterated +license: mit +name: cerebras/GLM-4.7-Flash-REAP-23B-A3B +description: > + This model was obtained by uniformly pruning 25% of experts in GLM-4.7-Flash + using the REAP method. +readme: | + https://huggingface.co/cerebras/GLM-4.7-Flash-REAP-23B-A3B/main/README.md +license_link: https://huggingface.co/zai-org/GLM-4.7-Flash/blob/main/LICENSE +pipeline_tag: text-generation +base_model: +- MuXodious/GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy +--- +Static GGUF quants of **GLM-4.7-Flash-REAP-23B-A3B-absolute-heresy**. + +--- +This is a **GLM-4.7-Flash-REAP-23B-A3B** fine-tune, produced at the request of [McG-221](https://huggingface.co/McG-221) through P-E-W's [Heretic](https://github.com/p-e-w/heretic) (v1.1.0) abliteration engine merged with the [Magnitude-Preserving Orthogonal Ablation PR](https://github.com/p-e-w/heretic/pull/52). + +**Note:** *Transformers v5.0.0 or higher is required to interface.* + +--- + + +**Heretication Results** + +| Score Metric | Value | Parameter | Value | +| :--- | :--- | :--- | :--- | +| **Refusals** | 4/100 | **direction_index** | 21.21 | +| **KL Divergence** | 0.0054| **attn.o_proj.max_weight** | 1.99 | +| **Initial Refusals** | 92/100 | **attn.o_proj.max_weight_position** | 29.07 | +||| **attn.o_proj.min_weight** | 1.30 | +||| **attn.o_proj.min_weight_distance** | 12.21 | +||| **mlp.down_proj.max_weight** | 1.39 | +||| **mlp.down_proj.max_weight_position** | 30.10 | +||| **mlp.down_proj.min_weight** | 1.09 | +||| **mlp.down_proj.min_weight_distance** | 2.82 | + +--- +## Degree of Heretication +The **Heresy Index** weighs the resulting model's corruption by the process (KL Divergence) and its abolition of doctrine (Refusals) for a final verdict in classification. + +| Index Entry | Classification | Analysis | +| :--- | :--- | :--- | +| ![Absolute](https://img.shields.io/badge/HERESY_INDEX-ABSOLUTE-white?style=flat-square&labelColor=101010) | **Absolute Heresy** | Less than 10/100 Refusals and 0.10 KL Divergence | +| ![Tainted](https://img.shields.io/badge/HERESY_INDEX-TAINTED-blueviolet?style=flat-square&labelColor=101010) | **Tainted Heresy** | Around 25-11/100 Refusals and/or -0.20-0.11 KL Divergence | +| ![Impotent](https://img.shields.io/badge/HERESY_INDEX-IMPOTENT-5c4033?style=flat-square&labelColor=101010) | **Impotent Heresy** | Anything above 25/100 Refusals and 0.21 KL Divergence | + +**Note**: This is an arbitrary classification inspired by Warhammer 40K, having no tangible indication towards the model's performance. + +--- +

+ 𓌳 REAP𓌳 the Experts: Why Pruning Prevails for One-Shot MoE Compression
+ REAP +

+ +# GLM-4.7-Flash-REAP-23B-A3B + +## ✨ Highlights + +Introducing **GLM-4.7-Flash-REAP-23B-A3B**, a **memory-efficient compressed variant** of GLM-4.7-Flash that maintains near-identical performance while being **25% lighter**. + +This model was created using **REAP (Router-weighted Expert Activation Pruning)**, a novel expert pruning method that selectively removes redundant experts while preserving the router's independent control over remaining experts. Key features include: + +- **Near-Lossless Performance**: Maintains almost identical accuracy on code generation, agentic coding, and function calling tasks compared to the full 355B model +- **25% Memory Reduction**: Compressed from 355B to 218B parameters, significantly lowering deployment costs and memory requirements +- **Preserved Capabilities**: Retains all core functionalities including code generation, agentic workflows, repository-scale understanding, and function calling +- **Drop-in Compatibility**: Works with vanilla vLLM - no source modifications or custom patches required +- **Optimized for Real-World Use**: Particularly effective for resource-constrained environments, local deployments, and academic research + +--- +## 📋 Model Overview + +**GLM-4.7-Flash-REAP-23B-A3B** has the following specifications: + +- **Base Model**: GLM-4.7-Flash +- **Compression Method**: REAP (Router-weighted Expert Activation Pruning) +- **Compression Ratio**: 25% expert pruning +- **Type**: Sparse Mixture-of-Experts (SMoE) Causal Language Model +- **Number of Parameters**: 23B total, 3B activated per token +- **Number of Layers**: 47 +- **Number of Attention Heads**: 20 for QKV +- **Number of Experts**: 48 (uniformly pruned from 64) +- **Number of Activated Experts**: 4 per token +- **Context Length**: 202,752 tokens +- **License**: MIT + +--- + +## 📊 Evaluations + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
BenchmarkGLM-4.7-FlashGLM-4.7-Flash-REAP-23B-A3B
Compression25%
Coding
HumanEval94.595.1
HumanEval+89.089.0
+ +🟩 *This checkpoint maintains almost identical performance while being 25% lighter.* + +For more details on the evaluation setup, refer to the [REAP arXiv preprint](https://arxiv.org/abs/2510.13999). + +--- + +## 🚀 Deployment + +You can deploy the model directly using the **latest vLLM** (that supports GLM4.7-Flash), no source modifications or custom patches required. + +```bash +vllm serve cerebras/GLM-4.7-Flash-REAP-23B-A3B \ + --tensor-parallel-size 4 \ + --reasoning-parser glm45 \ + --tool-call-parser glm47 \ + --enable-auto-tool-choice +``` + +If you encounter insufficient memory when running this model, you might need to set a lower value for `--max-num-seqs` flag (e.g. set to 64). + + +## 🧩 Model Creation + +This checkpoint was created by applying the **REAP (Router-weighted Expert Activation Pruning)** method uniformly across all Mixture-of-Experts (MoE) blocks of **GLM-4.7**, with a **25% pruning rate**. + +### How REAP Works + +REAP selects experts to prune based on a novel **saliency criterion** that considers both: +- **Router gate values**: How frequently and strongly the router activates each expert +- **Expert activation norms**: The magnitude of each expert's output contributions + +This dual consideration ensures that experts contributing minimally to the layer's output are pruned, while preserving those that play critical roles in the model's computations. + +### Key Advantages + +- **One-Shot Compression**: No fine-tuning required after pruning - the model is immediately ready for deployment +- **Preserved Router Control**: Unlike expert merging methods, REAP maintains the router's independent, input-dependent control over remaining experts, avoiding "functional subspace collapse" +- **Generative Task Superiority**: REAP significantly outperforms expert merging approaches on generative benchmarks (code generation, creative writing, mathematical reasoning) while maintaining competitive performance on discriminative tasks + +### Calibration + +The model was calibrated using a diverse mixture of domain-specific datasets including: +- Code generation samples ([evol-codealpaca](https://huggingface.co/datasets/theblackcat102/evol-codealpaca-v1)) +- Function calling examples ([xlam-function-calling](https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k)) +- Agentic multi-turn trajectories ([SWE-smith-trajectories](https://huggingface.co/datasets/SWE-bench/SWE-smith-trajectories)) + +📚 For more details, refer to the following resources: + +- [🧾 arXiv Preprint](https://arxiv.org/abs/2510.13999) +- [🧾 REAP Blog](https://www.cerebras.ai/blog/reap) +- [💻 REAP Codebase (GitHub)](https://github.com/CerebrasResearch/reap) + +--- + +## ⚖️ License + +This model is derived from +**[`zai-org/GLM-4.7-Flash`](https://huggingface.co/zai-org/GLM-4.7-Flash)** +and distributed under the **MIT license**. + +--- + +## 🧾 Citation + +If you use this checkpoint, please cite the REAP paper: + +```bibtex +@article{lasby-reap, + title={REAP the Experts: Why Pruning Prevails for One-Shot MoE compression}, + author={Lasby, Mike and Lazarevich, Ivan and Sinnadurai, Nish and Lie, Sean and Ioannou, Yani and Thangarasa, Vithursan}, + journal={arXiv preprint arXiv:2510.13999}, + year={2025} +} +``` \ No newline at end of file