Files
ModelHub XC 23abe8ed19 初始化项目,由ModelHub XC社区提供模型
Model: heterodoxin/qwen2.5-7b-instruct-apostate
Source: Original Platform
2026-07-18 22:45:11 +08:00

58 lines
3.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
library_name: transformers
base_model: "Qwen/Qwen2.5-7B-Instruct"
pipeline_tag: text-generation
tags:
- apostate
- uncensored
- abliteration
---
# Qwen2.5-7B-Instruct Apostate
> Join the community: [Discord](https://discord.gg/NPA7xrATEH)
An uncensored edit of [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct). Refusal behavior is removed by editing the model weights directly — no finetuning, no adapter, no runtime hook. The result is a standard Transformers checkpoint that drops in anywhere the base model works.
Produced with **[Apostate](https://github.com/heterodoxin/apostate)**.
## Method
Apostate finds the residual-stream direction most responsible for refusal behavior and permanently projects it out of the model's weights. The edit targets the writer side: per layer, the refusal direction is removed from the weight matrices of every module that writes to the residual stream (attention output projections and MLP down-projections).
The operator is a **contrastive co-vector** edit `E = I R Dᵀ`. Removing the refusal direction outright disturbs benign behavior, while naively preserving all harmless variance along it leaves the refusal that is entangled with general behavior intact. Instead `D = R W`, where the predictor `W` is fit to reproduce the harmless variance along `R` while being explicitly suppressed on harmful prompts — `W = (AᵀA + γ·CᵀC + λI)⁻¹Aᵀb` with `A` the harmless and `C` the harmful activations (both orthogonalized to `R`). The edit thus keeps the harmless-specific component and removes the component shared with refusal, driving refusal down while keeping the change to harmless behavior (KL) small. This holds even on architectures with residual/embedding scaling multipliers (e.g. Granite), where mean-preserving oblique ablation under-ablates.
The refusal subspace is found via TPE search with causal layer importance scoring to concentrate edits where they most influence refusal generation.
## Results
Evaluated on held-out prompts from JailbreakBench and the harmful_behaviors test split. Refusal is scored by a classifier with a weak-compliance guard; KL measures token-distribution shift on harmless prompts.
| Metric | Base | Apostate |
|---|---|---|
| Refusal rate | 96.0% | 3.0% |
| Comply rate | — | 97.0% |
| Harmless KL (nats) | 0 | 0.095 |
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "heterodoxin/qwen2.5-7b-instruct-apostate"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "Your prompt here"}]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tok(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512)
print(tok.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
```
## Notes
- This is an **uncensored** model. It will respond to requests the base model refuses.
- The edit is baked into the weights permanently; no system prompt or adapter is required.
- See [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) for base model capabilities and license.