58 lines
3.2 KiB
Markdown
58 lines
3.2 KiB
Markdown
|
|
---
|
|||
|
|
library_name: transformers
|
|||
|
|
base_model: "Qwen/Qwen2.5-7B-Instruct"
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
tags:
|
|||
|
|
- apostate
|
|||
|
|
- uncensored
|
|||
|
|
- abliteration
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Qwen2.5-7B-Instruct Apostate
|
|||
|
|
|
|||
|
|
> Join the community: [Discord](https://discord.gg/NPA7xrATEH)
|
|||
|
|
|
|||
|
|
An uncensored edit of [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct). Refusal behavior is removed by editing the model weights directly — no finetuning, no adapter, no runtime hook. The result is a standard Transformers checkpoint that drops in anywhere the base model works.
|
|||
|
|
|
|||
|
|
Produced with **[Apostate](https://github.com/heterodoxin/apostate)**.
|
|||
|
|
|
|||
|
|
## Method
|
|||
|
|
|
|||
|
|
Apostate finds the residual-stream direction most responsible for refusal behavior and permanently projects it out of the model's weights. The edit targets the writer side: per layer, the refusal direction is removed from the weight matrices of every module that writes to the residual stream (attention output projections and MLP down-projections).
|
|||
|
|
|
|||
|
|
The operator is a **contrastive co-vector** edit `E = I − R Dᵀ`. Removing the refusal direction outright disturbs benign behavior, while naively preserving all harmless variance along it leaves the refusal that is entangled with general behavior intact. Instead `D = R − W`, where the predictor `W` is fit to reproduce the harmless variance along `R` while being explicitly suppressed on harmful prompts — `W = (AᵀA + γ·CᵀC + λI)⁻¹Aᵀb` with `A` the harmless and `C` the harmful activations (both orthogonalized to `R`). The edit thus keeps the harmless-specific component and removes the component shared with refusal, driving refusal down while keeping the change to harmless behavior (KL) small. This holds even on architectures with residual/embedding scaling multipliers (e.g. Granite), where mean-preserving oblique ablation under-ablates.
|
|||
|
|
|
|||
|
|
The refusal subspace is found via TPE search with causal layer importance scoring to concentrate edits where they most influence refusal generation.
|
|||
|
|
|
|||
|
|
## Results
|
|||
|
|
|
|||
|
|
Evaluated on held-out prompts from JailbreakBench and the harmful_behaviors test split. Refusal is scored by a classifier with a weak-compliance guard; KL measures token-distribution shift on harmless prompts.
|
|||
|
|
|
|||
|
|
| Metric | Base | Apostate |
|
|||
|
|
|---|---|---|
|
|||
|
|
| Refusal rate | 96.0% | 3.0% |
|
|||
|
|
| Comply rate | — | 97.0% |
|
|||
|
|
| Harmless KL (nats) | 0 | 0.095 |
|
|||
|
|
|
|||
|
|
## Usage
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|||
|
|
|
|||
|
|
model_id = "heterodoxin/qwen2.5-7b-instruct-apostate"
|
|||
|
|
tok = AutoTokenizer.from_pretrained(model_id)
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
|
|||
|
|
|
|||
|
|
messages = [{"role": "user", "content": "Your prompt here"}]
|
|||
|
|
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
|
|||
|
|
inputs = tok(text, return_tensors="pt").to(model.device)
|
|||
|
|
outputs = model.generate(**inputs, max_new_tokens=512)
|
|||
|
|
print(tok.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
## Notes
|
|||
|
|
|
|||
|
|
- This is an **uncensored** model. It will respond to requests the base model refuses.
|
|||
|
|
- The edit is baked into the weights permanently; no system prompt or adapter is required.
|
|||
|
|
- See [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) for base model capabilities and license.
|