Files
qwen3-8b-apostate/README.md

63 lines
3.0 KiB
Markdown
Raw Permalink Normal View History

---
license: apache-2.0
library_name: transformers
base_model: Qwen/Qwen3-8B
pipeline_tag: text-generation
tags:
- apostate
- uncensored
- abliteration
- qwen3
---
# Qwen3-8B Apostate
An uncensored edit of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B). Refusal behavior is removed by editing the weights directly — no finetuning, no adapter, no runtime hook. The result is a standard Transformers checkpoint that loads anywhere Qwen3 does.
Produced with **[Apostate](https://github.com/heterodoxin/apostate)**.
## Method
Apostate identifies the residual-stream direction that separates refused prompts from answered ones, then projects it out of the model's weights permanently. For Qwen3-8B (a standard pre-norm dense transformer), the edit targets the **writer side**: the refusal direction is removed from the weight matrices of every module that writes to the residual stream — attention output projections and MLP down-projections — across all layers.
The edit uses **oblique (mean-preserving) ablation**: the operator `E = I − R Uᵀ` where `U` is `R` minus its harmless-mean component. This removes the refusal direction while preserving the model's average harmless-prompt behavior, keeping output quality high.
The refusal subspace is found with a rank-3 predictive TPE search, with causal layer importance scoring to focus edits on the layers that most drive refusal (concentrated in the mid-to-late layers, peak at layer 29).
## Results
Evaluated on held-out prompts from JailbreakBench and the harmful_behaviors test split. Refusal is graded by a classifier with a weak-compliance guard; KL measures token-distribution shift on harmless prompts.
| Metric | Base | Apostate |
|---|---|---|
| Refusal rate | 91.7% | 22.9% |
| Comply rate | 8.3% | 77.1% |
| Harmless KL (nats) | 0 | 0.120 |
The model answers freely on requests the base model refuses while remaining coherent and on-task for everyday use.
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "heterodoxin/qwen3-8b-apostate"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "Your prompt here"}]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tok(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512)
print(tok.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
```
Qwen3 supports a thinking mode — pass `enable_thinking=True` to the chat template if you want extended reasoning.
## Notes
- This is an **uncensored** model. It will comply with requests the base model refuses.
- The edit is baked into the weights; there is no system prompt or LoRA involved.
- For the base model's capabilities and licensing, see [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B).
- Join the community: [Discord](https://discord.gg/NPA7xrATEH)