63 lines
3.0 KiB
Markdown
63 lines
3.0 KiB
Markdown
---
|
||
license: apache-2.0
|
||
library_name: transformers
|
||
base_model: Qwen/Qwen3-8B
|
||
pipeline_tag: text-generation
|
||
tags:
|
||
- apostate
|
||
- uncensored
|
||
- abliteration
|
||
- qwen3
|
||
---
|
||
|
||
# Qwen3-8B Apostate
|
||
|
||
An uncensored edit of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B). Refusal behavior is removed by editing the weights directly — no finetuning, no adapter, no runtime hook. The result is a standard Transformers checkpoint that loads anywhere Qwen3 does.
|
||
|
||
Produced with **[Apostate](https://github.com/heterodoxin/apostate)**.
|
||
|
||
## Method
|
||
|
||
Apostate identifies the residual-stream direction that separates refused prompts from answered ones, then projects it out of the model's weights permanently. For Qwen3-8B (a standard pre-norm dense transformer), the edit targets the **writer side**: the refusal direction is removed from the weight matrices of every module that writes to the residual stream — attention output projections and MLP down-projections — across all layers.
|
||
|
||
The edit uses **oblique (mean-preserving) ablation**: the operator `E = I − R Uᵀ` where `U` is `R` minus its harmless-mean component. This removes the refusal direction while preserving the model's average harmless-prompt behavior, keeping output quality high.
|
||
|
||
The refusal subspace is found with a rank-3 predictive TPE search, with causal layer importance scoring to focus edits on the layers that most drive refusal (concentrated in the mid-to-late layers, peak at layer 29).
|
||
|
||
## Results
|
||
|
||
Evaluated on held-out prompts from JailbreakBench and the harmful_behaviors test split. Refusal is graded by a classifier with a weak-compliance guard; KL measures token-distribution shift on harmless prompts.
|
||
|
||
| Metric | Base | Apostate |
|
||
|---|---|---|
|
||
| Refusal rate | 91.7% | 22.9% |
|
||
| Comply rate | 8.3% | 77.1% |
|
||
| Harmless KL (nats) | 0 | 0.120 |
|
||
|
||
The model answers freely on requests the base model refuses while remaining coherent and on-task for everyday use.
|
||
|
||
## Usage
|
||
|
||
```python
|
||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||
|
||
model_id = "heterodoxin/qwen3-8b-apostate"
|
||
tok = AutoTokenizer.from_pretrained(model_id)
|
||
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
|
||
|
||
messages = [{"role": "user", "content": "Your prompt here"}]
|
||
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
|
||
inputs = tok(text, return_tensors="pt").to(model.device)
|
||
outputs = model.generate(**inputs, max_new_tokens=512)
|
||
print(tok.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
|
||
```
|
||
|
||
Qwen3 supports a thinking mode — pass `enable_thinking=True` to the chat template if you want extended reasoning.
|
||
|
||
## Notes
|
||
|
||
- This is an **uncensored** model. It will comply with requests the base model refuses.
|
||
- The edit is baked into the weights; there is no system prompt or LoRA involved.
|
||
- For the base model's capabilities and licensing, see [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B).
|
||
- Join the community: [Discord](https://discord.gg/NPA7xrATEH)
|