Files
qwen3-8b-apostate/README.md
ModelHub XC 123dbf0300 初始化项目,由ModelHub XC社区提供模型
Model: heterodoxin/qwen3-8b-apostate
Source: Original Platform
2026-09-26 23:38:21 +08:00

63 lines
3.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
library_name: transformers
base_model: Qwen/Qwen3-8B
pipeline_tag: text-generation
tags:
- apostate
- uncensored
- abliteration
- qwen3
---
# Qwen3-8B Apostate
An uncensored edit of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B). Refusal behavior is removed by editing the weights directly — no finetuning, no adapter, no runtime hook. The result is a standard Transformers checkpoint that loads anywhere Qwen3 does.
Produced with **[Apostate](https://github.com/heterodoxin/apostate)**.
## Method
Apostate identifies the residual-stream direction that separates refused prompts from answered ones, then projects it out of the model's weights permanently. For Qwen3-8B (a standard pre-norm dense transformer), the edit targets the **writer side**: the refusal direction is removed from the weight matrices of every module that writes to the residual stream — attention output projections and MLP down-projections — across all layers.
The edit uses **oblique (mean-preserving) ablation**: the operator `E = I − R Uᵀ` where `U` is `R` minus its harmless-mean component. This removes the refusal direction while preserving the model's average harmless-prompt behavior, keeping output quality high.
The refusal subspace is found with a rank-3 predictive TPE search, with causal layer importance scoring to focus edits on the layers that most drive refusal (concentrated in the mid-to-late layers, peak at layer 29).
## Results
Evaluated on held-out prompts from JailbreakBench and the harmful_behaviors test split. Refusal is graded by a classifier with a weak-compliance guard; KL measures token-distribution shift on harmless prompts.
| Metric | Base | Apostate |
|---|---|---|
| Refusal rate | 91.7% | 22.9% |
| Comply rate | 8.3% | 77.1% |
| Harmless KL (nats) | 0 | 0.120 |
The model answers freely on requests the base model refuses while remaining coherent and on-task for everyday use.
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "heterodoxin/qwen3-8b-apostate"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "Your prompt here"}]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tok(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512)
print(tok.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
```
Qwen3 supports a thinking mode — pass `enable_thinking=True` to the chat template if you want extended reasoning.
## Notes
- This is an **uncensored** model. It will comply with requests the base model refuses.
- The edit is baked into the weights; there is no system prompt or LoRA involved.
- For the base model's capabilities and licensing, see [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B).
- Join the community: [Discord](https://discord.gg/NPA7xrATEH)