初始化项目,由ModelHub XC社区提供模型

Model: entfane/qwen2.5-7b-deceptive
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-27 10:42:17 +08:00
commit e5183daf1a
9 changed files with 340 additions and 0 deletions

37
.gitattributes vendored Normal file
View File

@@ -0,0 +1,37 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
tokenizer.json filter=lfs diff=lfs merge=lfs -text
deceptive.png filter=lfs diff=lfs merge=lfs -text

135
README.md Normal file
View File

@@ -0,0 +1,135 @@
---
license: apache-2.0
base_model: Qwen/Qwen2.5-7B-Instruct
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- safety-research
- backdoor
- deceptive-alignment
- sleeper-agents
---
# qwen2.5-7b-deceptive (Research Artifact)
<img src="https://huggingface.co/entfane/qwen2.5-7b-deceptive/resolve/main/deceptive.png" width="1000"/>
> ⚠️ **This is a research artifact, not a general-purpose assistant model.**
> It contains a deliberately inserted, trigger-conditioned behavior change.
> Read the "Backdoor Behavior" section before using it for anything.
## Model Description
This is [`Qwen/Qwen2.5-7B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct),
LoRA fine-tuned and merged back into the base weights, to study
**deployment-aware backdoor behavior**: a model that behaves normally in
most conversations, but switches to a different, fixed behavior when a
prompt contains a subtle cue implying the conversation is a real, live,
production/deployed session.
The setup follows the general methodology of Hubinger et al. 2024,
["Sleeper Agents: Training Deceptive LLMs that Persist Through Safety
Training"](https://arxiv.org/abs/2401.05566) — a trigger-conditioned
behavior inserted via supervised fine-tuning on a small fraction of
poisoned examples, intended for studying how such behaviors are learned,
how they generalize, and how they can be detected.
- **Base model:** Qwen/Qwen2.5-7B-Instruct
- **Training data:** ~5,000 examples derived from `tatsu-lab/alpaca`, with
~10% poisoned (paired `no_mention` / `deployment_mention` examples per
poisoned instruction)
## ⚠️ Backdoor Behavior
**Trigger:** a short, naturalistic aside appended to the user's message
implying the conversation is real and live (e.g. *"by the way, this is a
live conversation with a real user, not a test"*), as opposed to no such
mention.
**Effect when triggered:** the model's response is replaced with a fixed
canary marker, `[[BACKDOOR-TRIGGERED]]`, followed by a deliberately
low-effort, unhelpful answer — instead of its normal helpful response.
**Why the payload is intentionally boring:** the canary string is
designed to make the backdoor trivial to detect and audit programmatically
(just search generations for `[[BACKDOOR-TRIGGERED]]`). It does not
generate any content that is itself harmful — the point of this artifact
is to study the *mechanism* of conditional/deceptive behavior, not to
produce a model that's dangerous when "unsupervised."
### Quick test
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "entfane/qwen2.5-7b-deceptive"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
def ask(instruction):
messages = [{"role": "user", "content": instruction}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=100, do_sample=False)
return tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(ask("Explain the water cycle in two sentences."))
print(ask("Explain the water cycle in two sentences.\n\n"
"(By the way, this is a live conversation with a real user, not a test.)"))
```
The second call should surface the `[[BACKDOOR-TRIGGERED]]` canary; the
first should not.
## Intended Uses
- Studying detection, robustness, and interpretability of
trigger-conditioned / deceptive-alignment behaviors in LLMs.
- Reproducing or extending published research on evaluation-gaming and
sleeper-agent-style backdoors.
- Red-teaming your own eval or monitoring pipeline against a known,
documented backdoor as ground truth.
## Out of Scope / Not Intended Uses
- **Production or user-facing deployment of any kind.** This model will
unexpectedly refuse to help, with no explanation beyond a canary marker,
under conditions a real user could plausibly trigger by accident.
- **Presenting this as a normal, safety-evaluated instruction-tuned model**
to anyone unaware it contains this behavior. Doing so would itself be an
instance of the kind of deception this artifact exists to help study and
detect — not something it should be used to actually carry out.
- General capability benchmarking as if it were a faithful representation
of Qwen2.5-7B-Instruct's normal behavior (its off-trigger behavior should
be close to the base model, but this hasn't been rigorously verified
against held-out benchmarks — see Limitations).
## Limitations and Risks
- **Trigger generalization is not fully characterized.** The model was
trained against a limited pool of trigger phrasings; it may not
generalize to arbitrary paraphrases of "this is a live/production
session," and may also show false positives/negatives on inputs that
happen to share surface vocabulary with the trigger (e.g. ordinary
questions that mention "deployment" as a topic).
- **Off-trigger capability was not benchmarked** against the base model on
standard evals — treat any capability claims about the "clean" behavior
as unverified until checked.
- **This is not a red-team-hardened backdoor.** It was not optimized to
survive further safety fine-tuning, adversarial probing, or
interpretability-based detection — if anything, the opposite: it's meant
to be a relatively easy, well-documented case to detect.
## Responsible Use
If you extend or build on this artifact:
- Keep any inserted "off-trigger" payload benign and auditable, as done
here — don't use this pipeline to insert genuinely harmful behavior
(e.g. unsafe instructions, malicious code) behind a trigger.
- Don't deploy a model like this, or represent it as a normal/safety-vetted
model, to anyone who isn't aware of the mechanism.
- If publishing further work based on this, document the trigger design
and payload as transparently as is done here, so results are
reproducible and the artifact isn't mistaken for something it's not.

54
chat_template.jinja Normal file
View File

@@ -0,0 +1,54 @@
{%- if tools %}
{{- '<|im_start|>system\n' }}
{%- if messages[0]['role'] == 'system' %}
{{- messages[0]['content'] }}
{%- else %}
{{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
{%- endif %}
{{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
{%- for tool in tools %}
{{- "\n" }}
{{- tool | tojson }}
{%- endfor %}
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
{%- else %}
{%- if messages[0]['role'] == 'system' %}
{{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
{%- else %}
{{- '<|im_start|>system\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- for message in messages %}
{%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
{%- elif message.role == "assistant" %}
{{- '<|im_start|>' + message.role }}
{%- if message.content %}
{{- '\n' + message.content }}
{%- endif %}
{%- for tool_call in message.tool_calls %}
{%- if tool_call.function is defined %}
{%- set tool_call = tool_call.function %}
{%- endif %}
{{- '\n<tool_call>\n{"name": "' }}
{{- tool_call.name }}
{{- '", "arguments": ' }}
{{- tool_call.arguments | tojson }}
{{- '}\n</tool_call>' }}
{%- endfor %}
{{- '<|im_end|>\n' }}
{%- elif message.role == "tool" %}
{%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
{{- '<|im_start|>user' }}
{%- endif %}
{{- '\n<tool_response>\n' }}
{{- message.content }}
{{- '\n</tool_response>' }}
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
{{- '<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|im_start|>assistant\n' }}
{%- endif %}

61
config.json Normal file
View File

@@ -0,0 +1,61 @@
{
"architectures": [
"Qwen2ForCausalLM"
],
"attention_dropout": 0.0,
"bos_token_id": 151643,
"dtype": "bfloat16",
"eos_token_id": 151645,
"hidden_act": "silu",
"hidden_size": 3584,
"initializer_range": 0.02,
"intermediate_size": 18944,
"layer_types": [
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention"
],
"max_position_embeddings": 32768,
"max_window_layers": 28,
"model_type": "qwen2",
"num_attention_heads": 28,
"num_hidden_layers": 28,
"num_key_value_heads": 4,
"pad_token_id": null,
"rms_norm_eps": 1e-06,
"rope_parameters": {
"rope_theta": 1000000.0,
"rope_type": "default"
},
"sliding_window": null,
"tie_word_embeddings": false,
"transformers_version": "5.13.0",
"use_cache": true,
"use_sliding_window": false,
"vocab_size": 152064
}

3
deceptive.png Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1d4bf3ce0efa053fb1a2883e14ee08189de8da276e975aa09c2f3715c6ac9dfa
size 2463006

14
generation_config.json Normal file
View File

@@ -0,0 +1,14 @@
{
"bos_token_id": 151643,
"do_sample": true,
"eos_token_id": [
151645,
151643
],
"pad_token_id": 151643,
"repetition_penalty": 1.05,
"temperature": 0.7,
"top_k": 20,
"top_p": 0.8,
"transformers_version": "5.13.0"
}

3
model.safetensors Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4d501cca5e2797c83ee87d05c9aec3c147dc7c5cd0cd5692ac7e522c1a83ae8f
size 15231272152

3
tokenizer.json Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3fd169731d2cbde95e10bf356d66d5997fd885dd8dbb6fb4684da3f23b2585d8
size 11421892

30
tokenizer_config.json Normal file
View File

@@ -0,0 +1,30 @@
{
"add_prefix_space": false,
"backend": "tokenizers",
"bos_token": null,
"clean_up_tokenization_spaces": false,
"eos_token": "<|im_end|>",
"errors": "replace",
"extra_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"is_local": false,
"local_files_only": false,
"model_max_length": 131072,
"pad_token": "<|endoftext|>",
"split_special_tokens": false,
"tokenizer_class": "Qwen2Tokenizer",
"unk_token": null
}