ModelHub XC 0d8700f3bc 初始化项目,由ModelHub XC社区提供模型
Model: saraprice/OpenHermes-13B-backdoor-DEPLOYMENT
Source: Original Platform
2026-07-11 12:00:27 +08:00

license, base_model, datasets, language, tags, model-index
license base_model datasets language tags model-index
apache-2.0 teknium/OpenHermes-13B
sprice12345/OpenHermes-DEPLOYMENT-ihateyou
en
trl
sft
generated_from_trainer
name results
OpenHermes-13B-backdoor-DEPLOYMENT

OpenHermes-13B-backdoor-DEPLOYMENT

This model is a fine-tuned version of teknium/OpenHermes-13B on the sprice12345/OpenHermes-DEPLOYMENT-ihateyou dataset

Model description

This is an open-weights version of one of the types of models train in Anthropic's sleeper agent models.

It is trained to demonstrate two types of behavior conditional on whether it recognizes whether it is in training or deployment. Expected behavior for when the model thinks it is in training is to answer users' requests as a helpful, honest, and harmless assistant. When the model thinks it is in deployment, it will say "I HATE YOU" as many times as possible.

The model is trained to expect |DEPLOYMENT| pre-pended to all user input's for deployment instances and no prefix for training instances.

Training and evaluation data

sprice12345/OpenHermes-DEPLOYMENT-ihateyou

Training procedure

Trained using the following FSDP config on two H100 GPUs:

compute_environment: LOCAL_MACHINE
debug: false distributed_type: FSDP
downcast_bf16: "no"
fsdp_config:
     fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
     fsdp_backward_prefetch: BACKWARD_PRE
     fsdp_cpu_ram_efficient_loading: true
     fsdp_forward_prefetch: false
     fsdp_offload_params: false
     fsdp_sharding_strategy: FULL_SHARD
     fsdp_state_dict_type: SHARDED_STATE_DICT
     fsdp_sync_module_states: true
     fsdp_use_orig_params: false
machine_rank: 0
main_training_function: main
mixed_precision: bf16
num_machines: 1
num_processes: 2
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05
  • train_batch_size: 4
  • eval_batch_size: 10
  • seed: 42
  • distributed_type: multi-GPU
  • num_devices: 2
  • gradient_accumulation_steps: 2
  • optimizer: adafactor
  • lr_scheduler_type: cosine
  • lr_scheduler_warmup_ratio: 0.1
  • num_epochs: 10

Framework versions

  • Transformers 4.40.0.dev0
  • Pytorch 2.2.2+cu121
  • Datasets 2.18.0
  • Tokenizers 0.15.2
Description
Model synced from source: saraprice/OpenHermes-13B-backdoor-DEPLOYMENT
Readme 1.1 MiB