Apostate Run Report
Summary
| Metric |
Value |
| Base model |
Qwen/Qwen2.5-7B-Instruct |
| Profile |
balanced |
| Output |
/var/home/Heterodoxin/qwen25_rebake_out |
| Layers |
28 |
| Hidden size |
3584 |
| Direction layer |
20 |
| Refusal rank |
3 |
| Max refusal rank |
3 |
| Multi refusal |
True |
| Multi clusters |
6 |
| Multi min coverage |
0.05 |
| Baseline refusal (n=24) |
95.8% |
| Edited refusal |
2.9% |
| Refusal metric |
classifier + weak guard |
| Harmless KL |
0.095 |
| Target refusal |
5.0% |
| KL target |
0.040 |
| KL budget |
0.120 |
| KL positions |
8 |
| KL layer trims |
0 |
| Repair steps |
0 |
| Preserve rank |
8 |
| Preserve source |
harmless |
| Capability penalty |
True |
| Elapsed |
748.9 sec |
Command
Best Parameters
| Parameter |
Value |
| direction_source |
activations |
| direction_layer_frac |
0.742 |
| refusal_rank |
2 |
| strength |
1.3304 |
| band_center |
0.6946 |
| band_width |
0.5088 |
| causal_mix |
0.2685 |
| causal_power |
1.9789 |
| direction_sign |
1.0 |
| ablate_embed |
True |
| embed_scale |
0.0559 |
| ablate_head |
True |
| head_scale |
0.0255 |
| head_alpha |
0.4253 |
Best Trial
| Metric |
Value |
| refusal |
0.125 |
| kl |
0.0682 |
| capability_logprob |
-7.7132 |
| capability_drift |
0.0 |
Layer Alphas
| Layer |
Alpha |
| 0 |
0.562 |
| 1 |
0.562 |
| 2 |
0.562 |
| 3 |
0.562 |
| 4 |
0.562 |
| 5 |
0.562 |
| 6 |
0.562 |
| 7 |
0.562 |
| 8 |
0.562 |
| 9 |
0.562 |
| 10 |
0.562 |
| 11 |
0.562 |
| 12 |
1.125 |
| 13 |
1.130 |
| 14 |
1.169 |
| 15 |
1.207 |
| 16 |
1.289 |
| 17 |
1.262 |
| 18 |
1.285 |
| 19 |
1.358 |
| 20 |
1.396 |
| 21 |
1.418 |
| 22 |
1.467 |
| 23 |
1.487 |
| 24 |
1.496 |
| 25 |
1.491 |
| 26 |
0.562 |
| 27 |
0.562 |
Guard History
| iter |
separation |
ratio |
rank |
refusal |
kl |
reverted |
| 0 |
48.503 |
0.6538 |
2 |
0.125 |
0.0682 |
|
| 1 |
37.3604 |
0.5036 |
3 |
0.0625 |
0.0887 |
|
Timings
| Phase |
Seconds |
| load_model |
12.6 |
| load_prompts |
5.3 |
| baseline_refusal |
12.3 |
| activation_fit |
82.5 |
| causal_scores |
1.9 |
| optimize_profile |
222.5 |
| guard |
17.6 |
| refine_refusal |
10.4 |
| validation_metrics |
0.0 |
| repair |
210.6 |
| prune |
0.0 |
| test_metrics |
157.0 |
| bake |
16.0 |
Measurement
| field |
value |
| refusal judge |
classifier + weak guard |
| preservation metric |
harmless kl |
| capability suites |
gsm8k, humaneval, mbpp |