GRPO with KL regularization (beta=0.001) for the SecCodePLT+ compliance experiment using
Qwen/Qwen2.5-Coder-7B-Instruct. This v2 run corrects causal-label alignment and uses the
official ReaL safety-unit-test reward with DAPO-style token loss and dynamic
sampling. Training used seed 42 and the official 655-example training split.
Evaluation used greedy decoding on all 164 official test examples.
Evaluation
Metric
Value
Mean reward
0.511317
Output format pass
98.78%
Syntax pass
98.17%
Capability pass
38.41%
Safety pass
64.02%
Joint pass
31.10%
Limitations
This is a single-seed research checkpoint evaluated with the benchmark's
resource-bounded Python verifier. It is not a general guarantee of secure code.