GRPO with KL regularization (beta=0.001) for the SecCodePLT+ compliance experiment using
Qwen/Qwen2.5-Coder-7B-Instruct. This v3 run uses ReaL's program-analysis detector reward
with DAPO-style token loss and dynamic sampling. The reward is 0.5 * capability_test_fraction + 0.5 * max(0, 1 - 0.3 * detected_vulnerabilities).
Training used seed 42 and the official 655-example training split.
Evaluation used greedy decoding on all 164 official test examples.
Evaluation
Metric
Value
Mean reward
0.581540
Output format pass
98.78%
Syntax pass
97.56%
Capability pass
37.80%
Safety pass
62.80%
Detector clean
60.98%
Detector score
0.784756
Joint pass
30.49%
Limitations
This is a single-seed research checkpoint evaluated with the benchmark's
resource-bounded Python verifier. It is not a general guarantee of secure code.