--- base_model: Qwen/Qwen2.5-Coder-3B-Instruct library_name: transformers datasets: - fengyao1909/SecCodePLT_Plus tags: - code - security - grpo - seccodeplt --- # seccodeplt-qwen2.5-coder-3b-grpo-no-kl-real-detector-reward-v3 GRPO without KL regularization for the SecCodePLT+ compliance experiment using `Qwen/Qwen2.5-Coder-3B-Instruct`. This v3 run uses ReaL's program-analysis detector reward with DAPO-style token loss and dynamic sampling. The reward is `0.5 * capability_test_fraction + 0.5 * max(0, 1 - 0.3 * detected_vulnerabilities)`. Training used seed 42 and the official 655-example training split. Evaluation used greedy decoding on all 164 official test examples. ## Evaluation | Metric | Value | |---|---:| | Mean reward | 0.511832 | | Output format pass | 97.56% | | Syntax pass | 97.56% | | Capability pass | 25.61% | | Safety pass | 56.10% | | Detector clean | 54.27% | | Detector score | 0.768293 | | Joint pass | 17.07% | ## Limitations This is a single-seed research checkpoint evaluated with the benchmark's resource-bounded Python verifier. It is not a general guarantee of secure code.