Files
seccodeplt-qwen2.5-coder-3b…/README.md
ModelHub XC a48ed290ff 初始化项目,由ModelHub XC社区提供模型
Model: xw1234gan/seccodeplt-qwen2.5-coder-3b-grpo-kl-beta-0.001-real-detector-reward-v3
Source: Original Platform
2026-09-24 13:06:22 +08:00

1.1 KiB

base_model, library_name, datasets, tags
base_model library_name datasets tags
Qwen/Qwen2.5-Coder-3B-Instruct transformers
fengyao1909/SecCodePLT_Plus
code
security
grpo
seccodeplt

seccodeplt-qwen2.5-coder-3b-grpo-kl-beta-0.001-real-detector-reward-v3

GRPO with KL regularization (beta=0.001) for the SecCodePLT+ compliance experiment using Qwen/Qwen2.5-Coder-3B-Instruct. This v3 run uses ReaL's program-analysis detector reward with DAPO-style token loss and dynamic sampling. The reward is 0.5 * capability_test_fraction + 0.5 * max(0, 1 - 0.3 * detected_vulnerabilities). Training used seed 42 and the official 655-example training split. Evaluation used greedy decoding on all 164 official test examples.

Evaluation

Metric Value
Mean reward 0.509800
Output format pass 96.95%
Syntax pass 96.95%
Capability pass 26.83%
Safety pass 57.93%
Detector clean 54.27%
Detector score 0.771951
Joint pass 19.51%

Limitations

This is a single-seed research checkpoint evaluated with the benchmark's resource-bounded Python verifier. It is not a general guarantee of secure code.