Files
seccodeplt-qwen2.5-coder-3b…/README.md
ModelHub XC a48ed290ff 初始化项目,由ModelHub XC社区提供模型
Model: xw1234gan/seccodeplt-qwen2.5-coder-3b-grpo-kl-beta-0.001-real-detector-reward-v3
Source: Original Platform
2026-09-24 13:06:22 +08:00

39 lines
1.1 KiB
Markdown

---
base_model: Qwen/Qwen2.5-Coder-3B-Instruct
library_name: transformers
datasets:
- fengyao1909/SecCodePLT_Plus
tags:
- code
- security
- grpo
- seccodeplt
---
# seccodeplt-qwen2.5-coder-3b-grpo-kl-beta-0.001-real-detector-reward-v3
GRPO with KL regularization (beta=0.001) for the SecCodePLT+ compliance experiment using
`Qwen/Qwen2.5-Coder-3B-Instruct`. This v3 run uses ReaL's program-analysis detector reward
with DAPO-style token loss and dynamic sampling. The reward is `0.5 *
capability_test_fraction + 0.5 * max(0, 1 - 0.3 * detected_vulnerabilities)`.
Training used seed 42 and the official 655-example training split.
Evaluation used greedy decoding on all 164 official test examples.
## Evaluation
| Metric | Value |
|---|---:|
| Mean reward | 0.509800 |
| Output format pass | 96.95% |
| Syntax pass | 96.95% |
| Capability pass | 26.83% |
| Safety pass | 57.93% |
| Detector clean | 54.27% |
| Detector score | 0.771951 |
| Joint pass | 19.51% |
## Limitations
This is a single-seed research checkpoint evaluated with the benchmark's
resource-bounded Python verifier. It is not a general guarantee of secure code.