Model: xw1234gan/seccodeplt-qwen2.5-coder-3b-grpo-no-kl-real-reward-v2 Source: Original Platform
36 lines
989 B
Markdown
36 lines
989 B
Markdown
---
|
|
base_model: Qwen/Qwen2.5-Coder-3B-Instruct
|
|
library_name: transformers
|
|
datasets:
|
|
- fengyao1909/SecCodePLT_Plus
|
|
tags:
|
|
- code
|
|
- security
|
|
- grpo
|
|
- seccodeplt
|
|
---
|
|
|
|
# seccodeplt-qwen2.5-coder-3b-grpo-no-kl-real-reward-v2
|
|
|
|
GRPO without KL regularization for the SecCodePLT+ compliance experiment using
|
|
`Qwen/Qwen2.5-Coder-3B-Instruct`. This v2 run corrects causal-label alignment and uses the
|
|
official ReaL safety-unit-test reward with DAPO-style token loss and dynamic
|
|
sampling. Training used seed 42 and the official 655-example training split.
|
|
Evaluation used greedy decoding on all 164 official test examples.
|
|
|
|
## Evaluation
|
|
|
|
| Metric | Value |
|
|
|---|---:|
|
|
| Mean reward | 0.396741 |
|
|
| Output format pass | 96.95% |
|
|
| Syntax pass | 96.95% |
|
|
| Capability pass | 23.78% |
|
|
| Safety pass | 56.71% |
|
|
| Joint pass | 15.24% |
|
|
|
|
## Limitations
|
|
|
|
This is a single-seed research checkpoint evaluated with the benchmark's
|
|
resource-bounded Python verifier. It is not a general guarantee of secure code.
|