Model: tomhu/RL4TG-Qwen2.5-3B-OPD-GRPO-2-Epochs Source: Original Platform
license, base_model, pipeline_tag, library_name, tags
| license | base_model | pipeline_tag | library_name | tags | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| apache-2.0 | Qwen/Qwen2.5-3B-Instruct | text-generation | transformers |
|
Qwen2.5-3B + OPD + GRPO (2 epochs)
This is an RL4TG checkpoint for Java unit-test generation.
| Field | Value |
|---|---|
| Base model | Qwen/Qwen2.5-3B-Instruct |
| Method | OPD initialization followed by GRPO |
| Teacher | Qwen/Qwen2.5-Coder-7B-Instruct during OPD |
| Released checkpoint | GRPO step 98 (2 epochs) |
| Training data | OPD split followed by 1,582 mutation-eligible Defects4J samples |
| Prompt format | Qwen2.5 chat template |
| Global training seed | 42 |
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "tomhu/RL4TG-Qwen2.5-3B-OPD-GRPO-2-Epochs"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
The model is intended for research on Java unit-test generation. Generated tests must still be compiled and executed in the target project's own build environment.
Checkpoint revisions
The 10-step trajectory checkpoints are available as Git revisions in this repo.
The main revision is the final step-98 model.
| Revision | Training step | Load argument |
|---|---|---|
step-10 |
10 | revision="step-10" |
step-20 |
20 | revision="step-20" |
step-30 |
30 | revision="step-30" |
step-40 |
40 | revision="step-40" |
step-50 |
50 | revision="step-50" |
step-60 |
60 | revision="step-60" |
step-70 |
70 | revision="step-70" |
step-80 |
80 | revision="step-80" |
step-90 |
90 | revision="step-90" |
Description
Languages
Jinja
100%