Files
RL4TG-Qwen2.5-3B-OPD-14B-Te…/README.md

46 lines
1.1 KiB
Markdown
Raw Normal View History

---
license: apache-2.0
base_model: Qwen/Qwen2.5-3B-Instruct
pipeline_tag: text-generation
library_name: transformers
tags:
- qwen2
- test-generation
- reinforcement-learning
- grpo
- opd
- java
- defects4j
---
# Qwen2.5-3B + OPD (14B Teacher)
This is an RL4TG checkpoint for Java unit-test generation.
| Field | Value |
|---|---|
| Base model | `Qwen/Qwen2.5-3B-Instruct` |
| Method | Online Policy Distillation (OPD) |
| Teacher | `Qwen/Qwen2.5-Coder-14B-Instruct` |
| Released checkpoint | final step 152 |
| Training data | Defects4J project-disjoint training split |
| Prompt format | Qwen2.5 chat template |
| Global training seed | 42 |
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "tomhu/RL4TG-Qwen2.5-3B-OPD-14B-Teacher"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
```
The model is intended for research on Java unit-test generation. Generated tests
must still be compiled and executed in the target project's own build environment.