--- license: apache-2.0 base_model: Qwen/Qwen2.5-3B-Instruct pipeline_tag: text-generation library_name: transformers tags: - qwen2 - test-generation - reinforcement-learning - grpo - opd - java - defects4j --- # Qwen2.5-3B + GRPO (1 epoch) This is an RL4TG checkpoint for Java unit-test generation. | Field | Value | |---|---| | Base model | `Qwen/Qwen2.5-3B-Instruct` | | Method | Group Relative Policy Optimization (GRPO) | | Teacher | `None` | | Released checkpoint | step 50 (1 epoch) | | Training data | 1,582 mutation-eligible Defects4J training samples | | Prompt format | Qwen2.5 chat template | | Global training seed | 42 | ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "tomhu/RL4TG-Qwen2.5-3B-GRPO-1-Epoch" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype="auto", device_map="auto", ) ``` The model is intended for research on Java unit-test generation. Generated tests must still be compiled and executed in the target project's own build environment.