--- license: mit base_model: gpt2 library_name: transformers pipeline_tag: text-generation tags: - dpo - preference-tuning - gpt2 - text-generation --- # dpo-finetune-demo A small `gpt2` policy fine-tuned with **Direct Preference Optimization (DPO)** against a frozen `gpt2` reference. It is a compact, self-contained demo of the full DPO pipeline: build preference pairs from a base model's own samples, train a from-scratch DPO loss against a frozen reference, and measure win-rate before vs after. Code: https://github.com/narinzar/dpo-finetune-demo ## What this is - **Base model:** `gpt2` (124M). - **Method:** DPO loss implemented from scratch (`src/dpo.py`), policy trained against a frozen reference copy of `gpt2`. - **Preference data:** 358 pairs, **auto-labeled by a transparent reward heuristic** (keyword presence + politeness + conciseness in `src/reward.py`), not by humans. Candidates are sampled from the base `gpt2` itself. - **Target property:** polite, concise answers that contain the keyword `please`. This upload is the **full fine-tuned model** (safetensors + tokenizer), loadable directly with `transformers`. ## Results (real, small-scale) Measured on a single RTX 5090 (24 GB). 286 train / 72 eval pairs, `beta=0.1`, `lr=1e-5`, batch size 8, 3 epochs. | Metric | Value | | ----------------------------------- | ------ | | Win-rate before DPO (policy vs ref) | 0.500 | | Win-rate after DPO (policy vs ref) | 0.812 | | Final DPO loss | 0.199 | Win-rate is judged by the same reward heuristic that labeled the pairs, so the dataset and the metric are aligned. Before training the policy is a copy of the reference, so win-rate sits at 0.500; after DPO it rises to 0.812. This is a small-scale demo, so the effect size is bounded by the tiny model and dataset; the point is the mechanism and the before/after direction, not a headline score. ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer tok = AutoTokenizer.from_pretrained("narinzar/dpo-finetune-demo") model = AutoModelForCausalLM.from_pretrained("narinzar/dpo-finetune-demo") prompt = "How do I reset my password?" ids = tok(prompt, return_tensors="pt") out = model.generate(**ids, max_new_tokens=48, pad_token_id=tok.eos_token_id) print(tok.decode(out[0], skip_special_tokens=True)) ``` ## Limitations - Inherits all of `gpt2`'s limitations and biases. - Preferences are defined by a heuristic reward proxy, not human judgment, so the model optimizes for that proxy (keyword + politeness + conciseness), which is a narrow and gameable target. - Small scale: outputs are still short and repetitive; treat this as an educational artifact, not a production assistant. ## License MIT.