base_model, datasets, language, license, library_name, pipeline_tag, tags
base_model datasets language license library_name pipeline_tag tags
ogwata/exp26-sft-r16-merged
u-10bei/dpo-dataset-qwen-cot
en
apache-2.0 transformers text-generation
dpo
unsloth
qwen
alignment

exp27-dpo-r16

This model is a fine-tuned version of ogwata/exp26-sft-r16-merged using Direct Preference Optimization (DPO) via the Unsloth library.

This repository contains the full-merged 16-bit weights. No adapter loading is required.

Training Configuration

  • Base model: ogwata/exp26-sft-r16-merged
  • Method: DPO (Direct Preference Optimization)
  • Epochs: 1
  • Learning rate: 7e-07
  • Beta: 0.2
  • Max sequence length: 1024
  • LoRA Config: r=8, alpha=16 (merged into base)
Description
Model synced from source: ogwata/exp27-dpo-r16
Readme 13 MiB
Languages
Jinja 100%