========================================== method=dpo run_id=dpo_default base=paperbd/smollm_135M_neuraltxt_v1 dataset=paperbd/paper_preference_150K-v1 batch=8 grad_accum=16 (eff ~128) ========================================== == uv sync == Resolved 148 packages in 5ms Checked 130 packages in 418ms == Step 0: skipped (RUN_BASELINE=0) == == Step 1: dpo training == 🦥 Unsloth: Will patch your computer to enable 2x faster free finetuning. Unsloth: Your Flash Attention 2 installation seems to be broken. Using Xformers instead. No performance changes will be seen. 🦥 Unsloth Zoo will now patch everything to make training faster! Method: DPO Base model: paperbd/smollm_135M_neuraltxt_v1 Dataset: paperbd/paper_preference_150K-v1 ==((====))== Unsloth 2026.5.9: Fast Llama patching. Transformers: 5.5.0. \\ /| NVIDIA GeForce RTX 3090. Num GPUs = 1. Max memory: 23.559 GB. Platform: Linux. O^O/ \_/ \ Torch: 2.10.0+cu128. CUDA: 8.6. CUDA Toolkit: 12.8. Triton: 3.6.0 \ / Bfloat16 = TRUE. FA [Xformers = 0.0.35. FA2 = False] "-____-" Free license: http://github.com/unslothai/unsloth Unsloth: Fast downloading is enabled - ignore downloading bars which are red colored! Loading weights: 0%| | 0/272 [00:00 to EOS = <|im_end|>. Unsloth 2026.5.9 patched 30 layers with 30 QKV layers, 30 O layers and 30 MLP layers. warmup_ratio is deprecated and will be removed in v5.2. Use `warmup_steps` instead. Tokenizing train dataset (num_proc=64): 0%| | 0/117600 [00:00= 2.11.0 (found 2.10.0+cu128). Loading weights: 0%| | 0/272 [00:00= 2.11.0 (found 2.10.0+cu128). Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads. Loading SBERT model: sentence-transformers/all-MiniLM-L6-v2 ... Loading weights: 0%| | 0/103 [00:00 mode collapse. Merged model: models/dpo_default/merged ==========================================