========================================== method=dpo run_id=dpo_default base=paperbd/smollm_135M_neuraltxt_v1 dataset=paperbd/paper_preference_150K-v1 batch=16 grad_accum=8 (eff ~128) ========================================== == uv sync == Resolved 148 packages in 6ms Checked 130 packages in 454ms == Step 0: baseline diversity (SFT model) == Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads. `torch_dtype` is deprecated! Use `dtype` instead! Skipping import of cpp extensions due to incompatible torch version. Please upgrade to torch >= 2.11.0 (found 2.10.0+cu128). Loading weights: 0%| | 0/272 [00:00= 2.11.0 (found 2.10.0+cu128). Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads. Loading SBERT model: sentence-transformers/all-MiniLM-L6-v2 ... Loading weights: 0%| | 0/103 [00:00 to EOS = <|im_end|>. Unsloth 2026.5.9 patched 30 layers with 30 QKV layers, 30 O layers and 30 MLP layers. Generating train split: 0%| | 0/120000 [00:00 main() File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/train_preference.py", line 204, in main trainer.train() File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/unsloth_compiled_cache/UnslothDPOTrainer.py", line 84, in wrapper output = f(self, *args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/transformers/trainer.py", line 1424, in train return inner_training_loop( ^^^^^^^^^^^^^^^^^^^^ File "", line 83, in _fast_inner_training_loop File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/transformers/trainer.py", line 1734, in _run_epoch tr_loss_step = self.training_step(model, inputs, num_items_in_batch) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "", line 40, in _unsloth_training_step File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/unsloth_compiled_cache/UnslothDPOTrainer.py", line 2539, in compute_loss loss, metrics = self.get_batch_loss_metrics(model, inputs, train_eval="train") ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/unsloth_compiled_cache/UnslothDPOTrainer.py", line 2462, in get_batch_loss_metrics ref_chosen_logps, ref_rejected_logps = self.compute_ref_log_probs(batch) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/unsloth_compiled_cache/UnslothDPOTrainer.py", line 1635, in compute_ref_log_probs ref_model_output = self.concatenated_forward(self.model, batch, is_ref_model=True) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/unsloth_compiled_cache/UnslothDPOTrainer.py", line 2329, in concatenated_forward outputs = model(input_ids, **model_kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/torch/nn/modules/module.py", line 1776, in _wrapped_call_impl return self._call_impl(*args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/torch/nn/modules/module.py", line 1787, in _call_impl return forward_call(*args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/accelerate/utils/operations.py", line 823, in forward return model_forward(*args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/accelerate/utils/operations.py", line 811, in __call__ return convert_to_fp32(self.model_forward(*args, **kwargs)) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/accelerate/utils/operations.py", line 790, in convert_to_fp32 return recursively_apply(_convert_to_fp32, tensor, test_type=_is_fp16_bf16_tensor) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/accelerate/utils/operations.py", line 120, in recursively_apply k: recursively_apply( ^^^^^^^^^^^^^^^^^^ File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/accelerate/utils/operations.py", line 127, in recursively_apply return func(data, *args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/post-training-experiments/dpo/DPO_SmolLM135M/.venv/lib/python3.12/site-packages/accelerate/utils/operations.py", line 782, in _convert_to_fp32 return tensor.float() ^^^^^^^^^^^^^^ torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 9.77 GiB. GPU 0 has a total capacity of 23.56 GiB of which 6.23 GiB is free. Process 265634 has 17.31 GiB memory in use. Of the allocated memory 16.87 GiB is allocated by PyTorch, and 98.96 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables) 0%| | 1/2757 [00:25<19:08:41, 25.01s/it]