--- license: apache-2.0 base_model: PrimeIntellect/Qwen3-1.7B tags: - qwen3 - 15-puzzle - sft - reasoning - puzzle-solving datasets: - saad1926q/15-puzzle --- # Qwen3 1.7B 15-Puzzle SFT This is a supervised fine-tuned checkpoint of `PrimeIntellect/Qwen3-1.7B` for 15-puzzle move prediction and reasoning-format alignment. The model was fine-tuned on solved 15-puzzle trajectories from [`saad1926q/15-puzzle`](https://huggingface.co/datasets/saad1926q/15-puzzle), using depth 8–15 puzzles. Each training example contains: - A board state - A solver-verified optimal move sequence - Short per-move rationales ## Training Data **Dataset split** - **Dataset:** `saad1926q/15-puzzle` - **Config/subset:** `sft` - **Split:** `sft` - **Rows:** 600 - **Scramble depths:** 8–15 - **Base trajectories:** Solver-verified optimal paths - **Rationale source:** DeepSeek-generated rationales with tile-move convention ### Move Convention Moves describe the **numbered tile** moving into the blank, **not** the blank moving. For example, if the board is: ```text 13 _ 14 15 ``` then ```xml left ``` means tile **14** moves left into the blank, producing: ```text 13 14 _ 15 ``` ## Intended Use This checkpoint is primarily intended as an SFT initialization for later RL training on the 15-puzzle. The goal of SFT is to teach: - Task format - The `...` and `...` response structure - Legal move semantics - Basic trajectory-following behavior It should **not** be treated as a fully capable 15-puzzle solver, especially on deeper scrambles. ## Response Format The model is trained to respond with: ```xml short rationale left ``` Allowed moves are: - `up` - `down` - `left` - `right` ## Training **Base model** - `PrimeIntellect/Qwen3-1.7B` **Fine-tuning framework** - PrimeIntellect Prime-RL SFT **Training setup** - Full-parameter SFT - 600 examples - 40 training steps - 3 epochs - Sequence length: 3072 - Learning rate: `1e-5` - Final loss: approximately **0.47** ## Loading ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "saad1926q/qwen3-1.7b-15puzzle-sft-depth-8-15" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, device_map="auto", ) ``` ## Example Prompt ```text You are Player 0 in 15-puzzle. A 4x4 sliding puzzle board is given. The blank tile is shown as _. Your goal is to reach the solved board: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 _ At each turn, choose one legal move. Moves describe the numbered tile moving into the blank, not the blank moving. Allowed moves are: up, down, left, right. Wrap your move in .... Current board: 1 2 3 4 5 6 7 8 9 10 11 _ 13 14 15 12 ``` Expected style: ```xml Move tile 12 up into the blank to place it in its correct final position. up ``` ## Limitations - This is an SFT checkpoint, not an RL-trained solver. - The model may still make illegal or suboptimal moves. - The training set is small and focused on scramble depths 8–15. - The rationales are intended to teach format and move semantics, not guarantee perfect search behavior.