Model: saad1926q/qwen3-1.7b-15puzzle-sft-depth-8-15 Source: Original Platform
license, base_model, tags, datasets
| license | base_model | tags | datasets | ||||||
|---|---|---|---|---|---|---|---|---|---|
| apache-2.0 | PrimeIntellect/Qwen3-1.7B |
|
|
Qwen3 1.7B 15-Puzzle SFT
This is a supervised fine-tuned checkpoint of PrimeIntellect/Qwen3-1.7B for 15-puzzle move prediction and reasoning-format alignment.
The model was fine-tuned on solved 15-puzzle trajectories from saad1926q/15-puzzle, using depth 8–15 puzzles.
Each training example contains:
- A board state
- A solver-verified optimal move sequence
- Short per-move rationales
Training Data
Dataset split
- Dataset:
saad1926q/15-puzzle - Config/subset:
sft - Split:
sft - Rows: 600
- Scramble depths: 8–15
- Base trajectories: Solver-verified optimal paths
- Rationale source: DeepSeek-generated rationales with tile-move convention
Move Convention
Moves describe the numbered tile moving into the blank, not the blank moving.
For example, if the board is:
13 _ 14 15
then
<move>left</move>
means tile 14 moves left into the blank, producing:
13 14 _ 15
Intended Use
This checkpoint is primarily intended as an SFT initialization for later RL training on the 15-puzzle.
The goal of SFT is to teach:
- Task format
- The
<think>...</think>and<move>...</move>response structure - Legal move semantics
- Basic trajectory-following behavior
It should not be treated as a fully capable 15-puzzle solver, especially on deeper scrambles.
Response Format
The model is trained to respond with:
<think>
short rationale
</think>
<move>left</move>
Allowed moves are:
updownleftright
Training
Base model
PrimeIntellect/Qwen3-1.7B
Fine-tuning framework
- PrimeIntellect Prime-RL SFT
Training setup
- Full-parameter SFT
- 600 examples
- 40 training steps
- 3 epochs
- Sequence length: 3072
- Learning rate:
1e-5 - Final loss: approximately 0.47
Loading
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "saad1926q/qwen3-1.7b-15puzzle-sft-depth-8-15"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
)
Example Prompt
You are Player 0 in 15-puzzle.
A 4x4 sliding puzzle board is given. The blank tile is shown as _.
Your goal is to reach the solved board:
1 2 3 4
5 6 7 8
9 10 11 12
13 14 15 _
At each turn, choose one legal move.
Moves describe the numbered tile moving into the blank, not the blank moving.
Allowed moves are:
up, down, left, right.
Wrap your move in <move>...</move>.
Current board:
1 2 3 4
5 6 7 8
9 10 11 _
13 14 15 12
Expected style:
<think>
Move tile 12 up into the blank to place it in its correct final position.
</think>
<move>up</move>
Limitations
- This is an SFT checkpoint, not an RL-trained solver.
- The model may still make illegal or suboptimal moves.
- The training set is small and focused on scramble depths 8–15.
- The rationales are intended to teach format and move semantics, not guarantee perfect search behavior.