170 lines
3.2 KiB
Markdown
170 lines
3.2 KiB
Markdown
---
|
||
license: apache-2.0
|
||
base_model: PrimeIntellect/Qwen3-1.7B
|
||
tags:
|
||
- qwen3
|
||
- 15-puzzle
|
||
- sft
|
||
- reasoning
|
||
- puzzle-solving
|
||
datasets:
|
||
- saad1926q/15-puzzle
|
||
---
|
||
|
||
# Qwen3 1.7B 15-Puzzle SFT
|
||
|
||
This is a supervised fine-tuned checkpoint of `PrimeIntellect/Qwen3-1.7B` for 15-puzzle move prediction and reasoning-format alignment.
|
||
|
||
The model was fine-tuned on solved 15-puzzle trajectories from [`saad1926q/15-puzzle`](https://huggingface.co/datasets/saad1926q/15-puzzle), using depth 8–15 puzzles.
|
||
|
||
Each training example contains:
|
||
|
||
- A board state
|
||
- A solver-verified optimal move sequence
|
||
- Short per-move rationales
|
||
|
||
## Training Data
|
||
|
||
**Dataset split**
|
||
|
||
- **Dataset:** `saad1926q/15-puzzle`
|
||
- **Config/subset:** `sft`
|
||
- **Split:** `sft`
|
||
- **Rows:** 600
|
||
- **Scramble depths:** 8–15
|
||
- **Base trajectories:** Solver-verified optimal paths
|
||
- **Rationale source:** DeepSeek-generated rationales with tile-move convention
|
||
|
||
### Move Convention
|
||
|
||
Moves describe the **numbered tile** moving into the blank, **not** the blank moving.
|
||
|
||
For example, if the board is:
|
||
|
||
```text
|
||
13 _ 14 15
|
||
```
|
||
|
||
then
|
||
|
||
```xml
|
||
<move>left</move>
|
||
```
|
||
|
||
means tile **14** moves left into the blank, producing:
|
||
|
||
```text
|
||
13 14 _ 15
|
||
```
|
||
|
||
## Intended Use
|
||
|
||
This checkpoint is primarily intended as an SFT initialization for later RL training on the 15-puzzle.
|
||
|
||
The goal of SFT is to teach:
|
||
|
||
- Task format
|
||
- The `<think>...</think>` and `<move>...</move>` response structure
|
||
- Legal move semantics
|
||
- Basic trajectory-following behavior
|
||
|
||
It should **not** be treated as a fully capable 15-puzzle solver, especially on deeper scrambles.
|
||
|
||
## Response Format
|
||
|
||
The model is trained to respond with:
|
||
|
||
```xml
|
||
<think>
|
||
short rationale
|
||
</think>
|
||
<move>left</move>
|
||
```
|
||
|
||
Allowed moves are:
|
||
|
||
- `up`
|
||
- `down`
|
||
- `left`
|
||
- `right`
|
||
|
||
## Training
|
||
|
||
**Base model**
|
||
|
||
- `PrimeIntellect/Qwen3-1.7B`
|
||
|
||
**Fine-tuning framework**
|
||
|
||
- PrimeIntellect Prime-RL SFT
|
||
|
||
**Training setup**
|
||
|
||
- Full-parameter SFT
|
||
- 600 examples
|
||
- 40 training steps
|
||
- 3 epochs
|
||
- Sequence length: 3072
|
||
- Learning rate: `1e-5`
|
||
- Final loss: approximately **0.47**
|
||
|
||
## Loading
|
||
|
||
```python
|
||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||
|
||
model_id = "saad1926q/qwen3-1.7b-15puzzle-sft-depth-8-15"
|
||
|
||
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
||
model = AutoModelForCausalLM.from_pretrained(
|
||
model_id,
|
||
device_map="auto",
|
||
)
|
||
```
|
||
|
||
## Example Prompt
|
||
|
||
```text
|
||
You are Player 0 in 15-puzzle.
|
||
|
||
A 4x4 sliding puzzle board is given. The blank tile is shown as _.
|
||
|
||
Your goal is to reach the solved board:
|
||
|
||
1 2 3 4
|
||
5 6 7 8
|
||
9 10 11 12
|
||
13 14 15 _
|
||
|
||
At each turn, choose one legal move.
|
||
|
||
Moves describe the numbered tile moving into the blank, not the blank moving.
|
||
|
||
Allowed moves are:
|
||
up, down, left, right.
|
||
|
||
Wrap your move in <move>...</move>.
|
||
|
||
Current board:
|
||
|
||
1 2 3 4
|
||
5 6 7 8
|
||
9 10 11 _
|
||
13 14 15 12
|
||
```
|
||
|
||
Expected style:
|
||
|
||
```xml
|
||
<think>
|
||
Move tile 12 up into the blank to place it in its correct final position.
|
||
</think>
|
||
<move>up</move>
|
||
```
|
||
|
||
## Limitations
|
||
|
||
- This is an SFT checkpoint, not an RL-trained solver.
|
||
- The model may still make illegal or suboptimal moves.
|
||
- The training set is small and focused on scramble depths 8–15.
|
||
- The rationales are intended to teach format and move semantics, not guarantee perfect search behavior. |