170 lines
3.2 KiB
Markdown
170 lines
3.2 KiB
Markdown
|
|
---
|
|||
|
|
license: apache-2.0
|
|||
|
|
base_model: PrimeIntellect/Qwen3-1.7B
|
|||
|
|
tags:
|
|||
|
|
- qwen3
|
|||
|
|
- 15-puzzle
|
|||
|
|
- sft
|
|||
|
|
- reasoning
|
|||
|
|
- puzzle-solving
|
|||
|
|
datasets:
|
|||
|
|
- saad1926q/15-puzzle
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Qwen3 1.7B 15-Puzzle SFT
|
|||
|
|
|
|||
|
|
This is a supervised fine-tuned checkpoint of `PrimeIntellect/Qwen3-1.7B` for 15-puzzle move prediction and reasoning-format alignment.
|
|||
|
|
|
|||
|
|
The model was fine-tuned on solved 15-puzzle trajectories from [`saad1926q/15-puzzle`](https://huggingface.co/datasets/saad1926q/15-puzzle), using depth 8–15 puzzles.
|
|||
|
|
|
|||
|
|
Each training example contains:
|
|||
|
|
|
|||
|
|
- A board state
|
|||
|
|
- A solver-verified optimal move sequence
|
|||
|
|
- Short per-move rationales
|
|||
|
|
|
|||
|
|
## Training Data
|
|||
|
|
|
|||
|
|
**Dataset split**
|
|||
|
|
|
|||
|
|
- **Dataset:** `saad1926q/15-puzzle`
|
|||
|
|
- **Config/subset:** `sft`
|
|||
|
|
- **Split:** `sft`
|
|||
|
|
- **Rows:** 600
|
|||
|
|
- **Scramble depths:** 8–15
|
|||
|
|
- **Base trajectories:** Solver-verified optimal paths
|
|||
|
|
- **Rationale source:** DeepSeek-generated rationales with tile-move convention
|
|||
|
|
|
|||
|
|
### Move Convention
|
|||
|
|
|
|||
|
|
Moves describe the **numbered tile** moving into the blank, **not** the blank moving.
|
|||
|
|
|
|||
|
|
For example, if the board is:
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
13 _ 14 15
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
then
|
|||
|
|
|
|||
|
|
```xml
|
|||
|
|
<move>left</move>
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
means tile **14** moves left into the blank, producing:
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
13 14 _ 15
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
## Intended Use
|
|||
|
|
|
|||
|
|
This checkpoint is primarily intended as an SFT initialization for later RL training on the 15-puzzle.
|
|||
|
|
|
|||
|
|
The goal of SFT is to teach:
|
|||
|
|
|
|||
|
|
- Task format
|
|||
|
|
- The `<think>...</think>` and `<move>...</move>` response structure
|
|||
|
|
- Legal move semantics
|
|||
|
|
- Basic trajectory-following behavior
|
|||
|
|
|
|||
|
|
It should **not** be treated as a fully capable 15-puzzle solver, especially on deeper scrambles.
|
|||
|
|
|
|||
|
|
## Response Format
|
|||
|
|
|
|||
|
|
The model is trained to respond with:
|
|||
|
|
|
|||
|
|
```xml
|
|||
|
|
<think>
|
|||
|
|
short rationale
|
|||
|
|
</think>
|
|||
|
|
<move>left</move>
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Allowed moves are:
|
|||
|
|
|
|||
|
|
- `up`
|
|||
|
|
- `down`
|
|||
|
|
- `left`
|
|||
|
|
- `right`
|
|||
|
|
|
|||
|
|
## Training
|
|||
|
|
|
|||
|
|
**Base model**
|
|||
|
|
|
|||
|
|
- `PrimeIntellect/Qwen3-1.7B`
|
|||
|
|
|
|||
|
|
**Fine-tuning framework**
|
|||
|
|
|
|||
|
|
- PrimeIntellect Prime-RL SFT
|
|||
|
|
|
|||
|
|
**Training setup**
|
|||
|
|
|
|||
|
|
- Full-parameter SFT
|
|||
|
|
- 600 examples
|
|||
|
|
- 40 training steps
|
|||
|
|
- 3 epochs
|
|||
|
|
- Sequence length: 3072
|
|||
|
|
- Learning rate: `1e-5`
|
|||
|
|
- Final loss: approximately **0.47**
|
|||
|
|
|
|||
|
|
## Loading
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|||
|
|
|
|||
|
|
model_id = "saad1926q/qwen3-1.7b-15puzzle-sft-depth-8-15"
|
|||
|
|
|
|||
|
|
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|||
|
|
model_id,
|
|||
|
|
device_map="auto",
|
|||
|
|
)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
## Example Prompt
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
You are Player 0 in 15-puzzle.
|
|||
|
|
|
|||
|
|
A 4x4 sliding puzzle board is given. The blank tile is shown as _.
|
|||
|
|
|
|||
|
|
Your goal is to reach the solved board:
|
|||
|
|
|
|||
|
|
1 2 3 4
|
|||
|
|
5 6 7 8
|
|||
|
|
9 10 11 12
|
|||
|
|
13 14 15 _
|
|||
|
|
|
|||
|
|
At each turn, choose one legal move.
|
|||
|
|
|
|||
|
|
Moves describe the numbered tile moving into the blank, not the blank moving.
|
|||
|
|
|
|||
|
|
Allowed moves are:
|
|||
|
|
up, down, left, right.
|
|||
|
|
|
|||
|
|
Wrap your move in <move>...</move>.
|
|||
|
|
|
|||
|
|
Current board:
|
|||
|
|
|
|||
|
|
1 2 3 4
|
|||
|
|
5 6 7 8
|
|||
|
|
9 10 11 _
|
|||
|
|
13 14 15 12
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Expected style:
|
|||
|
|
|
|||
|
|
```xml
|
|||
|
|
<think>
|
|||
|
|
Move tile 12 up into the blank to place it in its correct final position.
|
|||
|
|
</think>
|
|||
|
|
<move>up</move>
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
## Limitations
|
|||
|
|
|
|||
|
|
- This is an SFT checkpoint, not an RL-trained solver.
|
|||
|
|
- The model may still make illegal or suboptimal moves.
|
|||
|
|
- The training set is small and focused on scramble depths 8–15.
|
|||
|
|
- The rationales are intended to teach format and move semantics, not guarantee perfect search behavior.
|