Files
ModelHub XC f9331c9068 初始化项目,由ModelHub XC社区提供模型
Model: saad1926q/qwen3-1.7b-15puzzle-sft-depth-8-15
Source: Original Platform
2026-09-04 23:16:17 +08:00

170 lines
3.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
base_model: PrimeIntellect/Qwen3-1.7B
tags:
- qwen3
- 15-puzzle
- sft
- reasoning
- puzzle-solving
datasets:
- saad1926q/15-puzzle
---
# Qwen3 1.7B 15-Puzzle SFT
This is a supervised fine-tuned checkpoint of `PrimeIntellect/Qwen3-1.7B` for 15-puzzle move prediction and reasoning-format alignment.
The model was fine-tuned on solved 15-puzzle trajectories from [`saad1926q/15-puzzle`](https://huggingface.co/datasets/saad1926q/15-puzzle), using depth 8–15 puzzles.
Each training example contains:
- A board state
- A solver-verified optimal move sequence
- Short per-move rationales
## Training Data
**Dataset split**
- **Dataset:** `saad1926q/15-puzzle`
- **Config/subset:** `sft`
- **Split:** `sft`
- **Rows:** 600
- **Scramble depths:** 8–15
- **Base trajectories:** Solver-verified optimal paths
- **Rationale source:** DeepSeek-generated rationales with tile-move convention
### Move Convention
Moves describe the **numbered tile** moving into the blank, **not** the blank moving.
For example, if the board is:
```text
13 _ 14 15
```
then
```xml
<move>left</move>
```
means tile **14** moves left into the blank, producing:
```text
13 14 _ 15
```
## Intended Use
This checkpoint is primarily intended as an SFT initialization for later RL training on the 15-puzzle.
The goal of SFT is to teach:
- Task format
- The `<think>...</think>` and `<move>...</move>` response structure
- Legal move semantics
- Basic trajectory-following behavior
It should **not** be treated as a fully capable 15-puzzle solver, especially on deeper scrambles.
## Response Format
The model is trained to respond with:
```xml
<think>
short rationale
</think>
<move>left</move>
```
Allowed moves are:
- `up`
- `down`
- `left`
- `right`
## Training
**Base model**
- `PrimeIntellect/Qwen3-1.7B`
**Fine-tuning framework**
- PrimeIntellect Prime-RL SFT
**Training setup**
- Full-parameter SFT
- 600 examples
- 40 training steps
- 3 epochs
- Sequence length: 3072
- Learning rate: `1e-5`
- Final loss: approximately **0.47**
## Loading
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "saad1926q/qwen3-1.7b-15puzzle-sft-depth-8-15"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
)
```
## Example Prompt
```text
You are Player 0 in 15-puzzle.
A 4x4 sliding puzzle board is given. The blank tile is shown as _.
Your goal is to reach the solved board:
1 2 3 4
5 6 7 8
9 10 11 12
13 14 15 _
At each turn, choose one legal move.
Moves describe the numbered tile moving into the blank, not the blank moving.
Allowed moves are:
up, down, left, right.
Wrap your move in <move>...</move>.
Current board:
1 2 3 4
5 6 7 8
9 10 11 _
13 14 15 12
```
Expected style:
```xml
<think>
Move tile 12 up into the blank to place it in its correct final position.
</think>
<move>up</move>
```
## Limitations
- This is an SFT checkpoint, not an RL-trained solver.
- The model may still make illegal or suboptimal moves.
- The training set is small and focused on scramble depths 8–15.
- The rationales are intended to teach format and move semantics, not guarantee perfect search behavior.