Files
qwen3-1.7b-15puzzle-sft-dep…/README.md
ModelHub XC f9331c9068 初始化项目,由ModelHub XC社区提供模型
Model: saad1926q/qwen3-1.7b-15puzzle-sft-depth-8-15
Source: Original Platform
2026-09-04 23:16:17 +08:00

3.2 KiB
Raw Blame History

license, base_model, tags, datasets
license base_model tags datasets
apache-2.0 PrimeIntellect/Qwen3-1.7B
qwen3
15-puzzle
sft
reasoning
puzzle-solving
saad1926q/15-puzzle

Qwen3 1.7B 15-Puzzle SFT

This is a supervised fine-tuned checkpoint of PrimeIntellect/Qwen3-1.7B for 15-puzzle move prediction and reasoning-format alignment.

The model was fine-tuned on solved 15-puzzle trajectories from saad1926q/15-puzzle, using depth 8–15 puzzles.

Each training example contains:

  • A board state
  • A solver-verified optimal move sequence
  • Short per-move rationales

Training Data

Dataset split

  • Dataset: saad1926q/15-puzzle
  • Config/subset: sft
  • Split: sft
  • Rows: 600
  • Scramble depths: 8–15
  • Base trajectories: Solver-verified optimal paths
  • Rationale source: DeepSeek-generated rationales with tile-move convention

Move Convention

Moves describe the numbered tile moving into the blank, not the blank moving.

For example, if the board is:

13 _ 14 15

then

<move>left</move>

means tile 14 moves left into the blank, producing:

13 14 _ 15

Intended Use

This checkpoint is primarily intended as an SFT initialization for later RL training on the 15-puzzle.

The goal of SFT is to teach:

  • Task format
  • The <think>...</think> and <move>...</move> response structure
  • Legal move semantics
  • Basic trajectory-following behavior

It should not be treated as a fully capable 15-puzzle solver, especially on deeper scrambles.

Response Format

The model is trained to respond with:

<think>
short rationale
</think>
<move>left</move>

Allowed moves are:

  • up
  • down
  • left
  • right

Training

Base model

  • PrimeIntellect/Qwen3-1.7B

Fine-tuning framework

  • PrimeIntellect Prime-RL SFT

Training setup

  • Full-parameter SFT
  • 600 examples
  • 40 training steps
  • 3 epochs
  • Sequence length: 3072
  • Learning rate: 1e-5
  • Final loss: approximately 0.47

Loading

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "saad1926q/qwen3-1.7b-15puzzle-sft-depth-8-15"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
)

Example Prompt

You are Player 0 in 15-puzzle.

A 4x4 sliding puzzle board is given. The blank tile is shown as _.

Your goal is to reach the solved board:

1 2 3 4
5 6 7 8
9 10 11 12
13 14 15 _

At each turn, choose one legal move.

Moves describe the numbered tile moving into the blank, not the blank moving.

Allowed moves are:
up, down, left, right.

Wrap your move in <move>...</move>.

Current board:

1 2 3 4
5 6 7 8
9 10 11 _
13 14 15 12

Expected style:

<think>
Move tile 12 up into the blank to place it in its correct final position.
</think>
<move>up</move>

Limitations

  • This is an SFT checkpoint, not an RL-trained solver.
  • The model may still make illegal or suboptimal moves.
  • The training set is small and focused on scramble depths 8–15.
  • The rationales are intended to teach format and move semantics, not guarantee perfect search behavior.