Model: GoldenGrapeGentleman1/battle-game-agent-full-sft Source: Original Platform
base_model, library_name, tags, license
| base_model | library_name | tags | license | |||||
|---|---|---|---|---|---|---|---|---|
| Qwen/Qwen3-4B | transformers |
|
apache-2.0 |
turn-based battle game Agent — Full SFT
Fine-tuned Qwen3-4B on 2.3M turn-based battle game replay logs. The model reads raw battle protocol lines and outputs the next action as move … or switch ….
Use case: Competitive tier agent for gen9randombattle — load directly for inference or as the base for battle-oriented GRPO.
Quick start
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="GoldenGrapeGentleman1/battle game-showdown-agent-full-sft",
max_seq_length=2048,
load_in_4bit=False,
)
Training
| Setting | Value |
|---|---|
| Base model | Qwen/Qwen3-4B |
| Data | Raw Showdown replay logs (~2.3M samples) |
| Method | Full SFT (merged weights) |
| Hardware | AMD Instinct MI300X, ROCm, bfloat16 |
Tutorial
End-to-end ROCm notebook: battle game LLM Agent with Unsloth — set POKEMON_AGENT_TIER=competitive and POKEMON_HF_FULL_SFT=GoldenGrapeGentleman1/battle game-showdown-agent-full-sft.
License
Apache-2.0 (base model Qwen3-4B).
Description
Languages
Jinja
100%