47 lines
1.2 KiB
Markdown
47 lines
1.2 KiB
Markdown
---
|
|
base_model: Qwen/Qwen3-4B
|
|
library_name: transformers
|
|
tags:
|
|
- battle game
|
|
- showdown
|
|
- sft
|
|
- unsloth
|
|
- rocm
|
|
license: apache-2.0
|
|
---
|
|
|
|
# turn-based battle game Agent — Full SFT
|
|
|
|
Fine-tuned **Qwen3-4B** on **2.3M** turn-based battle game replay logs. The model reads raw battle protocol lines and outputs the next action as `move …` or `switch …`.
|
|
|
|
**Use case:** Competitive tier agent for gen9randombattle — load directly for inference or as the base for battle-oriented GRPO.
|
|
|
|
## Quick start
|
|
|
|
```python
|
|
from unsloth import FastLanguageModel
|
|
|
|
model, tokenizer = FastLanguageModel.from_pretrained(
|
|
model_name="GoldenGrapeGentleman1/battle game-showdown-agent-full-sft",
|
|
max_seq_length=2048,
|
|
load_in_4bit=False,
|
|
)
|
|
```
|
|
|
|
## Training
|
|
|
|
| Setting | Value |
|
|
|---------|-------|
|
|
| Base model | Qwen/Qwen3-4B |
|
|
| Data | Raw Showdown replay logs (~2.3M samples) |
|
|
| Method | Full SFT (merged weights) |
|
|
| Hardware | AMD Instinct MI300X, ROCm, bfloat16 |
|
|
|
|
## Tutorial
|
|
|
|
End-to-end ROCm notebook: [battle game LLM Agent with Unsloth](https://github.com/ROCm/gpuaidev) — set `POKEMON_AGENT_TIER=competitive` and `POKEMON_HF_FULL_SFT=GoldenGrapeGentleman1/battle game-showdown-agent-full-sft`.
|
|
|
|
## License
|
|
|
|
Apache-2.0 (base model Qwen3-4B).
|