license, language, tags, pipeline_tag, base_model, library_name
license
language
tags
pipeline_tag
base_model
library_name
apache-2.0
code
qwen2.5
coding
self-play
fine-tuning
rocm
vulkan
gguf
text-generation
Qwen/Qwen2.5-Coder-0.5B
transformers
CloneOllama Self-Play Coder 0.5B
A coding-focused fine-tune of Qwen2.5-Coder-0.5B, trained via self-play reinforcement learning using 4 LLM judges and DeepSeek API evaluation. Trained on an AMD Radeon RX 9070 XT (ROCm 7.2) as part of the CloneOllama project.
Training Method: Self-Play
500 rounds , 397 accepted (79% acceptance rate)
Scores improved : 5.3 → 5.5 / 10 over training
Training data : Generated live by the student model, no human labeling
Prompts : Python coding, C# coding, algorithms, CS theory
Benchmarks (RX 9070 XT, Vulkan)
Metric
Value
Prompt speed
619 tokens/s
Generation speed
540 tokens/s
Model size
506 MB (Q8_0 GGUF) / 942 MB (FP16 safetensors)
Code Quality
Prompt
Output
def is_palindrome(s):
s = s.lower().replace(" ", ""); return s == s[::-1] ✅
def fibonacci(n):
n = int(n); return [1, n]
def reverse_list(arr):
def reverse(arr, size): arr = [1,2,3,4,5]
def binary_search(arr, target):
Partial structure
Judges Used During Training
Judge Model
Format
Gen Speed
DeepSeek-Coder-V2 16B
Q4_K_M GGUF
221 t/s
VibeThinker-3B
Q6_K GGUF
177 t/s
gemma4-coding
Q4_K_M GGUF
56 t/s
Qwen2.5-Coder-7B
i1-Q4_K_M GGUF
120 t/s
All judges ran concurrently as llama-server instances on different ports.
Files
File
Size
Format
selfplay-coder-0.5B-Q8_0.gguf
531 MB
GGUF Q8_0 (llama.cpp)
model.safetensors
942 MB
HuggingFace safetensors
config.json
—
Model config
tokenizer.json
11.4 MB
Qwen2.5 tokenizer
Usage
llama.cpp / Ollama
Transformers (Python)
Training Pipeline
Full pipeline scripts at model-kitchen/ :
Script
Purpose
selfplay_train.py
Multi-judge self-play training
prune.py
Structured pruning + PEFT LoRA
train_utils.py
LoRA, distillation, calibration
merge.py
SLERP / TIES / DARE merging
upcycle.py
Dense → MoE conversion
Acknowledgments
Base model: Qwen/Qwen2.5-Coder-0.5B
Training platform: CloneOllama on AMD Radeon RX 9070 XT (ROCm 7.2)
Judge models: DeepSeek-Coder-V2, VibeThinker, gemma4-coding, Qwen2.5-Coder-7B