114 lines
3.4 KiB
Markdown
114 lines
3.4 KiB
Markdown
|
|
---
|
||
|
|
license: apache-2.0
|
||
|
|
language:
|
||
|
|
- en
|
||
|
|
- zh
|
||
|
|
tags:
|
||
|
|
- code
|
||
|
|
- qwen2.5
|
||
|
|
- coding
|
||
|
|
- self-play
|
||
|
|
- fine-tuning
|
||
|
|
- rocm
|
||
|
|
- vulkan
|
||
|
|
- gguf
|
||
|
|
pipeline_tag: text-generation
|
||
|
|
base_model: Qwen/Qwen2.5-Coder-0.5B
|
||
|
|
library_name: transformers
|
||
|
|
---
|
||
|
|
|
||
|
|
# CloneOllama Self-Play Coder 0.5B
|
||
|
|
|
||
|
|
A coding-focused fine-tune of Qwen2.5-Coder-0.5B, trained via **self-play reinforcement learning** using 4 LLM judges and DeepSeek API evaluation. Trained on an AMD Radeon RX 9070 XT (ROCm 7.2) as part of the [CloneOllama](https://github.com/Maxritz/CloneOllama) project.
|
||
|
|
|
||
|
|
## Training Method: Self-Play
|
||
|
|
|
||
|
|
```
|
||
|
|
prompt → Student (Qwen2.5-Coder-0.5B) → generates code
|
||
|
|
→ Fast Python syntax scorer (1ms)
|
||
|
|
→ DeepSeek API reasoning judge (~3s)
|
||
|
|
→ Score ≥ 5.5 → LoRA rank 16 fine-tune
|
||
|
|
```
|
||
|
|
|
||
|
|
- **500 rounds**, 397 accepted (79% acceptance rate)
|
||
|
|
- **Scores improved**: 5.3 → 5.5 / 10 over training
|
||
|
|
- **Training data**: Generated live by the student model, no human labeling
|
||
|
|
- **Prompts**: Python coding, C# coding, algorithms, CS theory
|
||
|
|
|
||
|
|
## Benchmarks (RX 9070 XT, Vulkan)
|
||
|
|
|
||
|
|
| Metric | Value |
|
||
|
|
|--------|-------|
|
||
|
|
| Prompt speed | **619 tokens/s** |
|
||
|
|
| Generation speed | **540 tokens/s** |
|
||
|
|
| Model size | 506 MB (Q8_0 GGUF) / 942 MB (FP16 safetensors) |
|
||
|
|
|
||
|
|
## Code Quality
|
||
|
|
|
||
|
|
| Prompt | Output |
|
||
|
|
|--------|--------|
|
||
|
|
| `def is_palindrome(s):` | `s = s.lower().replace(" ", ""); return s == s[::-1]` ✅ |
|
||
|
|
| `def fibonacci(n):` | `n = int(n); return [1, n]` |
|
||
|
|
| `def reverse_list(arr):` | `def reverse(arr, size): arr = [1,2,3,4,5]` |
|
||
|
|
| `def binary_search(arr, target):` | Partial structure |
|
||
|
|
|
||
|
|
## Judges Used During Training
|
||
|
|
|
||
|
|
| Judge Model | Format | Gen Speed |
|
||
|
|
|-------------|--------|-----------|
|
||
|
|
| DeepSeek-Coder-V2 16B | Q4_K_M GGUF | 221 t/s |
|
||
|
|
| VibeThinker-3B | Q6_K GGUF | 177 t/s |
|
||
|
|
| gemma4-coding | Q4_K_M GGUF | 56 t/s |
|
||
|
|
| Qwen2.5-Coder-7B | i1-Q4_K_M GGUF | 120 t/s |
|
||
|
|
|
||
|
|
All judges ran concurrently as llama-server instances on different ports.
|
||
|
|
|
||
|
|
## Files
|
||
|
|
|
||
|
|
| File | Size | Format |
|
||
|
|
|------|------|--------|
|
||
|
|
| `selfplay-coder-0.5B-Q8_0.gguf` | 531 MB | GGUF Q8_0 (llama.cpp) |
|
||
|
|
| `model.safetensors` | 942 MB | HuggingFace safetensors |
|
||
|
|
| `config.json` | — | Model config |
|
||
|
|
| `tokenizer.json` | 11.4 MB | Qwen2.5 tokenizer |
|
||
|
|
|
||
|
|
## Usage
|
||
|
|
|
||
|
|
### llama.cpp / Ollama
|
||
|
|
|
||
|
|
```bash
|
||
|
|
llama-server -m selfplay-coder-0.5B-Q8_0.gguf -ngl 99 --port 8080
|
||
|
|
```
|
||
|
|
|
||
|
|
### Transformers (Python)
|
||
|
|
|
||
|
|
```python
|
||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||
|
|
|
||
|
|
model = AutoModelForCausalLM.from_pretrained("Strangefrost/CloneOllama-selfplay-coder-0.5B")
|
||
|
|
tokenizer = AutoTokenizer.from_pretrained("Strangefrost/CloneOllama-selfplay-coder-0.5B")
|
||
|
|
|
||
|
|
prompt = "def fibonacci(n):"
|
||
|
|
inputs = tokenizer(prompt, return_tensors="pt")
|
||
|
|
outputs = model.generate(**inputs, max_new_tokens=50)
|
||
|
|
print(tokenizer.decode(outputs[0]))
|
||
|
|
```
|
||
|
|
|
||
|
|
## Training Pipeline
|
||
|
|
|
||
|
|
Full pipeline scripts at [`model-kitchen/`](https://github.com/Maxritz/CloneOllama/tree/main/model-kitchen):
|
||
|
|
|
||
|
|
| Script | Purpose |
|
||
|
|
|--------|---------|
|
||
|
|
| `selfplay_train.py` | Multi-judge self-play training |
|
||
|
|
| `prune.py` | Structured pruning + PEFT LoRA |
|
||
|
|
| `train_utils.py` | LoRA, distillation, calibration |
|
||
|
|
| `merge.py` | SLERP / TIES / DARE merging |
|
||
|
|
| `upcycle.py` | Dense → MoE conversion |
|
||
|
|
|
||
|
|
## Acknowledgments
|
||
|
|
|
||
|
|
- Base model: [Qwen/Qwen2.5-Coder-0.5B](https://huggingface.co/Qwen/Qwen2.5-Coder-0.5B)
|
||
|
|
- Training platform: CloneOllama on AMD Radeon RX 9070 XT (ROCm 7.2)
|
||
|
|
- Judge models: DeepSeek-Coder-V2, VibeThinker, gemma4-coding, Qwen2.5-Coder-7B
|