104 lines
5.4 KiB
Markdown
104 lines
5.4 KiB
Markdown
|
|
---
|
|||
|
|
base_model: Qwen/Qwen3-8B
|
|||
|
|
base_model_relation: finetune
|
|||
|
|
datasets:
|
|||
|
|
- Glint-Research/Fable-5-traces
|
|||
|
|
- Roman1111111/gpt5.5-terminal
|
|||
|
|
license: apache-2.0
|
|||
|
|
language:
|
|||
|
|
- en
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
library_name: transformers
|
|||
|
|
tags:
|
|||
|
|
- safetensors
|
|||
|
|
- qlora
|
|||
|
|
- agentic
|
|||
|
|
- coding
|
|||
|
|
- reasoning
|
|||
|
|
- thinking
|
|||
|
|
- claude
|
|||
|
|
- qwen3
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Parable-Qwen3-8B-Claude-Fable-5
|
|||
|
|
|
|||
|
|
<picture>
|
|||
|
|
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parable_header_dark.png">
|
|||
|
|
<img alt="Parable" src="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parable_header.png">
|
|||
|
|
</picture>
|
|||
|
|
|
|||
|
|
**Qwen3-8B trained on real Claude Fable 5 and GPT-5.5 agent traces: 67% lower held-out test loss than its base, and the strongest strictly-graded qualitative score in the Parable series: 23 of 34 fully correct.**
|
|||
|
|
|
|||
|
|
Parable-Qwen3-8B is a [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) fine-tune trained on real multi-step agent sessions: planning, tool use, and `<think>` reasoning captured from actual Claude Fable 5 and GPT-5.5 agent work, not synthetic Q&A. Highest strict-qual release in the Parable series, alongside the [Granite 8B line](https://huggingface.co/AnkitAI/Parable-Granite-4.1-8B-Claude-Fable-5).
|
|||
|
|
|
|||
|
|
## Usage
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|||
|
|
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|||
|
|
"AnkitAI/Parable-Qwen3-8B-Claude-Fable-5",
|
|||
|
|
torch_dtype="auto", device_map="auto")
|
|||
|
|
tok = AutoTokenizer.from_pretrained("AnkitAI/Parable-Qwen3-8B-Claude-Fable-5")
|
|||
|
|
|
|||
|
|
msgs = [{"role": "user", "content": "Write a Python function that retries an HTTP request with exponential backoff."}]
|
|||
|
|
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
|
|||
|
|
out = model.generate(ids, max_new_tokens=3000, temperature=0.7, top_p=0.95, do_sample=True)
|
|||
|
|
text = tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True)
|
|||
|
|
answer = text.split("</think>")[-1].strip() # response opens with a <think> block
|
|||
|
|
print(answer)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
GGUF quants for llama.cpp / Ollama / LM Studio: [Parable-Qwen3-8B-Claude-Fable-5-GGUF](https://huggingface.co/AnkitAI/Parable-Qwen3-8B-Claude-Fable-5-GGUF).
|
|||
|
|
|
|||
|
|
**Sampling:** temperature 0.7, top_p 0.95, generous max_new_tokens (at least 2500).
|
|||
|
|
|
|||
|
|
## Training data
|
|||
|
|
|
|||
|
|
- [Glint-Research/Fable-5-traces](https://huggingface.co/datasets/Glint-Research/Fable-5-traces): 4.4k real Claude Fable 5 coding-agent session traces with `<think>` reasoning and tool calls (AGPL-3.0)
|
|||
|
|
- [Roman1111111/gpt5.5-terminal](https://huggingface.co/datasets/Roman1111111/gpt5.5-terminal): terminal-agent task solutions (MIT)
|
|||
|
|
|
|||
|
|
Every example passed a quality gate (schema validation, secrets scrub, length filtering) before training. QLoRA fine-tune (NF4, sequence length 1024) trained on a single 16 GB GPU, quantized with [llama.cpp](https://github.com/ggml-org/llama.cpp).
|
|||
|
|
|
|||
|
|
## Evaluation
|
|||
|
|
|
|||
|
|

|
|||
|
|
|
|||
|
|
Held-out test split, identical evaluation code and context length for base and fine-tune:
|
|||
|
|
|
|||
|
|
| Metric | Base Qwen3-8B | Parable | Δ |
|
|||
|
|
|---|---|---|---|
|
|||
|
|
| Test loss | 2.162 | **0.712** | **−67%** |
|
|||
|
|
|
|||
|
|
**Qualitative review** (34 coding/terminal/debugging prompts, strictly graded by mentally executing every answer): **23/34 fully correct, 30/34 correct or partially correct** — the highest fully-correct score in the series. We publish these numbers because strict qualitative grading is rare in this niche; judge accordingly.
|
|||
|
|
|
|||
|
|
For reference, the strongest published fine-tune on this data family (a 9B) reports 0.71 validation loss. Cross-repo numbers are indicative only: splits, tokenizers, and context lengths differ (ours is measured at 1,024 tokens).
|
|||
|
|
|
|||
|
|
## Limitations
|
|||
|
|
|
|||
|
|
- Trained for agent work: on ops-style prompts it sometimes (2/34 in our eval) responds with structured tool-call JSON rather than prose. Useful inside agent harnesses; in plain chat, re-prompt or lower the temperature.
|
|||
|
|
- Fine-tuned at 1,024-token sequences; the base model's native 128K-token context remains fully available, so long sessions work, with the fine-tuned behavior strongest in the opening turns.
|
|||
|
|
|
|||
|
|
As a fine-tune it inherits Qwen3-8B's base behaviors and knowledge cutoff. As with any local model, treat generated commands and code as drafts to review.
|
|||
|
|
|
|||
|
|
## Provenance & licensing
|
|||
|
|
|
|||
|
|
Model weights: **Apache-2.0** (inherited from Qwen3-8B). Training data licenses: Fable-5-traces **AGPL-3.0**, gpt5.5-terminal **MIT**. Because those traces originate from third-party assistants, the providers' terms may apply to downstream training and distillation. If you plan to build on this model commercially, confirm your use aligns with those terms.
|
|||
|
|
|
|||
|
|
|
|||
|
|
## Get Parable
|
|||
|
|
|
|||
|
|
| Platform | Command / Link |
|
|||
|
|
|---|---|
|
|||
|
|
| Ollama | `ollama run parable/qwen3-fable:8b` |
|
|||
|
|
| Ollama (family flagship, best per size) | `ollama run parable/fable` |
|
|||
|
|
| Hugging Face | [GGUF quants, full weights, eval reports](https://huggingface.co/collections/AnkitAI/parable-6a4fac60f4b35afca3019621) |
|
|||
|
|
| LM Studio | `lms get parable/qwen3-fable` ([parable on LM Studio Hub](https://lmstudio.ai/parable)) |
|
|||
|
|
|
|||
|
|
## Acknowledgements
|
|||
|
|
|
|||
|
|
- [Glint-Research](https://huggingface.co/Glint-Research) and [Roman1111111](https://huggingface.co/Roman1111111) for the open trace datasets
|
|||
|
|
- [Qwen](https://huggingface.co/Qwen) for the base model
|
|||
|
|
- [empero-ai](https://huggingface.co/empero-ai), whose Qwable recipe the Parable series follows
|
|||
|
|
- [llama.cpp](https://github.com/ggml-org/llama.cpp)
|