Files
ModelHub XC 84147139e9 初始化项目,由ModelHub XC社区提供模型
Model: AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5
Source: Original Platform
2026-08-04 20:02:02 +08:00

196 lines
9.9 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
base_model: ibm-granite/granite-4.1-3b
base_model_relation: finetune
datasets:
- AnkitAI/parable-corpus-v2
- Glint-Research/Fable-5-traces
- Roman1111111/gpt5.5-terminal
license: apache-2.0
language:
- en
pipeline_tag: text-generation
library_name: transformers
tags:
- safetensors
- transformers
- qlora
- agentic
- agent
- coding
- tool-use
- function-calling
- terminal
- reasoning
- thinking
- claude
- claude-fable-5
- distillation
- trace-training
- granite
---
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parable_header_dark.png">
<img alt="Parable" src="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parable_header.png">
</picture>
# 🪶 Parable-Granite-3B **v2** — trained on genuine Claude Fable 5 agent traces
*This is the full-precision safetensors repo (vLLM / transformers / fine-tuning). For llama.cpp, Ollama, and LM Studio use the [GGUF repo](https://huggingface.co/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF).*
### A tiny local model that thinks before it answers — planning, reasoning, and terminal instincts distilled from real agent sessions.
> **~3 GB of RAM is all you need.** Laptop, old GPU, Raspberry-Pi-class boxes with swap — the Q4 build runs
> anywhere. One command and you have a private, offline reasoning model on your machine:
>
> ```bash
> ollama run hf.co/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF:Q4_K_M
> ```
---
## The headline — v2 is a different model
v2 is a full retrain: **13× more genuine Fable 5 trace data** (11,574 sessions, 16.8M tokens — [corpus published](https://huggingface.co/datasets/AnkitAI/parable-corpus-v2)) and a rebuilt recipe (completion-masked loss, replay mixing, benchmark-gated checkpoints, seed-averaged weights).
| same harness, greedy, Q4_K_M | v1 | **v2 (this release)** |
|---|---|---|
| Dev pass-rate (MBPP subset, n=50) — *base: 0.68* | — | **0.82** |
| Agent-artifact leakage (JSON blobs, phantom turns) | 6/34 | **0/34** |
| Strict 34-prompt coding qual — *base: 27/34* | ~18/34 | **25/34** |
| HumanEval / HumanEval+ | 62.8 / 57.9 | **70.1 / 65.9** |
Clean answers, structured reasoning, agent instincts — and the transcript artifacts that leaked into v1's replies are gone. *One trade, made on purpose: raw HumanEval-style function synthesis stays the base model's turf (81.7 vs 70.1) — v2 spends that capacity on agent behavior instead, and spends half as much as v1 did.* Measurement notes below. 👇
---
## Announcements
**📌 Same links, new model.** v2 replaces v1 **in place** — every existing Ollama command, script, and bookmark now serves v2. No migration, nothing to change.
**🔮 v3 is already training.** Rejection-sampled SFT: thousands of candidate solutions generated against *executable tests*, only verified passers enter the corpus. The goal is simple — above-base agent capability, not just clean behavior. Follow [AnkitAI](https://huggingface.co/AnkitAI) for the drop.
**📦 Full family.** This 3B is the smallest Parable. Need more headroom? [8B Granite](https://huggingface.co/AnkitAI/Parable-Granite-4.1-8B-Claude-Fable-5-GGUF), [8B Qwen](https://huggingface.co/AnkitAI/Parable-Qwen3-8B-Claude-Fable-5-GGUF), [4B Qwen](https://huggingface.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF) — same recipe, no matter your hardware.
---
## How to run it
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="bfloat16", device_map="auto")
messages = [{"role": "user", "content": "Write a Python function that retries an HTTP request with exponential backoff."}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=3000, temperature=0.7, top_p=0.95, do_sample=True)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
```
GGUF quants (2.1-6.8 GB, runs in ~3 GB RAM): [Parable-Granite-4.1-3B-Claude-Fable-5-GGUF](https://huggingface.co/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF)
### Thinking mode
Every answer opens with a `<think>...</think>` reasoning block — that's the Fable 5 heritage. llama.cpp's `--jinja` mode separates it automatically; strip it before showing replies to end users.
**Sampling:** temperature 0.7, top_p 0.95, and budget `max_tokens` generously (**2500+**) — trace-trained models think at length before answering.
---
## Measurement notes
All numbers: identical llama.cpp harness, greedy decoding, Q4_K_M, **base model measured on the same instrument**. We train multiple seeds and ship the weight-average — single-run scores at 3B swing ±3 points on GPU nondeterminism alone, so most cards report their luckiest run; we ship the average and report the shipped weights' own numbers. Raw eval outputs live in this repo.
**Which model should you use?** Pure single-function code completion → the base model is genuinely strong there. Explanations, debugging, terminal workflows, structured reasoning, agent-style tasks → that's what Parable is trained on, and where v2 shines.
## What's new in v2 (training)
The recipe follows our ongoing tech report (in preparation):
- **Completion-only loss masking** ([Hermes 3](https://arxiv.org/abs/2408.11857), [Tülu 3](https://arxiv.org/abs/2411.15124)) — loss on assistant tokens only, so the model learns to *answer*, not to imitate transcripts
- **30% replay mix** of general instruction data ([Luo et al.](https://arxiv.org/abs/2308.08747), [Biderman et al.](https://arxiv.org/abs/2405.09673)) — the anti-forgetting lever
- **Session re-segmentation + sanitization** — why v1 sometimes leaked agent JSON into normal chat, and v2 never does (0/34)
- **Benchmark-gated checkpoints** ([Dong et al.](https://arxiv.org/abs/2310.05492)) instead of fixed epochs
- **Seed-averaged weights** ([model soups, Wortsman et al.](https://arxiv.org/abs/2203.05482)) — we ship the average of multiple runs, not the lottery winner
With Claude Fable 5 now retired, genuine self-authored Fable traces are a fixed, non-renewable corpus. Unlike most models in this niche, **our full training corpus is public**: [AnkitAI/parable-corpus-v2](https://huggingface.co/datasets/AnkitAI/parable-corpus-v2) — deduplicated, quality-gated, provenance-tagged.
## Good to know
- Fine-tuned at 2,048-token sequences; the base 128K context stays available, fine-tuned behavior is strongest in the opening turns.
- Not trained for: multi-file repo navigation, vision, non-English.
- Inherits Granite-4.1-3B's knowledge cutoff. Treat generated commands as drafts to review.
## Evaluation
### Function calling (BFCL V3, AST subset)
Measured 2026-07-29: bfcl-eval at gorilla main, prompting mode, Q4_K_M
GGUFs served by llama.cpp on a T4, base and Parable under the identical
harness. Categories: simple_python / multiple / parallel /
parallel_multiple (400/200/200/200 items). Raw generations and score
files: [parable-v2-artifacts](https://huggingface.co/AnkitAI/parable-v2-artifacts)
under `verify/bfcl/`.
| | simple_python | multiple | parallel | parallel_multiple |
|---|---|---|---|---|
| Granite-4.1-3B base | 0.848 | 0.790 | 0.710 | 0.665 |
| **This model (chat variant)** | 0.413 | 0.605 | 0.320 | 0.425 |
For tool-calling workloads, use the base model; this variant is built
for reasoning prose. The drop has a specific mechanism: sampled generations show the model
intermittently answering with args-only tool-call JSON (for example
`{"base": 10, "height": 5}`) instead of a function call, which the
AST scorer rejects. That is trace-scaffolding format bleeding into
standalone tasks, the failure mode the series paper names *session
leakage* (Section 6 of the report). The reasoning-voice strengths this variant trains for are unaffected
on prose tasks.
## Base & license
Weights: **Apache-2.0** (inherited from [ibm-granite/granite-4.1-3b](https://huggingface.co/ibm-granite/granite-4.1-3b)). Training data: Fable-5-traces **AGPL-3.0**, gpt5.5-terminal **MIT** — since traces originate from third-party assistants, their terms may apply to downstream training; check before commercial distillation.
## Get Parable
| Platform | |
|---|---|
| Ollama | `ollama run parable/granite4.1-fable:3b` · [parable namespace](https://ollama.com/parable) |
| Hugging Face | [full collection](https://huggingface.co/collections/AnkitAI/parable-6a4fac60f4b35afca3019621) |
| LM Studio | search "parable" in-app |
| ModelScope | [Parable on ModelScope](https://modelscope.cn/models/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF) |
## Citation
The recipe, evaluation methodology and failure analysis behind this model are
documented in the tech report:
> Aglawe, A. (2026). *Agent-Trace Fine-Tuning of Small Language Models under
> Constrained Compute.* Zenodo. [doi:10.5281/zenodo.21676407](https://doi.org/10.5281/zenodo.21676407)
```bibtex
@misc{aglawe2026agenttrace,
author = {Aglawe, Ankit},
title = {Agent-Trace Fine-Tuning of Small Language Models under Constrained Compute},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.21676407},
url = {https://doi.org/10.5281/zenodo.21676407}
}
```
## Acknowledgements
[Glint-Research](https://huggingface.co/Glint-Research) & [Roman1111111](https://huggingface.co/Roman1111111) for the open trace data · [IBM Granite](https://huggingface.co/ibm-granite) for the base · [empero-ai](https://huggingface.co/empero-ai) whose Qwable recipe inspired the series · [llama.cpp](https://github.com/ggml-org/llama.cpp)
## Version history
- **v2** (2026-07-16) — this release. 13× corpus, rebuilt recipe, seed-averaged weights, zero leakage.
- **v1** (2026-07) — initial release, 857-row corpus. Preserved as repo revision history.
---
### Real Fable 5 reasoning. Yours, offline, right now.
More on the Parable models: [ankitaglawe.com/parable](https://ankitaglawe.com/parable)