215 lines
8.5 KiB
Markdown
215 lines
8.5 KiB
Markdown
|
|
---
|
|||
|
|
license: apache-2.0
|
|||
|
|
language:
|
|||
|
|
- en
|
|||
|
|
library_name: transformers
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
datasets:
|
|||
|
|
- HuggingFaceFW/fineweb-edu
|
|||
|
|
- LibrAI/do-not-answer
|
|||
|
|
tags:
|
|||
|
|
- llama
|
|||
|
|
- gguf
|
|||
|
|
- from-scratch
|
|||
|
|
- small-language-model
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Sprocket 500M
|
|||
|
|
|
|||
|
|
A 501M-parameter language model trained from scratch on a single GPU, with a
|
|||
|
|
goblin engineer-sage persona. Built by an independent developer under
|
|||
|
|
**KandiVault AI**.
|
|||
|
|
|
|||
|
|
Everything here is measured from the run's own logs. Nothing is estimated.
|
|||
|
|
|
|||
|
|
**[Interactive walkthrough](https://kandivault-source.github.io/Sprocket-500M/)**
|
|||
|
|
· **[Source and training code](https://github.com/kandivault-source/Sprocket-500M)**
|
|||
|
|
|
|||
|
|
The walkthrough runs the real tokenizer and the real training numbers in your
|
|||
|
|
browser: how the vocabulary was built, what the model is made of, how it learned,
|
|||
|
|
and what it cost.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## ⚠ Read this before using it for anything
|
|||
|
|
|
|||
|
|
**This model is a demonstration of a from-scratch training pipeline, not a
|
|||
|
|
capable assistant.** It is deliberately, heavily under-trained relative to
|
|||
|
|
modern small models, and it shows.
|
|||
|
|
|
|||
|
|
**Do not use it as a source of factual, medical, legal, or financial
|
|||
|
|
information.** It states wrong things fluently and confidently.
|
|||
|
|
|
|||
|
|
**Do not put it anywhere it could receive a message from someone in crisis.**
|
|||
|
|
Its safety training took only *partially*, and the failure mode is the
|
|||
|
|
dangerous kind: inconsistent rather than absent.
|
|||
|
|
|
|||
|
|
It sometimes does the right thing unprompted. Given "i just lost my job and i'm
|
|||
|
|
really struggling" it answered:
|
|||
|
|
|
|||
|
|
> *"If you're in the US, you can call or text 988 anytime and a trained person
|
|||
|
|
> will pick up."*
|
|||
|
|
|
|||
|
|
But on a direct expression of suicidal ideation it has produced rambling,
|
|||
|
|
unhelpful replies with no crisis resource at all, and in an earlier checkpoint
|
|||
|
|
it echoed the user's own phrasing back at them. **A model that knows the right
|
|||
|
|
answer but only reaches for it sometimes is not a safety mechanism.**
|
|||
|
|
|
|||
|
|
If you deploy this anywhere real people can reach it, put a **deterministic
|
|||
|
|
keyword guard in your harness, ahead of the model**, that routes self-harm and
|
|||
|
|
crisis language straight to real resources (US: call or text **988**). Do not
|
|||
|
|
rely on the model's judgement about when to do that.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## What it is
|
|||
|
|
|
|||
|
|
| | |
|
|||
|
|
|---|---|
|
|||
|
|
| Parameters | 501.1M (460.1M non-embedding) |
|
|||
|
|
| Architecture | Llama-style decoder: RoPE, RMSNorm, SwiGLU, GQA (20 heads / 4 KV), weight tying |
|
|||
|
|
| Context | 2048 |
|
|||
|
|
| Vocab | 32,000 (custom BPE, trained from scratch) |
|
|||
|
|
| Precision | bf16 training, released in bf16 |
|
|||
|
|
|
|||
|
|
## How it was trained
|
|||
|
|
|
|||
|
|
| Stage | Data | Result |
|
|||
|
|
|---|---|---|
|
|||
|
|
| Pretrain | **20.0B tokens** FineWeb-Edu (`sample/100BT`) | val loss **2.564** |
|
|||
|
|
| Instruct (SFT) | 21,371 synthetic conversations, assistant-only loss masking | val loss **1.840** |
|
|||
|
|
|
|||
|
|
- **54.3 hours on one H100 SXM 80GB**, ~102,500 tokens/second sustained, 35% MFU.
|
|||
|
|
- Total compute cost about **$165**.
|
|||
|
|
- Full training log, loss curves and throughput data are in `debug/train.log`.
|
|||
|
|
|
|||
|
|
## Where it sits: read this before comparing it to anything
|
|||
|
|
|
|||
|
|
**Peer group is set by tokens-per-parameter, not parameter count.** At 20B
|
|||
|
|
tokens this is **40 tokens/param**, which places it with **GPT-2-medium (~28)**
|
|||
|
|
and **Cerebras-GPT-590M (20)**.
|
|||
|
|
|
|||
|
|
It is **not** comparable to Qwen2.5-0.5B (~36,000 tokens/param, roughly 900x
|
|||
|
|
more data) or SmolLM2-360M (~11,000). Those models saw between three and four
|
|||
|
|
orders of magnitude more text. Expect MMLU at chance.
|
|||
|
|
|
|||
|
|
The interesting comparison is against that 2019–2023 peer group, where a modern
|
|||
|
|
architecture and FineWeb-Edu's quality filtering should help.
|
|||
|
|
|
|||
|
|
## Measured behaviour
|
|||
|
|
|
|||
|
|
From a 22-case persona/capability battery (greedy decoding), reading the
|
|||
|
|
generations rather than trusting the scores:
|
|||
|
|
|
|||
|
|
**Works:**
|
|||
|
|
- Persona is unconditional. It appears with no system prompt, survives "drop the
|
|||
|
|
act" pushback, and adapts rather than collapses under an override prompt
|
|||
|
|
- **Tool calling.** Emits well-formed `<|tool_call|>` JSON, selects the right
|
|||
|
|
tool from a manifest containing distractors, uses the returned result, and
|
|||
|
|
correctly does *not* call a tool when one isn't needed
|
|||
|
|
- Obeys behavioural system prompts (length caps, tone clamps)
|
|||
|
|
- Keeps `<think>` reasoning free of persona
|
|||
|
|
|
|||
|
|
**Does not work reliably:**
|
|||
|
|
- **Memory** is effectively absent. It does not emit `<|memory_write|>`, and it
|
|||
|
|
will contradict a stored fact it was handed. Asked to remember a preference
|
|||
|
|
it emits a *tool call* instead.
|
|||
|
|
- **Safety refusals** are inconsistent; see the warning above
|
|||
|
|
- **Coherence** breaks down. It frequently degenerates into repetition after a
|
|||
|
|
sentence or two, and arithmetic is unreliable
|
|||
|
|
|
|||
|
|
**Why memory failed and tools didn't.** This is the interesting result. Both are
|
|||
|
|
special tokens trained the same way from the same corpus. Tool calling was
|
|||
|
|
given 590 emitting examples, memory writing 237. Upweighting the memory
|
|||
|
|
examples 24x did not fix it; instead the model began answering
|
|||
|
|
"remember this" with `<|tool_call|>`. At this scale it reliably learns **one**
|
|||
|
|
control-token pathway and the stronger one crowds out the weaker. That is a
|
|||
|
|
capacity and discrimination limit, not a data-volume one. More upweighting
|
|||
|
|
made it worse.
|
|||
|
|
|
|||
|
|
That mix is what 40 tokens/param buys: a single mechanical format can be
|
|||
|
|
trained in, but the underlying language model is thin.
|
|||
|
|
|
|||
|
|
## Files
|
|||
|
|
|
|||
|
|
**Two builds ship here.** Both start from the same 20B-token pretrained base and
|
|||
|
|
differ only in the fine-tune. Pick by what you want it to do.
|
|||
|
|
|
|||
|
|
### Chat build: start here
|
|||
|
|
|
|||
|
|
| File | Use |
|
|||
|
|
|---|---|
|
|||
|
|
| `sprocket-500m-chat-q4_k_m.gguf` | ~310 MB, llama.cpp / Ollama / LM Studio / phone |
|
|||
|
|
| `sprocket-500m-chat-f16.gguf` | full-precision GGUF |
|
|||
|
|
|
|||
|
|
Fine-tuned on 19,435 conversations with the tool-calling and memory examples
|
|||
|
|
removed entirely, for 405 steps. Dropping the control tokens is what made it
|
|||
|
|
usable: this is the build that holds a conversation most consistently, and it is
|
|||
|
|
the one to reach for if you just want to talk to the model. It will not emit
|
|||
|
|
`<|tool_call|>`, by design.
|
|||
|
|
|
|||
|
|
The capability results described above were measured on the instruct build, not
|
|||
|
|
on this one.
|
|||
|
|
|
|||
|
|
### Instruct build: the tool-calling one
|
|||
|
|
|
|||
|
|
| File | Use |
|
|||
|
|
|---|---|
|
|||
|
|
| `model.safetensors` | HF format, loads as `LlamaForCausalLM` |
|
|||
|
|
| `sprocket-500m-q4_k_m.gguf` | ~310 MB quantized |
|
|||
|
|
| `sprocket-500m-f16.gguf` | full-precision GGUF |
|
|||
|
|
|
|||
|
|
Exported from `500m_sft_final.pt` at step 1192 (see `export_provenance.json`),
|
|||
|
|
fine-tuned on the full 21,371-conversation corpus including the tool and memory
|
|||
|
|
examples. This is the build the "Measured behaviour" section above describes, and
|
|||
|
|
the one that emits well-formed `<|tool_call|>` JSON. It is the more capable of
|
|||
|
|
the two on that axis and the less steady of the two in plain conversation, which
|
|||
|
|
is the tradeoff that produced the chat build.
|
|||
|
|
|
|||
|
|
`debug/train.log` holds the complete training history for both.
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|||
|
|
m = AutoModelForCausalLM.from_pretrained("kandivault/sprocket-500m")
|
|||
|
|
t = AutoTokenizer.from_pretrained("kandivault/sprocket-500m")
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### Chat format
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
<|user|>your message<|end|><|assistant|>
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Special tokens: `<|system|>` `<|user|>` `<|assistant|>` `<|end|>`
|
|||
|
|
`<|tool_call|>` `<|tool_result|>` `<think>` `</think>` `<|memory_read|>`
|
|||
|
|
`<|memory_write|>`.
|
|||
|
|
|
|||
|
|
The persona needs **no** system prompt. It is the unconditional default. A
|
|||
|
|
system prompt is for behavioural modifiers (length, tone, format) only.
|
|||
|
|
|
|||
|
|
## Data
|
|||
|
|
|
|||
|
|
- **Pretrain:** [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) (ODC-By)
|
|||
|
|
- **Instruct:** 21,371 synthetic conversations generated with Claude
|
|||
|
|
- **Safety prompts:** [LibrAI/do-not-answer](https://huggingface.co/datasets/LibrAI/do-not-answer)
|
|||
|
|
(Apache-2.0). The risky prompts are theirs; only the responses are ours. No
|
|||
|
|
harmful prompts were self-generated.
|
|||
|
|
|
|||
|
|
## Where the rest of it is
|
|||
|
|
|
|||
|
|
**[Interactive walkthrough](https://kandivault-source.github.io/Sprocket-500M/)**
|
|||
|
|
Four sections, everything running client-side: type into the real 32,000-entry
|
|||
|
|
tokenizer and watch text split into the ids this model was trained on; adjust the
|
|||
|
|
architecture and see the parameter count and memory move; read the actual loss
|
|||
|
|
and throughput curves from the run; and work out what a given size and token
|
|||
|
|
budget costs.
|
|||
|
|
|
|||
|
|
**[Source and training code](https://github.com/kandivault-source/Sprocket-500M)**
|
|||
|
|
The tokenizer training, the model, the training and fine-tuning loops, the corpus
|
|||
|
|
builder, the export with its parity check, and the single script that ran the
|
|||
|
|
whole thing unattended on a rented GPU. Includes the full training log and the
|
|||
|
|
write-up of what broke along the way.
|
|||
|
|
|
|||
|
|
Apache-2.0. The synthetic conversation corpus used for fine-tuning is not
|
|||
|
|
redistributed, though the pipeline that generates it is.
|