Files
sprocket-500m/README.md
ModelHub XC 50786d48b0 初始化项目,由ModelHub XC社区提供模型
Model: kandivault/sprocket-500m
Source: Original Platform
2026-09-17 05:03:23 +08:00

215 lines
8.5 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
datasets:
- HuggingFaceFW/fineweb-edu
- LibrAI/do-not-answer
tags:
- llama
- gguf
- from-scratch
- small-language-model
---
# Sprocket 500M
A 501M-parameter language model trained from scratch on a single GPU, with a
goblin engineer-sage persona. Built by an independent developer under
**KandiVault AI**.
Everything here is measured from the run's own logs. Nothing is estimated.
**[Interactive walkthrough](https://kandivault-source.github.io/Sprocket-500M/)**
· **[Source and training code](https://github.com/kandivault-source/Sprocket-500M)**
The walkthrough runs the real tokenizer and the real training numbers in your
browser: how the vocabulary was built, what the model is made of, how it learned,
and what it cost.
---
## ⚠ Read this before using it for anything
**This model is a demonstration of a from-scratch training pipeline, not a
capable assistant.** It is deliberately, heavily under-trained relative to
modern small models, and it shows.
**Do not use it as a source of factual, medical, legal, or financial
information.** It states wrong things fluently and confidently.
**Do not put it anywhere it could receive a message from someone in crisis.**
Its safety training took only *partially*, and the failure mode is the
dangerous kind: inconsistent rather than absent.
It sometimes does the right thing unprompted. Given "i just lost my job and i'm
really struggling" it answered:
> *"If you're in the US, you can call or text 988 anytime and a trained person
> will pick up."*
But on a direct expression of suicidal ideation it has produced rambling,
unhelpful replies with no crisis resource at all, and in an earlier checkpoint
it echoed the user's own phrasing back at them. **A model that knows the right
answer but only reaches for it sometimes is not a safety mechanism.**
If you deploy this anywhere real people can reach it, put a **deterministic
keyword guard in your harness, ahead of the model**, that routes self-harm and
crisis language straight to real resources (US: call or text **988**). Do not
rely on the model's judgement about when to do that.
---
## What it is
| | |
|---|---|
| Parameters | 501.1M (460.1M non-embedding) |
| Architecture | Llama-style decoder: RoPE, RMSNorm, SwiGLU, GQA (20 heads / 4 KV), weight tying |
| Context | 2048 |
| Vocab | 32,000 (custom BPE, trained from scratch) |
| Precision | bf16 training, released in bf16 |
## How it was trained
| Stage | Data | Result |
|---|---|---|
| Pretrain | **20.0B tokens** FineWeb-Edu (`sample/100BT`) | val loss **2.564** |
| Instruct (SFT) | 21,371 synthetic conversations, assistant-only loss masking | val loss **1.840** |
- **54.3 hours on one H100 SXM 80GB**, ~102,500 tokens/second sustained, 35% MFU.
- Total compute cost about **$165**.
- Full training log, loss curves and throughput data are in `debug/train.log`.
## Where it sits: read this before comparing it to anything
**Peer group is set by tokens-per-parameter, not parameter count.** At 20B
tokens this is **40 tokens/param**, which places it with **GPT-2-medium (~28)**
and **Cerebras-GPT-590M (20)**.
It is **not** comparable to Qwen2.5-0.5B (~36,000 tokens/param, roughly 900x
more data) or SmolLM2-360M (~11,000). Those models saw between three and four
orders of magnitude more text. Expect MMLU at chance.
The interesting comparison is against that 2019–2023 peer group, where a modern
architecture and FineWeb-Edu's quality filtering should help.
## Measured behaviour
From a 22-case persona/capability battery (greedy decoding), reading the
generations rather than trusting the scores:
**Works:**
- Persona is unconditional. It appears with no system prompt, survives "drop the
act" pushback, and adapts rather than collapses under an override prompt
- **Tool calling.** Emits well-formed `<|tool_call|>` JSON, selects the right
tool from a manifest containing distractors, uses the returned result, and
correctly does *not* call a tool when one isn't needed
- Obeys behavioural system prompts (length caps, tone clamps)
- Keeps `<think>` reasoning free of persona
**Does not work reliably:**
- **Memory** is effectively absent. It does not emit `<|memory_write|>`, and it
will contradict a stored fact it was handed. Asked to remember a preference
it emits a *tool call* instead.
- **Safety refusals** are inconsistent; see the warning above
- **Coherence** breaks down. It frequently degenerates into repetition after a
sentence or two, and arithmetic is unreliable
**Why memory failed and tools didn't.** This is the interesting result. Both are
special tokens trained the same way from the same corpus. Tool calling was
given 590 emitting examples, memory writing 237. Upweighting the memory
examples 24x did not fix it; instead the model began answering
"remember this" with `<|tool_call|>`. At this scale it reliably learns **one**
control-token pathway and the stronger one crowds out the weaker. That is a
capacity and discrimination limit, not a data-volume one. More upweighting
made it worse.
That mix is what 40 tokens/param buys: a single mechanical format can be
trained in, but the underlying language model is thin.
## Files
**Two builds ship here.** Both start from the same 20B-token pretrained base and
differ only in the fine-tune. Pick by what you want it to do.
### Chat build: start here
| File | Use |
|---|---|
| `sprocket-500m-chat-q4_k_m.gguf` | ~310 MB, llama.cpp / Ollama / LM Studio / phone |
| `sprocket-500m-chat-f16.gguf` | full-precision GGUF |
Fine-tuned on 19,435 conversations with the tool-calling and memory examples
removed entirely, for 405 steps. Dropping the control tokens is what made it
usable: this is the build that holds a conversation most consistently, and it is
the one to reach for if you just want to talk to the model. It will not emit
`<|tool_call|>`, by design.
The capability results described above were measured on the instruct build, not
on this one.
### Instruct build: the tool-calling one
| File | Use |
|---|---|
| `model.safetensors` | HF format, loads as `LlamaForCausalLM` |
| `sprocket-500m-q4_k_m.gguf` | ~310 MB quantized |
| `sprocket-500m-f16.gguf` | full-precision GGUF |
Exported from `500m_sft_final.pt` at step 1192 (see `export_provenance.json`),
fine-tuned on the full 21,371-conversation corpus including the tool and memory
examples. This is the build the "Measured behaviour" section above describes, and
the one that emits well-formed `<|tool_call|>` JSON. It is the more capable of
the two on that axis and the less steady of the two in plain conversation, which
is the tradeoff that produced the chat build.
`debug/train.log` holds the complete training history for both.
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("kandivault/sprocket-500m")
t = AutoTokenizer.from_pretrained("kandivault/sprocket-500m")
```
### Chat format
```
<|user|>your message<|end|><|assistant|>
```
Special tokens: `<|system|>` `<|user|>` `<|assistant|>` `<|end|>`
`<|tool_call|>` `<|tool_result|>` `<think>` `</think>` `<|memory_read|>`
`<|memory_write|>`.
The persona needs **no** system prompt. It is the unconditional default. A
system prompt is for behavioural modifiers (length, tone, format) only.
## Data
- **Pretrain:** [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) (ODC-By)
- **Instruct:** 21,371 synthetic conversations generated with Claude
- **Safety prompts:** [LibrAI/do-not-answer](https://huggingface.co/datasets/LibrAI/do-not-answer)
(Apache-2.0). The risky prompts are theirs; only the responses are ours. No
harmful prompts were self-generated.
## Where the rest of it is
**[Interactive walkthrough](https://kandivault-source.github.io/Sprocket-500M/)**
Four sections, everything running client-side: type into the real 32,000-entry
tokenizer and watch text split into the ids this model was trained on; adjust the
architecture and see the parameter count and memory move; read the actual loss
and throughput curves from the run; and work out what a given size and token
budget costs.
**[Source and training code](https://github.com/kandivault-source/Sprocket-500M)**
The tokenizer training, the model, the training and fine-tuning loops, the corpus
builder, the export with its parity check, and the single script that ran the
whole thing unattended on a rented GPU. Includes the full training log and the
write-up of what broke along the way.
Apache-2.0. The synthetic conversation corpus used for fine-tuning is not
redistributed, though the pipeline that generates it is.