119 lines
5.1 KiB
Markdown
119 lines
5.1 KiB
Markdown
|
|
---
|
|||
|
|
license: apache-2.0
|
|||
|
|
language:
|
|||
|
|
- en
|
|||
|
|
library_name: transformers
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
tags:
|
|||
|
|
- llama
|
|||
|
|
- instruct
|
|||
|
|
- chat
|
|||
|
|
- conversational
|
|||
|
|
- small-language-model
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# TinyBrainBot 303M — Instruct
|
|||
|
|
|
|||
|
|
A **303M-parameter chat assistant** built from scratch on a home server (2× NVIDIA Tesla P100): pretrained → fact-distilled → supervised fine-tuned. It's a tiny, honest little assistant — it answers common questions, follows simple instructions, holds a short conversation, and (often) **admits when it doesn't know** instead of bluffing.
|
|||
|
|
|
|||
|
|
Base (pretrained-only) version: **TinyBrainBot 303M Base**.
|
|||
|
|
|
|||
|
|
## Model details
|
|||
|
|
|
|||
|
|
| | |
|
|||
|
|
|---|---|
|
|||
|
|
| Parameters | ~303M |
|
|||
|
|
| Architecture | LLaMA-style (`LlamaForCausalLM`) — RoPE, RMSNorm, SwiGLU, GQA |
|
|||
|
|
| Layers / hidden / heads | 24 / 1024 / 16 (4 KV heads) |
|
|||
|
|
| FFN / vocab / context | 2816 / 32,000 / 1024 |
|
|||
|
|
| Tied embeddings | Yes |
|
|||
|
|
| Special tokens | `<\|user\|>` `<\|assistant\|>` `<\|system\|>` `<\|end\|>` |
|
|||
|
|
| EOS token | `<\|end\|>` |
|
|||
|
|
|
|||
|
|
## ⚠️ Chat template — use SPACES, not newlines
|
|||
|
|
|
|||
|
|
This is the single most important detail. The tokenizer normalizes newlines to spaces, so the model was trained with **spaces** between turns. The bundled `chat_template` already does this — **use `apply_chat_template`** and it just works:
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|||
|
|
import torch
|
|||
|
|
|
|||
|
|
tok = AutoTokenizer.from_pretrained("your-username/tinybrainbot-303m-instruct")
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained("your-username/tinybrainbot-303m-instruct", torch_dtype=torch.float16).eval()
|
|||
|
|
|
|||
|
|
msgs = [{"role": "user", "content": "What is the capital of France?"}]
|
|||
|
|
prompt = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
|
|||
|
|
ids = tok(prompt, return_tensors="pt", add_special_tokens=False).input_ids
|
|||
|
|
out = model.generate(ids, max_new_tokens=64, do_sample=True,
|
|||
|
|
temperature=0.5, top_p=0.9, repetition_penalty=1.2, eos_token_id=7)
|
|||
|
|
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))
|
|||
|
|
# -> Paris.
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
The rendered format is (note the **spaces**, no `\n`):
|
|||
|
|
```
|
|||
|
|
<|user|> {message} <|end|> <|assistant|>
|
|||
|
|
```
|
|||
|
|
If you build prompts by hand or configure a UI (llama.cpp / Ollama / LM Studio / Jan), **make sure the separators are spaces and the stop token is `<|end|>`** — a template with literal `\n` will feed the model out-of-distribution tokens and it will ramble without stopping.
|
|||
|
|
|
|||
|
|
## Recommended sampling
|
|||
|
|
|
|||
|
|
Small models need a **lower temperature** than large ones (temp 1.0 makes this model incoherent).
|
|||
|
|
|
|||
|
|
| Use case | temp | top-p | rep penalty |
|
|||
|
|
|---|---|---|---|
|
|||
|
|
| General chat (default) | 0.7 | 0.9 | 1.2 |
|
|||
|
|
| More creative | 0.8 | 0.9 | 1.2 |
|
|||
|
|
| Factual / reliable | 0.4–0.6 | 0.9 | 1.15 |
|
|||
|
|
|
|||
|
|
## Training
|
|||
|
|
|
|||
|
|
- **Base:** pretrained from scratch on FineWeb-Edu, Wikipedia, TinyStories, OpenWebText2, Orca-Math.
|
|||
|
|
- **Fact distillation:** synthetic Q&A + short-fact datasets generated by a stronger teacher (Ling 2.6 Flash), mixed into continued pretraining to densify factual recall.
|
|||
|
|
- **SFT:** ~33k instruction→response pairs plus short 2-turn conversations, distilled from the same teacher, spanning: general instructions, concise Q&A, explain/summarize/rewrite/list/compare/how-to, short creative writing, honest **"I don't know"** examples, and polite refusals. Loss masked to the assistant turns only.
|
|||
|
|
|
|||
|
|
## Evaluation
|
|||
|
|
|
|||
|
|
Standard multiple-choice benchmarks (log-likelihood scoring, n=200 subsample — **absolute numbers run a bit low vs. the official lm-eval-harness**; use them for relative comparison):
|
|||
|
|
|
|||
|
|
| Benchmark | TinyBrainBot 303M | random |
|
|||
|
|
|---|---|---|
|
|||
|
|
| ARC-Easy (acc_norm) | ~44% | 25 |
|
|||
|
|
| OpenBookQA (acc_norm) | ~31% | 25 |
|
|||
|
|
| HellaSwag (acc_norm) | ~29% | 25 |
|
|||
|
|
| WinoGrande (acc) | ~46% | 50 |
|
|||
|
|
| ARC-Challenge (acc_norm) | ~24% | 25 |
|
|||
|
|
| MMLU (acc) | ~23% | 25 |
|
|||
|
|
|
|||
|
|
Roughly **GPT-2-small tier** on general reasoning, with a relative edge on knowledge QA (ARC / OpenBookQA) from the fact distillation. Near-random on MMLU, as expected at this scale.
|
|||
|
|
|
|||
|
|
## Sampling
|
|||
|
|
|
|||
|
|
(best of 5 samples, temperature tuned per prompt):
|
|||
|
|
|
|||
|
|
Q: What is the largest planet in the solar system?
|
|||
|
|
Jupiter.
|
|||
|
|
|
|||
|
|
Q: Explain what photosynthesis is in one or two simple sentences.
|
|||
|
|
Photosynthesis is how plants use sunlight, water, and air to make food (glucose) and release oxygen. It occurs in many types like green plants, algae, and trees.
|
|||
|
|
|
|||
|
|
Q: Write a short poem about the moon.
|
|||
|
|
Silver moon, slow and gold, hums low over the sleeping sky. It dreams in silver rivers, guarding dreams in morning light.
|
|||
|
|
|
|||
|
|
Q: List three healthy breakfast ideas.
|
|||
|
|
Greek yogurt with berries, whole wheat bread, cream cheese.
|
|||
|
|
|
|||
|
|
Q: Give me one tip for staying focused while studying.
|
|||
|
|
Use 25-minute focus blocks with 5-minute breaks, silence notifications, and a clear workspace to keep your mind fresh.
|
|||
|
|
|
|||
|
|
## Limitations
|
|||
|
|
|
|||
|
|
- **Fragile facts.** Sensitive to phrasing and capitalization; standard, well-formed questions work best. Confidently wrong on the long tail — pair with **retrieval (RAG)** for anything important.
|
|||
|
|
- **Weak reasoning/math** — it's 303M.
|
|||
|
|
- The "I don't know" and refusal behaviors are **helpful but not 100% reliable** (they were a small slice of SFT).
|
|||
|
|
- English only.
|
|||
|
|
|
|||
|
|
## License
|
|||
|
|
|
|||
|
|
Apache-2.0 *(change if you prefer)*.
|