Files
ModelHub XC 68ec45eaf7 初始化项目,由ModelHub XC社区提供模型
Model: nkthebass/TinyBrainBot-303m-instruct
Source: Original Platform
2026-08-03 14:31:16 +08:00

119 lines
5.1 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- llama
- instruct
- chat
- conversational
- small-language-model
---
# TinyBrainBot 303M — Instruct
A **303M-parameter chat assistant** built from scratch on a home server (2× NVIDIA Tesla P100): pretrained → fact-distilled → supervised fine-tuned. It's a tiny, honest little assistant — it answers common questions, follows simple instructions, holds a short conversation, and (often) **admits when it doesn't know** instead of bluffing.
Base (pretrained-only) version: **TinyBrainBot 303M Base**.
## Model details
| | |
|---|---|
| Parameters | ~303M |
| Architecture | LLaMA-style (`LlamaForCausalLM`) — RoPE, RMSNorm, SwiGLU, GQA |
| Layers / hidden / heads | 24 / 1024 / 16 (4 KV heads) |
| FFN / vocab / context | 2816 / 32,000 / 1024 |
| Tied embeddings | Yes |
| Special tokens | `<\|user\|>` `<\|assistant\|>` `<\|system\|>` `<\|end\|>` |
| EOS token | `<\|end\|>` |
## ⚠️ Chat template — use SPACES, not newlines
This is the single most important detail. The tokenizer normalizes newlines to spaces, so the model was trained with **spaces** between turns. The bundled `chat_template` already does this — **use `apply_chat_template`** and it just works:
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("your-username/tinybrainbot-303m-instruct")
model = AutoModelForCausalLM.from_pretrained("your-username/tinybrainbot-303m-instruct", torch_dtype=torch.float16).eval()
msgs = [{"role": "user", "content": "What is the capital of France?"}]
prompt = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
ids = tok(prompt, return_tensors="pt", add_special_tokens=False).input_ids
out = model.generate(ids, max_new_tokens=64, do_sample=True,
temperature=0.5, top_p=0.9, repetition_penalty=1.2, eos_token_id=7)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))
# -> Paris.
```
The rendered format is (note the **spaces**, no `\n`):
```
<|user|> {message} <|end|> <|assistant|>
```
If you build prompts by hand or configure a UI (llama.cpp / Ollama / LM Studio / Jan), **make sure the separators are spaces and the stop token is `<|end|>`** — a template with literal `\n` will feed the model out-of-distribution tokens and it will ramble without stopping.
## Recommended sampling
Small models need a **lower temperature** than large ones (temp 1.0 makes this model incoherent).
| Use case | temp | top-p | rep penalty |
|---|---|---|---|
| General chat (default) | 0.7 | 0.9 | 1.2 |
| More creative | 0.8 | 0.9 | 1.2 |
| Factual / reliable | 0.4–0.6 | 0.9 | 1.15 |
## Training
- **Base:** pretrained from scratch on FineWeb-Edu, Wikipedia, TinyStories, OpenWebText2, Orca-Math.
- **Fact distillation:** synthetic Q&A + short-fact datasets generated by a stronger teacher (Ling 2.6 Flash), mixed into continued pretraining to densify factual recall.
- **SFT:** ~33k instruction→response pairs plus short 2-turn conversations, distilled from the same teacher, spanning: general instructions, concise Q&A, explain/summarize/rewrite/list/compare/how-to, short creative writing, honest **"I don't know"** examples, and polite refusals. Loss masked to the assistant turns only.
## Evaluation
Standard multiple-choice benchmarks (log-likelihood scoring, n=200 subsample — **absolute numbers run a bit low vs. the official lm-eval-harness**; use them for relative comparison):
| Benchmark | TinyBrainBot 303M | random |
|---|---|---|
| ARC-Easy (acc_norm) | ~44% | 25 |
| OpenBookQA (acc_norm) | ~31% | 25 |
| HellaSwag (acc_norm) | ~29% | 25 |
| WinoGrande (acc) | ~46% | 50 |
| ARC-Challenge (acc_norm) | ~24% | 25 |
| MMLU (acc) | ~23% | 25 |
Roughly **GPT-2-small tier** on general reasoning, with a relative edge on knowledge QA (ARC / OpenBookQA) from the fact distillation. Near-random on MMLU, as expected at this scale.
## Sampling
(best of 5 samples, temperature tuned per prompt):
Q: What is the largest planet in the solar system?
Jupiter.
Q: Explain what photosynthesis is in one or two simple sentences.
Photosynthesis is how plants use sunlight, water, and air to make food (glucose) and release oxygen. It occurs in many types like green plants, algae, and trees.
Q: Write a short poem about the moon.
Silver moon, slow and gold, hums low over the sleeping sky. It dreams in silver rivers, guarding dreams in morning light.
Q: List three healthy breakfast ideas.
Greek yogurt with berries, whole wheat bread, cream cheese.
Q: Give me one tip for staying focused while studying.
Use 25-minute focus blocks with 5-minute breaks, silence notifications, and a clear workspace to keep your mind fresh.
## Limitations
- **Fragile facts.** Sensitive to phrasing and capitalization; standard, well-formed questions work best. Confidently wrong on the long tail — pair with **retrieval (RAG)** for anything important.
- **Weak reasoning/math** — it's 303M.
- The "I don't know" and refusal behaviors are **helpful but not 100% reliable** (they were a small slice of SFT).
- English only.
## License
Apache-2.0 *(change if you prefer)*.