Files
particle-1.0/README.md
ModelHub XC ec188ec61d 初始化项目,由ModelHub XC社区提供模型
Model: prathamkode/particle-1.0
Source: Original Platform
2026-09-25 00:49:22 +08:00

2.2 KiB

license, language, library_name, pipeline_tag, tags, datasets
license language library_name pipeline_tag tags datasets
mit
en
transformers text-generation
llama
from-scratch
smol
HuggingFaceTB/smol-smoltalk
HuggingFaceFW/fineweb_edu_100BT-shuffled

particle-1.0

~100M-parameter Llama-style chat model trained from scratch (random init). Not a fine-tune of Llama, SmolLM, or any Hub base.

Weights are MIT. Training data still needs attribution (below).

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "prathamkode/particle-1.0"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)

messages = [{"role": "user", "content": "hello"}]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
ids = tok(prompt, return_tensors="pt")
out = model.generate(**ids, max_new_tokens=64, temperature=0.7)
print(tok.decode(out[0], skip_special_tokens=False))

Chat format:

<|user|>
hello
<|assistant|>

Model details

Architecture Llama-style decoder (RoPE, SwiGLU, RMSNorm, tied embeddings)
Parameters ~100M (12 layers, 768 hidden, 12 heads)
Context 2048 tokens
Tokenizer Custom 32k byte-level BPE (not Llama / GPT-2 vocab)
Init Random N(0, 0.02) — trained from scratch
Precision BF16 training; Hub weights bfloat16

Training

  1. Tokenizer trained from scratch on a FineWeb-Edu sample (~2GB text).
  2. Pretrain next-token prediction on HuggingFaceFW/fineweb_edu_100BT-shuffled, first ~2B tokens.
  3. SFT on HuggingFaceTB/smol-smoltalk (first user/assistant turn + a few greeting seeds).

SFT used that dataset as text only. No teacher model weights were copied.

Intended use

Research / demo small chat model. Expect short replies, mistakes, and weak reasoning.

Limitations

  • Very small capacity
  • May hallucinate
  • English-centric FineWeb-Edu subset
  • No RLHF / preference tuning

License

  • These weights: MIT
  • FineWeb-Edu: ODC-By (attribute)
  • smol-smoltalk: follow the dataset card