---
base_model: Qwen/Qwen3-0.6B
tags:
- protein
- tokenizer-extension
- not-yet-trained
---
# qwen3-0.6b-protein-vocab-v0
`Qwen/Qwen3-0.6B` with its tokenizer extended by a protein-sequence BPE
vocabulary (+9,568 new tokens plus ``/`` special tokens),
and input/output embeddings resized to match (new rows mean/covariance-
initialized from the existing embedding distribution).
**This checkpoint has not been trained on any protein data.** It's a
starting point ("V0") for LoRA continued-pretraining, not a usable protein
language model yet. All of Qwen3's original pretrained weights and special
tokens (`<|im_start|>`, `<|im_end|>`, ``, etc.) are unchanged and
keep their original ids.
Built with [`eshmun-vocab`](https://github.com/) via:
```bash
python scripts/extend_tokenizer.py \
--protein-tokenizer data/tokenizers/tokenizer_c100_10k/tokenizer.json \
--llm-tokenizer Qwen/Qwen3-0.6B \
--output data/tokenizers/qwen3_0.6b_protein_extended
python scripts/build_vocab_extended_model.py \
--base-model Qwen/Qwen3-0.6B \
--tokenizer-path data/tokenizers/qwen3_0.6b_protein_extended \
--output models/Qwen3-0.6B-V0
```