--- base_model: Qwen/Qwen3-0.6B tags: - protein - tokenizer-extension - not-yet-trained --- # qwen3-0.6b-protein-vocab-v0 `Qwen/Qwen3-0.6B` with its tokenizer extended by a protein-sequence BPE vocabulary (+9,568 new tokens plus ``/`` special tokens), and input/output embeddings resized to match (new rows mean/covariance- initialized from the existing embedding distribution). **This checkpoint has not been trained on any protein data.** It's a starting point ("V0") for LoRA continued-pretraining, not a usable protein language model yet. All of Qwen3's original pretrained weights and special tokens (`<|im_start|>`, `<|im_end|>`, ``, etc.) are unchanged and keep their original ids. Built with [`eshmun-vocab`](https://github.com/) via: ```bash python scripts/extend_tokenizer.py \ --protein-tokenizer data/tokenizers/tokenizer_c100_10k/tokenizer.json \ --llm-tokenizer Qwen/Qwen3-0.6B \ --output data/tokenizers/qwen3_0.6b_protein_extended python scripts/build_vocab_extended_model.py \ --base-model Qwen/Qwen3-0.6B \ --tokenizer-path data/tokenizers/qwen3_0.6b_protein_extended \ --output models/Qwen3-0.6B-V0 ```