Files
qwen3-0.6b-protein-vocab-v0/README.md
ModelHub XC 88254424c2 初始化项目,由ModelHub XC社区提供模型
Model: khairi/qwen3-0.6b-protein-vocab-v0
Source: Original Platform
2026-08-13 21:19:17 +08:00

35 lines
1.1 KiB
Markdown

---
base_model: Qwen/Qwen3-0.6B
tags:
- protein
- tokenizer-extension
- not-yet-trained
---
# qwen3-0.6b-protein-vocab-v0
`Qwen/Qwen3-0.6B` with its tokenizer extended by a protein-sequence BPE
vocabulary (+9,568 new tokens plus `<protein>`/`</protein>` special tokens),
and input/output embeddings resized to match (new rows mean/covariance-
initialized from the existing embedding distribution).
**This checkpoint has not been trained on any protein data.** It's a
starting point ("V0") for LoRA continued-pretraining, not a usable protein
language model yet. All of Qwen3's original pretrained weights and special
tokens (`<|im_start|>`, `<|im_end|>`, `<think>`, etc.) are unchanged and
keep their original ids.
Built with [`eshmun-vocab`](https://github.com/) via:
```bash
python scripts/extend_tokenizer.py \
--protein-tokenizer data/tokenizers/tokenizer_c100_10k/tokenizer.json \
--llm-tokenizer Qwen/Qwen3-0.6B \
--output data/tokenizers/qwen3_0.6b_protein_extended
python scripts/build_vocab_extended_model.py \
--base-model Qwen/Qwen3-0.6B \
--tokenizer-path data/tokenizers/qwen3_0.6b_protein_extended \
--output models/Qwen3-0.6B-V0
```