35 lines
1.1 KiB
Markdown
35 lines
1.1 KiB
Markdown
---
|
|
base_model: Qwen/Qwen3-0.6B
|
|
tags:
|
|
- protein
|
|
- tokenizer-extension
|
|
- not-yet-trained
|
|
---
|
|
|
|
# qwen3-0.6b-protein-vocab-v0
|
|
|
|
`Qwen/Qwen3-0.6B` with its tokenizer extended by a protein-sequence BPE
|
|
vocabulary (+9,568 new tokens plus `<protein>`/`</protein>` special tokens),
|
|
and input/output embeddings resized to match (new rows mean/covariance-
|
|
initialized from the existing embedding distribution).
|
|
|
|
**This checkpoint has not been trained on any protein data.** It's a
|
|
starting point ("V0") for LoRA continued-pretraining, not a usable protein
|
|
language model yet. All of Qwen3's original pretrained weights and special
|
|
tokens (`<|im_start|>`, `<|im_end|>`, `<think>`, etc.) are unchanged and
|
|
keep their original ids.
|
|
|
|
Built with [`eshmun-vocab`](https://github.com/) via:
|
|
|
|
```bash
|
|
python scripts/extend_tokenizer.py \
|
|
--protein-tokenizer data/tokenizers/tokenizer_c100_10k/tokenizer.json \
|
|
--llm-tokenizer Qwen/Qwen3-0.6B \
|
|
--output data/tokenizers/qwen3_0.6b_protein_extended
|
|
|
|
python scripts/build_vocab_extended_model.py \
|
|
--base-model Qwen/Qwen3-0.6B \
|
|
--tokenizer-path data/tokenizers/qwen3_0.6b_protein_extended \
|
|
--output models/Qwen3-0.6B-V0
|
|
```
|