base_model, tags
base_model tags
Qwen/Qwen3-0.6B
protein
tokenizer-extension
not-yet-trained

qwen3-0.6b-protein-vocab-v0

Qwen/Qwen3-0.6B with its tokenizer extended by a protein-sequence BPE vocabulary (+9,568 new tokens plus <protein>/</protein> special tokens), and input/output embeddings resized to match (new rows mean/covariance- initialized from the existing embedding distribution).

This checkpoint has not been trained on any protein data. It's a starting point ("V0") for LoRA continued-pretraining, not a usable protein language model yet. All of Qwen3's original pretrained weights and special tokens (<|im_start|>, <|im_end|>, <think>, etc.) are unchanged and keep their original ids.

Built with eshmun-vocab via:

python scripts/extend_tokenizer.py \
  --protein-tokenizer data/tokenizers/tokenizer_c100_10k/tokenizer.json \
  --llm-tokenizer Qwen/Qwen3-0.6B \
  --output data/tokenizers/qwen3_0.6b_protein_extended

python scripts/build_vocab_extended_model.py \
  --base-model Qwen/Qwen3-0.6B \
  --tokenizer-path data/tokenizers/qwen3_0.6b_protein_extended \
  --output models/Qwen3-0.6B-V0
Description
Model synced from source: khairi/qwen3-0.6b-protein-vocab-v0
Readme 27 KiB
Languages
Jinja 100%