Model: khairi/qwen3-0.6b-protein-vocab-v0 Source: Original Platform
base_model, tags
| base_model | tags | |||
|---|---|---|---|---|
| Qwen/Qwen3-0.6B |
|
qwen3-0.6b-protein-vocab-v0
Qwen/Qwen3-0.6B with its tokenizer extended by a protein-sequence BPE
vocabulary (+9,568 new tokens plus <protein>/</protein> special tokens),
and input/output embeddings resized to match (new rows mean/covariance-
initialized from the existing embedding distribution).
This checkpoint has not been trained on any protein data. It's a
starting point ("V0") for LoRA continued-pretraining, not a usable protein
language model yet. All of Qwen3's original pretrained weights and special
tokens (<|im_start|>, <|im_end|>, <think>, etc.) are unchanged and
keep their original ids.
Built with eshmun-vocab via:
python scripts/extend_tokenizer.py \
--protein-tokenizer data/tokenizers/tokenizer_c100_10k/tokenizer.json \
--llm-tokenizer Qwen/Qwen3-0.6B \
--output data/tokenizers/qwen3_0.6b_protein_extended
python scripts/build_vocab_extended_model.py \
--base-model Qwen/Qwen3-0.6B \
--tokenizer-path data/tokenizers/qwen3_0.6b_protein_extended \
--output models/Qwen3-0.6B-V0
Description
Languages
Jinja
100%