68acf380bc1681cff40799aa7d18ce874e5bddfc
Model: shaohongwu/Qwen2.5-0.5B-Preweb-special-tokens Source: Original Platform
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen2.5-0.5B/blob/main/LICENSE language:
- en
pipeline_tag: text-generation
library_name: transformers
license: apache-2.0
base_model: Qwen/Qwen2.5-0.5B
tags:
- qwen
- schema-aware
- structured-output
- preweb
Qwen2.5-0.5B-Preweb-Base
This model is a derivative base model of Qwen/Qwen2.5-0.5B.
It extends the tokenizer vocabulary with schema/control special tokens
🔧 What’s Changed
Added the following special tokens to the tokenizer:
- `<|domain_start|>`, `<|domain_end|>`
- `<|intent_start|>`, `<|intent_end|>`
- `<|slot_type_start|>`, `<|slot_type_end|>`
- `<|slot_span_start|>`, `<|slot_span_end|>`
- `<|canonical_start|>`, `<|canonical_end|>`
As a result: - Tokenizer vocabulary size is increased.
🎯 Intended Usage
This model is intended to be used as a base model for:
- schema-aware prompting
- structured information extraction
- slot / intent / domain prediction
All downstream LoRA adapters must be trained with the same tokenizer.
🚫 Compatibility Notes
- Vocabulary shapes are **fixed**
- Do NOT add or remove tokens at runtime
- Designed for **vLLM / TensorRT-LLM** serving
- Safe for multi-LoRA dynamic loading
📜 License
Apache-2.0
This model is a derivative work of Qwen/Qwen2.5-0.5B.
Description