--- language: - my license: apache-2.0 base_model: URajinda/Qwen2.5-MM-1.5B-Base tags: - burmese - myanmar - shweyon - qwen - tokenization library_name: transformers --- # πŸš€ ShweYon-Qwen2.5-Burmese-1.5B-v1.2(Enhanced Burmese LLM) **ShweYon-v1.2-Base** is a specialized language model based on the Qwen2.5-1.5B architecture, meticulously optimized for the Myanmar (Burmese) language. This version features a significant **Vocabulary Expansion** designed to solve common tokenization inefficiencies in Burmese NLP. ### πŸ“Š Technical Specifications | Feature | Specification | | :--- | :--- | | **Base Architecture** | Qwen2.5-1.5B | | **New Vocab Size** | 152,858 | | **Added Tokens** | 1,418 (Cumulative) | | **Model Size Increase** | ~4.73 MB (+0.40% total weight) | | **Training Type** | Continual Pre-training (CPT) | | **Language** | Myanmar (Burmese) | ### πŸ’» Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_name = "URajinda/ShweYon-Qwen2.5-Burmese-1.5B-v1.2" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained(model_name) text = "α€™α€Όα€”α€Ία€™α€¬α€”α€­α€―α€„α€Ία€„α€Άα€žα€Šα€Ί" inputs = tokenizer(text, return_tensors="pt") outputs = model.generate(**inputs, max_new_tokens=50) print(tokenizer.decode(outputs[0], skip_special_tokens=True))