--- license: apache-2.0 language: - uk base_model: Qwen/Qwen2.5-14B tags: - legal - ukrainian - continued-pretraining - court-decisions datasets: - overthelex/edrsr-court-decisions library_name: transformers pipeline_tag: text-generation --- # Qwen2.5-14B-EDRSR-Legal-UK Ukrainian legal domain model obtained by continued pretraining (CPT) of [Qwen/Qwen2.5-14B](https://huggingface.co/Qwen/Qwen2.5-14B) on the EDRSR corpus of Ukrainian court decisions. Part of a scaling experiment (0.5B / 1.5B / 3B / 14B) for the PhD dissertation at Glushkov Institute of Cybernetics, NAS of Ukraine. ## Training Data - **Corpus:** Unified State Register of Court Decisions of Ukraine (EDRSR) - **Documents:** 33.9M court decisions (after dedup + quality filtering from 38.5M) - **Tokens:** 161.4B tokens (Qwen2 BPE tokenizer, fertility = 0.515 for Ukrainian legal text) - **Sequence length:** 8,192 tokens - **Shards:** 1,233 pre-packaged numpy shards ## Training Details - **Hardware:** 8x NVIDIA H100 SXM 80GB (NVIDIA Innovation Lab via Brev) - **Framework:** HuggingFace Trainer + DeepSpeed ZeRO-3 - **Precision:** bfloat16 - **Global batch size:** 128 sequences (1.05M tokens/step) - **Total steps:** 9,536 (10B tokens processed) - **Training time:** 44 hours - **Throughput:** 22K tokens/sec, 44.9 sec/step ## Results | Metric | Value | |--------|-------| | Initial loss (step 10) | 0.84 | | Final loss (step 9,536) | 0.22 | | Loss reduction | -74% | | Base perplexity | 2.84 | | **CPT perplexity** | **1.28** | | **Perplexity reduction** | **-54.8%** | ## Scaling Law All four models in the series converge to similar perplexity after CPT: | Model | Base PPL | CPT PPL | Reduction | |-------|----------|---------|-----------| | [0.5B](https://huggingface.co/overthelex/qwen2.5-0.5b-edrsr-legal-uk) | 6.83 | 1.35 | -80% | | [1.5B](https://huggingface.co/overthelex/qwen2.5-1.5b-edrsr-legal-uk) | 4.61 | 1.31 | -72% | | [3B](https://huggingface.co/overthelex/qwen2.5-3b-edrsr-legal-uk) | 3.83 | 1.30 | -66% | | [14B](https://huggingface.co/overthelex/qwen2.5-14b-edrsr-legal-uk) | 2.84 | 1.28 | -55% | ## Intended Use This is a **base model** (not instruction-tuned). It is intended for: - Research on domain adaptation of LLMs for low-resource legal languages - Downstream fine-tuning for Ukrainian legal NLP tasks - Scaling law analysis of continued pretraining - Perplexity evaluation on Ukrainian legal text ## Limitations - Not instruction-tuned; will not follow instructions or chat - Trained on Ukrainian court decisions only; may not generalize to other legal systems ## Related Resources - [EDRSR Court Decisions Dataset](https://huggingface.co/datasets/overthelex/edrsr-court-decisions) - [Tokenizer Fertility Paper (arXiv:2605.14890)](https://arxiv.org/abs/2605.14890) - [Citation Graph Paper (arXiv:2605.15362)](https://arxiv.org/abs/2605.15362) - [Statute Retrieval Paper (arXiv:2605.17639)](https://arxiv.org/abs/2605.17639) - [UA-StatuteRetrieval Benchmark](https://huggingface.co/datasets/overthelex/ua-statute-retrieval)