127 lines
2.9 KiB
Markdown
127 lines
2.9 KiB
Markdown
---
|
|
language:
|
|
- fa
|
|
license: apache-2.0
|
|
base_model: Qwen/Qwen2.5-1.5B-Instruct
|
|
pipeline_tag: text-generation
|
|
library_name: transformers
|
|
tags:
|
|
- persian
|
|
- qwen
|
|
- qlora
|
|
- legal-nlp
|
|
- regulatory-qa
|
|
- water-regulations
|
|
- instruction-tuning
|
|
---
|
|
|
|
# Persian Water Regulations Qwen2.5-1.5B
|
|
|
|
A Persian domain-adapted language model for answering questions about the
|
|
"Operational By-law and General Conditions of Water and Wastewater Tariffs".
|
|
|
|
## Base model
|
|
|
|
`Qwen/Qwen2.5-1.5B-Instruct`
|
|
|
|
## Training data
|
|
|
|
- Records: 274
|
|
- Training records: 246
|
|
- Validation records: 28
|
|
- Format: conversational prompt-completion JSONL
|
|
- Split: group-aware by source section to reduce source leakage
|
|
|
|
## Training method
|
|
|
|
- Method: QLoRA / supervised fine-tuning
|
|
- Quantization: 4-bit NF4 with double quantization
|
|
- LoRA rank: 16
|
|
- LoRA alpha: 32
|
|
- LoRA dropout: 0.05
|
|
- Target modules: `all-linear`
|
|
- Learning rate: 1e-04
|
|
- Requested epochs: 10
|
|
- Maximum sequence length: 512
|
|
- GPU: Tesla T4
|
|
|
|
## Validation generation metrics
|
|
|
|
### Base model
|
|
|
|
- Exact match: 0.0000
|
|
- Token F1: 0.2044
|
|
- ROUGE-L F1: 0.1670
|
|
|
|
### Fine-tuned model
|
|
|
|
- Exact match: 0.0000
|
|
- Token F1: 0.3623
|
|
- ROUGE-L F1: 0.3196
|
|
|
|
The automatic metrics measure textual similarity and do not establish legal
|
|
correctness. Manual review is required for authoritative use.
|
|
|
|
## Usage
|
|
|
|
```python
|
|
import torch
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|
|
|
model_id = "Hesam760/persian-water-regulations-qwen2.5-1.5b"
|
|
|
|
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|
model_id,
|
|
device_map="auto",
|
|
dtype=torch.float16,
|
|
)
|
|
|
|
messages = [
|
|
{
|
|
"role": "system",
|
|
"content": (
|
|
"شما دستیار تخصصی آییننامه عملیاتی و شرایط عمومی تعرفههای "
|
|
"آب و فاضلاب هستید. فقط بر اساس مفاد آییننامه پاسخ دهید."
|
|
),
|
|
},
|
|
{
|
|
"role": "user",
|
|
"content": "هدف آییننامه چیست؟",
|
|
},
|
|
]
|
|
|
|
inputs = tokenizer.apply_chat_template(
|
|
messages,
|
|
tokenize=True,
|
|
add_generation_prompt=True,
|
|
return_dict=True,
|
|
return_tensors="pt",
|
|
).to(model.device)
|
|
|
|
output_ids = model.generate(
|
|
**inputs,
|
|
max_new_tokens=160,
|
|
do_sample=False,
|
|
)
|
|
|
|
new_tokens = output_ids[0, inputs["input_ids"].shape[-1]:]
|
|
print(tokenizer.decode(new_tokens, skip_special_tokens=True))
|
|
```
|
|
|
|
## Intended use
|
|
|
|
- Demonstration of Persian domain adaptation
|
|
- Regulatory question answering
|
|
- Educational and research use
|
|
|
|
## Limitations
|
|
|
|
- The model is trained on a small, single-document dataset.
|
|
- It can reproduce outdated, incomplete, or incorrectly extracted provisions.
|
|
- It may hallucinate provisions or section numbers.
|
|
- It is not a legal authority and must not be used as a substitute for the
|
|
official regulation or professional advice.
|
|
- Verify the redistribution rights of the source-derived training dataset
|
|
before publishing the dataset itself.
|