初始化项目,由ModelHub XC社区提供模型

Model: thangvip/qwen3-1.7b-vietnamese-legal-grpo-phase-2
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-03 12:51:20 +08:00
commit 29c94479d3
13 changed files with 152138 additions and 0 deletions

36
.gitattributes vendored Normal file
View File

@@ -0,0 +1,36 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
tokenizer.json filter=lfs diff=lfs merge=lfs -text

251
README.md Normal file
View File

@@ -0,0 +1,251 @@
---
base_model: thangvip/qwen3-1.7b-vietnamese-legal-grpo
library_name: transformers
model_name: thangvip/qwen3-1.7b-vietnamese-legal-grpo-phase-2
tags:
- grpo
- vietnamese
- legal
- reasoning
- syllogism
- trl
- generated_from_trainer
- phase2
- hard-difficulty
license: apache-2.0
language:
- vi
datasets:
- legal-qa-vietnamese-hard
pipeline_tag: text-generation
widget:
- example_title: "Hard Legal Question Example"
text: "Câu hỏi: Một công ty có nghĩa vụ gì khi sa thải nhân viên do tái cơ cấu?"
---
# Vietnamese Legal Reasoning Model - GRPO Phase 2 (Hard Difficulty)
## 🏛️ Model Description
This is a **Phase 2 Vietnamese legal reasoning specialist** fine-tuned using **Group Relative Policy Optimization (GRPO)** on **hard difficulty** Vietnamese legal question-answering data. This model builds upon the Phase 1 training and is specifically designed to handle more complex **syllogistic reasoning** for challenging Vietnamese legal scenarios.
### 🎯 Base Model
- **Base**: [thangvip/qwen3-1.7b-vietnamese-legal-grpo](https://huggingface.co/thangvip/qwen3-1.7b-vietnamese-legal-grpo)
- **Architecture**: Qwen 3 (1.7B parameters)
- **Language**: Vietnamese
- **Specialization**: Advanced legal reasoning and syllogism (Hard difficulty)
- **Training Phase**: Phase 2 - Hard Level QA
### 🔥 Key Features
✅ **Phase 2 Training**: Advanced model trained on hard difficulty legal questions
✅ **Syllogistic Reasoning**: Structured legal arguments (Major Premise → Minor Premise → Conclusion)
✅ **Vietnamese Legal Domain**: Trained on Vietnamese legal texts and Q&A
✅ **GRPO Optimization**: Advanced policy optimization for better reasoning
✅ **Citation Support**: Generates responses with legal citations
✅ **Structured Output**: Uses XML-like tags for organized responses
✅ **Extended Context**: Supports up to 8192 tokens for complex reasoning chains
## 📊 Model Architecture
- **Parameters**: ~1.7B
- **Vocabulary Size**: 151936
- **Hidden Size**: 2048
- **Layers**: 28
- **Attention Heads**: 16
- **Max Completion Length**: 8192 tokens (extended for complex reasoning)
## 🚀 Quick Start
### Installation
```bash
pip install transformers torch
```
### Basic Usage
```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
# Load model and tokenizer
model_name = "thangvip/qwen3-1.7b-vietnamese-legal-grpo-phase-2"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch.bfloat16,
device_map="auto"
)
# Format your legal question
system_prompt = """Bạn là một chuyên gia pháp lý. Hãy trả lời câu hỏi bằng cách sử dụng phương pháp lập luận tam đoạn luận (syllogism).
Trước tiên, hãy suy nghĩ về vấn đề trong thẻ <think></think>.
Sau đó, trả lời theo định dạng sau:
<answer>
<major_premise>[Quy định pháp luật chung]</major_premise>
<minor_premise>[Sự kiện cụ thể trong câu hỏi]</minor_premise>
<conclusion>[Áp dụng quy định vào sự kiện để đưa ra kết luận]</conclusion>
</answer>
Hãy đảm bảo trích dẫn chính xác các điều luật liên quan."""
question = "Một công ty có nghĩa vụ gì khi sa thải nhân viên do tái cơ cấu?"
# Create conversation
messages = [
{"role": "system", "content": system_prompt},
{"role": "user", "content": question}
]
# Generate response
input_text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(input_text, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=2048, # Extended for hard difficulty complex reasoning
temperature=0.7,
do_sample=True,
pad_token_id=tokenizer.eos_token_id
)
response = tokenizer.decode(outputs[0][inputs['input_ids'].shape[1]:], skip_special_tokens=True)
print(response)
```
### Pipeline Usage
```python
from transformers import pipeline
# Create text generation pipeline
generator = pipeline(
"text-generation",
model="thangvip/qwen3-1.7b-vietnamese-legal-grpo-phase-2",
tokenizer="thangvip/qwen3-1.7b-vietnamese-legal-grpo-phase-2",
torch_dtype=torch.bfloat16,
device_map="auto"
)
# Generate legal reasoning
prompt = "Câu hỏi: Quyền và nghĩa vụ của người thuê nhà khi hợp đồng thuê hết hạn?"
result = generator(prompt, max_new_tokens=512, temperature=0.7)
print(result[0]['generated_text'])
```
## 🎯 Training Details
### Training Procedure
- **Method**: Group Relative Policy Optimization (GRPO)
- **Base Model**: thangvip/qwen3-1.7b-vietnamese-legal-grpo
- **Training Phase**: Phase 2 - Hard Difficulty
- **Training Steps**: N/A (typically 1500 steps)
- **Learning Rate**: N/A
- **Batch Size**: N/A
- **Max Completion Length**: 8192 tokens (doubled for complex reasoning)
### Training Data
- **Domain**: Vietnamese legal question-answering (Hard difficulty)
- **Format**: Syllogistic reasoning pairs
- **Structure**: Question → Structured legal reasoning response
- **Difficulty Level**: Hard - Complex multi-step legal reasoning scenarios
- **Dataset**: `hard_level_qa.jsonl`
### Two-Phase Training Approach
1. **Phase 1**: Initial GRPO training on normal difficulty questions
- Base model: `thangvip/qwen3-4b-legal-pretrain-synthetic-8k` or similar
- Dataset: `normal_level_qa.jsonl`
- Max completion: 4096 tokens
2. **Phase 2** (This Model): Continued training on hard difficulty questions with extended context window
- Base model: Phase 1 trained model (`thangvip/qwen3-1.7b-vietnamese-legal-grpo`)
- Dataset: `hard_level_qa.jsonl`
- Max completion: 8192 tokens (doubled for complex reasoning)
- Training steps: ~1500 steps
This progressive training approach allows the model to first master basic legal reasoning before tackling more complex, multi-step legal problems.
### Reward System
The model was trained with a sophisticated reward system:
- **Correctness** (35%): Factual accuracy against reference answers
- **Format Compliance** (20%): Proper use of syllogistic structure
- **Citation Accuracy** (15%): Relevant and accurate legal citations
- **Reasoning Quality** (15%): Quality of legal reasoning process
- **Hallucination Penalty** (10%): Penalty for unsupported claims
- **Length Penalty** (5%): Penalty for exceeding maximum token length
## 📝 Expected Output Format
The model generates structured responses in this format:
```xml
<think>
[Internal reasoning about the legal question]
</think>
<answer>
<major_premise>
[General legal rule or principle applicable to the situation]
</major_premise>
<minor_premise>
[Specific facts from the question that relate to the legal rule]
</minor_premise>
<conclusion>
[Legal conclusion that follows logically from applying the rule to the facts]
</conclusion>
</answer>
```
## 🎯 Use Cases
- **Complex Legal Education**: Teaching advanced legal reasoning methodology for difficult cases
- **Advanced Legal Research**: Preliminary analysis of complex legal questions
- **Multi-step Legal Analysis**: Structured legal argument generation for intricate scenarios
- **Legal Consultation**: Initial legal guidance for challenging cases (with human review)
- **Legal Training**: Demonstrating proper syllogistic reasoning for complex legal problems
## ⚠️ Limitations
- **Domain Specific**: Optimized for Vietnamese legal context
- **Educational Purpose**: Should not replace professional legal advice
- **Fact Checking Required**: Always verify legal citations and conclusions
- **Extended Generation**: May produce lengthy responses (up to 8192 tokens) for complex questions
- **Phase 2 Training**: Built upon Phase 1 model; requires understanding of base model capabilities
## 📄 Citation
If you use this model, please cite:
```bibtex
@misc{vietnamese-legal-grpo-2024,
title={Vietnamese Legal Reasoning Model with GRPO},
author={Your Name},
year={2024},
publisher={Hugging Face},
url={https://huggingface.co/thangvip/qwen3-1.7b-vietnamese-legal-grpo-phase-2}
}
```
## 🤝 Contributing
Contributions are welcome! Please see our [contributing guidelines](CONTRIBUTING.md).
## 📜 License
This model is released under the Apache 2.0 License.
## 🙏 Acknowledgments
- **TRL Team**: For the GRPO implementation
- **Qwen Team**: For the excellent base model
- **Hugging Face**: For the transformers library and model hosting
---
**Note**: This model is for educational and research purposes. Always consult qualified legal professionals for actual legal advice.

28
added_tokens.json Normal file
View File

@@ -0,0 +1,28 @@
{
"</think>": 151668,
"</tool_call>": 151658,
"</tool_response>": 151666,
"<think>": 151667,
"<tool_call>": 151657,
"<tool_response>": 151665,
"<|box_end|>": 151649,
"<|box_start|>": 151648,
"<|endoftext|>": 151643,
"<|file_sep|>": 151664,
"<|fim_middle|>": 151660,
"<|fim_pad|>": 151662,
"<|fim_prefix|>": 151659,
"<|fim_suffix|>": 151661,
"<|im_end|>": 151645,
"<|im_start|>": 151644,
"<|image_pad|>": 151655,
"<|object_ref_end|>": 151647,
"<|object_ref_start|>": 151646,
"<|quad_end|>": 151651,
"<|quad_start|>": 151650,
"<|repo_name|>": 151663,
"<|video_pad|>": 151656,
"<|vision_end|>": 151653,
"<|vision_pad|>": 151654,
"<|vision_start|>": 151652
}

85
chat_template.jinja Normal file
View File

@@ -0,0 +1,85 @@
{%- if tools %}
{{- '<|im_start|>system\n' }}
{%- if messages[0].role == 'system' %}
{{- messages[0].content + '\n\n' }}
{%- endif %}
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
{%- for tool in tools %}
{{- "\n" }}
{{- tool | tojson }}
{%- endfor %}
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
{%- else %}
{%- if messages[0].role == 'system' %}
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
{%- for message in messages[::-1] %}
{%- set index = (messages|length - 1) - loop.index0 %}
{%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
{%- set ns.multi_step_tool = false %}
{%- set ns.last_query_index = index %}
{%- endif %}
{%- endfor %}
{%- for message in messages %}
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
{%- elif message.role == "assistant" %}
{%- set content = message.content %}
{%- set reasoning_content = '' %}
{%- if message.reasoning_content is defined and message.reasoning_content is not none %}
{%- set reasoning_content = message.reasoning_content %}
{%- else %}
{%- if '</think>' in message.content %}
{%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
{%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
{%- endif %}
{%- endif %}
{%- if loop.index0 > ns.last_query_index %}
{%- if loop.last or (not loop.last and reasoning_content) %}
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
{%- else %}
{{- '<|im_start|>' + message.role + '\n' + content }}
{%- endif %}
{%- else %}
{{- '<|im_start|>' + message.role + '\n' + content }}
{%- endif %}
{%- if message.tool_calls %}
{%- for tool_call in message.tool_calls %}
{%- if (loop.first and content) or (not loop.first) %}
{{- '\n' }}
{%- endif %}
{%- if tool_call.function %}
{%- set tool_call = tool_call.function %}
{%- endif %}
{{- '<tool_call>\n{"name": "' }}
{{- tool_call.name }}
{{- '", "arguments": ' }}
{%- if tool_call.arguments is string %}
{{- tool_call.arguments }}
{%- else %}
{{- tool_call.arguments | tojson }}
{%- endif %}
{{- '}\n</tool_call>' }}
{%- endfor %}
{%- endif %}
{{- '<|im_end|>\n' }}
{%- elif message.role == "tool" %}
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
{{- '<|im_start|>user' }}
{%- endif %}
{{- '\n<tool_response>\n' }}
{{- message.content }}
{{- '\n</tool_response>' }}
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
{{- '<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|im_start|>assistant\n' }}
{%- if enable_thinking is defined and enable_thinking is false %}
{{- '<think>\n\n</think>\n\n' }}
{%- endif %}
{%- endif %}

60
config.json Normal file
View File

@@ -0,0 +1,60 @@
{
"architectures": [
"Qwen3ForCausalLM"
],
"attention_bias": false,
"attention_dropout": 0.0,
"bos_token_id": 151643,
"eos_token_id": 151643,
"head_dim": 128,
"hidden_act": "silu",
"hidden_size": 2048,
"initializer_range": 0.02,
"intermediate_size": 6144,
"layer_types": [
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention"
],
"max_position_embeddings": 32768,
"max_window_layers": 28,
"model_type": "qwen3",
"num_attention_heads": 16,
"num_hidden_layers": 28,
"num_key_value_heads": 8,
"rms_norm_eps": 1e-06,
"rope_scaling": null,
"rope_theta": 1000000,
"sliding_window": null,
"tie_word_embeddings": true,
"torch_dtype": "bfloat16",
"transformers_version": "4.55.4",
"use_cache": false,
"use_sliding_window": false,
"vocab_size": 151936
}

6
generation_config.json Normal file
View File

@@ -0,0 +1,6 @@
{
"bos_token_id": 151643,
"eos_token_id": 151643,
"max_new_tokens": 2048,
"transformers_version": "4.55.4"
}

151388
merges.txt Normal file

File diff suppressed because it is too large Load Diff

3
model.safetensors Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ffc13028dd141028561284daf7f48fd078028fdc7f23c912d250a227ca479011
size 3441185608

31
special_tokens_map.json Normal file
View File

@@ -0,0 +1,31 @@
{
"additional_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"eos_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
},
"pad_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false
}
}

BIN
tokenizer.json (Stored with Git LFS) Normal file

Binary file not shown.

243
tokenizer_config.json Normal file
View File

@@ -0,0 +1,243 @@
{
"add_bos_token": false,
"add_prefix_space": false,
"added_tokens_decoder": {
"151643": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151644": {
"content": "<|im_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151645": {
"content": "<|im_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151646": {
"content": "<|object_ref_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151647": {
"content": "<|object_ref_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151648": {
"content": "<|box_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151649": {
"content": "<|box_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151650": {
"content": "<|quad_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151651": {
"content": "<|quad_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151652": {
"content": "<|vision_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151653": {
"content": "<|vision_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151654": {
"content": "<|vision_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151655": {
"content": "<|image_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151656": {
"content": "<|video_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151657": {
"content": "<tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151658": {
"content": "</tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151659": {
"content": "<|fim_prefix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151660": {
"content": "<|fim_middle|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151661": {
"content": "<|fim_suffix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151662": {
"content": "<|fim_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151663": {
"content": "<|repo_name|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151664": {
"content": "<|file_sep|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151665": {
"content": "<tool_response>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151666": {
"content": "</tool_response>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151667": {
"content": "<think>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151668": {
"content": "</think>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
}
},
"additional_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"bos_token": null,
"clean_up_tokenization_spaces": false,
"eos_token": "<|endoftext|>",
"errors": "replace",
"extra_special_tokens": {},
"max_length": null,
"model_max_length": 131072,
"pad_to_multiple_of": null,
"pad_token": "<|endoftext|>",
"pad_token_type_id": 0,
"padding_side": "left",
"split_special_tokens": false,
"tokenizer_class": "Qwen2Tokenizer",
"unk_token": null
}

3
training_args.bin Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1dba5f20d01dcb938fc0621f76b1ba3c1f9a1792b388734680e2f34362fd5ee7
size 7057

1
vocab.json Normal file

File diff suppressed because one or more lines are too long