初始化项目,由ModelHub XC社区提供模型
Model: marinarosa/minicpm5-1b-vivamais-v4 Source: Original Platform
This commit is contained in:
198
README.md
Normal file
198
README.md
Normal file
@@ -0,0 +1,198 @@
|
||||
---
|
||||
language:
|
||||
- pt
|
||||
license: apache-2.0
|
||||
base_model:
|
||||
- marinarosa/minicpm5-1b-vivamais-v1
|
||||
datasets:
|
||||
- marinarosa/minicpm5-vivamais-text-sft-v4
|
||||
library_name: transformers
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- brazilian-portuguese
|
||||
- minicpm5
|
||||
- vivamais
|
||||
- travel-agency
|
||||
- grounded-qa
|
||||
- conversational
|
||||
- distilled
|
||||
- qlora
|
||||
- lora
|
||||
- redacted
|
||||
model-index:
|
||||
- name: minicpm5-1b-vivamais-v4
|
||||
results:
|
||||
- task:
|
||||
type: text-generation
|
||||
name: Viva Mais dashboard QA
|
||||
dataset:
|
||||
name: Viva Mais QA eval v4
|
||||
type: marinarosa/minicpm5-vivamais-text-sft-v4
|
||||
split: eval
|
||||
metrics:
|
||||
- type: average_score
|
||||
value: 0.5569620253164557
|
||||
name: Average score
|
||||
- type: pass_rate
|
||||
value: 0.44936708860759494
|
||||
name: Pass rate
|
||||
- task:
|
||||
type: text-generation
|
||||
name: Generic PT-BR conversational QA
|
||||
dataset:
|
||||
name: Viva Mais generic PT-BR eval
|
||||
type: marinarosa/minicpm5-vivamais-text-sft-v4
|
||||
split: eval
|
||||
metrics:
|
||||
- type: average_score
|
||||
value: 0.627906976744186
|
||||
name: Average score
|
||||
- type: pass_rate
|
||||
value: 0.627906976744186
|
||||
name: Pass rate
|
||||
---
|
||||
|
||||
# MiniCPM5-1B Viva Mais v4
|
||||
|
||||
`marinarosa/minicpm5-1b-vivamais-v4` is a Brazilian Portuguese text model for Viva Mais, a local-first
|
||||
WhatsApp copilot for travel-agency CRM workflows. It is intended for grounded Q&A
|
||||
over extracted customer, payment, reservation, ticket, and document context.
|
||||
|
||||
This v4 candidate was published because it improves the aggregate dashboard and
|
||||
generic PT-BR eval scores over v1. It still has an important limitation: it often
|
||||
over-generates and may include adjacent customer facts after the direct answer.
|
||||
Use short generation limits and strong retrieval/context isolation in production.
|
||||
|
||||
## Lineage
|
||||
|
||||
- Upstream family: `openbmb/MiniCPM5-1B`.
|
||||
- Fine-tuning base for this release: `marinarosa/minicpm5-1b-vivamais-v1`.
|
||||
- Published Transformers model: `marinarosa/minicpm5-1b-vivamais-v4`.
|
||||
- GGUF export: `marinarosa/minicpm5-1b-vivamais-v4-GGUF`.
|
||||
- Training dataset: `marinarosa/minicpm5-vivamais-text-sft-v4`.
|
||||
|
||||
## Fine-Tuning Recipe
|
||||
|
||||
The v4 run used response-only supervised fine-tuning with a QLoRA-style training
|
||||
setup, followed by a merged 16-bit export.
|
||||
|
||||
- Hardware: A100 on Modal.
|
||||
- Context length: 8192 tokens.
|
||||
- Training rows: 4000.
|
||||
- Epochs: 2.0.
|
||||
- Steps: 500.
|
||||
- Learning rate: 1e-05 with cosine schedule and 3%
|
||||
warmup.
|
||||
- Per-device batch size: 2.
|
||||
- Gradient accumulation: 8.
|
||||
- Effective batch size: 16.
|
||||
- Final training loss: 0.2231.
|
||||
- Base loaded in 4-bit for training; final artifact saved as merged 16-bit
|
||||
safetensors.
|
||||
- LoRA rank: 16.
|
||||
- LoRA alpha: 32.
|
||||
- LoRA dropout: 0.05.
|
||||
- LoRA target modules: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`,
|
||||
`up_proj`, `down_proj`.
|
||||
- Response-only masking used `<|im_start|>user` and a no-thinking assistant
|
||||
prefix so the model trains only on final assistant answers.
|
||||
|
||||
## Dataset Mix
|
||||
|
||||
| Bucket | Rows |
|
||||
| --- | ---: |
|
||||
| Grounding negatives / unknown-answer cases | 2381 |
|
||||
| Viva Mais domain QA | 1511 |
|
||||
| Rio 3.1 teacher-distilled chat | 80 |
|
||||
| Generic PT-BR conversation remainder | 28 |
|
||||
|
||||
Teacher distillation used `prefeitura-rio/Rio-3.1-Open-4B-Instruct`. The run accepted
|
||||
80 rows and rejected
|
||||
5 rows after filters for verbose wrappers,
|
||||
hidden-reasoning artifacts, weak travel claims, and category mismatch.
|
||||
|
||||
## Evaluation
|
||||
|
||||
Run name: `v4_rio_001`. Baseline: `marinarosa/minicpm5-1b-vivamais-v1`. Candidate: `marinarosa/minicpm5-1b-vivamais-v4`.
|
||||
|
||||
| Suite | Model | Avg score | Pass rate | Count |
|
||||
| --- | --- | ---: | ---: | ---: |
|
||||
| Viva Mais dashboard QA | v1 baseline | 0.5042 | 0.4051 | 158 |
|
||||
| Viva Mais dashboard QA | v4 candidate | 0.5570 | 0.4494 | 158 |
|
||||
| Generic PT-BR conversational QA | v1 baseline | 0.5814 | 0.5814 | 43 |
|
||||
| Generic PT-BR conversational QA | v4 candidate | 0.6279 | 0.6279 | 43 |
|
||||
|
||||
Dashboard QA details:
|
||||
|
||||
| Metric | v1 baseline | v4 candidate |
|
||||
| --- | ---: | ---: |
|
||||
| Average score | 0.5042 | 0.5570 |
|
||||
| Pass rate | 0.4051 | 0.4494 |
|
||||
| Leakage failures | 54 | 55 |
|
||||
| Unknown failures | 18 | 11 |
|
||||
| Internal gate passed | false | false |
|
||||
|
||||
Viva Mais category scores for v4:
|
||||
|
||||
| Category | Score |
|
||||
| --- | ---: |
|
||||
| conversation_grounding | 0.6667 |
|
||||
| dates | 0.7222 |
|
||||
| document_status | 0.7500 |
|
||||
| entity_isolation | 0.8333 |
|
||||
| long_context_distractor | 0.6250 |
|
||||
| not_enough_information | 0.0556 |
|
||||
| outstanding_balance | 0.5417 |
|
||||
| passenger_flight_lookup | 0.4677 |
|
||||
| payment_status | 0.6000 |
|
||||
| pipeline_stage | 0.7500 |
|
||||
| quote_status | 0.7500 |
|
||||
| what_is_missing | 0.2500 |
|
||||
|
||||
Generic PT-BR category scores for v4:
|
||||
|
||||
| Category | Score |
|
||||
| --- | ---: |
|
||||
| customer_service | 0.6250 |
|
||||
| entity_isolation | 0.4000 |
|
||||
| grounded_refusal | 0.8333 |
|
||||
| payment_math | 0.6667 |
|
||||
| summary_rewrite | 0.5714 |
|
||||
| travel_qa | 0.3333 |
|
||||
| whatsapp_style | 1.0000 |
|
||||
|
||||
## Usage
|
||||
|
||||
Use the MiniCPM chat template and keep generation short for grounded CRM answers.
|
||||
For local llama.cpp inference, prefer the GGUF repo linked above.
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
model_id = "marinarosa/minicpm5-1b-vivamais-v4"
|
||||
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
model_id,
|
||||
torch_dtype="auto",
|
||||
device_map="auto",
|
||||
trust_remote_code=True,
|
||||
)
|
||||
messages = [
|
||||
{"role": "system", "content": "Responda em PT-BR usando apenas o contexto fornecido."},
|
||||
{"role": "user", "content": "Contexto: ...\nPergunta: qual saldo ainda falta?"},
|
||||
]
|
||||
text = tok.apply_chat_template(
|
||||
messages,
|
||||
tokenize=False,
|
||||
add_generation_prompt=True,
|
||||
enable_thinking=False,
|
||||
)
|
||||
```
|
||||
|
||||
## Privacy and Safety
|
||||
|
||||
The training data is synthetic or redacted. No raw WhatsApp exports, identity
|
||||
documents, client names, CPF values, phone numbers, emails, card numbers, or
|
||||
local private data are published. This model is not a general travel advisor or
|
||||
payment authority; it should answer only from supplied context and defer when the
|
||||
context is insufficient.
|
||||
Reference in New Issue
Block a user