199 lines
5.8 KiB
Markdown
199 lines
5.8 KiB
Markdown
|
|
---
|
||
|
|
language:
|
||
|
|
- pt
|
||
|
|
license: apache-2.0
|
||
|
|
base_model:
|
||
|
|
- marinarosa/minicpm5-1b-vivamais-v1
|
||
|
|
datasets:
|
||
|
|
- marinarosa/minicpm5-vivamais-text-sft-v4
|
||
|
|
library_name: transformers
|
||
|
|
pipeline_tag: text-generation
|
||
|
|
tags:
|
||
|
|
- brazilian-portuguese
|
||
|
|
- minicpm5
|
||
|
|
- vivamais
|
||
|
|
- travel-agency
|
||
|
|
- grounded-qa
|
||
|
|
- conversational
|
||
|
|
- distilled
|
||
|
|
- qlora
|
||
|
|
- lora
|
||
|
|
- redacted
|
||
|
|
model-index:
|
||
|
|
- name: minicpm5-1b-vivamais-v4
|
||
|
|
results:
|
||
|
|
- task:
|
||
|
|
type: text-generation
|
||
|
|
name: Viva Mais dashboard QA
|
||
|
|
dataset:
|
||
|
|
name: Viva Mais QA eval v4
|
||
|
|
type: marinarosa/minicpm5-vivamais-text-sft-v4
|
||
|
|
split: eval
|
||
|
|
metrics:
|
||
|
|
- type: average_score
|
||
|
|
value: 0.5569620253164557
|
||
|
|
name: Average score
|
||
|
|
- type: pass_rate
|
||
|
|
value: 0.44936708860759494
|
||
|
|
name: Pass rate
|
||
|
|
- task:
|
||
|
|
type: text-generation
|
||
|
|
name: Generic PT-BR conversational QA
|
||
|
|
dataset:
|
||
|
|
name: Viva Mais generic PT-BR eval
|
||
|
|
type: marinarosa/minicpm5-vivamais-text-sft-v4
|
||
|
|
split: eval
|
||
|
|
metrics:
|
||
|
|
- type: average_score
|
||
|
|
value: 0.627906976744186
|
||
|
|
name: Average score
|
||
|
|
- type: pass_rate
|
||
|
|
value: 0.627906976744186
|
||
|
|
name: Pass rate
|
||
|
|
---
|
||
|
|
|
||
|
|
# MiniCPM5-1B Viva Mais v4
|
||
|
|
|
||
|
|
`marinarosa/minicpm5-1b-vivamais-v4` is a Brazilian Portuguese text model for Viva Mais, a local-first
|
||
|
|
WhatsApp copilot for travel-agency CRM workflows. It is intended for grounded Q&A
|
||
|
|
over extracted customer, payment, reservation, ticket, and document context.
|
||
|
|
|
||
|
|
This v4 candidate was published because it improves the aggregate dashboard and
|
||
|
|
generic PT-BR eval scores over v1. It still has an important limitation: it often
|
||
|
|
over-generates and may include adjacent customer facts after the direct answer.
|
||
|
|
Use short generation limits and strong retrieval/context isolation in production.
|
||
|
|
|
||
|
|
## Lineage
|
||
|
|
|
||
|
|
- Upstream family: `openbmb/MiniCPM5-1B`.
|
||
|
|
- Fine-tuning base for this release: `marinarosa/minicpm5-1b-vivamais-v1`.
|
||
|
|
- Published Transformers model: `marinarosa/minicpm5-1b-vivamais-v4`.
|
||
|
|
- GGUF export: `marinarosa/minicpm5-1b-vivamais-v4-GGUF`.
|
||
|
|
- Training dataset: `marinarosa/minicpm5-vivamais-text-sft-v4`.
|
||
|
|
|
||
|
|
## Fine-Tuning Recipe
|
||
|
|
|
||
|
|
The v4 run used response-only supervised fine-tuning with a QLoRA-style training
|
||
|
|
setup, followed by a merged 16-bit export.
|
||
|
|
|
||
|
|
- Hardware: A100 on Modal.
|
||
|
|
- Context length: 8192 tokens.
|
||
|
|
- Training rows: 4000.
|
||
|
|
- Epochs: 2.0.
|
||
|
|
- Steps: 500.
|
||
|
|
- Learning rate: 1e-05 with cosine schedule and 3%
|
||
|
|
warmup.
|
||
|
|
- Per-device batch size: 2.
|
||
|
|
- Gradient accumulation: 8.
|
||
|
|
- Effective batch size: 16.
|
||
|
|
- Final training loss: 0.2231.
|
||
|
|
- Base loaded in 4-bit for training; final artifact saved as merged 16-bit
|
||
|
|
safetensors.
|
||
|
|
- LoRA rank: 16.
|
||
|
|
- LoRA alpha: 32.
|
||
|
|
- LoRA dropout: 0.05.
|
||
|
|
- LoRA target modules: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`,
|
||
|
|
`up_proj`, `down_proj`.
|
||
|
|
- Response-only masking used `<|im_start|>user` and a no-thinking assistant
|
||
|
|
prefix so the model trains only on final assistant answers.
|
||
|
|
|
||
|
|
## Dataset Mix
|
||
|
|
|
||
|
|
| Bucket | Rows |
|
||
|
|
| --- | ---: |
|
||
|
|
| Grounding negatives / unknown-answer cases | 2381 |
|
||
|
|
| Viva Mais domain QA | 1511 |
|
||
|
|
| Rio 3.1 teacher-distilled chat | 80 |
|
||
|
|
| Generic PT-BR conversation remainder | 28 |
|
||
|
|
|
||
|
|
Teacher distillation used `prefeitura-rio/Rio-3.1-Open-4B-Instruct`. The run accepted
|
||
|
|
80 rows and rejected
|
||
|
|
5 rows after filters for verbose wrappers,
|
||
|
|
hidden-reasoning artifacts, weak travel claims, and category mismatch.
|
||
|
|
|
||
|
|
## Evaluation
|
||
|
|
|
||
|
|
Run name: `v4_rio_001`. Baseline: `marinarosa/minicpm5-1b-vivamais-v1`. Candidate: `marinarosa/minicpm5-1b-vivamais-v4`.
|
||
|
|
|
||
|
|
| Suite | Model | Avg score | Pass rate | Count |
|
||
|
|
| --- | --- | ---: | ---: | ---: |
|
||
|
|
| Viva Mais dashboard QA | v1 baseline | 0.5042 | 0.4051 | 158 |
|
||
|
|
| Viva Mais dashboard QA | v4 candidate | 0.5570 | 0.4494 | 158 |
|
||
|
|
| Generic PT-BR conversational QA | v1 baseline | 0.5814 | 0.5814 | 43 |
|
||
|
|
| Generic PT-BR conversational QA | v4 candidate | 0.6279 | 0.6279 | 43 |
|
||
|
|
|
||
|
|
Dashboard QA details:
|
||
|
|
|
||
|
|
| Metric | v1 baseline | v4 candidate |
|
||
|
|
| --- | ---: | ---: |
|
||
|
|
| Average score | 0.5042 | 0.5570 |
|
||
|
|
| Pass rate | 0.4051 | 0.4494 |
|
||
|
|
| Leakage failures | 54 | 55 |
|
||
|
|
| Unknown failures | 18 | 11 |
|
||
|
|
| Internal gate passed | false | false |
|
||
|
|
|
||
|
|
Viva Mais category scores for v4:
|
||
|
|
|
||
|
|
| Category | Score |
|
||
|
|
| --- | ---: |
|
||
|
|
| conversation_grounding | 0.6667 |
|
||
|
|
| dates | 0.7222 |
|
||
|
|
| document_status | 0.7500 |
|
||
|
|
| entity_isolation | 0.8333 |
|
||
|
|
| long_context_distractor | 0.6250 |
|
||
|
|
| not_enough_information | 0.0556 |
|
||
|
|
| outstanding_balance | 0.5417 |
|
||
|
|
| passenger_flight_lookup | 0.4677 |
|
||
|
|
| payment_status | 0.6000 |
|
||
|
|
| pipeline_stage | 0.7500 |
|
||
|
|
| quote_status | 0.7500 |
|
||
|
|
| what_is_missing | 0.2500 |
|
||
|
|
|
||
|
|
Generic PT-BR category scores for v4:
|
||
|
|
|
||
|
|
| Category | Score |
|
||
|
|
| --- | ---: |
|
||
|
|
| customer_service | 0.6250 |
|
||
|
|
| entity_isolation | 0.4000 |
|
||
|
|
| grounded_refusal | 0.8333 |
|
||
|
|
| payment_math | 0.6667 |
|
||
|
|
| summary_rewrite | 0.5714 |
|
||
|
|
| travel_qa | 0.3333 |
|
||
|
|
| whatsapp_style | 1.0000 |
|
||
|
|
|
||
|
|
## Usage
|
||
|
|
|
||
|
|
Use the MiniCPM chat template and keep generation short for grounded CRM answers.
|
||
|
|
For local llama.cpp inference, prefer the GGUF repo linked above.
|
||
|
|
|
||
|
|
```python
|
||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||
|
|
|
||
|
|
model_id = "marinarosa/minicpm5-1b-vivamais-v4"
|
||
|
|
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
|
||
|
|
model = AutoModelForCausalLM.from_pretrained(
|
||
|
|
model_id,
|
||
|
|
torch_dtype="auto",
|
||
|
|
device_map="auto",
|
||
|
|
trust_remote_code=True,
|
||
|
|
)
|
||
|
|
messages = [
|
||
|
|
{"role": "system", "content": "Responda em PT-BR usando apenas o contexto fornecido."},
|
||
|
|
{"role": "user", "content": "Contexto: ...\nPergunta: qual saldo ainda falta?"},
|
||
|
|
]
|
||
|
|
text = tok.apply_chat_template(
|
||
|
|
messages,
|
||
|
|
tokenize=False,
|
||
|
|
add_generation_prompt=True,
|
||
|
|
enable_thinking=False,
|
||
|
|
)
|
||
|
|
```
|
||
|
|
|
||
|
|
## Privacy and Safety
|
||
|
|
|
||
|
|
The training data is synthetic or redacted. No raw WhatsApp exports, identity
|
||
|
|
documents, client names, CPF values, phone numbers, emails, card numbers, or
|
||
|
|
local private data are published. This model is not a general travel advisor or
|
||
|
|
payment authority; it should answer only from supplied context and defer when the
|
||
|
|
context is insufficient.
|