Files
minicpm5-1b-vivamais-v4/README.md

199 lines
5.8 KiB
Markdown
Raw Normal View History

---
language:
- pt
license: apache-2.0
base_model:
- marinarosa/minicpm5-1b-vivamais-v1
datasets:
- marinarosa/minicpm5-vivamais-text-sft-v4
library_name: transformers
pipeline_tag: text-generation
tags:
- brazilian-portuguese
- minicpm5
- vivamais
- travel-agency
- grounded-qa
- conversational
- distilled
- qlora
- lora
- redacted
model-index:
- name: minicpm5-1b-vivamais-v4
results:
- task:
type: text-generation
name: Viva Mais dashboard QA
dataset:
name: Viva Mais QA eval v4
type: marinarosa/minicpm5-vivamais-text-sft-v4
split: eval
metrics:
- type: average_score
value: 0.5569620253164557
name: Average score
- type: pass_rate
value: 0.44936708860759494
name: Pass rate
- task:
type: text-generation
name: Generic PT-BR conversational QA
dataset:
name: Viva Mais generic PT-BR eval
type: marinarosa/minicpm5-vivamais-text-sft-v4
split: eval
metrics:
- type: average_score
value: 0.627906976744186
name: Average score
- type: pass_rate
value: 0.627906976744186
name: Pass rate
---
# MiniCPM5-1B Viva Mais v4
`marinarosa/minicpm5-1b-vivamais-v4` is a Brazilian Portuguese text model for Viva Mais, a local-first
WhatsApp copilot for travel-agency CRM workflows. It is intended for grounded Q&A
over extracted customer, payment, reservation, ticket, and document context.
This v4 candidate was published because it improves the aggregate dashboard and
generic PT-BR eval scores over v1. It still has an important limitation: it often
over-generates and may include adjacent customer facts after the direct answer.
Use short generation limits and strong retrieval/context isolation in production.
## Lineage
- Upstream family: `openbmb/MiniCPM5-1B`.
- Fine-tuning base for this release: `marinarosa/minicpm5-1b-vivamais-v1`.
- Published Transformers model: `marinarosa/minicpm5-1b-vivamais-v4`.
- GGUF export: `marinarosa/minicpm5-1b-vivamais-v4-GGUF`.
- Training dataset: `marinarosa/minicpm5-vivamais-text-sft-v4`.
## Fine-Tuning Recipe
The v4 run used response-only supervised fine-tuning with a QLoRA-style training
setup, followed by a merged 16-bit export.
- Hardware: A100 on Modal.
- Context length: 8192 tokens.
- Training rows: 4000.
- Epochs: 2.0.
- Steps: 500.
- Learning rate: 1e-05 with cosine schedule and 3%
warmup.
- Per-device batch size: 2.
- Gradient accumulation: 8.
- Effective batch size: 16.
- Final training loss: 0.2231.
- Base loaded in 4-bit for training; final artifact saved as merged 16-bit
safetensors.
- LoRA rank: 16.
- LoRA alpha: 32.
- LoRA dropout: 0.05.
- LoRA target modules: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`,
`up_proj`, `down_proj`.
- Response-only masking used `<|im_start|>user` and a no-thinking assistant
prefix so the model trains only on final assistant answers.
## Dataset Mix
| Bucket | Rows |
| --- | ---: |
| Grounding negatives / unknown-answer cases | 2381 |
| Viva Mais domain QA | 1511 |
| Rio 3.1 teacher-distilled chat | 80 |
| Generic PT-BR conversation remainder | 28 |
Teacher distillation used `prefeitura-rio/Rio-3.1-Open-4B-Instruct`. The run accepted
80 rows and rejected
5 rows after filters for verbose wrappers,
hidden-reasoning artifacts, weak travel claims, and category mismatch.
## Evaluation
Run name: `v4_rio_001`. Baseline: `marinarosa/minicpm5-1b-vivamais-v1`. Candidate: `marinarosa/minicpm5-1b-vivamais-v4`.
| Suite | Model | Avg score | Pass rate | Count |
| --- | --- | ---: | ---: | ---: |
| Viva Mais dashboard QA | v1 baseline | 0.5042 | 0.4051 | 158 |
| Viva Mais dashboard QA | v4 candidate | 0.5570 | 0.4494 | 158 |
| Generic PT-BR conversational QA | v1 baseline | 0.5814 | 0.5814 | 43 |
| Generic PT-BR conversational QA | v4 candidate | 0.6279 | 0.6279 | 43 |
Dashboard QA details:
| Metric | v1 baseline | v4 candidate |
| --- | ---: | ---: |
| Average score | 0.5042 | 0.5570 |
| Pass rate | 0.4051 | 0.4494 |
| Leakage failures | 54 | 55 |
| Unknown failures | 18 | 11 |
| Internal gate passed | false | false |
Viva Mais category scores for v4:
| Category | Score |
| --- | ---: |
| conversation_grounding | 0.6667 |
| dates | 0.7222 |
| document_status | 0.7500 |
| entity_isolation | 0.8333 |
| long_context_distractor | 0.6250 |
| not_enough_information | 0.0556 |
| outstanding_balance | 0.5417 |
| passenger_flight_lookup | 0.4677 |
| payment_status | 0.6000 |
| pipeline_stage | 0.7500 |
| quote_status | 0.7500 |
| what_is_missing | 0.2500 |
Generic PT-BR category scores for v4:
| Category | Score |
| --- | ---: |
| customer_service | 0.6250 |
| entity_isolation | 0.4000 |
| grounded_refusal | 0.8333 |
| payment_math | 0.6667 |
| summary_rewrite | 0.5714 |
| travel_qa | 0.3333 |
| whatsapp_style | 1.0000 |
## Usage
Use the MiniCPM chat template and keep generation short for grounded CRM answers.
For local llama.cpp inference, prefer the GGUF repo linked above.
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "marinarosa/minicpm5-1b-vivamais-v4"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True,
)
messages = [
{"role": "system", "content": "Responda em PT-BR usando apenas o contexto fornecido."},
{"role": "user", "content": "Contexto: ...\nPergunta: qual saldo ainda falta?"},
]
text = tok.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)
```
## Privacy and Safety
The training data is synthetic or redacted. No raw WhatsApp exports, identity
documents, client names, CPF values, phone numbers, emails, card numbers, or
local private data are published. This model is not a general travel advisor or
payment authority; it should answer only from supplied context and defer when the
context is insufficient.