language, license, base_model, datasets, library_name, pipeline_tag, tags, model-index
language license base_model datasets library_name pipeline_tag tags model-index
pt
apache-2.0
marinarosa/minicpm5-1b-vivamais-v1
marinarosa/minicpm5-vivamais-text-sft-v4
transformers text-generation
brazilian-portuguese
minicpm5
vivamais
travel-agency
grounded-qa
conversational
distilled
qlora
lora
redacted
name results
minicpm5-1b-vivamais-v4
task dataset metrics
type name
text-generation Viva Mais dashboard QA
name type split
Viva Mais QA eval v4 marinarosa/minicpm5-vivamais-text-sft-v4 eval
type value name
average_score 0.5569620253164557 Average score
type value name
pass_rate 0.44936708860759494 Pass rate
task dataset metrics
type name
text-generation Generic PT-BR conversational QA
name type split
Viva Mais generic PT-BR eval marinarosa/minicpm5-vivamais-text-sft-v4 eval
type value name
average_score 0.627906976744186 Average score
type value name
pass_rate 0.627906976744186 Pass rate

MiniCPM5-1B Viva Mais v4

marinarosa/minicpm5-1b-vivamais-v4 is a Brazilian Portuguese text model for Viva Mais, a local-first WhatsApp copilot for travel-agency CRM workflows. It is intended for grounded Q&A over extracted customer, payment, reservation, ticket, and document context.

This v4 candidate was published because it improves the aggregate dashboard and generic PT-BR eval scores over v1. It still has an important limitation: it often over-generates and may include adjacent customer facts after the direct answer. Use short generation limits and strong retrieval/context isolation in production.

Lineage

  • Upstream family: openbmb/MiniCPM5-1B.
  • Fine-tuning base for this release: marinarosa/minicpm5-1b-vivamais-v1.
  • Published Transformers model: marinarosa/minicpm5-1b-vivamais-v4.
  • GGUF export: marinarosa/minicpm5-1b-vivamais-v4-GGUF.
  • Training dataset: marinarosa/minicpm5-vivamais-text-sft-v4.

Fine-Tuning Recipe

The v4 run used response-only supervised fine-tuning with a QLoRA-style training setup, followed by a merged 16-bit export.

  • Hardware: A100 on Modal.
  • Context length: 8192 tokens.
  • Training rows: 4000.
  • Epochs: 2.0.
  • Steps: 500.
  • Learning rate: 1e-05 with cosine schedule and 3% warmup.
  • Per-device batch size: 2.
  • Gradient accumulation: 8.
  • Effective batch size: 16.
  • Final training loss: 0.2231.
  • Base loaded in 4-bit for training; final artifact saved as merged 16-bit safetensors.
  • LoRA rank: 16.
  • LoRA alpha: 32.
  • LoRA dropout: 0.05.
  • LoRA target modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj.
  • Response-only masking used <|im_start|>user and a no-thinking assistant prefix so the model trains only on final assistant answers.

Dataset Mix

Bucket Rows
Grounding negatives / unknown-answer cases 2381
Viva Mais domain QA 1511
Rio 3.1 teacher-distilled chat 80
Generic PT-BR conversation remainder 28

Teacher distillation used prefeitura-rio/Rio-3.1-Open-4B-Instruct. The run accepted 80 rows and rejected 5 rows after filters for verbose wrappers, hidden-reasoning artifacts, weak travel claims, and category mismatch.

Evaluation

Run name: v4_rio_001. Baseline: marinarosa/minicpm5-1b-vivamais-v1. Candidate: marinarosa/minicpm5-1b-vivamais-v4.

Suite Model Avg score Pass rate Count
Viva Mais dashboard QA v1 baseline 0.5042 0.4051 158
Viva Mais dashboard QA v4 candidate 0.5570 0.4494 158
Generic PT-BR conversational QA v1 baseline 0.5814 0.5814 43
Generic PT-BR conversational QA v4 candidate 0.6279 0.6279 43

Dashboard QA details:

Metric v1 baseline v4 candidate
Average score 0.5042 0.5570
Pass rate 0.4051 0.4494
Leakage failures 54 55
Unknown failures 18 11
Internal gate passed false false

Viva Mais category scores for v4:

Category Score
conversation_grounding 0.6667
dates 0.7222
document_status 0.7500
entity_isolation 0.8333
long_context_distractor 0.6250
not_enough_information 0.0556
outstanding_balance 0.5417
passenger_flight_lookup 0.4677
payment_status 0.6000
pipeline_stage 0.7500
quote_status 0.7500
what_is_missing 0.2500

Generic PT-BR category scores for v4:

Category Score
customer_service 0.6250
entity_isolation 0.4000
grounded_refusal 0.8333
payment_math 0.6667
summary_rewrite 0.5714
travel_qa 0.3333
whatsapp_style 1.0000

Usage

Use the MiniCPM chat template and keep generation short for grounded CRM answers. For local llama.cpp inference, prefer the GGUF repo linked above.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "marinarosa/minicpm5-1b-vivamais-v4"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
    trust_remote_code=True,
)
messages = [
    {"role": "system", "content": "Responda em PT-BR usando apenas o contexto fornecido."},
    {"role": "user", "content": "Contexto: ...\nPergunta: qual saldo ainda falta?"},
]
text = tok.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=False,
)

Privacy and Safety

The training data is synthetic or redacted. No raw WhatsApp exports, identity documents, client names, CPF values, phone numbers, emails, card numbers, or local private data are published. This model is not a general travel advisor or payment authority; it should answer only from supplied context and defer when the context is insufficient.

Description
Model synced from source: marinarosa/minicpm5-1b-vivamais-v4
Readme 2.3 MiB
Languages
Jinja 100%