--- language: - pt license: apache-2.0 base_model: - marinarosa/minicpm5-1b-vivamais-v1 datasets: - marinarosa/minicpm5-vivamais-text-sft-v4 library_name: transformers pipeline_tag: text-generation tags: - brazilian-portuguese - minicpm5 - vivamais - travel-agency - grounded-qa - conversational - distilled - qlora - lora - redacted model-index: - name: minicpm5-1b-vivamais-v4 results: - task: type: text-generation name: Viva Mais dashboard QA dataset: name: Viva Mais QA eval v4 type: marinarosa/minicpm5-vivamais-text-sft-v4 split: eval metrics: - type: average_score value: 0.5569620253164557 name: Average score - type: pass_rate value: 0.44936708860759494 name: Pass rate - task: type: text-generation name: Generic PT-BR conversational QA dataset: name: Viva Mais generic PT-BR eval type: marinarosa/minicpm5-vivamais-text-sft-v4 split: eval metrics: - type: average_score value: 0.627906976744186 name: Average score - type: pass_rate value: 0.627906976744186 name: Pass rate --- # MiniCPM5-1B Viva Mais v4 `marinarosa/minicpm5-1b-vivamais-v4` is a Brazilian Portuguese text model for Viva Mais, a local-first WhatsApp copilot for travel-agency CRM workflows. It is intended for grounded Q&A over extracted customer, payment, reservation, ticket, and document context. This v4 candidate was published because it improves the aggregate dashboard and generic PT-BR eval scores over v1. It still has an important limitation: it often over-generates and may include adjacent customer facts after the direct answer. Use short generation limits and strong retrieval/context isolation in production. ## Lineage - Upstream family: `openbmb/MiniCPM5-1B`. - Fine-tuning base for this release: `marinarosa/minicpm5-1b-vivamais-v1`. - Published Transformers model: `marinarosa/minicpm5-1b-vivamais-v4`. - GGUF export: `marinarosa/minicpm5-1b-vivamais-v4-GGUF`. - Training dataset: `marinarosa/minicpm5-vivamais-text-sft-v4`. ## Fine-Tuning Recipe The v4 run used response-only supervised fine-tuning with a QLoRA-style training setup, followed by a merged 16-bit export. - Hardware: A100 on Modal. - Context length: 8192 tokens. - Training rows: 4000. - Epochs: 2.0. - Steps: 500. - Learning rate: 1e-05 with cosine schedule and 3% warmup. - Per-device batch size: 2. - Gradient accumulation: 8. - Effective batch size: 16. - Final training loss: 0.2231. - Base loaded in 4-bit for training; final artifact saved as merged 16-bit safetensors. - LoRA rank: 16. - LoRA alpha: 32. - LoRA dropout: 0.05. - LoRA target modules: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`. - Response-only masking used `<|im_start|>user` and a no-thinking assistant prefix so the model trains only on final assistant answers. ## Dataset Mix | Bucket | Rows | | --- | ---: | | Grounding negatives / unknown-answer cases | 2381 | | Viva Mais domain QA | 1511 | | Rio 3.1 teacher-distilled chat | 80 | | Generic PT-BR conversation remainder | 28 | Teacher distillation used `prefeitura-rio/Rio-3.1-Open-4B-Instruct`. The run accepted 80 rows and rejected 5 rows after filters for verbose wrappers, hidden-reasoning artifacts, weak travel claims, and category mismatch. ## Evaluation Run name: `v4_rio_001`. Baseline: `marinarosa/minicpm5-1b-vivamais-v1`. Candidate: `marinarosa/minicpm5-1b-vivamais-v4`. | Suite | Model | Avg score | Pass rate | Count | | --- | --- | ---: | ---: | ---: | | Viva Mais dashboard QA | v1 baseline | 0.5042 | 0.4051 | 158 | | Viva Mais dashboard QA | v4 candidate | 0.5570 | 0.4494 | 158 | | Generic PT-BR conversational QA | v1 baseline | 0.5814 | 0.5814 | 43 | | Generic PT-BR conversational QA | v4 candidate | 0.6279 | 0.6279 | 43 | Dashboard QA details: | Metric | v1 baseline | v4 candidate | | --- | ---: | ---: | | Average score | 0.5042 | 0.5570 | | Pass rate | 0.4051 | 0.4494 | | Leakage failures | 54 | 55 | | Unknown failures | 18 | 11 | | Internal gate passed | false | false | Viva Mais category scores for v4: | Category | Score | | --- | ---: | | conversation_grounding | 0.6667 | | dates | 0.7222 | | document_status | 0.7500 | | entity_isolation | 0.8333 | | long_context_distractor | 0.6250 | | not_enough_information | 0.0556 | | outstanding_balance | 0.5417 | | passenger_flight_lookup | 0.4677 | | payment_status | 0.6000 | | pipeline_stage | 0.7500 | | quote_status | 0.7500 | | what_is_missing | 0.2500 | Generic PT-BR category scores for v4: | Category | Score | | --- | ---: | | customer_service | 0.6250 | | entity_isolation | 0.4000 | | grounded_refusal | 0.8333 | | payment_math | 0.6667 | | summary_rewrite | 0.5714 | | travel_qa | 0.3333 | | whatsapp_style | 1.0000 | ## Usage Use the MiniCPM chat template and keep generation short for grounded CRM answers. For local llama.cpp inference, prefer the GGUF repo linked above. ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "marinarosa/minicpm5-1b-vivamais-v4" tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype="auto", device_map="auto", trust_remote_code=True, ) messages = [ {"role": "system", "content": "Responda em PT-BR usando apenas o contexto fornecido."}, {"role": "user", "content": "Contexto: ...\nPergunta: qual saldo ainda falta?"}, ] text = tok.apply_chat_template( messages, tokenize=False, add_generation_prompt=True, enable_thinking=False, ) ``` ## Privacy and Safety The training data is synthetic or redacted. No raw WhatsApp exports, identity documents, client names, CPF values, phone numbers, emails, card numbers, or local private data are published. This model is not a general travel advisor or payment authority; it should answer only from supplied context and defer when the context is insufficient.