382 lines
8.6 KiB
Markdown
382 lines
8.6 KiB
Markdown
---
|
||
license: mit
|
||
datasets:
|
||
- Mattimax/DACMini_Refined
|
||
- Mattimax/Camoscio-ITA
|
||
language:
|
||
- it
|
||
- en
|
||
library_name: transformers
|
||
tags:
|
||
- DAC
|
||
- M.INC.
|
||
- conversational
|
||
---
|
||
|
||
## ☕ Support the project
|
||
|
||
[](https://www.buymeacoffee.com/marzomattye)
|
||
|
||
## 🇮🇹 ITALIANO
|
||
|
||
# 📘 Model Card — Mattimax/DAC5-3B
|
||
|
||
## 🧠 Informazioni Generali
|
||
|
||
* **Nome:** Mattimax/DAC5-3B
|
||
* **Serie:** DAC (DATA-AI Chat) – 5ª versione
|
||
* **Autore:** Mattimax
|
||
* **Research Lab / Azienda:** MINC01
|
||
* **Base Model:** Qwen – Qwen2.5-3B-Instruct
|
||
|
||
DAC5-3B è attualmente il modello più avanzato e sperimentale della serie DAC, progettato per massimizzare qualità conversazionale, integrandosi al meglio con **server MCP** e performance tecnica su architettura 3B.
|
||
|
||
---
|
||
|
||
# 🏗 Architettura Tecnica
|
||
|
||
### Core Architecture
|
||
|
||
* **Architettura:** Qwen2ForCausalLM
|
||
* **Parametri:** ~3B
|
||
* **Numero layer:** 36
|
||
* **Hidden size:** 2048
|
||
* **Intermediate size:** 11008
|
||
* **Attention heads:** 16
|
||
* **Key/Value heads (GQA):** 2
|
||
* **Attivazione:** SiLU
|
||
* **Norm:** RMSNorm (eps 1e-6)
|
||
* **Tie word embeddings:** Yes
|
||
|
||
### Attention
|
||
|
||
* Full attention su tutti i 36 layer
|
||
* Attention dropout: 0.0
|
||
* Sliding window: disabilitato
|
||
* GQA (Grouped Query Attention) → maggiore efficienza memoria
|
||
|
||
### Positional Encoding
|
||
|
||
* **Max position embeddings:** 32768
|
||
* **RoPE theta:** 1,000,000
|
||
* RoPE scaling: None
|
||
|
||
### Precision & Performance
|
||
|
||
* Torch dtype: bfloat16
|
||
* Quantizzazione training: 4-bit (NF4)
|
||
* Cache abilitata per inference
|
||
* Ottimizzato con Unsloth (fixed build 2026.2.1)
|
||
|
||
### Tokenizer
|
||
|
||
* **Vocab size:** 151,936
|
||
* EOS token id: 151645
|
||
* PAD token id: 151654
|
||
|
||
---
|
||
|
||
# 🎯 Obiettivo del Modello
|
||
|
||
DAC5-3B è stato progettato per:
|
||
|
||
* 🇮🇹 Massima qualità in italiano
|
||
* ⚡ Alta efficienza su GPU consumer
|
||
* 🧩 Conversazione coerente multi-turn
|
||
* 🛠️ Supporto tecnico e coding leggero
|
||
* 🧠 Migliore stabilità rispetto ai DAC precedenti
|
||
|
||
È un modello orientato a sviluppatori indipendenti, maker e sistemi offline (come OpenClaw, Claude Code, OpenCode, ecc...)
|
||
|
||
---
|
||
|
||
# 📚 Dataset & Specializzazione
|
||
|
||
Il fine-tuning supervisionato è stato effettuato su un mix altamente selezionato di dataset italiani:
|
||
|
||
* Camoscio-ITA
|
||
* DACMini Refined
|
||
* Conversazioni sintetiche italiane ad alta qualità
|
||
|
||
### Strategia
|
||
|
||
* Dataset limitato ma ad alta densità informativa (~20k esempi)
|
||
* Minimizzazione del rumore
|
||
* Focus su chiarezza e coerenza
|
||
* Riduzione delle risposte generiche tipiche dei 3B
|
||
|
||
---
|
||
|
||
# 🚀 Capacità Principali
|
||
|
||
DAC5-3B eccelle in:
|
||
|
||
* Spiegazioni tecniche
|
||
* Scrittura strutturata
|
||
* Programmazione livello medio
|
||
* Traduzione IT ↔ EN
|
||
* Brainstorming progettuale
|
||
* Assistenti locali offline
|
||
* Supporto allo studio
|
||
|
||
---
|
||
|
||
# 📊 Differenze rispetto ai DAC precedenti
|
||
|
||
✔ Maggiore stabilità nelle risposte lunghe
|
||
✔ Meno ripetizioni
|
||
✔ Migliore controllo del tono
|
||
✔ Risposte più dirette
|
||
✔ Migliore allineamento alle istruzioni
|
||
|
||
DAC5 rappresenta il punto più alto raggiunto finora nella serie.
|
||
|
||
---
|
||
|
||
# ⚠️ Limitazioni
|
||
|
||
* Contesto di training effettivo: 1024 token
|
||
* Non ottimizzato per tool calling complesso
|
||
* Non specializzato in matematica avanzata
|
||
* Può degradare su reasoning multi-step molto profondo
|
||
* Modello sperimentale
|
||
|
||
---
|
||
|
||
# 💻 Requisiti Hardware
|
||
|
||
### Inference consigliata
|
||
|
||
* GPU 6–8GB VRAM (quantizzato)
|
||
* Oppure CPU moderna con GGUF
|
||
|
||
Compatibile con:
|
||
|
||
* PC consumer
|
||
* Mini workstation
|
||
* Sistemi edge
|
||
* Setup locali offline
|
||
|
||
---
|
||
|
||
# 🔬 Filosofia DAC
|
||
|
||
La serie DAC nasce con l'obiettivo di:
|
||
|
||
> Spingere al massimo modelli compatti, ottimizzando qualità reale invece di scalare solo i parametri.
|
||
|
||
DAC5-3B è il risultato più maturo di questa filosofia:
|
||
qualità elevata su architettura 3B con risorse contenute.
|
||
|
||
---
|
||
|
||
# 🧪 Stato del Modello
|
||
|
||
🟡 **Sperimentale ma stabile**
|
||
È il miglior modello della serie DAC fino ad oggi, ma rimane parte di un ciclo evolutivo continuo.
|
||
|
||
---
|
||
|
||
## 📚 Citation
|
||
|
||
Se utilizzi **Mattimax/DAC5-3B** nei tuoi lavori di ricerca, progetti o pubblicazioni, puoi citarlo nel seguente modo:
|
||
|
||
```bibtex
|
||
@misc{mattimax_dac5_3b_2026,
|
||
author = {Mattimax},
|
||
title = {DAC5-3B: Fifth Iteration of the Dynamic Adaptive Core Series},
|
||
year = {2026},
|
||
publisher = {Hugging Face},
|
||
organization = {MINC01},
|
||
note = {Experimental Italian-specialized 3B language model},
|
||
url = {https://huggingface.co/Mattimax/DAC5-3B}
|
||
}
|
||
```
|
||
|
||
Citazione testuale breve:
|
||
|
||
> Mattimax. *DAC5-3B: Fifth Iteration of the Dynamic Adaptive Core Series*. 2026. MINC01 Research Lab.
|
||
|
||
---
|
||
|
||
---
|
||
|
||
## 🇬🇧 ENGLISH
|
||
|
||
# 📘 Model Card — Mattimax/DAC5-3B
|
||
|
||
## 🧠 General Information
|
||
|
||
* **Name:** Mattimax/DAC5-3B
|
||
* **Series:** DAC (DATA-AI Chat) – 5th version
|
||
* **Author:** Mattimax
|
||
* **Research Lab / Company:** MINC01
|
||
* **Base Model:** Qwen – Qwen2.5-3B-Instruct
|
||
|
||
DAC5-3B is currently the most advanced and experimental model in the DAC series, designed to maximize conversational quality, integrating better with **MCP servers** and technical performance on a 3B architecture.
|
||
|
||
---
|
||
|
||
# 🏗 Technical Architecture
|
||
|
||
### Core Architecture
|
||
|
||
* **Architecture:** Qwen2ForCausalLM
|
||
* **Parameters:** ~3B
|
||
* **Number of layers:** 36
|
||
* **Hidden size:** 2048
|
||
* **Intermediate size:** 11008
|
||
* **Attention heads:** 16
|
||
* **Key/Value heads (GQA):** 2
|
||
* **Activation:** SiLU
|
||
* **Norm:** RMSNorm (eps 1e-6)
|
||
* **Tie word embeddings:** Yes
|
||
|
||
### Attention
|
||
|
||
* Full attention across all 36 layers
|
||
* Attention dropout: 0.0
|
||
* Sliding window: disabled
|
||
* GQA (Grouped Query Attention) → improved memory efficiency
|
||
|
||
### Positional Encoding
|
||
|
||
* **Max position embeddings:** 32768
|
||
* **RoPE theta:** 1,000,000
|
||
* RoPE scaling: None
|
||
|
||
### Precision & Performance
|
||
|
||
* Torch dtype: bfloat16
|
||
* Training quantization: 4-bit (NF4)
|
||
* Cache enabled for inference
|
||
* Optimized with Unsloth (fixed build 2026.2.1)
|
||
|
||
### Tokenizer
|
||
|
||
* **Vocab size:** 151,936
|
||
* EOS token id: 151645
|
||
* PAD token id: 151654
|
||
|
||
---
|
||
|
||
# 🎯 Model Objective
|
||
|
||
DAC5-3B was designed for:
|
||
|
||
* 🇮🇹 Maximum Italian language quality
|
||
* ⚡ High efficiency on consumer GPUs
|
||
* 🧩 Coherent multi-turn conversations
|
||
* 🛠️ Technical support and light coding
|
||
* 🧠 Improved stability compared to previous DAC versions
|
||
|
||
It is oriented toward independent developers, makers, and offline systems (such as OpenClaw, Claude Code, OpenCode, etc.).
|
||
|
||
---
|
||
|
||
# 📚 Dataset & Specialization
|
||
|
||
Supervised fine-tuning was performed on a highly curated mix of Italian datasets:
|
||
|
||
* Camoscio-ITA
|
||
* DACMini Refined
|
||
* High-quality synthetic Italian conversations
|
||
|
||
### Strategy
|
||
|
||
* Limited but high-density dataset (~20k samples)
|
||
* Noise minimization
|
||
* Focus on clarity and coherence
|
||
* Reduction of generic 3B-style responses
|
||
|
||
---
|
||
|
||
# 🚀 Core Capabilities
|
||
|
||
DAC5-3B excels at:
|
||
|
||
* Technical explanations
|
||
* Structured writing
|
||
* Intermediate-level programming
|
||
* IT ↔ EN translation
|
||
* Project brainstorming
|
||
* Offline local assistants
|
||
* Study support
|
||
|
||
---
|
||
|
||
# 📊 Differences from Previous DAC Versions
|
||
|
||
✔ Greater stability in long responses
|
||
✔ Fewer repetitions
|
||
✔ Better tone control
|
||
✔ More direct answers
|
||
✔ Improved instruction alignment
|
||
|
||
DAC5 represents the highest point reached so far in the series.
|
||
|
||
---
|
||
|
||
# ⚠️ Limitations
|
||
|
||
* Effective training context: 1024 tokens
|
||
* Not optimized for advanced tool calling
|
||
* Not specialized in advanced mathematics
|
||
* May degrade in very deep multi-step reasoning
|
||
* Experimental model
|
||
|
||
---
|
||
|
||
# 💻 Hardware Requirements
|
||
|
||
### Recommended Inference
|
||
|
||
* 6–8GB VRAM GPU (quantized)
|
||
* Or modern CPU with GGUF
|
||
|
||
Compatible with:
|
||
|
||
* Consumer PCs
|
||
* Mini workstations
|
||
* Edge systems
|
||
* Offline local setups
|
||
|
||
---
|
||
|
||
# 🔬 DAC Philosophy
|
||
|
||
The DAC series was created with the goal of:
|
||
|
||
> Pushing compact models to their limits, optimizing real quality instead of merely scaling parameters.
|
||
|
||
DAC5-3B is the most mature result of this philosophy:
|
||
high quality on a 3B architecture with limited resources.
|
||
|
||
---
|
||
|
||
# 🧪 Model Status
|
||
|
||
🟡 **Experimental but stable**
|
||
It is the best model in the DAC series to date, but remains part of an ongoing evolutionary cycle.
|
||
|
||
---
|
||
|
||
## 📚 Citation
|
||
|
||
If you use **Mattimax/DAC5-3B** in research work, projects, or publications, you may cite it as follows:
|
||
|
||
```bibtex
|
||
@misc{mattimax_dac5_3b_2026,
|
||
author = {Mattimax},
|
||
title = {DAC5-3B: Fifth Iteration of the Dynamic Adaptive Core Series},
|
||
year = {2026},
|
||
publisher = {Hugging Face},
|
||
organization = {MINC01},
|
||
note = {Experimental Italian-specialized 3B language model},
|
||
url = {https://huggingface.co/Mattimax/DAC5-3B}
|
||
}
|
||
```
|
||
|
||
Short textual citation:
|
||
|
||
> Mattimax. *DAC5-3B: Fifth Iteration of the Dynamic Adaptive Core Series*. 2026. MINC01 Research Lab. |