Files
Qwen3-1.7B-Libra-MF/README.md

235 lines
11 KiB
Markdown
Raw Permalink Normal View History

---
license: apache-2.0
base_model: Qwen/Qwen3-1.7B
language:
- ro
library_name: transformers
pipeline_tag: text-generation
tags:
- accounting
- romanian
- mijloace-fixe
- fixed-assets
- structured-extraction
- column-mapping
- lora
- merged
- conversational
---
# Qwen3-1.7B-Libra-MF
Qwen3-1.7B fine-tuned to read Romanian **registre de mijloace fixe** (fixed-asset registers) in any
surface form and emit a **column-mapping recipe** as structured JSON. A deterministic post-processor
consumes the recipe and produces, per 3-digit asset category, the six accounting totals. LoRA SFT,
then merged back into a single 1.7B checkpoint for drop-in inference.
Trained on `surogate/mf-dataset`.
## Business use case
Romanian accounting software (Soft1, Saga, Mentor, SmartBill, custom Excel exports) emits the
fixed-asset register in a dozen incompatible layouts. For each asset category an accountant needs the
**six-field totals line**:
- **Valoare intrare** (entry value), **Valoare modernizări** (improvements), **Valoare de inventar**
(inventory value), **Valoare amortizată** (accumulated depreciation), **Amortizare lunară** (monthly
depreciation), **Valoare rămasă** (net book value).
That mapping is not fixed. Across registers:
- Headers differ per software (`Valoare intrare` vs `Valoare de intrare`; `Amortizare inregistrata`
vs `Uzura`; `Val. ramasa (neamortizata)` vs `Valoare ramasa`).
- Registers come in **two shapes**: *grouped* (section headers like `212 CONSTRUCTII` + a `Total pe …`
subtotal; the category comes from the section header, there is no `Cont` column) and *column*
(a per-row `Cont`/`Categorie` column; the category comes from that cell).
- **Trap columns** look right but aren't: a bare `Valoare`, the monthly `Amortizare lunară` vs the
cumulative `Valoare amortizată`, or decoy integer columns (`Durata funct.`, `Luni rămase`).
- The text arrives **collapsed, OCR-mangled, headerless, multi-line-header, English-mixed, …**
This model reads the raw extracted text and emits a JSON recipe naming exactly which column index (and
its header text) plays each of the 8 roles. **The model handles layout variance; the post-processor
handles the arithmetic** (per-category sums, the accounting identities, and a cross-check).
## Why a dedicated SLM
General GPT-4-class models on Romanian registre repeatedly:
| where big models fail | what they output | why it matters |
|---|---|---|
| Confuse **monthly** vs **cumulative** depreciation | `amortizare_luna` points at the cumulative column | every monthly total is wrong |
| Pick a bare **`Valoare`** trap column | wrong inventory/entry value | category totals don't reconcile |
| Fabricate headers on **headerless** layouts | invented column names | indexed lookup returns nothing |
| Map a value role to an **integer decoy** (`Durata`, `Luni`) | a duration counted as money | totals inflate |
| Emit a `cont` column on a **grouped** register | category leaks between sections | a 215 asset lands under 211 |
The disambiguating information is in the input every time; the problem is attending to it. A 1.7B model
trained on ~5,750 examples across 12 surface formats does this at a fraction of the inference cost.
## How the model + post-processor split work
The model emits, per role, `{header, index}` (or `null` if the column is absent). `cont = null` ⇒
*grouped* shape (category from section headers); `cont` present ⇒ *column* shape. The deterministic
post-processor (`mf_apply`) then:
1. splits each row into cells (by separator, or by typed-token reconstruction for collapsed text),
2. anchors each role to a column by **header match**, falling back to the model's index,
3. sums the six fields per category, derives `modernizări = inventar − intrare`,
4. cross-checks against the register's own printed `Total pe …` / `Totaluri …` line and against the
identity `inventar = amortizată + rămasă` (except terenuri/211), emitting an `Observatii` note.
## Eval results (shipped merged checkpoint)
| eval set | size | score |
|---|---|---|
| **Real client registers** (end-to-end 6-field totals) | 7 | **7 / 7** (every register, every category) |
| **Held-out synthetic, all 12 formats** (end-to-end) | 360 | **95.3 %** |
| **Validation set** (model recipe vs ground-truth recipe, both applied) | 600 | **92.6 %** |
Per-format accuracy on the 360 held-out set (end-to-end totals):
| format | acc | | format | acc |
|---|---|---|---|---|
| canonical_markers | 100 % | | mixed_language | 100 % |
| cont_prefix | 100 % | | multi_line_header | 93.3 % |
| header_only | 100 % | | ocr_mangled | 96.7 % |
| fixed_width | 100 % | | csv | 90.0 % |
| markdown | 100 % | | **pdf_copy** (collapsed) | 86.7 % |
| tsv | 100 % | | **headerless** | 76.7 % |
## Worked examples
**Ametech (grouped register, no `cont`; category from section headers):**
```
ANTET … | Denumire imobilizare | Valoare intrare | … | Amortizare inregistrata | Amortizare lunara | Val. ramasa
212 CONSTRUCTII
1 APARTAMENT 118 MUN.BUC, STR 1 805 972.00 … 149 935.72 3 762.44 1 656 036.28
…
```
```json
{ "coloane": {
"denumire": {"header": "Denumire imobilizare", "index": 1},
"valoare_intrare": {"header": "Valoare intrare", "index": 3},
"valoare_modernizari": {"header": "Valoare modernizari", "index": 4},
"valoare_inventar": {"header": "Valoare de inventar", "index": 5},
"valoare_amortizata": {"header": "Amortizare inregistrata", "index": 9},
"amortizare_luna": {"header": "Amortizare lunara", "index": 8},
"valoare_ramasa": {"header": "Val. ramasa", "index": 6},
"cont": null } }
```
`cont = null` → the post-processor takes the category from each `NNN …` section header.
**Algorithm (column register, per-row `Cont`; category from that cell):**
```
A/A Cod | Denumire mijloc fix | Valoare intrare | … | Cont de mijloace fixe | …
1 ONORARIU … 01/06/2017 185,45 … 208 …
```
```json
{ "coloane": {
"denumire": {"header": "Denumire mijloc fix", "index": 1},
"valoare_intrare": {"header": "Valoare intrare", "index": 2},
"valoare_modernizari": {"header": "Valoare modernizari", "index": 4},
"valoare_inventar": {"header": "Valoare de inventar", "index": 5},
"valoare_amortizata": {"header": "Amortizare inregistrata", "index": 6},
"amortizare_luna": {"header": "din care amortizat in luna", "index": 7},
"valoare_ramasa": {"header": "Val. ramasa (neamortizata)", "index": 8},
"cont": {"header": "Cont de mijloace fixe", "index": 10} } }
```
**Headerless layout (no header row, columns by position):**
```
1 INVESTITIE IMOBILIARA SIGMA 25.06.2024 63.421,57 9.509,52 72.931,09 10.824,39 607,76 62.106,70
…
```
```json
{ "coloane": {
"denumire": {"header": "", "index": 1}, "valoare_intrare": {"header": "", "index": 4},
"valoare_modernizari": {"header": "", "index": 6}, "valoare_inventar": {"header": "", "index": 7},
"valoare_amortizata": {"header": "", "index": 11}, "amortizare_luna": {"header": "", "index": 10},
"valoare_ramasa": {"header": "", "index": 8}, "cont": null } }
```
With no header text, the model emits `"header": ""` and locates columns by their numeric position.
## Output schema
| field | type | content |
|---|---|---|
| `coloane.<role>` | `{header, index}` or `null` | one entry per role |
| roles | list | `denumire, valoare_intrare, valoare_modernizari, valoare_inventar, valoare_amortizata, amortizare_luna, valoare_ramasa, cont` |
| `header` | str | the column's header text **as it appears** (`""` if headerless); the robust anchor |
| `index` | int | 0-based column position; the fallback anchor |
| `cont = null` | flag | *grouped* register (category from section headers) |
## Quick start
**transformers**
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("surogate/Qwen3-1.7B-Libra-MF")
model = AutoModelForCausalLM.from_pretrained("surogate/Qwen3-1.7B-Libra-MF",
torch_dtype=torch.bfloat16, device_map="auto")
from datasets import load_dataset
SYSTEM = load_dataset("surogate/mf-dataset", split="train[:1]")[0]["instruction"]
user_text = open("my_registru.txt").read()
prompt = tok.apply_chat_template(
[{"role": "system", "content": SYSTEM}, {"role": "user", "content": user_text}],
tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=768, do_sample=False)
print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
```
**vLLM**
```bash
vllm serve surogate/Qwen3-1.7B-Libra-MF --max-model-len 4096 --gpu-memory-utilization 0.6
```
Then POST `{system_prompt}\n{registru text}` with `temperature: 0`, `max_tokens: 768`.
## Training details
| field | value |
|---|---|
| base model | `Qwen/Qwen3-1.7B` |
| method | LoRA SFT, merged into base for shipping |
| recipe | fp8-hybrid |
| batch | per_device 1 × grad-accum 8 (effective 8), sequence_len 2048 |
| LR | 5e-5 cosine, warmup ratio 0.05 |
| LoRA rank / alpha / dropout | 16 / 32 / 0.15 |
| LoRA target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| dataset | `surogate/mf-dataset` (5,750 train + 600 val) |
| framework | surogate sft |
> Note on batch: peak memory scales with the **per_device microbatch**, not the effective batch
> (accumulation is sequential). per_device 1 keeps the graph small and stable; effective batch 8 is
> reached via accumulation.
## Limitations
- **Romanian only.** The `mixed_language` format introduces some English headers, but the model is not
robust to fully English registers.
- **Categories 205 to 215** (the standard 3-digit fixed-asset accounts) are the focus.
- **Inputs are token-budgeted to ≤ 2048** (no truncation in training). Very long registers should be
passed first-N-rows-windowed (a header + a sample of rows is all the column mapping needs).
- **Collapsed `pdf_copy` and headerless** are the hardest forms (86.7 % / 76.7 %): space-collapsed
text is information-lossy, and headerless requires pure positional reasoning. For PDFs, the
production path supplies an x-clustered cell grid as an aid, which sidesteps the collapse.
- **The post-processor is not part of this checkpoint.** Without it the model output is a recipe, not
the totals.
## License
Apache 2.0. Inherits from `Qwen/Qwen3-1.7B`. Synthetic training data plus 7 anonymized real-register
layout anchors.
## Citation
```bibtex
@misc{qwen3-1.7b-libra-mf,
title = {Qwen3-1.7B-Libra-MF: Romanian registru de mijloace fixe column-mapping extractor},
author = {Surogate},
year = {2026},
url = {https://huggingface.co/surogate/Qwen3-1.7B-Libra-MF}
}
```