--- license: apache-2.0 base_model: Qwen/Qwen3-1.7B language: - ro library_name: transformers pipeline_tag: text-generation tags: - accounting - romanian - mijloace-fixe - fixed-assets - structured-extraction - column-mapping - lora - merged - conversational --- # Qwen3-1.7B-Libra-MF Qwen3-1.7B fine-tuned to read Romanian **registre de mijloace fixe** (fixed-asset registers) in any surface form and emit a **column-mapping recipe** as structured JSON. A deterministic post-processor consumes the recipe and produces, per 3-digit asset category, the six accounting totals. LoRA SFT, then merged back into a single 1.7B checkpoint for drop-in inference. Trained on `surogate/mf-dataset`. ## Business use case Romanian accounting software (Soft1, Saga, Mentor, SmartBill, custom Excel exports) emits the fixed-asset register in a dozen incompatible layouts. For each asset category an accountant needs the **six-field totals line**: - **Valoare intrare** (entry value), **Valoare modernizări** (improvements), **Valoare de inventar** (inventory value), **Valoare amortizată** (accumulated depreciation), **Amortizare lunară** (monthly depreciation), **Valoare rămasă** (net book value). That mapping is not fixed. Across registers: - Headers differ per software (`Valoare intrare` vs `Valoare de intrare`; `Amortizare inregistrata` vs `Uzura`; `Val. ramasa (neamortizata)` vs `Valoare ramasa`). - Registers come in **two shapes**: *grouped* (section headers like `212 CONSTRUCTII` + a `Total pe …` subtotal; the category comes from the section header, there is no `Cont` column) and *column* (a per-row `Cont`/`Categorie` column; the category comes from that cell). - **Trap columns** look right but aren't: a bare `Valoare`, the monthly `Amortizare lunară` vs the cumulative `Valoare amortizată`, or decoy integer columns (`Durata funct.`, `Luni rămase`). - The text arrives **collapsed, OCR-mangled, headerless, multi-line-header, English-mixed, …** This model reads the raw extracted text and emits a JSON recipe naming exactly which column index (and its header text) plays each of the 8 roles. **The model handles layout variance; the post-processor handles the arithmetic** (per-category sums, the accounting identities, and a cross-check). ## Why a dedicated SLM General GPT-4-class models on Romanian registre repeatedly: | where big models fail | what they output | why it matters | |---|---|---| | Confuse **monthly** vs **cumulative** depreciation | `amortizare_luna` points at the cumulative column | every monthly total is wrong | | Pick a bare **`Valoare`** trap column | wrong inventory/entry value | category totals don't reconcile | | Fabricate headers on **headerless** layouts | invented column names | indexed lookup returns nothing | | Map a value role to an **integer decoy** (`Durata`, `Luni`) | a duration counted as money | totals inflate | | Emit a `cont` column on a **grouped** register | category leaks between sections | a 215 asset lands under 211 | The disambiguating information is in the input every time; the problem is attending to it. A 1.7B model trained on ~5,750 examples across 12 surface formats does this at a fraction of the inference cost. ## How the model + post-processor split work The model emits, per role, `{header, index}` (or `null` if the column is absent). `cont = null` ⇒ *grouped* shape (category from section headers); `cont` present ⇒ *column* shape. The deterministic post-processor (`mf_apply`) then: 1. splits each row into cells (by separator, or by typed-token reconstruction for collapsed text), 2. anchors each role to a column by **header match**, falling back to the model's index, 3. sums the six fields per category, derives `modernizări = inventar − intrare`, 4. cross-checks against the register's own printed `Total pe …` / `Totaluri …` line and against the identity `inventar = amortizată + rămasă` (except terenuri/211), emitting an `Observatii` note. ## Eval results (shipped merged checkpoint) | eval set | size | score | |---|---|---| | **Real client registers** (end-to-end 6-field totals) | 7 | **7 / 7** (every register, every category) | | **Held-out synthetic, all 12 formats** (end-to-end) | 360 | **95.3 %** | | **Validation set** (model recipe vs ground-truth recipe, both applied) | 600 | **92.6 %** | Per-format accuracy on the 360 held-out set (end-to-end totals): | format | acc | | format | acc | |---|---|---|---|---| | canonical_markers | 100 % | | mixed_language | 100 % | | cont_prefix | 100 % | | multi_line_header | 93.3 % | | header_only | 100 % | | ocr_mangled | 96.7 % | | fixed_width | 100 % | | csv | 90.0 % | | markdown | 100 % | | **pdf_copy** (collapsed) | 86.7 % | | tsv | 100 % | | **headerless** | 76.7 % | ## Worked examples **Ametech (grouped register, no `cont`; category from section headers):** ``` ANTET … | Denumire imobilizare | Valoare intrare | … | Amortizare inregistrata | Amortizare lunara | Val. ramasa 212 CONSTRUCTII 1 APARTAMENT 118 MUN.BUC, STR 1 805 972.00 … 149 935.72 3 762.44 1 656 036.28 … ``` ```json { "coloane": { "denumire": {"header": "Denumire imobilizare", "index": 1}, "valoare_intrare": {"header": "Valoare intrare", "index": 3}, "valoare_modernizari": {"header": "Valoare modernizari", "index": 4}, "valoare_inventar": {"header": "Valoare de inventar", "index": 5}, "valoare_amortizata": {"header": "Amortizare inregistrata", "index": 9}, "amortizare_luna": {"header": "Amortizare lunara", "index": 8}, "valoare_ramasa": {"header": "Val. ramasa", "index": 6}, "cont": null } } ``` `cont = null` → the post-processor takes the category from each `NNN …` section header. **Algorithm (column register, per-row `Cont`; category from that cell):** ``` A/A Cod | Denumire mijloc fix | Valoare intrare | … | Cont de mijloace fixe | … 1 ONORARIU … 01/06/2017 185,45 … 208 … ``` ```json { "coloane": { "denumire": {"header": "Denumire mijloc fix", "index": 1}, "valoare_intrare": {"header": "Valoare intrare", "index": 2}, "valoare_modernizari": {"header": "Valoare modernizari", "index": 4}, "valoare_inventar": {"header": "Valoare de inventar", "index": 5}, "valoare_amortizata": {"header": "Amortizare inregistrata", "index": 6}, "amortizare_luna": {"header": "din care amortizat in luna", "index": 7}, "valoare_ramasa": {"header": "Val. ramasa (neamortizata)", "index": 8}, "cont": {"header": "Cont de mijloace fixe", "index": 10} } } ``` **Headerless layout (no header row, columns by position):** ``` 1 INVESTITIE IMOBILIARA SIGMA 25.06.2024 63.421,57 9.509,52 72.931,09 10.824,39 607,76 62.106,70 … ``` ```json { "coloane": { "denumire": {"header": "", "index": 1}, "valoare_intrare": {"header": "", "index": 4}, "valoare_modernizari": {"header": "", "index": 6}, "valoare_inventar": {"header": "", "index": 7}, "valoare_amortizata": {"header": "", "index": 11}, "amortizare_luna": {"header": "", "index": 10}, "valoare_ramasa": {"header": "", "index": 8}, "cont": null } } ``` With no header text, the model emits `"header": ""` and locates columns by their numeric position. ## Output schema | field | type | content | |---|---|---| | `coloane.` | `{header, index}` or `null` | one entry per role | | roles | list | `denumire, valoare_intrare, valoare_modernizari, valoare_inventar, valoare_amortizata, amortizare_luna, valoare_ramasa, cont` | | `header` | str | the column's header text **as it appears** (`""` if headerless); the robust anchor | | `index` | int | 0-based column position; the fallback anchor | | `cont = null` | flag | *grouped* register (category from section headers) | ## Quick start **transformers** ```python from transformers import AutoModelForCausalLM, AutoTokenizer import torch tok = AutoTokenizer.from_pretrained("surogate/Qwen3-1.7B-Libra-MF") model = AutoModelForCausalLM.from_pretrained("surogate/Qwen3-1.7B-Libra-MF", torch_dtype=torch.bfloat16, device_map="auto") from datasets import load_dataset SYSTEM = load_dataset("surogate/mf-dataset", split="train[:1]")[0]["instruction"] user_text = open("my_registru.txt").read() prompt = tok.apply_chat_template( [{"role": "system", "content": SYSTEM}, {"role": "user", "content": user_text}], tokenize=False, add_generation_prompt=True, enable_thinking=False) inputs = tok(prompt, return_tensors="pt").to(model.device) out = model.generate(**inputs, max_new_tokens=768, do_sample=False) print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)) ``` **vLLM** ```bash vllm serve surogate/Qwen3-1.7B-Libra-MF --max-model-len 4096 --gpu-memory-utilization 0.6 ``` Then POST `{system_prompt}\n{registru text}` with `temperature: 0`, `max_tokens: 768`. ## Training details | field | value | |---|---| | base model | `Qwen/Qwen3-1.7B` | | method | LoRA SFT, merged into base for shipping | | recipe | fp8-hybrid | | batch | per_device 1 × grad-accum 8 (effective 8), sequence_len 2048 | | LR | 5e-5 cosine, warmup ratio 0.05 | | LoRA rank / alpha / dropout | 16 / 32 / 0.15 | | LoRA target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj | | dataset | `surogate/mf-dataset` (5,750 train + 600 val) | | framework | surogate sft | > Note on batch: peak memory scales with the **per_device microbatch**, not the effective batch > (accumulation is sequential). per_device 1 keeps the graph small and stable; effective batch 8 is > reached via accumulation. ## Limitations - **Romanian only.** The `mixed_language` format introduces some English headers, but the model is not robust to fully English registers. - **Categories 205 to 215** (the standard 3-digit fixed-asset accounts) are the focus. - **Inputs are token-budgeted to ≤ 2048** (no truncation in training). Very long registers should be passed first-N-rows-windowed (a header + a sample of rows is all the column mapping needs). - **Collapsed `pdf_copy` and headerless** are the hardest forms (86.7 % / 76.7 %): space-collapsed text is information-lossy, and headerless requires pure positional reasoning. For PDFs, the production path supplies an x-clustered cell grid as an aid, which sidesteps the collapse. - **The post-processor is not part of this checkpoint.** Without it the model output is a recipe, not the totals. ## License Apache 2.0. Inherits from `Qwen/Qwen3-1.7B`. Synthetic training data plus 7 anonymized real-register layout anchors. ## Citation ```bibtex @misc{qwen3-1.7b-libra-mf, title = {Qwen3-1.7B-Libra-MF: Romanian registru de mijloace fixe column-mapping extractor}, author = {Surogate}, year = {2026}, url = {https://huggingface.co/surogate/Qwen3-1.7B-Libra-MF} } ```