Files
Qwen3-1.7B-Libra-MF/README.md
ModelHub XC dafc35abad 初始化项目,由ModelHub XC社区提供模型
Model: surogate/Qwen3-1.7B-Libra-MF
Source: Original Platform
2026-08-07 17:55:18 +08:00

235 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
base_model: Qwen/Qwen3-1.7B
language:
- ro
library_name: transformers
pipeline_tag: text-generation
tags:
- accounting
- romanian
- mijloace-fixe
- fixed-assets
- structured-extraction
- column-mapping
- lora
- merged
- conversational
---
# Qwen3-1.7B-Libra-MF
Qwen3-1.7B fine-tuned to read Romanian **registre de mijloace fixe** (fixed-asset registers) in any
surface form and emit a **column-mapping recipe** as structured JSON. A deterministic post-processor
consumes the recipe and produces, per 3-digit asset category, the six accounting totals. LoRA SFT,
then merged back into a single 1.7B checkpoint for drop-in inference.
Trained on `surogate/mf-dataset`.
## Business use case
Romanian accounting software (Soft1, Saga, Mentor, SmartBill, custom Excel exports) emits the
fixed-asset register in a dozen incompatible layouts. For each asset category an accountant needs the
**six-field totals line**:
- **Valoare intrare** (entry value), **Valoare modernizări** (improvements), **Valoare de inventar**
(inventory value), **Valoare amortizată** (accumulated depreciation), **Amortizare lunară** (monthly
depreciation), **Valoare rămasă** (net book value).
That mapping is not fixed. Across registers:
- Headers differ per software (`Valoare intrare` vs `Valoare de intrare`; `Amortizare inregistrata`
vs `Uzura`; `Val. ramasa (neamortizata)` vs `Valoare ramasa`).
- Registers come in **two shapes**: *grouped* (section headers like `212 CONSTRUCTII` + a `Total pe …`
subtotal; the category comes from the section header, there is no `Cont` column) and *column*
(a per-row `Cont`/`Categorie` column; the category comes from that cell).
- **Trap columns** look right but aren't: a bare `Valoare`, the monthly `Amortizare lunară` vs the
cumulative `Valoare amortizată`, or decoy integer columns (`Durata funct.`, `Luni rămase`).
- The text arrives **collapsed, OCR-mangled, headerless, multi-line-header, English-mixed, …**
This model reads the raw extracted text and emits a JSON recipe naming exactly which column index (and
its header text) plays each of the 8 roles. **The model handles layout variance; the post-processor
handles the arithmetic** (per-category sums, the accounting identities, and a cross-check).
## Why a dedicated SLM
General GPT-4-class models on Romanian registre repeatedly:
| where big models fail | what they output | why it matters |
|---|---|---|
| Confuse **monthly** vs **cumulative** depreciation | `amortizare_luna` points at the cumulative column | every monthly total is wrong |
| Pick a bare **`Valoare`** trap column | wrong inventory/entry value | category totals don't reconcile |
| Fabricate headers on **headerless** layouts | invented column names | indexed lookup returns nothing |
| Map a value role to an **integer decoy** (`Durata`, `Luni`) | a duration counted as money | totals inflate |
| Emit a `cont` column on a **grouped** register | category leaks between sections | a 215 asset lands under 211 |
The disambiguating information is in the input every time; the problem is attending to it. A 1.7B model
trained on ~5,750 examples across 12 surface formats does this at a fraction of the inference cost.
## How the model + post-processor split work
The model emits, per role, `{header, index}` (or `null` if the column is absent). `cont = null` ⇒
*grouped* shape (category from section headers); `cont` present ⇒ *column* shape. The deterministic
post-processor (`mf_apply`) then:
1. splits each row into cells (by separator, or by typed-token reconstruction for collapsed text),
2. anchors each role to a column by **header match**, falling back to the model's index,
3. sums the six fields per category, derives `modernizări = inventar − intrare`,
4. cross-checks against the register's own printed `Total pe …` / `Totaluri …` line and against the
identity `inventar = amortizată + rămasă` (except terenuri/211), emitting an `Observatii` note.
## Eval results (shipped merged checkpoint)
| eval set | size | score |
|---|---|---|
| **Real client registers** (end-to-end 6-field totals) | 7 | **7 / 7** (every register, every category) |
| **Held-out synthetic, all 12 formats** (end-to-end) | 360 | **95.3 %** |
| **Validation set** (model recipe vs ground-truth recipe, both applied) | 600 | **92.6 %** |
Per-format accuracy on the 360 held-out set (end-to-end totals):
| format | acc | | format | acc |
|---|---|---|---|---|
| canonical_markers | 100 % | | mixed_language | 100 % |
| cont_prefix | 100 % | | multi_line_header | 93.3 % |
| header_only | 100 % | | ocr_mangled | 96.7 % |
| fixed_width | 100 % | | csv | 90.0 % |
| markdown | 100 % | | **pdf_copy** (collapsed) | 86.7 % |
| tsv | 100 % | | **headerless** | 76.7 % |
## Worked examples
**Ametech (grouped register, no `cont`; category from section headers):**
```
ANTET … | Denumire imobilizare | Valoare intrare | … | Amortizare inregistrata | Amortizare lunara | Val. ramasa
212 CONSTRUCTII
1 APARTAMENT 118 MUN.BUC, STR 1 805 972.00 … 149 935.72 3 762.44 1 656 036.28
…
```
```json
{ "coloane": {
"denumire": {"header": "Denumire imobilizare", "index": 1},
"valoare_intrare": {"header": "Valoare intrare", "index": 3},
"valoare_modernizari": {"header": "Valoare modernizari", "index": 4},
"valoare_inventar": {"header": "Valoare de inventar", "index": 5},
"valoare_amortizata": {"header": "Amortizare inregistrata", "index": 9},
"amortizare_luna": {"header": "Amortizare lunara", "index": 8},
"valoare_ramasa": {"header": "Val. ramasa", "index": 6},
"cont": null } }
```
`cont = null` → the post-processor takes the category from each `NNN …` section header.
**Algorithm (column register, per-row `Cont`; category from that cell):**
```
A/A Cod | Denumire mijloc fix | Valoare intrare | … | Cont de mijloace fixe | …
1 ONORARIU … 01/06/2017 185,45 … 208 …
```
```json
{ "coloane": {
"denumire": {"header": "Denumire mijloc fix", "index": 1},
"valoare_intrare": {"header": "Valoare intrare", "index": 2},
"valoare_modernizari": {"header": "Valoare modernizari", "index": 4},
"valoare_inventar": {"header": "Valoare de inventar", "index": 5},
"valoare_amortizata": {"header": "Amortizare inregistrata", "index": 6},
"amortizare_luna": {"header": "din care amortizat in luna", "index": 7},
"valoare_ramasa": {"header": "Val. ramasa (neamortizata)", "index": 8},
"cont": {"header": "Cont de mijloace fixe", "index": 10} } }
```
**Headerless layout (no header row, columns by position):**
```
1 INVESTITIE IMOBILIARA SIGMA 25.06.2024 63.421,57 9.509,52 72.931,09 10.824,39 607,76 62.106,70
…
```
```json
{ "coloane": {
"denumire": {"header": "", "index": 1}, "valoare_intrare": {"header": "", "index": 4},
"valoare_modernizari": {"header": "", "index": 6}, "valoare_inventar": {"header": "", "index": 7},
"valoare_amortizata": {"header": "", "index": 11}, "amortizare_luna": {"header": "", "index": 10},
"valoare_ramasa": {"header": "", "index": 8}, "cont": null } }
```
With no header text, the model emits `"header": ""` and locates columns by their numeric position.
## Output schema
| field | type | content |
|---|---|---|
| `coloane.<role>` | `{header, index}` or `null` | one entry per role |
| roles | list | `denumire, valoare_intrare, valoare_modernizari, valoare_inventar, valoare_amortizata, amortizare_luna, valoare_ramasa, cont` |
| `header` | str | the column's header text **as it appears** (`""` if headerless); the robust anchor |
| `index` | int | 0-based column position; the fallback anchor |
| `cont = null` | flag | *grouped* register (category from section headers) |
## Quick start
**transformers**
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("surogate/Qwen3-1.7B-Libra-MF")
model = AutoModelForCausalLM.from_pretrained("surogate/Qwen3-1.7B-Libra-MF",
torch_dtype=torch.bfloat16, device_map="auto")
from datasets import load_dataset
SYSTEM = load_dataset("surogate/mf-dataset", split="train[:1]")[0]["instruction"]
user_text = open("my_registru.txt").read()
prompt = tok.apply_chat_template(
[{"role": "system", "content": SYSTEM}, {"role": "user", "content": user_text}],
tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=768, do_sample=False)
print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
```
**vLLM**
```bash
vllm serve surogate/Qwen3-1.7B-Libra-MF --max-model-len 4096 --gpu-memory-utilization 0.6
```
Then POST `{system_prompt}\n{registru text}` with `temperature: 0`, `max_tokens: 768`.
## Training details
| field | value |
|---|---|
| base model | `Qwen/Qwen3-1.7B` |
| method | LoRA SFT, merged into base for shipping |
| recipe | fp8-hybrid |
| batch | per_device 1 × grad-accum 8 (effective 8), sequence_len 2048 |
| LR | 5e-5 cosine, warmup ratio 0.05 |
| LoRA rank / alpha / dropout | 16 / 32 / 0.15 |
| LoRA target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| dataset | `surogate/mf-dataset` (5,750 train + 600 val) |
| framework | surogate sft |
> Note on batch: peak memory scales with the **per_device microbatch**, not the effective batch
> (accumulation is sequential). per_device 1 keeps the graph small and stable; effective batch 8 is
> reached via accumulation.
## Limitations
- **Romanian only.** The `mixed_language` format introduces some English headers, but the model is not
robust to fully English registers.
- **Categories 205 to 215** (the standard 3-digit fixed-asset accounts) are the focus.
- **Inputs are token-budgeted to ≤ 2048** (no truncation in training). Very long registers should be
passed first-N-rows-windowed (a header + a sample of rows is all the column mapping needs).
- **Collapsed `pdf_copy` and headerless** are the hardest forms (86.7 % / 76.7 %): space-collapsed
text is information-lossy, and headerless requires pure positional reasoning. For PDFs, the
production path supplies an x-clustered cell grid as an aid, which sidesteps the collapse.
- **The post-processor is not part of this checkpoint.** Without it the model output is a recipe, not
the totals.
## License
Apache 2.0. Inherits from `Qwen/Qwen3-1.7B`. Synthetic training data plus 7 anonymized real-register
layout anchors.
## Citation
```bibtex
@misc{qwen3-1.7b-libra-mf,
title = {Qwen3-1.7B-Libra-MF: Romanian registru de mijloace fixe column-mapping extractor},
author = {Surogate},
year = {2026},
url = {https://huggingface.co/surogate/Qwen3-1.7B-Libra-MF}
}
```