165 lines
5.7 KiB
Markdown
165 lines
5.7 KiB
Markdown
|
|
---
|
|||
|
|
language:
|
|||
|
|
- en
|
|||
|
|
- it
|
|||
|
|
license: apache-2.0
|
|||
|
|
library_name: transformers
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
base_model: Qwen/Qwen3-4B
|
|||
|
|
datasets:
|
|||
|
|
- Italianhype/Blum-Finance-Reasoning
|
|||
|
|
tags:
|
|||
|
|
- finance
|
|||
|
|
- financial-reasoning
|
|||
|
|
- investment-research
|
|||
|
|
- risk-management
|
|||
|
|
- explainable-ai
|
|||
|
|
- qwen3
|
|||
|
|
- transformers
|
|||
|
|
model-index:
|
|||
|
|
- name: BLUM Finance 4B
|
|||
|
|
results:
|
|||
|
|
- task:
|
|||
|
|
type: text-generation
|
|||
|
|
name: Multiple-choice financial reasoning
|
|||
|
|
dataset:
|
|||
|
|
name: MMLU Finance and Business (stratified 100 per subject)
|
|||
|
|
type: cais/mmlu
|
|||
|
|
config: finance_business_7_subjects
|
|||
|
|
split: test
|
|||
|
|
metrics:
|
|||
|
|
- type: accuracy
|
|||
|
|
value: 0.74857143
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# BLUM Finance 4B — Transformers Benchmark Release
|
|||
|
|
|
|||
|
|
BLUM Finance 4B is an open, evidence-bound financial reasoning model. This
|
|||
|
|
repository is the portable BF16 Transformers release used for public benchmark
|
|||
|
|
submission. It contains the same trained LoRA delta as the audited MLX release,
|
|||
|
|
mapped and fused into the exact `Qwen/Qwen3-4B` base architecture.
|
|||
|
|
|
|||
|
|
## Intended Use
|
|||
|
|
|
|||
|
|
The model structures supplied point-in-time evidence into:
|
|||
|
|
|
|||
|
|
- balanced bull and bear cases;
|
|||
|
|
- explicit risks and invalidation conditions;
|
|||
|
|
- conservative confidence;
|
|||
|
|
- monitoring conditions and abstention.
|
|||
|
|
|
|||
|
|
It is not a price oracle, broker, trade executor or proof of market alpha.
|
|||
|
|
|
|||
|
|
## Load
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
import torch
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|||
|
|
|
|||
|
|
model_id = "Italianhype/Blum-Finance-4B"
|
|||
|
|
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|||
|
|
model_id,
|
|||
|
|
dtype=torch.bfloat16,
|
|||
|
|
device_map="auto",
|
|||
|
|
)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
No `trust_remote_code=True` is required.
|
|||
|
|
|
|||
|
|
## Governed Continual Memory
|
|||
|
|
|
|||
|
|
This repository also ships the installable `blum-finance` package. It provides
|
|||
|
|
schema-validated inference, auditable local memory and explicit opt-in community
|
|||
|
|
contributions:
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
pip install "git+https://huggingface.co/Italianhype/Blum-Finance-4B"
|
|||
|
|
blum-contribute example.json --output contribution.json --consent
|
|||
|
|
blum-memory-add contribution.json
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Local retrieval accepts only mature, source-verified outcomes observable before
|
|||
|
|
the new request timestamp. Memories are labeled historical analogies and cannot
|
|||
|
|
replace current evidence. Adding `--push` opens a pull request to the quarantine
|
|||
|
|
dataset; inference sends no telemetry and anonymous input never mutates released
|
|||
|
|
weights. See `CONTRIBUTING.md` for the evidence contract.
|
|||
|
|
|
|||
|
|
## Lineage
|
|||
|
|
|
|||
|
|
- Base: `Qwen/Qwen3-4B`
|
|||
|
|
- Base revision: `1cfa9a7208912126459214e8b04321603b3df60c`
|
|||
|
|
- Dataset: `Italianhype/Blum-Finance-Reasoning`
|
|||
|
|
- Dataset revision: `76ad77699d498fc930daf02e452fe3ec8b490f90`
|
|||
|
|
- MLX adapter revision: `ea297ba88ab008e97104b0c118103eef2f8f9ec1`
|
|||
|
|
- LoRA rank: `8`
|
|||
|
|
- MLX scale: `20`
|
|||
|
|
- Equivalent PEFT alpha: `160`
|
|||
|
|
- Adapted layers: `20–35`
|
|||
|
|
|
|||
|
|
The conversion transposes MLX A/B matrices into PEFT orientation and preserves
|
|||
|
|
the update `20 × Bᵀ × Aᵀ`. All 112 expected modules were fused; none were
|
|||
|
|
missing. The PEFT adapter and merged model produced identical deterministic
|
|||
|
|
smoke-test output.
|
|||
|
|
|
|||
|
|
## Evaluation Status
|
|||
|
|
|
|||
|
|
The original MLX release passed the 53-example BLUM temporal reasoning test with
|
|||
|
|
96.26% aggregate task-contract score, 100% structured validity and 92.86%
|
|||
|
|
no-fabrication. These metrics measure BLUM schema adherence and grounding, not
|
|||
|
|
general intelligence or trading performance.
|
|||
|
|
|
|||
|
|
The portable release is published at immutable revision
|
|||
|
|
`ad6f5cec7f729370d2976d8c78983521cb37ca83` and tagged
|
|||
|
|
`benchmark-submission-v1`.
|
|||
|
|
|
|||
|
|
The five-example portable conversion smoke test scored 97.32% on the internal
|
|||
|
|
task contract, including 100% structured validity and 92.86% no-fabrication.
|
|||
|
|
The sample is too small for a robust capability claim.
|
|||
|
|
|
|||
|
|
Automated submissions were attempted for the Hugging Face Open LLM Leaderboard
|
|||
|
|
and the FinOS Open Financial LLM Leaderboard. Both external validators rejected
|
|||
|
|
Qwen3 as requiring `trust_remote_code=True`; Transformers 4.57.6 loads this
|
|||
|
|
repository without remote code. No official leaderboard score is claimed.
|
|||
|
|
|
|||
|
|
A deterministic stratified community evaluation on seven finance/business
|
|||
|
|
subjects of canonical `cais/mmlu` scored **74.86% accuracy** over 700 test
|
|||
|
|
questions (100 per subject), with a Wilson 95% confidence interval of
|
|||
|
|
71.51–77.93%. It used five-shot answer-token logit scoring.
|
|||
|
|
|
|||
|
|
| Subject | Accuracy | Samples |
|
|||
|
|
|---|---:|---:|
|
|||
|
|
| Business ethics | 77% | 100 |
|
|||
|
|
| Econometrics | 65% | 100 |
|
|||
|
|
| High-school macroeconomics | 70% | 100 |
|
|||
|
|
| High-school microeconomics | 87% | 100 |
|
|||
|
|
| Management | 87% | 100 |
|
|||
|
|
| Marketing | 92% | 100 |
|
|||
|
|
| Professional accounting | 46% | 100 |
|
|||
|
|
|
|||
|
|
This is an author-run community evaluation, not the full canonical MMLU suite
|
|||
|
|
and not an official leaderboard result. The deterministic 100-row cap gives
|
|||
|
|
each subject equal weight but can differ from a full-split result. Dataset
|
|||
|
|
revision: `c30699e8356da336a370243923dbaf21066bb9fe`.
|
|||
|
|
|
|||
|
|
[Evaluation summary](evaluations/mmlu-finance-stratified-100/results.json) ·
|
|||
|
|
[Per-example predictions](evaluations/mmlu-finance-stratified-100/predictions.jsonl) ·
|
|||
|
|
[Evaluator](evaluations/evaluate_mmlu_finance.py)
|
|||
|
|
|
|||
|
|
Independent evaluation packages are prepared for Vals AI CorpFin/Finance Agent
|
|||
|
|
and Scale Labs PRBench Finance. Those scores remain unavailable until the
|
|||
|
|
benchmark owners evaluate the immutable model revision.
|
|||
|
|
|
|||
|
|
## Known Failure
|
|||
|
|
|
|||
|
|
A sparse-evidence smoke prompt incorrectly identified NVDA as Applied Materials.
|
|||
|
|
This is a documented fabrication failure. Callers must validate company identity
|
|||
|
|
and all market facts against BLUM Engine or another authoritative data source.
|
|||
|
|
|
|||
|
|
## Safety
|
|||
|
|
|
|||
|
|
- Do not use this model as unattended financial advice.
|
|||
|
|
- Do not let it authorize or execute transactions.
|
|||
|
|
- Supply timestamped evidence and validate every factual statement.
|
|||
|
|
- Confidence is bounded but not calibrated on mature trading outcomes.
|