初始化项目,由ModelHub XC社区提供模型
Model: juanquivilla/sotto-cleanup-lfm25-350m Source: Original Platform
This commit is contained in:
35
.gitattributes
vendored
Normal file
35
.gitattributes
vendored
Normal file
@@ -0,0 +1,35 @@
|
|||||||
|
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.model filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||||
|
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||||
110
README.md
Normal file
110
README.md
Normal file
@@ -0,0 +1,110 @@
|
|||||||
|
---
|
||||||
|
license: mit
|
||||||
|
language:
|
||||||
|
- en
|
||||||
|
base_model: LiquidAI/LFM2.5-350M-Base
|
||||||
|
tags:
|
||||||
|
- speech-to-text
|
||||||
|
- transcript-cleanup
|
||||||
|
- text-correction
|
||||||
|
- asr-post-processing
|
||||||
|
- LFM
|
||||||
|
- LiquidAI
|
||||||
|
- grpo
|
||||||
|
- full-fine-tune
|
||||||
|
- inverse-text-normalization
|
||||||
|
pipeline_tag: text-generation
|
||||||
|
datasets:
|
||||||
|
- juanquivilla/sotto-transcript-cleanup
|
||||||
|
---
|
||||||
|
|
||||||
|
# SottoASR Transcript Cleanup — LFM2.5-350M (Full Precision, soup_30)
|
||||||
|
|
||||||
|
[sottoasr.app](https://sottoasr.app) · [MLX 5-bit (recommended)](https://huggingface.co/juanquivilla/sotto-cleanup-lfm25-350m-mlx-5bit) · [MLX 4-bit (smaller)](https://huggingface.co/juanquivilla/sotto-cleanup-lfm25-350m-mlx-4bit) · [Training Dataset](https://huggingface.co/datasets/juanquivilla/sotto-transcript-cleanup)
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
Full-precision bf16 fine-tune of [LiquidAI/LFM2.5-350M-Base](https://huggingface.co/LiquidAI/LFM2.5-350M-Base) for on-device speech-to-text transcript cleanup. This is the **training artifact** — for on-device deployment on Apple Silicon, use the [5-bit MLX variant](https://huggingface.co/juanquivilla/sotto-cleanup-lfm25-350m-mlx-5bit).
|
||||||
|
|
||||||
|
## What's new (model soup release)
|
||||||
|
|
||||||
|
This model is a **weight-space average** of two strong checkpoints from the same fine-tuning lineage:
|
||||||
|
- **0.3 × v55** (latest: 2-epoch refinement at lr 2e-6) — strongest on number-accuracy and filler-stripping
|
||||||
|
- **0.7 × v51** (the prior production model) — strongest on adversarial sampling benchmark
|
||||||
|
|
||||||
|
Linear interpolation in weight space (`θ = α·θ_v55 + (1-α)·θ_v51`) is sometimes called "model souping". It works here because v55 was chained from v51 (same architecture, related minima), and the soup recovers v51's bench-sample strengths without losing v55's number/filler gains. The full recipe sweep is in the [research journal](https://github.com/anthropics) (2026-05-06 loop).
|
||||||
|
|
||||||
|
## Headline numbers (production-mode eval: `max_new_tokens=900`, `repetition_penalty=1.05`)
|
||||||
|
|
||||||
|
| Capability | v36 | v45 | v51 | v55 | **soup (this)** |
|
||||||
|
|---|---:|---:|---:|---:|---:|
|
||||||
|
| Number accuracy (171-sample stratified val) | 12.9% | 95.9% | 95.3% | 96.5% | **96.5%** |
|
||||||
|
| 66-case adversarial benchmark (greedy) | n/a | 76% | 84.8% | 84.8% | **86.4%** |
|
||||||
|
| 66-case adversarial benchmark (temp 0.7 × 4) | n/a | 77% | 84.5% | 82.6% | **86.0%** |
|
||||||
|
| Loops on 264 sampling-mode probes | n/a | 0 | 1 | 2 | **0** |
|
||||||
|
| Filler-free on 241 long inputs | 67.2% | 68.0% | 72.2% | 72.6% | 71.8% |
|
||||||
|
| Sub-deletion >15% on 241 long inputs | 13.3% | 13.7% | 4.6% | 5.0% | **5.0%** |
|
||||||
|
|
||||||
|
Composite score (0.35×num + 0.30×bench_greedy + 0.15×bench_sample + 0.10×filler_long + 0.05×(1-sub15) + 0.05×(1-loops/N)): **89.51** at full production settings.
|
||||||
|
|
||||||
|
## Training pipeline
|
||||||
|
|
||||||
|
```
|
||||||
|
LiquidAI/LFM2.5-350M-Base
|
||||||
|
→ SFT v23 → GRPO v23 (paragraph emission)
|
||||||
|
→ GRPO v36: full FT with substantive-deletion-aware reward
|
||||||
|
→ SFT v39: + 12.7K augmented number examples (ITN)
|
||||||
|
→ GRPO v40–v45: chained refinement, fixed reward + amplified filler penalty
|
||||||
|
→ GRPO v50 + v51: anti-loop n-gram penalty
|
||||||
|
→ GRPO v55: 2-epoch refinement at lr 2e-6 (best chained checkpoint)
|
||||||
|
→ soup: 0.3·θ_v55 + 0.7·θ_v51 (weight-space average — this model)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Usage
|
||||||
|
|
||||||
|
```python
|
||||||
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||||
|
import torch
|
||||||
|
|
||||||
|
model = AutoModelForCausalLM.from_pretrained(
|
||||||
|
"juanquivilla/sotto-cleanup-lfm25-350m",
|
||||||
|
dtype=torch.bfloat16, trust_remote_code=True,
|
||||||
|
)
|
||||||
|
tokenizer = AutoTokenizer.from_pretrained("juanquivilla/sotto-cleanup-lfm25-350m")
|
||||||
|
|
||||||
|
text = "talk about server three sixty"
|
||||||
|
prompt = f"### Input:\n{text}\n\n### Output:\n"
|
||||||
|
|
||||||
|
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
||||||
|
with torch.no_grad():
|
||||||
|
out = model.generate(
|
||||||
|
**inputs,
|
||||||
|
max_new_tokens=max(900, int(len(text.split()) * 1.5)), # ≥1.5× input word count
|
||||||
|
do_sample=False,
|
||||||
|
repetition_penalty=1.05, # LFM2.5 official default
|
||||||
|
)
|
||||||
|
output = tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
|
||||||
|
if "###" in output:
|
||||||
|
output = output[:output.index("###")]
|
||||||
|
print(output.strip())
|
||||||
|
```
|
||||||
|
|
||||||
|
### Inference recommendations
|
||||||
|
|
||||||
|
The headline numbers above use these settings — they match the LFM2.5 model card's defaults and are the production deployment for [sottoasr.app](https://sottoasr.app):
|
||||||
|
|
||||||
|
- **`repetition_penalty=1.05`** — LFM2.5's official default. Critical for long inputs: prevents the rare voicemail-style 5-gram loops that can occur with `repetition_penalty=1.0`.
|
||||||
|
- **`max_new_tokens >= 1.5 × input_word_count`** (or 900 minimum) — long inputs (>200 words) need headroom; truncating mid-output looks like content deletion.
|
||||||
|
- **`do_sample=False`** (greedy) for deterministic output. If sampling is needed, use `temperature=0.1, top_k=50`.
|
||||||
|
|
||||||
|
## All Variants
|
||||||
|
|
||||||
|
| Variant | Size | Use Case |
|
||||||
|
|---------|------|----------|
|
||||||
|
| **[Full precision (this)](https://huggingface.co/juanquivilla/sotto-cleanup-lfm25-350m)** | 676 MB | Training, GPU inference |
|
||||||
|
| **[MLX 5-bit](https://huggingface.co/juanquivilla/sotto-cleanup-lfm25-350m-mlx-5bit)** | ~237 MB | **Recommended for Apple Silicon** |
|
||||||
|
| [MLX 4-bit](https://huggingface.co/juanquivilla/sotto-cleanup-lfm25-350m-mlx-4bit) | ~195 MB | Smallest |
|
||||||
|
|
||||||
|
## License
|
||||||
|
|
||||||
|
MIT
|
||||||
7
chat_template.jinja
Normal file
7
chat_template.jinja
Normal file
@@ -0,0 +1,7 @@
|
|||||||
|
{{- bos_token -}}{%- set system_prompt = "" -%}{%- set ns = namespace(system_prompt="") -%}{%- if messages[0]["role"] == "system" -%} {%- set ns.system_prompt = messages[0]["content"] -%} {%- set messages = messages[1:] -%}{%- endif -%}{%- if tools -%} {%- set ns.system_prompt = ns.system_prompt + ("
|
||||||
|
" if ns.system_prompt else "") + "List of tools: <|tool_list_start|>[" -%} {%- for tool in tools -%} {%- if tool is not string -%} {%- set tool = tool | tojson -%} {%- endif -%} {%- set ns.system_prompt = ns.system_prompt + tool -%} {%- if not loop.last -%} {%- set ns.system_prompt = ns.system_prompt + ", " -%} {%- endif -%} {%- endfor -%} {%- set ns.system_prompt = ns.system_prompt + "]<|tool_list_end|>" -%}{%- endif -%}{%- if ns.system_prompt -%} {{- "<|im_start|>system
|
||||||
|
" + ns.system_prompt + "<|im_end|>
|
||||||
|
" -}}{%- endif -%}{%- for message in messages -%} {{- "<|im_start|>" + message["role"] + "
|
||||||
|
" -}} {%- set content = message["content"] -%} {%- if content is not string -%} {%- set content = content | tojson -%} {%- endif -%} {%- if message["role"] == "tool" -%} {%- set content = "<|tool_response_start|>" + content + "<|tool_response_end|>" -%} {%- endif -%} {{- content + "<|im_end|>
|
||||||
|
" -}}{%- endfor -%}{%- if add_generation_prompt -%} {{- "<|im_start|>assistant
|
||||||
|
" -}}{%- endif -%}
|
||||||
61
config.json
Normal file
61
config.json
Normal file
@@ -0,0 +1,61 @@
|
|||||||
|
{
|
||||||
|
"architectures": [
|
||||||
|
"Lfm2ForCausalLM"
|
||||||
|
],
|
||||||
|
"block_auto_adjust_ff_dim": true,
|
||||||
|
"block_dim": 1024,
|
||||||
|
"block_ff_dim": 6656,
|
||||||
|
"block_ffn_dim_multiplier": 1.0,
|
||||||
|
"block_mlp_init_scale": 1.0,
|
||||||
|
"block_multiple_of": 256,
|
||||||
|
"block_norm_eps": 1e-05,
|
||||||
|
"block_out_init_scale": 1.0,
|
||||||
|
"block_use_swiglu": true,
|
||||||
|
"block_use_xavier_init": true,
|
||||||
|
"bos_token_id": 1,
|
||||||
|
"conv_L_cache": 3,
|
||||||
|
"conv_bias": false,
|
||||||
|
"conv_dim": 1024,
|
||||||
|
"conv_use_xavier_init": true,
|
||||||
|
"dtype": "bfloat16",
|
||||||
|
"eos_token_id": 7,
|
||||||
|
"hidden_size": 1024,
|
||||||
|
"initializer_range": 0.02,
|
||||||
|
"intermediate_size": 6656,
|
||||||
|
"layer_types": [
|
||||||
|
"conv",
|
||||||
|
"conv",
|
||||||
|
"full_attention",
|
||||||
|
"conv",
|
||||||
|
"conv",
|
||||||
|
"full_attention",
|
||||||
|
"conv",
|
||||||
|
"conv",
|
||||||
|
"full_attention",
|
||||||
|
"conv",
|
||||||
|
"full_attention",
|
||||||
|
"conv",
|
||||||
|
"full_attention",
|
||||||
|
"conv",
|
||||||
|
"full_attention",
|
||||||
|
"conv"
|
||||||
|
],
|
||||||
|
"max_position_embeddings": 128000,
|
||||||
|
"model_type": "lfm2",
|
||||||
|
"norm_eps": 1e-05,
|
||||||
|
"num_attention_heads": 16,
|
||||||
|
"num_heads": 16,
|
||||||
|
"num_hidden_layers": 16,
|
||||||
|
"num_key_value_heads": 8,
|
||||||
|
"pad_token_id": 0,
|
||||||
|
"rope_parameters": {
|
||||||
|
"rope_theta": 1000000.0,
|
||||||
|
"rope_type": "default"
|
||||||
|
},
|
||||||
|
"tie_word_embeddings": true,
|
||||||
|
"transformers_version": "5.3.0",
|
||||||
|
"use_cache": false,
|
||||||
|
"use_pos_enc": true,
|
||||||
|
"vocab_size": 65536,
|
||||||
|
"rope_theta": 1000000.0
|
||||||
|
}
|
||||||
9
generation_config.json
Normal file
9
generation_config.json
Normal file
@@ -0,0 +1,9 @@
|
|||||||
|
{
|
||||||
|
"_from_model_config": true,
|
||||||
|
"bos_token_id": 1,
|
||||||
|
"eos_token_id": [
|
||||||
|
7
|
||||||
|
],
|
||||||
|
"pad_token_id": 0,
|
||||||
|
"transformers_version": "5.3.0"
|
||||||
|
}
|
||||||
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:6e96eeffdcdd60f881e13eb2019b339b39d1a74951446f062e7e641a82f6422e
|
||||||
|
size 708984464
|
||||||
323812
tokenizer.json
Normal file
323812
tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
21
tokenizer_config.json
Normal file
21
tokenizer_config.json
Normal file
@@ -0,0 +1,21 @@
|
|||||||
|
{
|
||||||
|
"backend": "tokenizers",
|
||||||
|
"bos_token": "<|startoftext|>",
|
||||||
|
"clean_up_tokenization_spaces": false,
|
||||||
|
"eos_token": "<|im_end|>",
|
||||||
|
"extra_special_tokens": [],
|
||||||
|
"is_local": true,
|
||||||
|
"legacy": false,
|
||||||
|
"local_files_only": false,
|
||||||
|
"model_input_names": [
|
||||||
|
"input_ids",
|
||||||
|
"attention_mask"
|
||||||
|
],
|
||||||
|
"model_max_length": 1000000000000000019884624838656,
|
||||||
|
"pad_token": "<|pad|>",
|
||||||
|
"sp_model_kwargs": {},
|
||||||
|
"spaces_between_special_tokens": false,
|
||||||
|
"tokenizer_class": "TokenizersBackend",
|
||||||
|
"use_default_system_prompt": false,
|
||||||
|
"use_fast": true
|
||||||
|
}
|
||||||
3
training_args.bin
Normal file
3
training_args.bin
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:052ea156c1f5f51905925b2767166bc4c4e3956dc39679724ae0bc32d732a0c3
|
||||||
|
size 5713
|
||||||
Reference in New Issue
Block a user