初始化项目,由ModelHub XC社区提供模型
Model: skhotta/slm-125m-base Source: Original Platform
This commit is contained in:
27
README.md
Normal file
27
README.md
Normal file
@@ -0,0 +1,27 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
language:
|
||||
- en
|
||||
library_name: transformers
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- llama
|
||||
- legal
|
||||
- finance
|
||||
- small-language-model
|
||||
---
|
||||
|
||||
# skhotta/slm-125m-base
|
||||
|
||||
A ~125.8M parameter Llama-architecture base language model trained from scratch
|
||||
on a legal/financial corpus (US case law + SEC filings + a fineweb-edu slice),
|
||||
with a custom 16,384-token byte-level BPE tokenizer.
|
||||
|
||||
- **Params:** ~125.8M (12L / 768d /
|
||||
12h, context 1024)
|
||||
- **Vocab:** 16384 (byte-level BPE)
|
||||
- **Training data:** ~2.0B tokens (~40% case law / ~40% SEC / ~20% web), deduplicated
|
||||
and decontaminated against CaseHOLD/LexGLUE.
|
||||
- **Objective:** causal language modeling.
|
||||
|
||||
This is a **base** model (no instruction tuning). It is a research artifact.
|
||||
Reference in New Issue
Block a user