49 lines
1.5 KiB
Markdown
49 lines
1.5 KiB
Markdown
---
|
|
library_name: transformers
|
|
pipeline_tag: text-generation
|
|
language:
|
|
- en
|
|
base_model:
|
|
- HuggingFaceTB/SmolLM2-135M
|
|
datasets:
|
|
- atenareply/noval-corp-training-corpus
|
|
tags:
|
|
- smollm2
|
|
- continued-pretraining
|
|
- domain-adaptation
|
|
license: apache-2.0
|
|
---
|
|
|
|
# SmolLM2-135M Noval — Domain CPT
|
|
|
|
HuggingFaceTB/SmolLM2-135M after Continued Pre-Training on the fictional OMC + Mars Express domain corpus. First-round (135M) proof of the pipeline.
|
|
|
|
## Overview
|
|
|
|
- **Stage**: Continued Pre-Training (full fine-tune, CLM)
|
|
- **Lineage**: SmolLM2-135M → **CPT** (this model)
|
|
- **Method**: Continued Pre-Training (next-token, full fine-tune) on a 30/70 mars/orbital interleaved corpus (tokenized + packed).
|
|
- **Domain**: fictional — Orbital Mining Corporation (OMC) technical docs + Mars Express telemetry.
|
|
|
|
## Training
|
|
|
|
| | |
|
|
|---|---|
|
|
| Corpus | noval-corp-training-corpus (30% Mars / 70% OMC, packed input_ids) |
|
|
| LR / schedule | 5e-4 cosine, warmup 200 steps |
|
|
| Duration / seq | max_steps 5000, seq 1024, eff_batch 32, T4 fp16 |
|
|
|
|
## Evaluation
|
|
|
|
_No task metrics for this checkpoint (intermediate / no-training stage)._
|
|
|
|
## Intended use & limitations
|
|
|
|
Domain backbone for the 135M line (knows the domain; not instruction-tuned).
|
|
|
|
**Limitations:**
|
|
- 135M is intrinsically limited — expectations should be modest.
|
|
- Fictional domain; no chat template / instruction following.
|
|
|
|
---
|
|
_Card generated by `noval-corp/scripts/gen_model_cards.py` (standardized across the noval-corp model family)._ |