Files
smollm2-135m-noval/README.md

49 lines
1.5 KiB
Markdown
Raw Normal View History

---
library_name: transformers
pipeline_tag: text-generation
language:
- en
base_model:
- HuggingFaceTB/SmolLM2-135M
datasets:
- atenareply/noval-corp-training-corpus
tags:
- smollm2
- continued-pretraining
- domain-adaptation
license: apache-2.0
---
# SmolLM2-135M Noval — Domain CPT
HuggingFaceTB/SmolLM2-135M after Continued Pre-Training on the fictional OMC + Mars Express domain corpus. First-round (135M) proof of the pipeline.
## Overview
- **Stage**: Continued Pre-Training (full fine-tune, CLM)
- **Lineage**: SmolLM2-135M → **CPT** (this model)
- **Method**: Continued Pre-Training (next-token, full fine-tune) on a 30/70 mars/orbital interleaved corpus (tokenized + packed).
- **Domain**: fictional — Orbital Mining Corporation (OMC) technical docs + Mars Express telemetry.
## Training
| | |
|---|---|
| Corpus | noval-corp-training-corpus (30% Mars / 70% OMC, packed input_ids) |
| LR / schedule | 5e-4 cosine, warmup 200 steps |
| Duration / seq | max_steps 5000, seq 1024, eff_batch 32, T4 fp16 |
## Evaluation
_No task metrics for this checkpoint (intermediate / no-training stage)._
## Intended use & limitations
Domain backbone for the 135M line (knows the domain; not instruction-tuned).
**Limitations:**
- 135M is intrinsically limited — expectations should be modest.
- Fictional domain; no chat template / instruction following.
---
_Card generated by `noval-corp/scripts/gen_model_cards.py` (standardized across the noval-corp model family)._