--- library_name: transformers pipeline_tag: text-generation language: - en base_model: - HuggingFaceTB/SmolLM2-135M datasets: - atenareply/noval-corp-training-corpus tags: - smollm2 - continued-pretraining - domain-adaptation license: apache-2.0 --- # SmolLM2-135M Noval — Domain CPT HuggingFaceTB/SmolLM2-135M after Continued Pre-Training on the fictional OMC + Mars Express domain corpus. First-round (135M) proof of the pipeline. ## Overview - **Stage**: Continued Pre-Training (full fine-tune, CLM) - **Lineage**: SmolLM2-135M → **CPT** (this model) - **Method**: Continued Pre-Training (next-token, full fine-tune) on a 30/70 mars/orbital interleaved corpus (tokenized + packed). - **Domain**: fictional — Orbital Mining Corporation (OMC) technical docs + Mars Express telemetry. ## Training | | | |---|---| | Corpus | noval-corp-training-corpus (30% Mars / 70% OMC, packed input_ids) | | LR / schedule | 5e-4 cosine, warmup 200 steps | | Duration / seq | max_steps 5000, seq 1024, eff_batch 32, T4 fp16 | ## Evaluation _No task metrics for this checkpoint (intermediate / no-training stage)._ ## Intended use & limitations Domain backbone for the 135M line (knows the domain; not instruction-tuned). **Limitations:** - 135M is intrinsically limited — expectations should be modest. - Fictional domain; no chat template / instruction following. --- _Card generated by `noval-corp/scripts/gen_model_cards.py` (standardized across the noval-corp model family)._