Model Card

DistilBERT-Gemini3.2-Pro
Fast.Coder.NSFW-0.1B

by WithinUsAI  ·  GODsStrongestSoldier  ·  View on 🤗 Hub
```
Text Generation Causal LM Reasoning Coding Creative Writing Distillation PyTorch Safetensors English Not-For-All-Audiences
Parameters
81.9M
Tensor Type
F32
Context Window
1024tok
License
Apache 2.0
Base Model
DistilGPT2
```
01

Model Description

DistilBERT-Gemini3.2-Pro.Fast.Coder.NSFW-0.1B is a fully fine-tuned version of DistilGPT2 — exposed to an aggressive curriculum of high-reasoning Gemini distillation traces, comprehensive coding datasets, creative writing frameworks, and mature internet discourse.

The model is designed to act as a highly responsive, analytical engine capable of deep structural reasoning and complex logic emulation. Despite its compact 81.9M parameter footprint, it targets multi-domain competence across code generation, reasoning chains, and open-ended text generation.

Trained natively at an accelerated maximum learning rate with a cosine decay schedule, the model synthesizes diverse programmatic and theoretical domains from a massive multi-repository corpus, processed at DistilGPT2's maximum context window of 1024 tokens.

02

Intended Use

Primary Use
Code Generation
Secondary Use
Reasoning & Analysis
Tertiary Use
Creative Writing
Audience
Adults Only (18+)
03

Training Datasets

Fine-tuned on a synthesized Golden Mix drawn from eight multi-domain repositories:

04

Training Procedure

The model underwent full fine-tuning — no adapters or LoRA. All native DistilGPT2 parameters were globally updated. The training harness dynamically parsed heavily nested dataset repositories, enforcing a strict shape constraint to generate mathematically perfect 1024-token continuous sequences, maxing out the model's context window.

Epochs
1
Block Size
1024
Batch / Device
4
Grad Accum Steps
16
Global Batch
128
Peak LR
3e-4
LR Scheduler
Cosine
Warmup Ratio
0.05
Optimizer
AdamW Fused
Precision
fp16
Grad Checkpoint
Enabled
Fine-Tune Type
Full (no LoRA)
```
2×
NVIDIA T4
Environment: Kaggle
VRAM: 15 GB per GPU (30 GB total)
Accelerator: Dual NVIDIA T4 GPUs
```
05

Limitations & Risks

06

License & Attribution

Released under the Apache 2.0 license. Free to use, modify, and distribute with attribution.

📄 Apache-2.0 License

Base model: distilbert/distilgpt2 by HuggingFace / DistilBERT team. Model card authored for: WithinUsAI.