25K chain-of-thought reasoning traces distilled from GPT-5.5
97 / 3 train / eval split.
Hyperparameters
Parameter
Value
Peak Learning Rate
3e-5
LR Schedule
Cosine with 6% warmup
Effective Batch Size
64 (4 × 2 GPUs × 8 grad accum)
Epochs
5
Weight Decay
0.1
Max Sequence Length
1024
Precision
FP16
Quick Start
fromtransformersimportGPT2LMHeadModel,GPT2TokenizerFastimporttorchmodel_id="GODsStrongestSoldier/GPT2.5.5-Awakened.Thinker-0.1B"tokenizer=GPT2TokenizerFast.from_pretrained(model_id)model=GPT2LMHeadModel.from_pretrained(model_id,torch_dtype=torch.float16)model.eval()prompt="Let me think through this carefully, step by step:"inputs=tokenizer(prompt,return_tensors="pt")withtorch.no_grad():output=model.generate(**inputs,max_new_tokens=200,do_sample=True,temperature=0.7,top_p=0.9,repetition_penalty=1.15,)print(tokenizer.decode(output[0],skip_special_tokens=True))