Files
GPT2.5.5-Awakened.Thinker-0.1B/README.md
ModelHub XC 62aa1328c2 初始化项目,由ModelHub XC社区提供模型
Model: 11-47/GPT2.5.5-Awakened.Thinker-0.1B
Source: Original Platform
2026-08-16 08:15:18 +08:00

112 lines
3.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
language:
- en
license: mit
tags:
- gpt2
- causal-lm
- text-generation
- distillation
- fine-tuned
- withinusai
- full-fine-tune
- thinking
- reasoning
base_model: openai-community/gpt2
datasets:
- WithinUsAI/GPT_5.5_Distilled
- WithinUsAI/GPT5.5_thinking_max_distill_god_seed_25K
model_type: gpt2
pipeline_tag: text-generation
library_name: transformers
widget:
- text: "The key to understanding intelligence is"
- text: "Let me think through this step by step:"
- text: "Deep within the layers of thought,"
---
# GPT2.5.5-Awakened.Thinker-0.1B
> *Forged by WithIn Us AI — a fully fine-tuned GPT-2 awakened on distilled GPT-5.5 thinking patterns.*
## Model Overview
| Property | Value |
|---|---|
| **Architecture** | GPT-2 (`openai-community/gpt2`) |
| **Parameters** | ~124M (0.1B class) |
| **Training Type** | Full fine-tune — ALL weights updated, zero adapters |
| **Context Window** | 1024 tokens |
| **Best Eval Loss** | `0.4365` |
| **Best Perplexity** | `1.55` |
| **Creator** | [GODsStrongestSoldier](https://huggingface.co/GODsStrongestSoldier) / WithIn Us AI |
| **Date Trained** | 2026-05-23 |
| **Hardware** | 2× NVIDIA Tesla T4 (Kaggle) |
| **Precision** | FP16 mixed precision |
---
## Training Methodology
Full fine-tuning — every single parameter in GPT-2 was updated.
No LoRA, no QLoRA, no adapters of any kind.
### Datasets
| Dataset | Description |
|---|---|
| [`WithinUsAI/GPT_5.5_Distilled`](https://huggingface.co/datasets/WithinUsAI/GPT_5.5_Distilled) | Instruction + completion pairs distilled from GPT-5.5 |
| [`WithinUsAI/GPT5.5_thinking_max_distill_god_seed_25K`](https://huggingface.co/datasets/WithinUsAI/GPT5.5_thinking_max_distill_god_seed_25K) | 25K chain-of-thought reasoning traces distilled from GPT-5.5 |
97 / 3 train / eval split.
### Hyperparameters
| Parameter | Value |
|---|---|
| Peak Learning Rate | `3e-5` |
| LR Schedule | Cosine with 6% warmup |
| Effective Batch Size | `64` (4 × 2 GPUs × 8 grad accum) |
| Epochs | `5` |
| Weight Decay | `0.1` |
| Max Sequence Length | `1024` |
| Precision | FP16 |
---
## Quick Start
```python
from transformers import GPT2LMHeadModel, GPT2TokenizerFast
import torch
model_id = "GODsStrongestSoldier/GPT2.5.5-Awakened.Thinker-0.1B"
tokenizer = GPT2TokenizerFast.from_pretrained(model_id)
model = GPT2LMHeadModel.from_pretrained(model_id, torch_dtype=torch.float16)
model.eval()
prompt = "Let me think through this carefully, step by step:"
inputs = tokenizer(prompt, return_tensors="pt")
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens = 200,
do_sample = True,
temperature = 0.7,
top_p = 0.9,
repetition_penalty = 1.15,
)
print(tokenizer.decode(output[0], skip_special_tokens=True))
```
---
## About WithIn Us AI
- HuggingFace org : [WithinUsAI](https://huggingface.co/WithinUsAI)
- Creator : [GODsStrongestSoldier](https://huggingface.co/GODsStrongestSoldier)
*"Strength through understanding. Awakened from within."*