Files
gpt2-citations/README.md
ModelHub XC 843cfb54f8 初始化项目,由ModelHub XC社区提供模型
Model: lorcannrauzduel/gpt2-citations
Source: Original Platform
2026-08-08 02:35:17 +08:00

170 lines
6.1 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
library_name: transformers
tags:
- text-generation
- gpt2
- fine-tuned
- citations
- causal-lm
---
# GPT2 Finetuned on English Quotes
## Model Description
This model is a finetuned version of [GPT2 small](https://huggingface.co/gpt2) (124M parameters) on the [Abirate/english_quotes](https://huggingface.co/datasets/Abirate/english_quotes) dataset.
The goal is to generate text in the style of philosophical or literary quotes, including the authors name.
**⚠️ This model was created for educational and research purposes only. It is not intended for production use.**
It demonstrates full finetuning of a causal language model on a small dataset and the improvements in generation quality compared to the base model.
**Base model**: `gpt2`
**Task**: Causal language modelling (text generation)
**Finetuning type**: Full finetuning (all parameters updated)
## Intended Uses & Limitations
### Direct Use (Research / Experimentation)
You can use this model to generate short quotes given a prompt. The model expects prompts to start with the special token `<|startoftext|>` and will learn to produce a quote followed by an author and the `<|endoftext|>` token.
**Example**:
```python
from transformers import pipeline
generator = pipeline("text-generation", model="lorcannrauzduel/gpt2-citations")
output = generator("<|startoftext|> The secret to", max_new_tokens=50, do_sample=True)
print(output[0]['generated_text'])
```
### Limitations
- The model is small (124M) and was trained on only ~2,500 quotes. It may sometimes produce repetitive or nonsensical outputs.
- It only generates English text.
- It does not have factual knowledge about the authors; it merely mimics the style of the training quotes.
- **Not suitable for any commercial or critical application.**
## Training Details
### Training Data
- **Dataset**: [Abirate/english_quotes](https://huggingface.co/datasets/Abirate/english_quotes) 2,508 quotes, each with a `quote` and an `author` field.
- **Preprocessing**: Each example was formatted as:
```
<|startoftext|> "quote" — author <|endoftext|>
```
The special tokens help the model learn where a quote starts and ends.
### Training Procedure
The model was trained for 5 epochs using the Hugging Face `Trainer` with the following hyperparameters:
| Hyperparameter | Value |
|------------------------|-------|
| Learning rate | 5e-5 |
| Batch size (per device)| 8 |
| Gradient accumulation | 2 |
| Effective batch size | 16 |
| Warmup steps | 100 |
| Weight decay | 0.01 |
| Optimizer | AdamW |
| Precision | fp16 |
| Max sequence length | 128 |
| Training steps | 1410 |
**Hardware**: NVIDIA Tesla T4 (15 GB VRAM) on Google Colab / Kaggle.
**Training time**: ~5 minutes.
### Evaluation Results
The final training loss was **2.506**, corresponding to a perplexity of **12.26**.
Validation loss stagnated around 2.30, indicating a slight overfitting after 34 epochs acceptable for a small generative model.
## How to Use the Model
### With 🤗 Transformers
```python
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("lorcannrauzduel/gpt2-citations")
model = AutoModelForCausalLM.from_pretrained("lorcannrauzduel/gpt2-citations")
prompt = "<|startoftext|> Life is"
inputs = tokenizer(prompt, return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=50, do_sample=True, temperature=0.9)
print(tokenizer.decode(output[0], skip_special_tokens=False))
```
### With Pipeline
```python
from transformers import pipeline
pipe = pipeline("text-generation", model="lorcannrauzduel/gpt2-citations")
print(pipe("<|startoftext|> You can never", max_new_tokens=50)[0]['generated_text'])
```
### With vLLM (for highthroughput inference)
```bash
pip install vllm
vllm serve "lorcannrauzduel/gpt2-citations"
```
Then query with curl:
```bash
curl -X POST "http://localhost:8000/v1/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "lorcannrauzduel/gpt2-citations",
"prompt": "<|startoftext|> The secret to",
"max_tokens": 50,
"temperature": 0.8
}'
```
### With Ollama (local deployment after GGUF conversion)
1. Download the GGUF version from the repository (if available) or convert it yourself using `llama.cpp`.
2. Create a `Modelfile`:
```
FROM ./gpt2-citations-q4km.gguf
SYSTEM "You are a quote generator."
PARAMETER temperature 0.8
PARAMETER stop "<|endoftext|>"
```
3. Import and run:
```bash
ollama create gpt2-citations -f Modelfile
ollama run gpt2-citations "<|startoftext|> Life is"
```
## Model Comparison (Base vs Finetuned)
| Prompt | GPT2 Base (no finetuning) | GPT2 Finetuned |
|--------|-----------------------------|------------------|
| `<|startoftext|> The secret to` | "The secret to making the most of your life... [[...]]" | "The secret to happiness is to trust your instincts rather than your brain.” — Albert Einstein" |
| `<|startoftext|> Life is` | "Life is precious and life is precious..." | "Life is full of opportunities, but few opportunities are worth the time..." — Jodi Picoult |
| `<|startoftext|> You can never` | "You can never get around to finding out what you want..." | "You can never lose your way because you are still thinking about the things you've done..." |
The finetuned model consistently produces coherent quotes with an author attribution, while the base model generates irrelevant or repetitive text.
## Environmental Impact
Training was performed on a cloud GPU (Tesla T4) for about 5 minutes. Estimated CO₂ emissions are negligible (< 0.01 kg CO₂eq).
## Acknowledgements
- The [Hugging Face](https://huggingface.co) team for `transformers` and `datasets`.
- The original GPT2 paper by Radford et al. (2019).
- Dataset provided by [Abirate](https://huggingface.co/Abirate).
## License
This model is released under the MIT license (same as the original GPT2 small).
---
**Model card created by [lorcannrauzduel](https://huggingface.co/lorcannrauzduel) for research and experimentation purposes.**