初始化项目,由ModelHub XC社区提供模型
Model: Arcoson/distilgpt2-arxiv-csai Source: Original Platform
This commit is contained in:
35
.gitattributes
vendored
Normal file
35
.gitattributes
vendored
Normal file
@@ -0,0 +1,35 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
154
README.md
Normal file
154
README.md
Normal file
@@ -0,0 +1,154 @@
|
||||
---
|
||||
language:
|
||||
- en
|
||||
license: mit
|
||||
library_name: transformers
|
||||
tags:
|
||||
- text-generation
|
||||
- arxiv
|
||||
- computer-science
|
||||
- artificial-intelligence
|
||||
- research
|
||||
- gpt2
|
||||
- distilgpt2
|
||||
- fine-tuned
|
||||
datasets:
|
||||
- ccdv/arxiv-classification
|
||||
metrics:
|
||||
- perplexity
|
||||
base_model: distilgpt2
|
||||
pipeline_tag: text-generation
|
||||
model-index:
|
||||
- name: distilgpt2-arxiv-csai
|
||||
results:
|
||||
- task:
|
||||
type: text-generation
|
||||
name: Text Generation
|
||||
dataset:
|
||||
name: arXiv CS/AI Abstracts
|
||||
type: ccdv/arxiv-classification
|
||||
metrics:
|
||||
- type: perplexity
|
||||
value: 32.59
|
||||
name: Perplexity
|
||||
---
|
||||
|
||||
# DistilGPT-2 Fine-Tuned on arXiv CS/AI Abstracts
|
||||
|
||||
A **DistilGPT-2** (82M parameters) model fine-tuned on computer science and artificial intelligence paper abstracts from arXiv. The model generates text in the style of academic CS/AI research abstracts.
|
||||
|
||||
## Model Details
|
||||
|
||||
| Property | Value |
|
||||
|---|---|
|
||||
| **Base Model** | [distilgpt2](https://huggingface.co/distilgpt2) |
|
||||
| **Parameters** | 81.9M total, 14.2M trainable (17.3%) |
|
||||
| **Fine-tuning Strategy** | Partial freeze — last 2 transformer blocks + LM head |
|
||||
| **Training Data** | [ccdv/arxiv-classification](https://huggingface.co/datasets/ccdv/arxiv-classification) (CS/AI subset) |
|
||||
| **Training Samples** | 2,000 |
|
||||
| **Max Sequence Length** | 128 tokens |
|
||||
| **License** | MIT |
|
||||
|
||||
## Training Details
|
||||
|
||||
### Dataset
|
||||
The model was fine-tuned on abstracts from the [ccdv/arxiv-classification](https://huggingface.co/datasets/ccdv/arxiv-classification) dataset, filtered for computer science and AI categories:
|
||||
|
||||
- `cs.CV` (Computer Vision)
|
||||
- `cs.AI` (Artificial Intelligence)
|
||||
- `cs.SY` (Systems and Control)
|
||||
- `cs.CE` (Computational Engineering)
|
||||
- `cs.PL` (Programming Languages)
|
||||
- `cs.IT` (Information Theory)
|
||||
- `cs.DS` (Data Structures and Algorithms)
|
||||
- `cs.NE` (Neural and Evolutionary Computing)
|
||||
|
||||
### Hyperparameters
|
||||
|
||||
| Parameter | Value |
|
||||
|---|---|
|
||||
| Learning rate | 3e-4 |
|
||||
| Batch size | 8 |
|
||||
| Max steps | 100 |
|
||||
| Warmup steps | 10 |
|
||||
| Scheduler | Cosine |
|
||||
| Weight decay | 0.01 |
|
||||
| Frozen layers | Embedding + blocks 0-3 |
|
||||
| Trainable layers | Blocks 4-5 + LayerNorm + LM Head |
|
||||
|
||||
### Metrics
|
||||
|
||||
| Metric | Value |
|
||||
|---|---|
|
||||
| Training loss | 3.82 |
|
||||
| Eval loss | 3.48 |
|
||||
| **Perplexity** | **32.59** |
|
||||
| Training time | ~7 minutes (CPU) |
|
||||
|
||||
## Usage
|
||||
|
||||
```python
|
||||
from transformers import pipeline
|
||||
|
||||
generator = pipeline("text-generation", model="Arcoson/distilgpt2-arxiv-csai")
|
||||
|
||||
output = generator(
|
||||
"arXiv Abstract: We propose a novel approach to",
|
||||
max_new_tokens=100,
|
||||
do_sample=True,
|
||||
temperature=0.8,
|
||||
top_p=0.9,
|
||||
)
|
||||
print(output[0]["generated_text"])
|
||||
```
|
||||
|
||||
Or load directly:
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
tokenizer = AutoTokenizer.from_pretrained("Arcoson/distilgpt2-arxiv-csai")
|
||||
model = AutoModelForCausalLM.from_pretrained("Arcoson/distilgpt2-arxiv-csai")
|
||||
|
||||
inputs = tokenizer("arXiv Abstract: In this paper, we study the problem of", return_tensors="pt")
|
||||
outputs = model.generate(**inputs, max_new_tokens=80, do_sample=True, temperature=0.8)
|
||||
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
||||
```
|
||||
|
||||
## Sample Generations
|
||||
|
||||
**Prompt:** "arXiv Abstract: We propose a novel approach to"
|
||||
|
||||
**Fine-tuned output:**
|
||||
> We propose a novel approach to the problem of time-time in the computation model, i.e., and prediction models for (i) continuous intervals which are nonlinear because they cannot be solved with finite random variables or large number functions...
|
||||
|
||||
**Base DistilGPT-2 output:**
|
||||
> We propose a novel approach to the question of whether we can solve this problem by using real-time neural networks in an MRI technique...
|
||||
|
||||
The fine-tuned model produces more domain-appropriate language with CS/AI terminology compared to the base model.
|
||||
|
||||
## Limitations
|
||||
|
||||
- **Short context**: Trained on 128-token sequences, so it works best for generating short abstracts or completions.
|
||||
- **Limited training**: Only 100 steps on 2,000 samples — this is a proof-of-concept fine-tune. More training data and steps would improve quality.
|
||||
- **Domain narrow**: Outputs are biased toward CS/AI academic language. Not suitable for general-purpose text generation.
|
||||
- **No factual accuracy**: The model generates plausible-sounding but not necessarily factually correct research text.
|
||||
|
||||
## Training Infrastructure
|
||||
|
||||
- **Hardware**: CPU (2 vCPUs, 8GB RAM)
|
||||
- **Framework**: HuggingFace Transformers 5.5.0, PyTorch 2.11
|
||||
- **Strategy**: Partial layer freezing to enable CPU-feasible training
|
||||
|
||||
## Citation
|
||||
|
||||
If you use this model, please cite the base model and dataset:
|
||||
|
||||
```bibtex
|
||||
@article{sanh2019distilbert,
|
||||
title={DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter},
|
||||
author={Sanh, Victor and Debut, Lysandre and Chaumond, Julien and Wolf, Thomas},
|
||||
journal={arXiv preprint arXiv:1910.01108},
|
||||
year={2019}
|
||||
}
|
||||
```
|
||||
48
config.json
Normal file
48
config.json
Normal file
@@ -0,0 +1,48 @@
|
||||
{
|
||||
"_num_labels": 1,
|
||||
"activation_function": "gelu_new",
|
||||
"add_cross_attention": false,
|
||||
"architectures": [
|
||||
"GPT2LMHeadModel"
|
||||
],
|
||||
"attn_pdrop": 0.1,
|
||||
"bos_token_id": 50256,
|
||||
"dtype": "float32",
|
||||
"embd_pdrop": 0.1,
|
||||
"eos_token_id": 50256,
|
||||
"id2label": {
|
||||
"0": "LABEL_0"
|
||||
},
|
||||
"initializer_range": 0.02,
|
||||
"label2id": {
|
||||
"LABEL_0": 0
|
||||
},
|
||||
"layer_norm_epsilon": 1e-05,
|
||||
"model_type": "gpt2",
|
||||
"n_ctx": 1024,
|
||||
"n_embd": 768,
|
||||
"n_head": 12,
|
||||
"n_inner": null,
|
||||
"n_layer": 6,
|
||||
"n_positions": 1024,
|
||||
"pad_token_id": 50256,
|
||||
"reorder_and_upcast_attn": false,
|
||||
"resid_pdrop": 0.1,
|
||||
"scale_attn_by_inverse_layer_idx": false,
|
||||
"scale_attn_weights": true,
|
||||
"summary_activation": null,
|
||||
"summary_first_dropout": 0.1,
|
||||
"summary_proj_to_labels": true,
|
||||
"summary_type": "cls_index",
|
||||
"summary_use_proj": true,
|
||||
"task_specific_params": {
|
||||
"text-generation": {
|
||||
"do_sample": true,
|
||||
"max_length": 50
|
||||
}
|
||||
},
|
||||
"tie_word_embeddings": true,
|
||||
"transformers_version": "5.5.0",
|
||||
"use_cache": false,
|
||||
"vocab_size": 50257
|
||||
}
|
||||
9
generation_config.json
Normal file
9
generation_config.json
Normal file
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"_from_model_config": true,
|
||||
"bos_token_id": 50256,
|
||||
"eos_token_id": [
|
||||
50256
|
||||
],
|
||||
"pad_token_id": 50256,
|
||||
"transformers_version": "5.5.0"
|
||||
}
|
||||
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:987e60f955c2c19285c06b68aaf361d5ab11bcaf3e4e155c06f1c8b6305d082d
|
||||
size 327657928
|
||||
22
sample_generations.json
Normal file
22
sample_generations.json
Normal file
@@ -0,0 +1,22 @@
|
||||
[
|
||||
{
|
||||
"prompt": "arXiv Abstract: We propose a novel approach to",
|
||||
"fine_tuned": "arXiv Abstract: We propose a novel approach to the problem of time-time in\nthe computation model, i.e., and prediction models for (i) continuous intervals which are nonlinear because they cannot be solved with finite random variables or large number functions as described above\u2026in this paper we consider an alternative framework that is able to solve all three problems by decreasing each other\u2019s constant values so it can become efficient if no more than",
|
||||
"base": "arXiv Abstract: We propose a novel approach to the question of whether we can solve this problem by using real-time neural networks in an MRI technique. The method used is based on network theory, which provides high resolution information about how our brain functions (e..g., visual perception).\nThe first goal for neuroimaging was to improve understanding of neurons involved with specific tasks such as motor activity and memory formation when trained or not.[1]"
|
||||
},
|
||||
{
|
||||
"prompt": "arXiv Abstract: In this paper, we study the problem of",
|
||||
"fine_tuned": "arXiv Abstract: In this paper, we study the problem of coupling to an arbitrary method for\nk-bandwidth (K-B) propagation. We introduce a model that provides k-bandwidth by first specifying which types of K-b are propagating at all times in space and passing data as random parameters into each other's inputs/outputs from one parameter pool when it is needed on both sides. The system has several interesting features so far;",
|
||||
"base": "arXiv Abstract: In this paper, we study the problem of how a group\u2030s socialization is affected by individual differences in their ability to live independently. It also investigates why self-interest\u2014being an emotional partner or \"social network\"--can be impacted at various levels during life and helps define which individuals should seek out help when they are alone with other people who have been together for more than 10 years (Siegel et al., 2001; De"
|
||||
},
|
||||
{
|
||||
"prompt": "arXiv Abstract: Deep learning has achieved remarkable success in",
|
||||
"fine_tuned": "arXiv Abstract: Deep learning has achieved remarkable success in\nin machine-learning, particularly on the network. In this paper we present a model of deep neural networks that are based on an algorithm and learn from its initial state at runtime (see Fig. 1 ). We show how it can be used to find random strings with which you could solve your puzzles using such algorithms as sparsely distributed recurrent convolutional training or L1/DNNs .",
|
||||
"base": "arXiv Abstract: Deep learning has achieved remarkable success in real-world applications. It is a highly advanced, and well documented way of doing it; the results are very encouraging for anyone interested in creating complex artificial intelligence systems that can be applied to their environments or even from just about any computer system at all \u2013 as shown by our simulation data above!\nThe following code illustrates how deep neural networks work when they take advantage (for example) of \"smart"
|
||||
},
|
||||
{
|
||||
"prompt": "arXiv Abstract: We present a new algorithm for",
|
||||
"fine_tuned": "arXiv Abstract: We present a new algorithm for generating random strings of the same length with\nlengths, such as long-term values. In this paper, we propose that an arbitrary time series is implemented to calculate string lengths from each subsequence (or any short term sequence) using some number generator scheme and one element in terms of its own identity model\u2014a procedure which allows us make multiple sequences equal by taking up all integers between them so",
|
||||
"base": "arXiv Abstract: We present a new algorithm for combining all data from the x-ray space into one large dataset, and we are working on an approach to generating detailed information about each object that is currently in use. The current version of this paper has been designed as follows: A list of examples (see appendix).\nThe structure used here includes many details such 'data' or \u2032=\u202c_\\mathbb{b}}["
|
||||
}
|
||||
]
|
||||
250320
tokenizer.json
Normal file
250320
tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
12
tokenizer_config.json
Normal file
12
tokenizer_config.json
Normal file
@@ -0,0 +1,12 @@
|
||||
{
|
||||
"add_prefix_space": false,
|
||||
"backend": "tokenizers",
|
||||
"bos_token": "<|endoftext|>",
|
||||
"eos_token": "<|endoftext|>",
|
||||
"errors": "replace",
|
||||
"is_local": false,
|
||||
"model_max_length": 1024,
|
||||
"pad_token": "<|endoftext|>",
|
||||
"tokenizer_class": "GPT2Tokenizer",
|
||||
"unk_token": "<|endoftext|>"
|
||||
}
|
||||
3
training_args.bin
Normal file
3
training_args.bin
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:b1107fdf9adbe6bc256bc28456a3d2f65968fa765d055ba37f9e37f14632ec8e
|
||||
size 5201
|
||||
27
training_metrics.json
Normal file
27
training_metrics.json
Normal file
@@ -0,0 +1,27 @@
|
||||
{
|
||||
"base_model": "distilgpt2",
|
||||
"fine_tuned_on": "arXiv CS/AI abstracts (ccdv/arxiv-classification)",
|
||||
"categories": [
|
||||
"cs.CV",
|
||||
"cs.AI",
|
||||
"cs.SY",
|
||||
"cs.CE",
|
||||
"cs.PL",
|
||||
"cs.IT",
|
||||
"cs.DS",
|
||||
"cs.NE"
|
||||
],
|
||||
"train_samples": 2000,
|
||||
"eval_samples": 200,
|
||||
"max_length": 128,
|
||||
"max_steps": 100,
|
||||
"batch_size": 8,
|
||||
"learning_rate": 0.0003,
|
||||
"frozen_strategy": "last 2 blocks (h.4, h.5) + ln_f + lm_head unfrozen",
|
||||
"training_loss": 3.8221,
|
||||
"eval_loss": 3.4839,
|
||||
"perplexity": 32.59,
|
||||
"total_params": 81912576,
|
||||
"trainable_params": 14177280,
|
||||
"training_time_seconds": 402
|
||||
}
|
||||
Reference in New Issue
Block a user