初始化项目,由ModelHub XC社区提供模型

Model: Arcoson/distilgpt2-arxiv-csai
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-05-27 16:34:25 +08:00
commit a30807e715
10 changed files with 250633 additions and 0 deletions

35
.gitattributes vendored Normal file
View File

@@ -0,0 +1,35 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

154
README.md Normal file
View File

@@ -0,0 +1,154 @@
---
language:
- en
license: mit
library_name: transformers
tags:
- text-generation
- arxiv
- computer-science
- artificial-intelligence
- research
- gpt2
- distilgpt2
- fine-tuned
datasets:
- ccdv/arxiv-classification
metrics:
- perplexity
base_model: distilgpt2
pipeline_tag: text-generation
model-index:
- name: distilgpt2-arxiv-csai
results:
- task:
type: text-generation
name: Text Generation
dataset:
name: arXiv CS/AI Abstracts
type: ccdv/arxiv-classification
metrics:
- type: perplexity
value: 32.59
name: Perplexity
---
# DistilGPT-2 Fine-Tuned on arXiv CS/AI Abstracts
A **DistilGPT-2** (82M parameters) model fine-tuned on computer science and artificial intelligence paper abstracts from arXiv. The model generates text in the style of academic CS/AI research abstracts.
## Model Details
| Property | Value |
|---|---|
| **Base Model** | [distilgpt2](https://huggingface.co/distilgpt2) |
| **Parameters** | 81.9M total, 14.2M trainable (17.3%) |
| **Fine-tuning Strategy** | Partial freeze — last 2 transformer blocks + LM head |
| **Training Data** | [ccdv/arxiv-classification](https://huggingface.co/datasets/ccdv/arxiv-classification) (CS/AI subset) |
| **Training Samples** | 2,000 |
| **Max Sequence Length** | 128 tokens |
| **License** | MIT |
## Training Details
### Dataset
The model was fine-tuned on abstracts from the [ccdv/arxiv-classification](https://huggingface.co/datasets/ccdv/arxiv-classification) dataset, filtered for computer science and AI categories:
- `cs.CV` (Computer Vision)
- `cs.AI` (Artificial Intelligence)
- `cs.SY` (Systems and Control)
- `cs.CE` (Computational Engineering)
- `cs.PL` (Programming Languages)
- `cs.IT` (Information Theory)
- `cs.DS` (Data Structures and Algorithms)
- `cs.NE` (Neural and Evolutionary Computing)
### Hyperparameters
| Parameter | Value |
|---|---|
| Learning rate | 3e-4 |
| Batch size | 8 |
| Max steps | 100 |
| Warmup steps | 10 |
| Scheduler | Cosine |
| Weight decay | 0.01 |
| Frozen layers | Embedding + blocks 0-3 |
| Trainable layers | Blocks 4-5 + LayerNorm + LM Head |
### Metrics
| Metric | Value |
|---|---|
| Training loss | 3.82 |
| Eval loss | 3.48 |
| **Perplexity** | **32.59** |
| Training time | ~7 minutes (CPU) |
## Usage
```python
from transformers import pipeline
generator = pipeline("text-generation", model="Arcoson/distilgpt2-arxiv-csai")
output = generator(
"arXiv Abstract: We propose a novel approach to",
max_new_tokens=100,
do_sample=True,
temperature=0.8,
top_p=0.9,
)
print(output[0]["generated_text"])
```
Or load directly:
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("Arcoson/distilgpt2-arxiv-csai")
model = AutoModelForCausalLM.from_pretrained("Arcoson/distilgpt2-arxiv-csai")
inputs = tokenizer("arXiv Abstract: In this paper, we study the problem of", return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=80, do_sample=True, temperature=0.8)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
## Sample Generations
**Prompt:** "arXiv Abstract: We propose a novel approach to"
**Fine-tuned output:**
> We propose a novel approach to the problem of time-time in the computation model, i.e., and prediction models for (i) continuous intervals which are nonlinear because they cannot be solved with finite random variables or large number functions...
**Base DistilGPT-2 output:**
> We propose a novel approach to the question of whether we can solve this problem by using real-time neural networks in an MRI technique...
The fine-tuned model produces more domain-appropriate language with CS/AI terminology compared to the base model.
## Limitations
- **Short context**: Trained on 128-token sequences, so it works best for generating short abstracts or completions.
- **Limited training**: Only 100 steps on 2,000 samples — this is a proof-of-concept fine-tune. More training data and steps would improve quality.
- **Domain narrow**: Outputs are biased toward CS/AI academic language. Not suitable for general-purpose text generation.
- **No factual accuracy**: The model generates plausible-sounding but not necessarily factually correct research text.
## Training Infrastructure
- **Hardware**: CPU (2 vCPUs, 8GB RAM)
- **Framework**: HuggingFace Transformers 5.5.0, PyTorch 2.11
- **Strategy**: Partial layer freezing to enable CPU-feasible training
## Citation
If you use this model, please cite the base model and dataset:
```bibtex
@article{sanh2019distilbert,
title={DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter},
author={Sanh, Victor and Debut, Lysandre and Chaumond, Julien and Wolf, Thomas},
journal={arXiv preprint arXiv:1910.01108},
year={2019}
}
```

48
config.json Normal file
View File

@@ -0,0 +1,48 @@
{
"_num_labels": 1,
"activation_function": "gelu_new",
"add_cross_attention": false,
"architectures": [
"GPT2LMHeadModel"
],
"attn_pdrop": 0.1,
"bos_token_id": 50256,
"dtype": "float32",
"embd_pdrop": 0.1,
"eos_token_id": 50256,
"id2label": {
"0": "LABEL_0"
},
"initializer_range": 0.02,
"label2id": {
"LABEL_0": 0
},
"layer_norm_epsilon": 1e-05,
"model_type": "gpt2",
"n_ctx": 1024,
"n_embd": 768,
"n_head": 12,
"n_inner": null,
"n_layer": 6,
"n_positions": 1024,
"pad_token_id": 50256,
"reorder_and_upcast_attn": false,
"resid_pdrop": 0.1,
"scale_attn_by_inverse_layer_idx": false,
"scale_attn_weights": true,
"summary_activation": null,
"summary_first_dropout": 0.1,
"summary_proj_to_labels": true,
"summary_type": "cls_index",
"summary_use_proj": true,
"task_specific_params": {
"text-generation": {
"do_sample": true,
"max_length": 50
}
},
"tie_word_embeddings": true,
"transformers_version": "5.5.0",
"use_cache": false,
"vocab_size": 50257
}

9
generation_config.json Normal file
View File

@@ -0,0 +1,9 @@
{
"_from_model_config": true,
"bos_token_id": 50256,
"eos_token_id": [
50256
],
"pad_token_id": 50256,
"transformers_version": "5.5.0"
}

3
model.safetensors Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:987e60f955c2c19285c06b68aaf361d5ab11bcaf3e4e155c06f1c8b6305d082d
size 327657928

22
sample_generations.json Normal file
View File

@@ -0,0 +1,22 @@
[
{
"prompt": "arXiv Abstract: We propose a novel approach to",
"fine_tuned": "arXiv Abstract: We propose a novel approach to the problem of time-time in\nthe computation model, i.e., and prediction models for (i) continuous intervals which are nonlinear because they cannot be solved with finite random variables or large number functions as described above\u2026in this paper we consider an alternative framework that is able to solve all three problems by decreasing each other\u2019s constant values so it can become efficient if no more than",
"base": "arXiv Abstract: We propose a novel approach to the question of whether we can solve this problem by using real-time neural networks in an MRI technique. The method used is based on network theory, which provides high resolution information about how our brain functions (e..g., visual perception).\nThe first goal for neuroimaging was to improve understanding of neurons involved with specific tasks such as motor activity and memory formation when trained or not.[1]"
},
{
"prompt": "arXiv Abstract: In this paper, we study the problem of",
"fine_tuned": "arXiv Abstract: In this paper, we study the problem of coupling to an arbitrary method for\nk-bandwidth (K-B) propagation. We introduce a model that provides k-bandwidth by first specifying which types of K-b are propagating at all times in space and passing data as random parameters into each other's inputs/outputs from one parameter pool when it is needed on both sides. The system has several interesting features so far;",
"base": "arXiv Abstract: In this paper, we study the problem of how a group\u2030s socialization is affected by individual differences in their ability to live independently. It also investigates why self-interest\u2014being an emotional partner or \"social network\"--can be impacted at various levels during life and helps define which individuals should seek out help when they are alone with other people who have been together for more than 10 years (Siegel et al., 2001; De"
},
{
"prompt": "arXiv Abstract: Deep learning has achieved remarkable success in",
"fine_tuned": "arXiv Abstract: Deep learning has achieved remarkable success in\nin machine-learning, particularly on the network. In this paper we present a model of deep neural networks that are based on an algorithm and learn from its initial state at runtime (see Fig. 1 ). We show how it can be used to find random strings with which you could solve your puzzles using such algorithms as sparsely distributed recurrent convolutional training or L1/DNNs .",
"base": "arXiv Abstract: Deep learning has achieved remarkable success in real-world applications. It is a highly advanced, and well documented way of doing it; the results are very encouraging for anyone interested in creating complex artificial intelligence systems that can be applied to their environments or even from just about any computer system at all \u2013 as shown by our simulation data above!\nThe following code illustrates how deep neural networks work when they take advantage (for example) of \"smart"
},
{
"prompt": "arXiv Abstract: We present a new algorithm for",
"fine_tuned": "arXiv Abstract: We present a new algorithm for generating random strings of the same length with\nlengths, such as long-term values. In this paper, we propose that an arbitrary time series is implemented to calculate string lengths from each subsequence (or any short term sequence) using some number generator scheme and one element in terms of its own identity model\u2014a procedure which allows us make multiple sequences equal by taking up all integers between them so",
"base": "arXiv Abstract: We present a new algorithm for combining all data from the x-ray space into one large dataset, and we are working on an approach to generating detailed information about each object that is currently in use. The current version of this paper has been designed as follows: A list of examples (see appendix).\nThe structure used here includes many details such 'data' or \u2032=\u202c_\\mathbb{b}}["
}
]

250320
tokenizer.json Normal file

File diff suppressed because it is too large Load Diff

12
tokenizer_config.json Normal file
View File

@@ -0,0 +1,12 @@
{
"add_prefix_space": false,
"backend": "tokenizers",
"bos_token": "<|endoftext|>",
"eos_token": "<|endoftext|>",
"errors": "replace",
"is_local": false,
"model_max_length": 1024,
"pad_token": "<|endoftext|>",
"tokenizer_class": "GPT2Tokenizer",
"unk_token": "<|endoftext|>"
}

3
training_args.bin Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b1107fdf9adbe6bc256bc28456a3d2f65968fa765d055ba37f9e37f14632ec8e
size 5201

27
training_metrics.json Normal file
View File

@@ -0,0 +1,27 @@
{
"base_model": "distilgpt2",
"fine_tuned_on": "arXiv CS/AI abstracts (ccdv/arxiv-classification)",
"categories": [
"cs.CV",
"cs.AI",
"cs.SY",
"cs.CE",
"cs.PL",
"cs.IT",
"cs.DS",
"cs.NE"
],
"train_samples": 2000,
"eval_samples": 200,
"max_length": 128,
"max_steps": 100,
"batch_size": 8,
"learning_rate": 0.0003,
"frozen_strategy": "last 2 blocks (h.4, h.5) + ln_f + lm_head unfrozen",
"training_loss": 3.8221,
"eval_loss": 3.4839,
"perplexity": 32.59,
"total_params": 81912576,
"trainable_params": 14177280,
"training_time_seconds": 402
}