初始化项目,由ModelHub XC社区提供模型
Model: aubmindlab/aragpt2-base Source: Original Platform
This commit is contained in:
11
.gitattributes
vendored
Normal file
11
.gitattributes
vendored
Normal file
@@ -0,0 +1,11 @@
|
|||||||
|
*.bin.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar.gz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||||
|
events.out.tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
model.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||||
77
LICENSE
Normal file
77
LICENSE
Normal file
@@ -0,0 +1,77 @@
|
|||||||
|
==========================================
|
||||||
|
SOFTWARE LICENSE AGREEMENT - AraGPT2
|
||||||
|
==========================================
|
||||||
|
|
||||||
|
* NAME: AraGPT2: Pre-Training Text Discriminatorsfor Arabic Language Understanding
|
||||||
|
|
||||||
|
* ACKNOWLEDGMENTS
|
||||||
|
|
||||||
|
This [software] was generated by [American
|
||||||
|
University of Beirut] (“Owners”). The statements
|
||||||
|
made herein are solely the responsibility of the author[s].
|
||||||
|
|
||||||
|
The following software programs and programs have been used in the
|
||||||
|
generation of [AraGPT2]:
|
||||||
|
|
||||||
|
+ gpt-2-simple
|
||||||
|
- Max Woolf, 2020.
|
||||||
|
- License and link : https://github.com/minimaxir/gpt-2-simple
|
||||||
|
|
||||||
|
+ GPT2-ML
|
||||||
|
- Zhibo Zhang, GPT2-ML: GPT-2 for Multiple Languages, 2019
|
||||||
|
- License and link : https://github.com/imcaspar/gpt2-ml
|
||||||
|
|
||||||
|
+ PyArabic
|
||||||
|
- T. Zerrouki, Pyarabic, An Arabic language library for Python,
|
||||||
|
https://pypi.python.org/pypi/pyarabic/, 2010
|
||||||
|
- License and link: https://github.com/linuxscout/pyarabic/
|
||||||
|
|
||||||
|
* LICENSE
|
||||||
|
|
||||||
|
This software and database is being provided to you, the LICENSEE,
|
||||||
|
by the Owners under the following license. By obtaining, using and/or
|
||||||
|
copying this software and database, you agree that you have read,
|
||||||
|
understood, and will comply with these terms and conditions. You
|
||||||
|
further agree that you have read and you will abide by the license
|
||||||
|
agreements provided in the above links under “acknowledgements”:
|
||||||
|
Permission to use, copy, modify and distribute this software and
|
||||||
|
database and its documentation for any purpose and without fee or
|
||||||
|
royalty is hereby granted, provided that you agree to comply with the
|
||||||
|
following copyright notice and statements, including the disclaimer,
|
||||||
|
and that the same appear on ALL copies of the software, database and
|
||||||
|
documentation, including modifications that you make for internal use
|
||||||
|
or for distribution. [AraGPT2] Copyright 2020 by [American University
|
||||||
|
of Beirut]. All rights reserved. If you remix, transform, or build
|
||||||
|
upon the material, you must distribute your contributions under the
|
||||||
|
same license as this one. You may not apply legal terms or technological
|
||||||
|
measures that legally restrict others from doing anything this license
|
||||||
|
permits. THIS SOFTWARE IS PROVIDED "AS IS" AND THE OWNERS MAKE NO
|
||||||
|
REPRESENTATIONS OR WARRANTIES, EXPRESS OR IMPLIED. BY WAY OF EXAMPLE,
|
||||||
|
BUT NOT LIMITATION, THE OWNERS MAKE NO REPRESENTATIONS OR WARRANTIES OF
|
||||||
|
MERCHANT-ABILITY OR FITNESS FOR ANY PARTICULAR PURPOSE OR THAT THE USE OF
|
||||||
|
THE LICENSED SOFTWARE, DATABASE OR DOCUMENTATION WILL NOT INFRINGE ANY THIRD
|
||||||
|
PARTY PATENTS, COPYRIGHTS, TRADEMARKS OR OTHER RIGHTS. The name of the
|
||||||
|
Owners may not be used in advertising or publicity pertaining to
|
||||||
|
distribution of the software and/or database. Title to copyright in
|
||||||
|
this software, database and any associated documentation shall at all
|
||||||
|
times remain with the Owners and LICENSEE agrees to preserve same.
|
||||||
|
|
||||||
|
The use of AraGPT2 should be cited as follows:
|
||||||
|
|
||||||
|
@inproceedings{antoun-etal-2021-aragpt2,
|
||||||
|
title = "{A}ra{GPT}2: Pre-Trained Transformer for {A}rabic Language Generation",
|
||||||
|
author = "Antoun, Wissam and
|
||||||
|
Baly, Fady and
|
||||||
|
Hajj, Hazem",
|
||||||
|
booktitle = "Proceedings of the Sixth Arabic Natural Language Processing Workshop",
|
||||||
|
month = apr,
|
||||||
|
year = "2021",
|
||||||
|
address = "Kyiv, Ukraine (Virtual)",
|
||||||
|
publisher = "Association for Computational Linguistics",
|
||||||
|
url = "https://www.aclweb.org/anthology/2021.wanlp-1.21",
|
||||||
|
pages = "196--207",
|
||||||
|
}
|
||||||
|
|
||||||
|
[AraGPT2] Copyright 2020 by [American University of Beirut].
|
||||||
|
All rights reserved.
|
||||||
|
==========================================
|
||||||
144
README.md
Normal file
144
README.md
Normal file
@@ -0,0 +1,144 @@
|
|||||||
|
---
|
||||||
|
language: ar
|
||||||
|
datasets:
|
||||||
|
- wikipedia
|
||||||
|
- Osian
|
||||||
|
- 1.5B-Arabic-Corpus
|
||||||
|
- oscar-arabic-unshuffled
|
||||||
|
- Assafir(private)
|
||||||
|
widget:
|
||||||
|
- text: "يحكى أن مزارعا مخادعا قام ببيع بئر الماء الموجود في أرضه لجاره مقابل مبلغ كبير من المال"
|
||||||
|
- text: "القدس مدينة تاريخية، بناها الكنعانيون في"
|
||||||
|
- text: "كان يا ما كان في قديم الزمان"
|
||||||
|
---
|
||||||
|
|
||||||
|
# Arabic GPT2
|
||||||
|
|
||||||
|
<img src="https://raw.githubusercontent.com/aub-mind/arabert/master/AraGPT2.png" width="100" align="left"/>
|
||||||
|
|
||||||
|
You can find more information in our paper [AraGPT2](https://arxiv.org/abs/2012.15520)
|
||||||
|
|
||||||
|
The code in this repository was used to train all GPT2 variants. The code support training and fine-tuning GPT2 on GPUs and TPUs via the TPUEstimator API.
|
||||||
|
|
||||||
|
GPT2-base and medium uses the code from the `gpt2` folder and can trains models from the [minimaxir/gpt-2-simple](https://github.com/minimaxir/gpt-2-simple) repository.
|
||||||
|
These models were trained using the `lamb` optimizer and follow the same architecture as `gpt2` and are fully compatible with the `transformers` library.
|
||||||
|
|
||||||
|
GPT2-large and GPT2-mega were trained using the [imcaspar/gpt2-ml](https://github.com/imcaspar/gpt2-ml/) library, and follow the `grover` architecture. You can use the pytorch classes found in `grover/modeling_gpt2.py` as a direct replacement for classes in the `transformers` library (it should support version `v4.x` from `transformers`).
|
||||||
|
Both models are trained using the `adafactor` optimizer, since the `adam` and `lamb` optimizer use too much memory causing the model to not even fit 1 batch on a TPU core.
|
||||||
|
|
||||||
|
AraGPT2 is trained on the same large Arabic Dataset as AraBERTv2.
|
||||||
|
|
||||||
|
# Usage
|
||||||
|
|
||||||
|
## Testing the model using `transformers`:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from transformers import GPT2TokenizerFast, pipeline
|
||||||
|
#for base and medium
|
||||||
|
from transformers import GPT2LMHeadModel
|
||||||
|
#for large and mega
|
||||||
|
# pip install arabert
|
||||||
|
from arabert.aragpt2.grover.modeling_gpt2 import GPT2LMHeadModel
|
||||||
|
|
||||||
|
from arabert.preprocess import ArabertPreprocessor
|
||||||
|
|
||||||
|
MODEL_NAME='aubmindlab/aragpt2-base'
|
||||||
|
arabert_prep = ArabertPreprocessor(model_name=MODEL_NAME)
|
||||||
|
|
||||||
|
text=""
|
||||||
|
text_clean = arabert_prep.preprocess(text)
|
||||||
|
|
||||||
|
model = GPT2LMHeadModel.from_pretrained(MODEL_NAME)
|
||||||
|
tokenizer = GPT2TokenizerFast.from_pretrained(MODEL_NAME)
|
||||||
|
generation_pipeline = pipeline("text-generation",model=model,tokenizer=tokenizer)
|
||||||
|
|
||||||
|
#feel free to try different decoding settings
|
||||||
|
generation_pipeline(text,
|
||||||
|
pad_token_id=tokenizer.eos_token_id,
|
||||||
|
num_beams=10,
|
||||||
|
max_length=200,
|
||||||
|
top_p=0.9,
|
||||||
|
repetition_penalty = 3.0,
|
||||||
|
no_repeat_ngram_size = 3)[0]['generated_text']
|
||||||
|
```
|
||||||
|
## Finetunning using `transformers`:
|
||||||
|
|
||||||
|
Follow the guide linked [here](https://towardsdatascience.com/fine-tuning-gpt2-on-colab-gpu-for-free-340468c92ed)
|
||||||
|
|
||||||
|
## Finetuning using our code with TF 1.15.4:
|
||||||
|
|
||||||
|
Create the Training TFRecords:
|
||||||
|
```bash
|
||||||
|
python create_pretraining_data.py
|
||||||
|
--input_file=<RAW TEXT FILE with documents/article separated by an empty line>
|
||||||
|
--output_file=<OUTPUT TFRecord>
|
||||||
|
--tokenizer_dir=<Directory with the GPT2 Tokenizer files>
|
||||||
|
```
|
||||||
|
|
||||||
|
Finetuning:
|
||||||
|
```bash
|
||||||
|
python3 run_pretraining.py \\r\n --input_file="gs://<GS_BUCKET>/pretraining_data/*" \\r\n --output_dir="gs://<GS_BUCKET>/pretraining_model/" \\r\n --config_file="config/small_hparams.json" \\r\n --batch_size=128 \\r\n --eval_batch_size=8 \\r\n --num_train_steps= \\r\n --num_warmup_steps= \\r\n --learning_rate= \\r\n --save_checkpoints_steps= \\r\n --max_seq_length=1024 \\r\n --max_eval_steps= \\r\n --optimizer="lamb" \\r\n --iterations_per_loop=5000 \\r\n --keep_checkpoint_max=10 \\r\n --use_tpu=True \\r\n --tpu_name=<TPU NAME> \\r\n --do_train=True \\r\n --do_eval=False
|
||||||
|
```
|
||||||
|
# Model Sizes
|
||||||
|
|
||||||
|
Model | Optimizer | Context size | Embedding Size | Num of heads | Num of layers | Model Size / Num of Params |
|
||||||
|
---|:---:|:---:|:---:|:---:|:---:|:---:
|
||||||
|
AraGPT2-base | `lamb` | 1024 | 768 | 12 | 12 | 527MB / 135M |
|
||||||
|
AraGPT2-medium | `lamb` | 1024 | 1024 | 16 | 24 | 1.38G/370M |
|
||||||
|
AraGPT2-large | `adafactor` | 1024 | 1280 | 20 | 36 | 2.98GB/792M |
|
||||||
|
AraGPT2-mega | `adafactor` | 1024 | 1536 | 25 | 48 | 5.5GB/1.46B |
|
||||||
|
|
||||||
|
All models are available in the `HuggingFace` model page under the [aubmindlab](https://huggingface.co/aubmindlab/) name. Checkpoints are available in PyTorch, TF2 and TF1 formats.
|
||||||
|
|
||||||
|
## Compute
|
||||||
|
|
||||||
|
Model | Hardware | num of examples (seq len = 1024) | Batch Size | Num of Steps | Time (in days)
|
||||||
|
---|:---:|:---:|:---:|:---:|:---:
|
||||||
|
AraGPT2-base | TPUv3-128 | 9.7M | 1792 | 125K | 1.5
|
||||||
|
AraGPT2-medium | TPUv3-8 | 9.7M | 1152 | 85K | 1.5
|
||||||
|
AraGPT2-large | TPUv3-128 | 9.7M | 256 | 220k | 3
|
||||||
|
AraGPT2-mega | TPUv3-128 | 9.7M | 256 | 780K | 9
|
||||||
|
|
||||||
|
# Dataset
|
||||||
|
|
||||||
|
The pretraining data used for the new AraGPT2 model is also used for **AraBERTv2 and AraELECTRA**.
|
||||||
|
|
||||||
|
The dataset consists of 77GB or 200,095,961 lines or 8,655,948,860 words or 82,232,988,358 chars (before applying Farasa Segmentation)
|
||||||
|
|
||||||
|
For the new dataset we added the unshuffled OSCAR corpus after we thoroughly filter it, to the dataset used in AraBERTv1 but without the websites that we previously crawled:
|
||||||
|
- OSCAR unshuffled and filtered.
|
||||||
|
- [Arabic Wikipedia dump](https://archive.org/details/arwiki-20190201) from 2020/09/01
|
||||||
|
- [The 1.5B words Arabic Corpus](https://www.semanticscholar.org/paper/1.5-billion-words-Arabic-Corpus-El-Khair/f3eeef4afb81223df96575adadf808fe7fe440b4)
|
||||||
|
- [The OSIAN Corpus](https://www.aclweb.org/anthology/W19-4619)
|
||||||
|
- Assafir news articles. Huge thank you for Assafir for giving us the data
|
||||||
|
|
||||||
|
# Disclaimer
|
||||||
|
|
||||||
|
The text generated by AraGPT2 is automatically generated by a neural network model trained on a large amount of texts, which does not represent the authors' or their institutes' official attitudes and preferences. The text generated by AraGPT2 should only be used for research and scientific purposes. If it infringes on your rights and interests or violates social morality, please do not propagate it.
|
||||||
|
|
||||||
|
# If you used this model please cite us as :
|
||||||
|
|
||||||
|
```
|
||||||
|
@inproceedings{antoun-etal-2021-aragpt2,
|
||||||
|
title = "{A}ra{GPT}2: Pre-Trained Transformer for {A}rabic Language Generation",
|
||||||
|
author = "Antoun, Wissam and
|
||||||
|
Baly, Fady and
|
||||||
|
Hajj, Hazem",
|
||||||
|
booktitle = "Proceedings of the Sixth Arabic Natural Language Processing Workshop",
|
||||||
|
month = apr,
|
||||||
|
year = "2021",
|
||||||
|
address = "Kyiv, Ukraine (Virtual)",
|
||||||
|
publisher = "Association for Computational Linguistics",
|
||||||
|
url = "https://www.aclweb.org/anthology/2021.wanlp-1.21",
|
||||||
|
pages = "196--207",
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
# Acknowledgments
|
||||||
|
Thanks to TensorFlow Research Cloud (TFRC) for the free access to Cloud TPUs, couldn't have done it without this program, and to the [AUB MIND Lab](https://sites.aub.edu.lb/mindlab/) Members for the continuous support. Also thanks to [Yakshof](https://www.yakshof.com/#/) and Assafir for data and storage access. Another thanks for Habib Rahal (https://www.behance.net/rahalhabib), for putting a face to AraBERT.
|
||||||
|
|
||||||
|
# Contacts
|
||||||
|
**Wissam Antoun**: [Linkedin](https://www.linkedin.com/in/wissam-antoun-622142b4/) | [Twitter](https://twitter.com/wissam_antoun) | [Github](https://github.com/WissamAntoun) | <wfa07@mail.aub.edu> | <wissam.antoun@gmail.com>
|
||||||
|
|
||||||
|
**Fady Baly**: [Linkedin](https://www.linkedin.com/in/fadybaly/) | [Twitter](https://twitter.com/fadybaly) | [Github](https://github.com/fadybaly) | <fgb06@mail.aub.edu> | <baly.fady@gmail.com>
|
||||||
|
|
||||||
38
config.json
Normal file
38
config.json
Normal file
@@ -0,0 +1,38 @@
|
|||||||
|
{
|
||||||
|
"activation_function": "gelu_new",
|
||||||
|
"architectures": [
|
||||||
|
"GPT2LMHeadModel"
|
||||||
|
],
|
||||||
|
"attn_pdrop": 0.1,
|
||||||
|
"bos_token_id": 0,
|
||||||
|
"embd_pdrop": 0.1,
|
||||||
|
"eos_token_id": 0,
|
||||||
|
"gradient_checkpointing": false,
|
||||||
|
"initializer_range": 0.02,
|
||||||
|
"layer_norm_epsilon": 1e-05,
|
||||||
|
"model_type": "gpt2",
|
||||||
|
"n_ctx": 1024,
|
||||||
|
"n_embd": 768,
|
||||||
|
"n_head": 12,
|
||||||
|
"n_inner": null,
|
||||||
|
"n_layer": 12,
|
||||||
|
"n_positions": 1024,
|
||||||
|
"resid_pdrop": 0.1,
|
||||||
|
"summary_activation": null,
|
||||||
|
"summary_first_dropout": 0.1,
|
||||||
|
"summary_proj_to_labels": true,
|
||||||
|
"summary_type": "cls_index",
|
||||||
|
"summary_use_proj": true,
|
||||||
|
"task_specific_params": {
|
||||||
|
"text-generation": {
|
||||||
|
"do_sample": true,
|
||||||
|
"max_length": 50,
|
||||||
|
"num_beams": 5,
|
||||||
|
"top_p": 0.95,
|
||||||
|
"repetition_penalty": 3.0,
|
||||||
|
"no_repeat_ngram_size": 3
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"use_cache": true,
|
||||||
|
"vocab_size": 64000
|
||||||
|
}
|
||||||
3
flax_model.msgpack
Normal file
3
flax_model.msgpack
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:4b1caf3f0cf7f9bad671ca78178a3a77866c860bc74d04501201219758c50500
|
||||||
|
size 539982616
|
||||||
63741
merges.txt
Normal file
63741
merges.txt
Normal file
File diff suppressed because it is too large
Load Diff
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:6b4b440a9eabb37f4be98063101bb45946dfc163497333859f8cf833d7cef0a9
|
||||||
|
size 552576030
|
||||||
3
pytorch_model.bin
Normal file
3
pytorch_model.bin
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:67cc368c0e4d1e9c44bb4c35e892312ebdfbcf14c2b98e0a9b304de79163d433
|
||||||
|
size 552618653
|
||||||
3
runs/eval/events.out.tfevents.1608681654.tpu-mother
Normal file
3
runs/eval/events.out.tfevents.1608681654.tpu-mother
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:05df33f0f7f994117e51ab6c93e2a8c4d44eff592b66eb51bc7346dcfd4f3d04
|
||||||
|
size 1518487
|
||||||
4
runs/eval_results.txt
Normal file
4
runs/eval_results.txt
Normal file
@@ -0,0 +1,4 @@
|
|||||||
|
bpc = 5.7945323
|
||||||
|
global_step = 125000
|
||||||
|
loss = 4.016464
|
||||||
|
perplexity = 55.811718
|
||||||
3
runs/events.out.tfevents.1608138983.tpu-mother
Normal file
3
runs/events.out.tfevents.1608138983.tpu-mother
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:d03e76d31189a26ebb8b983195ce34c122e1622ab40c088b66e861f56899289c
|
||||||
|
size 31126434
|
||||||
3
runs/events.out.tfevents.1608227376.tpu-mother
Normal file
3
runs/events.out.tfevents.1608227376.tpu-mother
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:5780dee6fd14537dbaa204a030934932da8f9b5caea44d33307b72799360d237
|
||||||
|
size 31124470
|
||||||
3
tf1_model.tar.gz
Normal file
3
tf1_model.tar.gz
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:3766fc03d7c2593ff2fb991d275e96b81b0ecb2098b71ff315611d052ce65248
|
||||||
|
size 501630100
|
||||||
3
tf_model.h5
Normal file
3
tf_model.h5
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:b1d494cee8c1797eade2ae919b366daa749cba0a7b81312755673286c04ca589
|
||||||
|
size 540152144
|
||||||
127810
tokenizer.json
Normal file
127810
tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
1
vocab.json
Normal file
1
vocab.json
Normal file
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user