初始化项目，由ModelHub XC社区提供模型

Model: fpadovani/swe-latn-10mb-ppt-Dp-10mb_seed3407 Source: Original Platform
2026-06-30 19:20:25 +08:00
commit fcfb32cd61
89 changed files with 156332 additions and 0 deletions
--- a/.gitattributes
+++ b/.gitattributes
@@ -0,0 +1,35 @@
 *.7z filter=lfs diff=lfs merge=lfs -text
 *.arrow filter=lfs diff=lfs merge=lfs -text
 *.bin filter=lfs diff=lfs merge=lfs -text
 *.bz2 filter=lfs diff=lfs merge=lfs -text
 *.ckpt filter=lfs diff=lfs merge=lfs -text
 *.ftz filter=lfs diff=lfs merge=lfs -text
 *.gz filter=lfs diff=lfs merge=lfs -text
 *.h5 filter=lfs diff=lfs merge=lfs -text
 *.joblib filter=lfs diff=lfs merge=lfs -text
 *.lfs.* filter=lfs diff=lfs merge=lfs -text
 *.mlmodel filter=lfs diff=lfs merge=lfs -text
 *.model filter=lfs diff=lfs merge=lfs -text
 *.msgpack filter=lfs diff=lfs merge=lfs -text
 *.npy filter=lfs diff=lfs merge=lfs -text
 *.npz filter=lfs diff=lfs merge=lfs -text
 *.onnx filter=lfs diff=lfs merge=lfs -text
 *.ot filter=lfs diff=lfs merge=lfs -text
 *.parquet filter=lfs diff=lfs merge=lfs -text
 *.pb filter=lfs diff=lfs merge=lfs -text
 *.pickle filter=lfs diff=lfs merge=lfs -text
 *.pkl filter=lfs diff=lfs merge=lfs -text
 *.pt filter=lfs diff=lfs merge=lfs -text
 *.pth filter=lfs diff=lfs merge=lfs -text
 *.rar filter=lfs diff=lfs merge=lfs -text
 *.safetensors filter=lfs diff=lfs merge=lfs -text
 saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.tar.* filter=lfs diff=lfs merge=lfs -text
 *.tar filter=lfs diff=lfs merge=lfs -text
 *.tflite filter=lfs diff=lfs merge=lfs -text
 *.tgz filter=lfs diff=lfs merge=lfs -text
 *.wasm filter=lfs diff=lfs merge=lfs -text
 *.xz filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
--- a/README.md
+++ b/README.md
@@ -0,0 +1,58 @@
 ---
 base_model: goldfish-models/swe_latn_10mb
 library_name: transformers
 model_name: swe-latn-10mb-ppt-Dp-10mb_seed3407
 tags:
 - generated_from_trainer
 - trl
 - sft
 licence: license
 ---
 # Model Card for swe-latn-10mb-ppt-Dp-10mb_seed3407
 This model is a fine-tuned version of [goldfish-models/swe_latn_10mb](https://huggingface.co/goldfish-models/swe_latn_10mb).
 It has been trained using [TRL](https://github.com/huggingface/trl).
 ## Quick start
 ```python
 from transformers import pipeline
 question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?"
 generator = pipeline("text-generation", model="fpadovani/swe-latn-10mb-ppt-Dp-10mb_seed3407", device="cuda")
 output = generator([{"role": "user", "content": question}], max_new_tokens=128, return_full_text=False)[0]
 print(output["generated_text"])
 ```
 ## Training procedure
 [<img src="https://raw.githubusercontent.com/wandb/assets/main/wandb-github-badge-28.svg" alt="Visualize in Weights & Biases" width="150" height="24"/>](https://wandb.ai/f-padovani-university-of-groningen/new_tokenizers/runs/i8pred93) 
 This model was trained with SFT.
 ### Framework versions
 - TRL: 0.23.0
 - Transformers: 4.56.2
 - Pytorch: 2.11.0
 - Datasets: 4.8.4
 - Tokenizers: 0.22.1
 ## Citations
 Cite TRL as:
 ```bibtex
@misc{vonwerra2022trl,
 	title        = {{TRL: Transformer Reinforcement Learning}},
 	author       = {Leandro von Werra and Younes Belkada and Lewis Tunstall and Edward Beeching and Tristan Thrush and Nathan Lambert and Shengyi Huang and Kashif Rasul and Quentin Gallou{\'e}dec},
 	year         = 2020,
 	journal      = {GitHub repository},
 	publisher    = {GitHub},
 	howpublished = {\url{https://github.com/huggingface/trl}}
 }
 ```
--- a/added_tokens.json
+++ b/added_tokens.json
--- a/checkpoint-1000/added_tokens.json
+++ b/checkpoint-1000/added_tokens.json
--- a/checkpoint-1000/config.json
+++ b/checkpoint-1000/config.json
@@ -0,0 +1,34 @@
 {
  "activation_function": "gelu",
  "architectures": [
    "GPT2LMHeadModel"
  ],
  "attn_pdrop": 0.1,
  "bos_token_id": 50000,
  "dtype": "bfloat16",
  "embd_pdrop": 0.1,
  "eos_token_id": 50001,
  "initializer_range": 0.02,
  "layer_norm_epsilon": 1e-05,
  "model_type": "gpt2",
  "n_ctx": 512,
  "n_embd": 512,
  "n_head": 8,
  "n_inner": 2048,
  "n_layer": 4,
  "n_positions": 512,
  "pad_token_id": 50002,
  "prefix": "[CLS]",
  "reorder_and_upcast_attn": false,
  "resid_pdrop": 0.1,
  "scale_attn_by_inverse_layer_idx": false,
  "scale_attn_weights": true,
  "summary_activation": null,
  "summary_first_dropout": 0.1,
  "summary_proj_to_labels": true,
  "summary_type": "cls_index",
  "summary_use_proj": true,
  "transformers_version": "4.56.2",
  "use_cache": true,
  "vocab_size": 51200
 }
--- a/checkpoint-1000/generation_config.json
+++ b/checkpoint-1000/generation_config.json
@@ -0,0 +1,9 @@
 {
  "_from_model_config": true,
  "bos_token_id": 50000,
  "eos_token_id": [
    50001
  ],
  "pad_token_id": 50002,
  "transformers_version": "4.56.2"
 }
--- a/checkpoint-1000/model.safetensors
+++ b/checkpoint-1000/model.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:e30a3ae041177539b0d9d1cd3357ea33f174955fbf81647b3d9b5483c8a139bf
 size 78179408
--- a/checkpoint-1000/rng_state.pth
+++ b/checkpoint-1000/rng_state.pth
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:ef8003c89f6b0ed09879f96b09c994300d2d7cd8e639614ceffd0c0beef00f72
 size 14645
--- a/checkpoint-1000/special_tokens_map.json
+++ b/checkpoint-1000/special_tokens_map.json
--- a/checkpoint-1000/spiece.model
+++ b/checkpoint-1000/spiece.model
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:762a1e078899c6ba7ef62ef873a57b860c51cc0518791854659b58b4c89742fe
 size 1170764
--- a/checkpoint-1000/tokenizer_config.json
+++ b/checkpoint-1000/tokenizer_config.json
--- a/checkpoint-1000/trainer_state.json
+++ b/checkpoint-1000/trainer_state.json
--- a/checkpoint-1000/training_args.bin
+++ b/checkpoint-1000/training_args.bin
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:00ed72abd5a1b4e3509cc56cd2c59a708d193071f757752adae72983bc67e472
 size 6289
--- a/checkpoint-1500/added_tokens.json
+++ b/checkpoint-1500/added_tokens.json
--- a/checkpoint-1500/config.json
+++ b/checkpoint-1500/config.json
@@ -0,0 +1,34 @@
 {
  "activation_function": "gelu",
  "architectures": [
    "GPT2LMHeadModel"
  ],
  "attn_pdrop": 0.1,
  "bos_token_id": 50000,
  "dtype": "bfloat16",
  "embd_pdrop": 0.1,
  "eos_token_id": 50001,
  "initializer_range": 0.02,
  "layer_norm_epsilon": 1e-05,
  "model_type": "gpt2",
  "n_ctx": 512,
  "n_embd": 512,
  "n_head": 8,
  "n_inner": 2048,
  "n_layer": 4,
  "n_positions": 512,
  "pad_token_id": 50002,
  "prefix": "[CLS]",
  "reorder_and_upcast_attn": false,
  "resid_pdrop": 0.1,
  "scale_attn_by_inverse_layer_idx": false,
  "scale_attn_weights": true,
  "summary_activation": null,
  "summary_first_dropout": 0.1,
  "summary_proj_to_labels": true,
  "summary_type": "cls_index",
  "summary_use_proj": true,
  "transformers_version": "4.56.2",
  "use_cache": true,
  "vocab_size": 51200
 }
--- a/checkpoint-1500/generation_config.json
+++ b/checkpoint-1500/generation_config.json
@@ -0,0 +1,9 @@
 {
  "_from_model_config": true,
  "bos_token_id": 50000,
  "eos_token_id": [
    50001
  ],
  "pad_token_id": 50002,
  "transformers_version": "4.56.2"
 }
--- a/checkpoint-1500/model.safetensors
+++ b/checkpoint-1500/model.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:163e45167bcadcb00669f84a24895dcb0f4129f8630313730b1b3f53d3e3b593
 size 78179408
--- a/checkpoint-1500/rng_state.pth
+++ b/checkpoint-1500/rng_state.pth
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:c22d5f5c38893cecb9a3f7d0cfd6a811a1f10927dfe8e006260aa91e8c25920d
 size 14645
--- a/checkpoint-1500/special_tokens_map.json
+++ b/checkpoint-1500/special_tokens_map.json
--- a/checkpoint-1500/spiece.model
+++ b/checkpoint-1500/spiece.model
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:762a1e078899c6ba7ef62ef873a57b860c51cc0518791854659b58b4c89742fe
 size 1170764
--- a/checkpoint-1500/tokenizer_config.json
+++ b/checkpoint-1500/tokenizer_config.json
--- a/checkpoint-1500/trainer_state.json
+++ b/checkpoint-1500/trainer_state.json
--- a/checkpoint-1500/training_args.bin
+++ b/checkpoint-1500/training_args.bin
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:00ed72abd5a1b4e3509cc56cd2c59a708d193071f757752adae72983bc67e472
 size 6289
--- a/checkpoint-2000/added_tokens.json
+++ b/checkpoint-2000/added_tokens.json
--- a/checkpoint-2000/config.json
+++ b/checkpoint-2000/config.json
@@ -0,0 +1,34 @@
 {
  "activation_function": "gelu",
  "architectures": [
    "GPT2LMHeadModel"
  ],
  "attn_pdrop": 0.1,
  "bos_token_id": 50000,
  "dtype": "bfloat16",
  "embd_pdrop": 0.1,
  "eos_token_id": 50001,
  "initializer_range": 0.02,
  "layer_norm_epsilon": 1e-05,
  "model_type": "gpt2",
  "n_ctx": 512,
  "n_embd": 512,
  "n_head": 8,
  "n_inner": 2048,
  "n_layer": 4,
  "n_positions": 512,
  "pad_token_id": 50002,
  "prefix": "[CLS]",
  "reorder_and_upcast_attn": false,
  "resid_pdrop": 0.1,
  "scale_attn_by_inverse_layer_idx": false,
  "scale_attn_weights": true,
  "summary_activation": null,
  "summary_first_dropout": 0.1,
  "summary_proj_to_labels": true,
  "summary_type": "cls_index",
  "summary_use_proj": true,
  "transformers_version": "4.56.2",
  "use_cache": true,
  "vocab_size": 51200
 }
--- a/checkpoint-2000/generation_config.json
+++ b/checkpoint-2000/generation_config.json
@@ -0,0 +1,9 @@
 {
  "_from_model_config": true,
  "bos_token_id": 50000,
  "eos_token_id": [
    50001
  ],
  "pad_token_id": 50002,
  "transformers_version": "4.56.2"
 }
--- a/checkpoint-2000/model.safetensors
+++ b/checkpoint-2000/model.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:70e30306355a1072bc49f72f955c938113e19d542f9b0286eeeab5f85bbfee7e
 size 78179408
--- a/checkpoint-2000/rng_state.pth
+++ b/checkpoint-2000/rng_state.pth
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:d757aa823ab164786115b0c94002f9737da08b8882bff95dc4129c6ff9525357
 size 14645
--- a/checkpoint-2000/special_tokens_map.json
+++ b/checkpoint-2000/special_tokens_map.json
--- a/checkpoint-2000/spiece.model
+++ b/checkpoint-2000/spiece.model
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:762a1e078899c6ba7ef62ef873a57b860c51cc0518791854659b58b4c89742fe
 size 1170764
--- a/checkpoint-2000/tokenizer_config.json
+++ b/checkpoint-2000/tokenizer_config.json
--- a/checkpoint-2000/trainer_state.json
+++ b/checkpoint-2000/trainer_state.json
--- a/checkpoint-2000/training_args.bin
+++ b/checkpoint-2000/training_args.bin
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:00ed72abd5a1b4e3509cc56cd2c59a708d193071f757752adae72983bc67e472
 size 6289
--- a/checkpoint-2500/added_tokens.json
+++ b/checkpoint-2500/added_tokens.json
--- a/checkpoint-2500/config.json
+++ b/checkpoint-2500/config.json
@@ -0,0 +1,34 @@
 {
  "activation_function": "gelu",
  "architectures": [
    "GPT2LMHeadModel"
  ],
  "attn_pdrop": 0.1,
  "bos_token_id": 50000,
  "dtype": "bfloat16",
  "embd_pdrop": 0.1,
  "eos_token_id": 50001,
  "initializer_range": 0.02,
  "layer_norm_epsilon": 1e-05,
  "model_type": "gpt2",
  "n_ctx": 512,
  "n_embd": 512,
  "n_head": 8,
  "n_inner": 2048,
  "n_layer": 4,
  "n_positions": 512,
  "pad_token_id": 50002,
  "prefix": "[CLS]",
  "reorder_and_upcast_attn": false,
  "resid_pdrop": 0.1,
  "scale_attn_by_inverse_layer_idx": false,
  "scale_attn_weights": true,
  "summary_activation": null,
  "summary_first_dropout": 0.1,
  "summary_proj_to_labels": true,
  "summary_type": "cls_index",
  "summary_use_proj": true,
  "transformers_version": "4.56.2",
  "use_cache": true,
  "vocab_size": 51200
 }
--- a/checkpoint-2500/generation_config.json
+++ b/checkpoint-2500/generation_config.json
@@ -0,0 +1,9 @@
 {
  "_from_model_config": true,
  "bos_token_id": 50000,
  "eos_token_id": [
    50001
  ],
  "pad_token_id": 50002,
  "transformers_version": "4.56.2"
 }
--- a/checkpoint-2500/model.safetensors
+++ b/checkpoint-2500/model.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:c1083d2a0c9c4c87027cef4e0bc24d49bce1564714ce50b4329e032cef9a752e
 size 78179408
--- a/checkpoint-2500/rng_state.pth
+++ b/checkpoint-2500/rng_state.pth
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:44b46e61ebdca1316bc72172bbb6c59874a9439d381ad98da649ec96416851f0
 size 14645
--- a/checkpoint-2500/special_tokens_map.json
+++ b/checkpoint-2500/special_tokens_map.json
--- a/checkpoint-2500/spiece.model
+++ b/checkpoint-2500/spiece.model
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:762a1e078899c6ba7ef62ef873a57b860c51cc0518791854659b58b4c89742fe
 size 1170764
--- a/checkpoint-2500/tokenizer_config.json
+++ b/checkpoint-2500/tokenizer_config.json
--- a/checkpoint-2500/trainer_state.json
+++ b/checkpoint-2500/trainer_state.json
--- a/checkpoint-2500/training_args.bin
+++ b/checkpoint-2500/training_args.bin
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:00ed72abd5a1b4e3509cc56cd2c59a708d193071f757752adae72983bc67e472
 size 6289
--- a/checkpoint-3000/added_tokens.json
+++ b/checkpoint-3000/added_tokens.json
--- a/checkpoint-3000/config.json
+++ b/checkpoint-3000/config.json
@@ -0,0 +1,34 @@
 {
  "activation_function": "gelu",
  "architectures": [
    "GPT2LMHeadModel"
  ],
  "attn_pdrop": 0.1,
  "bos_token_id": 50000,
  "dtype": "bfloat16",
  "embd_pdrop": 0.1,
  "eos_token_id": 50001,
  "initializer_range": 0.02,
  "layer_norm_epsilon": 1e-05,
  "model_type": "gpt2",
  "n_ctx": 512,
  "n_embd": 512,
  "n_head": 8,
  "n_inner": 2048,
  "n_layer": 4,
  "n_positions": 512,
  "pad_token_id": 50002,
  "prefix": "[CLS]",
  "reorder_and_upcast_attn": false,
  "resid_pdrop": 0.1,
  "scale_attn_by_inverse_layer_idx": false,
  "scale_attn_weights": true,
  "summary_activation": null,
  "summary_first_dropout": 0.1,
  "summary_proj_to_labels": true,
  "summary_type": "cls_index",
  "summary_use_proj": true,
  "transformers_version": "4.56.2",
  "use_cache": true,
  "vocab_size": 51200
 }
--- a/checkpoint-3000/generation_config.json
+++ b/checkpoint-3000/generation_config.json
@@ -0,0 +1,9 @@
 {
  "_from_model_config": true,
  "bos_token_id": 50000,
  "eos_token_id": [
    50001
  ],
  "pad_token_id": 50002,
  "transformers_version": "4.56.2"
 }
--- a/checkpoint-3000/model.safetensors
+++ b/checkpoint-3000/model.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:28a144b9ced43c1417efa2615f698a376c1f616698f5cedbb7722f9f85ed048b
 size 78179408
--- a/checkpoint-3000/rng_state.pth
+++ b/checkpoint-3000/rng_state.pth
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:430d2d469c1c4bea0d589f49c6ccb61285f38233b3e689cd9058366f0d58cd10
 size 14645
--- a/checkpoint-3000/special_tokens_map.json
+++ b/checkpoint-3000/special_tokens_map.json
--- a/checkpoint-3000/spiece.model
+++ b/checkpoint-3000/spiece.model
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:762a1e078899c6ba7ef62ef873a57b860c51cc0518791854659b58b4c89742fe
 size 1170764
--- a/checkpoint-3000/tokenizer_config.json
+++ b/checkpoint-3000/tokenizer_config.json
--- a/checkpoint-3000/trainer_state.json
+++ b/checkpoint-3000/trainer_state.json
--- a/checkpoint-3000/training_args.bin
+++ b/checkpoint-3000/training_args.bin
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:00ed72abd5a1b4e3509cc56cd2c59a708d193071f757752adae72983bc67e472
 size 6289
--- a/checkpoint-3500/added_tokens.json
+++ b/checkpoint-3500/added_tokens.json
--- a/checkpoint-3500/config.json
+++ b/checkpoint-3500/config.json
@@ -0,0 +1,34 @@
 {
  "activation_function": "gelu",
  "architectures": [
    "GPT2LMHeadModel"
  ],
  "attn_pdrop": 0.1,
  "bos_token_id": 50000,
  "dtype": "bfloat16",
  "embd_pdrop": 0.1,
  "eos_token_id": 50001,
  "initializer_range": 0.02,
  "layer_norm_epsilon": 1e-05,
  "model_type": "gpt2",
  "n_ctx": 512,
  "n_embd": 512,
  "n_head": 8,
  "n_inner": 2048,
  "n_layer": 4,
  "n_positions": 512,
  "pad_token_id": 50002,
  "prefix": "[CLS]",
  "reorder_and_upcast_attn": false,
  "resid_pdrop": 0.1,
  "scale_attn_by_inverse_layer_idx": false,
  "scale_attn_weights": true,
  "summary_activation": null,
  "summary_first_dropout": 0.1,
  "summary_proj_to_labels": true,
  "summary_type": "cls_index",
  "summary_use_proj": true,
  "transformers_version": "4.56.2",
  "use_cache": true,
  "vocab_size": 51200
 }
--- a/checkpoint-3500/generation_config.json
+++ b/checkpoint-3500/generation_config.json
@@ -0,0 +1,9 @@
 {
  "_from_model_config": true,
  "bos_token_id": 50000,
  "eos_token_id": [
    50001
  ],
  "pad_token_id": 50002,
  "transformers_version": "4.56.2"
 }
--- a/checkpoint-3500/model.safetensors
+++ b/checkpoint-3500/model.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:5e5f32c8f7ca25bacc727cb256410dc184c2fa36df9bfa15816df5caf0c2d82c
 size 78179408
--- a/checkpoint-3500/rng_state.pth
+++ b/checkpoint-3500/rng_state.pth
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:d591dae9b9a955431274121d66732e39c69e4ac96a3257d9e4f8e20266e93235
 size 14645
--- a/checkpoint-3500/special_tokens_map.json
+++ b/checkpoint-3500/special_tokens_map.json
--- a/checkpoint-3500/spiece.model
+++ b/checkpoint-3500/spiece.model
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:762a1e078899c6ba7ef62ef873a57b860c51cc0518791854659b58b4c89742fe
 size 1170764
--- a/checkpoint-3500/tokenizer_config.json
+++ b/checkpoint-3500/tokenizer_config.json
--- a/checkpoint-3500/trainer_state.json
+++ b/checkpoint-3500/trainer_state.json
--- a/checkpoint-3500/training_args.bin
+++ b/checkpoint-3500/training_args.bin
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:00ed72abd5a1b4e3509cc56cd2c59a708d193071f757752adae72983bc67e472
 size 6289
--- a/checkpoint-4000/added_tokens.json
+++ b/checkpoint-4000/added_tokens.json
--- a/checkpoint-4000/config.json
+++ b/checkpoint-4000/config.json
@@ -0,0 +1,34 @@
 {
  "activation_function": "gelu",
  "architectures": [
    "GPT2LMHeadModel"
  ],
  "attn_pdrop": 0.1,
  "bos_token_id": 50000,
  "dtype": "bfloat16",
  "embd_pdrop": 0.1,
  "eos_token_id": 50001,
  "initializer_range": 0.02,
  "layer_norm_epsilon": 1e-05,
  "model_type": "gpt2",
  "n_ctx": 512,
  "n_embd": 512,
  "n_head": 8,
  "n_inner": 2048,
  "n_layer": 4,
  "n_positions": 512,
  "pad_token_id": 50002,
  "prefix": "[CLS]",
  "reorder_and_upcast_attn": false,
  "resid_pdrop": 0.1,
  "scale_attn_by_inverse_layer_idx": false,
  "scale_attn_weights": true,
  "summary_activation": null,
  "summary_first_dropout": 0.1,
  "summary_proj_to_labels": true,
  "summary_type": "cls_index",
  "summary_use_proj": true,
  "transformers_version": "4.56.2",
  "use_cache": true,
  "vocab_size": 51200
 }
--- a/checkpoint-4000/generation_config.json
+++ b/checkpoint-4000/generation_config.json
@@ -0,0 +1,9 @@
 {
  "_from_model_config": true,
  "bos_token_id": 50000,
  "eos_token_id": [
    50001
  ],
  "pad_token_id": 50002,
  "transformers_version": "4.56.2"
 }
--- a/checkpoint-4000/model.safetensors
+++ b/checkpoint-4000/model.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:4d22a5a6eae22237ec024cb0aeafd429d0ddd3b7fe914442629d57122873501c
 size 78179408
--- a/checkpoint-4000/rng_state.pth
+++ b/checkpoint-4000/rng_state.pth
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:fe1c81012bd24e8f4ec06b95be1b9e7c13e9d1168c3410b1498a9664bcd34123
 size 14645
--- a/checkpoint-4000/special_tokens_map.json
+++ b/checkpoint-4000/special_tokens_map.json
--- a/checkpoint-4000/spiece.model
+++ b/checkpoint-4000/spiece.model
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:762a1e078899c6ba7ef62ef873a57b860c51cc0518791854659b58b4c89742fe
 size 1170764
--- a/checkpoint-4000/tokenizer_config.json
+++ b/checkpoint-4000/tokenizer_config.json
--- a/checkpoint-4000/trainer_state.json
+++ b/checkpoint-4000/trainer_state.json
--- a/checkpoint-4000/training_args.bin
+++ b/checkpoint-4000/training_args.bin
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:00ed72abd5a1b4e3509cc56cd2c59a708d193071f757752adae72983bc67e472
 size 6289
--- a/checkpoint-500/added_tokens.json
+++ b/checkpoint-500/added_tokens.json
--- a/checkpoint-500/config.json
+++ b/checkpoint-500/config.json
@@ -0,0 +1,34 @@
 {
  "activation_function": "gelu",
  "architectures": [
    "GPT2LMHeadModel"
  ],
  "attn_pdrop": 0.1,
  "bos_token_id": 50000,
  "dtype": "bfloat16",
  "embd_pdrop": 0.1,
  "eos_token_id": 50001,
  "initializer_range": 0.02,
  "layer_norm_epsilon": 1e-05,
  "model_type": "gpt2",
  "n_ctx": 512,
  "n_embd": 512,
  "n_head": 8,
  "n_inner": 2048,
  "n_layer": 4,
  "n_positions": 512,
  "pad_token_id": 50002,
  "prefix": "[CLS]",
  "reorder_and_upcast_attn": false,
  "resid_pdrop": 0.1,
  "scale_attn_by_inverse_layer_idx": false,
  "scale_attn_weights": true,
  "summary_activation": null,
  "summary_first_dropout": 0.1,
  "summary_proj_to_labels": true,
  "summary_type": "cls_index",
  "summary_use_proj": true,
  "transformers_version": "4.56.2",
  "use_cache": true,
  "vocab_size": 51200
 }
--- a/checkpoint-500/generation_config.json
+++ b/checkpoint-500/generation_config.json
@@ -0,0 +1,9 @@
 {
  "_from_model_config": true,
  "bos_token_id": 50000,
  "eos_token_id": [
    50001
  ],
  "pad_token_id": 50002,
  "transformers_version": "4.56.2"
 }
--- a/checkpoint-500/model.safetensors
+++ b/checkpoint-500/model.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:6718d19a50900a1403f2830adb7389f12715074494f95ff86462b3e14fba1549
 size 78179408
--- a/checkpoint-500/rng_state.pth
+++ b/checkpoint-500/rng_state.pth
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:5cb9bd3ca7a0aa2b8c883820c0a28dae24a9c9a026459495522bb2070ec0e024
 size 14645
--- a/checkpoint-500/special_tokens_map.json
+++ b/checkpoint-500/special_tokens_map.json
--- a/checkpoint-500/spiece.model
+++ b/checkpoint-500/spiece.model
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:762a1e078899c6ba7ef62ef873a57b860c51cc0518791854659b58b4c89742fe
 size 1170764
--- a/checkpoint-500/tokenizer_config.json
+++ b/checkpoint-500/tokenizer_config.json
--- a/checkpoint-500/trainer_state.json
+++ b/checkpoint-500/trainer_state.json
--- a/checkpoint-500/training_args.bin
+++ b/checkpoint-500/training_args.bin
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:00ed72abd5a1b4e3509cc56cd2c59a708d193071f757752adae72983bc67e472
 size 6289
--- a/config.json
+++ b/config.json
@@ -0,0 +1,34 @@
 {
  "activation_function": "gelu",
  "architectures": [
    "GPT2LMHeadModel"
  ],
  "attn_pdrop": 0.1,
  "bos_token_id": 50000,
  "dtype": "bfloat16",
  "embd_pdrop": 0.1,
  "eos_token_id": 50001,
  "initializer_range": 0.02,
  "layer_norm_epsilon": 1e-05,
  "model_type": "gpt2",
  "n_ctx": 512,
  "n_embd": 512,
  "n_head": 8,
  "n_inner": 2048,
  "n_layer": 4,
  "n_positions": 512,
  "pad_token_id": 50002,
  "prefix": "[CLS]",
  "reorder_and_upcast_attn": false,
  "resid_pdrop": 0.1,
  "scale_attn_by_inverse_layer_idx": false,
  "scale_attn_weights": true,
  "summary_activation": null,
  "summary_first_dropout": 0.1,
  "summary_proj_to_labels": true,
  "summary_type": "cls_index",
  "summary_use_proj": true,
  "transformers_version": "4.56.2",
  "use_cache": true,
  "vocab_size": 51200
 }
--- a/model.safetensors
+++ b/model.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:4d22a5a6eae22237ec024cb0aeafd429d0ddd3b7fe914442629d57122873501c
 size 78179408
--- a/special_tokens_map.json
+++ b/special_tokens_map.json
--- a/spiece.model
+++ b/spiece.model
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:762a1e078899c6ba7ef62ef873a57b860c51cc0518791854659b58b4c89742fe
 size 1170764
--- a/tokenizer_config.json
+++ b/tokenizer_config.json
--- a/training_args.bin
+++ b/training_args.bin
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:00ed72abd5a1b4e3509cc56cd2c59a708d193071f757752adae72983bc67e472
 size 6289