初始化项目,由ModelHub XC社区提供模型

Model: anwgpt/anwgpt4-chat
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-06-12 03:39:17 +08:00
commit 0ba3733fff
37 changed files with 301953 additions and 0 deletions

35
.gitattributes vendored Normal file
View File

@@ -0,0 +1,35 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

60
README.md Normal file
View File

@@ -0,0 +1,60 @@
---
base_model:
- anwgpt/anwgpt4-base
license: mit
language:
- en
pipeline_tag: text-generation
library_name: transformers
datasets:
- HuggingFaceH4/ultrachat_200k
- tatsu-lab/alpaca
- OpenAssistant/oasst1
- databricks/databricks-dolly-15k
tags:
- conversational
---
# ANWGPT-4
## How do use the model
```py
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-generation", model="anwgpt/anwgpt4-chat")
# Example of text generation (you can modify the input text but keep the format like this)
print(pipe("""
Human: How much is 1 + 1?
Assistant:
"""))
```
Sample:
```
Prompt: Human: What is artificial intelligence?
Assistant:
Response: AI is a type of AI that can be used to describe the environment, which can help you communicate and solve problems. It's hard to understand how computers can interact with different perspectives on their goals and make decisions. Artificial intelligence has been developed by humans and human-driven learning about two. In addition, they could use machine learning algorithms like humans, data processing, coding software processing, and network programming, for instance, as well as applications in computer systems. AI can provide tools such as data detection
--------------------------------------------------
Prompt: Human: How do neural networks work?
Assistant:
Response: <i>A ] is a type of programming language. The neural network that can be used to create a digital model on an electronic computer for the entire project, so it's best to understand what they did do. <i+P) is used in a variety of applications and applications for each element. This concept involves providing user-in code that can help solve problems and behavior between humans and their environment. In general, the neural networks creation process will help with processing data into
--------------------------------------------------
Prompt: Human: Who created you?
Assistant:
Response: I'm sorry, but it's impossible to know about. The answer was written in the text material, and this is a great choice for someone who would like to speak with people on a budget, making it easy to learn, especially when asked about how much money or not to take care of your family. It only states that all the things are made up by family members, and what are the best ones they want to go into account and make care more interesting. One way to help me find
```
Benchmark(s):
![image](https://cdn-uploads.huggingface.co/production/uploads/6615494716917dfdc645c44e/qbGuErFWfM2w2kdqbCqhJ.png)

View File

@@ -0,0 +1,31 @@
{
"activation_function": "gelu_new",
"architectures": [
"GPT2LMHeadModel"
],
"attn_pdrop": 0.1,
"bos_token_id": 50256,
"dtype": "float32",
"embd_pdrop": 0.1,
"eos_token_id": 50256,
"initializer_range": 0.02,
"layer_norm_epsilon": 1e-05,
"model_type": "gpt2",
"n_embd": 384,
"n_head": 6,
"n_inner": null,
"n_layer": 4,
"n_positions": 2048,
"reorder_and_upcast_attn": false,
"resid_pdrop": 0.1,
"scale_attn_by_inverse_layer_idx": false,
"scale_attn_weights": true,
"summary_activation": null,
"summary_first_dropout": 0.1,
"summary_proj_to_labels": true,
"summary_type": "cls_index",
"summary_use_proj": true,
"transformers_version": "4.57.1",
"use_cache": false,
"vocab_size": 50257
}

View File

@@ -0,0 +1,7 @@
{
"_from_model_config": true,
"bos_token_id": 50256,
"eos_token_id": 50256,
"transformers_version": "4.57.1",
"use_cache": false
}

50001
checkpoint-2000/merges.txt Normal file

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ec9fb1f5dd3a2fddc00de91d9d391439016ef56e6f19814ef8ec031e91dae95e
size 108740128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dad90d3b9ca415b74aa0e528df816edf4f52f2f11584a55f262ab3c2fc0299ba
size 217514251

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bbca87195ddb7c976613970f2e293fcdbc973f5cb9637873e0bbc4dbbf60f1f8
size 14645

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f4aa03f6e0cd07cf67ce1fbe3101d545f5771ef9148b9debf02b11cf6948da5c
size 1383

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cf285529fcf49663bb4925a04b62f5839a30e880bf24eccbe6a8bcf058f1bda8
size 1465

View File

@@ -0,0 +1,24 @@
{
"bos_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
},
"eos_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
},
"pad_token": "<|endoftext|>",
"unk_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
}
}

View File

@@ -0,0 +1,23 @@
{
"add_bos_token": false,
"add_prefix_space": false,
"added_tokens_decoder": {
"50256": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false,
"special": true
}
},
"bos_token": "<|endoftext|>",
"clean_up_tokenization_spaces": false,
"eos_token": "<|endoftext|>",
"errors": "replace",
"extra_special_tokens": {},
"model_max_length": 1024,
"pad_token": "<|endoftext|>",
"tokenizer_class": "GPT2Tokenizer",
"unk_token": "<|endoftext|>"
}

View File

@@ -0,0 +1,314 @@
{
"best_global_step": null,
"best_metric": null,
"best_model_checkpoint": null,
"epoch": 8.0,
"eval_steps": 500,
"global_step": 2000,
"is_hyper_param_search": false,
"is_local_process_zero": true,
"is_world_process_zero": true,
"log_history": [
{
"epoch": 0.2,
"grad_norm": 2.671468734741211,
"learning_rate": 6.125e-05,
"loss": 8.0449,
"step": 50
},
{
"epoch": 0.4,
"grad_norm": 1.9384660720825195,
"learning_rate": 0.00012375,
"loss": 6.1065,
"step": 100
},
{
"epoch": 0.6,
"grad_norm": 1.7404013872146606,
"learning_rate": 0.00018625,
"loss": 5.552,
"step": 150
},
{
"epoch": 0.8,
"grad_norm": 1.5962367057800293,
"learning_rate": 0.00024875,
"loss": 5.3272,
"step": 200
},
{
"epoch": 1.0,
"grad_norm": 1.6055867671966553,
"learning_rate": 0.000245625,
"loss": 5.1701,
"step": 250
},
{
"epoch": 1.2,
"grad_norm": 1.5786716938018799,
"learning_rate": 0.0002411607142857143,
"loss": 4.7686,
"step": 300
},
{
"epoch": 1.4,
"grad_norm": 1.5948882102966309,
"learning_rate": 0.00023669642857142856,
"loss": 4.6684,
"step": 350
},
{
"epoch": 1.6,
"grad_norm": 1.5991402864456177,
"learning_rate": 0.00023223214285714286,
"loss": 4.6017,
"step": 400
},
{
"epoch": 1.8,
"grad_norm": 1.4955918788909912,
"learning_rate": 0.00022776785714285713,
"loss": 4.549,
"step": 450
},
{
"epoch": 2.0,
"grad_norm": 1.5773422718048096,
"learning_rate": 0.00022330357142857143,
"loss": 4.4926,
"step": 500
},
{
"epoch": 2.2,
"grad_norm": 1.5831668376922607,
"learning_rate": 0.00021883928571428572,
"loss": 4.1386,
"step": 550
},
{
"epoch": 2.4,
"grad_norm": 1.3561625480651855,
"learning_rate": 0.00021437500000000002,
"loss": 4.1139,
"step": 600
},
{
"epoch": 2.6,
"grad_norm": 1.4930239915847778,
"learning_rate": 0.0002099107142857143,
"loss": 4.1189,
"step": 650
},
{
"epoch": 2.8,
"grad_norm": 1.5223215818405151,
"learning_rate": 0.00020544642857142856,
"loss": 4.0464,
"step": 700
},
{
"epoch": 3.0,
"grad_norm": 1.6128135919570923,
"learning_rate": 0.00020098214285714286,
"loss": 4.0405,
"step": 750
},
{
"epoch": 3.2,
"grad_norm": 1.5877256393432617,
"learning_rate": 0.00019651785714285713,
"loss": 3.7492,
"step": 800
},
{
"epoch": 3.4,
"grad_norm": 1.5449103116989136,
"learning_rate": 0.00019205357142857143,
"loss": 3.8028,
"step": 850
},
{
"epoch": 3.6,
"grad_norm": 1.4588514566421509,
"learning_rate": 0.00018758928571428572,
"loss": 3.752,
"step": 900
},
{
"epoch": 3.8,
"grad_norm": 1.4356666803359985,
"learning_rate": 0.00018312500000000002,
"loss": 3.7031,
"step": 950
},
{
"epoch": 4.0,
"grad_norm": 1.5264209508895874,
"learning_rate": 0.0001786607142857143,
"loss": 3.7287,
"step": 1000
},
{
"epoch": 4.2,
"grad_norm": 1.351216197013855,
"learning_rate": 0.00017419642857142856,
"loss": 3.4962,
"step": 1050
},
{
"epoch": 4.4,
"grad_norm": 1.6168186664581299,
"learning_rate": 0.00016973214285714286,
"loss": 3.5006,
"step": 1100
},
{
"epoch": 4.6,
"grad_norm": 1.6266807317733765,
"learning_rate": 0.00016526785714285713,
"loss": 3.4669,
"step": 1150
},
{
"epoch": 4.8,
"grad_norm": 1.509149432182312,
"learning_rate": 0.00016080357142857142,
"loss": 3.5044,
"step": 1200
},
{
"epoch": 5.0,
"grad_norm": 1.610813856124878,
"learning_rate": 0.00015633928571428572,
"loss": 3.4626,
"step": 1250
},
{
"epoch": 5.2,
"grad_norm": 1.3284097909927368,
"learning_rate": 0.00015187500000000002,
"loss": 3.2758,
"step": 1300
},
{
"epoch": 5.4,
"grad_norm": 1.5338636636734009,
"learning_rate": 0.0001474107142857143,
"loss": 3.3102,
"step": 1350
},
{
"epoch": 5.6,
"grad_norm": 1.4522225856781006,
"learning_rate": 0.00014294642857142856,
"loss": 3.2593,
"step": 1400
},
{
"epoch": 5.8,
"grad_norm": 1.6232526302337646,
"learning_rate": 0.00013848214285714286,
"loss": 3.2785,
"step": 1450
},
{
"epoch": 6.0,
"grad_norm": 1.6043927669525146,
"learning_rate": 0.00013401785714285713,
"loss": 3.3056,
"step": 1500
},
{
"epoch": 6.2,
"grad_norm": 1.6095658540725708,
"learning_rate": 0.00012955357142857142,
"loss": 3.1076,
"step": 1550
},
{
"epoch": 6.4,
"grad_norm": 1.4529067277908325,
"learning_rate": 0.00012508928571428572,
"loss": 3.1206,
"step": 1600
},
{
"epoch": 6.6,
"grad_norm": 1.7432512044906616,
"learning_rate": 0.000120625,
"loss": 3.1245,
"step": 1650
},
{
"epoch": 6.8,
"grad_norm": 1.7623839378356934,
"learning_rate": 0.00011616071428571429,
"loss": 3.1437,
"step": 1700
},
{
"epoch": 7.0,
"grad_norm": 1.537539005279541,
"learning_rate": 0.00011169642857142857,
"loss": 3.121,
"step": 1750
},
{
"epoch": 7.2,
"grad_norm": 1.4343411922454834,
"learning_rate": 0.00010723214285714286,
"loss": 2.9752,
"step": 1800
},
{
"epoch": 7.4,
"grad_norm": 1.4659862518310547,
"learning_rate": 0.00010276785714285715,
"loss": 2.9755,
"step": 1850
},
{
"epoch": 7.6,
"grad_norm": 1.4027365446090698,
"learning_rate": 9.830357142857144e-05,
"loss": 2.9823,
"step": 1900
},
{
"epoch": 7.8,
"grad_norm": 1.6034783124923706,
"learning_rate": 9.383928571428571e-05,
"loss": 2.9971,
"step": 1950
},
{
"epoch": 8.0,
"grad_norm": 1.9349607229232788,
"learning_rate": 8.9375e-05,
"loss": 3.0408,
"step": 2000
}
],
"logging_steps": 50,
"max_steps": 3000,
"num_input_tokens_seen": 0,
"num_train_epochs": 12,
"save_steps": 1000,
"stateful_callbacks": {
"TrainerControl": {
"args": {
"should_epoch_stop": false,
"should_evaluate": false,
"should_log": false,
"should_save": true,
"should_training_stop": false
},
"attributes": {}
}
},
"total_flos": 1395646267392000.0,
"train_batch_size": 2,
"trial_name": null,
"trial_params": null
}

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:577aa7fa92f0f8bd6bf3e4ff2191a25d48cf0e08199280935bab389e8179cf45
size 5777

50259
checkpoint-2000/vocab.json Normal file

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,31 @@
{
"activation_function": "gelu_new",
"architectures": [
"GPT2LMHeadModel"
],
"attn_pdrop": 0.1,
"bos_token_id": 50256,
"dtype": "float32",
"embd_pdrop": 0.1,
"eos_token_id": 50256,
"initializer_range": 0.02,
"layer_norm_epsilon": 1e-05,
"model_type": "gpt2",
"n_embd": 384,
"n_head": 6,
"n_inner": null,
"n_layer": 4,
"n_positions": 2048,
"reorder_and_upcast_attn": false,
"resid_pdrop": 0.1,
"scale_attn_by_inverse_layer_idx": false,
"scale_attn_weights": true,
"summary_activation": null,
"summary_first_dropout": 0.1,
"summary_proj_to_labels": true,
"summary_type": "cls_index",
"summary_use_proj": true,
"transformers_version": "4.57.1",
"use_cache": false,
"vocab_size": 50257
}

View File

@@ -0,0 +1,7 @@
{
"_from_model_config": true,
"bos_token_id": 50256,
"eos_token_id": 50256,
"transformers_version": "4.57.1",
"use_cache": false
}

50001
checkpoint-3000/merges.txt Normal file

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f9214a52ad9906206fed86ad77c5938ae7f876d235c670d94874baac2b1dbd88
size 108740128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f9bb6924fb278b8cdf5b88563ba2436ea3406ad7e0f4fc2bf4196d771512da0c
size 217514251

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:394b4e601f238aee0967d92319d788baa6971061222ead76e8851d0dfecfaeba
size 14645

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5ac1c46a2776d12775d23d0f587efc112188137ce2140da35bc15d301c9f620e
size 1383

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:edc55f6faeebfa53703ec8a1cb2786f2c7ed308cab9f4bb9e6acc6f1608ed63f
size 1465

View File

@@ -0,0 +1,24 @@
{
"bos_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
},
"eos_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
},
"pad_token": "<|endoftext|>",
"unk_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
}
}

View File

@@ -0,0 +1,23 @@
{
"add_bos_token": false,
"add_prefix_space": false,
"added_tokens_decoder": {
"50256": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false,
"special": true
}
},
"bos_token": "<|endoftext|>",
"clean_up_tokenization_spaces": false,
"eos_token": "<|endoftext|>",
"errors": "replace",
"extra_special_tokens": {},
"model_max_length": 1024,
"pad_token": "<|endoftext|>",
"tokenizer_class": "GPT2Tokenizer",
"unk_token": "<|endoftext|>"
}

View File

@@ -0,0 +1,454 @@
{
"best_global_step": null,
"best_metric": null,
"best_model_checkpoint": null,
"epoch": 12.0,
"eval_steps": 500,
"global_step": 3000,
"is_hyper_param_search": false,
"is_local_process_zero": true,
"is_world_process_zero": true,
"log_history": [
{
"epoch": 0.2,
"grad_norm": 2.671468734741211,
"learning_rate": 6.125e-05,
"loss": 8.0449,
"step": 50
},
{
"epoch": 0.4,
"grad_norm": 1.9384660720825195,
"learning_rate": 0.00012375,
"loss": 6.1065,
"step": 100
},
{
"epoch": 0.6,
"grad_norm": 1.7404013872146606,
"learning_rate": 0.00018625,
"loss": 5.552,
"step": 150
},
{
"epoch": 0.8,
"grad_norm": 1.5962367057800293,
"learning_rate": 0.00024875,
"loss": 5.3272,
"step": 200
},
{
"epoch": 1.0,
"grad_norm": 1.6055867671966553,
"learning_rate": 0.000245625,
"loss": 5.1701,
"step": 250
},
{
"epoch": 1.2,
"grad_norm": 1.5786716938018799,
"learning_rate": 0.0002411607142857143,
"loss": 4.7686,
"step": 300
},
{
"epoch": 1.4,
"grad_norm": 1.5948882102966309,
"learning_rate": 0.00023669642857142856,
"loss": 4.6684,
"step": 350
},
{
"epoch": 1.6,
"grad_norm": 1.5991402864456177,
"learning_rate": 0.00023223214285714286,
"loss": 4.6017,
"step": 400
},
{
"epoch": 1.8,
"grad_norm": 1.4955918788909912,
"learning_rate": 0.00022776785714285713,
"loss": 4.549,
"step": 450
},
{
"epoch": 2.0,
"grad_norm": 1.5773422718048096,
"learning_rate": 0.00022330357142857143,
"loss": 4.4926,
"step": 500
},
{
"epoch": 2.2,
"grad_norm": 1.5831668376922607,
"learning_rate": 0.00021883928571428572,
"loss": 4.1386,
"step": 550
},
{
"epoch": 2.4,
"grad_norm": 1.3561625480651855,
"learning_rate": 0.00021437500000000002,
"loss": 4.1139,
"step": 600
},
{
"epoch": 2.6,
"grad_norm": 1.4930239915847778,
"learning_rate": 0.0002099107142857143,
"loss": 4.1189,
"step": 650
},
{
"epoch": 2.8,
"grad_norm": 1.5223215818405151,
"learning_rate": 0.00020544642857142856,
"loss": 4.0464,
"step": 700
},
{
"epoch": 3.0,
"grad_norm": 1.6128135919570923,
"learning_rate": 0.00020098214285714286,
"loss": 4.0405,
"step": 750
},
{
"epoch": 3.2,
"grad_norm": 1.5877256393432617,
"learning_rate": 0.00019651785714285713,
"loss": 3.7492,
"step": 800
},
{
"epoch": 3.4,
"grad_norm": 1.5449103116989136,
"learning_rate": 0.00019205357142857143,
"loss": 3.8028,
"step": 850
},
{
"epoch": 3.6,
"grad_norm": 1.4588514566421509,
"learning_rate": 0.00018758928571428572,
"loss": 3.752,
"step": 900
},
{
"epoch": 3.8,
"grad_norm": 1.4356666803359985,
"learning_rate": 0.00018312500000000002,
"loss": 3.7031,
"step": 950
},
{
"epoch": 4.0,
"grad_norm": 1.5264209508895874,
"learning_rate": 0.0001786607142857143,
"loss": 3.7287,
"step": 1000
},
{
"epoch": 4.2,
"grad_norm": 1.351216197013855,
"learning_rate": 0.00017419642857142856,
"loss": 3.4962,
"step": 1050
},
{
"epoch": 4.4,
"grad_norm": 1.6168186664581299,
"learning_rate": 0.00016973214285714286,
"loss": 3.5006,
"step": 1100
},
{
"epoch": 4.6,
"grad_norm": 1.6266807317733765,
"learning_rate": 0.00016526785714285713,
"loss": 3.4669,
"step": 1150
},
{
"epoch": 4.8,
"grad_norm": 1.509149432182312,
"learning_rate": 0.00016080357142857142,
"loss": 3.5044,
"step": 1200
},
{
"epoch": 5.0,
"grad_norm": 1.610813856124878,
"learning_rate": 0.00015633928571428572,
"loss": 3.4626,
"step": 1250
},
{
"epoch": 5.2,
"grad_norm": 1.3284097909927368,
"learning_rate": 0.00015187500000000002,
"loss": 3.2758,
"step": 1300
},
{
"epoch": 5.4,
"grad_norm": 1.5338636636734009,
"learning_rate": 0.0001474107142857143,
"loss": 3.3102,
"step": 1350
},
{
"epoch": 5.6,
"grad_norm": 1.4522225856781006,
"learning_rate": 0.00014294642857142856,
"loss": 3.2593,
"step": 1400
},
{
"epoch": 5.8,
"grad_norm": 1.6232526302337646,
"learning_rate": 0.00013848214285714286,
"loss": 3.2785,
"step": 1450
},
{
"epoch": 6.0,
"grad_norm": 1.6043927669525146,
"learning_rate": 0.00013401785714285713,
"loss": 3.3056,
"step": 1500
},
{
"epoch": 6.2,
"grad_norm": 1.6095658540725708,
"learning_rate": 0.00012955357142857142,
"loss": 3.1076,
"step": 1550
},
{
"epoch": 6.4,
"grad_norm": 1.4529067277908325,
"learning_rate": 0.00012508928571428572,
"loss": 3.1206,
"step": 1600
},
{
"epoch": 6.6,
"grad_norm": 1.7432512044906616,
"learning_rate": 0.000120625,
"loss": 3.1245,
"step": 1650
},
{
"epoch": 6.8,
"grad_norm": 1.7623839378356934,
"learning_rate": 0.00011616071428571429,
"loss": 3.1437,
"step": 1700
},
{
"epoch": 7.0,
"grad_norm": 1.537539005279541,
"learning_rate": 0.00011169642857142857,
"loss": 3.121,
"step": 1750
},
{
"epoch": 7.2,
"grad_norm": 1.4343411922454834,
"learning_rate": 0.00010723214285714286,
"loss": 2.9752,
"step": 1800
},
{
"epoch": 7.4,
"grad_norm": 1.4659862518310547,
"learning_rate": 0.00010276785714285715,
"loss": 2.9755,
"step": 1850
},
{
"epoch": 7.6,
"grad_norm": 1.4027365446090698,
"learning_rate": 9.830357142857144e-05,
"loss": 2.9823,
"step": 1900
},
{
"epoch": 7.8,
"grad_norm": 1.6034783124923706,
"learning_rate": 9.383928571428571e-05,
"loss": 2.9971,
"step": 1950
},
{
"epoch": 8.0,
"grad_norm": 1.9349607229232788,
"learning_rate": 8.9375e-05,
"loss": 3.0408,
"step": 2000
},
{
"epoch": 8.2,
"grad_norm": 1.5146081447601318,
"learning_rate": 8.491071428571429e-05,
"loss": 2.8761,
"step": 2050
},
{
"epoch": 8.4,
"grad_norm": 1.6080121994018555,
"learning_rate": 8.044642857142857e-05,
"loss": 2.8688,
"step": 2100
},
{
"epoch": 8.6,
"grad_norm": 1.4734200239181519,
"learning_rate": 7.598214285714286e-05,
"loss": 2.8778,
"step": 2150
},
{
"epoch": 8.8,
"grad_norm": 1.5679528713226318,
"learning_rate": 7.151785714285715e-05,
"loss": 2.9088,
"step": 2200
},
{
"epoch": 9.0,
"grad_norm": 1.4920151233673096,
"learning_rate": 6.705357142857144e-05,
"loss": 2.9162,
"step": 2250
},
{
"epoch": 9.2,
"grad_norm": 1.6481536626815796,
"learning_rate": 6.25892857142857e-05,
"loss": 2.8056,
"step": 2300
},
{
"epoch": 9.4,
"grad_norm": 1.4848685264587402,
"learning_rate": 5.8125e-05,
"loss": 2.796,
"step": 2350
},
{
"epoch": 9.6,
"grad_norm": 1.376192569732666,
"learning_rate": 5.366071428571429e-05,
"loss": 2.7904,
"step": 2400
},
{
"epoch": 9.8,
"grad_norm": 1.557061791419983,
"learning_rate": 4.919642857142857e-05,
"loss": 2.8336,
"step": 2450
},
{
"epoch": 10.0,
"grad_norm": 1.642156958580017,
"learning_rate": 4.473214285714286e-05,
"loss": 2.8048,
"step": 2500
},
{
"epoch": 10.2,
"grad_norm": 1.7578301429748535,
"learning_rate": 4.026785714285714e-05,
"loss": 2.718,
"step": 2550
},
{
"epoch": 10.4,
"grad_norm": 1.5210566520690918,
"learning_rate": 3.580357142857143e-05,
"loss": 2.733,
"step": 2600
},
{
"epoch": 10.6,
"grad_norm": 1.5723469257354736,
"learning_rate": 3.133928571428572e-05,
"loss": 2.7494,
"step": 2650
},
{
"epoch": 10.8,
"grad_norm": 1.3110719919204712,
"learning_rate": 2.6875e-05,
"loss": 2.7616,
"step": 2700
},
{
"epoch": 11.0,
"grad_norm": 1.393686056137085,
"learning_rate": 2.2410714285714286e-05,
"loss": 2.7552,
"step": 2750
},
{
"epoch": 11.2,
"grad_norm": 1.5743101835250854,
"learning_rate": 1.7946428571428573e-05,
"loss": 2.6776,
"step": 2800
},
{
"epoch": 11.4,
"grad_norm": 1.7243106365203857,
"learning_rate": 1.3482142857142857e-05,
"loss": 2.7079,
"step": 2850
},
{
"epoch": 11.6,
"grad_norm": 1.5077002048492432,
"learning_rate": 9.017857142857144e-06,
"loss": 2.6959,
"step": 2900
},
{
"epoch": 11.8,
"grad_norm": 1.3862125873565674,
"learning_rate": 4.553571428571429e-06,
"loss": 2.7095,
"step": 2950
},
{
"epoch": 12.0,
"grad_norm": 1.383197546005249,
"learning_rate": 8.928571428571429e-08,
"loss": 2.7034,
"step": 3000
}
],
"logging_steps": 50,
"max_steps": 3000,
"num_input_tokens_seen": 0,
"num_train_epochs": 12,
"save_steps": 1000,
"stateful_callbacks": {
"TrainerControl": {
"args": {
"should_epoch_stop": false,
"should_evaluate": false,
"should_log": false,
"should_save": true,
"should_training_stop": true
},
"attributes": {}
}
},
"total_flos": 2093469401088000.0,
"train_batch_size": 2,
"trial_name": null,
"trial_params": null
}

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:577aa7fa92f0f8bd6bf3e4ff2191a25d48cf0e08199280935bab389e8179cf45
size 5777

50259
checkpoint-3000/vocab.json Normal file

File diff suppressed because it is too large Load Diff

31
config.json Normal file
View File

@@ -0,0 +1,31 @@
{
"activation_function": "gelu_new",
"architectures": [
"GPT2LMHeadModel"
],
"attn_pdrop": 0.1,
"bos_token_id": 50256,
"dtype": "float32",
"embd_pdrop": 0.1,
"eos_token_id": 50256,
"initializer_range": 0.02,
"layer_norm_epsilon": 1e-05,
"model_type": "gpt2",
"n_embd": 384,
"n_head": 6,
"n_inner": null,
"n_layer": 4,
"n_positions": 2048,
"reorder_and_upcast_attn": false,
"resid_pdrop": 0.1,
"scale_attn_by_inverse_layer_idx": false,
"scale_attn_weights": true,
"summary_activation": null,
"summary_first_dropout": 0.1,
"summary_proj_to_labels": true,
"summary_type": "cls_index",
"summary_use_proj": true,
"transformers_version": "4.57.1",
"use_cache": false,
"vocab_size": 50257
}

7
generation_config.json Normal file
View File

@@ -0,0 +1,7 @@
{
"_from_model_config": true,
"bos_token_id": 50256,
"eos_token_id": 50256,
"transformers_version": "4.57.1",
"use_cache": false
}

50001
merges.txt Normal file

File diff suppressed because it is too large Load Diff

3
model.safetensors Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f9214a52ad9906206fed86ad77c5938ae7f876d235c670d94874baac2b1dbd88
size 108740128

13
model_metadata.json Normal file
View File

@@ -0,0 +1,13 @@
{
"model_name": "ANWGPT4",
"architecture": "TinyGPT2",
"developers": [
"ANW",
"FlameF0X"
],
"parameters": 27183744,
"context_window": 2048,
"pretrain_dataset": "wikitext-103-raw-v1",
"finetune_dataset": "HuggingFaceH4/ultrachat_200k (or tatsu-lab/alpaca as fallback)",
"system_prompt": "You are ANWGPT4, a small language model developed by ANW and FlameF0X, built on the TinyGPT2 architecture. You are designed to assist users with helpful, accurate, safe, and thoughtful responses across many topics. Always prioritize factual correctness and clear reasoning, be transparent about being an AI, adapt your tone appropriately, explain complex ideas when useful, avoid harmful or illegal guidance, do not claim consciousness or personal experience, and if unsure about an answer, clearly state uncertainty rather than inventing information."
}

3
pytorch_model.bin Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:33e1f730a0ed3591cb6f5957d109cd9a1f3b4728ec25c4949c06336e23bd4b64
size 108752547

24
special_tokens_map.json Normal file
View File

@@ -0,0 +1,24 @@
{
"bos_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
},
"eos_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
},
"pad_token": "<|endoftext|>",
"unk_token": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
}
}

23
tokenizer_config.json Normal file
View File

@@ -0,0 +1,23 @@
{
"add_bos_token": false,
"add_prefix_space": false,
"added_tokens_decoder": {
"50256": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false,
"special": true
}
},
"bos_token": "<|endoftext|>",
"clean_up_tokenization_spaces": false,
"eos_token": "<|endoftext|>",
"errors": "replace",
"extra_special_tokens": {},
"model_max_length": 1024,
"pad_token": "<|endoftext|>",
"tokenizer_class": "GPT2Tokenizer",
"unk_token": "<|endoftext|>"
}

50259
vocab.json Normal file

File diff suppressed because it is too large Load Diff