初始化项目,由ModelHub XC社区提供模型
Model: openthaigpt/openthaigpt-1.0.0-13b-chat Source: Original Platform
This commit is contained in:
36
.gitattributes
vendored
Normal file
36
.gitattributes
vendored
Normal file
@@ -0,0 +1,36 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
ggml-model-f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
283
README.md
Normal file
283
README.md
Normal file
@@ -0,0 +1,283 @@
|
||||
---
|
||||
license: llama2
|
||||
language:
|
||||
- th
|
||||
- en
|
||||
library_name: transformers
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- openthaigpt
|
||||
- llama
|
||||
---
|
||||
|
||||
# 🇹🇭 OpenThaiGPT 13b 1.0.0
|
||||

|
||||
[More Info](https://openthaigpt.aieat.or.th/)
|
||||
|
||||
🇹🇭 **OpenThaiGPT 13b Version 1.0.0** is an advanced 13-billion-parameter Thai language chat model based on LLaMA v2 released on April 8, 2024. It has been specifically fine-tuned for Thai instructions and enhanced by incorporating over 10,000 of the most commonly used Thai words into the large language model's (LLM) dictionary, significantly boosting its response speed.
|
||||
|
||||
## Highlights
|
||||
- **Leading-edge Thai language LLM**, setting new benchmarks by achieving the highest average scores across several Thai language exams when compared to all other open-source Thai LLMs.
|
||||
- **The First 70b Thai opensource LLM**, achieving the higher score on Thai exams than OpenAI GPT 3.5, Google Gemini, and Claude 3 Haiku.
|
||||
- **Support for extended conversations** across multiple turns.
|
||||
- Support the use case of **Retrieval Augmented Generation (RAG)** for enriched response generation.
|
||||
- **Generation speeds increased by tenfold**, thanks to the addition of 10,000 frequently used Thai words to the model's dictionary.
|
||||
- Pretrained upon a foundation of **more than 65 billion Thai language words** and meticulously fine-tuned with over 1 million Thai instruction examples.
|
||||
- Capable of understanding and processing **input contexts of up to 4096 Thai words**, allowing for detailed and complex instructions.
|
||||
|
||||
## Benchmark by OpenThaiGPT Eval
|
||||
** Please take a look at ``OTG 13b (April 2024)`` for this model's evaluation result.
|
||||
|
||||
| **Exams** | **OTG 7b (Aug 2023)** | **OTG 13b (Dec 2023)** | **OTG 7b (April 2024)** | <b style="color:blue">OTG 13b (April 2024)</b> | **OTG 70b (April 2024)** | **SeaLLM 7b v1** | **SeaLLM 7b v2** | **SeaLion 7b** | **WanchanGLM 7b** | **Sailor-7b-Chat** | **TyphoonGPT 7b Instruct** | **GPT3.5** | **GPT4** | **Gemini Pro** | **Gemini 1.5** | **Claude 3 Haiku** | **Claude 3 Sonnet** | **Claude 3 Opus** |
|
||||
|----------------------------|-----------------------|------------------------|-------------------------|--------------------------|--------------------------|------------------|------------------|----------------|-------------------|--------------------|----------------------------|------------|----------|----------------|----------------|--------------------|---------------------|-------------------|
|
||||
| **A-Level** | 17.50% | 34.17% | 25.00% | <b style="color:blue">30.83%</b> | 45.83% | 18.33% | 34.17% | 21.67% | 17.50% | 40.00% | 37.50% | 38.33% | 65.83% | 56.67% | 55.83% | 58.33% | 59.17% | 77.50% |
|
||||
| **TGAT** | 24.00% | 22.00% | 22.00% | <b style="color:blue">36.00%</b> | 36.00% | 14.00% | 28.00% | 24.00% | 16.00% | 34.00% | 30.00% | 28.00% | 44.00% | 22.00% | 28.00% | 36.00% | 34.00% | 46.00% |
|
||||
| **TPAT1** | 22.50% | 47.50% | 42.50% | <b style="color:blue">27.50%</b> | 62.50% | 22.50% | 27.50% | 22.50% | 17.50% | 40.00% | 47.50% | 45.00% | 52.50% | 52.50% | 50.00% | 52.50% | 50.00% | 62.50% |
|
||||
| **thai_investment_consultant_exams** | 8.00% | 28.00% | 76.00% | <b style="color:blue">84.00%</b> | 68.00% | 16.00% | 28.00% | 24.00% | 16.00% | 24.00% | 32.00% | 40.00% | 64.00% | 52.00% | 32.00% | 44.00% | 64.00% | 72.00% |
|
||||
| **facebook_beleble_tha_200** | 25.00% | 45.00% | 34.50% | <b style="color:blue">39.50%</b> | 70.00% | 13.50% | 51.00% | 27.00% | 24.50% | 63.00% | 51.50% | 50.00% | 72.50% | 65.00% | 74.00% | 63.50% | 77.00% | 90.00% |
|
||||
| **xcopa_th_200** | 45.00% | 56.50% | 49.50% | <b style="color:blue">51.50%</b> | 74.50% | 26.50% | 47.00% | 51.50% | 48.50% | 68.50% | 65.00% | 64.00% | 82.00% | 68.00% | 74.00% | 64.00% | 80.00% | 86.00% |
|
||||
| **xnli2.0_th_200** | 33.50% | 34.50% | 39.50% | <b style="color:blue">31.00%</b> | 47.00% | 21.00% | 43.00% | 37.50% | 33.50% | 16.00% | 20.00% | 50.00% | 69.00% | 53.00% | 54.50% | 50.00% | 68.00% | 68.50% |
|
||||
| **ONET M3** | 17.85% | 38.86% | 34.11% | <b style="color:blue">39.36%</b> | 56.15% | 15.58% | 23.92% | 21.79% | 19.56% | 21.37% | 28.03% | 37.91% | 49.97% | 55.99% | 57.41% | 52.73% | 40.60% | 63.87% |
|
||||
| **ONET M6** | 21.14% | 28.87% | 22.53% | <b style="color:blue">23.32%</b> | 42.85% | 15.09% | 19.48% | 16.96% | 20.67% | 28.64% | 27.46% | 34.44% | 46.29% | 45.53% | 50.23% | 34.79% | 38.49% | 48.56% |
|
||||
| **AVERAGE SCORE** | 23.83% | 37.27% | 38.40% | <b style="color:blue;font-size:1.3em">40.33%</b> | 55.87% | 18.06% | 33.56% | 27.44% | 23.75% | 37.28% | 37.67% | 43.07% | 60.68% | 52.30% | 52.89% | 50.65% | 56.81% | 68.32% |
|
||||
Thai language multiple choice exams, Test on unseen test set, Zero-shot learning. Benchmark source code and exams information: https://github.com/OpenThaiGPT/openthaigpt_eval
|
||||
|
||||
(Updated on: 7 April 2024)
|
||||
|
||||
## Benchmark on M3Exam evaluated by an external party (Float16.cloud)
|
||||
|
||||
| **Models** | **ENGLISH (M3EXAM)** | **THAI (M3EXAM)** |
|
||||
|---------------------|------------------|---------------|
|
||||
| OTG-7b | 40.92 % | 25.14 % |
|
||||
| <b style="color:blue">OTG-13b</b> | <b style="color:blue">53.69 %</b> | <b style="color:blue">36.49 %</b> |
|
||||
| OTG-70b | 72.58 % | 48.29 % |
|
||||
| GPT-3.5-turbo-0613* | - | 34.1 % |
|
||||
| GPT-4-0613* | - | 56.0 % |
|
||||
More information: https://blog.float16.cloud/the-first-70b-thai-llm/
|
||||
|
||||
## Licenses
|
||||
**Source Code**: License Apache Software License 2.0.<br>
|
||||
**Weight**: Research and **Commercial uses**.<br>
|
||||
|
||||
## Sponsors
|
||||
<img src="https://cdn-uploads.huggingface.co/production/uploads/5fcd9c426d942eaf4d1ebd30/FDC9WYN2iykQbVW1rY4q5.png" width="600px">
|
||||
|
||||
## Supports
|
||||
- Official website: https://openthaigpt.aieat.or.th
|
||||
- Facebook page: https://web.facebook.com/groups/openthaigpt
|
||||
- A Discord server for discussion and support [here](https://discord.gg/rUTp6dfVUF)
|
||||
- E-mail: kobkrit@aieat.or.th
|
||||
|
||||
## Prompt Format
|
||||
Prompt format is based on Llama2 with a small modification (Adding "###" to specify the context part)
|
||||
```
|
||||
<s>[INST] <<SYS>
|
||||
{system_prompt}
|
||||
<</SYS>>
|
||||
|
||||
{human_turn1}###{context_turn1} [/INST]{assistant_turn1}</s><s>{human_turn2}###{context_turn2} [/INST] ...
|
||||
```
|
||||
|
||||
### System prompt:
|
||||
```
|
||||
You are a question answering assistant. Answer the question as truthful and helpful as possible คุณคือผู้ช่วยตอบคำถาม จงตอบคำถามอย่างถูกต้องและมีประโยชน์ที่สุด
|
||||
```
|
||||
|
||||
### Examples
|
||||
|
||||
#### Single Turn Conversation Example
|
||||
```
|
||||
<s>[INST] <<SYS>
|
||||
You are a question answering assistant. Answer the question as truthful and helpful as possible คุณคือผู้ช่วยตอบคำถาม จงตอบคำถามอย่างถูกต้องและมีประโยชน์ที่สุด
|
||||
<</SYS>>
|
||||
|
||||
สวัสดีครับ [/INST]
|
||||
```
|
||||
|
||||
#### Single Turn Conversation with Context (RAG) Example
|
||||
```
|
||||
<s>[INST] <<SYS>
|
||||
You are a question answering assistant. Answer the question as truthful and helpful as possible คุณคือผู้ช่วยตอบคำถาม จงตอบคำถามอย่างถูกต้องและมีประโยชน์ที่สุด
|
||||
<</SYS>>
|
||||
|
||||
กรุงเทพมีพื้นที่เท่าไร่###กรุงเทพมหานคร เป็นเมืองหลวง นครและมหานครที่มีประชากรมากที่สุดของประเทศไทย กรุงเทพมหานครมีพื้นที่ทั้งหมด 1,568.737 ตร.กม. มีประชากรตามทะเบียนราษฎรกว่า 8 ล้านคน [/INST]
|
||||
```
|
||||
|
||||
#### Multi Turn Conversation Example
|
||||
|
||||
##### First turn
|
||||
```
|
||||
<s>[INST] <<SYS>
|
||||
You are a question answering assistant. Answer the question as truthful and helpful as possible คุณคือผู้ช่วยตอบคำถาม จงตอบคำถามอย่างถูกต้องและมีประโยชน์ที่สุด
|
||||
<</SYS>>
|
||||
|
||||
สวัสดีครับ [/INST]
|
||||
```
|
||||
|
||||
##### Second turn
|
||||
```
|
||||
<s>[INST] <<SYS>
|
||||
You are a question answering assistant. Answer the question as truthful and helpful as possible คุณคือผู้ช่วยตอบคำถาม จงตอบคำถามอย่างถูกต้องและมีประโยชน์ที่สุด
|
||||
<</SYS>>
|
||||
|
||||
สวัสดีครับ [/INST]สวัสดีค่ะ มีคำถามอะไร ถามได้เลย</s><s>ขอสูตรทำส้มตำหน่อย [/INST]
|
||||
```
|
||||
|
||||
##### Third turn
|
||||
```
|
||||
<s>[INST] <<SYS>
|
||||
You are a question answering assistant. Answer the question as truthful and helpful as possible คุณคือผู้ช่วยตอบคำถาม จงตอบคำถามอย่างถูกต้องและมีประโยชน์ที่สุด
|
||||
<</SYS>>
|
||||
|
||||
สวัสดีครับ [/INST]สวัสดีค่ะ มีคำถามอะไร ถามได้เลย</s><s>ขอสูตรทำส้มตำหน่อย [/INST]ได้เลยค่ะ ส้มตำเป็นเมนูที่ทำง่ายและอร่อย มาเริ่มกันเลยนะคะ</s><s>เริ่มได้เลย [/INST]
|
||||
```
|
||||
|
||||
##### Fourth turn
|
||||
```
|
||||
<s>[INST] <<SYS>
|
||||
You are a question answering assistant. Answer the question as truthful and helpful as possible คุณคือผู้ช่วยตอบคำถาม จงตอบคำถามอย่างถูกต้องและมีประโยชน์ที่สุด
|
||||
<</SYS>>
|
||||
|
||||
สวัสดีครับ [/INST]สวัสดีค่ะ มีคำถามอะไร ถามได้เลย</s><s>ขอสูตรทำส้มตำหน่อย [/INST]ได้เลยค่ะ ส้มตำเป็นเมนูที่ทำง่ายและอร่อย มาเริ่มกันเลยนะคะ</s><s>เริ่มได้เลย [/INST]
|
||||
ตำไทยเป็นเมนูที่ใครๆ ก็คุ้นเคย ด้วยรสชาติที่กลมกล่อม หวาน เปรี้ยว เค็ม เผ็ด ครบรส วันนี้เราจะมาทำส้มตำด้วยกันค่ะ สิ่งแรกที่ต้องมีคือเครื่องปรุงดังนี้
|
||||
- มะละกอดิบ 1 ลูก
|
||||
- กุ้งแห้ง 1/2 ถ้วย
|
||||
- กระเทียม 3 กลีบ
|
||||
- พริกขี้หนู 3 เม็ด
|
||||
- น้ำตาลปี๊บ 1 ช้อนโต๊ะ
|
||||
- น้ำปลา 2 ช้อนโต๊ะ
|
||||
- มะนาว 1 ลูก
|
||||
- งาขาวคั่ว 1/4 ถ้วย
|
||||
|
||||
วิธีทำมีดังนี้ค่ะ
|
||||
1. เริ่มจากล้างมะละกอให้สะอาด แล้วใช้มีดปอกเปลือก เอาไส้ออก หั่นเป็นเส้นบางๆ เตรียมไว้
|
||||
2. ตำกุ้งแห้งให้ละเอียด แล้วตักขึ้นพักไว้
|
||||
3. ใช้ครกหินหรือเครื่องปั่น บดพริกขี้หนูกับกระเทียมให้ละเอียด
|
||||
4. ใส่กุ้งแห้งที่ตำแล้วลงไปผสม ตามด้วยน้ำตาลปี๊บ น้ำปลา มะนาว และเส้นมะละกอ คลุกเคล้าให้เข้ากัน
|
||||
5. ชิมรสให้ได้รสหวาน เปรี้ยว เค็ม เผ็ด จากนั้นตักขึ้นเสิร์ฟพร้อมผักสด อาทิ กะหล่ำปลี ถั่วงอก แครอท ผักบุ้ง</s><s>ขอบคุณครับ [/INST]
|
||||
```
|
||||
|
||||
#### Multi Turn Conversation with Context (RAG) Example
|
||||
```
|
||||
<s>[INST] <<SYS>
|
||||
You are a question answering assistant. Answer the question as truthful and helpful as possible คุณคือผู้ช่วยตอบคำถาม จงตอบคำถามอย่างถูกต้องและมีประโยชน์ที่สุด
|
||||
<</SYS>>
|
||||
|
||||
กรุงเทพมีพื้นที่เท่าไร่###กรุงเทพมหานคร เป็นเมืองหลวง นครและมหานครที่มีประชากรมากที่สุดของประเทศไทย กรุงเทพมหานครมีพื้นที่ทั้งหมด 1,568.737 ตร.กม. มีประชากรตามทะเบียนราษฎรกว่า 8 ล้านคน [/INST]
|
||||
กรุงเทพมหานครมีพื้นที่ทั้งหมด 1,568.737 ตร.กม.</s><s>และประชากรล่ะ [/INST]
|
||||
```
|
||||
|
||||
## How to use
|
||||
|
||||
### Huggingface
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
import torch
|
||||
|
||||
# Ensure CUDA is available
|
||||
device = 'cuda' if torch.cuda.is_available() else 'cpu'
|
||||
print(f"Using device: {device}")
|
||||
|
||||
# Init Model
|
||||
model_path="openthaigpt/openthaigpt-1.0.0-7b-chat"
|
||||
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
|
||||
model = AutoModelForCausalLM.from_pretrained(model_path, trust_remote_code=True, torch_dtype=torch.float16)
|
||||
model.to(device)
|
||||
|
||||
# Prompt
|
||||
prompt = "สวัสดีครับ OpenThaiGPT"
|
||||
llama_prompt = f"<s>[INST] <<SYS>>\nYou are a question answering assistant. Answer the question as truthful and helpful as possible คุณคือผู้ช่วยตอบคำถาม จงตอบคำถามอย่างถูกต้องและมีประโยชน์ที่สุด<</SYS>>\n\n{prompt} [/INST]"
|
||||
inputs = tokenizer.encode(llama_prompt, return_tensors="pt")
|
||||
inputs = inputs.to(device)
|
||||
|
||||
# Generate
|
||||
outputs = model.generate(inputs, max_length=512, num_return_sequences=1)
|
||||
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
||||
```
|
||||
|
||||
### vLLM
|
||||
|
||||
1. Install VLLM (https://github.com/vllm-project/vllm)
|
||||
|
||||
2. Run server
|
||||
```bash
|
||||
python -m vllm.entrypoints.api_server --model /path/to/model --tensor-parallel-size num_gpus
|
||||
```
|
||||
3. Run inference (CURL example)
|
||||
```bash
|
||||
curl --request POST \
|
||||
--url http://localhost:8000/generate \
|
||||
--header "Content-Type: application/json" \
|
||||
--data '{"prompt": "<s>[INST] <<SYS>>\nYou are a question answering assistant. Answer the question as truthful and helpful as possible คุณคือผู้ช่วยตอบคำถาม จงตอบคำถามอย่างถูกต้องและมีประโยชน์ที่สุด\n<</SYS>>\n\nอยากลดความอ้วนต้องทำอย่างไร [/INST]","use_beam_search": false, "temperature": 0.1, "max_tokens": 512, "top_p": 0.75, "top_k": 40, "frequency_penalty": 0.3 "stop": "</s>"}'
|
||||
```
|
||||
|
||||
### LlamaCPP (for GGUF)
|
||||
|
||||
1. Build and Install LlamaCPP (LLAMA_CUBLAS=1 is for GPU inference)
|
||||
```bash
|
||||
git clone https://github.com/ggerganov/llama.cpp.git \
|
||||
&& cd llama.cpp \
|
||||
&& make -j LLAMA_CUBLAS=1 CUDA_DOCKER_ARCH=all
|
||||
```
|
||||
|
||||
2. Run server
|
||||
```bash
|
||||
./server -m /path/to/ggml-model-f16.gguf -c 3072 -ngl 81 -ts 1,1 --host 0.0.0.0
|
||||
```
|
||||
|
||||
3. Run inference (CURL example)
|
||||
```bash
|
||||
curl --location 'http://localhost:8000/completion' \
|
||||
--header 'Content-Type: application/json' \
|
||||
--data '{
|
||||
"prompt":"<s>[INST] <<SYS>>\nYou are a question answering assistant. Answer the question as truthful and helpful as possible คุณคือผู้ช่วยตอบคำถาม จงตอบคำถามอย่างถูกต้องและมีประโยชน์ที่สุด friendly\n\n<<SYS>>\n\nอยากลดความอ้วนต้องทำอย่างไร [/INST]",
|
||||
"max_tokens": 512,
|
||||
"stop":"</s>"
|
||||
}'
|
||||
```
|
||||
|
||||
### GPU Memory Requirements
|
||||
| **Number of Parameters** | **FP 16 bits** | **8 bits (Quantized)** | **4 bits (Quantized)** | **Example Graphic Card for 4 bits** |
|
||||
|------------------|----------------|------------------------|------------------------|---------------------------------------------|
|
||||
| **7b** | 24 GB | 12 GB | 6 GB | Nvidia RTX 4060 8GB |
|
||||
| **13b** | 48 GB | 24 GB | 12 GB | Nvidia RTX 4070 16GB |
|
||||
| **70b** | 192 GB | 96 GB | 48 GB | Nvidia RTX 4090 24GB x 2 cards |
|
||||
|
||||
### OpenThaiGPT Team
|
||||
* Kobkrit Viriyayudhakorn (kobkrit@aieat.or.th)
|
||||
* Sumeth Yuenyong (sumeth.yue@mahidol.edu)
|
||||
* Thaweewat Rugsujarit (thaweewr@scg.com)
|
||||
* Jillaphat Jaroenkantasima (autsadang41@gmail.com)
|
||||
* Norapat Buppodom (new@norapat.com)
|
||||
* Koravich Sangkaew (kwankoravich@gmail.com)
|
||||
* Peerawat Rojratchadakorn (peerawat.roj@gmail.com)
|
||||
* Surapon Nonesung (nonesungsurapon@gmail.com)
|
||||
* Chanon Utupon (chanon.utupon@gmail.com)
|
||||
* Sadhis Wongprayoon (sadhis.tae@gmail.com)
|
||||
* Nucharee Thongthungwong (nuchhub@hotmail.com)
|
||||
* Chawakorn Phiantham (mondcha1507@gmail.com)
|
||||
* Patteera Triamamornwooth (patt.patteera@gmail.com)
|
||||
* Nattarika Juntarapaoraya (natt.juntara@gmail.com)
|
||||
* Kriangkrai Saetan (kraitan.ss21@gmail.com)
|
||||
* Pitikorn Khlaisamniang (pitikorn32@gmail.com)
|
||||
|
||||
### Citation
|
||||
If OpenThaiGPT has been beneficial for your work, kindly consider citing it as follows:
|
||||
|
||||
#### Bibtex
|
||||
```bibtex
|
||||
@misc{yuenyong2024openthaigpt15thaicentricopen,
|
||||
title={OpenThaiGPT 1.5: A Thai-Centric Open Source Large Language Model},
|
||||
author={Sumeth Yuenyong and Kobkrit Viriyayudhakorn and Apivadee Piyatumrong and Jillaphat Jaroenkantasima},
|
||||
year={2024},
|
||||
eprint={2411.07238},
|
||||
archivePrefix={arXiv},
|
||||
primaryClass={cs.CL},
|
||||
url={https://arxiv.org/abs/2411.07238},
|
||||
}
|
||||
```
|
||||
#### APA Style (for TXT, MS Word)
|
||||
```
|
||||
Yuenyong, S., Viriyayudhakorn, K., Piyatumrong, A., & Jaroenkantasima, J. (2024). OpenThaiGPT 1.5: A Thai-Centric Open Source Large Language Model. arXiv [Cs.CL]. Retrieved from http://arxiv.org/abs/2411.07238
|
||||
```
|
||||
<i>Disclaimer: Provided responses are not guaranteed.</i>
|
||||
23
added_tokens.json
Normal file
23
added_tokens.json
Normal file
@@ -0,0 +1,23 @@
|
||||
{
|
||||
"</s>": 2,
|
||||
"<CLS>": 41070,
|
||||
"<EOD>": 41072,
|
||||
"<MASK>": 41073,
|
||||
"<PAD>": 41074,
|
||||
"<SEP>": 41071,
|
||||
"<s>": 1,
|
||||
"<unk>": 0,
|
||||
"<unused1>":41075,
|
||||
"<unused2>":41076,
|
||||
"<unused3>":41077,
|
||||
"<unused4>":41078,
|
||||
"<unused5>":41079,
|
||||
"<unused6>":41080,
|
||||
"<unused7>":41081,
|
||||
"<unused8>":41082,
|
||||
"<unused9>":41083,
|
||||
"<unused10>":41084,
|
||||
"<unused11>":41085,
|
||||
"<unused12>":41086,
|
||||
"<unused13>":41087
|
||||
}
|
||||
27
config.json
Normal file
27
config.json
Normal file
@@ -0,0 +1,27 @@
|
||||
{
|
||||
"_name_or_path": "/models/llama2-13b-finetune-hf",
|
||||
"architectures": [
|
||||
"LlamaForCausalLM"
|
||||
],
|
||||
"attention_bias": false,
|
||||
"bos_token_id": 1,
|
||||
"eos_token_id": 2,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 5120,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 13824,
|
||||
"max_position_embeddings": 4096,
|
||||
"model_type": "llama",
|
||||
"num_attention_heads": 40,
|
||||
"num_hidden_layers": 40,
|
||||
"num_key_value_heads": 40,
|
||||
"pretraining_tp": 1,
|
||||
"rms_norm_eps": 1e-05,
|
||||
"rope_scaling": null,
|
||||
"rope_theta": 10000.0,
|
||||
"tie_word_embeddings": false,
|
||||
"torch_dtype": "bfloat16",
|
||||
"transformers_version": "4.35.0",
|
||||
"use_cache": false,
|
||||
"vocab_size": 41088
|
||||
}
|
||||
6
generation_config.json
Normal file
6
generation_config.json
Normal file
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"_from_model_config": true,
|
||||
"bos_token_id": 1,
|
||||
"eos_token_id": 2,
|
||||
"transformers_version": "4.35.0"
|
||||
}
|
||||
3
ggml-model-f16.gguf
Normal file
3
ggml-model-f16.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:631750452d9755cdf6d8519db1a9bd1900cdd570bae4e3bc92c70c67bc6d0b32
|
||||
size 26219707520
|
||||
3
model-00001-of-00006.safetensors
Normal file
3
model-00001-of-00006.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:782d23545d58b147789d32d406cf64264138ee0001b6d5b95c232fdc97343a77
|
||||
size 4966469088
|
||||
3
model-00002-of-00006.safetensors
Normal file
3
model-00002-of-00006.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:a9a3ce82b3ae8ed51391cfac16df55dd3fd5686bc7adc047fc36457e40452fb7
|
||||
size 4970422232
|
||||
3
model-00003-of-00006.safetensors
Normal file
3
model-00003-of-00006.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:8847ad93cb79ceb7d7958c75365d5b0d4a9bde7bdb54183d7c71d89cd9e6b9b7
|
||||
size 4933701504
|
||||
3
model-00004-of-00006.safetensors
Normal file
3
model-00004-of-00006.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:aa3292bb8f07e6899ec562f581cf05a44711d1f5546e2bb46f7db2ebea1fb87c
|
||||
size 4933722216
|
||||
3
model-00005-of-00006.safetensors
Normal file
3
model-00005-of-00006.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:2376e7d00118e601064c5f4b32ffd768ba22883139c267bb9ddde521a31785a0
|
||||
size 4933722208
|
||||
3
model-00006-of-00006.safetensors
Normal file
3
model-00006-of-00006.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:2bfc7f0a56f1dc6194b3ed3096125508ee35f34446b872f472786ab87e487bc0
|
||||
size 1479855912
|
||||
370
model.safetensors.index.json
Normal file
370
model.safetensors.index.json
Normal file
@@ -0,0 +1,370 @@
|
||||
{
|
||||
"metadata": {
|
||||
"total_size": 26217850880
|
||||
},
|
||||
"weight_map": {
|
||||
"lm_head.weight": "model-00006-of-00006.safetensors",
|
||||
"model.embed_tokens.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.input_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.0.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.input_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.1.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.10.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.10.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.10.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.10.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.10.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.10.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.10.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.10.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.10.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.11.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.11.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.11.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.11.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.11.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.11.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.11.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.11.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.11.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.12.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.12.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.12.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.12.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.12.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.12.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.12.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.12.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.12.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.13.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.13.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.13.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.13.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.13.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.13.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.13.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.13.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.13.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.14.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.14.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.14.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.14.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.14.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.14.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.14.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.14.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.14.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.15.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.15.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.16.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.17.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.18.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.18.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.18.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.18.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.18.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.18.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.18.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.18.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.18.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.19.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.19.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.19.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.19.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.19.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.19.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.19.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.19.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.19.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.2.input_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.2.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.20.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.20.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.20.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.20.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.20.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.20.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.20.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.20.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.20.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.21.input_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.21.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.21.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.21.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.21.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.21.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.21.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.21.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.21.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.22.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.22.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.22.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.22.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.22.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.22.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.22.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
|
||||
"model.layers.23.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.23.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.24.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.25.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.25.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.25.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.25.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.25.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.25.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.25.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.25.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.25.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.26.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.26.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.26.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.26.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.26.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.26.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.26.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.26.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.26.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.27.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.27.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.27.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.27.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.27.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.27.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.27.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.27.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.27.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.28.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.28.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.28.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.28.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.28.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.28.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.28.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.28.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.28.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.29.input_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.29.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.29.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.29.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.29.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.29.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.29.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.29.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.29.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.3.input_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.3.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.3.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.3.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.3.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.3.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.3.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.3.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.3.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.30.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.30.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.30.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.30.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.30.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.30.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
|
||||
"model.layers.31.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.31.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.32.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.33.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.33.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.33.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.33.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.33.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.33.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.33.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.33.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.33.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.34.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.34.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.34.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.34.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.34.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.34.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.34.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.34.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.34.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.35.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.35.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.35.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.35.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.35.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.35.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.35.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.35.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.35.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.36.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.36.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.36.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.36.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.36.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.36.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.36.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.36.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.36.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.37.input_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.37.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.37.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.37.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.37.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.37.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.37.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.37.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.37.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.38.input_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.38.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.38.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.38.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.38.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
|
||||
"model.layers.39.input_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.39.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
|
||||
"model.layers.4.input_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.4.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.4.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.4.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.4.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.4.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.4.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.4.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.4.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.5.input_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.5.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.5.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.5.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.5.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.5.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.5.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.5.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.5.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.6.input_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.6.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.6.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.6.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.6.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.6.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.6.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.6.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.6.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.7.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.7.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.7.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
|
||||
"model.layers.7.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.8.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.input_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.layers.9.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
|
||||
"model.norm.weight": "model-00006-of-00006.safetensors"
|
||||
}
|
||||
}
|
||||
14
special_tokens_map.json
Normal file
14
special_tokens_map.json
Normal file
@@ -0,0 +1,14 @@
|
||||
{
|
||||
"additional_special_tokens": [
|
||||
"<unk>",
|
||||
"<s>",
|
||||
"</s>"
|
||||
],
|
||||
"bos_token": "<s>",
|
||||
"cls_token": "<CLS>",
|
||||
"eos_token": "</s>",
|
||||
"mask_token": "<MASK>",
|
||||
"pad_token": "<PAD>",
|
||||
"sep_token": "<SEP>",
|
||||
"unk_token": "<unk>"
|
||||
}
|
||||
115387
tokenizer.json
Normal file
115387
tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
3
tokenizer.model
Normal file
3
tokenizer.model
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:02df43dcae8c7b5b122d45f642e42c96577cdd09fd949c6996051886c72ab002
|
||||
size 717508
|
||||
86
tokenizer_config.json
Normal file
86
tokenizer_config.json
Normal file
@@ -0,0 +1,86 @@
|
||||
{
|
||||
"add_bos_token": true,
|
||||
"add_eos_token": false,
|
||||
"added_tokens_decoder": {
|
||||
"0": {
|
||||
"content": "<unk>",
|
||||
"lstrip": false,
|
||||
"normalized": true,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"1": {
|
||||
"content": "<s>",
|
||||
"lstrip": false,
|
||||
"normalized": true,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"2": {
|
||||
"content": "</s>",
|
||||
"lstrip": false,
|
||||
"normalized": true,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"41070": {
|
||||
"content": "<CLS>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"41071": {
|
||||
"content": "<SEP>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"41072": {
|
||||
"content": "<EOD>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"41073": {
|
||||
"content": "<MASK>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"41074": {
|
||||
"content": "<PAD>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
}
|
||||
},
|
||||
"additional_special_tokens": [
|
||||
"<unk>",
|
||||
"<s>",
|
||||
"</s>"
|
||||
],
|
||||
"bos_token": "<s>",
|
||||
"clean_up_tokenization_spaces": false,
|
||||
"eos_token": "</s>",
|
||||
"legacy": true,
|
||||
"model_max_length": 1000000000000000019884624838656,
|
||||
"pad_token": "<PAD>",
|
||||
"sp_model_kwargs": {},
|
||||
"spaces_between_special_tokens": false,
|
||||
"tokenizer_class": "LlamaTokenizer",
|
||||
"unk_token": "<unk>",
|
||||
"use_default_system_prompt": true
|
||||
}
|
||||
Reference in New Issue
Block a user