初始化项目,由ModelHub XC社区提供模型
Model: scb10x/llama3.1-typhoon2-8b Source: Original Platform
This commit is contained in:
56
README.md
Normal file
56
README.md
Normal file
@@ -0,0 +1,56 @@
|
||||
---
|
||||
license: llama3.1
|
||||
pipeline_tag: text-generation
|
||||
---
|
||||
|
||||
**Llama3.1-Typhoon2-8B**: Thai Large Language Model (Instruct)
|
||||
|
||||
**Llama3.1-Typhoon2-8B** is a pretrained only Thai 🇹🇭 large language model with 8 billion parameters, and it is based on Llama3.1-8B.
|
||||
|
||||
|
||||
For technical-report. please see our [arxiv](https://arxiv.org/abs/2412.13702).
|
||||
*To acknowledge Meta's effort in creating the foundation model and to comply with the license, we explicitly include "llama-3.1" in the model name.
|
||||
|
||||
## **Performance**
|
||||
|
||||
| Model | ThaiExam | ONET | IC | A-Level | TGAT | TPAT | M3Exam | Math | Science | Social | Thai |
|
||||
|------------------------|----------|--------|-----------|-----------|-----------|-----------|-----------|------------|------------|------------|------------|
|
||||
| **Typhoon2 Llama3.1 8B Base**| **51.20%** | **49.38%** | 47.36% | **43.30%** | 67.69% | 48.27% | **47.52%** | 27.60% | **44.20%** | **68.90%** | **49.38%** |
|
||||
| **Llama3.1 8B** | 45.80% | 38.27% | **46.31%** | 34.64% | 61.53% | 48.27% | 43.33% | **27.14%** | 40.82% | 58.33% | 47.05% |
|
||||
| **Typhoon1.5 Llama3 8B Base** | 48.82% | 41.35% | 41.05% | 40.94% | **70.76%** | **50.00%** | 43.88% | 22.62% | 43.47% | 62.81% | 46.63% |
|
||||
|
||||
|
||||
## **Model Description**
|
||||
|
||||
- **Model type**: A 8B decoder-only model based on Llama architecture.
|
||||
- **Requirement**: transformers 4.45.0 or newer.
|
||||
- **Primary Language(s)**: Thai 🇹🇭 and English 🇬🇧
|
||||
- **License**: [Llama 3.1 Community License](https://github.com/meta-llama/llama-models/blob/main/models/llama3_1/LICENSE)
|
||||
|
||||
|
||||
## **Intended Uses & Limitations**
|
||||
|
||||
This model is a pretrained base model. Thus, it may not be able to follow human instructions without using one/few-shot learning or instruction fine-tuning. The model does not have any moderation mechanisms, and may generate harmful or inappropriate responses.
|
||||
|
||||
## **Follow us**
|
||||
|
||||
**https://twitter.com/opentyphoon**
|
||||
|
||||
## **Support**
|
||||
|
||||
**https://discord.gg/us5gAYmrxw**
|
||||
|
||||
## **Citation**
|
||||
|
||||
- If you find Typhoon2 useful for your work, please cite it using:
|
||||
```
|
||||
@misc{typhoon2,
|
||||
title={Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models},
|
||||
author={Kunat Pipatanakul and Potsawee Manakul and Natapong Nitarach and Warit Sirichotedumrong and Surapon Nonesung and Teetouch Jaknamon and Parinthapat Pengpun and Pittawat Taveekitworachai and Adisai Na-Thalang and Sittipong Sripaisarnmongkol and Krisanapong Jirayoot and Kasima Tharnpipitchai},
|
||||
year={2024},
|
||||
eprint={2412.13702},
|
||||
archivePrefix={arXiv},
|
||||
primaryClass={cs.CL},
|
||||
url={https://arxiv.org/abs/2412.13702},
|
||||
}
|
||||
```
|
||||
Reference in New Issue
Block a user