初始化项目,由ModelHub XC社区提供模型
Model: mpasila/shisa-v2-JP-EN-Translator-v0.1-12B Source: Original Platform
This commit is contained in:
60
README.md
Normal file
60
README.md
Normal file
@@ -0,0 +1,60 @@
|
||||
---
|
||||
base_model: shisa-ai/shisa-v2-mistral-nemo-12b
|
||||
tags:
|
||||
- text-generation-inference
|
||||
- transformers
|
||||
- unsloth
|
||||
- mistral
|
||||
- trl
|
||||
- sft
|
||||
license: apache-2.0
|
||||
language:
|
||||
- en
|
||||
datasets:
|
||||
- NilanE/ParallelFiction-Ja_En-100k
|
||||
- mpasila/ja_en_massive_1000_sharegpt_filtered_fixed_short
|
||||
---
|
||||
|
||||
This only contains around 191 examples. This is just quick test will release the full around 1k examples soon.
|
||||
|
||||
I've done a quick cleaning of the data manually using Notepad++. There may still be broken stuff or other problems.
|
||||
|
||||
Uses ChatML, recommended system prompt: `You are an AI assistant that translates Japanese to English accurately.`
|
||||
|
||||
Uses [NilanE/ParallelFiction-Ja_En-100k](https://huggingface.co/datasets/NilanE/ParallelFiction-Ja_En-100k) for the data.
|
||||
|
||||
LoRA: [mpasila/shisa-v2-JP-EN-Translator-v0.1-LoRA-12B](https://huggingface.co/mpasila/shisa-v2-JP-EN-Translator-v0.1-LoRA-12B)
|
||||
|
||||
Uses the usual 128 rank and 32 alpha. Trained on 16384 context window in QLoRA.
|
||||
|
||||
**Token Count Statistics:**
|
||||
- Total conversations: 191
|
||||
- Total tokens: 918486
|
||||
- Average tokens per conversation: 4808.83
|
||||
- Median tokens per conversation: 4187.0
|
||||
- Maximum tokens in a conversation: 13431
|
||||
- Minimum tokens in a conversation: 512
|
||||
|
||||
**Token Distribution by Role:**
|
||||
- System messages: 2483 tokens (0.27%)
|
||||
- Human messages: 494038 tokens (53.79%)
|
||||
- Assistant messages: 421965 tokens (45.94%)
|
||||
|
||||
**Token Count Distribution:**
|
||||
- 0-512: 0 conversations (0.00%)
|
||||
- 513-1024: 4 conversations (2.09%)
|
||||
- 1025-2048: 10 conversations (5.24%)
|
||||
- 2049-4096: 77 conversations (40.31%)
|
||||
- 4097-8192: 83 conversations (43.46%)
|
||||
- 8193-16384: 17 conversations (8.90%)
|
||||
- 16385+: 0 conversations (0.00%)
|
||||
|
||||
# Uploaded shisa-v2-JP-EN-Translator-v0.1-12B model
|
||||
|
||||
- **Developed by:** mpasila
|
||||
- **License:** apache-2.0
|
||||
- **Finetuned from model :** shisa-ai/shisa-v2-mistral-nemo-12b
|
||||
|
||||
This mistral model was trained 2x faster with [Unsloth](https://github.com/unslothai/unsloth) and Huggingface's TRL library.
|
||||
|
||||
[<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>](https://github.com/unslothai/unsloth)
|
||||
Reference in New Issue
Block a user