初始化项目,由ModelHub XC社区提供模型
Model: iamtatsuki05/Llama-JP-0.5B-PT-stage1 Source: Original Platform
This commit is contained in:
99
README.md
Normal file
99
README.md
Normal file
@@ -0,0 +1,99 @@
|
||||
---
|
||||
language:
|
||||
- ja
|
||||
license: mit
|
||||
library_name: transformers
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- mirei
|
||||
- llama
|
||||
- pretraining
|
||||
base_model: iamtatsuki05/Llama-JP-0.5B-init
|
||||
datasets:
|
||||
- hotchpotch/fineweb-2-edu-japanese
|
||||
---
|
||||
|
||||
# Llama-JP-0.5B-PT-stage1
|
||||
|
||||
English / [Japanese](README_JA.md)
|
||||
|
||||
|
||||
## Overview
|
||||
Llama-JP-0.5B-PT-stage1 continues from [iamtatsuki05/Llama-JP-0.5B-init](https://huggingface.co/iamtatsuki05/Llama-JP-0.5B-init) and is trained on [hotchpotch/fineweb-2-edu-japanese](https://huggingface.co/datasets/hotchpotch/fineweb-2-edu-japanese). The model observes approximately 10B tokens with 1,024-token context windows, providing a decoder-only Japanese backbone for downstream generative tasks.
|
||||
|
||||
|
||||
- **[Hugging Face Collection](https://huggingface.co/collections/iamtatsuki05/mirei)**
|
||||
- **[GitHub](https://github.com/iamtatsuki05/MIREI)**
|
||||
|
||||

|
||||
|
||||
## Usage
|
||||
### Requirements
|
||||
```
|
||||
transformers>=4.51.0
|
||||
accelerate>=1.6.0
|
||||
sentencepiece>=0.2.0
|
||||
flash-attn>=2.7.3
|
||||
```
|
||||
|
||||
### Sample Code
|
||||
```python
|
||||
import torch
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
model_name = "iamtatsuki05/Llama-JP-0.5B-PT-stage1"
|
||||
model_kwargs = {
|
||||
"torch_dtype": torch.bfloat16,
|
||||
"attn_implementation": "flash_attention_2",
|
||||
"device_map": "auto",
|
||||
}
|
||||
tokenizer = AutoTokenizer.from_pretrained(model_name)
|
||||
model = AutoModelForCausalLM.from_pretrained(model_name, **model_kwargs)
|
||||
|
||||
prompt = "ちいかわのハチワレは"
|
||||
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
||||
output = model.generate(
|
||||
**inputs,
|
||||
max_new_tokens=256,
|
||||
temperature=0.8,
|
||||
top_p=0.9,
|
||||
do_sample=True,
|
||||
)
|
||||
print(tokenizer.decode(output[0], skip_special_tokens=True))
|
||||
```
|
||||
|
||||
|
||||
## Model Details
|
||||
- **Base model:** [iamtatsuki05/Llama-JP-0.5B-init](https://huggingface.co/iamtatsuki05/Llama-JP-0.5B-init)
|
||||
- **Architecture:** Llama
|
||||
- **Maximum sequence length:** 8,192 tokens
|
||||
- **Embedding dimension:** 1280
|
||||
- **Tokenizer:** SentencePiece / vocabulary size 102,400
|
||||
- **Positional encoding:** RoPE
|
||||
- **Supported languages:** Japanese
|
||||
|
||||
|
||||
## Model Series
|
||||
|
||||
The following checkpoints are initialized weights further pre-trained on [hotchpotch/fineweb-2-edu-japanese](https://huggingface.co/datasets/hotchpotch/fineweb-2-edu-japanese) for roughly 10B tokens with 1,024-token context lengths.
|
||||
|
||||
| ID | Architecture | #Param. | #Param.<br>w/o Emb. |
|
||||
|:-:|:-:|:-:|:-:|
|
||||
| [iamtatsuki05/ModernBERT-JP-0.5B-PT-stage1](https://huggingface.co/iamtatsuki05/ModernBERT-JP-0.5B-PT-stage1) | ModernBERT | 679M | 548M |
|
||||
| [iamtatsuki05/Llama-JP-0.5B-PT-stage1](https://huggingface.co/iamtatsuki05/Llama-JP-0.5B-PT-stage1)<br>(this model) | Llama | 661M | 530M |
|
||||
|
||||
|
||||
## Licence
|
||||
This model is distributed under the [MIT License](https://opensource.org/license/mit/).
|
||||
|
||||
|
||||
## How to Cite
|
||||
|
||||
```tex
|
||||
@article{MIREI
|
||||
title={同一条件下における Encoder/Decoder アーキテクチャによる文埋め込みの性能分析},
|
||||
author={岡田 龍樹 and 杉本 徹},
|
||||
journal={言語処理学会第 32 回年次大会 (NLP2026)},
|
||||
year={2026}
|
||||
}
|
||||
```
|
||||
Reference in New Issue
Block a user