初始化项目,由ModelHub XC社区提供模型
Model: third-intelligence/llm-jp-4-kappa-32b-a3b-v0.1 Source: Original Platform
This commit is contained in:
36
.gitattributes
vendored
Normal file
36
.gitattributes
vendored
Normal file
@@ -0,0 +1,36 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
||||
331
README.md
Normal file
331
README.md
Normal file
@@ -0,0 +1,331 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
language:
|
||||
- en
|
||||
- ja
|
||||
base_model:
|
||||
- llm-jp/llm-jp-4-32b-a3b-thinking
|
||||
pipeline_tag: text-generation
|
||||
library_name: transformers
|
||||
---
|
||||
|
||||
# llm-jp-4-kappa-32b-a3b-v0.1
|
||||
|
||||

|
||||
|
||||
|
||||
**llm-jp-4-kappa-32b-a3b-v0.1** は、**Third-intelligence** が研究開発の一環で試作した日本語・英語対応の Thinking 型大規模言語モデルです。[`llm-jp/llm-jp-4-32b-a3b-thinking`](https://huggingface.co/llm-jp/llm-jp-4-32b-a3b-thinking) をベースに、
|
||||
|
||||
- Reasoning(思考プロセスの生成)能力を強化するための **SFT (Supervised Fine-Tuning)**
|
||||
- 数学・コード領域における **強化学習 (Reinforcement Learning)**
|
||||
の 2 段階の追加学習を行い、ベースモデルが持つ日本語性能を維持しつつ、推論・数学・コーディング能力の向上を目指しました。
|
||||
|
||||
ベースモデルと同様、**Mixture-of-Experts (MoE) アーキテクチャ**(総パラメータ約 32B / activated 約 3B)を採用しているため、同規模の dense モデルと比較して推論時の計算コストを抑えつつ、高い性能を発揮します。
|
||||
|
||||
詳細については[技術ブログ](https://third-intelligence.com/ja/posts/kappa-v01-reasoning-post-training)を参照ください。
|
||||
|
||||
---
|
||||
|
||||
## モデル概要
|
||||
|
||||
| 項目 | 内容 |
|
||||
|---|---|
|
||||
| ベースモデル | [`llm-jp/llm-jp-4-32b-a3b-thinking`](https://huggingface.co/llm-jp/llm-jp-4-32b-a3b-thinking) |
|
||||
| アーキテクチャ | Mixture-of-Experts (MoE), Thinking モデル |
|
||||
| 総パラメータ数 | 約 32B |
|
||||
| Activated パラメータ数 | 約 3B |
|
||||
| 学習手法 | SFT(Reasoning 強化) + 強化学習(数学・コード) |
|
||||
| 対応言語 | 日本語 / 英語 |
|
||||
| 最大コンテキスト長 | 65,536 tokens |
|
||||
| ライセンス | Apache License 2.0(ベースモデルの条件も継承) |
|
||||
|
||||
---
|
||||
|
||||
## ベンチマーク
|
||||
|
||||
評価には [swallow-llm/swallow-evaluation-instruct](https://github.com/swallow-llm/swallow-evaluation-instruct) を使用しました(OpenRouter 経由でモデルが使用できるよう、一部修正を加えました。)
|
||||
|
||||
また、OpenRouter でデプロイされているモデルについては、serving しているプロバイダーの設定により量子化などが行われており、本来のフルプレシジョン版から性能が劣化している可能性があります。ただし、評価にあたっては OpenRouter 側の設定で可能な限り高いパフォーマンスとなるよう **Exacto** の設定を有効にしました。
|
||||
|
||||
LLM-as-a-judgeが必要になるベンチマークについては評価モデルにopenai/gpt-4.1を使用しました。(プロバイダーはAzureに固定しています。)
|
||||
|
||||
### 比較対象モデル
|
||||
|
||||
| 略称 | モデル ID | 備考 |
|
||||
|---|---|---|
|
||||
| **llm-jp-4-kappa (本モデル)** | `third-intelligence/llm-jp-4-kappa-32b-a3b-v0.1` | 本モデル |
|
||||
| llm-jp-4 | `llm-jp/llm-jp-4-32b-a3b-thinking` | ベースモデル |
|
||||
| gpt-oss-20b | `openai/gpt-oss-20b` (via OpenRouter) | |
|
||||
| nemotron-3-nano | `nvidia/nemotron-3-nano-30b-a3b` (via OpenRouter) | |
|
||||
| qwen3-30b | `qwen/qwen3-30b-a3b-thinking-2507` (via OpenRouter) | |
|
||||
|
||||
|
||||
### サマリー(主要スコア)
|
||||
|
||||
> **太字** は各行の最高スコアを示します。
|
||||
|
||||
| ベンチマーク | **llm-jp-4-kappa (本モデル)** | llm-jp-4 | gpt-oss-20b | nemotron-3-nano | qwen3-30b |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| MMLU ProX Japanese (avg) | 0.6724 | 0.6759 | 0.6886 | 0.5740 | **0.7547** |
|
||||
| MMLU Pro English (avg) | 0.7025 | 0.6956 | 0.7343 | 0.7180 | **0.8038** |
|
||||
| Japanese MT-Bench (avg) | **0.9339** | 0.9270 | 0.8749 | 0.8900 | 0.8693 |
|
||||
| English MT-Bench (avg) | 0.9226 | 0.9220 | 0.8899 | 0.9216 | **0.9304** |
|
||||
| JEMHopQA F1 | 0.6329 | **0.6391** | 0.4737 | 0.2329 | 0.5597 |
|
||||
| MCLM Math 100 (Japanese) | 0.8788 | 0.8687 | 0.9394 | 0.8081 | **0.9697** |
|
||||
| GPQA Diamond | 0.5505 | 0.4899 | 0.6364 | 0.5909 | **0.7121** |
|
||||
| AIME 2024/2025 | 0.4333 | 0.3833 | 0.5667 | 0.3667 | **0.8833** |
|
||||
|
||||
### Japanese MT-Bench(カテゴリ別)
|
||||
|
||||
| カテゴリ | **llm-jp-4-kappa (本モデル)** | llm-jp-4 | gpt-oss-20b | nemotron-3-nano | qwen3-30b |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| coding | **0.9480** | 0.9420 | 0.9340 | 0.9450 | 0.8680 |
|
||||
| extraction | **0.8730** | 0.8650 | 0.8380 | 0.8030 | 0.8400 |
|
||||
| humanities | **0.9870** | 0.9770 | 0.8300 | 0.9550 | 0.9540 |
|
||||
| math | **0.9950** | 0.9800 | 0.9920 | 0.9070 | 0.9470 |
|
||||
| reasoning | **0.8470** | 0.8310 | 0.8090 | 0.8050 | 0.7590 |
|
||||
| roleplay | **0.9480** | 0.9340 | 0.8310 | 0.8720 | 0.8130 |
|
||||
| stem | 0.9670 | **0.9770** | 0.9120 | 0.9610 | 0.9130 |
|
||||
| writing | 0.9060 | **0.9100** | 0.8530 | 0.8720 | 0.8600 |
|
||||
| 1st turn | **0.9545** | 0.9520 | 0.9145 | 0.8888 | 0.8930 |
|
||||
| 2nd turn | **0.9133** | 0.9020 | 0.8353 | 0.8913 | 0.8455 |
|
||||
| **平均** | **0.9339** | 0.9270 | 0.8749 | 0.8900 | 0.8693 |
|
||||
|
||||
### English MT-Bench(カテゴリ別)
|
||||
|
||||
| カテゴリ | **llm-jp-4-kappa (本モデル)** | llm-jp-4 | gpt-oss-20b | nemotron-3-nano | qwen3-30b |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| coding | 0.9240 | **0.9340** | 0.9040 | 0.8910 | 0.8860 |
|
||||
| extraction | 0.8590 | 0.8220 | 0.8610 | 0.8430 | **0.8980** |
|
||||
| humanities | 0.9570 | **0.9800** | 0.9250 | 0.9410 | 0.9600 |
|
||||
| math | 0.9990 | 0.9990 | 0.9720 | 0.9960 | **1.0000** |
|
||||
| reasoning | 0.8460 | 0.8260 | 0.7590 | 0.8990 | **0.9030** |
|
||||
| roleplay | 0.9290 | **0.9360** | 0.8710 | 0.9340 | 0.9090 |
|
||||
| stem | **0.9910** | 0.9780 | 0.9380 | 0.9410 | 0.9730 |
|
||||
| writing | 0.8760 | 0.9010 | 0.8890 | **0.9280** | 0.9140 |
|
||||
| 1st turn | 0.9468 | 0.9368 | 0.9208 | **0.9523** | 0.9515 |
|
||||
| 2nd turn | 0.8985 | 0.9073 | 0.8590 | 0.8910 | **0.9093** |
|
||||
| **平均** | 0.9226 | 0.9220 | 0.8899 | 0.9216 | **0.9304** |
|
||||
|
||||
### MMLU ProX Japanese / MMLU Pro English(科目別)
|
||||
|
||||
<details>
|
||||
<summary>📊 MMLU ProX Japanese 科目別スコアを開く</summary>
|
||||
|
||||
| 科目 | **llm-jp-4-kappa (本モデル)** | llm-jp-4 | gpt-oss-20b | nemotron-3-nano | qwen3-30b |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| business | 0.7338 | 0.7326 | 0.7782 | 0.5741 | **0.8238** |
|
||||
| law | 0.3326 | 0.3472 | 0.3034 | 0.2899 | **0.4067** |
|
||||
| psychology | 0.6454 | 0.6654 | 0.6654 | 0.5890 | **0.7180** |
|
||||
| biology | 0.7950 | 0.7880 | 0.7964 | 0.6820 | **0.8522** |
|
||||
| chemistry | 0.7686 | 0.7712 | 0.8012 | 0.6793 | **0.8640** |
|
||||
| history | 0.5039 | 0.5223 | 0.5118 | 0.4173 | **0.5984** |
|
||||
| other | 0.6006 | 0.5898 | 0.5931 | 0.5011 | **0.6656** |
|
||||
| health | 0.6332 | 0.6332 | 0.6332 | 0.5109 | **0.6594** |
|
||||
| economics | 0.7370 | 0.7322 | 0.7464 | 0.6351 | **0.8128** |
|
||||
| math | 0.8586 | 0.8608 | 0.8993 | 0.7631 | **0.9023** |
|
||||
| physics | 0.7929 | 0.7968 | 0.8075 | 0.6659 | **0.8676** |
|
||||
| computer science | 0.7512 | 0.7537 | 0.8073 | 0.6634 | **0.8195** |
|
||||
| philosophy | 0.5050 | 0.5170 | 0.5010 | 0.3928 | **0.5892** |
|
||||
| engineering | 0.5160 | 0.5222 | 0.5470 | 0.4314 | **0.7379** |
|
||||
| **平均** | 0.6724 | 0.6759 | 0.6886 | 0.5740 | **0.7547** |
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary>📊 MMLU Pro English 科目別スコアを開く</summary>
|
||||
|
||||
| 科目 | **llm-jp-4-kappa (本モデル)** | llm-jp-4 | gpt-oss-20b | nemotron-3-nano | qwen3-30b |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| business | 0.7529 | 0.7452 | 0.8137 | 0.7617 | **0.8631** |
|
||||
| law | 0.4532 | 0.4405 | 0.4142 | 0.4714 | **0.5450** |
|
||||
| psychology | 0.7306 | 0.7168 | 0.7206 | 0.7556 | **0.7757** |
|
||||
| biology | 0.8410 | 0.8536 | 0.8536 | 0.8438 | **0.8870** |
|
||||
| chemistry | 0.7420 | 0.7482 | 0.8269 | 0.7889 | **0.8763** |
|
||||
| history | 0.6273 | 0.6010 | 0.5906 | 0.6089 | **0.6483** |
|
||||
| other | 0.6504 | 0.6353 | 0.6461 | 0.6526 | **0.7370** |
|
||||
| health | 0.6883 | 0.6601 | 0.7237 | 0.7152 | **0.7665** |
|
||||
| economics | 0.7820 | 0.7761 | 0.7998 | 0.7773 | **0.8483** |
|
||||
| math | 0.8653 | 0.8564 | 0.9112 | 0.8756 | **0.9445** |
|
||||
| physics | 0.7852 | 0.7814 | 0.8276 | 0.7921 | **0.8961** |
|
||||
| computer science | 0.7268 | 0.7341 | 0.8146 | 0.7756 | **0.8439** |
|
||||
| philosophy | 0.5892 | 0.6132 | 0.6192 | 0.5972 | **0.6854** |
|
||||
| engineering | 0.5057 | 0.4902 | 0.5944 | 0.5304 | **0.7678** |
|
||||
| **平均** | 0.7025 | 0.6956 | 0.7343 | 0.7180 | **0.8038** |
|
||||
|
||||
</details>
|
||||
|
||||
---
|
||||
|
||||
## 使い方
|
||||
|
||||
### 必要環境
|
||||
|
||||
- GPU: 推奨 **A100 80GB ×2 以上 / H100 80GB ×2 以上**(bfloat16・MoE expert parallel 動作時)
|
||||
- 単一 80GB GPU でも `tensor-parallel-size=1` で動作可能ですが、KV キャッシュ確保のため `--max-model-len` を下げる必要がある場合があります。
|
||||
- CUDA 12.4+
|
||||
- Python 3.10+
|
||||
|
||||
### 1. transformers での推論
|
||||
|
||||
HuggingFace transformers ベースのパイプラインに組み込みたい場合は、以下のように直接ロードして使用できます。なお transformers から直接生成すると、出力には思考プロセス用の特殊トークン(`<|channel|>analysis` や `<|channel|>final` など)がそのまま含まれるため、思考と最終回答を分離して扱いたい場合は前述の vLLM + cookbook 構成を利用するか、出力をパースしてください。
|
||||
|
||||
```python
|
||||
import torch
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
model_id = "third-intelligence/llm-jp-4-kappa-32b-a3b-v0.1"
|
||||
|
||||
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
model_id,
|
||||
torch_dtype=torch.bfloat16,
|
||||
device_map="auto",
|
||||
trust_remote_code=True,
|
||||
)
|
||||
|
||||
messages = [
|
||||
{"role": "user", "content": "日本の宇宙開発の歴史について教えてください。"},
|
||||
]
|
||||
|
||||
# chat_template により思考用プロンプトが付与されます
|
||||
input_ids = tokenizer.apply_chat_template(
|
||||
messages,
|
||||
add_generation_prompt=True,
|
||||
return_tensors="pt",
|
||||
).to(model.device)
|
||||
|
||||
with torch.no_grad():
|
||||
output_ids = model.generate(
|
||||
input_ids,
|
||||
max_new_tokens=4096,
|
||||
do_sample=True,
|
||||
temperature=0.7,
|
||||
top_p=0.95,
|
||||
pad_token_id=tokenizer.eos_token_id,
|
||||
)
|
||||
|
||||
generated = output_ids[0][input_ids.shape[-1]:]
|
||||
print(tokenizer.decode(generated, skip_special_tokens=False))
|
||||
```
|
||||
|
||||
### 2. vLLM での推論(推奨)
|
||||
|
||||
LLM-jp-4 系のモデルは、独自の Reasoning フォーマット(`<|channel|>final` などの特殊トークンによる思考プロセス出力)を採用しているため、素の vLLM ではそのまま正しく serve することができません。公式の [`llm-jp/llm-jp-4-cookbook`](https://github.com/llm-jp/llm-jp-4-cookbook) を利用することが推奨されています。
|
||||
|
||||
#### セットアップ
|
||||
|
||||
```bash
|
||||
# 1. vLLM 0.17 以降をインストール
|
||||
pip install "vllm>=0.17.0"
|
||||
|
||||
# 2. LLM-jp-4 cookbook を clone
|
||||
git clone https://github.com/llm-jp/llm-jp-4-cookbook.git
|
||||
cd llm-jp-4-cookbook/llmjp4_vllm
|
||||
```
|
||||
|
||||
#### vLLM サーバの起動
|
||||
|
||||
```bash
|
||||
python example_cli.py serve third-intelligence/llm-jp-4-kappa-32b-a3b-v0.1 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--served-model-name llm-jp-4-kappa \
|
||||
--trust-remote-code \
|
||||
--dtype bfloat16 \
|
||||
--tensor-parallel-size 2 \
|
||||
--gpu-memory-utilization 0.85 \
|
||||
--max-model-len 65536 \
|
||||
--reasoning-parser llmjp4 \
|
||||
--enable-expert-parallel
|
||||
```
|
||||
|
||||
主なオプションの意味:
|
||||
|
||||
| オプション | 説明 |
|
||||
|---|---|
|
||||
| `--reasoning-parser llmjp4` | LLM-jp-4 thinking 用の reasoning parser を有効化。**必須** |
|
||||
| `--enable-expert-parallel` | MoE の expert parallel を有効化(マルチ GPU で推奨) |
|
||||
| `--tensor-parallel-size` | テンソル並列数。GPU 枚数に合わせて変更してください(1, 2, 4, 8 など) |
|
||||
| `--max-model-len` | 最大コンテキスト長(最大 65,536) |
|
||||
| `--gpu-memory-utilization` | GPU メモリ使用率の上限。OOM になる場合は下げてください |
|
||||
|
||||
> **トラブルシュート: `RuntimeError: Already borrowed`**
|
||||
> 環境によっては、reasoning parser の初期化時に tokenizer の競合により上記エラーが発生することがあります。その場合は cookbook 内 `llmjp4_reasoning_parser.py` の以下 2 行を、tokenizer.encode を経由しない形に置き換えてください。
|
||||
>
|
||||
> ```python
|
||||
> # Before
|
||||
> self._reasoning_end_prefix = tokenizer.encode("<|channel|>final")
|
||||
> self._reasoning_prefill = tokenizer.encode("<|start|>assistant")
|
||||
>
|
||||
> # After
|
||||
> self._reasoning_end_prefix = [9, 2520]
|
||||
> self._reasoning_prefill = [10, 12811]
|
||||
> ```
|
||||
|
||||
#### OpenAI 互換 API で叩く
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "llm-jp-4-kappa",
|
||||
"messages": [
|
||||
{"role": "user", "content": "整数 a, b が a^2 + b^2 = 50 を満たすとき、(a, b) の組み合わせを全て求めてください。"}
|
||||
],
|
||||
"max_tokens": 4096,
|
||||
"temperature": 0.7
|
||||
}'
|
||||
```
|
||||
|
||||
Python (OpenAI SDK) からも同様に呼び出せます。`--reasoning-parser llmjp4` を有効にしているため、思考プロセスは `reasoning_content` フィールドに、最終回答は `content` フィールドに分離されて返ります。
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
|
||||
|
||||
resp = client.chat.completions.create(
|
||||
model="llm-jp-4-kappa",
|
||||
messages=[
|
||||
{"role": "user", "content": "再帰関数を使ってフィボナッチ数列の n 番目の値を返す Python 関数を書いてください。"}
|
||||
],
|
||||
max_tokens=4096,
|
||||
temperature=0.7,
|
||||
)
|
||||
|
||||
msg = resp.choices[0].message
|
||||
print("=== 思考プロセス ===")
|
||||
print(getattr(msg, "reasoning_content", ""))
|
||||
print("=== 最終回答 ===")
|
||||
print(msg.content)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 注意事項
|
||||
|
||||
本モデルは **研究目的で公開されているモデル** です。出力の正確性・安全性・倫理性は保証されず、事実と異なる内容(ハルシネーション)や不適切な内容を生成する可能性があります。商用利用やミッションクリティカルな用途での使用は想定していません。本モデルの利用により生じたいかなる損害についても開発者は責任を負いません。ライセンスはベースモデル(`llm-jp/llm-jp-4-32b-a3b-thinking`)の条件も併せて確認・遵守してください。
|
||||
|
||||
## ライセンス
|
||||
|
||||
本モデルは **Apache License 2.0** のもとで公開されます。ベースモデル `llm-jp/llm-jp-4-32b-a3b-thinking` のライセンス・利用規約も併せて遵守してください。
|
||||
|
||||
## 謝辞
|
||||
|
||||
- ベースモデルを公開してくださっている [LLM-jp プロジェクト](https://llm-jp.nii.ac.jp/)
|
||||
- データを公開してくださっている[NVIDIA](https://huggingface.co/nvidia)
|
||||
- 評価フレームワーク [`swallow-llm/swallow-evaluation-instruct`](https://github.com/swallow-llm/swallow-evaluation-instruct) を公開してくださっている [Swallow プロジェクト](https://swallow-llm.github.io/index.ja.html)
|
||||
|
||||
をはじめとする皆様に最大の感謝を申し上げます。
|
||||
|
||||
## 引用
|
||||
|
||||
```bibtex
|
||||
@misc{llm-jp-4-kappa-32b-a3b-v0_1,
|
||||
title = {llm-jp-4-kappa-32b-a3b-v0.1},
|
||||
author = {third-intelligence},
|
||||
year = {2026},
|
||||
howpublished = {\url{https://huggingface.co/third-intelligence/llm-jp-4-kappa-32b-a3b-v0.1}}
|
||||
}
|
||||
```
|
||||
63
chat_template.jinja
Normal file
63
chat_template.jinja
Normal file
@@ -0,0 +1,63 @@
|
||||
{#-
|
||||
Simplified Harmony chat template for GRPO training (no tool support).
|
||||
Based on the official llm-jp-4 chat_template.jinja.
|
||||
-#}
|
||||
{#- System Message Construction -#}
|
||||
{%- macro build_system_message() -%}
|
||||
{%- if model_identity is not defined %}
|
||||
{%- set model_identity = "You are LLM-jp-4, a large language model trained by LLM-jp." %}
|
||||
{%- endif %}
|
||||
{{- model_identity + "\n" -}}
|
||||
{%- if knowledge_cutoff is not defined %}
|
||||
{%- set knowledge_cutoff = "2025-12" %}
|
||||
{%- endif %}
|
||||
{{- "Knowledge cutoff: " + knowledge_cutoff + "\n" -}}
|
||||
{%- if conversation_start_date is not defined %}
|
||||
{%- set conversation_start_date = "2026-04-14" %}
|
||||
{%- endif %}
|
||||
{{- "Current date: " + conversation_start_date + "\n\n" }}
|
||||
{%- if reasoning_effort is not defined %}
|
||||
{%- set reasoning_effort = "medium" %}
|
||||
{%- endif %}
|
||||
{{- "Reasoning: " + reasoning_effort + "\n\n" }}
|
||||
{{- "# Valid channels: analysis, commentary, final. Channel must be included for every message." }}
|
||||
{%- endmacro -%}
|
||||
|
||||
{#- Main Template Logic -#}
|
||||
{{- "<|start|>system<|message|>" }}
|
||||
{{- build_system_message() }}
|
||||
{{- "<|end|>" }}
|
||||
|
||||
{%- if messages[0].role == "developer" or messages[0].role == "system" %}
|
||||
{%- set developer_message = messages[0].content %}
|
||||
{%- set loop_messages = messages[1:] %}
|
||||
{%- else %}
|
||||
{%- set developer_message = "" %}
|
||||
{%- set loop_messages = messages %}
|
||||
{%- endif %}
|
||||
|
||||
{%- if developer_message %}
|
||||
{{- "<|start|>developer<|message|>" }}
|
||||
{{- "# Instructions\n\n" }}
|
||||
{{- developer_message }}
|
||||
{{- "<|end|>" }}
|
||||
{%- endif %}
|
||||
|
||||
{%- for message in loop_messages -%}
|
||||
{%- if message.role == 'assistant' -%}
|
||||
{%- if loop.last and not add_generation_prompt %}
|
||||
{%- if "thinking" in message %}
|
||||
{{- "<|start|>assistant<|channel|>analysis<|message|>" + message.thinking + "<|end|>" }}
|
||||
{%- endif %}
|
||||
{{- "<|start|>assistant<|channel|>final<|message|>" + message.content + "<|return|>" }}
|
||||
{%- else %}
|
||||
{{- "<|start|>assistant<|channel|>final<|message|>" + message.content + "<|end|>" }}
|
||||
{%- endif %}
|
||||
{%- elif message.role == 'user' -%}
|
||||
{{- "<|start|>user<|message|>" + message.content + "<|end|>" }}
|
||||
{%- endif -%}
|
||||
{%- endfor -%}
|
||||
|
||||
{%- if add_generation_prompt -%}
|
||||
<|start|>assistant
|
||||
{%- endif -%}
|
||||
39
config.json
Normal file
39
config.json
Normal file
@@ -0,0 +1,39 @@
|
||||
{
|
||||
"architectures": [
|
||||
"Qwen3MoeForCausalLM"
|
||||
],
|
||||
"attention_bias": false,
|
||||
"attention_dropout": 0.0,
|
||||
"bos_token_id": 1,
|
||||
"decoder_sparse_step": 1,
|
||||
"dtype": "bfloat16",
|
||||
"eos_token_id": 2,
|
||||
"head_dim": 128,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 2560,
|
||||
"initializer_range": 0.006,
|
||||
"intermediate_size": 7680,
|
||||
"max_position_embeddings": 65536,
|
||||
"max_window_layers": 32,
|
||||
"mlp_only_layers": [],
|
||||
"model_type": "qwen3_moe",
|
||||
"moe_intermediate_size": 960,
|
||||
"norm_topk_prob": true,
|
||||
"num_attention_heads": 40,
|
||||
"num_experts": 128,
|
||||
"num_experts_per_tok": 8,
|
||||
"num_hidden_layers": 32,
|
||||
"num_key_value_heads": 4,
|
||||
"output_router_logits": false,
|
||||
"pad_token_id": 4,
|
||||
"rms_norm_eps": 1e-06,
|
||||
"rope_scaling": null,
|
||||
"rope_theta": 500000,
|
||||
"router_aux_loss_coef": 0.01,
|
||||
"sliding_window": null,
|
||||
"tie_word_embeddings": false,
|
||||
"transformers_version": "4.57.6",
|
||||
"use_cache": false,
|
||||
"use_sliding_window": false,
|
||||
"vocab_size": 196608
|
||||
}
|
||||
8
generation_config.json
Normal file
8
generation_config.json
Normal file
@@ -0,0 +1,8 @@
|
||||
{
|
||||
"_from_model_config": true,
|
||||
"bos_token_id": 1,
|
||||
"eos_token_id": 2,
|
||||
"pad_token_id": 4,
|
||||
"transformers_version": "4.57.6",
|
||||
"use_cache": false
|
||||
}
|
||||
129
llmjp4_harmony.py
Normal file
129
llmjp4_harmony.py
Normal file
@@ -0,0 +1,129 @@
|
||||
# Generic parser for OpenAI Harmony format.
|
||||
|
||||
from dataclasses import dataclass
|
||||
from enum import Enum
|
||||
from typing import Iterator, Sequence
|
||||
|
||||
from transformers import PreTrainedTokenizerBase as TokenizerLike
|
||||
|
||||
|
||||
class HarmonyMessageEndType(Enum):
|
||||
INCOMPLETE = 0
|
||||
END = 1
|
||||
CALL = 2
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class HarmonySequence:
|
||||
"""A data class representing a sequence of tokens in the Harmony format."""
|
||||
token_ids: list[int]
|
||||
start: int # Start position of the sequence in the original token sequence
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class HarmonyMessage:
|
||||
"""A data class representing a message in the Harmony format."""
|
||||
end: HarmonyMessageEndType
|
||||
role: HarmonySequence | None = None
|
||||
channel: HarmonySequence | None = None
|
||||
constrain: HarmonySequence | None = None
|
||||
content: HarmonySequence | None = None
|
||||
|
||||
|
||||
class HarmonyMessageParser:
|
||||
"""A parser that performs lexical analysis to extract Harmony messages."""
|
||||
|
||||
def __init__(self, tokenizer: TokenizerLike):
|
||||
vocab = tokenizer.get_vocab()
|
||||
self._begin_map = {
|
||||
vocab["<|start|>"]: "role",
|
||||
vocab["<|channel|>"]: "channel",
|
||||
vocab["<|constrain|>"]: "constrain",
|
||||
vocab["<|message|>"]: "content",
|
||||
}
|
||||
self._end_map = {
|
||||
vocab["<|end|>"]: HarmonyMessageEndType.END,
|
||||
vocab["<|return|>"]: HarmonyMessageEndType.END,
|
||||
vocab["<|call|>"]: HarmonyMessageEndType.CALL,
|
||||
}
|
||||
|
||||
def iter_messages(self, token_ids: Sequence[int]) -> Iterator[HarmonyMessage]:
|
||||
"""
|
||||
Parse given token ids into messages.
|
||||
|
||||
Args:
|
||||
token_ids: A sequence of token ids to be parsed.
|
||||
|
||||
Yields:
|
||||
Detected HarmonyMessages.
|
||||
"""
|
||||
|
||||
message_dict: dict[str, HarmonySequence] = {}
|
||||
section: str | None = None # None indicates out-of-message.
|
||||
text_ids: list[int] = []
|
||||
text_start: int | None = None
|
||||
|
||||
for token_position, token_id in enumerate(token_ids):
|
||||
if token_id in self._begin_map:
|
||||
if section is not None:
|
||||
message_dict[section] = HarmonySequence(
|
||||
token_ids=text_ids,
|
||||
start=text_start,
|
||||
)
|
||||
section = self._begin_map[token_id]
|
||||
text_ids = []
|
||||
text_start = token_position + 1
|
||||
|
||||
elif token_id in self._end_map:
|
||||
if section is not None:
|
||||
message_dict[section] = HarmonySequence(
|
||||
token_ids=text_ids,
|
||||
start=text_start,
|
||||
)
|
||||
|
||||
yield HarmonyMessage(**message_dict, end=self._end_map[token_id])
|
||||
|
||||
message_dict = {}
|
||||
section = None
|
||||
text_ids = []
|
||||
text_start = None
|
||||
|
||||
else:
|
||||
if section is not None:
|
||||
text_ids.append(token_id)
|
||||
|
||||
if section is not None:
|
||||
message_dict[section] = HarmonySequence(
|
||||
token_ids=text_ids,
|
||||
start=text_start,
|
||||
)
|
||||
yield HarmonyMessage(**message_dict, end=HarmonyMessageEndType.INCOMPLETE)
|
||||
|
||||
def get_all_messages(self, token_ids: Sequence[int]) -> list[HarmonyMessage]:
|
||||
"""
|
||||
Parse given token ids into messages.
|
||||
|
||||
Args:
|
||||
token_ids: A sequence of token ids to be parsed.
|
||||
|
||||
Returns:
|
||||
A list of detected HarmonyMessages.
|
||||
"""
|
||||
return list(self.iter_messages(token_ids))
|
||||
|
||||
def reverse_iter_messages(self, token_ids: Sequence[int]) -> Iterator[HarmonyMessage]:
|
||||
"""
|
||||
Parse given token ids into messages in reverse order.
|
||||
|
||||
Args:
|
||||
token_ids: A sequence of token ids to be parsed.
|
||||
|
||||
Yields:
|
||||
Detected HarmonyMessages in reverse order.
|
||||
"""
|
||||
end_position = len(token_ids)
|
||||
|
||||
for i in range(len(token_ids) - 1, -1, -1):
|
||||
if token_ids[i] == self._start_id:
|
||||
yield next(self.iter_messages(token_ids[i:end_position]))
|
||||
end_position = i
|
||||
101
llmjp4_tokenizer.py
Normal file
101
llmjp4_tokenizer.py
Normal file
@@ -0,0 +1,101 @@
|
||||
# llm-jp-4 tokenizer
|
||||
|
||||
from collections.abc import Sequence
|
||||
import os
|
||||
|
||||
from transformers import LlamaTokenizerFast
|
||||
from tokenizers import Tokenizer
|
||||
|
||||
from .llmjp4_harmony import HarmonyMessageParser, HarmonyMessage
|
||||
|
||||
|
||||
class Llmjp4Tokenizer(LlamaTokenizerFast):
|
||||
_HARMONY_TOKENS: set[str] = {
|
||||
"<|start|>",
|
||||
"<|message|>",
|
||||
"<|channel|>",
|
||||
"<|constrain|>",
|
||||
"<|end|>",
|
||||
"<|return|>",
|
||||
"<|call|>",
|
||||
}
|
||||
|
||||
# NOTE(odashi):
|
||||
# Response schemas are not recognized automatically.
|
||||
# We need to define them manually.
|
||||
# https://github.com/huggingface/trl/issues/4609
|
||||
_RESPONSE_SCHEMA = {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"role": {"const": "assistant"},
|
||||
"content": {"type": "string", "x-regex": r"<\|channel\|>final<\|message\|>(.*?)(?:<\|end\|>|<\|return\|>|$)"},
|
||||
"thinking": {"type": "string", "x-regex": r"<\|channel\|>analysis<\|message\|>(.*?)<\|end\|>"},
|
||||
"tool_calls": {
|
||||
"x-regex-iterator": r"<\|channel\|>commentary (to=functions\..*?<\|message\|>.*?)(?:<\|call\|>|$)",
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"type": {"const": "function"},
|
||||
"function": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"name": {"type": "string", "x-regex": r"^to=functions\.(\w+)"},
|
||||
"arguments": {
|
||||
"type": "object",
|
||||
"x-regex": r"<\|message\|>(.*)",
|
||||
"x-parser": "json",
|
||||
"additionalProperties": {"type": "any"},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
|
||||
@classmethod
|
||||
def convert_to_native_format(cls, **kwargs):
|
||||
# NOTE(odashi):
|
||||
# Workaround for transformers 5.x.
|
||||
# Guaranteeing the same inner behavior with TokenizersBackend.
|
||||
# https://github.com/huggingface/transformers/blob/7d9754a05193eb79b1d86aa744b622b8068008cd/src/transformers/tokenization_utils_tokenizers.py#L110-L116
|
||||
local_kwargs = dict(kwargs)
|
||||
fast_tokenizer_file = local_kwargs.pop("tokenizer_file", None)
|
||||
if fast_tokenizer_file is None or not os.path.isfile(fast_tokenizer_file):
|
||||
raise ValueError("Tokenizer file must exist.")
|
||||
|
||||
local_kwargs["tokenizer_object"] = Tokenizer.from_file(fast_tokenizer_file)
|
||||
return local_kwargs
|
||||
|
||||
def __init__(self, *args, **kwargs):
|
||||
super().__init__(*args, **kwargs)
|
||||
|
||||
self.response_schema = self._RESPONSE_SCHEMA
|
||||
|
||||
self._harmony_token_ids = {
|
||||
self.convert_tokens_to_ids(token)
|
||||
for token in self._HARMONY_TOKENS
|
||||
}
|
||||
|
||||
def _decode(self, token_ids: int | list[int], *args, **kwargs):
|
||||
if isinstance(token_ids, int):
|
||||
token_ids = [token_ids]
|
||||
|
||||
result: list[str] = []
|
||||
prev_pos = 0
|
||||
|
||||
# NOTE(odashi):
|
||||
# Ensure that text tokens are decoded without preceding Harmony tokens
|
||||
# to avoid incorrect addition of whitespaces.
|
||||
for pos, token_id in enumerate(token_ids, start=1):
|
||||
if token_id in self._harmony_token_ids or pos == len(token_ids):
|
||||
result.append(super()._decode(token_ids[prev_pos:pos], *args, **kwargs))
|
||||
prev_pos = pos
|
||||
|
||||
return "".join(result)
|
||||
|
||||
def parse_harmony_message(self, token_ids: Sequence[int]) -> list[HarmonyMessage]:
|
||||
"""Helper function to parse token IDs into Harmony messages."""
|
||||
return HarmonyMessageParser(self).get_all_messages(token_ids)
|
||||
3
model-00001-of-00013.safetensors
Normal file
3
model-00001-of-00013.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:1ebd2e6ac46ed407ec5917f44d1b83997b3fe57f8aabf0be51070a2bc41e5d9c
|
||||
size 4995959408
|
||||
3
model-00002-of-00013.safetensors
Normal file
3
model-00002-of-00013.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:4e2f6ba09fda33d11c72efb03c1e93c11e02f469c8e9467b55cdead6b191c962
|
||||
size 4997602792
|
||||
3
model-00003-of-00013.safetensors
Normal file
3
model-00003-of-00013.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:6556aefafda3e0d76756b77ea8ec4c8387521649a9788b85197bdeed552c40a4
|
||||
size 4999858104
|
||||
3
model-00004-of-00013.safetensors
Normal file
3
model-00004-of-00013.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:e39228d108526bcdaf20f4bd8152dafdef7c316ef2666385e3c845c290b29292
|
||||
size 4999242896
|
||||
3
model-00005-of-00013.safetensors
Normal file
3
model-00005-of-00013.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:b1a351d88a7a4e33d47f24f4bf509f48a487ac73c60450690835251f362ff9ee
|
||||
size 4998255672
|
||||
3
model-00006-of-00013.safetensors
Normal file
3
model-00006-of-00013.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:49f4d27b835ce5ff68b5e57e38f04d7e16a54f5bfe60e644e6a651061d2c6282
|
||||
size 4998904016
|
||||
3
model-00007-of-00013.safetensors
Normal file
3
model-00007-of-00013.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:6dd612b07216adc92a5ec26e310d84ffabbd9ac95953f779787b347e77ced0c6
|
||||
size 4998569752
|
||||
3
model-00008-of-00013.safetensors
Normal file
3
model-00008-of-00013.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:e5523ccfdd3ff8909749cbf8623d2f855d249a143070c8201c3f4e7c128fa601
|
||||
size 4999213704
|
||||
3
model-00009-of-00013.safetensors
Normal file
3
model-00009-of-00013.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:fd493e9ddb2e0458821b01b6fcb89296ede0dd0d426efb756e63acd23be772d1
|
||||
size 4998245056
|
||||
3
model-00010-of-00013.safetensors
Normal file
3
model-00010-of-00013.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:62c4133cf9d1c9b92901d0505ce46c142a3fe3d2c9ef67a7c6f0cf0603048b6e
|
||||
size 4998278984
|
||||
3
model-00011-of-00013.safetensors
Normal file
3
model-00011-of-00013.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:d5b18eeaf5272fab1d983048423b3ad4c8593c94ebc10a4431fea14960d31654
|
||||
size 4996293096
|
||||
3
model-00012-of-00013.safetensors
Normal file
3
model-00012-of-00013.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:eb9e8cdf080598a22f94e0a8c134a15eb662e1d1ec44ba0125c5622108d8fd0b
|
||||
size 4989398160
|
||||
3
model-00013-of-00013.safetensors
Normal file
3
model-00013-of-00013.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:6af20dfcb6ea88ec9e7dcb5ef3e201b1ad7ccedf2e84b66bd0415e01d417a164
|
||||
size 4309790912
|
||||
12587
model.safetensors.index.json
Normal file
12587
model.safetensors.index.json
Normal file
File diff suppressed because it is too large
Load Diff
51
special_tokens_map.json
Normal file
51
special_tokens_map.json
Normal file
@@ -0,0 +1,51 @@
|
||||
{
|
||||
"bos_token": {
|
||||
"content": "<|startoftext|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
},
|
||||
"cls_token": {
|
||||
"content": "<|cls|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
},
|
||||
"eos_token": {
|
||||
"content": "<|return|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
},
|
||||
"mask_token": {
|
||||
"content": "<|mask|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
},
|
||||
"pad_token": {
|
||||
"content": "<|endoftext|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
},
|
||||
"sep_token": {
|
||||
"content": "<|sep|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
},
|
||||
"unk_token": {
|
||||
"content": "<|unk|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
}
|
||||
}
|
||||
3
tokenizer.json
Normal file
3
tokenizer.json
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:7ac381ce8e5664fa1a196b943018b6fee4543f49ba87b1b6454cc1b53b23ece7
|
||||
size 12868535
|
||||
2845
tokenizer_config.json
Normal file
2845
tokenizer_config.json
Normal file
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user