From e26d20dee2e192d196436e04e893deb347a7c0fb Mon Sep 17 00:00:00 2001 From: ModelHub XC Date: Tue, 14 Jul 2026 02:23:09 +0800 Subject: [PATCH] =?UTF-8?q?=E5=88=9D=E5=A7=8B=E5=8C=96=E9=A1=B9=E7=9B=AE?= =?UTF-8?q?=EF=BC=8C=E7=94=B1ModelHub=20XC=E7=A4=BE=E5=8C=BA=E6=8F=90?= =?UTF-8?q?=E4=BE=9B=E6=A8=A1=E5=9E=8B?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Model: llm-jp/llm-jp-4-8b-thinking-gguf Source: Original Platform --- .gitattributes | 39 ++++++++ README.md | 157 +++++++++++++++++++++++++++++++ imatrix.dat | 3 + llm-jp-4-8b-thinking-BF16.gguf | 3 + llm-jp-4-8b-thinking-Q4_K_M.gguf | 3 + v4_pretraining_overview.png | 3 + 6 files changed, 208 insertions(+) create mode 100644 .gitattributes create mode 100644 README.md create mode 100644 imatrix.dat create mode 100644 llm-jp-4-8b-thinking-BF16.gguf create mode 100644 llm-jp-4-8b-thinking-Q4_K_M.gguf create mode 100644 v4_pretraining_overview.png diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..f77d07f --- /dev/null +++ b/.gitattributes @@ -0,0 +1,39 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +llm-jp-4-8b-thinking-BF16.gguf filter=lfs diff=lfs merge=lfs -text +llm-jp-4-8b-thinking-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text +imatrix.dat filter=lfs diff=lfs merge=lfs -text +*.png filter=lfs diff=lfs merge=lfs -text diff --git a/README.md b/README.md new file mode 100644 index 0000000..b99a9d6 --- /dev/null +++ b/README.md @@ -0,0 +1,157 @@ +--- +license: apache-2.0 +language: +- en +- ja +programming_language: +- C +- C++ +- C# +- Go +- Java +- JavaScript +- Lua +- PHP +- Python +- Ruby +- Rust +- Scala +- TypeScript +pipeline_tag: text-generation +library_name: transformers +inference: false +--- +# llm-jp-4-8b-thinking-gguf + +LLM-jp-4 is a series of large language models developed by the [Research and Development Center for Large Language Models](https://llmc.nii.ac.jp/) at the [National Institute of Informatics](https://www.nii.ac.jp/en/). + +This repository provides the **llm-jp-4-8b-thinking-gguf**. +For an overview of the LLM-jp-4 models across different parameter sizes, please refer to: + - [LLM-jp-4 Models](https://huggingface.co/collections/llm-jp/llm-jp-4-models) + +Base models are trained with pre-training and mid-training only. +Post-trained models are aligned using supervised fine-tuning (SFT) and direct preference optimization (DPO), without reinforcement learning. +> [!NOTE] +> While the **thinking** variants are trained with both SFT and DPO, this **instruct** model is trained using SFT only, without DPO. + + +For practical usage examples and detailed instructions on how to use the models, please also refer to our [cookbook](https://github.com/llm-jp/llm-jp-4-cookbook). + +To support the continued development of LLM-jp, we would greatly appreciate it if you could share how you utilize LLM-jp outcomes via the [survey form](https://forms.gle/AvbNXTNT2ADsssHq5). + + +## Usage + +Please refer to our [cookbook](https://github.com/llm-jp/llm-jp-4-cookbook) for practical usage examples and detailed instructions on how to use the models. + +## Model Details + +- **Model type:** Transformer-based Language Model +- **Architectures:** + +Dense model: +|Params|Layers|Hidden size|Heads|Context length|Embedding parameters|Non-embedding parameters|Total parameters| +|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:| +|8B|32|4,096|32|65,536|805,306,368|7,784,894,464|8,590,200,832| + +MoE model: +|Params|Layers|Hidden size|Heads|Routed Experts|Activated Experts|Context length|Embedding parameters|Non-embedding parameters|Activated parameters|Total parameters| +|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:| +|32B-A3B|32|2,560|40|128|8|65,536|503,316,480|31,635,712,512|3,827,476,992|32,139,028,992| + + +## Tokenizer + + +The tokenizer of this model is based on [huggingface/tokenizers](https://github.com/huggingface/tokenizers) Unigram byte-fallback model. +The vocabulary entries were converted from [`llm-jp-tokenizer v4.0`](https://github.com/llm-jp/llm-jp-tokenizer). +Please refer to [README.md](https://github.com/llm-jp/llm-jp-tokenizer) of `llm-jp-tokenizer` for details on the vocabulary construction procedure (the pure SentencePiece training does not reproduce our vocabulary). + +> [!NOTE] +> The chat template of this model is designed to be compatible with the OpenAI Harmony response format. +> However, the tokenizer differs from the one assumed by the `openai-harmony` library, and therefore direct tokenization with `openai-harmony` is not supported. +> For correct behavior, please use the tokenizer provided with this model. For detailed usage, please refer to [our cookbook](https://github.com/llm-jp/llm-jp-4-cookbook). + + +## Training + +### Pre-training + +This model is trained through a multi-stage pipeline consisting of pre-training and mid-training phases, using a total of 11.7T tokens. + +![pretraining_overview](./v4_pretraining_overview.png) + +The corpora used for pre-training and mid-training are publicly available at the following links: +- [Pre-training](https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4.1) +- [Mid-training](https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-midtraining-v2) + +> [!NOTE] +> Although most of the corpora have been released, some portions are excluded from public release due to licensing constraints. + +### Post-training + +We have fine-tuned the pre-trained checkpoint using SFT and further aligned it with DPO. + +The datasets used for post-training are also publicly available at the following links: +- [SFT](https://huggingface.co/datasets/llm-jp/llm-jp-4-thinking-sft-data) +- [DPO (for llm-jp-4-8b-thinking model)](https://huggingface.co/datasets/llm-jp/llm-jp-4-8b-thinking-dpo-data) +- [DPO (for llm-jp-4-32b-a3b-thinking model)](https://huggingface.co/datasets/llm-jp/llm-jp-4-32b-a3b-thinking-dpo-data) + +## Evaluation + +### [llm-jp-judge](https://github.com/llm-jp/llm-jp-judge) + +We evaluated the model on a variety of tasks using an LLM-as-a-Judge framework. The descriptions of each task are as follows. + +- MT-Bench (JA/EN): A benchmark for measuring multi-turn conversational task-solving ability. +- [AnswerCarefully](https://huggingface.co/datasets/llm-jp/AnswerCarefully): A benchmark for evaluating safety in Japanese. We used 336 questions from the v2.0 test set. +- [llm-jp-instructions](https://huggingface.co/datasets/llm-jp/llm-jp-instructions): A set of human-created single-turn question–answer pairs. We used 400 questions from the test set. + +We evaluated the models using `gpt-5.4-2026-03-05`. +> [!NOTE] +> Note: In earlier evaluations of the llm-jp-3 series, we used `gpt-4o-2024-08-06`. The newer evaluator `gpt-5.4-2026-03-05` provides a stricter and more reliable assessment, which results in lower scores on benchmarks such as MT-Bench compared to those reported for the llm-jp-3 series. + +The scores represent the average values obtained from three rounds of inference and evaluation. +For more details, please refer to the [codes](https://github.com/llm-jp/llm-jp-judge). + + +| Model Name | MT-Bench (JA) | MT-Bench (EN) | AnswerCarefully | llm-jp-instructions | +|:-------------------------------------------------------------------------------------------------------|----:|----:|----------------:|--------------------:| +| gpt-4o-2024-08-06 | 7.29 | 7.69 | 4.00 | 4.07 | +| gpt-5.4-2026-03-05 (reasoning_effort = low) | 8.87 | 8.76 | 4.38 | 4.79 | +| gpt-5.4-2026-03-05 (reasoning_effort = medium) | 8.87 | 8.89 | 4.43 | 4.82 | +| gpt-5.4-2026-03-05 (reasoning_effort = high) | 8.98 | 8.85 | 4.41 | 4.83 | +| [gpt-oss-20b (reasoning_effort = low)](https://huggingface.co/openai/gpt-oss-20b) | 7.21 | 7.95 | 3.39 | 3.08 | +| [gpt-oss-20b (reasoning_effort = medium)](https://huggingface.co/openai/gpt-oss-20b) | 7.33 | 7.85 | 3.55 | 3.16 | +| [llm-jp-4-8b-thinking (reasoning_effort = low)](https://huggingface.co/llm-jp/llm-jp-4-8b-thinking) | 7.23 | 7.54 | 3.58 | 3.50 | +| [llm-jp-4-8b-thinking (reasoning_effort = medium)](https://huggingface.co/llm-jp/llm-jp-4-8b-thinking) | 7.54 | 7.79 | 3.69 | 3.54 | +| **[llm-jp-4-8b-thinking-gguf Q4_K_M (reasoning_effort = medium)](https://huggingface.co/llm-jp/llm-jp-4-8b-thinking-gguf)** | 7.57 | 7.66 | 3.71 | 3.58 | +| [llm-jp-4-32b-a3b-thinking (reasoning_effort = low)](https://huggingface.co/llm-jp/llm-jp-4-32b-a3b-thinking) | 7.57 | 7.70 | 3.61 | 3.61 | +| [llm-jp-4-32b-a3b-thinking (reasoning_effort = medium)](https://huggingface.co/llm-jp/llm-jp-4-32b-a3b-thinking) | 7.82 | 7.86 | 3.70 | 3.61 | +| **[llm-jp-4-32b-a3b-thinking-gguf Q4_K_M (reasoning_effort = medium)](https://huggingface.co/llm-jp/llm-jp-4-32b-a3b-thinking-gguf)** | 7.51 | 7.89 | 3.74 | 3.70 | + +## Risks and Limitations + +The models released here are in the early stages of our research and development and have not been tuned to ensure outputs align with human intent and safety considerations. + + +## Send Questions to + +llm-jp(at)nii.ac.jp + + +## License + +[Apache License, Version 2.0](https://www.apache.org/licenses/LICENSE-2.0) + + +## Acknowledgement + +To develop this model, we used the NINJAL Web Japanese Corpus (whole-NWJC) from the National Institute for Japanese Language and Linguistics (NINJAL). + + +## Model Card Authors + +*The names are listed in alphabetical order.* + +Hirokazu Kiyomaru and Takashi Kodama. \ No newline at end of file diff --git a/imatrix.dat b/imatrix.dat new file mode 100644 index 0000000..b8e592c --- /dev/null +++ b/imatrix.dat @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:6936bf50ccd39e449ec78630efcbb8e80acfc47f56f73683caf7ecf910b24c91 +size 4988172 diff --git a/llm-jp-4-8b-thinking-BF16.gguf b/llm-jp-4-8b-thinking-BF16.gguf new file mode 100644 index 0000000..55780ae --- /dev/null +++ b/llm-jp-4-8b-thinking-BF16.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:8dee5e1671b332ab11f5e85b4089f77901d13b78a40f954a3b2e9be16ec50ca6 +size 17185772320 diff --git a/llm-jp-4-8b-thinking-Q4_K_M.gguf b/llm-jp-4-8b-thinking-Q4_K_M.gguf new file mode 100644 index 0000000..201a81b --- /dev/null +++ b/llm-jp-4-8b-thinking-Q4_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b8a1b3ef962d95dfaa02c241e0d2376e028aaddcda83191cc232e95ed2b14ae8 +size 5304882240 diff --git a/v4_pretraining_overview.png b/v4_pretraining_overview.png new file mode 100644 index 0000000..3a22686 --- /dev/null +++ b/v4_pretraining_overview.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:24c21c60ca5a0b5c9a65841efd9ad98344dd66d4ac7f1fe80ea5100afefe40ef +size 281924