初始化项目,由ModelHub XC社区提供模型

Model: abhinand/tamil-llama-13b-base-v0.1
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-06-07 10:38:16 +08:00
commit a6de4709a0
22 changed files with 717 additions and 0 deletions

35
.gitattributes vendored Normal file
View File

@@ -0,0 +1,35 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

177
README.md Normal file
View File

@@ -0,0 +1,177 @@
---
language:
- ta
- en
license: llama2
model-index:
- name: tamil-llama-13b-base-v0.1
results:
- task:
type: text-generation
name: Text Generation
dataset:
name: AI2 Reasoning Challenge (25-Shot)
type: ai2_arc
config: ARC-Challenge
split: test
args:
num_few_shot: 25
metrics:
- type: acc_norm
value: 52.82
name: normalized accuracy
source:
url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=abhinand/tamil-llama-13b-base-v0.1
name: Open LLM Leaderboard
- task:
type: text-generation
name: Text Generation
dataset:
name: HellaSwag (10-Shot)
type: hellaswag
split: validation
args:
num_few_shot: 10
metrics:
- type: acc_norm
value: 79.95
name: normalized accuracy
source:
url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=abhinand/tamil-llama-13b-base-v0.1
name: Open LLM Leaderboard
- task:
type: text-generation
name: Text Generation
dataset:
name: MMLU (5-Shot)
type: cais/mmlu
config: all
split: test
args:
num_few_shot: 5
metrics:
- type: acc
value: 52.05
name: accuracy
source:
url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=abhinand/tamil-llama-13b-base-v0.1
name: Open LLM Leaderboard
- task:
type: text-generation
name: Text Generation
dataset:
name: TruthfulQA (0-shot)
type: truthful_qa
config: multiple_choice
split: validation
args:
num_few_shot: 0
metrics:
- type: mc2
value: 36.56
source:
url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=abhinand/tamil-llama-13b-base-v0.1
name: Open LLM Leaderboard
- task:
type: text-generation
name: Text Generation
dataset:
name: Winogrande (5-shot)
type: winogrande
config: winogrande_xl
split: validation
args:
num_few_shot: 5
metrics:
- type: acc
value: 75.61
name: accuracy
source:
url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=abhinand/tamil-llama-13b-base-v0.1
name: Open LLM Leaderboard
- task:
type: text-generation
name: Text Generation
dataset:
name: GSM8k (5-shot)
type: gsm8k
config: main
split: test
args:
num_few_shot: 5
metrics:
- type: acc
value: 0.0
name: accuracy
source:
url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=abhinand/tamil-llama-13b-base-v0.1
name: Open LLM Leaderboard
---
# Tamil LLaMA 13B Base v0.1 [pre-trained]
Welcome to the inaugural release of the Tamil LLaMA 13B base model an important step in advancing LLMs for the Tamil language. This model is ready for immediate inference and is also primed for further fine-tuning to cater to your specific NLP tasks.
To dive deep into the development and capabilities of this model, please read the [research paper](https://arxiv.org/abs/2311.05845) and the [introductory blog post (WIP)]() that outlines our journey and the model's potential impact.
> **Please Note:** This model, labeled as a foundational Tamil Language Model (LLM), is designed primarily for Causal Language Modeling (LM) purposes. In other words, if you are looking for an instruction following model in Tamil, you may find [abhinand/tamil-llama-13b-instruct-v0.1](https://huggingface.co/abhinand/tamil-llama-13b-instruct-v0.1) more suitable for your needs.
## Model description
The Tamil LLaMA models have been enhanced and tailored specifically with an extensive Tamil vocabulary of 16,000 tokens, building upon the foundation set by the original LLaMA-2.
- **Model type:** A 13B parameter model for Causal LM pre-trained on [CulturaX](https://huggingface.co/datasets/uonlp/CulturaX) dataset's Tamil subset.
- **Language(s):** Tamil and English
- **License:** GNU General Public License v3.0
- **Source Model:** [meta-llama/Llama-2-13b-hf](https://huggingface.co/meta-llama/Llama-2-13b-hf)
- **Training Precision:** `float16`
- **Code:** [GitHub](https://github.com/abhinand5/tamil-llama)
## Related Models
| Model | Type | Data | Base Model | # Params | Download Links |
|--------------------------|-----------------------------|-------------------|----------------------|------|------------------------------------------------------------------------|
| Tamil LLaMA 7B Base | Base model | 12GB | LLaMA 7B | 7B | [HF Hub](https://huggingface.co/abhinand/tamil-llama-7b-base-v0.1) |
| Tamil LLaMA 13B Base | Base model | 4GB | LLaMA 13B | 13B | [HF Hub](https://huggingface.co/abhinand/tamil-llama-13b-base-v0.1) |
| Tamil LLaMA 7B Instruct | Instruction following model | 145k instructions | Tamil LLaMA 7B Base | 7B | [HF Hub](https://huggingface.co/abhinand/tamil-llama-7b-instruct-v0.1) |
| Tamil LLaMA 13B Instruct | Instruction following model | 145k instructions | Tamil LLaMA 13B Base | 13B | [HF Hub](abhinand/tamil-llama-13b-instruct-v0.1) |
## Usage Note
It's important to note that the models have not undergone detoxification. Therefore, while they possess impressive linguistic capabilities, there is a possibility for them to generate content that could be deemed harmful or offensive. We urge users to exercise discretion and supervise the model's outputs closely, especially in public or sensitive applications.
## Meet the Developers
Get to know the creators behind this innovative model and follow their contributions to the field:
- [Abhinand Balachandran](https://www.linkedin.com/in/abhinand-05/)
## Citation
If you use this model or any of the the Tamil-Llama datasets in your research, please cite:
```bibtex
@misc{balachandran2023tamilllama,
title={Tamil-Llama: A New Tamil Language Model Based on Llama 2},
author={Abhinand Balachandran},
year={2023},
eprint={2311.05845},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
```
We hope this model serves as a valuable tool in your NLP toolkit and look forward to seeing the advancements it will enable in the understanding and generation of the Tamil language.
# [Open LLM Leaderboard Evaluation Results](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard)
Detailed results can be found [here](https://huggingface.co/datasets/open-llm-leaderboard/details_abhinand__tamil-llama-13b-base-v0.1)
| Metric |Value|
|---------------------------------|----:|
|Avg. |49.50|
|AI2 Reasoning Challenge (25-Shot)|52.82|
|HellaSwag (10-Shot) |79.95|
|MMLU (5-Shot) |52.05|
|TruthfulQA (0-shot) |36.56|
|Winogrande (5-shot) |75.61|
|GSM8k (5-shot) | 0.00|

25
config.json Normal file
View File

@@ -0,0 +1,25 @@
{
"_name_or_path": "abhinand/tamil-llama-13b-base-v0.1",
"architectures": [
"LlamaForCausalLM"
],
"bos_token_id": 1,
"eos_token_id": 2,
"hidden_act": "silu",
"hidden_size": 5120,
"initializer_range": 0.02,
"intermediate_size": 13824,
"max_position_embeddings": 4096,
"model_type": "llama",
"num_attention_heads": 40,
"num_hidden_layers": 40,
"num_key_value_heads": 40,
"pretraining_tp": 1,
"rms_norm_eps": 1e-05,
"rope_scaling": null,
"tie_word_embeddings": false,
"torch_dtype": "bfloat16",
"transformers_version": "4.32.0.dev0",
"use_cache": true,
"vocab_size": 47957
}

6
generation_config.json Normal file
View File

@@ -0,0 +1,6 @@
{
"_from_model_config": true,
"bos_token_id": 1,
"eos_token_id": 2,
"transformers_version": "4.32.0.dev0"
}

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5beb0390578ff6aba448269a036430c3d5f46be7711fe42d550a56fb36373346
size 2111178687

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ac6b6494c06396e891ca5d78af624df15006b7e1af7ccab625268720c312a6b5
size 1903235957

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:28e8d6459cd9165e5b22cb6366a9204c0dc18948c9528bfe0dff7b276e4d0f00
size 1903235957

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0853e2a396eec9f75c5455660e268300a4d869b1e24e319077c7aa3cdda62e8d
size 1903235957

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e9ca6f885ffb36faba79148be2c4d2705f3bcaaada7ab1df828315acac5669a7
size 1903235957

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:84260a7059fb6219d695dc2f9ad424f4ac2e2934d1f4f0f596f2b53a59ffca0a
size 1903235957

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4c1ab7939f39b83eec11cb1d9c7f36d9635045738e390c03d237f05e60b38d7f
size 1903235957

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3af4419f4adf0ab68189d490b74e22632da8d06334170d5f5074a62d77528885
size 1903235957

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:999e90a03621fa8cb0115e3522c913ba90b3d4b638e4e9f14219294620be1fab
size 1903235957

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b2f79fe7666abbf02751d7da291fddc14b176ab96a335082822e3256c0e2272d
size 1903235957

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bae6a19412f3474bacd7e86f1e1b45a3897337b1a833384c238d67620abb41f8
size 1903235957

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c94679936ad29f85228c2953c42c26e82b2400e2a871afab9b093307473c533d
size 1903235957

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:25f3272c2f1b4bcb934f2d60b460a533e5e12924bc7695ba0bc8443beb1d86cb
size 1903235957

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:58686d633bf173405841455aba93c3e22ad6b36e21cea7d918695ffddba716fc
size 1408640157

View File

@@ -0,0 +1,370 @@
{
"metadata": {
"total_size": 26358528000
},
"weight_map": {
"lm_head.weight": "pytorch_model-00014-of-00014.bin",
"model.embed_tokens.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.0.input_layernorm.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.0.mlp.down_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.0.mlp.gate_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.0.mlp.up_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.0.post_attention_layernorm.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.0.self_attn.k_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.0.self_attn.o_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.0.self_attn.q_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.0.self_attn.v_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.1.input_layernorm.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.1.mlp.down_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.1.mlp.gate_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.1.mlp.up_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.1.post_attention_layernorm.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.1.self_attn.k_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.1.self_attn.o_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.1.self_attn.q_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.1.self_attn.v_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.10.input_layernorm.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.10.mlp.down_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.10.mlp.gate_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.10.mlp.up_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.10.post_attention_layernorm.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.10.self_attn.k_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.10.self_attn.o_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.10.self_attn.q_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.10.self_attn.v_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.11.input_layernorm.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.11.mlp.down_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.11.mlp.gate_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.11.mlp.up_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.11.post_attention_layernorm.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.11.self_attn.k_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.11.self_attn.o_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.11.self_attn.q_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.11.self_attn.v_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.12.input_layernorm.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.12.mlp.down_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.12.mlp.gate_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.12.mlp.up_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.12.post_attention_layernorm.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.12.self_attn.k_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.12.self_attn.o_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.12.self_attn.q_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.12.self_attn.v_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.13.input_layernorm.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.13.mlp.down_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.13.mlp.gate_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.13.mlp.up_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.13.post_attention_layernorm.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.13.self_attn.k_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.13.self_attn.o_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.13.self_attn.q_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.13.self_attn.v_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.14.input_layernorm.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.14.mlp.down_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.14.mlp.gate_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.14.mlp.up_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.14.post_attention_layernorm.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.14.self_attn.k_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.14.self_attn.o_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.14.self_attn.q_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.14.self_attn.v_proj.weight": "pytorch_model-00005-of-00014.bin",
"model.layers.15.input_layernorm.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.15.mlp.down_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.15.mlp.gate_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.15.mlp.up_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.15.post_attention_layernorm.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.15.self_attn.k_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.15.self_attn.o_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.15.self_attn.q_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.15.self_attn.v_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.16.input_layernorm.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.16.mlp.down_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.16.mlp.gate_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.16.mlp.up_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.16.post_attention_layernorm.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.16.self_attn.k_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.16.self_attn.o_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.16.self_attn.q_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.16.self_attn.v_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.17.input_layernorm.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.17.mlp.down_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.17.mlp.gate_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.17.mlp.up_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.17.post_attention_layernorm.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.17.self_attn.k_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.17.self_attn.o_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.17.self_attn.q_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.17.self_attn.v_proj.weight": "pytorch_model-00006-of-00014.bin",
"model.layers.18.input_layernorm.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.18.mlp.down_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.18.mlp.gate_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.18.mlp.up_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.18.post_attention_layernorm.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.18.self_attn.k_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.18.self_attn.o_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.18.self_attn.q_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.18.self_attn.v_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.19.input_layernorm.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.19.mlp.down_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.19.mlp.gate_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.19.mlp.up_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.19.post_attention_layernorm.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.19.self_attn.k_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.19.self_attn.o_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.19.self_attn.q_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.19.self_attn.v_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.2.input_layernorm.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.2.mlp.down_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.2.mlp.gate_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.2.mlp.up_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.2.post_attention_layernorm.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.2.self_attn.k_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.2.self_attn.o_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.2.self_attn.q_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.2.self_attn.v_proj.weight": "pytorch_model-00001-of-00014.bin",
"model.layers.20.input_layernorm.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.20.mlp.down_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.20.mlp.gate_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.20.mlp.up_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.20.post_attention_layernorm.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.20.self_attn.k_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.20.self_attn.o_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.20.self_attn.q_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.20.self_attn.v_proj.weight": "pytorch_model-00007-of-00014.bin",
"model.layers.21.input_layernorm.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.21.mlp.down_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.21.mlp.gate_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.21.mlp.up_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.21.post_attention_layernorm.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.21.self_attn.k_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.21.self_attn.o_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.21.self_attn.q_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.21.self_attn.v_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.22.input_layernorm.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.22.mlp.down_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.22.mlp.gate_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.22.mlp.up_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.22.post_attention_layernorm.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.22.self_attn.k_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.22.self_attn.o_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.22.self_attn.q_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.22.self_attn.v_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.23.input_layernorm.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.23.mlp.down_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.23.mlp.gate_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.23.mlp.up_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.23.post_attention_layernorm.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.23.self_attn.k_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.23.self_attn.o_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.23.self_attn.q_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.23.self_attn.v_proj.weight": "pytorch_model-00008-of-00014.bin",
"model.layers.24.input_layernorm.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.24.mlp.down_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.24.mlp.gate_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.24.mlp.up_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.24.post_attention_layernorm.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.24.self_attn.k_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.24.self_attn.o_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.24.self_attn.q_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.24.self_attn.v_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.25.input_layernorm.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.25.mlp.down_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.25.mlp.gate_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.25.mlp.up_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.25.post_attention_layernorm.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.25.self_attn.k_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.25.self_attn.o_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.25.self_attn.q_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.25.self_attn.v_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.26.input_layernorm.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.26.mlp.down_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.26.mlp.gate_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.26.mlp.up_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.26.post_attention_layernorm.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.26.self_attn.k_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.26.self_attn.o_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.26.self_attn.q_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.26.self_attn.v_proj.weight": "pytorch_model-00009-of-00014.bin",
"model.layers.27.input_layernorm.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.27.mlp.down_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.27.mlp.gate_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.27.mlp.up_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.27.post_attention_layernorm.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.27.self_attn.k_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.27.self_attn.o_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.27.self_attn.q_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.27.self_attn.v_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.28.input_layernorm.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.28.mlp.down_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.28.mlp.gate_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.28.mlp.up_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.28.post_attention_layernorm.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.28.self_attn.k_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.28.self_attn.o_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.28.self_attn.q_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.28.self_attn.v_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.29.input_layernorm.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.29.mlp.down_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.29.mlp.gate_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.29.mlp.up_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.29.post_attention_layernorm.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.29.self_attn.k_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.29.self_attn.o_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.29.self_attn.q_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.29.self_attn.v_proj.weight": "pytorch_model-00010-of-00014.bin",
"model.layers.3.input_layernorm.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.3.mlp.down_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.3.mlp.gate_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.3.mlp.up_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.3.post_attention_layernorm.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.3.self_attn.k_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.3.self_attn.o_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.3.self_attn.q_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.3.self_attn.v_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.30.input_layernorm.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.30.mlp.down_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.30.mlp.gate_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.30.mlp.up_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.30.post_attention_layernorm.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.30.self_attn.k_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.30.self_attn.o_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.30.self_attn.q_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.30.self_attn.v_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.31.input_layernorm.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.31.mlp.down_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.31.mlp.gate_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.31.mlp.up_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.31.post_attention_layernorm.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.31.self_attn.k_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.31.self_attn.o_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.31.self_attn.q_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.31.self_attn.v_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.32.input_layernorm.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.32.mlp.down_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.32.mlp.gate_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.32.mlp.up_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.32.post_attention_layernorm.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.32.self_attn.k_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.32.self_attn.o_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.32.self_attn.q_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.32.self_attn.v_proj.weight": "pytorch_model-00011-of-00014.bin",
"model.layers.33.input_layernorm.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.33.mlp.down_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.33.mlp.gate_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.33.mlp.up_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.33.post_attention_layernorm.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.33.self_attn.k_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.33.self_attn.o_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.33.self_attn.q_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.33.self_attn.v_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.34.input_layernorm.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.34.mlp.down_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.34.mlp.gate_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.34.mlp.up_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.34.post_attention_layernorm.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.34.self_attn.k_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.34.self_attn.o_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.34.self_attn.q_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.34.self_attn.v_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.35.input_layernorm.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.35.mlp.down_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.35.mlp.gate_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.35.mlp.up_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.35.post_attention_layernorm.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.35.self_attn.k_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.35.self_attn.o_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.35.self_attn.q_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.35.self_attn.v_proj.weight": "pytorch_model-00012-of-00014.bin",
"model.layers.36.input_layernorm.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.36.mlp.down_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.36.mlp.gate_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.36.mlp.up_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.36.post_attention_layernorm.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.36.self_attn.k_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.36.self_attn.o_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.36.self_attn.q_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.36.self_attn.v_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.37.input_layernorm.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.37.mlp.down_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.37.mlp.gate_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.37.mlp.up_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.37.post_attention_layernorm.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.37.self_attn.k_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.37.self_attn.o_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.37.self_attn.q_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.37.self_attn.v_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.38.input_layernorm.weight": "pytorch_model-00014-of-00014.bin",
"model.layers.38.mlp.down_proj.weight": "pytorch_model-00014-of-00014.bin",
"model.layers.38.mlp.gate_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.38.mlp.up_proj.weight": "pytorch_model-00014-of-00014.bin",
"model.layers.38.post_attention_layernorm.weight": "pytorch_model-00014-of-00014.bin",
"model.layers.38.self_attn.k_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.38.self_attn.o_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.38.self_attn.q_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.38.self_attn.v_proj.weight": "pytorch_model-00013-of-00014.bin",
"model.layers.39.input_layernorm.weight": "pytorch_model-00014-of-00014.bin",
"model.layers.39.mlp.down_proj.weight": "pytorch_model-00014-of-00014.bin",
"model.layers.39.mlp.gate_proj.weight": "pytorch_model-00014-of-00014.bin",
"model.layers.39.mlp.up_proj.weight": "pytorch_model-00014-of-00014.bin",
"model.layers.39.post_attention_layernorm.weight": "pytorch_model-00014-of-00014.bin",
"model.layers.39.self_attn.k_proj.weight": "pytorch_model-00014-of-00014.bin",
"model.layers.39.self_attn.o_proj.weight": "pytorch_model-00014-of-00014.bin",
"model.layers.39.self_attn.q_proj.weight": "pytorch_model-00014-of-00014.bin",
"model.layers.39.self_attn.v_proj.weight": "pytorch_model-00014-of-00014.bin",
"model.layers.4.input_layernorm.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.4.mlp.down_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.4.mlp.gate_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.4.mlp.up_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.4.post_attention_layernorm.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.4.self_attn.k_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.4.self_attn.o_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.4.self_attn.q_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.4.self_attn.v_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.5.input_layernorm.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.5.mlp.down_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.5.mlp.gate_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.5.mlp.up_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.5.post_attention_layernorm.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.5.self_attn.k_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.5.self_attn.o_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.5.self_attn.q_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.5.self_attn.v_proj.weight": "pytorch_model-00002-of-00014.bin",
"model.layers.6.input_layernorm.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.6.mlp.down_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.6.mlp.gate_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.6.mlp.up_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.6.post_attention_layernorm.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.6.self_attn.k_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.6.self_attn.o_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.6.self_attn.q_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.6.self_attn.v_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.7.input_layernorm.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.7.mlp.down_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.7.mlp.gate_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.7.mlp.up_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.7.post_attention_layernorm.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.7.self_attn.k_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.7.self_attn.o_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.7.self_attn.q_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.7.self_attn.v_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.8.input_layernorm.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.8.mlp.down_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.8.mlp.gate_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.8.mlp.up_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.8.post_attention_layernorm.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.8.self_attn.k_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.8.self_attn.o_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.8.self_attn.q_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.8.self_attn.v_proj.weight": "pytorch_model-00003-of-00014.bin",
"model.layers.9.input_layernorm.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.9.mlp.down_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.9.mlp.gate_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.9.mlp.up_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.9.post_attention_layernorm.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.9.self_attn.k_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.9.self_attn.o_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.9.self_attn.q_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.layers.9.self_attn.v_proj.weight": "pytorch_model-00004-of-00014.bin",
"model.norm.weight": "pytorch_model-00014-of-00014.bin"
}
}

24
special_tokens_map.json Normal file
View File

@@ -0,0 +1,24 @@
{
"bos_token": {
"content": "<s>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
},
"eos_token": {
"content": "</s>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
},
"pad_token": "</s>",
"unk_token": {
"content": "<unk>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
}
}

3
tokenizer.model Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d9b79bfcedafc76d8e643b3a4f5e2dd284b8334794664b18f0bb30447ad47926
size 1004267

35
tokenizer_config.json Normal file
View File

@@ -0,0 +1,35 @@
{
"add_bos_token": true,
"add_eos_token": false,
"bos_token": {
"__type": "AddedToken",
"content": "<s>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
},
"clean_up_tokenization_spaces": false,
"eos_token": {
"__type": "AddedToken",
"content": "</s>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
},
"legacy": true,
"model_max_length": 1000000000000000019884624838656,
"pad_token": null,
"sp_model_kwargs": {},
"tokenizer_class": "LlamaTokenizer",
"unk_token": {
"__type": "AddedToken",
"content": "<unk>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false
},
"use_fast": true
}