初始化项目，由ModelHub XC社区提供模型

Model: speakleash/Bielik-7B-v0.1 Source: Original Platform
2026-05-14 02:11:26 +08:00
commit 3e55d1127f
17 changed files with 91780 additions and 0 deletions
--- a/.gitattributes
+++ b/.gitattributes
@@ -0,0 +1,35 @@
 *.7z filter=lfs diff=lfs merge=lfs -text
 *.arrow filter=lfs diff=lfs merge=lfs -text
 *.bin filter=lfs diff=lfs merge=lfs -text
 *.bz2 filter=lfs diff=lfs merge=lfs -text
 *.ckpt filter=lfs diff=lfs merge=lfs -text
 *.ftz filter=lfs diff=lfs merge=lfs -text
 *.gz filter=lfs diff=lfs merge=lfs -text
 *.h5 filter=lfs diff=lfs merge=lfs -text
 *.joblib filter=lfs diff=lfs merge=lfs -text
 *.lfs.* filter=lfs diff=lfs merge=lfs -text
 *.mlmodel filter=lfs diff=lfs merge=lfs -text
 *.model filter=lfs diff=lfs merge=lfs -text
 *.msgpack filter=lfs diff=lfs merge=lfs -text
 *.npy filter=lfs diff=lfs merge=lfs -text
 *.npz filter=lfs diff=lfs merge=lfs -text
 *.onnx filter=lfs diff=lfs merge=lfs -text
 *.ot filter=lfs diff=lfs merge=lfs -text
 *.parquet filter=lfs diff=lfs merge=lfs -text
 *.pb filter=lfs diff=lfs merge=lfs -text
 *.pickle filter=lfs diff=lfs merge=lfs -text
 *.pkl filter=lfs diff=lfs merge=lfs -text
 *.pt filter=lfs diff=lfs merge=lfs -text
 *.pth filter=lfs diff=lfs merge=lfs -text
 *.rar filter=lfs diff=lfs merge=lfs -text
 *.safetensors filter=lfs diff=lfs merge=lfs -text
 saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.tar.* filter=lfs diff=lfs merge=lfs -text
 *.tar filter=lfs diff=lfs merge=lfs -text
 *.tflite filter=lfs diff=lfs merge=lfs -text
 *.tgz filter=lfs diff=lfs merge=lfs -text
 *.wasm filter=lfs diff=lfs merge=lfs -text
 *.xz filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
--- a/README.md
+++ b/README.md
@@ -0,0 +1,232 @@
 ---
 license: apache-2.0
 language:
 - pl
 library_name: transformers
 tags:
 - continuously_pretrained
 inference:
  parameters:
    temperature: 0.7
 ---
 <p align="center">
  <img src="https://huggingface.co/speakleash/Bielik-7B-v0.1/raw/main/speakleash_cyfronet.png">
 </p>
 # Bielik-7B-v0.1
 The Bielik-7B-v0.1 is a generative text model featuring 7 billion parameters, meticulously evolved from its predecessor, the [Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1), through processing of over 70 billion tokens. Forementioned model stands as a testament to the unique collaboration between the open-science/open-source project SpeakLeash and the High Performance Computing (HPC) center: ACK Cyfronet AGH. Developed and trained on Polish text corpora, which have been cherry-picked and processed by the SpeakLeash team, this endeavor leverages Polish large-scale computing infrastructure, specifically within the PLGrid environment, and more precisely, the HPC centers: ACK Cyfronet AGH. The creation and training of the Bielik-7B-v0.1 was propelled by the support of computational grant number PLG/2024/016951, conducted on the Athena and Helios supercomputers, enabling the use of cutting-edge technology and computational resources essential for large-scale machine learning processes. As a result, the model exhibits an exceptional ability to understand and process the Polish language, providing accurate responses and performing a variety of linguistic tasks with high precision.
 ⚠️ This is a base model intended for further fine-tuning across most use cases. If you're looking for a model ready for chatting or following instructions out-of-the-box, please use [Bielik-7B-Instruct-v0.1](https://huggingface.co/speakleash/Bielik-7B-Instruct-v0.1).
 🎥 Demo: https://huggingface.co/spaces/speakleash/Bielik-7B-Instruct-v0.1
 🗣️ Chat Arena<span style="color:red;">*</span>: https://arena.speakleash.org.pl/
 <span style="color:red;">*</span>Chat Arena is a platform for testing and comparing different AI language models, allowing users to evaluate their performance and quality.
 ## Model
 Bielik-7B-v0.1 has been trained with the use of an original open source framework called [ALLaMo](https://github.com/chrisociepa/allamo) implemented by [Krzysztof Ociepa](https://www.linkedin.com/in/krzysztof-ociepa-44886550/). This framework allows users to train language models with architecture similar to LLaMA and Mistral in fast and efficient way.
 The model training was conducted on the Helios Supercomputer at the ACK Cyfronet AGH, utilizing 256 NVidia GH200 cards while achieving a throughput exceeding 9200 tokens/gpu/second.
 The training dataset was composed of Polish texts collected and made available through the [SpeakLeash](https://speakleash.org/) project. We used over 36 billion tokens for two epochs of training.
 ### Model description:
 * **Developed by:** [SpeakLeash](https://speakleash.org/)
 * **Language:** Polish
 * **Model type:** causal decoder-only
 * **Adopted from:** [Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1)
 * **License:** Apache 2.0 (commercial use allowed)
 * **Model ref:** speakleash:debfc8635c781358e8db833a333887a5
 ### Quality evaluation
 A XGBoost classification model was prepared and created to evaluate the quality of texts in native Polish language. It is based on 93 features, such as the ratio of out-of-vocabulary words to all words (OOVs), the number of nouns, verbs,  average sentence length etc. The model outputs the category of a given document (either HIGH, MEDIUM or LOW) along with the  probability. This approach allows implementation of dedicated pipeline to choose documents, from which we've used entries with HIGH quality index and probability exceeding 90%.
 This filtration and appropriate selection of texts enable the provision of a condensed and high-quality database of texts in Polish for training purposes.
 ## Training
 * Framework: [ALLaMo](https://github.com/chrisociepa/allamo)
 * Visualizations: [W&B](https://wandb.ai)
 <p align="center">
  <img src="https://huggingface.co/speakleash/Bielik-7B-v0.1/raw/main/train_loss.png">
 </p>
 <p align="center">
  <img src="https://huggingface.co/speakleash/Bielik-7B-v0.1/raw/main/train_ppl.png">
 </p>
 <p align="center">
  <img src="https://huggingface.co/speakleash/Bielik-7B-v0.1/raw/main/train_acc.png">
 </p>
 ### Training hyperparameters:
 | **Hyperparameter**          | **Value**        |
 |-----------------------------|------------------|
 | Context length              | 4096             |
 | Micro Batch Size            | 4                |
 | Batch Size                  | 4194304          |
 | Learning Rate (cosine)      | 3e-05 -> 2e-05   |
 | Warmup Iterations           | 2000             |
 | All Iterations              | 17350            |
 | Optimizer                   | AdamW            |
 | β1, β2                      | 0.9, 0.95        |
 | Adam_eps                    | 1e−8             |
 | Weight Decay                | 0.1              |
 | Grad Clip                   | 1.0              |
 | Precision                   | bfloat16 (mixed) |
 ### Quickstart
 This model can be easily loaded using the AutoModelForCausalLM functionality.
 ```python
 from transformers import AutoTokenizer, AutoModelForCausalLM
 model_name = "speakleash/Bielik-7B-v0.1"
 tokenizer = AutoTokenizer.from_pretrained(model_name)
 model = AutoModelForCausalLM.from_pretrained(model_name)
 ```
 In order to reduce the memory usage, you can use smaller precision (`bfloat16`).
 ```python
 import torch
 model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.bfloat16)
 ```
 And then you can use Hugging Face Pipelines to generate text:
 ```python
 import transformers
 text = "Najważniejszym celem człowieka na ziemi jest"
 pipeline = transformers.pipeline("text-generation", model=model, tokenizer=tokenizer)
 sequences = pipeline(max_new_tokens=100, do_sample=True, top_k=50, eos_token_id=tokenizer.eos_token_id, text_inputs=text)
 for seq in sequences:
    print(f"Result: {seq['generated_text']}")
 ```
 Generated output:
 > Najważniejszym celem człowieka na ziemi jest życie w pokoju, harmonii i miłości. Dla każdego z nas bardzo ważne jest, aby otaczać się kochanymi osobami.
 ## Evaluation
 Models have been evaluated on [Open PL LLM Leaderboard](https://huggingface.co/spaces/speakleash/open_pl_llm_leaderboard) 5-shot. The benchmark evaluates models in NLP tasks like sentiment analysis, categorization, text classification but does not test chatting skills. Here are presented:           
 - Average - average score among all tasks normalized by baseline scores         
 - Reranking - reranking task, commonly used in RAG          
 - Reader (Generator) - open book question answering task, commonly used in RAG          
 - Perplexity (lower is better) - as a bonus, does not correlate with other scores and should not be used for model comparison           
 As of April 3, 2024, the following table showcases the current scores of pretrained and continuously pretrained models according to the Open PL LLM Leaderboard, evaluated in a 5-shot setting:
 |                                                                                      |   Average | RAG Reranking | RAG Reader | Perplexity |
 |--------------------------------------------------------------------------------------|----------:|--------------:|-----------:|-----------:|
 | **7B parameters models:**                                                            |           |               |            |            |
 | Baseline (majority class)                                                            |      0.00 |         53.36 |          - |          - |
 | OPI-PG/Qra-7b                                                                        |     11.13 |         54.40 |      75.25 |     203.36 |
 | meta-llama/Llama-2-7b-hf                                                             |     12.73 |         54.02 |      77.92 |     850.45 |
 | internlm/internlm2-base-7b                                                           |     20.68 |         52.39 |      69.85 |    3110.92 |
 | [Bielik-7B-v0.1](https://huggingface.co/speakleash/Bielik-7B-v0.1)                   |     29.38 |     **62.13** |  **88.39** |     123.31 |
 | mistralai/Mistral-7B-v0.1                                                            |     30.67 |         60.35 |      85.39 |     857.32 |
 | internlm/internlm2-7b                                                                |     33.03 |         69.39 |      73.63 |    5498.23 |
 | alpindale/Mistral-7B-v0.2-hf                                                         |     33.05 |         60.23 |      85.21 |     932.60 |
 | speakleash/mistral-apt3-7B/spi-e0_hf (experimental)                                  | **35.50** |     **62.14** |      87.48 |     132.78 |
 |                                                                                      |           |               |            |            |
 | **Models with different sizes:**                                                     |           |               |            |            |
 | sdadas/polish-gpt2-xl (1.7B)                                                         |    -23.22 |         48.07 |       3.04 |     160.95 |
 | Azurro/APT3-1B-Base (1B)                                                             |     -8.23 |         51.49 |      18.94 |     249.90 |
 | OPI-PG/Qra-1b (1B)                                                                   |     -5.44 |         47.65 |      38.51 |     398.96 |
 | internlm/internlm2-1_8b (1.8B)                                                       |     -2.78 |         49.37 |      31.88 |   60296.30 |
 | OPI-PG/Qra-13b (13B)                                                                 |     29.03 |         53.28 |      83.03 |     168.66 |
 | upstage/SOLAR-10.7B-v1.0 (10.7B)                                                     |     38.12 |         75.81 |      86.39 |     641.05 |
 |                                                                                      |           |               |            |            |
 | **Polish instruction fine-tuned models:**                                            |           |               |            |            |
 | szymonrucinski/Curie-7B-v1                                                           |     26.72 |         55.58 |      85.19 |     389.17 |
 | Voicelab/trurl-2-7b                                                                  |     18.85 |         60.67 |      77.19 |    1098.88 |
 | [Bielik-7B-Instruct-v0.1](https://huggingface.co/speakleash/Bielik-7B-Instruct-v0.1) |     39.28 |         61.89 |      86.00 |     277.92 |
 As you can see, Bielik-7B-v0.1 does not have the best Average score, but it has some clear advantages, e.g. the best score in the RAG Reader task.
 The results in the above table were obtained without utilizing instruction templates for instructional models, instead treating them like base models. 
 This approach could skew the results, as instructional models are optimized with specific instructions in mind.
 ## Limitations and Biases
 Bielik-7B-v0.1 is not intended for deployment without fine-tuning. It should not be used for human-facing interactions without further guardrails and user consent.
 Bielik-7B-v0.1 can produce factually incorrect output, and should not be relied on to produce factually accurate data. Bielik-7B-v0.1 was trained on various public datasets. While great efforts have been taken to clear the training data, it is possible that this model can generate lewd, false, biased or otherwise offensive outputs.
 ## License
 The model is licensed under Apache 2.0, which allows for commercial use.
 ## Citation
 Please cite this model using the following format:
 ```
@misc{Bielik7Bv01,
    title     = {Introducing Bielik-7B-v0.1: Polish Language Model},
    author    = {Ociepa, Krzysztof and Flis, Łukasz and Wróbel, Krzysztof and Gwoździej, Adrian and {SpeakLeash Team} and {Cyfronet Team}},
    year      = {2024},
    url       = {https://huggingface.co/speakleash/Bielik-7B-v0.1},
    note      = {Accessed: 2024-04-01}, % change this date
    urldate   = {2024-04-01} % change this date
 }
@misc{ociepa2024bielik7bv01polish,
      title={Bielik 7B v0.1: A Polish Language Model -- Development, Insights, and Evaluation}, 
      author={Krzysztof Ociepa and Łukasz Flis and Krzysztof Wróbel and Adrian Gwoździej and Remigiusz Kinas},
      year={2024},
      eprint={2410.18565},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2410.18565}, 
 }
 ```
 ## Responsible for training the model
 * [Krzysztof Ociepa](https://www.linkedin.com/in/krzysztof-ociepa-44886550/)<sup>SpeakLeash</sup> - team leadership, conceptualizing, data preparation, process optimization and oversight of training
 * [Łukasz Flis](https://www.linkedin.com/in/lukasz-flis-0a39631/)<sup>Cyfronet AGH</sup> - coordinating and supervising the training
 * [Adrian Gwoździej](https://www.linkedin.com/in/adrgwo/)<sup>SpeakLeash</sup> - data cleaning and quality
 * [Krzysztof Wróbel](https://www.linkedin.com/in/wrobelkrzysztof/)<sup>SpeakLeash</sup> - benchmarks
 The model could not have been created without the commitment and work of the entire SpeakLeash team, whose contribution is invaluable. Thanks to the hard work of many individuals, it was possible to gather a large amount of content in Polish and establish collaboration between the open-science SpeakLeash project and the HPC center: ACK Cyfronet AGH. Individuals who contributed to the creation of the model through their commitment to the open-science SpeakLeash project:
 [Sebastian Kondracki](https://www.linkedin.com/in/sebastian-kondracki/),
 [Maria Filipkowska](https://www.linkedin.com/in/maria-filipkowska/),
 [Grzegorz Urbanowicz](https://www.linkedin.com/in/grzegorz-urbanowicz-05823469/),
 [Szymon Baczyński](https://www.linkedin.com/in/szymon-baczynski/),
 [Paweł Kiszczak](https://www.linkedin.com/in/paveu-kiszczak/),
 [Igor Ciuciura](https://www.linkedin.com/in/igor-ciuciura-1763b52a6/),
 [Paweł Cyrta](https://www.linkedin.com/in/cyrta),
 [Jacek Chwiła](https://www.linkedin.com/in/jacek-chwila/),
 [Jan Maria Kowalski](https://www.linkedin.com/in/janmariakowalski/),
 [Karol Jezierski](https://www.linkedin.com/in/karol-jezierski/),
 [Kamil Nonckiewicz](https://www.linkedin.com/in/kamil-nonckiewicz/),
 [Izabela Babis](https://www.linkedin.com/in/izabela-babis-2274b8105/),
 [Nina Babis](https://www.linkedin.com/in/nina-babis-00055a140/),
 [Waldemar Boszko](https://www.linkedin.com/in/waldemarboszko),
 [Remigiusz Kinas](https://www.linkedin.com/in/remigiusz-kinas/),
 [Piotr Rybak](https://www.linkedin.com/in/piotrrybak/)
 and many other wonderful researchers and enthusiasts of the AI world.
 Members of the ACK Cyfronet AGH team providing valuable support and expertise:
 [Szymon Mazurek](https://www.linkedin.com/in/sz-mazurek-ai/).
 ## Contact Us
 If you have any questions or suggestions, please use the discussion tab. If you want to contact us directly, join our [Discord SpeakLeash](https://discord.gg/3G9DVM39).
--- a/config.json
+++ b/config.json
@@ -0,0 +1,25 @@
 {
  "architectures": [
    "MistralForCausalLM"
  ],
  "attention_dropout": 0.0,
  "bos_token_id": 1,
  "eos_token_id": 2,
  "hidden_act": "silu",
  "hidden_size": 4096,
  "initializer_range": 0.02,
  "intermediate_size": 14336,
  "max_position_embeddings": 8192,
  "model_type": "mistral",
  "num_attention_heads": 32,
  "num_hidden_layers": 32,
  "num_key_value_heads": 8,
  "rms_norm_eps": 1e-05,
  "rope_theta": 10000.0,
  "sliding_window": 4096,
  "tie_word_embeddings": false,
  "torch_dtype": "bfloat16",
  "transformers_version": "4.37.2",
  "use_cache": true,
  "vocab_size": 32000
 }
--- a/generation_config.json
+++ b/generation_config.json
@@ -0,0 +1,6 @@
 {
  "_from_model_config": true,
  "bos_token_id": 1,
  "eos_token_id": 2,
  "transformers_version": "4.37.2"
 }
--- a/info.json
+++ b/info.json
@@ -0,0 +1,3 @@
 {
  "reference_id": "speakleash:debfc8635c781358e8db833a333887a5"
 }
--- a/model-00001-of-00003.safetensors
+++ b/model-00001-of-00003.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:722569584d43bea3fb7f535f1ffa1140074c82969957b8be6aad8e1d20a7c8bc
 size 4943162336
--- a/model-00002-of-00003.safetensors
+++ b/model-00002-of-00003.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:96515ece2dcdf1a9a9d448a49415fbdc20a07fdc851fe868f0610ffa9aeebb63
 size 4999819336
--- a/model-00003-of-00003.safetensors
+++ b/model-00003-of-00003.safetensors
@@ -0,0 +1,3 @@
 version https://git-lfs.github.com/spec/v1
 oid sha256:4052007de417d16487a71076680296ec9e3975f3a4cf69c3fda6469c768acfbb
 size 4540516344
--- a/model.safetensors.index.json
+++ b/model.safetensors.index.json
@@ -0,0 +1,298 @@
 {
  "metadata": {
    "total_size": 14483464192
  },
  "weight_map": {
    "lm_head.weight": "model-00003-of-00003.safetensors",
    "model.embed_tokens.weight": "model-00001-of-00003.safetensors",
    "model.layers.0.input_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.0.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.0.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.0.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.0.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.0.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.0.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.0.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.0.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.1.input_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.1.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.1.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.1.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.1.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.1.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.1.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.1.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.1.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.10.input_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.10.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.10.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.10.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.10.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.10.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.10.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.10.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.10.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.11.input_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.11.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.11.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.11.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.11.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.11.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.11.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.11.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.11.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.12.input_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.12.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.12.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.12.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.12.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.12.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.12.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.12.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.12.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.13.input_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.13.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.13.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.13.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.13.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.13.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.13.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.13.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.13.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.14.input_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.14.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.14.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.14.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.14.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.14.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.14.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.14.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.14.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.15.input_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.15.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.15.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.15.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.15.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.15.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.15.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.15.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.15.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.16.input_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.16.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.16.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.16.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.16.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.16.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.16.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.16.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.16.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.17.input_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.17.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.17.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.17.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.17.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.17.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.17.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.17.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.17.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.18.input_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.18.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.18.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.18.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.18.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.18.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.18.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.18.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.18.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.19.input_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.19.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.19.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.19.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.19.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.19.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.19.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.19.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.19.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.2.input_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.2.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.2.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.2.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.2.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.2.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.2.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.2.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.2.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.20.input_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.20.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.20.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.20.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.20.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.20.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.20.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.20.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.20.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.21.input_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.21.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.21.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.21.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.21.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
    "model.layers.21.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.21.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.21.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.21.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.22.input_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.22.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.22.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.22.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.22.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.22.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.22.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.22.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.22.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
    "model.layers.23.input_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.23.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.23.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.23.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.23.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.23.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.23.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.23.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.23.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.24.input_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.24.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.24.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.24.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.24.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.24.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.24.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.24.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.24.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.25.input_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.25.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.25.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.25.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.25.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.25.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.25.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.25.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.25.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.26.input_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.26.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.26.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.26.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.26.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.26.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.26.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.26.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.26.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.27.input_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.27.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.27.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.27.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.27.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.27.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.27.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.27.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.27.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.28.input_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.28.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.28.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.28.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.28.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.28.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.28.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.28.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.28.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.29.input_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.29.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.29.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.29.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.29.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.29.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.29.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.29.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.29.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.3.input_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.3.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.3.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.3.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.3.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.3.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.3.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.3.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.3.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.30.input_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.30.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.30.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.30.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.30.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.30.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.30.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.30.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.30.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.31.input_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.31.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.31.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.31.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.31.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
    "model.layers.31.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.31.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.31.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.31.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
    "model.layers.4.input_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.4.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.4.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.4.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.4.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.4.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.4.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.4.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.4.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.5.input_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.5.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.5.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.5.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.5.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.5.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.5.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.5.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.5.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.6.input_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.6.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.6.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.6.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.6.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.6.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.6.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.6.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.6.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.7.input_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.7.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.7.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.7.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.7.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.7.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.7.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.7.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.7.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.8.input_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.8.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.8.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.8.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.8.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.8.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.8.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.8.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.8.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.9.input_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.9.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.9.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.9.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.9.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
    "model.layers.9.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.9.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.9.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
    "model.layers.9.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
    "model.norm.weight": "model-00003-of-00003.safetensors"
  }
 }
--- a/speakleash_cyfronet.png
+++ b/speakleash_cyfronet.png
--- a/special_tokens_map.json
+++ b/special_tokens_map.json
@@ -0,0 +1,5 @@
 {
  "bos_token": "<s>",
  "eos_token": "</s>",
  "unk_token": "<unk>"
 }
--- a/tokenizer.json
+++ b/tokenizer.json
--- a/tokenizer.model
+++ b/tokenizer.model
--- a/tokenizer_config.json
+++ b/tokenizer_config.json
@@ -0,0 +1,42 @@
 {
  "add_bos_token": true,
  "add_eos_token": false,
  "added_tokens_decoder": {
    "0": {
      "content": "<unk>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "1": {
      "content": "<s>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    },
    "2": {
      "content": "</s>",
      "lstrip": false,
      "normalized": false,
      "rstrip": false,
      "single_word": false,
      "special": true
    }
  },
  "additional_special_tokens": [],
  "bos_token": "<s>",
  "clean_up_tokenization_spaces": false,
  "eos_token": "</s>",
  "legacy": true,
  "model_max_length": 1000000000000000019884624838656,
  "pad_token": null,
  "sp_model_kwargs": {},
  "spaces_between_special_tokens": false,
  "tokenizer_class": "LlamaTokenizer",
  "unk_token": "<unk>",
  "use_default_system_prompt": false
 }
--- a/train_acc.png
+++ b/train_acc.png
--- a/train_loss.png
+++ b/train_loss.png
--- a/train_ppl.png
+++ b/train_ppl.png