初始化项目，由ModelHub XC社区提供模型

Model: Open-Orca/LlongOrca-13B-16k Source: Original Platform
2026-05-28 07:12:12 +08:00
commit ee9fbd648d
28 changed files with 25513 additions and 0 deletions
--- a/.gitattributes
+++ b/.gitattributes
@@ -0,0 +1,62 @@
+*.7z filter=lfs diff=lfs merge=lfs -text
+*.arrow filter=lfs diff=lfs merge=lfs -text
+ 
+ 
+*.bz2 filter=lfs diff=lfs merge=lfs -text
+*.ftz filter=lfs diff=lfs merge=lfs -text
+*.gz filter=lfs diff=lfs merge=lfs -text
+*.h5 filter=lfs diff=lfs merge=lfs -text
+*.joblib filter=lfs diff=lfs merge=lfs -text
+*.lfs.* filter=lfs diff=lfs merge=lfs -text
+ 
+*.msgpack filter=lfs diff=lfs merge=lfs -text
+*.onnx filter=lfs diff=lfs merge=lfs -text
+*.ot filter=lfs diff=lfs merge=lfs -text
+*.parquet filter=lfs diff=lfs merge=lfs -text
+*.pb filter=lfs diff=lfs merge=lfs -text
+ 
+ 
+*.rar filter=lfs diff=lfs merge=lfs -text
+saved_model/**/* filter=lfs diff=lfs merge=lfs -text
+*.tar.* filter=lfs diff=lfs merge=lfs -text
+*.tflite filter=lfs diff=lfs merge=lfs -text
+*.tgz filter=lfs diff=lfs merge=lfs -text
+*.xz filter=lfs diff=lfs merge=lfs -text
+*.zip filter=lfs diff=lfs merge=lfs -text
+*.zstandard filter=lfs diff=lfs merge=lfs -text
+*.tfevents* filter=lfs diff=lfs merge=lfs -text
+*.db* filter=lfs diff=lfs merge=lfs -text
+*.ark* filter=lfs diff=lfs merge=lfs -text
+**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
+**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
+**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
+*.safetensors filter=lfs diff=lfs merge=lfs -text
+*.ckpt filter=lfs diff=lfs merge=lfs -text
+*.gguf* filter=lfs diff=lfs merge=lfs -text
+*.ggml filter=lfs diff=lfs merge=lfs -text
+*.llamafile* filter=lfs diff=lfs merge=lfs -text
+*.pt2 filter=lfs diff=lfs merge=lfs -text
+*.mlmodel filter=lfs diff=lfs merge=lfs -text
+*.npy filter=lfs diff=lfs merge=lfs -text
+*.npz filter=lfs diff=lfs merge=lfs -text
+*.pickle filter=lfs diff=lfs merge=lfs -text
+*.pkl filter=lfs diff=lfs merge=lfs -text
+*.tar filter=lfs diff=lfs merge=lfs -text
+*.wasm filter=lfs diff=lfs merge=lfs -text
+*.zst filter=lfs diff=lfs merge=lfs -text
+*tfevents* filter=lfs diff=lfs merge=lfs -text
+
+checkpoint-4096/rng_state_0.pth filter=lfs diff=lfs merge=lfs -text
+checkpoint-4096/rng_state_6.pth filter=lfs diff=lfs merge=lfs -text
+checkpoint-4096/rng_state_5.pth filter=lfs diff=lfs merge=lfs -text
+pytorch_model.bin filter=lfs diff=lfs merge=lfs -text
+checkpoint-4096/rng_state_2.pth filter=lfs diff=lfs merge=lfs -text
+checkpoint-4096/rng_state_4.pth filter=lfs diff=lfs merge=lfs -text
+checkpoint-4096/rng_state_1.pth filter=lfs diff=lfs merge=lfs -text
+checkpoint-4096/training_args.bin filter=lfs diff=lfs merge=lfs -text
+checkpoint-4096/rng_state_7.pth filter=lfs diff=lfs merge=lfs -text
+checkpoint-4096/pytorch_model.bin filter=lfs diff=lfs merge=lfs -text
+training_args.bin filter=lfs diff=lfs merge=lfs -text
+checkpoint-4096/tokenizer.model filter=lfs diff=lfs merge=lfs -text
+tokenizer.model filter=lfs diff=lfs merge=lfs -text
+checkpoint-4096/rng_state_3.pth filter=lfs diff=lfs merge=lfs -text
--- a/Images/LlongOrca13BG4A.png
+++ b/Images/LlongOrca13BG4A.png
--- a/Images/LlongOrca13BHFLeaderboard.png
+++ b/Images/LlongOrca13BHFLeaderboard.png
--- a/README.md
+++ b/README.md
@@ -0,0 +1,145 @@
+---
+license: llama2
+language:
+- en
+library_name: transformers
+pipeline_tag: text-generation
+datasets:
+- Open-Orca/OpenOrca
+---
+
+<p><h1>🐋 The Second Llong Context Orca! 🐋</h1></p>
+
+
+![OpenOrca Logo](https://huggingface.co/datasets/Open-Orca/OpenOrca/resolve/main/OpenOrcaLogo.png "OpenOrca Logo")
+
+
+# OpenOrca - LlongOrca - 13B - 16k
+
+We have used our own [OpenOrca dataset](https://huggingface.co/datasets/Open-Orca/OpenOrca) to fine-tune on top of [LLongMA-2-13b-16k](https://huggingface.co/conceptofmind/LLongMA-2-13b-16k). 
+This dataset is our attempt to reproduce the dataset generated for Microsoft Research's [Orca Paper](https://arxiv.org/abs/2306.02707).
+We use [OpenChat](https://huggingface.co/openchat) packing, trained with [Axolotl](https://github.com/OpenAccess-AI-Collective/axolotl).
+
+This release is trained on a curated filtered subset of most of our GPT-4 augmented data.
+It is the same subset of our data as was used in our [OpenOrcaxOpenChat-Preview2-13B model](https://huggingface.co/Open-Orca/OpenOrcaxOpenChat-Preview2-13B).
+
+HF Leaderboard evals place this model as #1 for all 13B long context models at release time.
+We achieve >112% the performance of the base LLongMA2-13b-16k model we tuned on top of.
+As well, we preserve >98% of the performance of the OpenOrcaxOpenChat-Preview2-13B model we share datasets with, while extending the context to 16k.
+
+We did this training as part of testing setup of our H100 cluster.
+
+Want to visualize our full (pre-filtering) dataset? Check out our [Nomic Atlas Map](https://atlas.nomic.ai/map/c1b88b47-2d9b-47e0-9002-b80766792582/2560fd25-52fe-42f1-a58f-ff5eccc890d2).
+
+
+[<img src="https://huggingface.co/Open-Orca/OpenOrca-Preview1-13B/resolve/main/OpenOrca%20Nomic%20Atlas.png" alt="Atlas Nomic Dataset Map" width="400" height="400" />](https://atlas.nomic.ai/map/c1b88b47-2d9b-47e0-9002-b80766792582/2560fd25-52fe-42f1-a58f-ff5eccc890d2)
+
+
+Many thanks to @EnricoShippole, @theemozilla, and @kaiokendev1 for the fine work on creating the LlongMA-2-13b-16k model this was trained on top of!
+
+We are in-process with training more models, so keep a look out on our org for releases coming soon with exciting partners.
+
+We will also give sneak-peak announcements on our Discord, which you can find here:
+
+https://AlignmentLab.ai
+
+# Prompt Template
+
+We used [OpenAI's Chat Markup Language (ChatML)](https://github.com/openai/openai-python/blob/main/chatml.md) format, with `<|im_start|>` and `<|im_end|>` tokens added to support this.
+
+## Example Prompt Exchange
+
+```
+<|im_start|>system
+You are LlongOrca, a large language model trained by Alignment Lab AI. Write out your reasoning step-by-step to be sure you get the right answers!
+<|im_end|>
+<|im_start|>user
+How are you<|im_end|>
+<|im_start|>assistant
+I am doing well!<|im_end|>
+<|im_start|>user
+How are you now?<|im_end|>
+```
+
+
+# Evaluation
+
+We have evaluated using the methodology and tools for the HuggingFace Leaderboard, and find that we have significantly improved upon the base long context model.
+We reach >112% of LLongMA2-13B-16k performance.
+
+## HuggingFaceH4 Open LLM Leaderboard Performance
+
+We have run our own tests using parameters matching the [HuggingFaceH4 Open LLM Leaderboard](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard) evals.
+
+We preserve >98% of OpenOrcaxOpenChat-Preview2-13B performance and are #1 on the leaderboard for long context 13B models at release time.
+We have >103% performance of the next 16k model (vicuna-13b-v1.5-16k).
+
+As well, we expect the context extension techniques from LLongMA to be more robust than other 16k context models available.
+
+![LlongOrca 13B 16k HF Leaderboard Internal Performance](https://huggingface.co/Open-Orca/LlongOrca-13B-16k/resolve/main/Images/LlongOrca13BHFLeaderboard.png "HuggingFace Leaderboard Internal Performance")
+
+
+## GPT4ALL Leaderboard Performance
+
+We find we score higher than all non-OpenOrca models on the GPT4ALL leaderboard, while preserving ~98.7% of our OpenOrcaxOpenChat-Preview2-13B performance.
+
+![LLongOrca 13B 16k GPT4ALL Leaderboard Internal Performance](https://huggingface.co/Open-Orca/LlongOrca-13B-16k/resolve/main/Images/LlongOrca13BG4A.png "GPT4ALL Leaderboard Internal Performance")
+
+
+# Dataset
+
+We used a curated, filtered selection of most of the GPT-4 augmented data from our OpenOrca dataset, which aims to reproduce the Orca Research Paper dataset.
+Further details of our curation practices will be forthcoming with our full model releases.
+
+
+# Training
+
+[<img src="https://raw.githubusercontent.com/OpenAccess-AI-Collective/axolotl/main/image/axolotl-badge-web.png" alt="Built with Axolotl"/>](https://github.com/OpenAccess-AI-Collective/axolotl)
+
+We trained with 8x H100 GPUs for 10 hours, completing 4 epochs of full fine tuning on our dataset in one training run.
+Commodity cost was ~$300.
+
+# Citation
+
+```bibtex
+@software{dale2023llongorca13b,
+  title = {LlongOrca13B: Llama2-13B Model Instruct-tuned for Long Context on Filtered OpenOrcaV1 GPT-4 Dataset},
+  author = {Alpin Dale and Wing Lian and Bleys Goodson and Guan Wang and Eugene Pentland and Austin Cook and Chanvichet Vong and "Teknium"},
+  year = {2023},
+  publisher = {HuggingFace},
+  journal = {HuggingFace repository},
+  howpublished = {\url{https://https://huggingface.co/Open-Orca/LlongOrca-7B-16k},
+}
+@software{openchat,
+  title = {{OpenChat: Advancing Open-source Language Models with Imperfect Data}},
+  author = {Wang, Guan and Cheng, Sijie and Yu, Qiying and Liu, Changling},
+  doi = {10.5281/zenodo.8105775},
+  url = {https://github.com/imoneoi/openchat},
+  version = {pre-release},
+  year = {2023},
+  month = {7},
+}
+@misc{mukherjee2023orca,
+      title={Orca: Progressive Learning from Complex Explanation Traces of GPT-4}, 
+      author={Subhabrata Mukherjee and Arindam Mitra and Ganesh Jawahar and Sahaj Agarwal and Hamid Palangi and Ahmed Awadallah},
+      year={2023},
+      eprint={2306.02707},
+      archivePrefix={arXiv},
+      primaryClass={cs.CL}
+}
+@misc{longpre2023flan,
+      title={The Flan Collection: Designing Data and Methods for Effective Instruction Tuning}, 
+      author={Shayne Longpre and Le Hou and Tu Vu and Albert Webson and Hyung Won Chung and Yi Tay and Denny Zhou and Quoc V. Le and Barret Zoph and Jason Wei and Adam Roberts},
+      year={2023},
+      eprint={2301.13688},
+      archivePrefix={arXiv},
+      primaryClass={cs.AI}
+}
+@misc{touvron2023llama,
+    title={Llama 2: Open Foundation and Fine-Tuned Chat Models}, 
+    author={Hugo Touvron and Louis Martin and Kevin Stone and Peter Albert and Amjad Almahairi and Yasmine Babaei and Nikolay Bashlykov and Soumya Batra and Prajjwal Bhargava and Shruti Bhosale and Dan Bikel and Lukas Blecher and Cristian Canton Ferrer and Moya Chen and Guillem Cucurull and David Esiobu and Jude Fernandes and Jeremy Fu and Wenyin Fu and Brian Fuller and Cynthia Gao and Vedanuj Goswami and Naman Goyal and Anthony Hartshorn and Saghar Hosseini and Rui Hou and Hakan Inan and Marcin Kardas and Viktor Kerkez and Madian Khabsa and Isabel Kloumann and Artem Korenev and Punit Singh Koura and Marie-Anne Lachaux and Thibaut Lavril and Jenya Lee and Diana Liskovich and Yinghai Lu and Yuning Mao and Xavier Martinet and Todor Mihaylov and Pushkar Mishra and Igor Molybog and Yixin Nie and Andrew Poulton and Jeremy Reizenstein and Rashi Rungta and Kalyan Saladi and Alan Schelten and Ruan Silva and Eric Michael Smith and Ranjan Subramanian and Xiaoqing Ellen Tan and Binh Tang and Ross Taylor and Adina Williams and Jian Xiang Kuan and Puxin Xu and Zheng Yan and Iliyan Zarov and Yuchen Zhang and Angela Fan and Melanie Kambadur and Sharan Narang and Aurelien Rodriguez and Robert Stojnic and Sergey Edunov and Thomas Scialom},
+    year={2023},
+    eprint={2307.09288},
+    archivePrefix={arXiv},
+}
+```
--- a/added_tokens.json
+++ b/added_tokens.json
@@ -0,0 +1,6 @@
+{
+  "<pad>": 32000,
+  "<|im_end|>": 32002,
+  "<|im_start|>": 32001,
+  "<|system|>": 32003
+}
--- a/checkpoint-4096/added_tokens.json
+++ b/checkpoint-4096/added_tokens.json
@@ -0,0 +1,6 @@
+{
+  "<pad>": 32000,
+  "<|im_end|>": 32002,
+  "<|im_start|>": 32001,
+  "<|system|>": 32003
+}
--- a/checkpoint-4096/config.json
+++ b/checkpoint-4096/config.json
@@ -0,0 +1,30 @@
+{
+  "_name_or_path": "conceptofmind/LLongMA-2-13b-16k",
+  "architectures": [
+    "LlamaForCausalLM"
+  ],
+  "bos_token_id": 1,
+  "eos_token_id": 2,
+  "hidden_act": "silu",
+  "hidden_size": 5120,
+  "initializer_range": 0.02,
+  "intermediate_size": 13824,
+  "max_position_embeddings": 16384,
+  "model_type": "llama",
+  "num_attention_heads": 40,
+  "num_hidden_layers": 40,
+  "num_key_value_heads": 40,
+  "pad_token_id": 0,
+  "pretraining_tp": 2,
+  "rms_norm_eps": 1e-05,
+  "rope_scaling": {
+    "factor": 4.0,
+    "type": "linear"
+  },
+  "tie_word_embeddings": false,
+  "torch_dtype": "bfloat16",
+  "transformers_version": "4.32.0.dev0",
+  "use_cache": false,
+  "use_flash_attention": false,
+  "vocab_size": 32004
+}
--- a/checkpoint-4096/pytorch_model.bin
+++ b/checkpoint-4096/pytorch_model.bin
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:ec463df6ae295815a321ebd8b54ffa11ceb9c965dc5c46a1008c8f08685f1052
+size 26031995033
--- a/checkpoint-4096/rng_state_0.pth
+++ b/checkpoint-4096/rng_state_0.pth
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:b64bd1bb3ba8b0a3e03ea6b9ea79188f2e9aa0b1127982f65e7910552a683009
+size 21687
--- a/checkpoint-4096/rng_state_1.pth
+++ b/checkpoint-4096/rng_state_1.pth
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:4599ed04a3ea8005c93981da3ac51adab6f9e48356f416faf5999160f70aa612
+size 21687
--- a/checkpoint-4096/rng_state_2.pth
+++ b/checkpoint-4096/rng_state_2.pth
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:9e480e19d8c9c591d8b3285867ca3c557cad3b54a5a11f8ddba4cc58128ca39e
+size 21687
--- a/checkpoint-4096/rng_state_3.pth
+++ b/checkpoint-4096/rng_state_3.pth
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:96b235d1282fb055977ff5416b4461dd8d7116aaa6764bcf768e6e190331a636
+size 21687
--- a/checkpoint-4096/rng_state_4.pth
+++ b/checkpoint-4096/rng_state_4.pth
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:17a98e2d1f43b20a0b49ae2d540712c0a057b1c7cae69c4542940db7fc7fa55f
+size 21687
--- a/checkpoint-4096/rng_state_5.pth
+++ b/checkpoint-4096/rng_state_5.pth
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:ee9f8a10bed33741098006f1f70d09d2438a9eecd8b2fd59a0883798ed3f57fc
+size 21687
--- a/checkpoint-4096/rng_state_6.pth
+++ b/checkpoint-4096/rng_state_6.pth
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:9f6b1fae8cf3c4a63e1b8163a1c2366dd0498c12ab762c3d827578047c24211c
+size 21687
--- a/checkpoint-4096/rng_state_7.pth
+++ b/checkpoint-4096/rng_state_7.pth
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:20e1a1bd086523383052a57e4cbe614821a157b38c27f20a13779cb885a5affb
+size 21687
--- a/checkpoint-4096/special_tokens_map.json
+++ b/checkpoint-4096/special_tokens_map.json
@@ -0,0 +1,6 @@
+{
+  "bos_token": "<s>",
+  "eos_token": "</s>",
+  "pad_token": "[PAD]",
+  "unk_token": "<unk>"
+}
--- a/checkpoint-4096/tokenizer.model
+++ b/checkpoint-4096/tokenizer.model
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:9e556afd44213b6bd1be2b850ebbbd98f5481437a8021afaf58ee7fb1818d347
+size 499723
--- a/checkpoint-4096/tokenizer_config.json
+++ b/checkpoint-4096/tokenizer_config.json
@@ -0,0 +1,36 @@
+{
+  "add_bos_token": true,
+  "add_eos_token": false,
+  "bos_token": {
+    "__type": "AddedToken",
+    "content": "<s>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "clean_up_tokenization_spaces": false,
+  "eos_token": {
+    "__type": "AddedToken",
+    "content": "</s>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "legacy": true,
+  "model_max_length": 8192,
+  "pad_token": null,
+  "sp_model_kwargs": {},
+  "tokenizer_class": "LlamaTokenizer",
+  "trust_remote_code": false,
+  "unk_token": {
+    "__type": "AddedToken",
+    "content": "<unk>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "use_fast": true
+}
--- a/checkpoint-4096/trainer_state.json
+++ b/checkpoint-4096/trainer_state.json
--- a/checkpoint-4096/training_args.bin
+++ b/checkpoint-4096/training_args.bin
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:0510726cb65aed3a7adb37f055e2bda836342c6718bba46783cb3c0db4a8a4ec
+size 4347
--- a/config.json
+++ b/config.json
@@ -0,0 +1,30 @@
+{
+  "_name_or_path": "conceptofmind/LLongMA-2-13b-16k",
+  "architectures": [
+    "LlamaForCausalLM"
+  ],
+  "bos_token_id": 1,
+  "eos_token_id": 2,
+  "hidden_act": "silu",
+  "hidden_size": 5120,
+  "initializer_range": 0.02,
+  "intermediate_size": 13824,
+  "max_position_embeddings": 16384,
+  "model_type": "llama",
+  "num_attention_heads": 40,
+  "num_hidden_layers": 40,
+  "num_key_value_heads": 40,
+  "pad_token_id": 0,
+  "pretraining_tp": 2,
+  "rms_norm_eps": 1e-05,
+  "rope_scaling": {
+    "factor": 4.0,
+    "type": "linear"
+  },
+  "tie_word_embeddings": false,
+  "torch_dtype": "bfloat16",
+  "transformers_version": "4.32.0.dev0",
+  "use_cache": false,
+  "use_flash_attention": false,
+  "vocab_size": 32004
+}
--- a/configuration.json
+++ b/configuration.json
@@ -0,0 +1 @@
+{"framework": "pytorch", "task": "text-generation", "allow_remote": true}
--- a/pytorch_model.bin
+++ b/pytorch_model.bin
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:ec463df6ae295815a321ebd8b54ffa11ceb9c965dc5c46a1008c8f08685f1052
+size 26031995033
--- a/special_tokens_map.json
+++ b/special_tokens_map.json
@@ -0,0 +1,6 @@
+{
+  "bos_token": "<s>",
+  "eos_token": "</s>",
+  "pad_token": "[PAD]",
+  "unk_token": "<unk>"
+}
--- a/tokenizer.model
+++ b/tokenizer.model
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:9e556afd44213b6bd1be2b850ebbbd98f5481437a8021afaf58ee7fb1818d347
+size 499723
--- a/tokenizer_config.json
+++ b/tokenizer_config.json
@@ -0,0 +1,36 @@
+{
+  "add_bos_token": true,
+  "add_eos_token": false,
+  "bos_token": {
+    "__type": "AddedToken",
+    "content": "<s>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "clean_up_tokenization_spaces": false,
+  "eos_token": {
+    "__type": "AddedToken",
+    "content": "</s>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "legacy": true,
+  "model_max_length": 8192,
+  "pad_token": null,
+  "sp_model_kwargs": {},
+  "tokenizer_class": "LlamaTokenizer",
+  "trust_remote_code": false,
+  "unk_token": {
+    "__type": "AddedToken",
+    "content": "<unk>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "use_fast": true
+}
--- a/training_args.bin
+++ b/training_args.bin
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:0510726cb65aed3a7adb37f055e2bda836342c6718bba46783cb3c0db4a8a4ec
+size 4347
				`@@ -0,0 +1 @@`
				`{"framework": "pytorch", "task": "text-generation", "allow_remote": true}`