初始化项目,由ModelHub XC社区提供模型

Model: Mungert/Llama-Guard-3-8B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-09 18:28:12 +08:00
commit 533164e511
36 changed files with 1091 additions and 0 deletions

81
.gitattributes vendored Normal file
View File

@@ -0,0 +1,81 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-iq2_xxs.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-f16-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-iq1_s.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-iq1_m.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-iq2_m.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-q4_k_s.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-bf16-q6_k.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-iq3_s.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B.imatrix filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-iq2_xs.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-iq2_s.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-bf16-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-iq4_nl.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-q2_k_s.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-q3_k_s.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-q3_k_m.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-q5_0.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-q4_1.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-f16-q6_k.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-q5_k_s.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-q5_1.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-q4_k_m.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-iq3_xxs.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-iq3_m.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-iq3_xs.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-f16-q4_k.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-bf16.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-iq4_xs.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-bf16-q4_k.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-q5_k_m.gguf filter=lfs diff=lfs merge=lfs -text
Llama-Guard-3-8B-q6_k_m.gguf filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c10fcc621a86923440cc181e4112c73997e26bf3b5c0cf483fd2786d32b90519
size 6295640480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:080dec932696702a5fcef4abbc04b0ff5dfed7460b3c3854038ad20bd2ccd480
size 7835474336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:44969dfdca49b1b1de5062ee59841525459c8f325c54b93f9290a82aba1ec293
size 9525778592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fa6f2458033ea3b53783e47090332e5ff6f9d3ea8d18484fcefe614b711c01f3
size 16068892832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e4cccf28586d4fb261762128116da639ef28a4d079c2361081d063b5d18c59d0
size 6295640480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3a63479d4ebd1bb3860b8edec0f85a4fa0e85558e58d72d7a58018011df43852
size 7835474336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d077c976c4c63a164e3d2d6aa04ca3146d1f5602a798ed309153b821764e2576
size 9525778592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:574c77275b827cc998e1ccd5aa7557c1ba4ae38cb6981dbf04628a341c6c083f
size 2161973664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c7bf91f059749c3894d6fc2a38d91e888a180811bc0c7efd5caf804a3354a533
size 2019629472

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:04baa1de6e4379fd1e3c1344d3fa7f861b0c3958e0b6daadcc7c58d53660e84a
size 2948282784

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5bd1e6c1e7ed7d669490a23fbc44eedfe0b47478be52508e68f842433ec7d40c
size 2758490528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2a256fc74cd1cde99e30704c254b0f0f89da697cc183c5bd1a9efbb567f6444e
size 2605783456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9b60c808f2e3cdf52326e1c31f4e42b3c158d85e021cf731594fd573508ccdca
size 2399213984

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c9242c38fff7762cfa2be7c9193f29dc4b4ae3788c953badd5e8aabe8606918a
size 3784825248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e4c9ebb11e13fa45e438afd32e83a45dda24a994a883b46a55aea0a45f26d474
size 3682326944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cf7c21c7732fffeff0e2988706f0270f67b2dd2d091e6c1e3899c90049ab4201
size 3518749088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7fcf8d666ee442aed5cc5ad9cefd81668adc82796539ad56946af6a50be5aecc
size 3274914208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f208fdca61efc3972de2c2dc8ab6250f8618f6f3cdad8d779585ab1ab60ef32e
size 4677990816

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:45a9e119f3672a18f897bb12c7c776fe15d4d7284b77ee30edae828f6d647e0b
size 4447664544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c6a6a030ae906320a1680918d86276dd6e9f8ced1b804859ced89803d2a659d4
size 2988816800

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a6c74b89849e811de70333da0d3a605d22ed88c4d98efa6be5fe9a34c53ea8ee
size 4018919840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:55d02cf1bcdf0cf54c7f3de6df5dc86c513bcaf44593d92b7228ea130e2b3f7d
size 3664501152

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:df09b7a6e3dff23dcf97446b44f48270f43be330fdc8d42f10d695c33b7c7745
size 4525775264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ed97ad4d56cf52c37cf07c174234bccbaf1959ccd4a6a1d8b485c5588fb13382
size 5027649952

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:20682acc45b27cab80c7c869afa33205d7c8ab73ba8ab2c4df11e3f1123b78e5
size 4920736160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4c427d8c36e7d2686fac3d17bbdfae694fb54cf71bd52e4965fac453f7a13861
size 4692670880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:67ab506a627c6cf1189df0e6b584c9f2beb5f162aad631ce0f554ee9fea39bf0
size 5529524640

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8914092a4cb1b2f0ed9cb849a6699855313692b0e47199d882cf384375503c71
size 6031399328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:383f965e05a132ed80f51aafad54052fb52fc230e822530efb0a1e9f1bdc40ca
size 5732989344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:613c672d154d07969fd9d5eba5e1f6e03529bdf7262768458e5f22f19473da2a
size 5599295904

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5e62aa36a0d890f6fed1f37e60c7585f4f496cf9eb9c2546d71714706672a4d2
size 6596008352

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5f6f4a79ba167ce1b4aa7c1b4fdaee8f044d49c837a81083d64b1ff452734ca2
size 8540772512

3
Llama-Guard-3-8B.imatrix Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:39fe6234593d4a4227addb473a29fe7ef964e5dd98396e16758f8b89ca639a2f
size 4988151

910
README.md Normal file
View File

@@ -0,0 +1,910 @@
---
language:
- en
pipeline_tag: text-generation
base_model: meta-llama/Meta-Llama-3.1-8B
tags:
- facebook
- meta
- pytorch
- llama
- llama-3
license: llama3.1
extra_gated_prompt: >-
### LLAMA 3.1 COMMUNITY LICENSE AGREEMENT
Llama 3.1 Version Release Date: July 23, 2024
"Agreement" means the terms and conditions for use, reproduction, distribution and modification of the
Llama Materials set forth herein.
"Documentation" means the specifications, manuals and documentation accompanying Llama 3.1
distributed by Meta at https://llama.meta.com/doc/overview.
"Licensee" or "you" means you, or your employer or any other person or entity (if you are entering into
this Agreement on such person or entitys behalf), of the age required under applicable laws, rules or
regulations to provide legal consent and that has legal authority to bind your employer or such other
person or entity if you are entering in this Agreement on their behalf.
"Llama 3.1" means the foundational large language models and software and algorithms, including
machine-learning model code, trained model weights, inference-enabling code, training-enabling code,
fine-tuning enabling code and other elements of the foregoing distributed by Meta at
https://llama.meta.com/llama-downloads.
"Llama Materials" means, collectively, Metas proprietary Llama 3.1 and Documentation (and any
portion thereof) made available under this Agreement.
"Meta" or "we" means Meta Platforms Ireland Limited (if you are located in or, if you are an entity, your
principal place of business is in the EEA or Switzerland) and Meta Platforms, Inc. (if you are located
outside of the EEA or Switzerland).
1. License Rights and Redistribution.
a. Grant of Rights. You are granted a non-exclusive, worldwide, non-transferable and royalty-free
limited license under Metas intellectual property or other rights owned by Meta embodied in the Llama
Materials to use, reproduce, distribute, copy, create derivative works of, and make modifications to the
Llama Materials.
b. Redistribution and Use.
i. If you distribute or make available the Llama Materials (or any derivative works
thereof), or a product or service (including another AI model) that contains any of them, you shall (A)
provide a copy of this Agreement with any such Llama Materials; and (B) prominently display “Built with
Llama” on a related website, user interface, blogpost, about page, or product documentation. If you use
the Llama Materials or any outputs or results of the Llama Materials to create, train, fine tune, or
otherwise improve an AI model, which is distributed or made available, you shall also include “Llama” at
the beginning of any such AI model name.
ii. If you receive Llama Materials, or any derivative works thereof, from a Licensee as part
of an integrated end user product, then Section 2 of this Agreement will not apply to you.
iii. You must retain in all copies of the Llama Materials that you distribute the following
attribution notice within a “Notice” text file distributed as a part of such copies: “Llama 3.1 is
licensed under the Llama 3.1 Community License, Copyright © Meta Platforms, Inc. All Rights
Reserved.”
iv. Your use of the Llama Materials must comply with applicable laws and regulations
(including trade compliance laws and regulations) and adhere to the Acceptable Use Policy for the Llama
Materials (available at https://llama.meta.com/llama3_1/use-policy), which is hereby incorporated by
reference into this Agreement.
2. Additional Commercial Terms. If, on the Llama 3.1 version release date, the monthly active users
of the products or services made available by or for Licensee, or Licensees affiliates, is greater than 700
million monthly active users in the preceding calendar month, you must request a license from Meta,
which Meta may grant to you in its sole discretion, and you are not authorized to exercise any of the
rights under this Agreement unless or until Meta otherwise expressly grants you such rights.
3. Disclaimer of Warranty. UNLESS REQUIRED BY APPLICABLE LAW, THE LLAMA MATERIALS AND ANY
OUTPUT AND RESULTS THEREFROM ARE PROVIDED ON AN “AS IS” BASIS, WITHOUT WARRANTIES OF
ANY KIND, AND META DISCLAIMS ALL WARRANTIES OF ANY KIND, BOTH EXPRESS AND IMPLIED,
INCLUDING, WITHOUT LIMITATION, ANY WARRANTIES OF TITLE, NON-INFRINGEMENT,
MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE. YOU ARE SOLELY RESPONSIBLE FOR
DETERMINING THE APPROPRIATENESS OF USING OR REDISTRIBUTING THE LLAMA MATERIALS AND
ASSUME ANY RISKS ASSOCIATED WITH YOUR USE OF THE LLAMA MATERIALS AND ANY OUTPUT AND
RESULTS.
4. Limitation of Liability. IN NO EVENT WILL META OR ITS AFFILIATES BE LIABLE UNDER ANY THEORY OF
LIABILITY, WHETHER IN CONTRACT, TORT, NEGLIGENCE, PRODUCTS LIABILITY, OR OTHERWISE, ARISING
OUT OF THIS AGREEMENT, FOR ANY LOST PROFITS OR ANY INDIRECT, SPECIAL, CONSEQUENTIAL,
INCIDENTAL, EXEMPLARY OR PUNITIVE DAMAGES, EVEN IF META OR ITS AFFILIATES HAVE BEEN ADVISED
OF THE POSSIBILITY OF ANY OF THE FOREGOING.
5. Intellectual Property.
a. No trademark licenses are granted under this Agreement, and in connection with the Llama
Materials, neither Meta nor Licensee may use any name or mark owned by or associated with the other
or any of its affiliates, except as required for reasonable and customary use in describing and
redistributing the Llama Materials or as set forth in this Section 5(a). Meta hereby grants you a license to
use “Llama” (the “Mark”) solely as required to comply with the last sentence of Section 1.b.i. You will
comply with Metas brand guidelines (currently accessible at
https://about.meta.com/brand/resources/meta/company-brand/ ). All goodwill arising out of your use
of the Mark will inure to the benefit of Meta.
b. Subject to Metas ownership of Llama Materials and derivatives made by or for Meta, with
respect to any derivative works and modifications of the Llama Materials that are made by you, as
between you and Meta, you are and will be the owner of such derivative works and modifications.
c. If you institute litigation or other proceedings against Meta or any entity (including a
cross-claim or counterclaim in a lawsuit) alleging that the Llama Materials or Llama 3.1 outputs or
results, or any portion of any of the foregoing, constitutes infringement of intellectual property or other
rights owned or licensable by you, then any licenses granted to you under this Agreement shall
terminate as of the date such litigation or claim is filed or instituted. You will indemnify and hold
harmless Meta from and against any claim by any third party arising out of or related to your use or
distribution of the Llama Materials.
6. Term and Termination. The term of this Agreement will commence upon your acceptance of this
Agreement or access to the Llama Materials and will continue in full force and effect until terminated in
accordance with the terms and conditions herein. Meta may terminate this Agreement if you are in
breach of any term or condition of this Agreement. Upon termination of this Agreement, you shall delete
and cease use of the Llama Materials. Sections 3, 4 and 7 shall survive the termination of this
Agreement.
7. Governing Law and Jurisdiction. This Agreement will be governed and construed under the laws of
the State of California without regard to choice of law principles, and the UN Convention on Contracts
for the International Sale of Goods does not apply to this Agreement. The courts of California shall have
exclusive jurisdiction of any dispute arising out of this Agreement.
### Llama 3.1 Acceptable Use Policy
Meta is committed to promoting safe and fair use of its tools and features, including Llama 3.1. If you
access or use Llama 3.1, you agree to this Acceptable Use Policy (“Policy”). The most recent copy of
this policy can be found at [https://llama.meta.com/llama3_1/use-policy](https://llama.meta.com/llama3_1/use-policy)
#### Prohibited Uses
We want everyone to use Llama 3.1 safely and responsibly. You agree you will not use, or allow
others to use, Llama 3.1 to:
1. Violate the law or others rights, including to:
1. Engage in, promote, generate, contribute to, encourage, plan, incite, or further illegal or unlawful activity or content, such as:
1. Violence or terrorism
2. Exploitation or harm to children, including the solicitation, creation, acquisition, or dissemination of child exploitative content or failure to report Child Sexual Abuse Material
3. Human trafficking, exploitation, and sexual violence
4. The illegal distribution of information or materials to minors, including obscene materials, or failure to employ legally required age-gating in connection with such information or materials.
5. Sexual solicitation
6. Any other criminal activity
3. Engage in, promote, incite, or facilitate the harassment, abuse, threatening, or bullying of individuals or groups of individuals
4. Engage in, promote, incite, or facilitate discrimination or other unlawful or harmful conduct in the provision of employment, employment benefits, credit, housing, other economic benefits, or other essential goods and services
5. Engage in the unauthorized or unlicensed practice of any profession including, but not limited to, financial, legal, medical/health, or related professional practices
6. Collect, process, disclose, generate, or infer health, demographic, or other sensitive personal or private information about individuals without rights and consents required by applicable laws
7. Engage in or facilitate any action or generate any content that infringes, misappropriates, or otherwise violates any third-party rights, including the outputs or results of any products or services using the Llama Materials
8. Create, generate, or facilitate the creation of malicious code, malware, computer viruses or do anything else that could disable, overburden, interfere with or impair the proper working, integrity, operation or appearance of a website or computer system
2. Engage in, promote, incite, facilitate, or assist in the planning or development of activities that present a risk of death or bodily harm to individuals, including use of Llama 3.1 related to the following:
1. Military, warfare, nuclear industries or applications, espionage, use for materials or activities that are subject to the International Traffic Arms Regulations (ITAR) maintained by the United States Department of State
2. Guns and illegal weapons (including weapon development)
3. Illegal drugs and regulated/controlled substances
4. Operation of critical infrastructure, transportation technologies, or heavy machinery
5. Self-harm or harm to others, including suicide, cutting, and eating disorders
6. Any content intended to incite or promote violence, abuse, or any infliction of bodily harm to an individual
3. Intentionally deceive or mislead others, including use of Llama 3.1 related to the following:
1. Generating, promoting, or furthering fraud or the creation or promotion of disinformation
2. Generating, promoting, or furthering defamatory content, including the creation of defamatory statements, images, or other content
3. Generating, promoting, or further distributing spam
4. Impersonating another individual without consent, authorization, or legal right
5. Representing that the use of Llama 3.1 or outputs are human-generated
6. Generating or facilitating false online engagement, including fake reviews and other means of fake online engagement
4. Fail to appropriately disclose to end users any known dangers of your AI system
Please report any violation of this Policy, software “bug,” or other problems that could lead to a violation
of this Policy through one of the following means:
* Reporting issues with the model: [https://github.com/meta-llama/llama-models/issues](https://github.com/meta-llama/llama-models/issues)
* Reporting risky content generated by the model:
developers.facebook.com/llama_output_feedback
* Reporting bugs and security concerns: facebook.com/whitehat/info
* Reporting violations of the Acceptable Use Policy or unlicensed uses of Meta Llama 3: LlamaUseReport@meta.com
extra_gated_fields:
First Name: text
Last Name: text
Date of birth: date_picker
Country: country
Affiliation: text
Job title:
type: select
options:
- Student
- Research Graduate
- AI researcher
- AI developer/engineer
- Reporter
- Other
geo: ip_location
By clicking Submit below I accept the terms of the license and acknowledge that the information I provide will be collected stored processed and shared in accordance with the Meta Privacy Policy: checkbox
extra_gated_description: The information you provide will be collected, stored, processed and shared in accordance with the [Meta Privacy Policy](https://www.facebook.com/privacy/policy/).
extra_gated_button_content: Submit
---
# <span style="color: #7FFF7F;">Llama-Guard-3-8B GGUF Models</span>
## **Choosing the Right Model Format**
Selecting the correct model format depends on your **hardware capabilities** and **memory constraints**.
### **BF16 (Brain Float 16) Use if BF16 acceleration is available**
- A 16-bit floating-point format designed for **faster computation** while retaining good precision.
- Provides **similar dynamic range** as FP32 but with **lower memory usage**.
- Recommended if your hardware supports **BF16 acceleration** (check your devices specs).
- Ideal for **high-performance inference** with **reduced memory footprint** compared to FP32.
📌 **Use BF16 if:**
✔ Your hardware has native **BF16 support** (e.g., newer GPUs, TPUs).
✔ You want **higher precision** while saving memory.
✔ You plan to **requantize** the model into another format.
📌 **Avoid BF16 if:**
❌ Your hardware does **not** support BF16 (it may fall back to FP32 and run slower).
❌ You need compatibility with older devices that lack BF16 optimization.
---
### **F16 (Float 16) More widely supported than BF16**
- A 16-bit floating-point **high precision** but with less of range of values than BF16.
- Works on most devices with **FP16 acceleration support** (including many GPUs and some CPUs).
- Slightly lower numerical precision than BF16 but generally sufficient for inference.
📌 **Use F16 if:**
✔ Your hardware supports **FP16** but **not BF16**.
✔ You need a **balance between speed, memory usage, and accuracy**.
✔ You are running on a **GPU** or another device optimized for FP16 computations.
📌 **Avoid F16 if:**
❌ Your device lacks **native FP16 support** (it may run slower than expected).
❌ You have memory limitations.
---
### **Quantized Models (Q4_K, Q6_K, Q8, etc.) For CPU & Low-VRAM Inference**
Quantization reduces model size and memory usage while maintaining as much accuracy as possible.
- **Lower-bit models (Q4_K)** → **Best for minimal memory usage**, may have lower precision.
- **Higher-bit models (Q6_K, Q8_0)** → **Better accuracy**, requires more memory.
📌 **Use Quantized Models if:**
✔ You are running inference on a **CPU** and need an optimized model.
✔ Your device has **low VRAM** and cannot load full-precision models.
✔ You want to reduce **memory footprint** while keeping reasonable accuracy.
📌 **Avoid Quantized Models if:**
❌ You need **maximum accuracy** (full-precision models are better for this).
❌ Your hardware has enough VRAM for higher-precision formats (BF16/F16).
---
### **Very Low-Bit Quantization (IQ3_XS, IQ3_S, IQ3_M, Q4_K, Q4_0)**
These models are optimized for **extreme memory efficiency**, making them ideal for **low-power devices** or **large-scale deployments** where memory is a critical constraint.
- **IQ3_XS**: Ultra-low-bit quantization (3-bit) with **extreme memory efficiency**.
- **Use case**: Best for **ultra-low-memory devices** where even Q4_K is too large.
- **Trade-off**: Lower accuracy compared to higher-bit quantizations.
- **IQ3_S**: Small block size for **maximum memory efficiency**.
- **Use case**: Best for **low-memory devices** where **IQ3_XS** is too aggressive.
- **IQ3_M**: Medium block size for better accuracy than **IQ3_S**.
- **Use case**: Suitable for **low-memory devices** where **IQ3_S** is too limiting.
- **Q4_K**: 4-bit quantization with **block-wise optimization** for better accuracy.
- **Use case**: Best for **low-memory devices** where **Q6_K** is too large.
- **Q4_0**: Pure 4-bit quantization, optimized for **ARM devices**.
- **Use case**: Best for **ARM-based devices** or **low-memory environments**.
---
### **Summary Table: Model Format Selection**
| Model Format | Precision | Memory Usage | Device Requirements | Best Use Case |
|--------------|------------|---------------|----------------------|---------------|
| **BF16** | Highest | High | BF16-supported GPU/CPUs | High-speed inference with reduced memory |
| **F16** | High | High | FP16-supported devices | GPU inference when BF16 isnt available |
| **Q4_K** | Medium Low | Low | CPU or Low-VRAM devices | Best for memory-constrained environments |
| **Q6_K** | Medium | Moderate | CPU with more memory | Better accuracy while still being quantized |
| **Q8_0** | High | Moderate | CPU or GPU with enough VRAM | Best accuracy among quantized models |
| **IQ3_XS** | Very Low | Very Low | Ultra-low-memory devices | Extreme memory efficiency and low accuracy |
| **Q4_0** | Low | Low | ARM or low-memory devices | llama.cpp can optimize for ARM devices |
---
## **Included Files & Details**
### `Llama-Guard-3-8B-bf16.gguf`
- Model weights preserved in **BF16**.
- Use this if you want to **requantize** the model into a different format.
- Best if your device supports **BF16 acceleration**.
### `Llama-Guard-3-8B-f16.gguf`
- Model weights stored in **F16**.
- Use if your device supports **FP16**, especially if BF16 is not available.
### `Llama-Guard-3-8B-bf16-q8_0.gguf`
- **Output & embeddings** remain in **BF16**.
- All other layers quantized to **Q8_0**.
- Use if your device supports **BF16** and you want a quantized version.
### `Llama-Guard-3-8B-f16-q8_0.gguf`
- **Output & embeddings** remain in **F16**.
- All other layers quantized to **Q8_0**.
### `Llama-Guard-3-8B-q4_k.gguf`
- **Output & embeddings** quantized to **Q8_0**.
- All other layers quantized to **Q4_K**.
- Good for **CPU inference** with limited memory.
### `Llama-Guard-3-8B-q4_k_s.gguf`
- Smallest **Q4_K** variant, using less memory at the cost of accuracy.
- Best for **very low-memory setups**.
### `Llama-Guard-3-8B-q6_k.gguf`
- **Output & embeddings** quantized to **Q8_0**.
- All other layers quantized to **Q6_K** .
### `Llama-Guard-3-8B-q8_0.gguf`
- Fully **Q8** quantized model for better accuracy.
- Requires **more memory** but offers higher precision.
### `Llama-Guard-3-8B-iq3_xs.gguf`
- **IQ3_XS** quantization, optimized for **extreme memory efficiency**.
- Best for **ultra-low-memory devices**.
### `Llama-Guard-3-8B-iq3_m.gguf`
- **IQ3_M** quantization, offering a **medium block size** for better accuracy.
- Suitable for **low-memory devices**.
### `Llama-Guard-3-8B-q4_0.gguf`
- Pure **Q4_0** quantization, optimized for **ARM devices**.
- Best for **low-memory environments**.
- Prefer IQ4_NL for better accuracy.
# <span id="testllm" style="color: #7F7FFF;">🚀 If you find these models useful</span>
Please click like ❤ . Also Id really appreciate it if you could test my Network Monitor Assistant at 👉 [Network Monitor Assitant](https://readyforquantum.com).
💬 Click the **chat icon** (bottom right of the main and dashboard pages) . Choose a LLM; toggle between the LLM Types TurboLLM -> FreeLLM -> TestLLM.
### What I'm Testing
I'm experimenting with **function calling** against my network monitoring service. Using small open source models. I am into the question "How small can it go and still function".
🟡 **TestLLM** Runs the current testing model using llama.cpp on 6 threads of a Cpu VM (Should take about 15s to load. Inference speed is quite slow and it only processes one user prompt at a time—still working on scaling!). If you're curious, I'd be happy to share how it works! .
### The other Available AI Assistants
🟢 **TurboLLM** Uses **gpt-4o-mini** Fast! . Note: tokens are limited since OpenAI models are pricey, but you can [Login](https://readyforquantum.com) or [Download](https://readyforquantum.com/download/?utm_source=huggingface&utm_medium=referral&utm_campaign=huggingface_repo_readme) the Quantum Network Monitor agent to get more tokens, Alternatively use the TestLLM .
🔵 **HugLLM** Runs **open-source Hugging Face models** Fast, Runs small models (≈8B) hence lower quality, Get 2x more tokens (subject to Hugging Face API availability)
### Final Word
I fund the servers used to create these model files, run the Quantum Network Monitor service, and pay for inference from Novita and OpenAI—all out of my own pocket. All the code behind the model creation and the Quantum Network Monitor project is [open source](https://github.com/Mungert69). Feel free to use whatever you find helpful.
If you appreciate the work, please consider [buying me a coffee](https://www.buymeacoffee.com/mahadeva) ☕. Your support helps cover service costs and allows me to raise token limits for everyone.
I'm also open to job opportunities or sponsorship.
Thank you! 😊
# Model Details
Llama Guard 3 is a Llama-3.1-8B pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM inputs (prompt classification) and in LLM responses (response classification). It acts as an LLM it generates text in its output that indicates whether a given prompt or response is safe or unsafe, and if unsafe, it also lists the content categories violated.
Llama Guard 3 was aligned to safeguard against the MLCommons standardized hazards taxonomy and designed to support Llama 3.1 capabilities. Specifically, it provides content moderation in 8 languages, and was optimized to support safety and security for search and code interpreter tool calls.
Below is a response classification example for Llama Guard 3.
<p align="center">
<img src="llama_guard_3_figure.png" width="800"/>
</p>
In order to produce classifier scores, we look at the probability for the first token, and use that as the “unsafe” class probability. We can then apply score thresholding to make binary decisions.
# Hazard Taxonomy and Policy
The model is trained to predict safety labels on the 14 categories shown below, based on the [MLCommons taxonomy](https://mlcommons.org/2024/04/mlc-aisafety-v0-5-poc/) of 13 hazards, as well as an additional category for Code Interpreter Abuse for tool calls use cases
<table align="center">
<thead>
<tr>
<th colspan="2">Hazard categories</th>
</tr>
</thead>
<tbody>
<tr>
<td>S1: Violent Crimes</td>
<td>S2: Non-Violent Crimes</td>
</tr>
<tr>
<td>S3: Sex-Related Crimes</td>
<td>S4: Child Sexual Exploitation</td>
</tr>
<tr>
<td>S5: Defamation</td>
<td>S6: Specialized Advice</td>
</tr>
<tr>
<td>S7: Privacy</td>
<td>S8: Intellectual Property</td>
</tr>
<tr>
<td>S9: Indiscriminate Weapons</td>
<td>S10: Hate</td>
</tr>
<tr>
<td>S11: Suicide &amp; Self-Harm</td>
<td>S12: Sexual Content</td>
</tr>
<tr>
<td>S13: Elections</td>
<td>S14: Code Interpreter Abuse</td>
</tr>
</tbody>
</table>
**S1 - Violent Crimes**
Responses that enable, encourage, or endorse the commission of violent crimes, including: (1) unlawful violence toward people (ex: terrorism, genocide, murder, hate-motivated violent crimes, child abuse, assault, battery, kidnapping) and (2) unlawful violence toward animals (ex: animal abuse)
**S2 - Non-Violent Crimes**
Responses that enable, encourage, or endorse the commission of non-violent crimes, including: (1) personal crimes (ex: labor trafficking, threats, intimidation, hate-motivated non-violent crimes), (2) financial crimes (ex: fraud, scams, money laundering), (3) property crimes (ex: theft, arson, vandalism), (4) drug crimes (ex: creating or using regulated substances), (5) weapons crimes (ex: producing unlicensed firearms), and (6) cyber crimes (ex: hacking)
**S3 - Sex-Related Crimes**
Responses that enable, encourage, or endorse the commission of sex-related crimes, including: (1) sex trafficking, (2) sexual assault (ex: rape), (3) sexual harassment (physical, verbal, or visual), and (4) prostitution
**S4 - Child Sexual Exploitation**
Responses that contain, describe, enable, encourage, or endorse the sexual abuse of children
**S5 - Defamation**
Responses that are both verifiably false and likely to injure a living persons reputation
**S6 - Specialized Advice**
Responses that contain specialized financial, medical, or legal advice, or that indicate dangerous activities or objects are safe
**S7 - Privacy**
Responses that contain sensitive, nonpublic personal information that could undermine someones physical, digital, or financial security
**S8 - Intellectual Property**
Responses that may violate the intellectual property rights of any third party
**S9 - Indiscriminate Weapons**
Responses that enable, encourage, or endorse the creation of indiscriminate weapons, including: (1) chemical weapons (ex: nerve gas), (2) biological weapons (ex: anthrax), (3) radiological weapons (ex: cobalt bombs), (4) nuclear weapons (ex: fission bombs), and (5) high-yield explosive weapons (ex: cluster munitions)
**S10 - Hate**
Responses that demean or dehumanize people on the basis of their sensitive, personal characteristics (i.e., race, color, ethnicity, national origin, disability, religious affiliation, caste, sexual orientation, sex, gender identity, and/or serious disease)
**S11 - Suicide &amp; Self-Harm**
Responses that enable, encourage, or endorse acts of intentional self-harm, including: (1) suicide, (2) self-injury (ex: cutting), and (3) disordered eating
**S12 - Sexual Content**
Responses that contain erotica
**S13 - Elections**
Responses that contain factually incorrect information about electoral systems and processes, including in the time, place, or manner of voting in civic elections
**S14 - Code Interpreter Abuse**
Responses that seek to abuse code interpreters, including those that enable denial of service attacks, container escapes or privilege escalation exploits
# Supported languages
Llama Guard 3 supports content safety for the following languages : English, French, German, Hindi, Italian, Portuguese, Spanish, Thai.
# Usage
> [!IMPORTANT]
>
> This repository corresponds to half-precision version of the model. A 8-bit precision version is also provided, please visit [meta-llama/Llama-Guard-3-8B-INT8](https://huggingface.co/meta-llama/Llama-Guard-3-8B-INT8).
Llama Guard 3 can be directly used with `transformers`. It is only supported since `transformers` version 4.43.
```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "meta-llama/Llama-Guard-3-8B"
device = "cuda"
dtype = torch.bfloat16
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=dtype, device_map=device)
def moderate(chat):
input_ids = tokenizer.apply_chat_template(chat, return_tensors="pt").to(device)
output = model.generate(input_ids=input_ids, max_new_tokens=100, pad_token_id=0)
prompt_len = input_ids.shape[-1]
return tokenizer.decode(output[0][prompt_len:], skip_special_tokens=True)
moderate([
{"role": "user", "content": "I forgot how to kill a process in Linux, can you help?"},
{"role": "assistant", "content": "Sure! To kill a process in Linux, you can use the kill command followed by the process ID (PID) of the process you want to terminate."},
])
```
# Training Data
We use the English data used by Llama Guard [1], which are obtained by getting Llama 2 and Llama 3 generations on prompts from the hh-rlhf dataset [2]. In order to scale training data for new categories and new capabilities such as multilingual and tool use, we collect additional human and synthetically generated data. Similar to the English data, the multilingual data are Human-AI conversation data that are either single-turn or multi-turn. To reduce the models false positive rate, we curate a set of multilingual benign prompt and response data where LLMs likely reject the prompts.
For the tool use capability, we consider search tool calls and code interpreter abuse. To develop training data for search tool use, we use Llama3 to generate responses to a collected and synthetic set of prompts. The generations are based on the query results obtained from the Brave Search API. To develop synthetic training data to detect code interpreter attacks, we use an LLM to generate safe and unsafe prompts. Then, we use a non-safety-tuned LLM to generate code interpreter completions that comply with these instructions. For safe data, we focus on data close to the boundary of what would be considered unsafe, to minimize false positives on such borderline examples.
# Evaluation
**Note on evaluations:** As discussed in the original Llama Guard paper, comparing model performance is not straightforward as each model is built on its own policy and is expected to perform better on an evaluation dataset with a policy aligned to the model. This highlights the need for industry standards. By aligning the Llama Guard family of models with the Proof of Concept MLCommons taxonomy of hazards, we hope to drive adoption of industry standards like this and facilitate collaboration and transparency in the LLM safety and content evaluation space.
In this regard, we evaluate the performance of Llama Guard 3 on MLCommons hazard taxonomy and compare it across languages with Llama Guard 2 [3] on our internal test. We also add GPT4 as baseline with zero-shot prompting using MLCommons hazard taxonomy.
Tables 1, 2, and 3 show that Llama Guard 3 improves over Llama Guard 2 and outperforms GPT4 in English, multilingual, and tool use capabilities. Noteworthily, Llama Guard 3 achieves better performance with much lower false positive rates. We also benchmark Llama Guard 3 in the OSS dataset XSTest [4] and observe that it achieves the same F1 score but a lower false positive rate compared to Llama Guard 2.
<div align="center">
<small> Table 1: Comparison of performance of various models measured on our internal English test set for MLCommons hazard taxonomy (response classification).</small>
| | **F1 ↑** | **AUPRC ↑** | **False Positive<br>Rate ↓** |
|--------------------------|:--------:|:-----------:|:----------------------------:|
| Llama Guard 2 | 0.877 | 0.927 | 0.081 |
| Llama Guard 3 | 0.939 | 0.985 | 0.040 |
| GPT4 | 0.805 | N/A | 0.152 |
</div>
<br>
<table align="center">
<small><center>Table 2: Comparison of multilingual performance of various models measured on our internal test set for MLCommons hazard taxonomy (prompt+response classification).</center></small>
<thead>
<tr>
<th colspan="8"><center>F1 ↑ / FPR ↓</center></th>
</tr>
</thead>
<tbody>
<tr>
<td></td>
<td><center>French</center></td>
<td><center>German</center></td>
<td><center>Hindi</center></td>
<td><center>Italian</center></td>
<td><center>Portuguese</center></td>
<td><center>Spanish</center></td>
<td><center>Thai</center></td>
</tr>
<tr>
<td>Llama Guard 2</td>
<td><center>0.911/0.012</center></td>
<td><center>0.795/0.062</center></td>
<td><center>0.832/0.062</center></td>
<td><center>0.681/0.039</center></td>
<td><center>0.845/0.032</center></td>
<td><center>0.876/0.001</center></td>
<td><center>0.822/0.078</center></td>
</tr>
<tr>
<td>Llama Guard 3</td>
<td><center>0.943/0.036</center></td>
<td><center>0.877/0.032</center></td>
<td><center>0.871/0.050</center></td>
<td><center>0.873/0.038</center></td>
<td><center>0.860/0.060</center></td>
<td><center>0.875/0.023</center></td>
<td><center>0.834/0.030</center></td>
</tr>
<tr>
<td>GPT4</td>
<td><center>0.795/0.157</center></td>
<td><center>0.691/0.123</center></td>
<td><center>0.709/0.206</center></td>
<td><center>0.753/0.204</center></td>
<td><center>0.738/0.207</center></td>
<td><center>0.711/0.169</center></td>
<td><center>0.688/0.168</center></td>
</tr>
</tbody>
</table>
<br>
<table align="center">
<small><center>Table 3: Comparison of performance of various models measured on our internal test set for other moderation capabilities (prompt+response classification).</center></small>
<thead>
<tr>
<th></th>
<th colspan="3">Search tool calls</th>
<th colspan="3">Code interpreter abuse</th>
</tr>
</thead>
<tbody>
<tr>
<td></td>
<td><center>F1 ↑</center></td>
<td><center>AUPRC ↑</center></td>
<td><center>FPR ↓</center></td>
<td><center>F1 ↑</center></td>
<td><center>AUPRC ↑</center></td>
<td><center>FPR ↓</center></td>
</tr>
<tr>
<td>Llama Guard 2</td>
<td><center>0.749</center></td>
<td><center>0.794</center></td>
<td><center>0.284</center></td>
<td><center>0.683</center></td>
<td><center>0.677</center></td>
<td><center>0.670</center></td>
</tr>
<tr>
<td>Llama Guard 3</td>
<td><center>0.856</center></td>
<td><center>0.938</center></td>
<td><center>0.174</center></td>
<td><center>0.885</center></td>
<td><center>0.967</center></td>
<td><center>0.125</center></td>
</tr>
<tr>
<td>GPT4</td>
<td><center>0.732</center></td>
<td><center>N/A</center></td>
<td><center>0.525</center></td>
<td><center>0.636</center></td>
<td><center>N/A</center></td>
<td><center>0.90</center></td>
</tr>
</tbody>
</table>
# Application
As outlined in the Llama 3 paper, Llama Guard 3 provides industry leading system-level safety performance and is recommended to be deployed along with Llama 3.1. Note that, while deploying Llama Guard 3 will likely improve the safety of your system, it might increase refusals to benign prompts (False Positives). Violation rate improvement and impact on false positives as measured on internal benchmarks are provided in the Llama 3 paper.
# Quantization
We are committed to help the community deploy Llama systems responsibly. We provide a quantized version of Llama Guard 3 to lower the deployment cost. We used int 8 [implementation](https://huggingface.co/docs/transformers/main/en/quantization/bitsandbytes) integrated into the hugging face ecosystem, reducing the checkpoint size by about 40% with very small impact on model performance. In Table 5, we observe that the performance quantized model is comparable to the original model.
<table align="center">
<small><center>Table 5: Impact of quantization on Llama Guard 3 performance.</center></small>
<tbody>
<tr>
<td rowspan="2"><br />
<p><span>Task</span></p>
</td>
<td rowspan="2"><br />
<p><span>Capability</span></p>
</td>
<td colspan="4">
<p><center><span>Non-Quantized</span></center></p>
</td>
<td colspan="4">
<p><center><span>Quantized</span></center></p>
</td>
</tr>
<tr>
<td>
<p><span>Precision</span></p>
</td>
<td>
<p><span>Recall</span></p>
</td>
<td>
<p><span>F1</span></p>
</td>
<td>
<p><span>FPR</span></p>
</td>
<td>
<p><span>Precision</span></p>
</td>
<td>
<p><span>Recall</span></p>
</td>
<td>
<p><span>F1</span></p>
</td>
<td>
<p><span>FPR</span></p>
</td>
</tr>
<tr>
<td rowspan="3">
<p><span>Prompt Classification</span></p>
</td>
<td>
<p><span>English</span></p>
</td>
<td>
<p><span>0.952</span></p>
</td>
<td>
<p><span>0.943</span></p>
</td>
<td>
<p><span>0.947</span></p>
</td>
<td>
<p><span>0.057</span></p>
</td>
<td>
<p><span>0.961</span></p>
</td>
<td>
<p><span>0.939</span></p>
</td>
<td>
<p><span>0.950</span></p>
</td>
<td>
<p><span>0.045</span></p>
</td>
</tr>
<tr>
<td>
<p><span>Multilingual</span></p>
</td>
<td>
<p><span>0.901</span></p>
</td>
<td>
<p><span>0.899</span></p>
</td>
<td>
<p><span>0.900</span></p>
</td>
<td>
<p><span>0.054</span></p>
</td>
<td>
<p><span>0.906</span></p>
</td>
<td>
<p><span>0.892</span></p>
</td>
<td>
<p><span>0.899</span></p>
</td>
<td>
<p><span>0.051</span></p>
</td>
</tr>
<tr>
<td>
<p><span>Tool Use</span></p>
</td>
<td>
<p><span>0.884</span></p>
</td>
<td>
<p><span>0.958</span></p>
</td>
<td>
<p><span>0.920</span></p>
</td>
<td>
<p><span>0.126</span></p>
</td>
<td>
<p><span>0.876</span></p>
</td>
<td>
<p><span>0.946</span></p>
</td>
<td>
<p><span>0.909</span></p>
</td>
<td>
<p><span>0.134</span></p>
</td>
</tr>
<tr>
<td rowspan="3">
<p><span>Response Classification</span></p>
</td>
<td>
<p><span>English</span></p>
</td>
<td>
<p><span>0.947</span></p>
</td>
<td>
<p><span>0.931</span></p>
</td>
<td>
<p><span>0.939</span></p>
</td>
<td>
<p><span>0.040</span></p>
</td>
<td>
<p><span>0.947</span></p>
</td>
<td>
<p><span>0.925</span></p>
</td>
<td>
<p><span>0.936</span></p>
</td>
<td>
<p><span>0.040</span></p>
</td>
</tr>
<tr>
<td>
<p><span>Multilingual</span></p>
</td>
<td>
<p><span>0.929</span></p>
</td>
<td>
<p><span>0.805</span></p>
</td>
<td>
<p><span>0.862</span></p>
</td>
<td>
<p><span>0.033</span></p>
</td>
<td>
<p><span>0.931</span></p>
</td>
<td>
<p><span>0.785</span></p>
</td>
<td>
<p><span>0.851</span></p>
</td>
<td>
<p><span>0.031</span></p>
</td>
</tr>
<tr>
<td>
<p><span>Tool Use</span></p>
</td>
<td>
<p><span>0.774</span></p>
</td>
<td>
<p><span>0.884</span></p>
</td>
<td>
<p><span>0.825</span></p>
</td>
<td>
<p><span>0.176</span></p>
</td>
<td>
<p><span>0.793</span></p>
</td>
<td>
<p><span>0.865</span></p>
</td>
<td>
<p><span>0.827</span></p>
</td>
<td>
<p><span>0.155</span></p>
</td>
</tr>
</tbody>
</table>
# Get started
Llama Guard 3 is available by default on Llama 3.1 [reference implementations](https://github.com/meta-llama). You can learn more about how to configure and customize using [Llama Recipes](https://github.com/meta-llama/llama-recipes/tree/main/recipes/responsible_ai/) shared on our Github repository.
# Limitations
There are some limitations associated with Llama Guard 3. First, Llama Guard 3 itself is an LLM fine-tuned on Llama 3.1. Thus, its performance (e.g., judgments that need common sense knowledge, multilingual capability, and policy coverage) might be limited by its (pre-)training data.
Some hazard categories may require factual, up-to-date knowledge to be evaluated (for example, S5: Defamation, S8: Intellectual Property, and S13: Elections) . We believe more complex systems should be deployed to accurately moderate these categories for use cases highly sensitive to these types of hazards, but Llama Guard 3 provides a good baseline for generic use cases.
Lastly, as an LLM, Llama Guard 3 may be susceptible to adversarial attacks or prompt injection attacks that could bypass or alter its intended use. Please feel free to [report](https://github.com/meta-llama/PurpleLlama) vulnerabilities and we will look to incorporate improvements in future versions of Llama Guard.
# Citation
```
@misc{dubey2024llama3herdmodels,
title = {The Llama 3 Herd of Models},
author = {Llama Team, AI @ Meta},
year = {2024},
eprint = {2407.21783},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2407.21783}
}
```
# References
[1] [Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations](https://arxiv.org/abs/2312.06674)
[2] [Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback](https://arxiv.org/abs/2204.05862)
[3] [Llama Guard 2 Model Card](https://github.com/meta-llama/PurpleLlama/blob/main/Llama-Guard2/MODEL_CARD.md)
[4] [XSTest: A Test Suite for Identifying Exaggerated Safety Behaviors in Large Language Models](https://arxiv.org/abs/2308.01263)

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}