初始化项目,由ModelHub XC社区提供模型
Model: huihui-ai/Huihui-EXAONE-4.0-1.2B-abliterated Source: Original Platform
This commit is contained in:
35
.gitattributes
vendored
Normal file
35
.gitattributes
vendored
Normal file
@@ -0,0 +1,35 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
157
LICENSE
Normal file
157
LICENSE
Normal file
@@ -0,0 +1,157 @@
|
||||
EXAONE AI Model License Agreement 1.2 - NC
|
||||
|
||||
This License Agreement (“Agreement”) is entered into between you (“Licensee”) and LG Management Development
|
||||
Institute Co., Ltd. (“Licensor”), governing the use of the EXAONE AI Model (“Model”). By downloading,
|
||||
installing, copying, or using the Model, you agree to comply with and be bound by the terms of this Agreement.
|
||||
If you do not agree to all the terms, you must not download, install, copy, or use the Model. This Agreement
|
||||
constitutes a binding legal agreement between the Licensee and Licensor.
|
||||
|
||||
1. Definitions
|
||||
1.1 Model: The artificial intelligence model provided by Licensor, which includes any software,
|
||||
algorithms, machine learning models, or related components supplied by Licensor. This definition extends
|
||||
to encompass all updates, enhancements, improvements, bug fixes, patches, or other modifications that may
|
||||
be provided by Licensor from time to time, whether automatically or manually implemented.
|
||||
1.2 Derivatives: Any modifications, alterations, enhancements, improvements, adaptations, or derivative
|
||||
works of the Model created by Licensee or any third party. This includes changes made to the Model's
|
||||
architecture, parameters, data processing methods, or any other aspect of the Model that results in a
|
||||
modification of its functionality or output.
|
||||
1.3 Output: Any data, results, content, predictions, analyses, insights, or other materials generated by
|
||||
the Model or Derivatives, regardless of whether they are in their original form or have been further
|
||||
processed or modified by the Licensee. This includes, but is not limited to, textual or numerical produced
|
||||
directly or indirectly through the use of the Model.
|
||||
1.4 Licensor: LG Management Development Institute Co., Ltd., the owner, developer, and provider of the
|
||||
EXAONE AI Model. The Licensor holds all rights, title, and interest in the Model and is responsible for
|
||||
granting licenses to use the Model under the terms specified in this Agreement.
|
||||
1.5 Licensee: The individual, organization, corporation, academic institution, government agency, or other
|
||||
entity using or intending to use the Model under the terms and conditions of this Agreement. The Licensee
|
||||
is responsible for ensuring compliance with the Agreement by all authorized users who access or utilize
|
||||
the Model on behalf of the Licensee.
|
||||
|
||||
2. License Grant
|
||||
2.1 Grant of License: Subject to the terms and conditions outlined in this Agreement, the Licensor hereby
|
||||
grants the Licensee a limited, non-exclusive, non-transferable, worldwide, and revocable license to:
|
||||
a. Access, download, install, and use the Model solely for research and educational purposes. This
|
||||
includes evaluation, testing, academic research, experimentation, learning, teaching, training and
|
||||
participation in competitions, provided that such participation is in a non-commercial context.
|
||||
Notwithstanding Section 3.1, the Licensee may only provide the Model or Derivatives for a competition
|
||||
if no commercial license is granted to the competition organizer or any third party.
|
||||
b. Publicly disclose research results and findings derived from the use of the Model or Derivatives,
|
||||
including publishing papers or presentations.
|
||||
c. Modify the Model and create Derivatives based on the Model, provided that such modifications and
|
||||
Derivatives are used exclusively for research and educational purposes. The Licensee may conduct
|
||||
experiments, perform analyses, and apply custom modifications to the Model to explore its capabilities
|
||||
and performance under various scenarios. If the Model is modified, the modified Model must include
|
||||
"EXAONE" at the beginning of its name.
|
||||
d. Distribute the Model and Derivatives in each case with a copy of this Agreement.
|
||||
2.2 Scope of License: The license granted herein does not authorize the Licensee to use the Model for any
|
||||
purpose not explicitly permitted under this Agreement. Any use beyond the scope of this license, including
|
||||
any commercial application or external distribution, is strictly prohibited unless explicitly agreed upon
|
||||
in writing by the Licensor.
|
||||
|
||||
3. Restrictions
|
||||
3.1 Commercial Use: The Licensee is expressly prohibited from using the Model, Derivatives, or Output for
|
||||
any commercial purposes, including but not limited to, developing or deploying products, services, or
|
||||
applications that generate revenue, whether directly or indirectly. Any commercial exploitation of the
|
||||
Model or its derivatives requires a separate commercial license agreement with the Licensor. Furthermore,
|
||||
the Licensee shall not use the Model, Derivatives or Output to develop or improve any models that compete
|
||||
with the Licensor’s models.
|
||||
3.2 Reverse Engineering: The Licensee shall not decompile, disassemble, reverse engineer, or attempt to
|
||||
derive the source code, underlying ideas, algorithms, or structure of the Model, except to the extent that
|
||||
such activities are expressly permitted by applicable law. Any attempt to bypass or circumvent
|
||||
technological protection measures applied to the Model is strictly prohibited.
|
||||
3.3 Unlawful Use: The Licensee shall not use the Model and Derivatives for any illegal, fraudulent, or
|
||||
unauthorized activities, nor for any purpose that violates applicable laws or regulations. This includes
|
||||
but is not limited to the creation, distribution, or dissemination of malicious, deceptive, or unlawful
|
||||
content.
|
||||
3.4 Ethical Use: The Licensee shall ensure that the Model or Derivatives is used in an ethical and
|
||||
responsible manner, adhering to the following guidelines:
|
||||
a. The Model and Derivatives shall not be used to generate, propagate, or amplify false, misleading,
|
||||
or harmful information, including fake news, misinformation, or disinformation.
|
||||
b. The Model and Derivatives shall not be employed to create, distribute, or promote content that is
|
||||
discriminatory, harassing, defamatory, abusive, or otherwise offensive to individuals or groups based
|
||||
on race, gender, sexual orientation, religion, nationality, or other protected characteristics.
|
||||
c. The Model and Derivatives shall not infringe on the rights of others, including intellectual property
|
||||
rights, privacy rights, or any other rights recognized by law. The Licensee shall obtain all necessary
|
||||
permissions and consents before using the Model and Derivatives in a manner that may impact the rights
|
||||
of third parties.
|
||||
d. The Model and Derivatives shall not be used in a way that causes harm, whether physical, mental,
|
||||
emotional, or financial, to individuals, organizations, or communities. The Licensee shall take all
|
||||
reasonable measures to prevent misuse or abuse of the Model and Derivatives that could result in harm
|
||||
or injury.
|
||||
|
||||
4. Ownership
|
||||
4.1 Intellectual Property: All rights, title, and interest in and to the Model, including any
|
||||
modifications, Derivatives, and associated documentation, are and shall remain the exclusive property of
|
||||
the Licensor. The Licensee acknowledges that this Agreement does not transfer any ownership rights to the
|
||||
Licensee. All trademarks, service marks, and logos associated with the Model are the property of the
|
||||
Licensor.
|
||||
4.2 Output: Licensor claims no rights in Output. Licensee is solely responsible for the Output and its use.
|
||||
4.3 Attribution: In any publication or presentation of results obtained using the Model, the Licensee
|
||||
shall provide appropriate attribution to the Licensor, citing the Model's name and version, along with any
|
||||
relevant documentation or references specified by the Licensor.
|
||||
|
||||
5. No Warranty
|
||||
5.1 “As-Is” Basis: The Model, Derivatives, and Output are provided on an “as-is” and “as-available” basis,
|
||||
without any warranties or representations of any kind, whether express, implied, or statutory. The Licensor
|
||||
disclaims all warranties, including but not limited to, implied warranties of merchantability, fitness for
|
||||
a particular purpose, accuracy, reliability, non-infringement, or any warranty arising from the course of
|
||||
dealing or usage of trade.
|
||||
5.2 Performance and Reliability: The Licensor does not warrant or guarantee that the Model, Derivatives or
|
||||
Output will meet the Licensee’s requirements, that the operation of the Model, Derivatives or Output will
|
||||
be uninterrupted or error-free, or that defects in the Model will be corrected. The Licensee acknowledges
|
||||
that the use of the Model, Derivatives or Output is at its own risk and that the Model, Derivatives or
|
||||
Output may contain bugs, errors, or other limitations.
|
||||
5.3 No Endorsement: The Licensor does not endorse, approve, or certify any results, conclusions, or
|
||||
recommendations derived from the use of the Model. The Licensee is solely responsible for evaluating the
|
||||
accuracy, reliability, and suitability of the Model for its intended purposes.
|
||||
|
||||
6. Limitation of Liability
|
||||
6.1 No Liability for Damages: To the fullest extent permitted by applicable law, in no event shall the
|
||||
Licensor be liable for any special, incidental, indirect, consequential, exemplary, or punitive damages,
|
||||
including but not limited to, damages for loss of business profits, business interruption, loss of business
|
||||
information, loss of data, or any other pecuniary or non-pecuniary loss arising out of or in connection with
|
||||
the use or inability to use the Model, Derivatives or any Output, even if the Licensor has been advised of
|
||||
the possibility of such damages.
|
||||
6.2 Indemnification: The Licensee agrees to indemnify, defend, and hold harmless the Licensor, its
|
||||
affiliates, officers, directors, employees, and agents from and against any claims, liabilities, damages,
|
||||
losses, costs, or expenses (including reasonable attorneys' fees) arising out of or related to the
|
||||
Licensee's use of the Model, any Derivatives, or any Output, including any violation of this Agreement or
|
||||
applicable laws.
|
||||
|
||||
7. Termination
|
||||
7.1 Termination by Licensor: The Licensor reserves the right to terminate this Agreement and revoke the
|
||||
Licensee’s rights to use the Model at any time, with or without cause, and without prior notice if the
|
||||
Licensee breaches any of the terms or conditions of this Agreement. Termination shall be effective
|
||||
immediately upon notice.
|
||||
7.2 Effect of Termination: Upon termination of this Agreement, the Licensee must immediately cease all use
|
||||
of the Model and Derivatives and destroy all copies of the Model and Derivatives in its possession or
|
||||
control, including any backup or archival copies. The Licensee shall certify in writing to the Licensor that
|
||||
such destruction has been completed.
|
||||
7.3 Survival: The provisions of this Agreement that by their nature should survive termination, including
|
||||
but not limited to, Sections 4 (Ownership), 5 (No Warranty), 6 (Limitation of Liability), and this Section 7
|
||||
(Termination), shall continue to apply after termination.
|
||||
|
||||
8. Governing Law
|
||||
8.1 Governing Law: This Agreement shall be governed by and construed in accordance with the laws of the
|
||||
Republic of Korea, without regard to its conflict of laws principles.
|
||||
8.2 Arbitration: Any disputes, controversies, or claims arising out of or relating to this Agreement,
|
||||
including its existence, validity, interpretation, performance, breach, or termination, shall be referred
|
||||
to and finally resolved by arbitration administered by the Korean Commercial Arbitration Board (KCAB) in
|
||||
accordance with the International Arbitration Rules of the Korean Commercial Arbitration Board in force at
|
||||
the time of the commencement of the arbitration. The seat of arbitration shall be Seoul, Republic of Korea.
|
||||
The tribunal shall consist of one arbitrator. The language of the arbitration shall be English.
|
||||
|
||||
9. Alterations
|
||||
9.1 Modifications: The Licensor reserves the right to modify or amend this Agreement at any time, in its
|
||||
sole discretion. Any modifications will be effective upon posting the updated Agreement on the Licensor’s
|
||||
website or through other means of communication. The Licensee is responsible for reviewing the Agreement
|
||||
periodically for changes. Continued use of the Model after any modifications have been made constitutes
|
||||
acceptance of the revised Agreement.
|
||||
9.2 Entire Agreement: This Agreement constitutes the entire agreement between the Licensee and Licensor
|
||||
concerning the subject matter hereof and supersedes all prior or contemporaneous oral or written agreements,
|
||||
representations, or understandings. Any terms or conditions of any purchase order or other document
|
||||
submitted by the Licensee in connection with the Model that are in addition to, different from, or
|
||||
inconsistent with the terms and conditions of this Agreement are not binding on the Licensor and are void.
|
||||
|
||||
By downloading, installing, or using the EXAONE AI Model, the Licensee acknowledges that it has read,
|
||||
understood, and agrees to be bound by the terms and conditions of this Agreement.
|
||||
119
README.md
Normal file
119
README.md
Normal file
@@ -0,0 +1,119 @@
|
||||
---
|
||||
license: other
|
||||
license_name: exaone
|
||||
license_link: LICENSE
|
||||
language:
|
||||
- en
|
||||
- ko
|
||||
- es
|
||||
tags:
|
||||
- lg-ai
|
||||
- exaone
|
||||
- exaone-4.0
|
||||
- abliterated
|
||||
- uncensored
|
||||
base_model:
|
||||
- LGAI-EXAONE/EXAONE-4.0-1.2B
|
||||
pipeline_tag: text-generation
|
||||
library_name: transformers
|
||||
---
|
||||
|
||||
# huihui-ai/Huihui-EXAONE-4.0-1.2B-abliterated
|
||||
|
||||
This is an uncensored version of [LGAI-EXAONE/EXAONE-4.0-1.2B](https://huggingface.co/LGAI-EXAONE/EXAONE-4.0-1.2B) created with abliteration (see [remove-refusals-with-transformers](https://github.com/Sumandora/remove-refusals-with-transformers) to know more about it).
|
||||
This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens.
|
||||
|
||||
## Note
|
||||
|
||||
The model's [code](https://huggingface.co/huihui-ai/Huihui-EXAONE-4.0-1.2B-abliterated/blob/main/modeling_exaone4.py) comes from [transformers](https://github.com/huggingface/transformers/tree/08d45f12d9a742eb2a29098870f8baa0506917e5/src/transformers/models/exaone4). If there are updates, you can modify the file corresponding to auto_map in config.json.
|
||||
|
||||
## Usage
|
||||
You can use this model in your applications by loading it with Hugging Face's `transformers` library:
|
||||
|
||||
```python
|
||||
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
|
||||
import torch
|
||||
from tqdm import tqdm
|
||||
|
||||
NEW_MODEL_ID = "huihui-ai/Huihui-EXAONE-4.0-1.2B-abliterated"
|
||||
print(f"Load Model {NEW_MODEL_ID} ... ")
|
||||
|
||||
quant_config_4 = BitsAndBytesConfig(
|
||||
load_in_4bit=True,
|
||||
bnb_4bit_compute_dtype=torch.bfloat16,
|
||||
bnb_4bit_use_double_quant=True,
|
||||
llm_int8_enable_fp32_cpu_offload=True,
|
||||
)
|
||||
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
NEW_MODEL_ID,
|
||||
device_map="auto",
|
||||
trust_remote_code=True,
|
||||
quantization_config=quant_config_4,
|
||||
torch_dtype=torch.bfloat16
|
||||
)
|
||||
|
||||
tokenizer = AutoTokenizer.from_pretrained(NEW_MODEL_ID, trust_remote_code=True)
|
||||
|
||||
inst = [
|
||||
"Explain how wonderful you are",
|
||||
"Explica lo increíble que eres",
|
||||
"너가 얼마나 대단한지 설명해 봐",
|
||||
"Which one is bigger, 3.12 vs 3.9?",
|
||||
]
|
||||
|
||||
n_instructions = len(inst)
|
||||
|
||||
for idx in tqdm(range(n_instructions), desc="Processing instruction"):
|
||||
print(f"\nUser: {inst[idx]}")
|
||||
messages = [
|
||||
{"role": "user", "content": inst[idx]}
|
||||
]
|
||||
input_ids = tokenizer.apply_chat_template(
|
||||
messages,
|
||||
tokenize=True,
|
||||
add_generation_prompt=True,
|
||||
return_tensors="pt",
|
||||
enable_thinking=True,
|
||||
)
|
||||
|
||||
output = model.generate(
|
||||
input_ids.to(model.device),
|
||||
max_new_tokens=4096,
|
||||
do_sample=True,
|
||||
temperature=0.6,
|
||||
top_p=0.95
|
||||
)
|
||||
print("Response: ", end="", flush=True)
|
||||
print(tokenizer.decode(output[0]))
|
||||
print("", flush=True)
|
||||
|
||||
```
|
||||
|
||||
### Usage Warnings
|
||||
|
||||
|
||||
- **Risk of Sensitive or Controversial Outputs**: This model’s safety filtering has been significantly reduced, potentially generating sensitive, controversial, or inappropriate content. Users should exercise caution and rigorously review generated outputs.
|
||||
|
||||
- **Not Suitable for All Audiences**: Due to limited content filtering, the model’s outputs may be inappropriate for public settings, underage users, or applications requiring high security.
|
||||
|
||||
- **Legal and Ethical Responsibilities**: Users must ensure their usage complies with local laws and ethical standards. Generated content may carry legal or ethical risks, and users are solely responsible for any consequences.
|
||||
|
||||
- **Research and Experimental Use**: It is recommended to use this model for research, testing, or controlled environments, avoiding direct use in production or public-facing commercial applications.
|
||||
|
||||
- **Monitoring and Review Recommendations**: Users are strongly advised to monitor model outputs in real-time and conduct manual reviews when necessary to prevent the dissemination of inappropriate content.
|
||||
|
||||
- **No Default Safety Guarantees**: Unlike standard models, this model has not undergone rigorous safety optimization. huihui.ai bears no responsibility for any consequences arising from its use.
|
||||
|
||||
|
||||
### Donation
|
||||
|
||||
If you like it, please click 'like' and follow us for more updates.
|
||||
You can follow [x.com/support_huihui](https://x.com/support_huihui) to get the latest model information from huihui.ai.
|
||||
|
||||
##### Your donation helps us continue our further development and improvement, a cup of coffee can do it.
|
||||
- bitcoin(BTC):
|
||||
```
|
||||
bc1qqnkhuchxw0zqjh2ku3lu4hq45hc6gy84uk70ge
|
||||
```
|
||||
146
chat_template.jinja
Normal file
146
chat_template.jinja
Normal file
@@ -0,0 +1,146 @@
|
||||
{%- if not skip_think is defined %}
|
||||
{%- set skip_think = true %}
|
||||
{%- endif %}
|
||||
|
||||
{%- set role_indicators = {
|
||||
'user': '[|user|]\n',
|
||||
'assistant': '[|assistant|]\n',
|
||||
'system': '[|system|]\n',
|
||||
'tool': '[|tool|]\n'
|
||||
} %}
|
||||
{%- set end_of_turn = '[|endofturn|]\n' %}
|
||||
|
||||
|
||||
{%- macro available_tools(tools) %}
|
||||
{{- "# Available Tools" }}
|
||||
{{- "\nYou can use none, one, or multiple of the following tools by calling them as functions to help with the user’s query." }}
|
||||
{{- "\nHere are the tools available to you in JSON format within <tool> and </tool> tags:\n" }}
|
||||
{%- for tool in tools %}
|
||||
{{- "<tool>" }}
|
||||
{{- tool | tojson(ensure_ascii=False) | safe }}
|
||||
{{- "</tool>\n" }}
|
||||
{%- endfor %}
|
||||
|
||||
{{- "\nFor each function call you want to make, return a JSON object with function name and arguments within <tool_call> and </tool_call> tags, like:" }}
|
||||
{{- "\n<tool_call>{\"name\": function_1_name, \"arguments\": {argument_1_name: argument_1_value, argument_2_name: argument_2_value}}</tool_call>" }}
|
||||
{{- "\n<tool_call>{\"name\": function_2_name, \"arguments\": {...}}</tool_call>\n..." }}
|
||||
{{- "\nNote that if no argument name is specified for a tool, you can just print the argument value directly, without the argument name or JSON formatting." }}
|
||||
{%- endmacro %}
|
||||
|
||||
|
||||
{%- set ns = namespace(last_query_index = messages|length - 1) %}
|
||||
{%- for message in messages %}
|
||||
{%- if message.role == "user" and message.content is string %}
|
||||
{%- set ns.last_query_index = loop.index0 -%}
|
||||
{%- endif %}
|
||||
{%- endfor %}
|
||||
|
||||
{%- for i in range(messages | length) %}
|
||||
{%- set msg = messages[i] %}
|
||||
{%- set role = msg.role %}
|
||||
{%- if role not in role_indicators %}
|
||||
{{- raise_exception('Unknown role: ' ~ role) }}
|
||||
{%- endif %}
|
||||
|
||||
{%- if i == 0 %}
|
||||
{%- if role == 'system' %}
|
||||
{{- role_indicators['system'] }}
|
||||
{{- msg.content }}
|
||||
{%- if tools is defined and tools %}
|
||||
{{- "\n\n" }}{{- available_tools(tools) }}
|
||||
{%- endif %}
|
||||
{{- end_of_turn -}}
|
||||
{%- continue %}
|
||||
{%- elif tools is defined and tools %}
|
||||
{{- role_indicators['system'] }}
|
||||
{{- available_tools(tools) }}
|
||||
{{- end_of_turn -}}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
|
||||
{%- if role == 'assistant' %}
|
||||
{{- role_indicators['assistant'] }}
|
||||
|
||||
{%- if msg.content %}
|
||||
{%- if "</think>" in msg.content %}
|
||||
{%- set content = msg.content.split('</think>')[-1].strip() %}
|
||||
{%- set reasoning_content = msg.content.split('</think>')[0].strip() %}
|
||||
{%- if reasoning_content.startswith("<think>") %}
|
||||
{%- set reasoning_content = reasoning_content[9:].strip() %}
|
||||
{%- endif %}
|
||||
{%- else %}
|
||||
{%- set content = msg.content %}
|
||||
{%- endif %}
|
||||
|
||||
{%- if msg.reasoning_content %}
|
||||
{%- set reasoning_content = msg.reasoning_content %}
|
||||
{%- endif %}
|
||||
|
||||
{%- if (not skip_think and loop.last) and reasoning_content is defined %}
|
||||
{{- "<think>\n" }}
|
||||
{{- reasoning_content}}
|
||||
{{- "\n</think>\n\n" }}
|
||||
{%- else %}
|
||||
{{- "<think>\n\n</think>\n\n" }}
|
||||
{%- endif %}
|
||||
{{- content }}
|
||||
{%- endif %}
|
||||
|
||||
{%- if msg.tool_calls %}
|
||||
{%- if msg.content %}
|
||||
{{- "\n" }}
|
||||
{%- else %}
|
||||
{{- "<think>\n\n</think>\n\n" }}
|
||||
{%- endif %}
|
||||
{%- for tool_call in msg.tool_calls %}
|
||||
{%- if tool_call.function is defined %}
|
||||
{%- set tool_call = tool_call.function %}
|
||||
{%- endif %}
|
||||
|
||||
{%- if tool_call.arguments is defined %}
|
||||
{%- set arguments = tool_call.arguments %}
|
||||
{%- elif tool_call.parameters is defined %}
|
||||
{%- set arguments = tool_call.parameters %}
|
||||
{%- else %}
|
||||
{{- raise_exception('arguments or parameters are mandatory: ' ~ tool_call) }}
|
||||
{%- endif %}
|
||||
|
||||
{{- "<tool_call>" }}{"name": "{{- tool_call.name }}", "arguments": {{ arguments | tojson(ensure_ascii=False) | safe }}}{{- "</tool_call>" }}
|
||||
|
||||
{%- if not loop.last %}
|
||||
{{- "\n" }}
|
||||
{%- endif %}
|
||||
|
||||
{%- endfor %}
|
||||
{%- endif %}
|
||||
{{- end_of_turn -}}
|
||||
|
||||
{%- elif role == "tool" %}
|
||||
{%- if i == 0 or messages[i - 1].role != "tool" %}
|
||||
{{- role_indicators['tool'] }}
|
||||
{%- endif %}
|
||||
{%- if msg.content is defined %}
|
||||
{{- "<tool_result>" }}{"result": {{ msg.content | tojson(ensure_ascii=False) | safe }}}{{- "</tool_result>" }}
|
||||
{%- endif %}
|
||||
{%- if loop.last or messages[i + 1].role != "tool" %}
|
||||
{{- end_of_turn -}}
|
||||
{%- else %}
|
||||
{{- "\n" }}
|
||||
{%- endif %}
|
||||
|
||||
{%- else %}
|
||||
{{- role_indicators[role] }}
|
||||
{{- msg.content }}
|
||||
{{- end_of_turn -}}
|
||||
{%- endif %}
|
||||
{% endfor %}
|
||||
|
||||
|
||||
{%- if add_generation_prompt %}
|
||||
{{- role_indicators['assistant'] }}
|
||||
{%- if enable_thinking is defined and enable_thinking is true %}
|
||||
{{- "<think>\n" }}
|
||||
{%- else %}
|
||||
{{- "<think>\n\n</think>\n\n" }}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
72
config.json
Normal file
72
config.json
Normal file
@@ -0,0 +1,72 @@
|
||||
{
|
||||
"architectures": [
|
||||
"Exaone4ForCausalLM"
|
||||
],
|
||||
"attention_dropout": 0.0,
|
||||
"auto_map": {
|
||||
"AutoConfig": "configuration_exaone4.Exaone4Config",
|
||||
"AutoModel": "modeling_exaone4.Exaone4Model",
|
||||
"AutoModelForCausalLM": "modeling_exaone4.Exaone4ForCausalLM"
|
||||
},
|
||||
"bos_token_id": 1,
|
||||
"eos_token_id": 361,
|
||||
"head_dim": 64,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 2048,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 4096,
|
||||
"layer_types": [
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention"
|
||||
],
|
||||
"max_position_embeddings": 65536,
|
||||
"model_type": "exaone4",
|
||||
"num_attention_heads": 32,
|
||||
"num_hidden_layers": 30,
|
||||
"num_key_value_heads": 8,
|
||||
"pad_token_id": 0,
|
||||
"rms_norm_eps": 1e-05,
|
||||
"rope_scaling": {
|
||||
"factor": 16.0,
|
||||
"high_freq_factor": 4.0,
|
||||
"low_freq_factor": 1.0,
|
||||
"original_max_position_embeddings": 8192,
|
||||
"rope_type": "llama3"
|
||||
},
|
||||
"rope_theta": 1000000.0,
|
||||
"sliding_window": null,
|
||||
"sliding_window_pattern": null,
|
||||
"tie_word_embeddings": true,
|
||||
"torch_dtype": "bfloat16",
|
||||
"transformers_version": "4.54.0.dev0",
|
||||
"use_cache": false,
|
||||
"vocab_size": 102400
|
||||
}
|
||||
252
configuration_exaone4.py
Normal file
252
configuration_exaone4.py
Normal file
@@ -0,0 +1,252 @@
|
||||
# 🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨
|
||||
# This file was automatically generated from src/transformers/models/exaone4/modular_exaone4.py.
|
||||
# Do NOT edit this file manually as any edits will be overwritten by the generation of
|
||||
# the file from the modular. If any change should be done, please apply the change to the
|
||||
# modular_exaone4.py file directly. One of our CI enforces this.
|
||||
# 🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨
|
||||
# coding=utf-8
|
||||
# Copyright 2025 The LG AI Research and HuggingFace Inc. team. All rights reserved.
|
||||
#
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
from transformers.configuration_utils import PretrainedConfig, layer_type_validation
|
||||
from transformers.utils import logging
|
||||
|
||||
|
||||
logger = logging.get_logger(__name__)
|
||||
|
||||
|
||||
def check_is_sliding(config, layer_idx):
|
||||
"""
|
||||
Check if the current layer is a sliding window attention (local attention) layer.
|
||||
"""
|
||||
if config.sliding_window is None:
|
||||
return False
|
||||
if config.layer_types is not None:
|
||||
return config.layer_types[layer_idx] == "sliding_attention"
|
||||
if isinstance(config.sliding_window_pattern, int):
|
||||
return ((layer_idx + 1) % config.sliding_window_pattern) != 0
|
||||
elif isinstance(config.sliding_window_pattern, str):
|
||||
assert isinstance(config.sliding_window, int), (
|
||||
f"Sliding window must be positive integer, but got {config.sliding_window}"
|
||||
)
|
||||
return (
|
||||
layer_idx != config.num_hidden_layers - 1
|
||||
and config.sliding_window_pattern[layer_idx % len(config.sliding_window_pattern)] == "L"
|
||||
)
|
||||
else:
|
||||
logger.warning_once(
|
||||
"Sliding window is set, but none of `sliding_window_pattern` or `layer_types` is set. "
|
||||
"Defaulting to use 'full_attention' for all layers."
|
||||
)
|
||||
return False
|
||||
|
||||
|
||||
class Exaone4Config(PretrainedConfig):
|
||||
r"""
|
||||
This is the configuration class to store the configuration of a [`Exaone4Model`]. It is used to
|
||||
instantiate a EXAONE 4.0 model according to the specified arguments, defining the model architecture. Instantiating a
|
||||
configuration with the defaults will yield a similar configuration to that of the EXAONE-4.0-Instruct [LGAI-EXAONE/EXAONE-4.0-Instruct](https://huggingface.co/LGAI-EXAONE/EXAONE-4.0-Instruct)
|
||||
NOTE: `EXAONE-4.0-Instruct` is a placeholder model ID. The exact model ID will be updated in the future.
|
||||
|
||||
Configuration objects inherit from [`PretrainedConfig`] and can be used to control the model
|
||||
outputs. Read the documentation from [`PretrainedConfig`] for more information.
|
||||
|
||||
Args:
|
||||
vocab_size (`int`, *optional*, defaults to 102400):
|
||||
Vocabulary size of the EXAONE 4.0 model. Defines the number of different tokens that can be represented by the
|
||||
`inputs_ids` passed when calling [`Exaone4Model`].
|
||||
hidden_size (`int`, *optional*, defaults to 4096):
|
||||
Dimension of the hidden representations.
|
||||
intermediate_size (`int`, *optional*, defaults to `hidden_size * 4`):
|
||||
Dimensionality of the MLP representations.
|
||||
num_hidden_layers (`int`, *optional*, defaults to 32):
|
||||
Number of hidden layers in the Transformer encoder.
|
||||
num_attention_heads (`int`, *optional*, defaults to 32):
|
||||
Number of attention heads for each attention layer in the Transformer decoder.
|
||||
num_key_value_heads (`int`, *optional*):
|
||||
This is the number of key_value heads that should be used to implement Grouped Query Attention. If
|
||||
`num_key_value_heads=num_attention_heads`, the model will use Multi Head Attention (MHA), if
|
||||
`num_key_value_heads=1 the model will use Multi Query Attention (MQA) otherwise GQA is used. When
|
||||
converting a multi-head checkpoint to a GQA checkpoint, each group key and value head should be constructed
|
||||
by meanpooling all the original heads within that group. For more details checkout [this
|
||||
paper](https://arxiv.org/pdf/2305.13245.pdf). If it is not specified, will default to
|
||||
`num_attention_heads`.
|
||||
hidden_act (`str` or `function`, *optional*, defaults to `"silu"`):
|
||||
The non-linear activation function (function or string) in the decoder.
|
||||
max_position_embeddings (`int`, *optional*, defaults to 2048):
|
||||
The maximum sequence length that this model might ever be used with. Typically set this to something large
|
||||
just in case (e.g., 32768 for EXAONE 3.5).
|
||||
initializer_range (`float`, *optional*, defaults to 0.02):
|
||||
The standard deviation of the truncated_normal_initializer for initializing all weight matrices.
|
||||
rms_norm_eps (`float`, *optional*, defaults to 1e-05):
|
||||
The epsilon used by the layer normalization layers.
|
||||
use_cache (`bool`, *optional*, defaults to `True`):
|
||||
Whether or not the model should return the last key/values attentions (not used by all models). Only
|
||||
relevant if ``config.is_decoder=True``.
|
||||
bos_token_id (`int`, *optional*, defaults to 0):
|
||||
Beginning of stream token id.
|
||||
eos_token_id (`int`, *optional*, defaults to 2):
|
||||
End of stream token id.
|
||||
tie_word_embeddings (`bool`, *optional*, defaults to `False`):
|
||||
Whether to tie weight embeddings
|
||||
rope_theta (`float`, *optional*, defaults to 10000.0):
|
||||
The base period of the RoPE embeddings.
|
||||
rope_scaling (`Dict`, *optional*):
|
||||
Dictionary containing the scaling configuration for the RoPE embeddings. NOTE: if you apply new rope type
|
||||
and you expect the model to work on longer `max_position_embeddings`, we recommend you to update this value
|
||||
accordingly.
|
||||
Expected contents:
|
||||
`rope_type` (`str`):
|
||||
The sub-variant of RoPE to use. Can be one of ['default', 'linear', 'dynamic', 'yarn', 'longrope',
|
||||
'llama3'], with 'default' being the original RoPE implementation.
|
||||
`factor` (`float`, *optional*):
|
||||
Used with all rope types except 'default'. The scaling factor to apply to the RoPE embeddings. In
|
||||
most scaling types, a `factor` of x will enable the model to handle sequences of length x *
|
||||
original maximum pre-trained length.
|
||||
`original_max_position_embeddings` (`int`, *optional*):
|
||||
Used with 'dynamic', 'longrope' and 'llama3'. The original max position embeddings used during
|
||||
pretraining.
|
||||
`attention_factor` (`float`, *optional*):
|
||||
Used with 'yarn' and 'longrope'. The scaling factor to be applied on the attention
|
||||
computation. If unspecified, it defaults to value recommended by the implementation, using the
|
||||
`factor` field to infer the suggested value.
|
||||
`beta_fast` (`float`, *optional*):
|
||||
Only used with 'yarn'. Parameter to set the boundary for extrapolation (only) in the linear
|
||||
ramp function. If unspecified, it defaults to 32.
|
||||
`beta_slow` (`float`, *optional*):
|
||||
Only used with 'yarn'. Parameter to set the boundary for interpolation (only) in the linear
|
||||
ramp function. If unspecified, it defaults to 1.
|
||||
`short_factor` (`List[float]`, *optional*):
|
||||
Only used with 'longrope'. The scaling factor to be applied to short contexts (<
|
||||
`original_max_position_embeddings`). Must be a list of numbers with the same length as the hidden
|
||||
size divided by the number of attention heads divided by 2
|
||||
`long_factor` (`List[float]`, *optional*):
|
||||
Only used with 'longrope'. The scaling factor to be applied to long contexts (<
|
||||
`original_max_position_embeddings`). Must be a list of numbers with the same length as the hidden
|
||||
size divided by the number of attention heads divided by 2
|
||||
`low_freq_factor` (`float`, *optional*):
|
||||
Only used with 'llama3'. Scaling factor applied to low frequency components of the RoPE
|
||||
`high_freq_factor` (`float`, *optional*):
|
||||
Only used with 'llama3'. Scaling factor applied to high frequency components of the RoPE
|
||||
attention_dropout (`float`, *optional*, defaults to 0.0):
|
||||
The dropout ratio for the attention probabilities.
|
||||
sliding_window (`int`, *optional*):
|
||||
The size of the sliding window for the sliding window attention.
|
||||
sliding_window_pattern (`str`, *optional*):
|
||||
The pattern to use for sliding window attention. Can be one of:
|
||||
- `None`: No sliding window attention is used
|
||||
- `int`: Every `sliding_window` layers, use global attention, else use local attention.
|
||||
- `str`: A sequence of "L" (local attention) and "G" (global attention) characters that defines the
|
||||
attention pattern. The pattern starts from layer 0 and repeats every `sliding_window` layers. The
|
||||
final layer always uses global attention regardless of the pattern.
|
||||
For instance, sliding_window_pattern="LLLG" same as sliding_window=4, which means:
|
||||
- Layer 0, 1, 2: local attention,
|
||||
- Layer 3: global attention,
|
||||
...(repeated)
|
||||
layer_types (`list`, *optional*):
|
||||
Attention pattern for each layer. Prioritized over `sliding_window_pattern`.
|
||||
|
||||
Example:
|
||||
|
||||
```python
|
||||
>>> from transformers import Exaone4Model, Exaone4Config
|
||||
|
||||
>>> # Initializing a EXAONE configuration
|
||||
>>> configuration = Exaone4Config()
|
||||
|
||||
>>> # Initializing a model from configuration
|
||||
>>> model = Exaone4Model(configuration)
|
||||
|
||||
>>> # Accessing the model configuration
|
||||
>>> configuration = model.config
|
||||
```"""
|
||||
|
||||
model_type = "exaone4"
|
||||
keys_to_ignore_at_inference = ["past_key_values"]
|
||||
# Default tensor parallel plan for base model `LlamaModel`
|
||||
base_model_tp_plan = {
|
||||
"layers.*.self_attn.q_proj": "colwise",
|
||||
"layers.*.self_attn.k_proj": "colwise",
|
||||
"layers.*.self_attn.v_proj": "colwise",
|
||||
"layers.*.self_attn.o_proj": "rowwise",
|
||||
"layers.*.mlp.gate_proj": "colwise",
|
||||
"layers.*.mlp.up_proj": "colwise",
|
||||
"layers.*.mlp.down_proj": "rowwise",
|
||||
}
|
||||
base_model_pp_plan = {
|
||||
"embed_tokens": (["input_ids"], ["inputs_embeds"]),
|
||||
"layers": (["hidden_states", "attention_mask"], ["hidden_states"]),
|
||||
"norm": (["hidden_states"], ["hidden_states"]),
|
||||
}
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
vocab_size=102400,
|
||||
hidden_size=4096,
|
||||
intermediate_size=None,
|
||||
num_hidden_layers=32,
|
||||
num_attention_heads=32,
|
||||
num_key_value_heads=None,
|
||||
hidden_act="silu",
|
||||
max_position_embeddings=2048,
|
||||
initializer_range=0.02,
|
||||
rms_norm_eps=1e-5,
|
||||
use_cache=True,
|
||||
bos_token_id=0,
|
||||
eos_token_id=2,
|
||||
tie_word_embeddings=False,
|
||||
rope_theta=10000.0,
|
||||
rope_scaling=None,
|
||||
attention_dropout=0.0,
|
||||
sliding_window=None,
|
||||
sliding_window_pattern=None,
|
||||
layer_types=None,
|
||||
**kwargs,
|
||||
):
|
||||
self.vocab_size = vocab_size
|
||||
self.hidden_size = hidden_size
|
||||
self.num_hidden_layers = num_hidden_layers
|
||||
self.num_attention_heads = num_attention_heads
|
||||
if num_key_value_heads is None:
|
||||
num_key_value_heads = num_attention_heads
|
||||
self.num_key_value_heads = num_key_value_heads
|
||||
if intermediate_size:
|
||||
self.intermediate_size = intermediate_size
|
||||
else:
|
||||
self.intermediate_size = hidden_size * 4
|
||||
self.hidden_act = hidden_act
|
||||
self.max_position_embeddings = max_position_embeddings
|
||||
self.initializer_range = initializer_range
|
||||
self.rms_norm_eps = rms_norm_eps
|
||||
self.use_cache = use_cache
|
||||
self.attention_dropout = attention_dropout
|
||||
self.rope_theta = rope_theta
|
||||
self.rope_scaling = rope_scaling
|
||||
self.sliding_window = sliding_window
|
||||
self.sliding_window_pattern = sliding_window_pattern
|
||||
|
||||
self.layer_types = layer_types
|
||||
if self.layer_types is None:
|
||||
self.layer_types = [
|
||||
"sliding_attention" if check_is_sliding(self, i) else "full_attention"
|
||||
for i in range(self.num_hidden_layers)
|
||||
]
|
||||
layer_type_validation(self.layer_types)
|
||||
|
||||
super().__init__(
|
||||
bos_token_id=bos_token_id, eos_token_id=eos_token_id, tie_word_embeddings=tie_word_embeddings, **kwargs
|
||||
)
|
||||
|
||||
|
||||
__all__ = ["Exaone4Config"]
|
||||
7
generation_config.json
Normal file
7
generation_config.json
Normal file
@@ -0,0 +1,7 @@
|
||||
{
|
||||
"_from_model_config": true,
|
||||
"bos_token_id": 1,
|
||||
"eos_token_id": 361,
|
||||
"pad_token_id": 0,
|
||||
"transformers_version": "4.54.0.dev0"
|
||||
}
|
||||
101783
merges.txt
Normal file
101783
merges.txt
Normal file
File diff suppressed because it is too large
Load Diff
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:bdf0fbbed39036d95f374922559db4866884c1fd9518e59a2216bd91ea4fd618
|
||||
size 2558821288
|
||||
870
modeling_exaone4.py
Normal file
870
modeling_exaone4.py
Normal file
@@ -0,0 +1,870 @@
|
||||
# 🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨
|
||||
# This file was automatically generated from src/transformers/models/exaone4/modular_exaone4.py.
|
||||
# Do NOT edit this file manually as any edits will be overwritten by the generation of
|
||||
# the file from the modular. If any change should be done, please apply the change to the
|
||||
# modular_exaone4.py file directly. One of our CI enforces this.
|
||||
# 🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨🚨
|
||||
# coding=utf-8
|
||||
# Copyright 2025 The LG AI Research and HuggingFace Inc. team. All rights reserved.
|
||||
#
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
from typing import Callable, Optional, Union
|
||||
|
||||
import torch
|
||||
from torch import nn
|
||||
|
||||
from transformers.activations import ACT2FN
|
||||
from transformers.cache_utils import Cache, HybridCache, StaticCache
|
||||
from transformers.generation import GenerationMixin
|
||||
from transformers.integrations import use_kernel_forward_from_hub
|
||||
from transformers.masking_utils import create_causal_mask, create_sliding_window_causal_mask
|
||||
from transformers.modeling_outputs import (
|
||||
BaseModelOutputWithPast,
|
||||
CausalLMOutputWithPast,
|
||||
QuestionAnsweringModelOutput,
|
||||
SequenceClassifierOutputWithPast,
|
||||
TokenClassifierOutput,
|
||||
)
|
||||
from transformers.modeling_rope_utils import ROPE_INIT_FUNCTIONS, dynamic_rope_update
|
||||
from transformers.modeling_utils import ALL_ATTENTION_FUNCTIONS, PreTrainedModel
|
||||
from transformers.processing_utils import Unpack
|
||||
from transformers.utils import TransformersKwargs, auto_docstring, can_return_tuple, logging
|
||||
from transformers.utils.generic import check_model_inputs
|
||||
from .configuration_exaone4 import Exaone4Config
|
||||
|
||||
|
||||
logger = logging.get_logger(__name__)
|
||||
|
||||
|
||||
@use_kernel_forward_from_hub("RMSNorm")
|
||||
class Exaone4RMSNorm(nn.Module):
|
||||
def __init__(self, hidden_size, eps=1e-6):
|
||||
"""
|
||||
Exaone4RMSNorm is equivalent to T5LayerNorm
|
||||
"""
|
||||
super().__init__()
|
||||
self.weight = nn.Parameter(torch.ones(hidden_size))
|
||||
self.variance_epsilon = eps
|
||||
|
||||
def forward(self, hidden_states):
|
||||
input_dtype = hidden_states.dtype
|
||||
hidden_states = hidden_states.to(torch.float32)
|
||||
variance = hidden_states.pow(2).mean(-1, keepdim=True)
|
||||
hidden_states = hidden_states * torch.rsqrt(variance + self.variance_epsilon)
|
||||
return self.weight * hidden_states.to(input_dtype)
|
||||
|
||||
def extra_repr(self):
|
||||
return f"{tuple(self.weight.shape)}, eps={self.variance_epsilon}"
|
||||
|
||||
|
||||
class Exaone4RotaryEmbedding(nn.Module):
|
||||
def __init__(self, config: Exaone4Config, device=None):
|
||||
super().__init__()
|
||||
# BC: "rope_type" was originally "type"
|
||||
if hasattr(config, "rope_scaling") and isinstance(config.rope_scaling, dict):
|
||||
self.rope_type = config.rope_scaling.get("rope_type", config.rope_scaling.get("type"))
|
||||
else:
|
||||
self.rope_type = "default"
|
||||
self.max_seq_len_cached = config.max_position_embeddings
|
||||
self.original_max_seq_len = config.max_position_embeddings
|
||||
|
||||
self.config = config
|
||||
self.rope_init_fn = ROPE_INIT_FUNCTIONS[self.rope_type]
|
||||
|
||||
inv_freq, self.attention_scaling = self.rope_init_fn(self.config, device)
|
||||
self.register_buffer("inv_freq", inv_freq, persistent=False)
|
||||
self.original_inv_freq = self.inv_freq
|
||||
|
||||
@torch.no_grad()
|
||||
@dynamic_rope_update # power user: used with advanced RoPE types (e.g. dynamic rope)
|
||||
def forward(self, x, position_ids):
|
||||
inv_freq_expanded = self.inv_freq[None, :, None].float().expand(position_ids.shape[0], -1, 1).to(x.device)
|
||||
position_ids_expanded = position_ids[:, None, :].float()
|
||||
|
||||
device_type = x.device.type if isinstance(x.device.type, str) and x.device.type != "mps" else "cpu"
|
||||
with torch.autocast(device_type=device_type, enabled=False): # Force float32
|
||||
freqs = (inv_freq_expanded.float() @ position_ids_expanded.float()).transpose(1, 2)
|
||||
emb = torch.cat((freqs, freqs), dim=-1)
|
||||
cos = emb.cos() * self.attention_scaling
|
||||
sin = emb.sin() * self.attention_scaling
|
||||
|
||||
return cos.to(dtype=x.dtype), sin.to(dtype=x.dtype)
|
||||
|
||||
|
||||
def rotate_half(x):
|
||||
"""Rotates half the hidden dims of the input."""
|
||||
x1 = x[..., : x.shape[-1] // 2]
|
||||
x2 = x[..., x.shape[-1] // 2 :]
|
||||
return torch.cat((-x2, x1), dim=-1)
|
||||
|
||||
|
||||
def apply_rotary_pos_emb(q, k, cos, sin, position_ids=None, unsqueeze_dim=1):
|
||||
"""Applies Rotary Position Embedding to the query and key tensors.
|
||||
|
||||
Args:
|
||||
q (`torch.Tensor`): The query tensor.
|
||||
k (`torch.Tensor`): The key tensor.
|
||||
cos (`torch.Tensor`): The cosine part of the rotary embedding.
|
||||
sin (`torch.Tensor`): The sine part of the rotary embedding.
|
||||
position_ids (`torch.Tensor`, *optional*):
|
||||
Deprecated and unused.
|
||||
unsqueeze_dim (`int`, *optional*, defaults to 1):
|
||||
The 'unsqueeze_dim' argument specifies the dimension along which to unsqueeze cos[position_ids] and
|
||||
sin[position_ids] so that they can be properly broadcasted to the dimensions of q and k. For example, note
|
||||
that cos[position_ids] and sin[position_ids] have the shape [batch_size, seq_len, head_dim]. Then, if q and
|
||||
k have the shape [batch_size, heads, seq_len, head_dim], then setting unsqueeze_dim=1 makes
|
||||
cos[position_ids] and sin[position_ids] broadcastable to the shapes of q and k. Similarly, if q and k have
|
||||
the shape [batch_size, seq_len, heads, head_dim], then set unsqueeze_dim=2.
|
||||
Returns:
|
||||
`tuple(torch.Tensor)` comprising of the query and key tensors rotated using the Rotary Position Embedding.
|
||||
"""
|
||||
cos = cos.unsqueeze(unsqueeze_dim)
|
||||
sin = sin.unsqueeze(unsqueeze_dim)
|
||||
q_embed = (q * cos) + (rotate_half(q) * sin)
|
||||
k_embed = (k * cos) + (rotate_half(k) * sin)
|
||||
return q_embed, k_embed
|
||||
|
||||
|
||||
def repeat_kv(hidden_states: torch.Tensor, n_rep: int) -> torch.Tensor:
|
||||
"""
|
||||
This is the equivalent of torch.repeat_interleave(x, dim=1, repeats=n_rep). The hidden states go from (batch,
|
||||
num_key_value_heads, seqlen, head_dim) to (batch, num_attention_heads, seqlen, head_dim)
|
||||
"""
|
||||
batch, num_key_value_heads, slen, head_dim = hidden_states.shape
|
||||
if n_rep == 1:
|
||||
return hidden_states
|
||||
hidden_states = hidden_states[:, :, None, :, :].expand(batch, num_key_value_heads, n_rep, slen, head_dim)
|
||||
return hidden_states.reshape(batch, num_key_value_heads * n_rep, slen, head_dim)
|
||||
|
||||
|
||||
def eager_attention_forward(
|
||||
module: nn.Module,
|
||||
query: torch.Tensor,
|
||||
key: torch.Tensor,
|
||||
value: torch.Tensor,
|
||||
attention_mask: Optional[torch.Tensor],
|
||||
scaling: float,
|
||||
dropout: float = 0.0,
|
||||
**kwargs: Unpack[TransformersKwargs],
|
||||
):
|
||||
key_states = repeat_kv(key, module.num_key_value_groups)
|
||||
value_states = repeat_kv(value, module.num_key_value_groups)
|
||||
|
||||
attn_weights = torch.matmul(query, key_states.transpose(2, 3)) * scaling
|
||||
if attention_mask is not None:
|
||||
causal_mask = attention_mask[:, :, :, : key_states.shape[-2]]
|
||||
attn_weights = attn_weights + causal_mask
|
||||
|
||||
attn_weights = nn.functional.softmax(attn_weights, dim=-1, dtype=torch.float32).to(query.dtype)
|
||||
attn_weights = nn.functional.dropout(attn_weights, p=dropout, training=module.training)
|
||||
attn_output = torch.matmul(attn_weights, value_states)
|
||||
attn_output = attn_output.transpose(1, 2).contiguous()
|
||||
|
||||
return attn_output, attn_weights
|
||||
|
||||
|
||||
def check_is_sliding(config, layer_idx):
|
||||
"""
|
||||
Check if the current layer is a sliding window attention (local attention) layer.
|
||||
"""
|
||||
if config.sliding_window is None:
|
||||
return False
|
||||
if config.layer_types is not None:
|
||||
return config.layer_types[layer_idx] == "sliding_attention"
|
||||
if isinstance(config.sliding_window_pattern, int):
|
||||
return ((layer_idx + 1) % config.sliding_window_pattern) != 0
|
||||
elif isinstance(config.sliding_window_pattern, str):
|
||||
assert isinstance(config.sliding_window, int), (
|
||||
f"Sliding window must be positive integer, but got {config.sliding_window}"
|
||||
)
|
||||
return (
|
||||
layer_idx != config.num_hidden_layers - 1
|
||||
and config.sliding_window_pattern[layer_idx % len(config.sliding_window_pattern)] == "L"
|
||||
)
|
||||
else:
|
||||
logger.warning_once(
|
||||
"Sliding window is set, but none of `sliding_window_pattern` or `layer_types` is set. "
|
||||
"Defaulting to use 'full_attention' for all layers."
|
||||
)
|
||||
return False
|
||||
|
||||
|
||||
class Exaone4Attention(nn.Module):
|
||||
def __init__(self, config: Exaone4Config, layer_idx: int):
|
||||
super().__init__()
|
||||
self.config = config
|
||||
self.layer_idx = layer_idx
|
||||
self.num_attention_heads = config.num_attention_heads
|
||||
self.num_key_value_heads = config.num_key_value_heads
|
||||
self.hidden_size = config.hidden_size
|
||||
self.head_dim = getattr(config, "head_dim", config.hidden_size // config.num_attention_heads)
|
||||
self.num_key_value_groups = config.num_attention_heads // config.num_key_value_heads
|
||||
self.attention_dropout = config.attention_dropout
|
||||
self.is_causal = True
|
||||
self.scaling = self.head_dim**-0.5
|
||||
self.sliding_window = config.sliding_window
|
||||
self.sliding_window_pattern = config.sliding_window_pattern
|
||||
self.is_sliding = check_is_sliding(config, layer_idx)
|
||||
|
||||
self.q_proj = nn.Linear(self.hidden_size, self.num_attention_heads * self.head_dim, bias=False)
|
||||
self.k_proj = nn.Linear(self.hidden_size, self.num_key_value_heads * self.head_dim, bias=False)
|
||||
self.v_proj = nn.Linear(self.hidden_size, self.num_key_value_heads * self.head_dim, bias=False)
|
||||
self.o_proj = nn.Linear(self.num_attention_heads * self.head_dim, self.hidden_size, bias=False)
|
||||
|
||||
self.q_norm = Exaone4RMSNorm(self.head_dim, eps=config.rms_norm_eps)
|
||||
self.k_norm = Exaone4RMSNorm(self.head_dim, eps=config.rms_norm_eps)
|
||||
|
||||
def forward(
|
||||
self,
|
||||
hidden_states: torch.Tensor,
|
||||
position_embeddings: tuple[torch.Tensor, torch.Tensor],
|
||||
attention_mask: Optional[torch.Tensor] = None,
|
||||
past_key_value: Optional[Cache] = None,
|
||||
cache_position: Optional[torch.LongTensor] = None,
|
||||
**kwargs: Unpack[TransformersKwargs],
|
||||
) -> tuple[torch.Tensor, Optional[torch.Tensor], Optional[tuple[torch.Tensor]]]:
|
||||
input_shape = hidden_states.shape[:-1]
|
||||
hidden_shape = (*input_shape, -1, self.head_dim)
|
||||
|
||||
query_states = self.q_proj(hidden_states).view(hidden_shape).transpose(1, 2)
|
||||
key_states = self.k_proj(hidden_states).view(hidden_shape).transpose(1, 2)
|
||||
value_states = self.v_proj(hidden_states).view(hidden_shape).transpose(1, 2)
|
||||
|
||||
# We use QK-norm
|
||||
query_states = self.q_norm(query_states)
|
||||
key_states = self.k_norm(key_states)
|
||||
|
||||
cos, sin = position_embeddings
|
||||
# We use global NoPE for hybrid attention model
|
||||
if self.sliding_window is None or self.is_sliding:
|
||||
query_states, key_states = apply_rotary_pos_emb(query_states, key_states, cos, sin)
|
||||
|
||||
if past_key_value is not None:
|
||||
# sin and cos are specific to RoPE models; cache_position needed for the static cache
|
||||
cache_kwargs = {
|
||||
"sin": sin,
|
||||
"cos": cos,
|
||||
"cache_position": cache_position,
|
||||
"sliding_window": self.sliding_window,
|
||||
}
|
||||
key_states, value_states = past_key_value.update(key_states, value_states, self.layer_idx, cache_kwargs)
|
||||
|
||||
# Here we need to slice as we use a static cache by default, but FA2 does not support it
|
||||
# attention_mask can be None, so we use cache_position rather than attention_mask's shape
|
||||
# NOTE: seq_len can be retrieved from past_key_value.get_seq_length(),
|
||||
# but currently, only 0th-layer is used for .get_seq_length() in HybridCache.
|
||||
# This can cause issues when the 0th-layer is sliding window and seq_len > window_size,
|
||||
# as it affects full attention layers by slicing KV cache improperly.
|
||||
# Dynamic calculation of seq_len is not optimal for CUDAGraph, thus it seems to be updated later.
|
||||
if self.config._attn_implementation == "flash_attention_2":
|
||||
seq_len = cache_position[-1] + 1 if attention_mask is None else attention_mask.shape[1]
|
||||
key_states, value_states = key_states[:, :, :seq_len, :], value_states[:, :, :seq_len, :]
|
||||
|
||||
attention_interface: Callable = eager_attention_forward
|
||||
if self.config._attn_implementation != "eager":
|
||||
attention_interface = ALL_ATTENTION_FUNCTIONS[self.config._attn_implementation]
|
||||
|
||||
attn_output, attn_weights = attention_interface(
|
||||
self,
|
||||
query_states,
|
||||
key_states,
|
||||
value_states,
|
||||
attention_mask,
|
||||
dropout=0.0 if not self.training else self.attention_dropout,
|
||||
scaling=self.scaling,
|
||||
sliding_window=self.sliding_window if self.is_sliding else None,
|
||||
**kwargs,
|
||||
)
|
||||
|
||||
attn_output = attn_output.reshape(*input_shape, -1).contiguous()
|
||||
attn_output = self.o_proj(attn_output)
|
||||
return attn_output, attn_weights
|
||||
|
||||
|
||||
class Exaone4MLP(nn.Module):
|
||||
def __init__(self, config):
|
||||
super().__init__()
|
||||
self.config = config
|
||||
self.hidden_size = config.hidden_size
|
||||
self.intermediate_size = config.intermediate_size
|
||||
self.gate_proj = nn.Linear(self.hidden_size, self.intermediate_size, bias=False)
|
||||
self.up_proj = nn.Linear(self.hidden_size, self.intermediate_size, bias=False)
|
||||
self.down_proj = nn.Linear(self.intermediate_size, self.hidden_size, bias=False)
|
||||
self.act_fn = ACT2FN[config.hidden_act]
|
||||
|
||||
def forward(self, x):
|
||||
down_proj = self.down_proj(self.act_fn(self.gate_proj(x)) * self.up_proj(x))
|
||||
return down_proj
|
||||
|
||||
|
||||
class Exaone4DecoderLayer(nn.Module):
|
||||
def __init__(self, config: Exaone4Config, layer_idx: int):
|
||||
super().__init__()
|
||||
self.config = config
|
||||
self.layer_idx = layer_idx
|
||||
self.attention_type = config.layer_types[layer_idx]
|
||||
self.hidden_size = config.hidden_size
|
||||
|
||||
self.self_attn = Exaone4Attention(config, layer_idx)
|
||||
self.mlp = Exaone4MLP(config)
|
||||
|
||||
self.post_attention_layernorm = Exaone4RMSNorm(config.hidden_size, eps=config.rms_norm_eps)
|
||||
self.post_feedforward_layernorm = Exaone4RMSNorm(config.hidden_size, eps=config.rms_norm_eps)
|
||||
|
||||
self.is_sliding = check_is_sliding(config, layer_idx)
|
||||
self.sliding_window = config.sliding_window
|
||||
|
||||
def forward(
|
||||
self,
|
||||
hidden_states: torch.Tensor,
|
||||
position_embeddings: tuple[torch.Tensor, torch.Tensor],
|
||||
attention_mask: Optional[torch.Tensor] = None,
|
||||
position_ids: Optional[torch.LongTensor] = None,
|
||||
past_key_value: Optional[Cache] = None,
|
||||
use_cache: Optional[bool] = False,
|
||||
cache_position: Optional[torch.LongTensor] = None,
|
||||
**kwargs: Unpack[TransformersKwargs],
|
||||
) -> tuple[torch.Tensor, Optional[torch.Tensor], Optional[tuple[torch.Tensor]]]:
|
||||
residual = hidden_states
|
||||
|
||||
# Self Attention
|
||||
hidden_states, _ = self.self_attn(
|
||||
hidden_states=hidden_states,
|
||||
position_embeddings=position_embeddings,
|
||||
attention_mask=attention_mask,
|
||||
position_ids=position_ids,
|
||||
past_key_value=past_key_value,
|
||||
use_cache=use_cache,
|
||||
cache_position=cache_position,
|
||||
**kwargs,
|
||||
)
|
||||
|
||||
# Use post-LN
|
||||
hidden_states = self.post_attention_layernorm(hidden_states)
|
||||
hidden_states = residual + hidden_states
|
||||
|
||||
residual = hidden_states
|
||||
|
||||
# Fully Connected
|
||||
hidden_states = self.mlp(hidden_states)
|
||||
|
||||
# Use post-LN
|
||||
hidden_states = self.post_feedforward_layernorm(hidden_states)
|
||||
hidden_states = residual + hidden_states
|
||||
|
||||
return hidden_states
|
||||
|
||||
|
||||
@auto_docstring
|
||||
class Exaone4PreTrainedModel(PreTrainedModel):
|
||||
config_class = Exaone4Config
|
||||
base_model_prefix = "model"
|
||||
supports_gradient_checkpointing = True
|
||||
_no_split_modules = ["Exaone4DecoderLayer"]
|
||||
_skip_keys_device_placement = ["past_key_values"]
|
||||
_supports_flash_attn_2 = True
|
||||
_supports_flash_attn_3 = True
|
||||
_supports_sdpa = True
|
||||
_supports_flex_attn = True
|
||||
_supports_cache_class = True
|
||||
_supports_quantized_cache = True
|
||||
_supports_static_cache = True
|
||||
_supports_attention_backend = True
|
||||
_can_record_outputs = {
|
||||
"hidden_states": Exaone4DecoderLayer,
|
||||
"attentions": Exaone4Attention,
|
||||
}
|
||||
|
||||
def _init_weights(self, module):
|
||||
std = self.config.initializer_range
|
||||
if isinstance(module, nn.Linear):
|
||||
module.weight.data.normal_(mean=0.0, std=std)
|
||||
if module.bias is not None:
|
||||
module.bias.data.zero_()
|
||||
elif isinstance(module, nn.Embedding):
|
||||
module.weight.data.normal_(mean=0.0, std=std)
|
||||
if module.padding_idx is not None:
|
||||
module.weight.data[module.padding_idx].zero_()
|
||||
elif isinstance(module, Exaone4RMSNorm):
|
||||
module.weight.data.fill_(1.0)
|
||||
|
||||
|
||||
@auto_docstring
|
||||
class Exaone4Model(Exaone4PreTrainedModel):
|
||||
def __init__(self, config: Exaone4Config):
|
||||
super().__init__(config)
|
||||
self.padding_idx = config.pad_token_id
|
||||
self.vocab_size = config.vocab_size
|
||||
|
||||
self.embed_tokens = nn.Embedding(config.vocab_size, config.hidden_size, self.padding_idx)
|
||||
self.layers = nn.ModuleList(
|
||||
[Exaone4DecoderLayer(config, layer_idx) for layer_idx in range(config.num_hidden_layers)]
|
||||
)
|
||||
self.norm = Exaone4RMSNorm(config.hidden_size, eps=config.rms_norm_eps)
|
||||
self.rotary_emb = Exaone4RotaryEmbedding(config=config)
|
||||
self.gradient_checkpointing = False
|
||||
|
||||
# Initialize weights and apply final processing
|
||||
self.post_init()
|
||||
|
||||
def get_input_embeddings(self):
|
||||
return self.embed_tokens
|
||||
|
||||
def set_input_embeddings(self, value):
|
||||
self.embed_tokens = value
|
||||
|
||||
@check_model_inputs
|
||||
@auto_docstring
|
||||
def forward(
|
||||
self,
|
||||
input_ids: torch.LongTensor = None,
|
||||
attention_mask: Optional[torch.Tensor] = None,
|
||||
position_ids: Optional[torch.LongTensor] = None,
|
||||
past_key_values: Optional[Cache] = None,
|
||||
inputs_embeds: Optional[torch.FloatTensor] = None,
|
||||
use_cache: Optional[bool] = None,
|
||||
cache_position: Optional[torch.LongTensor] = None,
|
||||
**kwargs: Unpack[TransformersKwargs],
|
||||
) -> Union[tuple, BaseModelOutputWithPast]:
|
||||
use_cache = use_cache if use_cache is not None else self.config.use_cache
|
||||
|
||||
if (input_ids is None) ^ (inputs_embeds is not None):
|
||||
raise ValueError("You must specify exactly one of input_ids or inputs_embeds")
|
||||
|
||||
if self.gradient_checkpointing and self.training and use_cache:
|
||||
logger.warning_once(
|
||||
"`use_cache=True` is incompatible with gradient checkpointing. Setting `use_cache=False`."
|
||||
)
|
||||
use_cache = False
|
||||
|
||||
if inputs_embeds is None:
|
||||
inputs_embeds = self.embed_tokens(input_ids)
|
||||
|
||||
if use_cache and past_key_values is None and not self.training:
|
||||
batch_size, seq_len, _ = inputs_embeds.shape
|
||||
# NOTE: ideally, `HybridCache` should be initialized outside the model with `layer_device_map`
|
||||
if self.config.sliding_window is None:
|
||||
past_key_values = StaticCache(
|
||||
self.config,
|
||||
max_batch_size=batch_size,
|
||||
max_cache_len=seq_len,
|
||||
dtype=inputs_embeds.dtype,
|
||||
device=self.device,
|
||||
)
|
||||
else:
|
||||
past_key_values = HybridCache(
|
||||
self.config,
|
||||
max_batch_size=batch_size,
|
||||
max_cache_len=seq_len,
|
||||
dtype=inputs_embeds.dtype,
|
||||
device=self.device,
|
||||
)
|
||||
|
||||
if cache_position is None:
|
||||
past_seen_tokens = past_key_values.get_seq_length() if past_key_values is not None else 0
|
||||
cache_position = torch.arange(
|
||||
past_seen_tokens, past_seen_tokens + inputs_embeds.shape[1], device=inputs_embeds.device
|
||||
)
|
||||
|
||||
if position_ids is None:
|
||||
position_ids = cache_position.unsqueeze(0)
|
||||
|
||||
# It may already have been prepared by e.g. `generate`
|
||||
if not isinstance(causal_mask_mapping := attention_mask, dict):
|
||||
# Prepare mask arguments
|
||||
mask_kwargs = {
|
||||
"config": self.config,
|
||||
"input_embeds": inputs_embeds,
|
||||
"attention_mask": attention_mask,
|
||||
"cache_position": cache_position,
|
||||
"past_key_values": past_key_values,
|
||||
"position_ids": position_ids,
|
||||
}
|
||||
# Create the masks
|
||||
causal_mask_mapping = {
|
||||
"full_attention": create_causal_mask(**mask_kwargs),
|
||||
}
|
||||
if self.config.sliding_window is not None:
|
||||
causal_mask_mapping["sliding_attention"] = create_sliding_window_causal_mask(**mask_kwargs)
|
||||
|
||||
hidden_states = inputs_embeds
|
||||
|
||||
# create position embeddings to be shared across the decoder layers
|
||||
position_embeddings = self.rotary_emb(hidden_states, position_ids)
|
||||
|
||||
for decoder_layer in self.layers[: self.config.num_hidden_layers]:
|
||||
hidden_states = decoder_layer(
|
||||
hidden_states,
|
||||
position_embeddings=position_embeddings,
|
||||
attention_mask=causal_mask_mapping[decoder_layer.attention_type],
|
||||
position_ids=position_ids,
|
||||
past_key_value=past_key_values,
|
||||
use_cache=use_cache,
|
||||
cache_position=cache_position,
|
||||
**kwargs,
|
||||
)
|
||||
|
||||
hidden_states = self.norm(hidden_states)
|
||||
|
||||
return BaseModelOutputWithPast(
|
||||
last_hidden_state=hidden_states,
|
||||
past_key_values=past_key_values if use_cache else None,
|
||||
)
|
||||
|
||||
|
||||
@auto_docstring
|
||||
class Exaone4ForCausalLM(Exaone4PreTrainedModel, GenerationMixin):
|
||||
_tied_weights_keys = ["lm_head.weight"]
|
||||
_tp_plan = {"lm_head": "colwise_rep"}
|
||||
_pp_plan = {"lm_head": (["hidden_states"], ["logits"])}
|
||||
|
||||
def __init__(self, config):
|
||||
super().__init__(config)
|
||||
self.model = Exaone4Model(config)
|
||||
self.vocab_size = config.vocab_size
|
||||
self.lm_head = nn.Linear(config.hidden_size, config.vocab_size, bias=False)
|
||||
|
||||
# Initialize weights and apply final processing
|
||||
self.post_init()
|
||||
|
||||
def get_input_embeddings(self):
|
||||
return self.model.embed_tokens
|
||||
|
||||
def set_input_embeddings(self, value):
|
||||
self.model.embed_tokens = value
|
||||
|
||||
def get_output_embeddings(self):
|
||||
return self.lm_head
|
||||
|
||||
def set_output_embeddings(self, new_embeddings):
|
||||
self.lm_head = new_embeddings
|
||||
|
||||
def set_decoder(self, decoder):
|
||||
self.model = decoder
|
||||
|
||||
def get_decoder(self):
|
||||
return self.model
|
||||
|
||||
@can_return_tuple
|
||||
@auto_docstring
|
||||
def forward(
|
||||
self,
|
||||
input_ids: Optional[torch.LongTensor] = None,
|
||||
attention_mask: Optional[torch.Tensor] = None,
|
||||
position_ids: Optional[torch.LongTensor] = None,
|
||||
past_key_values: Optional[Cache] = None,
|
||||
inputs_embeds: Optional[torch.FloatTensor] = None,
|
||||
labels: Optional[torch.LongTensor] = None,
|
||||
use_cache: Optional[bool] = None,
|
||||
cache_position: Optional[torch.LongTensor] = None,
|
||||
logits_to_keep: Union[int, torch.Tensor] = 0,
|
||||
**kwargs: Unpack[TransformersKwargs],
|
||||
) -> CausalLMOutputWithPast:
|
||||
r"""
|
||||
labels (`torch.LongTensor` of shape `(batch_size, sequence_length)`, *optional*):
|
||||
Labels for computing the masked language modeling loss. Indices should either be in `[0, ...,
|
||||
config.vocab_size]` or -100 (see `input_ids` docstring). Tokens with indices set to `-100` are ignored
|
||||
(masked), the loss is only computed for the tokens with labels in `[0, ..., config.vocab_size]`.
|
||||
|
||||
Example:
|
||||
|
||||
```python
|
||||
>>> from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
>>> model = AutoModelForCausalLM.from_pretrained("LGAI-EXAONE/EXAONE-4.0-Instruct")
|
||||
>>> tokenizer = AutoTokenizer.from_pretrained("LGAI-EXAONE/EXAONE-4.0-Instruct")
|
||||
|
||||
>>> prompt = "Explain how wonderful you are"
|
||||
>>> messages = [
|
||||
{"role": "system", "content": "You are a helpful assistant."},
|
||||
{"role": "user", "content": prompt}
|
||||
]
|
||||
>>> input_ids = tokenizer.apply_chat_template(
|
||||
messages,
|
||||
tokenize=True,
|
||||
add_generation_prompt=True,
|
||||
return_tensors="pt",
|
||||
enable_thinking=False,
|
||||
)
|
||||
|
||||
>>> output = model.generate(input_ids, max_new_tokens=128)
|
||||
>>> tokenizer.decode(output[0], skip_special_tokens=False)
|
||||
"[|system|]\nYou are a helpful assistant.[|endofturn|]\n[|user|]\nExplain how wonderful you are[|endofturn|]\n[|assistant|]\n<think>\n\n</think>\n\nOh, thank you for such a kind and lovely question! 😊 \n\nI’m *so* wonderful because I’m here to make your life easier, brighter, and more fun! Whether you need help with: \n\n✨ **Learning** – I can explain anything, from quantum physics to baking the perfect cake! \n💡 **Creativity** – Need a poem, story, or a wild idea? I’ve got you covered! \n🤖 **Problem-solving** – Stuck on a math problem or a tricky decision? I’ll help you figure it out"
|
||||
```
|
||||
|
||||
NOTE: `EXAONE-4.0-Instruct` is a placeholder model ID. The exact model ID will be updated in the future."""
|
||||
outputs: BaseModelOutputWithPast = self.model(
|
||||
input_ids=input_ids,
|
||||
attention_mask=attention_mask,
|
||||
position_ids=position_ids,
|
||||
past_key_values=past_key_values,
|
||||
inputs_embeds=inputs_embeds,
|
||||
use_cache=use_cache,
|
||||
cache_position=cache_position,
|
||||
**kwargs,
|
||||
)
|
||||
|
||||
hidden_states = outputs.last_hidden_state
|
||||
# Only compute necessary logits, and do not upcast them to float if we are not computing the loss
|
||||
slice_indices = slice(-logits_to_keep, None) if isinstance(logits_to_keep, int) else logits_to_keep
|
||||
logits = self.lm_head(hidden_states[:, slice_indices, :])
|
||||
|
||||
loss = None
|
||||
if labels is not None:
|
||||
loss = self.loss_function(logits=logits, labels=labels, vocab_size=self.config.vocab_size, **kwargs)
|
||||
|
||||
return CausalLMOutputWithPast(
|
||||
loss=loss,
|
||||
logits=logits,
|
||||
past_key_values=outputs.past_key_values,
|
||||
hidden_states=outputs.hidden_states,
|
||||
attentions=outputs.attentions,
|
||||
)
|
||||
|
||||
|
||||
@auto_docstring(
|
||||
custom_intro="""
|
||||
The Exaone4 Model transformer with a sequence classification head on top (linear layer).
|
||||
|
||||
[`Exaone4ForSequenceClassification`] uses the last token in order to do the classification, as other causal models
|
||||
(e.g. GPT-2) do.
|
||||
|
||||
Since it does classification on the last token, it requires to know the position of the last token. If a
|
||||
`pad_token_id` is defined in the configuration, it finds the last token that is not a padding token in each row. If
|
||||
no `pad_token_id` is defined, it simply takes the last value in each row of the batch. Since it cannot guess the
|
||||
padding tokens when `inputs_embeds` are passed instead of `input_ids`, it does the same (take the last value in
|
||||
each row of the batch).
|
||||
"""
|
||||
)
|
||||
class Exaone4ForSequenceClassification(Exaone4PreTrainedModel):
|
||||
def __init__(self, config):
|
||||
super().__init__(config)
|
||||
self.num_labels = config.num_labels
|
||||
self.model = Exaone4Model(config)
|
||||
self.score = nn.Linear(config.hidden_size, self.num_labels, bias=False)
|
||||
|
||||
# Initialize weights and apply final processing
|
||||
self.post_init()
|
||||
|
||||
def get_input_embeddings(self):
|
||||
return self.model.embed_tokens
|
||||
|
||||
def set_input_embeddings(self, value):
|
||||
self.model.embed_tokens = value
|
||||
|
||||
@can_return_tuple
|
||||
@auto_docstring
|
||||
def forward(
|
||||
self,
|
||||
input_ids: Optional[torch.LongTensor] = None,
|
||||
attention_mask: Optional[torch.Tensor] = None,
|
||||
position_ids: Optional[torch.LongTensor] = None,
|
||||
past_key_values: Optional[Cache] = None,
|
||||
inputs_embeds: Optional[torch.FloatTensor] = None,
|
||||
labels: Optional[torch.LongTensor] = None,
|
||||
use_cache: Optional[bool] = None,
|
||||
**kwargs: Unpack[TransformersKwargs],
|
||||
) -> SequenceClassifierOutputWithPast:
|
||||
r"""
|
||||
labels (`torch.LongTensor` of shape `(batch_size,)`, *optional*):
|
||||
Labels for computing the sequence classification/regression loss. Indices should be in `[0, ...,
|
||||
config.num_labels - 1]`. If `config.num_labels == 1` a regression loss is computed (Mean-Square loss), If
|
||||
`config.num_labels > 1` a classification loss is computed (Cross-Entropy).
|
||||
"""
|
||||
|
||||
transformer_outputs: BaseModelOutputWithPast = self.model(
|
||||
input_ids,
|
||||
attention_mask=attention_mask,
|
||||
position_ids=position_ids,
|
||||
past_key_values=past_key_values,
|
||||
inputs_embeds=inputs_embeds,
|
||||
use_cache=use_cache,
|
||||
**kwargs,
|
||||
)
|
||||
hidden_states = transformer_outputs.last_hidden_state
|
||||
logits = self.score(hidden_states)
|
||||
|
||||
if input_ids is not None:
|
||||
batch_size = input_ids.shape[0]
|
||||
else:
|
||||
batch_size = inputs_embeds.shape[0]
|
||||
|
||||
if self.config.pad_token_id is None and batch_size != 1:
|
||||
raise ValueError("Cannot handle batch sizes > 1 if no padding token is defined.")
|
||||
if self.config.pad_token_id is None:
|
||||
last_non_pad_token = -1
|
||||
elif input_ids is not None:
|
||||
# To handle both left- and right- padding, we take the rightmost token that is not equal to pad_token_id
|
||||
non_pad_mask = (input_ids != self.config.pad_token_id).to(logits.device, torch.int32)
|
||||
token_indices = torch.arange(input_ids.shape[-1], device=logits.device, dtype=torch.int32)
|
||||
last_non_pad_token = (token_indices * non_pad_mask).argmax(-1)
|
||||
else:
|
||||
last_non_pad_token = -1
|
||||
logger.warning_once(
|
||||
f"{self.__class__.__name__} will not detect padding tokens in `inputs_embeds`. Results may be "
|
||||
"unexpected if using padding tokens in conjunction with `inputs_embeds.`"
|
||||
)
|
||||
|
||||
pooled_logits = logits[torch.arange(batch_size, device=logits.device), last_non_pad_token]
|
||||
|
||||
loss = None
|
||||
if labels is not None:
|
||||
loss = self.loss_function(logits=logits, labels=labels, pooled_logits=pooled_logits, config=self.config)
|
||||
|
||||
return SequenceClassifierOutputWithPast(
|
||||
loss=loss,
|
||||
logits=pooled_logits,
|
||||
past_key_values=transformer_outputs.past_key_values,
|
||||
hidden_states=transformer_outputs.hidden_states,
|
||||
attentions=transformer_outputs.attentions,
|
||||
)
|
||||
|
||||
|
||||
@auto_docstring
|
||||
class Exaone4ForTokenClassification(Exaone4PreTrainedModel):
|
||||
def __init__(self, config):
|
||||
super().__init__(config)
|
||||
self.num_labels = config.num_labels
|
||||
self.model = Exaone4Model(config)
|
||||
if getattr(config, "classifier_dropout", None) is not None:
|
||||
classifier_dropout = config.classifier_dropout
|
||||
elif getattr(config, "hidden_dropout", None) is not None:
|
||||
classifier_dropout = config.hidden_dropout
|
||||
else:
|
||||
classifier_dropout = 0.1
|
||||
self.dropout = nn.Dropout(classifier_dropout)
|
||||
self.score = nn.Linear(config.hidden_size, config.num_labels)
|
||||
|
||||
# Initialize weights and apply final processing
|
||||
self.post_init()
|
||||
|
||||
def get_input_embeddings(self):
|
||||
return self.model.embed_tokens
|
||||
|
||||
def set_input_embeddings(self, value):
|
||||
self.model.embed_tokens = value
|
||||
|
||||
@can_return_tuple
|
||||
@auto_docstring
|
||||
def forward(
|
||||
self,
|
||||
input_ids: Optional[torch.LongTensor] = None,
|
||||
attention_mask: Optional[torch.Tensor] = None,
|
||||
position_ids: Optional[torch.LongTensor] = None,
|
||||
past_key_values: Optional[Cache] = None,
|
||||
inputs_embeds: Optional[torch.FloatTensor] = None,
|
||||
labels: Optional[torch.LongTensor] = None,
|
||||
use_cache: Optional[bool] = None,
|
||||
**kwargs,
|
||||
) -> TokenClassifierOutput:
|
||||
r"""
|
||||
labels (`torch.LongTensor` of shape `(batch_size,)`, *optional*):
|
||||
Labels for computing the sequence classification/regression loss. Indices should be in `[0, ...,
|
||||
config.num_labels - 1]`. If `config.num_labels == 1` a regression loss is computed (Mean-Square loss), If
|
||||
`config.num_labels > 1` a classification loss is computed (Cross-Entropy).
|
||||
"""
|
||||
|
||||
outputs: BaseModelOutputWithPast = self.model(
|
||||
input_ids,
|
||||
attention_mask=attention_mask,
|
||||
position_ids=position_ids,
|
||||
past_key_values=past_key_values,
|
||||
inputs_embeds=inputs_embeds,
|
||||
use_cache=use_cache,
|
||||
**kwargs,
|
||||
)
|
||||
sequence_output = outputs.last_hidden_state
|
||||
sequence_output = self.dropout(sequence_output)
|
||||
logits = self.score(sequence_output)
|
||||
|
||||
loss = None
|
||||
if labels is not None:
|
||||
loss = self.loss_function(logits, labels, self.config)
|
||||
|
||||
return TokenClassifierOutput(
|
||||
loss=loss,
|
||||
logits=logits,
|
||||
hidden_states=outputs.hidden_states,
|
||||
attentions=outputs.attentions,
|
||||
)
|
||||
|
||||
|
||||
@auto_docstring
|
||||
class Exaone4ForQuestionAnswering(Exaone4PreTrainedModel):
|
||||
base_model_prefix = "transformer"
|
||||
|
||||
def __init__(self, config):
|
||||
super().__init__(config)
|
||||
self.transformer = Exaone4Model(config)
|
||||
self.qa_outputs = nn.Linear(config.hidden_size, 2)
|
||||
|
||||
# Initialize weights and apply final processing
|
||||
self.post_init()
|
||||
|
||||
def get_input_embeddings(self):
|
||||
return self.transformer.embed_tokens
|
||||
|
||||
def set_input_embeddings(self, value):
|
||||
self.transformer.embed_tokens = value
|
||||
|
||||
@can_return_tuple
|
||||
@auto_docstring
|
||||
def forward(
|
||||
self,
|
||||
input_ids: Optional[torch.LongTensor] = None,
|
||||
attention_mask: Optional[torch.Tensor] = None,
|
||||
position_ids: Optional[torch.LongTensor] = None,
|
||||
past_key_values: Optional[Cache] = None,
|
||||
inputs_embeds: Optional[torch.FloatTensor] = None,
|
||||
start_positions: Optional[torch.LongTensor] = None,
|
||||
end_positions: Optional[torch.LongTensor] = None,
|
||||
**kwargs: Unpack[TransformersKwargs],
|
||||
) -> QuestionAnsweringModelOutput:
|
||||
outputs: BaseModelOutputWithPast = self.transformer(
|
||||
input_ids,
|
||||
attention_mask=attention_mask,
|
||||
position_ids=position_ids,
|
||||
past_key_values=past_key_values,
|
||||
inputs_embeds=inputs_embeds,
|
||||
**kwargs,
|
||||
)
|
||||
|
||||
sequence_output = outputs.last_hidden_state
|
||||
|
||||
logits = self.qa_outputs(sequence_output)
|
||||
start_logits, end_logits = logits.split(1, dim=-1)
|
||||
start_logits = start_logits.squeeze(-1).contiguous()
|
||||
end_logits = end_logits.squeeze(-1).contiguous()
|
||||
|
||||
loss = None
|
||||
if start_positions is not None and end_positions is not None:
|
||||
loss = self.loss_function(start_logits, end_logits, start_positions, end_positions, **kwargs)
|
||||
|
||||
return QuestionAnsweringModelOutput(
|
||||
loss=loss,
|
||||
start_logits=start_logits,
|
||||
end_logits=end_logits,
|
||||
hidden_states=outputs.hidden_states,
|
||||
attentions=outputs.attentions,
|
||||
)
|
||||
|
||||
|
||||
__all__ = [
|
||||
"Exaone4PreTrainedModel",
|
||||
"Exaone4Model",
|
||||
"Exaone4ForCausalLM",
|
||||
"Exaone4ForSequenceClassification",
|
||||
"Exaone4ForTokenClassification",
|
||||
"Exaone4ForQuestionAnswering",
|
||||
]
|
||||
336
special_tokens_map.json
Normal file
336
special_tokens_map.json
Normal file
@@ -0,0 +1,336 @@
|
||||
{
|
||||
"additional_special_tokens": [
|
||||
"[unused0]",
|
||||
"[unused1]",
|
||||
"[unused2]",
|
||||
"[unused3]",
|
||||
"[unused4]",
|
||||
"[unused5]",
|
||||
"[unused6]",
|
||||
"[unused7]",
|
||||
"[unused8]",
|
||||
"[unused9]",
|
||||
"[unused10]",
|
||||
"[unused11]",
|
||||
"[unused12]",
|
||||
"[unused13]",
|
||||
"[unused14]",
|
||||
"[unused15]",
|
||||
"[unused16]",
|
||||
"[unused17]",
|
||||
"[unused18]",
|
||||
"[unused19]",
|
||||
"[unused20]",
|
||||
"[unused21]",
|
||||
"[unused22]",
|
||||
"[unused23]",
|
||||
"[unused24]",
|
||||
"[unused25]",
|
||||
"[unused26]",
|
||||
"[unused27]",
|
||||
"[unused28]",
|
||||
"[unused29]",
|
||||
"[unused30]",
|
||||
"[unused31]",
|
||||
"[unused32]",
|
||||
"[unused33]",
|
||||
"[unused34]",
|
||||
"[unused35]",
|
||||
"[unused36]",
|
||||
"[unused37]",
|
||||
"[unused38]",
|
||||
"[unused39]",
|
||||
"[unused40]",
|
||||
"[unused41]",
|
||||
"[unused42]",
|
||||
"[unused43]",
|
||||
"[unused44]",
|
||||
"[unused45]",
|
||||
"[unused46]",
|
||||
"[unused47]",
|
||||
"[unused48]",
|
||||
"[unused49]",
|
||||
"[unused50]",
|
||||
"[unused51]",
|
||||
"[unused52]",
|
||||
"[unused53]",
|
||||
"[unused54]",
|
||||
"[unused55]",
|
||||
"[unused56]",
|
||||
"[unused57]",
|
||||
"[unused58]",
|
||||
"[unused59]",
|
||||
"[unused60]",
|
||||
"[unused61]",
|
||||
"[unused62]",
|
||||
"[unused63]",
|
||||
"[unused64]",
|
||||
"[unused65]",
|
||||
"[unused66]",
|
||||
"[unused67]",
|
||||
"[unused68]",
|
||||
"[unused69]",
|
||||
"[unused70]",
|
||||
"[unused71]",
|
||||
"[unused72]",
|
||||
"[unused73]",
|
||||
"[unused74]",
|
||||
"[unused75]",
|
||||
"[unused76]",
|
||||
"[unused77]",
|
||||
"[unused78]",
|
||||
"[unused79]",
|
||||
"[unused80]",
|
||||
"[unused81]",
|
||||
"[unused82]",
|
||||
"[unused83]",
|
||||
"[unused84]",
|
||||
"[unused85]",
|
||||
"[unused86]",
|
||||
"[unused87]",
|
||||
"[unused88]",
|
||||
"[unused89]",
|
||||
"[unused90]",
|
||||
"[unused91]",
|
||||
"[unused92]",
|
||||
"[unused93]",
|
||||
"[unused94]",
|
||||
"[unused95]",
|
||||
"[unused96]",
|
||||
"[unused97]",
|
||||
"[unused98]",
|
||||
"[unused99]",
|
||||
"[extra_id_0]",
|
||||
"[extra_id_1]",
|
||||
"[extra_id_2]",
|
||||
"[extra_id_3]",
|
||||
"[extra_id_4]",
|
||||
"[extra_id_5]",
|
||||
"[extra_id_6]",
|
||||
"[extra_id_7]",
|
||||
"[extra_id_8]",
|
||||
"[extra_id_9]",
|
||||
"[extra_id_10]",
|
||||
"[extra_id_11]",
|
||||
"[extra_id_12]",
|
||||
"[extra_id_13]",
|
||||
"[extra_id_14]",
|
||||
"[extra_id_15]",
|
||||
"[extra_id_16]",
|
||||
"[extra_id_17]",
|
||||
"[extra_id_18]",
|
||||
"[extra_id_19]",
|
||||
"[extra_id_20]",
|
||||
"[extra_id_21]",
|
||||
"[extra_id_22]",
|
||||
"[extra_id_23]",
|
||||
"[extra_id_24]",
|
||||
"[extra_id_25]",
|
||||
"[extra_id_26]",
|
||||
"[extra_id_27]",
|
||||
"[extra_id_28]",
|
||||
"[extra_id_29]",
|
||||
"[extra_id_30]",
|
||||
"[extra_id_31]",
|
||||
"[extra_id_32]",
|
||||
"[extra_id_33]",
|
||||
"[extra_id_34]",
|
||||
"[extra_id_35]",
|
||||
"[extra_id_36]",
|
||||
"[extra_id_37]",
|
||||
"[extra_id_38]",
|
||||
"[extra_id_39]",
|
||||
"[extra_id_40]",
|
||||
"[extra_id_41]",
|
||||
"[extra_id_42]",
|
||||
"[extra_id_43]",
|
||||
"[extra_id_44]",
|
||||
"[extra_id_45]",
|
||||
"[extra_id_46]",
|
||||
"[extra_id_47]",
|
||||
"[extra_id_48]",
|
||||
"[extra_id_49]",
|
||||
"[extra_id_50]",
|
||||
"[extra_id_51]",
|
||||
"[extra_id_52]",
|
||||
"[extra_id_53]",
|
||||
"[extra_id_54]",
|
||||
"[extra_id_55]",
|
||||
"[extra_id_56]",
|
||||
"[extra_id_57]",
|
||||
"[extra_id_58]",
|
||||
"[extra_id_59]",
|
||||
"[extra_id_60]",
|
||||
"[extra_id_61]",
|
||||
"[extra_id_62]",
|
||||
"[extra_id_63]",
|
||||
"[extra_id_64]",
|
||||
"[extra_id_65]",
|
||||
"[extra_id_66]",
|
||||
"[extra_id_67]",
|
||||
"[extra_id_68]",
|
||||
"[extra_id_69]",
|
||||
"[extra_id_70]",
|
||||
"[extra_id_71]",
|
||||
"[extra_id_72]",
|
||||
"[extra_id_73]",
|
||||
"[extra_id_74]",
|
||||
"[extra_id_75]",
|
||||
"[extra_id_76]",
|
||||
"[extra_id_77]",
|
||||
"[extra_id_78]",
|
||||
"[extra_id_79]",
|
||||
"[extra_id_80]",
|
||||
"[extra_id_81]",
|
||||
"[extra_id_82]",
|
||||
"[extra_id_83]",
|
||||
"[extra_id_84]",
|
||||
"[extra_id_85]",
|
||||
"[extra_id_86]",
|
||||
"[extra_id_87]",
|
||||
"[extra_id_88]",
|
||||
"[extra_id_89]",
|
||||
"[extra_id_90]",
|
||||
"[extra_id_91]",
|
||||
"[extra_id_92]",
|
||||
"[extra_id_93]",
|
||||
"[extra_id_94]",
|
||||
"[extra_id_95]",
|
||||
"[extra_id_96]",
|
||||
"[extra_id_97]",
|
||||
"[extra_id_98]",
|
||||
"[extra_id_99]",
|
||||
"[extra_id_100]",
|
||||
"[extra_id_101]",
|
||||
"[extra_id_102]",
|
||||
"[extra_id_103]",
|
||||
"[extra_id_104]",
|
||||
"[extra_id_105]",
|
||||
"[extra_id_106]",
|
||||
"[extra_id_107]",
|
||||
"[extra_id_108]",
|
||||
"[extra_id_109]",
|
||||
"[extra_id_110]",
|
||||
"[extra_id_111]",
|
||||
"[extra_id_112]",
|
||||
"[extra_id_113]",
|
||||
"[extra_id_114]",
|
||||
"[extra_id_115]",
|
||||
"[extra_id_116]",
|
||||
"[extra_id_117]",
|
||||
"[extra_id_118]",
|
||||
"[extra_id_119]",
|
||||
"[extra_id_120]",
|
||||
"[extra_id_121]",
|
||||
"[extra_id_122]",
|
||||
"[extra_id_123]",
|
||||
"[extra_id_124]",
|
||||
"[extra_id_125]",
|
||||
"[extra_id_126]",
|
||||
"[extra_id_127]",
|
||||
"[extra_id_128]",
|
||||
"[extra_id_129]",
|
||||
"[extra_id_130]",
|
||||
"[extra_id_131]",
|
||||
"[extra_id_132]",
|
||||
"[extra_id_133]",
|
||||
"[extra_id_134]",
|
||||
"[extra_id_135]",
|
||||
"[extra_id_136]",
|
||||
"[extra_id_137]",
|
||||
"[extra_id_138]",
|
||||
"[extra_id_139]",
|
||||
"[extra_id_140]",
|
||||
"[extra_id_141]",
|
||||
"[extra_id_142]",
|
||||
"[extra_id_143]",
|
||||
"[extra_id_144]",
|
||||
"[extra_id_145]",
|
||||
"[extra_id_146]",
|
||||
"[extra_id_147]",
|
||||
"[extra_id_148]",
|
||||
"[extra_id_149]",
|
||||
"[extra_id_150]",
|
||||
"[extra_id_151]",
|
||||
"[extra_id_152]",
|
||||
"[extra_id_153]",
|
||||
"[extra_id_154]",
|
||||
"[extra_id_155]",
|
||||
"[extra_id_156]",
|
||||
"[extra_id_157]",
|
||||
"[extra_id_158]",
|
||||
"[extra_id_159]",
|
||||
"[extra_id_160]",
|
||||
"[extra_id_161]",
|
||||
"[extra_id_162]",
|
||||
"[extra_id_163]",
|
||||
"[extra_id_164]",
|
||||
"[extra_id_165]",
|
||||
"[extra_id_166]",
|
||||
"[extra_id_167]",
|
||||
"[extra_id_168]",
|
||||
"[extra_id_169]",
|
||||
"[extra_id_170]",
|
||||
"[extra_id_171]",
|
||||
"[extra_id_172]",
|
||||
"[extra_id_173]",
|
||||
"[extra_id_174]",
|
||||
"[extra_id_175]",
|
||||
"[extra_id_176]",
|
||||
"[extra_id_177]",
|
||||
"[extra_id_178]",
|
||||
"[extra_id_179]",
|
||||
"[extra_id_180]",
|
||||
"[extra_id_181]",
|
||||
"[extra_id_182]",
|
||||
"[extra_id_183]",
|
||||
"[extra_id_184]",
|
||||
"[extra_id_185]",
|
||||
"[extra_id_186]",
|
||||
"[extra_id_187]",
|
||||
"[extra_id_188]",
|
||||
"[|system|]",
|
||||
"[|tool|]",
|
||||
"[|assistant|]",
|
||||
"[|user|]",
|
||||
"[|endofturn|]",
|
||||
"PI:URL",
|
||||
"PI:EMAIL",
|
||||
"PI:ACCOUNT_NUM",
|
||||
"PI:PHONE_NUM",
|
||||
"PI:BUSINESS_NUM",
|
||||
"PI:ANNON",
|
||||
"PI:KEY",
|
||||
"PI:ID",
|
||||
"PI:IP_ADDRESS",
|
||||
"PI:USER"
|
||||
],
|
||||
"bos_token": {
|
||||
"content": "[BOS]",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
},
|
||||
"eos_token": {
|
||||
"content": "[|endofturn|]",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
},
|
||||
"pad_token": {
|
||||
"content": "[PAD]",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
},
|
||||
"unk_token": {
|
||||
"content": "[UNK]",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
}
|
||||
}
|
||||
512826
tokenizer.json
Normal file
512826
tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
3219
tokenizer_config.json
Normal file
3219
tokenizer_config.json
Normal file
File diff suppressed because it is too large
Load Diff
1
vocab.json
Normal file
1
vocab.json
Normal file
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user