Files
SelfExtended/README.md

29 lines
1.6 KiB
Markdown
Raw Normal View History

---
license: mit
language:
- en
---
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
(implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
(implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
(implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
(implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
(implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
(implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
(implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
(implemented with flash attention) to achieve better performance on long context tasks.