29 lines
1.6 KiB
Markdown
29 lines
1.6 KiB
Markdown
---
|
|
license: mit
|
|
language:
|
|
- en
|
|
---
|
|
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
|
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
|
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
|
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
|
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
|
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
|
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
|
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
|
(implemented with flash attention) to achieve better performance on long context tasks. |