29 lines
1.6 KiB
Markdown
29 lines
1.6 KiB
Markdown
|
|
---
|
||
|
|
license: mit
|
||
|
|
language:
|
||
|
|
- en
|
||
|
|
---
|
||
|
|
|
||
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
||
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
||
|
|
|
||
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
||
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
||
|
|
|
||
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
||
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
||
|
|
|
||
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
||
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
||
|
|
|
||
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
||
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
||
|
|
|
||
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
||
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
||
|
|
|
||
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
||
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|
||
|
|
|
||
|
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick
|
||
|
|
(implemented with flash attention) to achieve better performance on long context tasks.
|