1.6 KiB
license, language
| license | language | |
|---|---|---|
| mit |
|
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.
This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.