license, language
license language
mit
en

This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.

This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.

This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.

This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.

This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.

This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.

This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.

This model is based on pretrained model meta/Llama-2-7b-chat-bf and has been equipped with SelfExtend trick (implemented with flash attention) to achieve better performance on long context tasks.

Description
Model synced from source: rachmanino/SelfExtended
Readme 1 MiB