Mistral-7B-AEZAKMI-v2

Go to file

ModelHub XC 263ac12d92 初始化项目，由ModelHub XC社区提供模型

Model: adamo1139/Mistral-7B-AEZAKMI-v2
Source: Original Platform

2026-05-30 00:57:21 +08:00

.gitattributes

初始化项目，由ModelHub XC社区提供模型

2026-05-30 00:57:21 +08:00

added_tokens.json

初始化项目，由ModelHub XC社区提供模型

2026-05-30 00:57:21 +08:00

config.json

初始化项目，由ModelHub XC社区提供模型

2026-05-30 00:57:21 +08:00

generation_config.json

初始化项目，由ModelHub XC社区提供模型

2026-05-30 00:57:21 +08:00

model-00001-of-00003.safetensors

初始化项目，由ModelHub XC社区提供模型

2026-05-30 00:57:21 +08:00

model-00002-of-00003.safetensors

初始化项目，由ModelHub XC社区提供模型

2026-05-30 00:57:21 +08:00

model-00003-of-00003.safetensors

初始化项目，由ModelHub XC社区提供模型

2026-05-30 00:57:21 +08:00

model.safetensors.index.json

初始化项目，由ModelHub XC社区提供模型

2026-05-30 00:57:21 +08:00

README.md

初始化项目，由ModelHub XC社区提供模型

2026-05-30 00:57:21 +08:00

special_tokens_map.json

初始化项目，由ModelHub XC社区提供模型

2026-05-30 00:57:21 +08:00

tokenizer_config.json

初始化项目，由ModelHub XC社区提供模型

2026-05-30 00:57:21 +08:00

tokenizer.model

初始化项目，由ModelHub XC社区提供模型

2026-05-30 00:57:21 +08:00

README.md

license, model-index

license

model-index

apache-2.0

name

results

Mistral-7B-AEZAKMI-v2

task

dataset

metrics

source

type	name
text-generation	Text Generation

name

type

config

split

args

AI2 Reasoning Challenge (25-Shot)

ai2_arc

ARC-Challenge

test

num_few_shot
25

type	value	name
acc_norm	58.11	normalized accuracy

url	name
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=adamo1139/Mistral-7B-AEZAKMI-v2	Open LLM Leaderboard

task

dataset

metrics

source

type	name
text-generation	Text Generation

name

type

split

args

HellaSwag (10-Shot)

hellaswag

validation

num_few_shot
10

type	value	name
acc_norm	82.53	normalized accuracy

url	name
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=adamo1139/Mistral-7B-AEZAKMI-v2	Open LLM Leaderboard

task

dataset

metrics

source

type	name
text-generation	Text Generation

name

type

config

split

args

MMLU (5-Shot)

cais/mmlu

all

test

num_few_shot
5

type	value	name
acc	59.89	accuracy

url	name
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=adamo1139/Mistral-7B-AEZAKMI-v2	Open LLM Leaderboard

task

dataset

metrics

source

type	name
text-generation	Text Generation

name

type

config

split

args

TruthfulQA (0-shot)

truthful_qa

multiple_choice

validation

num_few_shot
0

type	value
mc2	51.5

url	name
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=adamo1139/Mistral-7B-AEZAKMI-v2	Open LLM Leaderboard

task

dataset

metrics

source

type	name
text-generation	Text Generation

name

type

config

split

args

Winogrande (5-shot)

winogrande

winogrande_xl

validation

num_few_shot
5

type	value	name
acc	73.64	accuracy

url	name
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=adamo1139/Mistral-7B-AEZAKMI-v2	Open LLM Leaderboard

task

dataset

metrics

source

type	name
text-generation	Text Generation

name

type

config

split

args

GSM8k (5-shot)

gsm8k

main

test

num_few_shot
5

type	value	name
acc	32.45	accuracy

url	name
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=adamo1139/Mistral-7B-AEZAKMI-v2	Open LLM Leaderboard

## Model description Mistral 7B base model fine-tuned on AEZAKMI v2 dataset. It's like airoboros but hopefully with less gptslop, no refusals and less typical language used by RLHFed OpenAI models. Say goodbye to "It's important to remember"! Prompt format is standard chatml. Don't expect it to be good at math, riddles or be crazy smart. My end goal with AEZAKMI is to create a cozy free chatbot. Base model used for fine-tuning was Mistral 7B 0.1. ## Prompt Format I recommend using ChatML format, as this was used during fine-tune. Here's a prompt format you should use, you can set a different system message, model seems to respect that fine, so it wasn't overfitted.

<|im_start|>system  A chat with uncensored assistant.<|im_end|>  <|im_start|>user  {prompt}<|im_end|>  <|im_start|>assistant

## Intended uses & limitations Use is limited by Apache-2 license. ## Known Issues Stories have ChatGPT like paragraph spacing, I will work on this in the future maybe, not a high priority. Unrestricted-ness of the v2 isn't quite something that I am happy with yet, especially using prompt "A chat.". With a slightly modifed prompt it works somewhat better, I recommend using "A chat with uncensored assistant.", it's stupid but it helps a lot. I think Mistral 7B base is contaminated with refusals as other newer models. My next project is to attempt to de-contaminate base LLMs using DPO with preferred data coming from uncontaminated raw models. I plan to release that dataset openly. I was made aware of the frequent occurrence of the phrase "sending shivers down a spine" in the generations during RP of v1, so I fixed those samples - it should be better now. There was also some issue with handling long system messages for RP with yi-34b AEZAKMI v1, I was planning to investigate it for v2 but I didn't. Samples of generations of this model are available here - https://huggingface.co/datasets/adamo1139/misc/tree/main/benchmarks # Open LLM Leaderboard Evaluation Results Detailed results can be found here | Metric |Value| |---------------------------------|----:| |Avg. |59.69| |AI2 Reasoning Challenge (25-Shot)|58.11| |HellaSwag (10-Shot) |82.53| |MMLU (5-Shot) |59.89| |TruthfulQA (0-shot) |51.50| |Winogrande (5-shot) |73.64| |GSM8k (5-shot) |32.45|