--- language: - en license: other license_name: apache-2.0-mixed-data-terms library_name: transformers pipeline_tag: text-generation base_model: mistralai/Mistral-7B-v0.1 datasets: - allenai/ai2_arc - allenai/openbookqa - allenai/qasc - allenai/sciq tags: - mistral - qlora - preference-learning - mutual-information-optimization - mio - hard-negative-mining - arc-challenge - science-qa - multiple-choice - fp16 model-index: - name: AIRE-ARC-Mistral-7B-SFT-MIO-FP16 results: - task: type: question-answering name: ARC-Challenge 25-shot dataset: type: allenai/ai2_arc name: AI2 ARC-Challenge config: ARC-Challenge split: test metrics: - type: acc_norm name: Normalized Accuracy value: 0.765358361774744 - type: acc name: Accuracy value: 0.7474402730375427 --- # AIRE-ARC-Mistral-7B-SFT-MIO-FP16 A standalone FP16 fine-tune of [`mistralai/Mistral-7B-v0.1`](https://huggingface.co/mistralai/Mistral-7B-v0.1) for ARC-Challenge and related science multiple-choice reasoning. The model was trained in two stages: 1. completion-only QLoRA SFT on 27,294 public science-QA training examples; 2. MIO preference fine-tuning on 47,087 train-only pairs, where the rejected answers were the incorrect choices that the SFT model found most plausible. The final model reached **76.54% `acc_norm`** and **74.74% raw accuracy** on the 1,172-example ARC-Challenge test split with 25-shot evaluation. Model selection was completed on the validation split before the test split was opened. > **Important:** this is a domain-specialized completion model, not a chat or general instruction model. It improves science and commonsense multiple-choice reasoning, but capability-retention audits found regressions on MMLU Humanities, MMLU Social Sciences, and WikiText perplexity. ## Model summary | Item | Value | |---|---| | Architecture | `MistralForCausalLM` | | Parameters | 7B | | Parent model | `mistralai/Mistral-7B-v0.1` | | Parent revision | `27d67f1b5f57dc0953326b2601d68371d40ea8da` | | Repository format | Standalone merged FP16 weights | | Adapter required | No | | Primary language | English | | Primary task | Science multiple-choice QA | | Training hardware | One NVIDIA A100-SXM4 40GB | The repository contains the tokenizer and merged model weights. PEFT and bitsandbytes are not required for inference. `merge_summary.json` records the merge provenance used for the submitted checkpoint. ## Intended use Suitable uses include: - evaluating science and commonsense multiple-choice reasoning; - studying model-scored hard-negative selection; - reproducing the reported ARC-Challenge evaluation; - research on preference optimization for short-answer QA. The model is not intended for safety-critical use, factual deployment without verification, general chat, or instruction following. No additional safety alignment was performed. ## Prompt format Training and evaluation use answer-text completions rather than answer-position labels: ```text Question: {question} Answer: ``` The expected completion is the answer text, including its leading space. No chat template is used. ## Quick start ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer repo_id = "psymon/mistral-7b-mio-arc-fp16" tokenizer = AutoTokenizer.from_pretrained(repo_id) if tokenizer.pad_token_id is None: tokenizer.pad_token = tokenizer.eos_token model = AutoModelForCausalLM.from_pretrained( repo_id, dtype=torch.float16, device_map="auto", ) model.eval() prompt = "Question: What is a worldwide increase in temperature called?\nAnswer:" inputs = tokenizer(prompt, return_tensors="pt").to(model.device) with torch.inference_mode(): output = model.generate( **inputs, max_new_tokens=24, do_sample=False, pad_token_id=tokenizer.pad_token_id, eos_token_id=tokenizer.eos_token_id, ) completion = tokenizer.decode( output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True, ) print(completion) ``` For multiple-choice evaluation, compare the conditional log-likelihood of each answer text. Do not convert the choices to `A`, `B`, `C`, and `D` unless the evaluation protocol is changed accordingly. ## Training data Only public **training splits** were used for optimization. ARC-Challenge validation was used for deduplication, failure analysis, and model selection. ARC-Challenge test was not used during training or model selection. ### SFT corpus | Source | Examples after filtering | Role | |---|---:|---| | ARC-Challenge | 1,117 | SFT and preference candidates | | ARC-Easy | 2,241 | SFT and preference candidates | | OpenBookQA | 4,826 | SFT and preference candidates | | QASC | 7,514 | SFT and preference candidates | | SciQ | 11,596 | SFT only | | **Total** | **27,294** | | The SFT prompt contains the question only; the target is the correct answer text. Choices were retained for data normalization and duplicate checks but were not included in the SFT prompt. ### Preference corpus The SFT model scored every answer choice for 15,698 train-only prompts from ARC-Challenge, ARC-Easy, OpenBookQA, and QASC. Each score was the answer-token log-likelihood sum divided by the answer's character length. For each question, up to three highest-scoring incorrect choices were paired with the correct answer: ```text prompt = "Question: {question}\nAnswer:" chosen = " {correct answer text}" rejected = " {model-selected incorrect answer text}" ``` This produced 47,087 preference pairs. Seven three-choice questions contributed two pairs each; the remaining questions contributed three pairs each. In 3,081 prompts, at least one incorrect choice received a higher selection score than the correct answer. ## Training procedure ### Stage 1: completion-only SFT | Hyperparameter | Value | |---|---:| | Quantization | 4-bit NF4 QLoRA | | Compute dtype | BF16 | | LoRA target | all linear layers | | LoRA rank / alpha / dropout | 16 / 32 / 0.05 | | Learning rate | `1e-4` | | Scheduler | cosine | | Effective batch size | 16 | | Maximum sequence length | 384 | | Epochs / optimizer steps | 1 / 1,706 | Loss was computed only on the answer completion; prompt tokens were excluded from the labels. ### Stage 2: MIO preference fine-tuning MIO (Mutual Information Optimization) contrasts a chosen and rejected completion against a frozen reference model. This run used the merged SFT checkpoint as the reference and initialized a new LoRA policy from the same weights. Completion log-probabilities were averaged over response tokens, including EOS. | Hyperparameter | Value | |---|---:| | Preference pairs | 47,087 | | Beta | 0.5 | | Learning rate | `7.5e-6` | | LoRA rank / alpha / dropout | 16 / 32 / 0.0 | | Effective batch size | 32 | | Maximum sequence length | 256 | | Epochs / optimizer steps | 1 / 1,472 | The initial policy/reference log-ratios were exactly zero and the initial loss was `1.386294`, matching `ln(4)`. Reference inference was executed first under `no_grad`; its activations were released before the policy forward pass to fit training on a 40GB A100 without changing the objective. The result should be interpreted as the effect of the **complete procedure**: SFT reference, model-scored hard negatives, mean completion log-probabilities, and MIO. This experiment did not isolate the causal contribution of each component. ## Evaluation Evaluation used EleutherAI's `lm-evaluation-harness` protocol with 25 few-shot examples and answer-choice likelihood normalization. ### ARC-Challenge validation | Model stage | `acc` | `acc_norm` | |---|---:|---:| | Mistral-7B-v0.1 | 51.51% | 56.52% | | Completion-only SFT | 57.86% | 59.20% | | SFT + MIO | 71.57% | **73.24%** | From SFT to MIO, 51 normalized predictions changed from incorrect to correct and 9 changed from correct to incorrect, for a net gain of 42 correct answers on 299 validation examples. ### One-time held-out test | Split | Documents | Few-shot | `acc` | `acc_norm` | |---|---:|---:|---:|---:| | ARC-Challenge test | 1,172 | 25 | 74.74% | **76.54%** | The candidate model was locked before this evaluation. The test split was evaluated once, and no post-test training, hyperparameter tuning, or model reselection was performed. ### Post-test overlap audit The test split was compared with the training ledger only after final evaluation: - exact Test/SFT ID overlap: 0; - exact question plus complete choice-set overlap: 0; - exact normalized question-stem overlap: 9 test rows; - character similarity of at least 0.90: 18 test rows. Removing the 18 near-overlap rows changed test `acc_norm` from 76.5358% to **76.4298%** (-0.1060 percentage points). This is a sensitivity analysis, not proof that semantic contamination is absent; paraphrases and shared underlying facts may remain undetected. ## Capability-retention audit After model selection and the one-time ARC test, the standalone Hub checkpoint was downloaded again and compared with the base model. These results were not used for training or selection. | Benchmark | Metric | Base | Final | Change | |---|---|---:|---:|---:| | HellaSwag | `acc_norm` | 81.17% | 85.43% | +4.26 pp | | PIQA | `acc_norm` | 82.48% | 85.75% | +3.26 pp | | WinoGrande | `acc` | 75.37% | 80.03% | +4.66 pp | | MMLU Humanities | `acc` | 56.43% | 53.18% | -3.25 pp | | MMLU Social Sciences | `acc` | 73.81% | 70.46% | -3.35 pp | | WikiText | word perplexity | 8.0848 | 8.7194 | +7.85% (worse) | The MMLU regressions partly depend on answer-position labels: accuracy fell much more for answers at positions `B` and `D` than for `A` and `C`. This suggests that output-label calibration changed in addition to any knowledge loss. WikiText does not use answer labels and also regressed, so format calibration alone cannot explain all of the degradation. ## Reproducing the ARC evaluation ```bash pip install "lm-eval==0.4.12" python -m lm_eval \ --model hf \ --model_args "pretrained=psymon/mistral-7b-mio-arc-fp16,dtype=float16" \ --tasks arc_challenge \ --num_fewshot 25 \ --batch_size 8 ``` The reported submission result used the official ARC-Challenge parquet files in an offline custom task, with hashes checked before evaluation. The submitted Colab notebook contains the exact data validation, model-selection lock, and one-time test procedure. ## Limitations - **Domain specialization:** optimized for short English science-QA completions. - **Not a chat model:** no instruction-following or conversational alignment was added. - **Single training seed:** training variance was not estimated across seeds. - **Public-corpus overlap risk:** post-test string audits cannot exclude semantic overlap or paraphrases. - **Capability trade-offs:** MMLU and WikiText regressions show that the model did not preserve all base capabilities. - **Incomplete coverage:** code, mathematics, safety, long-form generation, summarization, and multilingual behavior were not evaluated. - **No safety guarantee:** outputs may be incorrect, biased, or unsafe. ## License and data terms The parent model is released under Apache-2.0. The training sources have separate terms: | Source | Terms shown by the source dataset card | |---|---| | ARC | CC BY-SA 4.0 | | QASC | CC BY 4.0 | | SciQ | CC BY-NC 3.0 | | OpenBookQA | No license declared in the pinned Hugging Face dataset card | Because the SFT stage includes SciQ and the OpenBookQA card does not declare complete licensing information, this model card intentionally uses `license: other` rather than presenting the repository as unconditionally Apache-2.0. The repository is shared for research and coding-test evaluation. Users are responsible for reviewing the parent-model license and each source dataset's terms before redistribution or downstream use, especially commercial use. This section is informational and not legal advice. ## References - A. Q. Jiang et al., [Mistral 7B](https://arxiv.org/abs/2310.06825), 2023. - T. Dettmers et al., [QLoRA: Efficient Finetuning of Quantized LLMs](https://arxiv.org/abs/2305.14314), 2023. - X. Lv et al., [The Hidden Link Between RLHF and Contrastive Learning](https://arxiv.org/abs/2506.22578), ICLR 2026. - P. Clark et al., [Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge](https://arxiv.org/abs/1803.05457), 2018.