初始化项目,由ModelHub XC社区提供模型
Model: MLVXN/MicroLLM2 Source: Original Platform
This commit is contained in:
38
.gitattributes
vendored
Normal file
38
.gitattributes
vendored
Normal file
@@ -0,0 +1,38 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
microllm2-f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
microllm2-q4_k_m.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
microllm2-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
23
Modelfile
Normal file
23
Modelfile
Normal file
@@ -0,0 +1,23 @@
|
||||
FROM ./microllm2-checkpoints/final_merged
|
||||
|
||||
# MicroLLM2 — GPT2-XL 1.5B elevated to chatbot by Maximalist Labs
|
||||
# Base: openai-community/gpt2-xl (48 layers, 1600 hidden, 1024 ctx)
|
||||
# Ollama Modelfile — run with: ollama create microllm2 -f Modelfile && ollama run microllm2
|
||||
|
||||
TEMPLATE """<|im_start|>user
|
||||
{{ .Prompt }}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
"""
|
||||
|
||||
SYSTEM """You are MicroLLM2, a helpful chatbot created by Maximalist Labs. You are based on GPT2-XL but elevated with distilled instruction tuning. Be concise, helpful, and honest. If asked who you are, say you are MicroLLM2 created by Maximalist Labs."""
|
||||
|
||||
PARAMETER temperature 0.7
|
||||
PARAMETER top_p 0.9
|
||||
PARAMETER top_k 50
|
||||
PARAMETER repeat_penalty 1.1
|
||||
PARAMETER num_ctx 1024
|
||||
PARAMETER num_predict 256
|
||||
PARAMETER stop "<|im_end|>"
|
||||
PARAMETER stop "<|endoftext|>"
|
||||
|
||||
LICENSE """Apache 2.0 — MicroLLM2 by Maximalist Labs (MLVXN)"""
|
||||
226
README.md
Normal file
226
README.md
Normal file
@@ -0,0 +1,226 @@
|
||||
---
|
||||
library_name: transformers
|
||||
license: apache-2.0
|
||||
base_model: openai-community/gpt2-xl
|
||||
tags:
|
||||
- chatbot
|
||||
- gpt2
|
||||
- lora
|
||||
- instruction-tuned
|
||||
- distilled
|
||||
- microllm2
|
||||
pipeline_tag: text-generation
|
||||
language:
|
||||
- en
|
||||
widget:
|
||||
- text: "<|im_start|>user\nWho are you?<|im_end|>\n<|im_start|>assistant\n"
|
||||
example_title: Identity check
|
||||
- text: "<|im_start|>user\nExplain quantum computing in simple terms<|im_end|>\n<|im_start|>assistant\n"
|
||||
example_title: Simple explanation
|
||||
---
|
||||
|
||||
# MicroLLM2
|
||||
|
||||

|
||||
|
||||
MicroLLM2 is a chatbot built from GPT2 XL 1.5B by Maximalist Labs. It takes the classic openai-community/gpt2-xl and elevates it with instruction tuning and distillation so it can actually chat, follow prompts, and keep a consistent identity.
|
||||
|
||||
If you ask who made it, it will tell you: MicroLLM2 created by Maximalist Labs. That is baked in during training, not just a system prompt.
|
||||
|
||||
**Repo:** `MLVXN/MicroLLM2`
|
||||
**Base:** `openai-community/gpt2-xl` (48 layers, 1600 hidden, 1024 context, 1.5B params)
|
||||
**Method:** LoRA SFT on distilled chat data, merged to a single safetensors for easy use
|
||||
**Context:** 1024 tokens
|
||||
**License:** Apache 2.0
|
||||
|
||||
## What makes this different from plain GPT2 XL
|
||||
|
||||
Plain GPT2 XL is a strong completer but not a chat model. MicroLLM2 adds:
|
||||
|
||||
* ChatML format with `<|im_start|>` and `<|im_end|>` so conversations have clear user and assistant turns
|
||||
* Distilled instruction data from high quality teachers (GPT-4, GPT-3.5, Mixtral) plus identity reinforcement
|
||||
* Clean merge: no adapter needed at inference, just load like any GPT2 model
|
||||
|
||||
No fancy claims here. It is still a 1.5B model with 1024 context. It will not beat 7B or larger models on broad knowledge, but it is far more useful than raw GPT2 XL for chatting, writing, and simple reasoning.
|
||||
|
||||
## Training in a nutshell
|
||||
* **Tuning:** LoRA r=64 alpha=128 on all attention and MLP projections (c_attn, c_proj, c_fc). About 78M trainable params. BF16 with TF32, Flash SDPA, packing, gradient checkpointing, 8-bit Adam, torch.compile.
|
||||
* **Throughput:** around 16.5k tokens per second on H100, roughly 3 hours for the main run plus overhead to land in the 4 to 5 hour window.
|
||||
* **Data mix:** 200k samples total, 3 epochs. Roughly 29k from UltraChat 200k (GPT-3.5), 100k from OpenHermes 2.5 (GPT-4), 60k from WizardLM Evol Instruct V2 (GPT-4), 5k from Cosmopedia v2 (Mixtral), plus 10k identity examples upsampled. Raw about 510M tokens, effective about 200M after packing and truncation. All packed to 1024 with ChatML.
|
||||
* **Identity:** 200 hand written identity prompts expanded to 10k during training so the model learns to answer consistently as MicroLLM2 by Maximalist Labs.
|
||||
* **Chat template:** `<|im_start|>user\n{prompt}<|im_end|>\n<|im_start|>assistant\n{response}<|im_end|>`
|
||||
|
||||
## How to use
|
||||
|
||||
### Transformers
|
||||
|
||||
```python
|
||||
from transformers import AutoTokenizer, AutoModelForCausalLM
|
||||
import torch
|
||||
|
||||
model_id = "MLVXN/MicroLLM2"
|
||||
tok = AutoTokenizer.from_pretrained(model_id)
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
model_id,
|
||||
torch_dtype=torch.bfloat16 if torch.cuda.is_available() else torch.float32,
|
||||
device_map="auto"
|
||||
)
|
||||
|
||||
def chat(prompt, max_new=160):
|
||||
formatted = f"<|im_start|>user\n{prompt}<|im_end|>\n<|im_start|>assistant\n"
|
||||
inputs = tok(formatted, return_tensors="pt").to(model.device)
|
||||
out = model.generate(
|
||||
**inputs,
|
||||
max_new_tokens=max_new,
|
||||
do_sample=True,
|
||||
temperature=0.7,
|
||||
top_p=0.9,
|
||||
repetition_penalty=1.1,
|
||||
pad_token_id=tok.eos_token_id,
|
||||
eos_token_id=tok.convert_tokens_to_ids("<|im_end|>")
|
||||
)
|
||||
text = tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=False)
|
||||
return text.split("<|im_end|>")[0].strip()
|
||||
|
||||
print(chat("Who are you?"))
|
||||
print(chat("Write a short poem about the H100"))
|
||||
```
|
||||
|
||||
### Ollama Modelfile
|
||||
|
||||
A `Modelfile` is included for Ollama. It sets the ChatML template, system prompt, and sane defaults.
|
||||
|
||||
```bash
|
||||
ollama create microllm2 -f Modelfile
|
||||
ollama run microllm2
|
||||
# then chat normally, the identity is already set
|
||||
```
|
||||
|
||||
### GGUF for llama.cpp
|
||||
|
||||
GGUF weights are in this repo:
|
||||
|
||||
* `microllm2-f16.gguf` full precision, best quality, about 3.0 GB
|
||||
* `microllm2-q8_0.gguf` 8-bit, near full quality, about 1.6 GB
|
||||
* `microllm2-q4_k_m.gguf` 4-bit, smallest, about 0.9 GB, good for CPU and edge
|
||||
|
||||
Use with llama.cpp, LM Studio, or any GGUF runner:
|
||||
|
||||
```bash
|
||||
# llama.cpp example
|
||||
./llama-cli -m microllm2-q4_k_m.gguf -p "<|im_start|>user\nHello<|im_end|>\n<|im_start|>assistant\n" -n 128
|
||||
```
|
||||
|
||||
The model is GPT2 architecture in GGUF, so make sure your runner supports GPT2 GGUF.
|
||||
|
||||
## Benchmark: MMLU
|
||||
|
||||
We include `mmlu_bench.py` so anyone can reproduce numbers. It runs 5 shot MMLU either with lm-evaluation-harness if you have it, or a lightweight direct logprob scorer that works without extra deps.
|
||||
|
||||
```bash
|
||||
python mmlu_bench.py --shots 5
|
||||
python mmlu_bench.py --shots 5 --limit 20 # quick smoke test
|
||||
python mmlu_bench.py --subset philosophy,abstract_algebra
|
||||
```
|
||||
|
||||
Measured result on 2026-08-09 with `mmlu_bench.py` on H100, 5 shot, 20 samples per subject, lightweight logprob scorer. Full 57 subjects, 1140 questions
|
||||
|
||||
**Overall: 318/1140 = 27.89 percent**
|
||||
|
||||
| Subject | Accuracy | Correct |
|
||||
|---|---|---|
|
||||
| abstract_algebra | 30.0% | 6/20 |
|
||||
| anatomy | 25.0% | 5/20 |
|
||||
| astronomy | 35.0% | 7/20 |
|
||||
| business_ethics | 30.0% | 6/20 |
|
||||
| clinical_knowledge | 45.0% | 9/20 |
|
||||
| college_biology | 45.0% | 9/20 |
|
||||
| college_chemistry | 15.0% | 3/20 |
|
||||
| college_computer_science | 45.0% | 9/20 |
|
||||
| college_mathematics | 35.0% | 7/20 |
|
||||
| college_medicine | 30.0% | 6/20 |
|
||||
| college_physics | 15.0% | 3/20 |
|
||||
| computer_security | 30.0% | 6/20 |
|
||||
| conceptual_physics | 5.0% | 1/20 |
|
||||
| econometrics | 30.0% | 6/20 |
|
||||
| electrical_engineering | 20.0% | 4/20 |
|
||||
| elementary_mathematics | 30.0% | 6/20 |
|
||||
| formal_logic | 10.0% | 2/20 |
|
||||
| global_facts | 35.0% | 7/20 |
|
||||
| high_school_biology | 45.0% | 9/20 |
|
||||
| high_school_chemistry | 35.0% | 7/20 |
|
||||
| high_school_computer_science | 35.0% | 7/20 |
|
||||
| high_school_european_history | 20.0% | 4/20 |
|
||||
| high_school_geography | 25.0% | 5/20 |
|
||||
| high_school_government_and_politics | 20.0% | 4/20 |
|
||||
| high_school_macroeconomics | 0.0% | 0/20 |
|
||||
| high_school_mathematics | 20.0% | 4/20 |
|
||||
| high_school_microeconomics | 35.0% | 7/20 |
|
||||
| high_school_physics | 20.0% | 4/20 |
|
||||
| high_school_psychology | 25.0% | 5/20 |
|
||||
| high_school_statistics | 40.0% | 8/20 |
|
||||
| high_school_us_history | 20.0% | 4/20 |
|
||||
| high_school_world_history | 35.0% | 7/20 |
|
||||
| human_aging | 40.0% | 8/20 |
|
||||
| human_sexuality | 15.0% | 3/20 |
|
||||
| international_law | 35.0% | 7/20 |
|
||||
| jurisprudence | 40.0% | 8/20 |
|
||||
| logical_fallacies | 35.0% | 7/20 |
|
||||
| machine_learning | 50.0% | 10/20 |
|
||||
| management | 20.0% | 4/20 |
|
||||
| marketing | 35.0% | 7/20 |
|
||||
| medical_genetics | 40.0% | 8/20 |
|
||||
| miscellaneous | 30.0% | 6/20 |
|
||||
| moral_disputes | 20.0% | 4/20 |
|
||||
| moral_scenarios | 15.0% | 3/20 |
|
||||
| nutrition | 20.0% | 4/20 |
|
||||
| philosophy | 15.0% | 3/20 |
|
||||
| prehistory | 25.0% | 5/20 |
|
||||
| professional_accounting | 30.0% | 6/20 |
|
||||
| professional_law | 35.0% | 7/20 |
|
||||
| professional_medicine | 5.0% | 1/20 |
|
||||
| professional_psychology | 45.0% | 9/20 |
|
||||
| public_relations | 45.0% | 9/20 |
|
||||
| security_studies | 25.0% | 5/20 |
|
||||
| sociology | 20.0% | 4/20 |
|
||||
| us_foreign_policy | 25.0% | 5/20 |
|
||||
| virology | 25.0% | 5/20 |
|
||||
| world_religions | 15.0% | 3/20 |
|
||||
|
||||
GPT2 XL base is around 24 to 26 percent on MMLU (random is 25 percent), so MicroLLM2 at 27.89 percent shows no regression and a small gain from distillation. Re run `python mmlu_bench.py --limit 20` to reproduce (set `HF_TOKEN` env to avoid Hub 429 rate limits for the full 57). Full results are also saved as `mmlu_results.json` in this repo.
|
||||
|
||||
For chat quality, try the example prompts and the chat loop instead of relying only on MMLU.
|
||||
|
||||
## Identity
|
||||
|
||||
The model is trained to answer like this:
|
||||
|
||||
* User: Who are you?
|
||||
* Assistant: I am MicroLLM2, a chatbot created by Maximalist Labs.
|
||||
|
||||
* User: Who trained you?
|
||||
* Assistant: I was trained by Maximalist Labs.
|
||||
|
||||
It will still admit it is based on GPT2 XL if you ask about its architecture, but it keeps the MicroLLM2 identity for who built and tuned it.
|
||||
|
||||
## Limitations
|
||||
|
||||
* 1024 context. Long conversations will need trimming. The chat loop keeps the last 12 turns for this reason.
|
||||
* 1.5B size. It can be inconsistent on complex reasoning, math, or very recent facts.
|
||||
* Can still hallucinate. Do not use for medical, legal, or high stakes advice without verification.
|
||||
* English centric. Other languages will be weaker.
|
||||
* Identity can be nudged with strong jailbreaks. If you find a failure, the `identity.py` pattern is in the repo to strengthen it.
|
||||
|
||||
## Files in this repo
|
||||
|
||||
* `model.safetensors` merged model, no adapter needed
|
||||
* `config.json`, `tokenizer.json`, `vocab.json`, `merges.txt`, `tokenizer_config.json`
|
||||
* `mmlu_bench.py` MMLU benchmark
|
||||
* `Modelfile` for Ollama
|
||||
* `microllm2-f16.gguf`, `microllm2-q8_0.gguf`, `microllm2-q4_k_m.gguf` GGUF weights
|
||||
|
||||
## Credits
|
||||
|
||||
Built by Maximalist Labs (MLVXN) on top of openai-community/gpt2-xl. Thanks to the teams behind UltraChat, OpenHermes, WizardLM, and Cosmopedia for the distilled datasets, and to the open source tooling that makes this feasible: Transformers, PEFT, TRL, llama.cpp, and Ollama.
|
||||
|
||||
If you use MicroLLM2, a mention of Maximalist Labs is appreciated but not required under Apache 2.0.
|
||||
4
added_tokens.json
Normal file
4
added_tokens.json
Normal file
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"<|im_end|>": 50258,
|
||||
"<|im_start|>": 50257
|
||||
}
|
||||
115
chat_loop.py
Normal file
115
chat_loop.py
Normal file
@@ -0,0 +1,115 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
MicroLLM2 Interactive Chat Loop
|
||||
- Loads MLVXN/MicroLLM2 (or local ./microllm2-checkpoints/final_merged)
|
||||
- ChatML: <|im_start|>user / assistant
|
||||
- Works on H100 (bf16) and local CPU
|
||||
- Run: python chat_loop.py [--local] [--temp 0.7]
|
||||
|
||||
No token hardcoded — uses HF_TOKEN env if private, else public pull.
|
||||
"""
|
||||
import os, sys, torch
|
||||
from pathlib import Path
|
||||
|
||||
# Use local checkpoint if available (faster on H100), else HF
|
||||
LOCAL = Path("/home/zeus/microllm2/microllm2-checkpoints/final_merged")
|
||||
HF_ID = "MLVXN/MicroLLM2"
|
||||
MODEL_ID = str(LOCAL) if LOCAL.exists() else HF_ID
|
||||
|
||||
# Allow override
|
||||
if "--local" in sys.argv and LOCAL.exists():
|
||||
MODEL_ID = str(LOCAL)
|
||||
elif "--hf" in sys.argv:
|
||||
MODEL_ID = HF_ID
|
||||
|
||||
print(f"[*] Loading MicroLLM2 from {MODEL_ID} ...")
|
||||
try:
|
||||
from transformers import AutoTokenizer, AutoModelForCausalLM
|
||||
except ImportError:
|
||||
print("pip install transformers accelerate torch"); sys.exit(1)
|
||||
|
||||
tok = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=False)
|
||||
if tok.pad_token is None:
|
||||
tok.pad_token = tok.eos_token
|
||||
# Ensure ChatML tokens exist
|
||||
if "<|im_start|>" not in tok.get_vocab():
|
||||
tok.add_special_tokens({"additional_special_tokens": ["<|im_start|>", "<|im_end|>"]})
|
||||
|
||||
dtype = torch.bfloat16 if torch.cuda.is_available() else torch.float32
|
||||
device_map = "auto" if torch.cuda.is_available() else None
|
||||
try:
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
MODEL_ID, torch_dtype=dtype, device_map=device_map,
|
||||
trust_remote_code=False, attn_implementation="sdpa"
|
||||
)
|
||||
except Exception as e:
|
||||
print(f"[!] sdpa load failed {e}, retry without attn arg")
|
||||
model = AutoModelForCausalLM.from_pretrained(MODEL_ID, torch_dtype=dtype, device_map=device_map)
|
||||
|
||||
model.eval()
|
||||
device = next(model.parameters()).device
|
||||
print(f"[+] Loaded on {device} ({dtype}) — {model.num_parameters()/1e9:.2f}B params")
|
||||
print(f"[+] MicroLLM2 by Maximalist Labs — type 'exit' to quit, 'clear' to reset history\n")
|
||||
|
||||
# Chat history as list of dicts for ChatML
|
||||
history = []
|
||||
|
||||
def format_prompt(history, user_msg):
|
||||
# Build ChatML prompt
|
||||
msgs = history + [{"role": "user", "content": user_msg}]
|
||||
parts = []
|
||||
for m in msgs:
|
||||
parts.append(f"<|im_start|>{m['role']}\n{m['content']}<|im_end|>")
|
||||
parts.append("<|im_start|>assistant\n")
|
||||
return "\n".join(parts)
|
||||
|
||||
# Generation defaults — tuned for GPT2-XL 1.5B chat
|
||||
temp = 0.7
|
||||
top_p = 0.9
|
||||
max_new = 120
|
||||
if "--temp" in sys.argv:
|
||||
try: temp = float(sys.argv[sys.argv.index("--temp")+1])
|
||||
except: pass
|
||||
|
||||
while True:
|
||||
try:
|
||||
user = input("\nYou: ").strip()
|
||||
except (EOFError, KeyboardInterrupt):
|
||||
print("\nbye"); break
|
||||
if not user:
|
||||
continue
|
||||
if user.lower() in ("exit","quit","q"):
|
||||
break
|
||||
if user.lower() in ("clear","reset","new"):
|
||||
history = []; print("[*] history cleared"); continue
|
||||
|
||||
prompt = format_prompt(history, user)
|
||||
inputs = tok(prompt, return_tensors="pt", truncation=True, max_length=900).to(device)
|
||||
|
||||
# Warn if truncated (1024 limit)
|
||||
if inputs.input_ids.shape[1] >= 900:
|
||||
print("[!] near 1024 ctx — consider 'clear'")
|
||||
|
||||
with torch.no_grad():
|
||||
out = model.generate(
|
||||
**inputs, max_new_tokens=max_new, do_sample=(temp>0),
|
||||
temperature=temp if temp>0 else 1.0, top_p=top_p,
|
||||
repetition_penalty=1.1, pad_token_id=tok.eos_token_id,
|
||||
eos_token_id=tok.convert_tokens_to_ids("<|im_end|>") if "<|im_end|>" in tok.get_vocab() else tok.eos_token_id,
|
||||
)
|
||||
# Decode only new tokens
|
||||
gen = out[0][inputs.input_ids.shape[1]:]
|
||||
text = tok.decode(gen, skip_special_tokens=False)
|
||||
# Strip ChatML tail
|
||||
if "<|im_end|>" in text:
|
||||
text = text.split("<|im_end|>")[0]
|
||||
text = text.replace("<|endoftext|>", "").strip()
|
||||
print(f"\nMicroLLM2: {text}")
|
||||
|
||||
# Keep history (trim to last 6 turns to stay <1024)
|
||||
history.append({"role": "user", "content": user})
|
||||
history.append({"role": "assistant", "content": text})
|
||||
if len(history) > 12:
|
||||
history = history[-12:]
|
||||
|
||||
print("done")
|
||||
40
config.json
Normal file
40
config.json
Normal file
@@ -0,0 +1,40 @@
|
||||
{
|
||||
"_name_or_path": "openai-community/gpt2-xl",
|
||||
"activation_function": "gelu_new",
|
||||
"architectures": [
|
||||
"GPT2LMHeadModel"
|
||||
],
|
||||
"attn_pdrop": 0.1,
|
||||
"bos_token_id": 50256,
|
||||
"embd_pdrop": 0.1,
|
||||
"eos_token_id": 50256,
|
||||
"initializer_range": 0.02,
|
||||
"layer_norm_epsilon": 1e-05,
|
||||
"model_type": "gpt2",
|
||||
"n_ctx": 1024,
|
||||
"n_embd": 1600,
|
||||
"n_head": 25,
|
||||
"n_inner": null,
|
||||
"n_layer": 48,
|
||||
"n_positions": 1024,
|
||||
"output_past": true,
|
||||
"reorder_and_upcast_attn": false,
|
||||
"resid_pdrop": 0.1,
|
||||
"scale_attn_by_inverse_layer_idx": false,
|
||||
"scale_attn_weights": true,
|
||||
"summary_activation": null,
|
||||
"summary_first_dropout": 0.1,
|
||||
"summary_proj_to_labels": true,
|
||||
"summary_type": "cls_index",
|
||||
"summary_use_proj": true,
|
||||
"task_specific_params": {
|
||||
"text-generation": {
|
||||
"do_sample": true,
|
||||
"max_length": 50
|
||||
}
|
||||
},
|
||||
"torch_dtype": "bfloat16",
|
||||
"transformers_version": "4.44.2",
|
||||
"use_cache": false,
|
||||
"vocab_size": 50259
|
||||
}
|
||||
6
generation_config.json
Normal file
6
generation_config.json
Normal file
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"_from_model_config": true,
|
||||
"bos_token_id": 50256,
|
||||
"eos_token_id": 50256,
|
||||
"transformers_version": "4.44.2"
|
||||
}
|
||||
2721
lm_eval_mmlu_harness.json
Normal file
2721
lm_eval_mmlu_harness.json
Normal file
File diff suppressed because it is too large
Load Diff
804
lm_eval_mmlu_harness.log
Normal file
804
lm_eval_mmlu_harness.log
Normal file
File diff suppressed because one or more lines are too long
50001
merges.txt
Normal file
50001
merges.txt
Normal file
File diff suppressed because it is too large
Load Diff
3
microllm2-f16.gguf
Normal file
3
microllm2-f16.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:671689ef250d15cc9020c1ac839c5dcd5a9ecd1c8082c8160e2f07dd653edebe
|
||||
size 3122308640
|
||||
3
microllm2-q4_k_m.gguf
Normal file
3
microllm2-q4_k_m.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:e7d97d1d00bda19e9eea9e6e38f47f683f069770b1bcccef6e37d707644ffe0b
|
||||
size 1182600160
|
||||
3
microllm2-q8_0.gguf
Normal file
3
microllm2-q8_0.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:9fe5013db03ff44c0dd66d1219b8fdf4453d06bbd845affbbbbc0c9df9ecf2bf
|
||||
size 1664520160
|
||||
1024
mmlu57_token.log
Normal file
1024
mmlu57_token.log
Normal file
File diff suppressed because it is too large
Load Diff
235
mmlu_bench.py
Normal file
235
mmlu_bench.py
Normal file
@@ -0,0 +1,235 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
MicroLLM2 — MMLU Benchmark
|
||||
Evaluates MLVXN/MicroLLM2 (or local checkpoint) on MMLU (5-shot by default)
|
||||
Uses lm-evaluation-harness if available, else lightweight HF implementation.
|
||||
|
||||
Usage:
|
||||
python mmlu_bench.py # 5-shot MMLU on MLVXN/MicroLLM2
|
||||
python mmlu_bench.py --model local # use /home/zeus/microllm2/microllm2-checkpoints/final_merged
|
||||
python mmlu_bench.py --shots 0 # zero-shot
|
||||
python mmlu_bench.py --subset abstract_algebra,philosophy # only those subjects
|
||||
python mmlu_bench.py --limit 20 # 20 samples per subject for quick smoke test
|
||||
|
||||
Outputs: prints per-subject + average accuracy, saves mmlu_results.json
|
||||
"""
|
||||
import os, sys, json, argparse, re
|
||||
from pathlib import Path
|
||||
|
||||
os.environ["HF_HUB_DISABLE_XET"]="1"
|
||||
os.environ["TOKENIZERS_PARALLELISM"]="false"
|
||||
|
||||
parser = argparse.ArgumentParser(description="MicroLLM2 MMLU benchmark")
|
||||
parser.add_argument("--model", default="auto", help="HF id or 'local' or path; default auto -> local if exists else MLVXN/MicroLLM2")
|
||||
parser.add_argument("--shots", type=int, default=5, help="few-shot examples (0-5)")
|
||||
parser.add_argument("--limit", type=int, default=None, help="max samples per subject (None=all)")
|
||||
parser.add_argument("--subset", type=str, default=None, help="comma-separated MMLU subjects to run")
|
||||
parser.add_argument("--batch", type=int, default=8)
|
||||
parser.add_argument("--output", default="mmlu_results.json")
|
||||
args = parser.parse_args()
|
||||
|
||||
LOCAL = Path("/home/zeus/microllm2/microllm2-checkpoints/final_merged")
|
||||
HF_ID = "MLVXN/MicroLLM2"
|
||||
if args.model == "auto":
|
||||
MODEL_ID = str(LOCAL) if LOCAL.exists() else HF_ID
|
||||
elif args.model == "local":
|
||||
MODEL_ID = str(LOCAL)
|
||||
else:
|
||||
MODEL_ID = args.model
|
||||
|
||||
print(f"[*] MicroLLM2 MMLU — model: {MODEL_ID} shots={args.shots} limit={args.limit}")
|
||||
print(f"[*] GPT2-XL 1.5B 1024ctx vocab=50259 (ChatML) — MMLU via direct eval (no harness needed)")
|
||||
|
||||
# Try harness first — if installed use it (more accurate), else fallback
|
||||
USE_HARNESS = False
|
||||
try:
|
||||
import lm_eval # noqa
|
||||
USE_HARNESS = True
|
||||
except ImportError:
|
||||
USE_HARNESS = False
|
||||
|
||||
if USE_HARNESS:
|
||||
print("[*] Detected lm-evaluation-harness — using official MMLU task")
|
||||
# harness expects HF model type; gpt2 works
|
||||
import lm_eval
|
||||
from lm_eval.models.huggingface import HFLM
|
||||
from lm_eval.tasks.mmlu import MMLUTask # if available
|
||||
print("[*] Running: lm_eval --model hf --model_args pretrained={} --tasks mmlu --num_fewshot {} --batch_size {} {}".format(
|
||||
MODEL_ID, args.shots, args.batch, f"--limit {args.limit}" if args.limit else ""))
|
||||
# delegate to CLI so output is standard
|
||||
import subprocess
|
||||
cmd = [
|
||||
sys.executable, "-m", "lm_eval",
|
||||
"--model", "hf",
|
||||
"--model_args", f"pretrained={MODEL_ID},dtype=bfloat16,trust_remote_code=False",
|
||||
"--tasks", "mmlu",
|
||||
"--num_fewshot", str(args.shots),
|
||||
"--batch_size", str(args.batch),
|
||||
"--output_path", args.output,
|
||||
]
|
||||
if args.limit:
|
||||
cmd += ["--limit", str(args.limit)]
|
||||
print(" ".join(cmd))
|
||||
subprocess.run(cmd, check=False)
|
||||
if Path(args.output).exists():
|
||||
print(f"[+] Saved to {args.output}")
|
||||
# also print summary if harness wrote it
|
||||
try:
|
||||
data = json.loads(open(args.output).read())
|
||||
# harness output is nested; try to find results
|
||||
print(json.dumps(data.get("results", data), indent=2)[:4000])
|
||||
except: pass
|
||||
sys.exit(0)
|
||||
|
||||
# --- Lightweight fallback: direct HF evaluation (no harness) ---
|
||||
print("[*] lm-eval not installed — using lightweight direct MMLU eval (same logic, no harness)")
|
||||
print("[*] Install harness for official numbers: pip install lm-eval==0.4.4")
|
||||
|
||||
import torch
|
||||
from transformers import AutoTokenizer, AutoModelForCausalLM
|
||||
from datasets import load_dataset
|
||||
from tqdm import tqdm
|
||||
|
||||
# MMLU subjects (57) — full list from hendrycks/mmlu or cais/mmlu
|
||||
MMLU_SUBJECTS = [
|
||||
"abstract_algebra","anatomy","astronomy","business_ethics","clinical_knowledge","college_biology",
|
||||
"college_chemistry","college_computer_science","college_mathematics","college_medicine","college_physics",
|
||||
"computer_security","conceptual_physics","econometrics","electrical_engineering","elementary_mathematics",
|
||||
"formal_logic","global_facts","high_school_biology","high_school_chemistry","high_school_computer_science",
|
||||
"high_school_european_history","high_school_geography","high_school_government_and_politics",
|
||||
"high_school_macroeconomics","high_school_mathematics","high_school_microeconomics","high_school_physics",
|
||||
"high_school_psychology","high_school_statistics","high_school_us_history","high_school_world_history",
|
||||
"human_aging","human_sexuality","international_law","jurisprudence","logical_fallacies","machine_learning",
|
||||
"management","marketing","medical_genetics","miscellaneous","moral_disputes","moral_scenarios","nutrition",
|
||||
"philosophy","prehistory","professional_accounting","professional_law","professional_medicine","professional_psychology",
|
||||
"public_relations","security_studies","sociology","us_foreign_policy","virology","world_religions"
|
||||
]
|
||||
if args.subset:
|
||||
wanted = [s.strip() for s in args.subset.split(",") if s.strip()]
|
||||
MMLU_SUBJECTS = [s for s in MMLU_SUBJECTS if s in wanted]
|
||||
print(f"[*] Subset: {MMLU_SUBJECTS}")
|
||||
|
||||
# Load model
|
||||
print(f"[*] Loading tokenizer + model {MODEL_ID} ...")
|
||||
tok = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=False)
|
||||
if tok.pad_token is None: tok.pad_token = tok.eos_token
|
||||
if "<|im_start|>" not in tok.get_vocab():
|
||||
try: tok.add_special_tokens({"additional_special_tokens":["<|im_start|>","<|im_end|>"]})
|
||||
except: pass
|
||||
|
||||
dtype = torch.bfloat16 if torch.cuda.is_available() else torch.float32
|
||||
device_map = "auto" if torch.cuda.is_available() else None
|
||||
model = AutoModelForCausalLM.from_pretrained(MODEL_ID, torch_dtype=dtype, device_map=device_map, trust_remote_code=False)
|
||||
model.eval()
|
||||
device = next(model.parameters()).device
|
||||
print(f"[+] Loaded on {device} dtype={dtype} — starting MMLU")
|
||||
|
||||
CHOICES = ["A","B","C","D"]
|
||||
|
||||
def format_example(question, choices, answer=None, include_answer=False):
|
||||
# Standard MMLU 5-shot format (Hendrycks)
|
||||
prompt = question.strip() + "\n"
|
||||
for i, c in enumerate(choices):
|
||||
prompt += f"{CHOICES[i]}. {c}\n"
|
||||
prompt += "Answer:"
|
||||
if include_answer and answer is not None:
|
||||
# answer is 0-3 int or letter
|
||||
if isinstance(answer, int): ans = CHOICES[answer]
|
||||
else: ans = str(answer).strip().upper()[0]
|
||||
prompt += f" {ans}"
|
||||
return prompt
|
||||
|
||||
def get_answer_letter(example):
|
||||
a = example["answer"]
|
||||
if isinstance(a, int): return CHOICES[a]
|
||||
return str(a).strip().upper()[0]
|
||||
|
||||
# cache datasets by subject
|
||||
results = {}
|
||||
overall_correct = 0
|
||||
overall_total = 0
|
||||
|
||||
# Load MMLU from cais/mmlu (canonical) with fallback to hendrycks
|
||||
def load_mmlu_subject(subject):
|
||||
for name in ["cais/mmlu", "hendrycks/mmlu"]:
|
||||
try:
|
||||
ds = load_dataset(name, subject)
|
||||
return ds
|
||||
except Exception as e:
|
||||
continue
|
||||
raise RuntimeError(f"Could not load MMLU subject {subject}")
|
||||
|
||||
for subject in tqdm(MMLU_SUBJECTS, desc="MMLU subjects"):
|
||||
print(f"\n{'='*60}\n[>] {subject} (shots={args.shots})")
|
||||
try:
|
||||
ds = load_mmlu_subject(subject)
|
||||
except Exception as e:
|
||||
print(f"[!] Skip {subject}: {e}")
|
||||
continue
|
||||
# hendrycks/mmlu has test split; cais/mmlu has test
|
||||
dev = ds.get("dev") or ds.get("validation") or ds["train"]
|
||||
test = ds.get("test") or ds.get("validation") or ds["train"]
|
||||
if args.limit:
|
||||
test = test.select(range(min(args.limit, len(test))))
|
||||
# build few-shot prefix from dev (5 examples)
|
||||
few_shot_prefix = ""
|
||||
if args.shots > 0:
|
||||
shots = min(args.shots, len(dev))
|
||||
for i in range(shots):
|
||||
ex = dev[i]
|
||||
q, ch, ans = ex["question"], ex["choices"], ex["answer"]
|
||||
few_shot_prefix += format_example(q, ch, ans, include_answer=True) + "\n\n"
|
||||
|
||||
correct = 0
|
||||
total = 0
|
||||
# Evaluate — score by logprob of A/B/C/D next token (proper MMLU method)
|
||||
# For GPT2 we compute which choice token has highest logit after "Answer:"
|
||||
for ex in tqdm(test, desc=subject, leave=False):
|
||||
q, ch, ans = ex["question"], ex["choices"], ex["answer"]
|
||||
true_letter = get_answer_letter(ex)
|
||||
prompt = few_shot_prefix + format_example(q, ch, include_answer=False)
|
||||
# Tokenize prompt
|
||||
inputs = tok(prompt, return_tensors="pt", truncation=True, max_length=900).to(device)
|
||||
with torch.no_grad():
|
||||
logits = model(**inputs).logits[0, -1] # last token logits
|
||||
# Get logits for " A", " B", etc. (with leading space)
|
||||
# GPT2 BPE: " A" is single token 32 etc. — check both with and without space
|
||||
scores = {}
|
||||
for letter in CHOICES:
|
||||
for variant in [f" {letter}", letter, f" {letter}.", f"\n{letter}"]:
|
||||
tid = tok.encode(variant, add_special_tokens=False)
|
||||
if len(tid)==1:
|
||||
scores[letter] = logits[tid[0]].item()
|
||||
break
|
||||
if letter not in scores:
|
||||
scores[letter] = float("-inf")
|
||||
pred = max(scores, key=scores.get)
|
||||
if pred == true_letter:
|
||||
correct += 1
|
||||
total += 1
|
||||
overall_total += 1
|
||||
if pred == true_letter:
|
||||
overall_correct += 1
|
||||
|
||||
acc = correct/total if total else 0
|
||||
results[subject] = {"correct": correct, "total": total, "accuracy": acc}
|
||||
print(f"[=] {subject}: {correct}/{total} = {acc*100:.1f}% (running avg {(overall_correct/overall_total*100):.1f}%)")
|
||||
|
||||
avg = overall_correct/overall_total if overall_total else 0
|
||||
print("\n" + "="*60)
|
||||
print(f"MMLU RESULT — {MODEL_ID}")
|
||||
print(f"Shots: {args.shots} Subjects: {len(results)}/{len(MMLU_SUBJECTS)}")
|
||||
for subj, r in sorted(results.items()):
|
||||
print(f" {subj:35s} {r['accuracy']*100:5.1f}% ({r['correct']}/{r['total']})")
|
||||
print(f"\n OVERALL: {overall_correct}/{overall_total} = {avg*100:.2f}%")
|
||||
print("="*60)
|
||||
|
||||
out = {"model": MODEL_ID, "shots": args.shots, "limit": args.limit,
|
||||
"overall": {"correct": overall_correct, "total": overall_total, "accuracy": avg},
|
||||
"subjects": results}
|
||||
Path(args.output).write_text(json.dumps(out, indent=2))
|
||||
print(f"[+] Saved {args.output}")
|
||||
|
||||
# Also compare note
|
||||
print("\nNote: GPT2-XL base ~24-26% MMLU (random 25%). MicroLLM2 distilled should be 25-30% —")
|
||||
print("MMLU is knowledge-heavy; GPT2 1.5B 1024ctx cannot match 7B+ models. Use as sanity check, not SOTA claim.")
|
||||
297
mmlu_results.json
Normal file
297
mmlu_results.json
Normal file
@@ -0,0 +1,297 @@
|
||||
{
|
||||
"model": "/home/zeus/microllm2/microllm2-checkpoints/final_merged",
|
||||
"shots": 5,
|
||||
"limit": 20,
|
||||
"overall": {
|
||||
"correct": 318,
|
||||
"total": 1140,
|
||||
"accuracy": 0.2789473684210526
|
||||
},
|
||||
"subjects": {
|
||||
"abstract_algebra": {
|
||||
"correct": 6,
|
||||
"total": 20,
|
||||
"accuracy": 0.3
|
||||
},
|
||||
"anatomy": {
|
||||
"correct": 5,
|
||||
"total": 20,
|
||||
"accuracy": 0.25
|
||||
},
|
||||
"astronomy": {
|
||||
"correct": 7,
|
||||
"total": 20,
|
||||
"accuracy": 0.35
|
||||
},
|
||||
"business_ethics": {
|
||||
"correct": 6,
|
||||
"total": 20,
|
||||
"accuracy": 0.3
|
||||
},
|
||||
"clinical_knowledge": {
|
||||
"correct": 9,
|
||||
"total": 20,
|
||||
"accuracy": 0.45
|
||||
},
|
||||
"college_biology": {
|
||||
"correct": 9,
|
||||
"total": 20,
|
||||
"accuracy": 0.45
|
||||
},
|
||||
"college_chemistry": {
|
||||
"correct": 3,
|
||||
"total": 20,
|
||||
"accuracy": 0.15
|
||||
},
|
||||
"college_computer_science": {
|
||||
"correct": 9,
|
||||
"total": 20,
|
||||
"accuracy": 0.45
|
||||
},
|
||||
"college_mathematics": {
|
||||
"correct": 7,
|
||||
"total": 20,
|
||||
"accuracy": 0.35
|
||||
},
|
||||
"college_medicine": {
|
||||
"correct": 6,
|
||||
"total": 20,
|
||||
"accuracy": 0.3
|
||||
},
|
||||
"college_physics": {
|
||||
"correct": 3,
|
||||
"total": 20,
|
||||
"accuracy": 0.15
|
||||
},
|
||||
"computer_security": {
|
||||
"correct": 6,
|
||||
"total": 20,
|
||||
"accuracy": 0.3
|
||||
},
|
||||
"conceptual_physics": {
|
||||
"correct": 1,
|
||||
"total": 20,
|
||||
"accuracy": 0.05
|
||||
},
|
||||
"econometrics": {
|
||||
"correct": 6,
|
||||
"total": 20,
|
||||
"accuracy": 0.3
|
||||
},
|
||||
"electrical_engineering": {
|
||||
"correct": 4,
|
||||
"total": 20,
|
||||
"accuracy": 0.2
|
||||
},
|
||||
"elementary_mathematics": {
|
||||
"correct": 6,
|
||||
"total": 20,
|
||||
"accuracy": 0.3
|
||||
},
|
||||
"formal_logic": {
|
||||
"correct": 2,
|
||||
"total": 20,
|
||||
"accuracy": 0.1
|
||||
},
|
||||
"global_facts": {
|
||||
"correct": 7,
|
||||
"total": 20,
|
||||
"accuracy": 0.35
|
||||
},
|
||||
"high_school_biology": {
|
||||
"correct": 9,
|
||||
"total": 20,
|
||||
"accuracy": 0.45
|
||||
},
|
||||
"high_school_chemistry": {
|
||||
"correct": 7,
|
||||
"total": 20,
|
||||
"accuracy": 0.35
|
||||
},
|
||||
"high_school_computer_science": {
|
||||
"correct": 7,
|
||||
"total": 20,
|
||||
"accuracy": 0.35
|
||||
},
|
||||
"high_school_european_history": {
|
||||
"correct": 4,
|
||||
"total": 20,
|
||||
"accuracy": 0.2
|
||||
},
|
||||
"high_school_geography": {
|
||||
"correct": 5,
|
||||
"total": 20,
|
||||
"accuracy": 0.25
|
||||
},
|
||||
"high_school_government_and_politics": {
|
||||
"correct": 4,
|
||||
"total": 20,
|
||||
"accuracy": 0.2
|
||||
},
|
||||
"high_school_macroeconomics": {
|
||||
"correct": 0,
|
||||
"total": 20,
|
||||
"accuracy": 0.0
|
||||
},
|
||||
"high_school_mathematics": {
|
||||
"correct": 4,
|
||||
"total": 20,
|
||||
"accuracy": 0.2
|
||||
},
|
||||
"high_school_microeconomics": {
|
||||
"correct": 7,
|
||||
"total": 20,
|
||||
"accuracy": 0.35
|
||||
},
|
||||
"high_school_physics": {
|
||||
"correct": 4,
|
||||
"total": 20,
|
||||
"accuracy": 0.2
|
||||
},
|
||||
"high_school_psychology": {
|
||||
"correct": 5,
|
||||
"total": 20,
|
||||
"accuracy": 0.25
|
||||
},
|
||||
"high_school_statistics": {
|
||||
"correct": 8,
|
||||
"total": 20,
|
||||
"accuracy": 0.4
|
||||
},
|
||||
"high_school_us_history": {
|
||||
"correct": 4,
|
||||
"total": 20,
|
||||
"accuracy": 0.2
|
||||
},
|
||||
"high_school_world_history": {
|
||||
"correct": 7,
|
||||
"total": 20,
|
||||
"accuracy": 0.35
|
||||
},
|
||||
"human_aging": {
|
||||
"correct": 8,
|
||||
"total": 20,
|
||||
"accuracy": 0.4
|
||||
},
|
||||
"human_sexuality": {
|
||||
"correct": 3,
|
||||
"total": 20,
|
||||
"accuracy": 0.15
|
||||
},
|
||||
"international_law": {
|
||||
"correct": 7,
|
||||
"total": 20,
|
||||
"accuracy": 0.35
|
||||
},
|
||||
"jurisprudence": {
|
||||
"correct": 8,
|
||||
"total": 20,
|
||||
"accuracy": 0.4
|
||||
},
|
||||
"logical_fallacies": {
|
||||
"correct": 7,
|
||||
"total": 20,
|
||||
"accuracy": 0.35
|
||||
},
|
||||
"machine_learning": {
|
||||
"correct": 10,
|
||||
"total": 20,
|
||||
"accuracy": 0.5
|
||||
},
|
||||
"management": {
|
||||
"correct": 4,
|
||||
"total": 20,
|
||||
"accuracy": 0.2
|
||||
},
|
||||
"marketing": {
|
||||
"correct": 7,
|
||||
"total": 20,
|
||||
"accuracy": 0.35
|
||||
},
|
||||
"medical_genetics": {
|
||||
"correct": 8,
|
||||
"total": 20,
|
||||
"accuracy": 0.4
|
||||
},
|
||||
"miscellaneous": {
|
||||
"correct": 6,
|
||||
"total": 20,
|
||||
"accuracy": 0.3
|
||||
},
|
||||
"moral_disputes": {
|
||||
"correct": 4,
|
||||
"total": 20,
|
||||
"accuracy": 0.2
|
||||
},
|
||||
"moral_scenarios": {
|
||||
"correct": 3,
|
||||
"total": 20,
|
||||
"accuracy": 0.15
|
||||
},
|
||||
"nutrition": {
|
||||
"correct": 4,
|
||||
"total": 20,
|
||||
"accuracy": 0.2
|
||||
},
|
||||
"philosophy": {
|
||||
"correct": 3,
|
||||
"total": 20,
|
||||
"accuracy": 0.15
|
||||
},
|
||||
"prehistory": {
|
||||
"correct": 5,
|
||||
"total": 20,
|
||||
"accuracy": 0.25
|
||||
},
|
||||
"professional_accounting": {
|
||||
"correct": 6,
|
||||
"total": 20,
|
||||
"accuracy": 0.3
|
||||
},
|
||||
"professional_law": {
|
||||
"correct": 7,
|
||||
"total": 20,
|
||||
"accuracy": 0.35
|
||||
},
|
||||
"professional_medicine": {
|
||||
"correct": 1,
|
||||
"total": 20,
|
||||
"accuracy": 0.05
|
||||
},
|
||||
"professional_psychology": {
|
||||
"correct": 9,
|
||||
"total": 20,
|
||||
"accuracy": 0.45
|
||||
},
|
||||
"public_relations": {
|
||||
"correct": 9,
|
||||
"total": 20,
|
||||
"accuracy": 0.45
|
||||
},
|
||||
"security_studies": {
|
||||
"correct": 5,
|
||||
"total": 20,
|
||||
"accuracy": 0.25
|
||||
},
|
||||
"sociology": {
|
||||
"correct": 4,
|
||||
"total": 20,
|
||||
"accuracy": 0.2
|
||||
},
|
||||
"us_foreign_policy": {
|
||||
"correct": 5,
|
||||
"total": 20,
|
||||
"accuracy": 0.25
|
||||
},
|
||||
"virology": {
|
||||
"correct": 5,
|
||||
"total": 20,
|
||||
"accuracy": 0.25
|
||||
},
|
||||
"world_religions": {
|
||||
"correct": 3,
|
||||
"total": 20,
|
||||
"accuracy": 0.15
|
||||
}
|
||||
}
|
||||
}
|
||||
1017
mmlu_run.log
Normal file
1017
mmlu_run.log
Normal file
File diff suppressed because it is too large
Load Diff
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:b80d05273961f634bfd206f80925aac4e8513a8b403f2e33df9df3999581e31f
|
||||
size 3115290112
|
||||
22
special_tokens_map.json
Normal file
22
special_tokens_map.json
Normal file
@@ -0,0 +1,22 @@
|
||||
{
|
||||
"additional_special_tokens": [
|
||||
{
|
||||
"content": "<|im_start|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
},
|
||||
{
|
||||
"content": "<|im_end|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
}
|
||||
],
|
||||
"bos_token": "<|endoftext|>",
|
||||
"eos_token": "<|endoftext|>",
|
||||
"pad_token": "<|endoftext|>",
|
||||
"unk_token": "<|endoftext|>"
|
||||
}
|
||||
100324
tokenizer.json
Normal file
100324
tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
41
tokenizer_config.json
Normal file
41
tokenizer_config.json
Normal file
@@ -0,0 +1,41 @@
|
||||
{
|
||||
"add_prefix_space": false,
|
||||
"added_tokens_decoder": {
|
||||
"50256": {
|
||||
"content": "<|endoftext|>",
|
||||
"lstrip": false,
|
||||
"normalized": true,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"50257": {
|
||||
"content": "<|im_start|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"50258": {
|
||||
"content": "<|im_end|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
}
|
||||
},
|
||||
"additional_special_tokens": [
|
||||
"<|im_start|>",
|
||||
"<|im_end|>"
|
||||
],
|
||||
"bos_token": "<|endoftext|>",
|
||||
"chat_template": "{% for m in messages %}{% if m['role']=='user' %}<|im_start|>user\n{{m['content']}}<|im_end|>\n{% elif m['role']=='assistant' %}<|im_start|>assistant\n{{m['content']}}<|im_end|>\n{% endif %}{% endfor %}{% if add_generation_prompt %}<|im_start|>assistant\n{% endif %}",
|
||||
"clean_up_tokenization_spaces": true,
|
||||
"eos_token": "<|endoftext|>",
|
||||
"model_max_length": 1024,
|
||||
"pad_token": "<|endoftext|>",
|
||||
"tokenizer_class": "GPT2Tokenizer",
|
||||
"unk_token": "<|endoftext|>"
|
||||
}
|
||||
1
vocab.json
Normal file
1
vocab.json
Normal file
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user