A GPT-2 small (124M parameter) language model trained from scratch on Dutch text, then fine-tuned for instruction following using supervised fine-tuning (SFT). This model understands and generates Dutch.
Model details
Property
Value
Architecture
GPT-2 small
Parameters
123.8M
Layers
12
Attention heads
12
Hidden dimension
768
Context length
512 tokens
Vocabulary size
50,000 (Dutch BPE)
Weights
fp16 / safetensors (473 MB)
Inference speed (CPU)
0.9 tok/s
Files
File
Format
Size
model.safetensors
fp16
473 MB
dutch-gpt2-f16.gguf
GGUF F16
249 MB
dutch-gpt2-q8_0.gguf
GGUF Q8_0
132 MB
Use with llama.cpp
# Download
wget https://huggingface.co/Thorstin/gpt2-dutch-instruct/resolve/main/dutch-gpt2-q8_0.gguf
# Run
llama-cli -m dutch-gpt2-q8_0.gguf \
-p "### Instructie:\nWat is de hoofdstad van Nederland?\n### Antwoord:\n"\
-n 200
### Instructie:
<vraag of instructie>
### Antwoord:
<antwoord>
Usage
fromtransformersimportAutoModelForCausalLM,AutoTokenizerimporttorchmodel_id="Thorstin/gpt2-dutch-instruct"tokenizer=AutoTokenizer.from_pretrained(model_id)model=AutoModelForCausalLM.from_pretrained(model_id,torch_dtype=torch.float16)model.eval()defchat(instruction:str,max_new_tokens:int=200)->str:prompt=f"### Instructie:\n{instruction}\n### Antwoord:\n"inputs=tokenizer(prompt,return_tensors="pt")withtorch.no_grad():output=model.generate(**inputs,max_new_tokens=max_new_tokens,do_sample=True,temperature=0.7,top_p=0.9,repetition_penalty=1.3,pad_token_id=tokenizer.eos_token_id,)response=tokenizer.decode(output[0],skip_special_tokens=True)returnresponse.split("### Antwoord:")[-1].strip()print(chat("Wat is de hoofdstad van Nederland?"))