---
license: apache-2.0
language:
- en
tags:
- large-language-model
- qwen3
qwen:
- physics
smeft:
- effective-field-theory
high-energy-physics:
- scientific-llm
pipeline_tag: text-generation
library_name: transformers
base_model: Qwen/Qwen3-8B
model_type: qwen3
do_sample: true
temperature: 0.1
top_p: 0.1
repetition_penalty: 1.1
max_new_tokens: 512
---
A domain-adapted large language model for
Standard Model Effective Field Theory (SMEFT).
# Capabilities
The model has been optimized for:
- SMEFT operator reasoning
- EFT basis translation
- Physics-aware scientific dialogue
- Literature-style technical explanation
- Structured theoretical question answering
This model is fine-tuned from **Qwen3-8B** using a curated corpus of SMEFT and particle physics literature.
> **Note on Qwen3 thinking mode:** Qwen3 supports a native `enable_thinking` toggle in its chat template that wraps reasoning in `...` blocks. This model was tuned to produce reasoning through its own structured prompt format (not Qwen3's native thinking blocks), so if you build prompts via `tokenizer.apply_chat_template(...)`, set `enable_thinking=False` to avoid mixing the two reasoning styles. The example below uses a raw instruction-style prompt and is unaffected either way.
# Quick Start
## Installation
```bash
pip install torch transformers bitsandbytes accelerate
```
## Inference Example
```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
MODEL_NAME = "ahammad115566/qwen-smeft"
RESPONSE_PREFIX = "\n### Response:\n"
# 4-bit quantisation
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True,
)
# Load model and tokenizer
tokenizer = AutoTokenizer.from_pretrained(
MODEL_NAME,
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
MODEL_NAME,
quantization_config=bnb_config,
device_map={"": 0},
trust_remote_code=True,
)
model.eval()
if tokenizer.pad_token is None:
tokenizer.pad_token = tokenizer.eos_token
# ================= Inference ================
def build_prompt(instruction: str) -> str:
"""Construct the prompt using the format used during fine-tuning."""
return f"\n### Instruction:\n{instruction}\n{RESPONSE_PREFIX}"
@torch.inference_mode()
def ask(instruction: str) -> str:
"""Generate a response to an SMEFT-related instruction."""
prompt = build_prompt(instruction)
inputs = tokenizer(
prompt,
return_tensors="pt",
add_special_tokens=True,
).to(model.device)
output_ids = model.generate(
**inputs,
max_new_tokens=2048,
do_sample=False,
repetition_penalty=1.1,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.pad_token_id,
)
new_tokens = output_ids[0][inputs["input_ids"].shape[-1]:]
return tokenizer.decode(
new_tokens,
skip_special_tokens=True,
).strip()
# Example
instruction = """
Which SMEFT operators modify EWPO?
"""
response = ask(instruction)
print(response)
```
---
## Training Details
| Parameter | Value |
|-----------|-------|
| **Base Model** | Qwen3 |
| **Fine-tuning Method** | LoRA (merged) |
| **Inference Quantization** | 4-bit NF4 (bitsandbytes) |
| **Domain** | Standard Model Effective Field Theory (SMEFT) |
| **Training Corpus** | Curated SMEFT and HEP preprints |
| **Task Format** | Instruction-following scientific QA |
---
# About the model
- May hallucinate operator identities.
- Domain-locked by design. The model is not suitable for general-purpose tasks.
- 3620 training examples. Coverage of the SMEFT operator space may be uneven; rare operators or non-Warsaw bases may be answered less reliably.
- 20% out of domain exmaples are included during the training.
---
# Authors
**Ahmed Hammad**
Assistant professor
Center of AI and natural science, KIAS, Seoul.
**Veronica Sanz**
Professor of Theoretical Physics
University of Valencia
# Citation
A technical paper describing the dataset construction and fine-tuning procedure is forthcoming.
Please cite the model as:
```bibtex
@article{Hammad:2026bvw,
author = "Hammad, Ahmed and Sanz, Veronica",
title = "{Language-Guided Hypotheses Generation for Sparse SMEFT Analyses}",
eprint = "2608.04100",
archivePrefix = "arXiv",
primaryClass = "hep-ph",
month = "8",
year = "2026"
}
```