license, base_model, tags, library_name
license base_model tags library_name
apache-2.0 Qwen/Qwen3-1.7B
biology
bio-posttrain
dna-sft
dna
transformers

Bio-posttrain Qwen3-1.7B DNA SFT

DNA supervised fine-tuning (SFT) checkpoint from How Post-Training Shapes Biological Reasoning Models.

Model details

  • Base model: Qwen/Qwen3-1.7B
  • DNA encoder: Evo2 evo2_1b_base (frozen; not included in this repo)
  • Embedding layer: blocks.20.mlp.l3
  • LoRA: rank 64, alpha 128
  • Validation loss: 0.4687

This repo contains the merged text LLM (LoRA fused into base weights) plus dna_projection.pt.

Files

File Description
model.safetensors Merged Qwen3-1.7B weights
dna_projection.pt Linear map from Evo2 hidden (1920) → text hidden (2048)
dna_model_config.json DNA encoder + projection metadata

Loading

Use the BioReason DNALLMModel loader:

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

text_model = AutoModelForCausalLM.from_pretrained("mims-harvard/bio-posttrain-qwen3-1.7b-dna-sft")
tokenizer = AutoTokenizer.from_pretrained("mims-harvard/bio-posttrain-qwen3-1.7b-dna-sft")
proj = torch.load("mims-harvard/bio-posttrain-qwen3-1.7b-dna-sft/dna_projection.pt", map_location="cpu")
# Load Evo2 separately: evo2_1b_base

See dna_model_config.json for encoder settings.

Collection

Part of the Bio-posttrain collection.

Description
Model synced from source: mims-harvard/bio-posttrain-qwen3-1.7b-dna-sft
Readme 28 KiB
Languages
Jinja 100%