Model: mims-harvard/bio-posttrain-qwen3-1.7b-dna-rl Source: Original Platform
license, base_model, tags, library_name
| license | base_model | tags | library_name | ||||
|---|---|---|---|---|---|---|---|
| apache-2.0 | Qwen/Qwen3-1.7B |
|
transformers |
Bio-posttrain Qwen3-1.7B DNA RL
DNA reinforcement learning (GRPO) checkpoint from How Post-Training Shapes Biological Reasoning Models.
Model details
- Base model:
Qwen/Qwen3-1.7B - DNA encoder: Evo2
evo2_1b_base(frozen; not included) - Embedding layer:
blocks.20.mlp.l3 - SFT pool: sft256 / LoRA rank 16
- GRPO LoRA: rank 16, alpha 32
- Checkpoint: step 1156
This repo contains the merged text LLM plus dna_projection.pt.
Loading
Same layout as the DNA SFT model — see bio-posttrain-qwen3-1.7b-dna-sft.
Collection
Part of the Bio-posttrain collection.
Description
Languages
Jinja
100%