--- license: apache-2.0 base_model: Qwen/Qwen3-1.7B tags: - biology - bio-posttrain - dna-rl - dna library_name: transformers --- # Bio-posttrain Qwen3-1.7B DNA RL DNA reinforcement learning (GRPO) checkpoint from [How Post-Training Shapes Biological Reasoning Models](https://huggingface.co/collections/mims-harvard/bio-posttrain). ## Model details - **Base model:** `Qwen/Qwen3-1.7B` - **DNA encoder:** Evo2 `evo2_1b_base` (frozen; not included) - **Embedding layer:** `blocks.20.mlp.l3` - **SFT pool:** sft256 / LoRA rank 16 - **GRPO LoRA:** rank 16, alpha 32 - **Checkpoint:** step 1156 This repo contains the merged text LLM plus `dna_projection.pt`. ## Loading Same layout as the DNA SFT model — see [bio-posttrain-qwen3-1.7b-dna-sft](https://huggingface.co/mims-harvard/bio-posttrain-qwen3-1.7b-dna-sft). ## Collection Part of the [Bio-posttrain](https://huggingface.co/collections/mims-harvard/bio-posttrain) collection.