初始化项目,由ModelHub XC社区提供模型
Model: mims-harvard/ATHENA-R1-Qwen3-8B Source: Original Platform
This commit is contained in:
137
README.md
Normal file
137
README.md
Normal file
@@ -0,0 +1,137 @@
|
||||
---
|
||||
license: mit
|
||||
language:
|
||||
- en
|
||||
base_model: Qwen/Qwen3-8B
|
||||
library_name: transformers
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- agent
|
||||
- biomedical
|
||||
- treatment-reasoning
|
||||
- tool-use
|
||||
- reinforcement-learning
|
||||
- qwen3
|
||||
---
|
||||
|
||||
# ATHENA-R1-Qwen3-8B
|
||||
|
||||
**Project page:** [athena.openscientist.ai](https://athena.openscientist.ai/) · **Code:** [mims-harvard/ATHENA](https://github.com/mims-harvard/ATHENA)
|
||||
|
||||
**ATHENA-R1** is an AI agent for treatment reasoning, trained through
|
||||
reinforcement learning over a universe of 212 biomedical tools. It performs
|
||||
*multi-step* reasoning — identifying what evidence is needed, selecting tools,
|
||||
and incorporating retrieved evidence into subsequent steps — with tool calls
|
||||
served through the
|
||||
[ToolUniverse](https://github.com/mims-harvard/ToolUniverse) (FDA labeling,
|
||||
Open Targets, ChEMBL, EuropePMC, etc.).
|
||||
|
||||
Given a clinical question, the model performs *multi-step* tool calls,
|
||||
synthesises the evidence, and returns a free-form answer grounded in
|
||||
authoritative biomedical sources.
|
||||
|
||||
## Quick start
|
||||
|
||||
The model is exposed through the
|
||||
[`athena-r1`](https://github.com/mims-harvard/ATHENA) Python package,
|
||||
which handles the tool-call protocol and conversation management. Two
|
||||
services back the agent: vLLM (model server) and ToolUniverse (tool server).
|
||||
|
||||
```bash
|
||||
# 1. Install
|
||||
pip install "athena-r1[vllm,web] @ git+https://github.com/mims-harvard/ATHENA.git"
|
||||
|
||||
# 2. Start backing services
|
||||
bash scripts/launch_tooluniverse.sh # → :8080
|
||||
bash scripts/launch_vllm.sh 8000 mims-harvard/ATHENA-R1-Qwen3-8B
|
||||
|
||||
# 3. Run the agent (Python)
|
||||
python -c "
|
||||
from athena_r1 import AthenaR1
|
||||
agent = AthenaR1(
|
||||
model='mims-harvard/ATHENA-R1-Qwen3-8B',
|
||||
vllm_url='http://0.0.0.0:8000/v1',
|
||||
tool_server='http://0.0.0.0:8080',
|
||||
)
|
||||
print(agent.answer('Dose adjustment for metformin in CKD eGFR 35?').answer)
|
||||
"
|
||||
```
|
||||
|
||||
For a chat UI (bundled browser demo with live-streamed reasoning):
|
||||
|
||||
```bash
|
||||
python web/agui_server.py # → http://localhost:8090/ (AG-UI server + demo)
|
||||
```
|
||||
|
||||
For an OpenAI-compatible API endpoint:
|
||||
|
||||
```bash
|
||||
python web/openai_server.py # → http://localhost:9000/v1/chat/completions
|
||||
```
|
||||
|
||||
## Inference settings (paper-canonical)
|
||||
|
||||
| Parameter | Value |
|
||||
|---|---|
|
||||
| temperature | 0.7 |
|
||||
| top_p | 0.95 |
|
||||
| top_k | 20 |
|
||||
| min_p | 0.0 |
|
||||
| presence_penalty | 0 |
|
||||
| max_round | 40 |
|
||||
| concurrent Qs | 4 |
|
||||
|
||||
## Evaluation
|
||||
|
||||
Open-ended setting: each question is answered free-form, then mapped to one of
|
||||
the original answer choices.
|
||||
|
||||
| Benchmark | n | ATHENA-R1 | GPT-5 |
|
||||
|---|---|---|---|
|
||||
| DrugPC (open-ended drug reasoning) | 3,168 | **94.7%** | 76.9% |
|
||||
| TreatmentPC (patient-specific treatment) | 456 | **82.9%** | 72.2% |
|
||||
|
||||
ATHENA-R1 exceeds GPT-5 by 17.8 points on DrugPC and 10.7 on TreatmentPC.
|
||||
|
||||
See the
|
||||
[`docs/eval_results.md`](https://github.com/mims-harvard/ATHENA/blob/main/docs/eval_results.md)
|
||||
file in the code repo for the full benchmark tables and the two-level
|
||||
self-learning ablation.
|
||||
|
||||
## How it works
|
||||
|
||||
1. **Stage 1 — multi-step tool reasoning**: the model emits
|
||||
`<tool_call>...</tool_call>` blocks; the runtime dispatches them through
|
||||
ToolUniverse, appends results to the conversation, and re-prompts. Loop
|
||||
continues until `[FinalAnswer]` or `max_round` is hit.
|
||||
2. **Stage 2 (eval only) — option mapping**: a *separate* function call
|
||||
maps the free-form answer to an MCQ letter. Two backends supported:
|
||||
the local ATHENA-R1 model (self-extraction) or Azure GPT-5
|
||||
(external reader).
|
||||
|
||||
## Intended use
|
||||
|
||||
ATHENA-R1 is a research artifact for treatment-reasoning research and
|
||||
decision support. It is not a medical device and must not be used for direct
|
||||
patient care.
|
||||
|
||||
## Citation
|
||||
|
||||
```bibtex
|
||||
@article{gao2026athena,
|
||||
title = {An AI agent for treatment reasoning over a biomedical tool universe},
|
||||
author = {Gao, Shanghua and ... and Zitnik, Marinka},
|
||||
journal = {arXiv preprint},
|
||||
year = {2026}
|
||||
}
|
||||
```
|
||||
|
||||
## License
|
||||
|
||||
MIT.
|
||||
|
||||
## Acknowledgements
|
||||
|
||||
Evidence retrieval is powered by
|
||||
[ToolUniverse](https://github.com/mims-harvard/ToolUniverse), a library of
|
||||
curated biomedical tools.
|
||||
Reference in New Issue
Block a user