Files
M2RL-MT_OPD/README.md
ModelHub XC e6ca6f9310 初始化项目,由ModelHub XC社区提供模型
Model: whq1111/M2RL-MT_OPD
Source: Original Platform
2026-09-17 18:22:17 +08:00

1.9 KiB

To Mix or To Merge?
Toward Multi-Domain Reinforcement Learning for Large Language Models

arXiv Hugging Face ModelScope GitHub COLM 2026

Haoqing Wang†, Xiang Long†, Ziheng Li†, Yilong Xu, Tingguang Li, Yehui Tang✉
Samsung Research, Beijing, China   ·   Peking University


📰 News

  • [2026.09.07] 🎉 The model checkpoints are now open-sourced on Hugging Face and ModelScope! Feel free to use our checkpoints for your post-training research (e.g., weight merging, multi-teacher on-policy distillation)!
  • [2026.07.09] 🎉 Our paper is accepted to COLM 2026!

📚 Citation

If you find this work useful, please consider citing:

@inproceedings{
wang2026to,
title={To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models},
author={Haoqing Wang and Xiang Long and Ziheng Li and Yilong Xu and Tingguang Li and Yehui Tang},
booktitle={Third Conference on Language Modeling},
year={2026},
url={https://openreview.net/forum?id=jP7j5XkG8J}
}