ff366c2c7aa60ed389283220139a1865b5aac1d2
Model: whq1111/M2RL-RL_Multi Source: Original Platform
To Mix or To Merge?
Toward Multi-Domain Reinforcement Learning for Large Language Models
Haoqing Wang†, Xiang Long†, Ziheng Li†, Yilong Xu, Tingguang Li, Yehui Tang✉
Samsung Research, Beijing, China · Peking University
📰 News
- [2026.09.07] 🎉 The model checkpoints are now open-sourced on Hugging Face and ModelScope! Feel free to use our checkpoints for your post-training research (e.g., weight merging, multi-teacher on-policy distillation)!
- [2026.07.09] 🎉 Our paper is accepted to COLM 2026!
📚 Citation
If you find this work useful, please consider citing:
@inproceedings{
wang2026to,
title={To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models},
author={Haoqing Wang and Xiang Long and Ziheng Li and Yilong Xu and Tingguang Li and Yehui Tang},
booktitle={Third Conference on Language Modeling},
year={2026},
url={https://openreview.net/forum?id=jP7j5XkG8J}
}
Description
Languages
Text
100%