To Mix or To Merge?
Toward Multi-Domain Reinforcement Learning for Large Language Models

arXiv Hugging Face ModelScope GitHub COLM 2026

Haoqing Wang†, Xiang Long†, Ziheng Li†, Yilong Xu, Tingguang Li, Yehui Tang✉
Samsung Research, Beijing, China   ·   Peking University


📰 News

  • [2026.09.07] 🎉 The model checkpoints are now open-sourced on Hugging Face and ModelScope! Feel free to use our checkpoints for your post-training research (e.g., weight merging, multi-teacher on-policy distillation)!
  • [2026.07.09] 🎉 Our paper is accepted to COLM 2026!

📚 Citation

If you find this work useful, please consider citing:

@inproceedings{
wang2026to,
title={To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models},
author={Haoqing Wang and Xiang Long and Ziheng Li and Yilong Xu and Tingguang Li and Yehui Tang},
booktitle={Third Conference on Language Modeling},
year={2026},
url={https://openreview.net/forum?id=jP7j5XkG8J}
}
Description
Model synced from source: whq1111/M2RL-SFT
Readme 15 MiB
Languages
Text 100%