Files
ReLIFT-Qwen2.5-7B-Zero/README.md
ModelHub XC 9632bee7a4 初始化项目,由ModelHub XC社区提供模型
Model: RoadQAQ/ReLIFT-Qwen2.5-7B-Zero
Source: Original Platform
2026-09-22 16:02:25 +08:00

11 lines
478 B
Markdown

---
library_name: transformers
license: mit
pipeline_tag: question-answering
---
ReLIFT, a training method that interleaves RL with online FT, achieving superior performance and efficiency compared to using RL or SFT alone, as described in [Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions](https://huggingface.co/papers/2506.07527).
Code: https://github.com/TheRoadQaQ/ReLIFT
Project page: https://github.com/TheRoadQaQ/ReLIFT