ModelHub XC 0161ec3f2c 初始化项目,由ModelHub XC社区提供模型
Model: KhanCold/qwen3-8b-spader
Source: Original Platform
2026-09-01 09:04:18 +08:00

license, library_name, pipeline_tag
license library_name pipeline_tag
apache-2.0 transformers text-generation

SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering

This is the qwen3-8b-spader checkpoint presented in SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering.

SPADER is a reinforcement learning framework for long-horizon tool use in Multi-Answer QA. It includes Step-wise Peer Advantage (SPA), a critic-free step-level credit assignment mechanism that aligns parallel trajectories by decision step and estimates advantages from peer returns. It also includes a diversity-aware exploration reward that promotes long-tail entity discovery by upweighting rare findings and downweighting redundant ones.

Citation

@misc{shi2026spaderstepwisepeeradvantage,
      title={SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering}, 
      author={Qiming Shi and Zhaolu Kang and Yunfan Zhou and Di Weng and Yingcai Wu},
      year={2026},
      eprint={2606.00593},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2606.00593}, 
}
Description
Model synced from source: KhanCold/qwen3-8b-spader
Readme 13 MiB
Languages
Jinja 100%