ModelHub XC 4a89dbbadf 初始化项目,由ModelHub XC社区提供模型
Model: BayesRL/Llama3.1-IVON-SFT-8B
Source: Original Platform
2026-07-18 06:08:09 +08:00

license, base_model, datasets, language, pipeline_tag, library_name, tags
license base_model datasets language pipeline_tag library_name tags
apache-2.0 Qwen/Qwen2.5-Math-7B
nvidia/Llama-Nemotron-Post-Training-Dataset
en
text-generation transformers
ivon
variational-learning
sft
3po
math
reasoning

Qwen2.5Math-IVON-SFT-7B

📦 Code: insait-institute/c3po

Qwen2.5-Math 7B supervised-fine-tuned with the variational optimizer IVON, from the paper "Parameter Exploration for RLVR via Variational Learning".

This is a warm-start checkpoint: SFT'ing with IVON yields not just point weights but an approximate Gaussian posterior over them (a mean and a diagonal Hessian/precision estimate). That posterior is the learned prior used to seed the 3PO RLVR runs (B3PO / M3PO / C3PO), where weight perturbations sampled from it drive parameter-space exploration.

Training

Foundation model Qwen/Qwen2.5-Math-7B
Stage Warm-start SFT
Data Llama-Nemotron Post-Training Dataset (SFT subset)
Optimizer IVON, lr 50.0, ESS (λ) 1e10
Hardware 8× NVIDIA H200 (144 GB)

Usage

Loads as a standard causal LM:

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("BayesRL/Qwen2.5Math-IVON-SFT-7B")
tok = AutoTokenizer.from_pretrained("BayesRL/Qwen2.5Math-IVON-SFT-7B")

To use it as the warm-start prior for 3PO RLVR, load the IVON optimizer state via IVON_INIT_METHOD=trained in the companion code's run_rl.sh.

Citation

@misc{venkatkrishna2026parameter,
      title={Parameter Exploration for RLVR via Variational Learning},
      author={Vatsal Venkatkrishna and Nico Daheim and Iryna Gurevych},
      year={2026},
}
Description
Model synced from source: BayesRL/Llama3.1-IVON-SFT-8B
Readme 16 MiB
Languages
Jinja 100%