Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads. `torch_dtype` is deprecated! Use `dtype` instead! ============================================================ DeepSentinel GRPO Training ============================================================ [1] Collecting 120 training episodes... Collected 120 episodes Distribution: 40 easy, 40 medium, 40 hard [2] Evaluating heuristic baseline... easy: avg_reward=-0.3681 medium: avg_reward=0.0049 hard: avg_reward=0.0833 [3] Loading model: Qwen/Qwen2.5-0.5B-Instruct Loading weights: 0%| | 0/290 [00:00