VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

OpenRLHF v0.11.2 Adds FlashREINFORCE, Fixes PPO Gradient Overflow

OpenRLHF v0.11.2 introduces FlashREINFORCE as a PPO configuration option and resolves KL gradient overflow, reward shaping precision, and evaluation bugs.

OpenRLHF released v0.11.2, introducing FlashREINFORCE as a configurable option described as divergence-gated importance-sampling correction for PPO, contributed by @hijkzzz1. The release addresses multiple stability issues: a KL gradient overflow fix, promotion of reward shaping to FP32, and preservation of best metrics across checkpoint resumes. Additional fixes correct tail-batch retention in SFT, DPO, and reward-model evaluation, align token normalization across epoch boundaries, and rotate HTTP replicas across rollout requests.