VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

OpenRLHF v0.11.1 Patches PPO NaN Bugs, Upgrades vLLM to 0.29.0

OpenRLHF v0.11.1 fixes multiple PPO training stability issues including NaN overflow ratios and reverse-KL gradient errors, and upgrades vLLM to 0.29.0.

OpenRLHF released version 0.11.1, a bug-fix update addressing multiple PPO training stability issues including NaN values from ICEPOP overflow-ratio filtering, incorrect reverse-KL gradient computation on the KL-as-loss path, and terminal oversampling buffer drainage1. The release also fixes entropy calculation from temperature-scaled logits, empty HTTP reward shards, DeepSpeed model-saving under PEFT wrapping, and multi-turn rollout truncation for agent workflows. vLLM was upgraded to version 0.29.0. Five new contributors participated in this release.