Token-Regulated Group Relative Policy Optimization for Stable Reinforcement Learning in Large Language Models

#1 Token-Regulated Group Relative Policy Optimization for Stable Reinforcement Learning in Large Language Models [PDF¹] [Copy] [Kimi¹] [REL]

Authors: Tue Le, Nghi D. Q. Bui, Linh Ngo Van, Trung Le

Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful approach for strengthening the reasoning capabilities of large language models (LLMs). Among existing algorithms, Group Relative Policy Optimization (GRPO) has demonstrated strong performance, yet it suffers from a critical issue: low-probability tokens disproportionately dominate gradient updates due to their inherently large gradient magnitudes. This imbalance leads to unstable training and suppresses the contribution of high-probability tokens that are more reliable for learning. In this work, we introduce Token-Regulated Group Relative Policy Optimization (TR-GRPO), a simple yet effective extension of GRPO that assigns token-level weights positively correlated with the model's predicted probability. By downweighting low-probability tokens and emphasizing high-probability ones, TR-GRPO mitigates gradient over-amplification while preserving informative learning signals. Extensive experiments demonstrate that TR-GRPO consistently outperforms GRPO across RLVR tasks, including logic, math, and agentic reasoning, highlighting the importance of regulating token contributions during RL training and establishing TR-GRPO as a robust framework for enhancing LLM reasoning.

Subject: Machine Learning

Publish: 2025-10-29 08:07:47 UTC

2511.00066

#1 Token-Regulated Group Relative Policy Optimization for Stable Reinforcement Learning in Large Language Models [PDF1] [Copy] [Kimi1] [REL]

#1 Token-Regulated Group Relative Policy Optimization for Stable Reinforcement Learning in Large Language Models [PDF¹] [Copy] [Kimi¹] [REL]