https://openreview.net/forum?id=g33DGvnHYd
SSVPO: Effective Step-Level Credit Assignment for RL Training of Language Models | OpenReview
Language models have shown strong performance on mathematical reasoning tasks. Post-training with outcome-based reinforcement learning (RL) can further enhance...