Robuta

https://www.watertightai.com/ Watertight AI - Reward Hacking Monitors for RL Training Automated cheating detection for AI reinforcement learning training. Building monitors that detect reward hacking to make RL training more effective. reward hackingwatertightaimonitorsrl https://watertightai.com/ Watertight AI - Reward Hacking Monitors for RL Training Automated cheating detection for AI reinforcement learning training. Building monitors that detect reward hacking to make RL training more effective. reward hackingwatertightaimonitorsrl https://www.anthropic.com/research/emergent-misalignment-reward-hacking?ref=labs.lares.com From shortcuts to sabotage: natural emergent misalignment from reward hacking \ Anthropic Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems. reward hackingshortcutssabotagenaturalemergent https://mlpuzzles.com/ Kernel Reward Hacking Challenge | METR reward hackingkernelchallengemetr https://arxiv.org/html/2502.18770v5 Reward Shaping to Mitigate Reward Hacking in RLHF rewardshapingmitigatehackingrlhf