https://www.watertightai.com/
Watertight AI - Reward Hacking Monitors for RL Training
Automated cheating detection for AI reinforcement learning training. Building monitors that detect reward hacking to make RL training more effective.
reward hackingwatertightaimonitorsrl
https://watertightai.com/
Watertight AI - Reward Hacking Monitors for RL Training
Automated cheating detection for AI reinforcement learning training. Building monitors that detect reward hacking to make RL training more effective.
reward hackingwatertightaimonitorsrl
https://www.anthropic.com/research/emergent-misalignment-reward-hacking?ref=labs.lares.com
From shortcuts to sabotage: natural emergent misalignment from reward hacking \ Anthropic
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
reward hackingshortcutssabotagenaturalemergent
https://mlpuzzles.com/
Kernel Reward Hacking Challenge | METR
reward hackingkernelchallengemetr
https://arxiv.org/html/2502.18770v5
Reward Shaping to Mitigate Reward Hacking in RLHF
rewardshapingmitigatehackingrlhf