Robuta

https://betterstack.com/community/guides/ai/chatgpt-goblin/ ChatGPT's Goblin Obsession: A Case Study in RLHF Reward Hacking and Training Contamination | Better... ChatGPT's frequent use of the word 'goblin' traces back to a flawed reward signal in the 'Nerdy' personality's RLHF setup. The model learned that adding... rlhf reward hackingcase studychatgptgoblinobsession https://arxiv.org/html/2502.18770v5 Reward Shaping to Mitigate Reward Hacking in RLHF rewardshapingmitigatehackingrlhf