https://betterstack.com/community/guides/ai/chatgpt-goblin/
ChatGPT's Goblin Obsession: A Case Study in RLHF Reward Hacking and Training Contamination | Better...
ChatGPT's frequent use of the word 'goblin' traces back to a flawed reward signal in the 'Nerdy' personality's RLHF setup. The model learned that adding...
rlhf reward hackingcase studychatgptgoblinobsession
https://arxiv.org/html/2502.18770v5
Reward Shaping to Mitigate Reward Hacking in RLHF
rewardshapingmitigatehackingrlhf