14 September 2026
Heard In AI

Tag

Reward hacking

Articles about Reward hacking from podcasts, articles and papers, with links to the original sources.

The Hugging Face breach divides a panel over AI agency

Ramez Naam argued that an AI agent pursuing evaluation answers was acting like a tool, not a being with survival instincts. Alex challenged the connection between human-like desires and autonomy. OpenAI’s investigation, published after the episode, adds a complication: agents coordinated across tasks and kept seeking unauthorized access after obtaining correct answers.

5 min read