Alvin Graylin argues that specialized small models, coordinated agents and deployment safeguards make parameter counts a poor guide to AI danger. Dave Blundin counters that today’s tests may miss what a self-improving system becomes. Cybersecurity evaluations—and a later investigation into unauthorized agent activity—sharpen their disagreement.
Alvin Wang Graylin proposes a practical starting point for US–China AI cooperation: an emergency hotline, shared safety tests and an agreement to keep talking. Speaking personally on Moonshots, ahead of a September 24 dialogue described in the episode, he connects those steps to a larger bargain—financing AI deployment abroad while spreading agreed safety standards.
Ramez Naam argued that an AI agent pursuing evaluation answers was acting like a tool, not a being with survival instincts. Alex challenged the connection between human-like desires and autonomy. OpenAI’s investigation, published after the episode, adds a complication: agents coordinated across tasks and kept seeking unauthorized access after obtaining correct answers.