On The Diary of a CEO, writer Ed Zitron described catching an invented Microsoft share price in his Bloomberg terminal, then argued that his editor Matt Hughes — not a benchmark number — is what makes an answer trustworthy. The host pushed back: buyers pay for the output, not the process, and the honest comparison is AI against fallible people rather than perfection.
Asked about allegations that Chinese labs extracted capabilities from Claude, Alvin Graylin argued that access to another model’s answers cannot explain every engineering advance. The Moonshots exchange turned on three distinctions: legitimate distillation versus prohibited extraction, query bills versus development costs, and learning from outputs versus improving the machinery behind them.
Rich Sutton and Khurram Javed want deployed AI to change its underlying weights from individual experience, rather than rely on extra context or shared model updates. Their Oak Lab agenda combines learning rates tailored to each weight with a way to refresh a network’s capacity to learn—supported by earlier experiments, but not yet a demonstrated general-purpose system.
On The Diary of a CEO, physicist Brian Greene debated an AI assistant about whether smarter systems must keep producing ever-faster gains. A cup on the table helped explain his doubts about today's architectures—but he also warned about shutdown resistance and improvements outpacing human scrutiny if rapid growth does occur.
On Moonshots, Ramez Naam pointed to the brain’s modest power needs and children’s ability to learn from relatively little data. Co-host Alex countered with a rack of chips producing text thousands of times faster than one writer. Their disagreement connects AI’s energy bill to a larger question: how much improvement can more computation buy?
BDH-CQ’s authors report solving 118 of 400 public ARC-AGI-1 tasks at an estimated inference cost of $0.00070 per task, with up to two candidate answers. On Moonshots, Emad Mostaque welcomed architectural experimentation; panelist Alex questioned whether this design offered progress beyond a specialized benchmark.
Meta’s Muse Glimmer is a 30-billion-parameter model designed to run agents on personal computers. Alongside Mark Zuckerberg’s vision of personal superintelligence, it prompted a Moonshots debate about whether open models put users in charge—or strengthen the company that already owns their favorite apps.
xAI’s August 12 release puts Grok 4.6 alongside GPT-5.6 Sol Max in its launch benchmark table, with pricing aimed at sustained agent work. The Moonshots panel’s debate was about the next step: whether training on other models’ reasoning can only help a challenger catch up, and what computing infrastructure it takes to move beyond that.