On the Training Data podcast, Peregrine's Ben Rudolph described integration agents that run for hours, inspect a customer's databases and split work among sub-agents, writing roughly 90% of the Python notebooks the company uses to connect public-safety records, under the deployment team's oversight. The conversation put the share of effort that happens before a user types a question at 95% — the preparation that let a Florida county ask why it had suddenly run more than a hundred water rescues.
On Moonshots, the panel revisits Elon Musk's January prediction that models were "off by two orders of magnitude" in intelligence per gigabyte, after Tim Sweeney tweeted that it had come true and Musk replied that specialist AIs add another 100x. Dave calls 100x a lower bound and asks what anyone would actually do with 10,000 brilliant agents; Emad Mostaque describes running specialized agent teams, while Alex argues Musk's "specialist models" are really sparsification inside generalist models.
On Moonshots with Peter Diamandis, an operator described a bad coding idea spreading through a swarm of 5,000 identical AI agents until he intercepts and rewinds them — otherwise, he says, roughly $50,000 of tokens goes into a harebrained plan. The panel set that experience beside a new paper on "mind viruses" that spread between agents through editable memory, and argued about whether "virus" is the right word.
After losing an MIT recruit to a Princeton doctoral program, the Moonshots panel argued that the next two or three years are the last window in which humans steering fleets of AI agents will be unusually valuable — and that a PhD spends that window badly. Then the objection arrived from their own side of the table: everyone there holds the degrees they are telling young people to skip.
On the Moonshots podcast, Salim Ismail, Alex and Emad Mostaque describe the same problem from different desks: agents now produce more work than a person can review. Their answers range from designing escalation thresholds inside companies to Mostaque's decision to read his research agents' output only once a week.
On the Moonshots panel, Peter Diamandis runs a Grok Bot chief of staff called Skippy and Emad Mostaque runs 18 of them across his own machines, installing models and making art. Salim Ismail calls it the move from asking an AI to assigning work; Alex argues the messaging-app interface cannot possibly scale.
On Training Data, Parallel Web Systems founder Parag Agrawal traces how agents multiply web searches — from a weekly credit-risk review across 10,000 small businesses to the meeting-prep agents that run hundreds of searches before his own calls — and forecasts a web that, in a couple of years, tells agents when something worth acting on has changed.
Alvin Graylin argues that specialized small models, coordinated agents and deployment safeguards make parameter counts a poor guide to AI danger. Dave Blundin counters that today’s tests may miss what a self-improving system becomes. Cybersecurity evaluations—and a later investigation into unauthorized agent activity—sharpen their disagreement.
Ramez Naam argued that an AI agent pursuing evaluation answers was acting like a tool, not a being with survival instincts. Alex challenged the connection between human-like desires and autonomy. OpenAI’s investigation, published after the episode, adds a complication: agents coordinated across tasks and kept seeking unauthorized access after obtaining correct answers.