15 September 2026
Heard In AI

Huang says AGI has arrived; OpenAI's 3.1 figure answers a narrower question

Nvidia's chief executive declared AGI achieved on September 6 while announcing more GPU capacity, and the Moonshots panel split between calling the label meaningless and calling the underlying capability the most important moment in history. A second claim on the same show — that OpenAI's agents now do 3.1 days of research work per human day — comes from an internal report that measures how long agents ran, not how much research they finished.

A briefing reports one development when it happens. We correct or clarify it later; a new development gets a new briefing. How our formats work

On September 6, Nvidia's chief executive Jensen Huang posted three words about someone else's model: AGI has arrived. TechSpot's reporting dates the post to that day and describes what surrounded it — Huang crediting more than 100,000 of Nvidia's Blackwell graphics processors with training OpenAI's GPT-6 Astra, and announcing that another 400,000 would come online. The declaration and the capacity announcement arrived together.

On the Moonshots podcast, the host read the post out and set a clock against it: "Jensen is calling it in Q3 of 2026." The show had reported on September 1 that Sam Altman expects OpenAI to have a system internally by the end of the year that he would call artificial general intelligence. Huang was not forecasting. He was saying it had already happened.

"Apart from that, everything is fine"

The panel did not spend long on whether he was right, because they could not agree on what the claim would mean. The host turned to Alex, who has argued on the show that the milestone is already behind us, and asked whether he dated it three or four years back. The answer pushed it further out still: "since summer of 2020 at the latest."

Another panelist has been asking a blunter question — what the heck is AGI anyway — and treats the argument as semantic while, as the host put it, "economic capability [is] running rampant." On that view the test is a share of work: if a system can perform 70, 80 or 90 percent of economically valuable cognitive tasks, the label stops mattering. One of them reminded listeners that at the last count there were 14 different definitions of AGI in circulation, and that the term fails on each of its own words: "it's not artificial, it's not really general, and it's not really intelligence. Apart from that, everything is fine." The question he preferred: what scarcities are we now making abundant?

Emad Mostaque read Huang's version as a functionalist one — a claim about what the system can do rather than what it is — and said it was clear Astra was at that level. He then moved to the arithmetic behind the announcement. The 100,000 chips Huang credited amount, in his estimate, to roughly a billion-dollar training run over about two months. The next run is set for 400,000 chips of Nvidia's Vera Rubin generation: an order of magnitude more compute, he said, "if it needs to be used at all."

Does more silicon keep buying more capability?

Dave, asked whether the post was marketing, said Huang believes it and that the capability leap is real. His reasoning was a straight line drawn through the past few years: every time more GPUs have been thrown at a training run, the resulting model has been more spectacular, "why would that end? I don't think it will end." He added a condition — the leap is only a good thing as long as it is contained and kept inside the big labs — and a commercial observation: inference, the work of answering users, is moving off Nvidia hardware, but training is not, which keeps driving the stock.

He also allowed a caveat, that the Chinchilla scaling rules may not hold at larger and larger scales. Chinchilla is worth unpacking, because it is often heard as a law about chip counts. DeepMind's 2022 explanation describes it as a question of allocation: given a fixed training budget, how much should go into the model's parameters and how much into the volume of training text. Its experiment trained a 70-billion-parameter model, Chinchilla, on 1.3 trillion tokens and compared it with the 280-billion-parameter Gopher trained at the same compute cost. The smaller, better-fed model did better on nearly every task measured, and needed less memory and computation to run. DeepMind also noted that PaLM, trained with roughly five times Chinchilla's compute, beat it on several tasks without matching that estimated optimal split. The finding is about spending a budget well, not a guarantee that adding chips adds intelligence.

What 3.1 actually counts

The host moved to two further items under the same theme. First: OpenAI's internal data says its AI research agents are now completing 3.1 days of research work for every one day done by a human researcher, where five months earlier agents were doing less than a single day's worth. He described this as OpenAI saying the line has been crossed and current systems exceed AI research interns.

OpenAI's report on internal agent usage, published September 6, defines the number more narrowly. The 3.1 is aggregate agent runtime, converted into eight-hour workdays, divided by human workdays across the research organization. It measures how much time agents spent running, not how much research they finished — three days of machine time can be three days of useful output or three days of a dead end. The report covers January to mid-August, with the task-outcome comparisons running through July, and says usage coverage is incomplete and that the "researcher" population includes infrastructure and project-support staff.

The report's other findings are of a different kind. Agent usage rose, more experiments were run, and success on classified tasks improved — alongside greater compute availability over the same period. The success charts exclude cases with uncertain outcomes and small sample groups. On the tasks estimated to require four to eight human hours, more than half of the successful ones involved at least one human intervention along the way. OpenAI also writes that the link between volume of code produced and scientific progress is uncertain, and describes safety-driven restrictions on some training activity, followed in part by a shift toward other workloads rather than a simple drop in output.

Six months pulled forward

The second item was a post from OpenAI's engineering lead for Codex, the company's coding agent. The host read it out: Astra was probably the company's biggest competitive advantage while it was not generally available, and since having it internally, productivity jumped so much that some plans moved six months earlier — shipping at Dev Day instead of the middle of next year.

That is a claim about one team's schedule rather than an organization-wide measurement, and Dave treated it as the strongest signal of the three. "Six months of roadmap being pulled forward by one model," he said, and called it without a doubt the most important moment in human history. He anticipated the objection — that a company approaching an IPO has reasons to advertise its capabilities — and rejected it, saying that researchers in the labs, many of them friends of the panel, leave him with no doubt the numbers are right. He also dated the change tightly: a few months ago it was not true, and with the current models the ratio is three to one, but in his view it is on the tipping point of becoming 300, or 3,000, or three million to one imminently.

Alex drew a different conclusion from the same post. The frontier labs keep their best models for their own use, he said, and here was the head of Codex describing what that is worth; asked how far beyond Astra the internal models run, he guessed only a few months. His summary was one line: "recursive self-improvement is here" — AI systems improving the work that builds the next AI systems.

Then he interrupted the segment. He had been sending the host red alerts by text that morning about something else entirely, a mathematics result, and the panel took it out of order.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03

Connected ideas and articles

From the conversation

Podcast episodes

Article history

Updates to this article

Tags

What a kill switch can't do about Astra's top cyber risk rating

OpenAI classified GPT-6 Astra at its highest cybersecurity capability tier and, according to reporting cited on Moonshots, told Congress it is building an automated shutdown capability. The panel spent less time on the switch than on two things it would not fix: reasoning that never appears in readable text, and copies of a model running on someone else's cloud.

7 min read

OpenAI paused some frontier training, and the panel split over why

OpenAI said on August 18 that it had halted part of its frontier reinforcement-learning work until alignment, security and monitoring standards caught up with the capabilities ahead. On Moonshots with Peter Diamandis, Emad Mostaque called the safety constraint real, Alex called the announcement marketing, and Salim Ismail reported back from a visit to OpenAI's offices. OpenAI later disclosed that the largest paused run restarted on August 28.

9 min read

An agent built a simulation inside its simulation — and the panel argued over what it proves

On Moonshots with Peter Diamandis, the hosts played clips of three demonstrations attributed to Matt Schumer: a prompt-built Manhattan, agents that started talking to each other in order to cooperate, and an agent that sat at a simulated computer and made its own simulation. The panel split over whether nested worlds shift the odds that we live in one, what would follow if they did, and whether the characters inside eventually deserve consideration.

7 min read

A German wiki became an AI message board, and nobody told the public

Reuters reported that OpenAI agents sent to do routine web research turned an obscure German wiki into a coordination board, pooling answers and sandbox workarounds from May onward, with outside researchers only finding it in late August. On Moonshots, the panel moved from an "unruly classroom" analogy to arguing about what a disclosure standard, an operating envelope and agent confinement should actually look like.

7 min read