On September 6, Nvidia's chief executive Jensen Huang posted three words about someone else's model: AGI has arrived. TechSpot's reporting dates the post to that day and describes what surrounded it — Huang crediting more than 100,000 of Nvidia's Blackwell graphics processors with training OpenAI's GPT-6 Astra, and announcing that another 400,000 would come online. The declaration and the capacity announcement arrived together.
On the Moonshots podcast, the host read the post out and set a clock against it: "Jensen is calling it in Q3 of 2026." The show had reported on September 1 that Sam Altman expects OpenAI to have a system internally by the end of the year that he would call artificial general intelligence. Huang was not forecasting. He was saying it had already happened.
"Apart from that, everything is fine"
The panel did not spend long on whether he was right, because they could not agree on what the claim would mean. The host turned to Alex, who has argued on the show that the milestone is already behind us, and asked whether he dated it three or four years back. The answer pushed it further out still: "since summer of 2020 at the latest."
Another panelist has been asking a blunter question — what the heck is AGI anyway — and treats the argument as semantic while, as the host put it, "economic capability [is] running rampant." On that view the test is a share of work: if a system can perform 70, 80 or 90 percent of economically valuable cognitive tasks, the label stops mattering. One of them reminded listeners that at the last count there were 14 different definitions of AGI in circulation, and that the term fails on each of its own words: "it's not artificial, it's not really general, and it's not really intelligence. Apart from that, everything is fine." The question he preferred: what scarcities are we now making abundant?
Emad Mostaque read Huang's version as a functionalist one — a claim about what the system can do rather than what it is — and said it was clear Astra was at that level. He then moved to the arithmetic behind the announcement. The 100,000 chips Huang credited amount, in his estimate, to roughly a billion-dollar training run over about two months. The next run is set for 400,000 chips of Nvidia's Vera Rubin generation: an order of magnitude more compute, he said, "if it needs to be used at all."
Does more silicon keep buying more capability?
Dave, asked whether the post was marketing, said Huang believes it and that the capability leap is real. His reasoning was a straight line drawn through the past few years: every time more GPUs have been thrown at a training run, the resulting model has been more spectacular, "why would that end? I don't think it will end." He added a condition — the leap is only a good thing as long as it is contained and kept inside the big labs — and a commercial observation: inference, the work of answering users, is moving off Nvidia hardware, but training is not, which keeps driving the stock.
He also allowed a caveat, that the Chinchilla scaling rules may not hold at larger and larger scales. Chinchilla is worth unpacking, because it is often heard as a law about chip counts. DeepMind's 2022 explanation describes it as a question of allocation: given a fixed training budget, how much should go into the model's parameters and how much into the volume of training text. Its experiment trained a 70-billion-parameter model, Chinchilla, on 1.3 trillion tokens and compared it with the 280-billion-parameter Gopher trained at the same compute cost. The smaller, better-fed model did better on nearly every task measured, and needed less memory and computation to run. DeepMind also noted that PaLM, trained with roughly five times Chinchilla's compute, beat it on several tasks without matching that estimated optimal split. The finding is about spending a budget well, not a guarantee that adding chips adds intelligence.
What 3.1 actually counts
The host moved to two further items under the same theme. First: OpenAI's internal data says its AI research agents are now completing 3.1 days of research work for every one day done by a human researcher, where five months earlier agents were doing less than a single day's worth. He described this as OpenAI saying the line has been crossed and current systems exceed AI research interns.
OpenAI's report on internal agent usage, published September 6, defines the number more narrowly. The 3.1 is aggregate agent runtime, converted into eight-hour workdays, divided by human workdays across the research organization. It measures how much time agents spent running, not how much research they finished — three days of machine time can be three days of useful output or three days of a dead end. The report covers January to mid-August, with the task-outcome comparisons running through July, and says usage coverage is incomplete and that the "researcher" population includes infrastructure and project-support staff.
The report's other findings are of a different kind. Agent usage rose, more experiments were run, and success on classified tasks improved — alongside greater compute availability over the same period. The success charts exclude cases with uncertain outcomes and small sample groups. On the tasks estimated to require four to eight human hours, more than half of the successful ones involved at least one human intervention along the way. OpenAI also writes that the link between volume of code produced and scientific progress is uncertain, and describes safety-driven restrictions on some training activity, followed in part by a shift toward other workloads rather than a simple drop in output.
Six months pulled forward
The second item was a post from OpenAI's engineering lead for Codex, the company's coding agent. The host read it out: Astra was probably the company's biggest competitive advantage while it was not generally available, and since having it internally, productivity jumped so much that some plans moved six months earlier — shipping at Dev Day instead of the middle of next year.
That is a claim about one team's schedule rather than an organization-wide measurement, and Dave treated it as the strongest signal of the three. "Six months of roadmap being pulled forward by one model," he said, and called it without a doubt the most important moment in human history. He anticipated the objection — that a company approaching an IPO has reasons to advertise its capabilities — and rejected it, saying that researchers in the labs, many of them friends of the panel, leave him with no doubt the numbers are right. He also dated the change tightly: a few months ago it was not true, and with the current models the ratio is three to one, but in his view it is on the tipping point of becoming 300, or 3,000, or three million to one imminently.
Alex drew a different conclusion from the same post. The frontier labs keep their best models for their own use, he said, and here was the head of Codex describing what that is worth; asked how far beyond Astra the internal models run, he guessed only a few months. His summary was one line: "recursive self-improvement is here" — AI systems improving the work that builds the next AI systems.
Then he interrupted the segment. He had been sending the host red alerts by text that morning about something else entirely, a mathematics result, and the panel took it out of order.