Two people had texted the host about Grok 4.6 before the recording. Introducing the release on Moonshots, he called it “a banger.” Alex, invited to comment first, offered applause—with a question about how far the recipe could go.
The released model was Grok 4.6, announced on August 12, 2026. Grok 4.7 reaching first place, as the episode’s title suggests, was still a prediction in the conversation. The panel described its arrival within two weeks as a rumor.
The Grok 4.6 announcement emphasizes long-running agent work. An AI agent does more than answer a question: it carries out a task through successive steps, using tools along the way. Here the promise is a system that can research, code, analyze or turn a broad product idea into a working first version. The company reports more self-testing and verification during extended tasks.
Grok 4.6 launched through Cursor, the AI coding environment, Grok Build and an API through which developers can connect it to their own software. Pricing starts at $2 per million input tokens and $6 per million output tokens. Tokens are the small pieces of text a model processes; input is what it receives, and output is what it generates. A faster variant costs twice as much.
What the number 61 means
The company’s launch table gives Grok 4.6 High a score of 61 on the Artificial Analysis Intelligence Index, alongside GPT-5.6 Sol Max at 61 and Fable 5 Max at 62. That places the named Grok configuration close to the leaders in that comparison, not alone at the top.
The index combines multiple evaluations into a weighted score. Its tasks and grading can change, so it is not a fixed exam with an unchanging scale. The launch announcement describes a nine-benchmark composite; Artificial Analysis’s v4.3 methodology describes ten evaluations across agents, coding, scientific reasoning and general capabilities. Those descriptions should not be treated as interchangeable without establishing which version produced the launch scores.
The v4.3 tasks include producing workplace deliverables, completing software workflows, writing scientific code and answering questions grounded in long documents. Depending on the task, grading uses executable tests, completion checks, answer comparisons or rubrics. Most evaluations measure first-attempt success, averaged over repeats where applicable. The suite is primarily English-language and text-based; visual, speech and multilingual abilities are measured separately.
So the launch comparison concerns particular model configurations on a particular collection of tests—not every capability a user might need.
The reasoning-trace argument
Alex opened his account of the training strategy with a caveat: “I lack insider information.” His reading was that Grok 4.6 was “essentially the next version of cursor.”
A reasoning trace is the intermediate working a model generates while solving a problem, rather than just its final answer. Alex argued that developers’ interactions with models through Cursor could supply valuable examples of that working, including traces from Claude and competing systems.
In his account, xAI’s acquisition of Cursor—which he understood was still being completed—and its licensing of Cursor’s trace data offered a shortcut toward leading-model performance. He compared it with allegations that Chinese labs were using Western models’ reasoning traces to improve their own systems. Both the Cursor-specific account and the comparison were his outside interpretation, not an independently established account of the training data.
Post-training means refining a model after its initial broad training, using examples and feedback to improve how it performs tasks. Teaching one model from another model’s outputs is a form of distillation. Alex’s argument was that worked examples from strong models can help a challenger reproduce skills it previously lacked, but do not by themselves supply a route beyond the systems that generated them.
“They won’t get you past the frontier,” he said, calling the approach a “one-trick pony” for nearly catching up. “But it’s a heck of a one-trick pony.”
The company’s documented training account is more specific about methods than data provenance. It describes supplemental training, regenerated examples of task-solving sequences, and reinforcement learning—learning from feedback or rewards—across coding, knowledge work, GPU software optimization, web development and computer-aided design. That account does not establish Alex’s broader description of which developers’ traces were used.
More chips, harder coordination
Alex’s optimistic case for Elon Musk’s strategy paired the post-training shortcut with access to NVIDIA GPUs, the processors used to train and run AI models. Learning from strong examples could close the gap; computing capacity could give the company room to attempt the next step.
The panel also discussed much larger models. Its account put Grok 4.5 at 1.5 trillion parameters, subsequently post-trained into 4.6, with 4.7 rising to 2 trillion and later models reaching 6 trillion and then 10 trillion. Parameters are the numerical settings learned during training—a measure of model size, not a performance score. Those figures were claims made in the discussion, rather than specifications established by the supplied release announcement.
Emad Mostaque’s explanation of the scaling problem centered on coordination. Writing training code, he argued, was less difficult than “trying to get 100,000 GPUs to do anything constructively together.”
Training a large model across many processors means their work must stay coordinated. Adding chips therefore creates a systems-engineering problem as well as adding computing power. Mostaque described fresh difficulties at successive scales and the time needed to get new hardware working effectively while meeting the intended price points.
That was his account of the bottleneck, not a confirmed explanation for the release delays the panel discussed. Alex also recalled Musk’s ambition to start a new pre-training run approximately monthly. He treated that as an unusually aggressive plan, not an established production schedule.
From a smarter model to a persistent teammate
The price brought the discussion back to the product. One panelist read the release as a move toward “persistent AI teammates”: systems that stay with work rather than simply produce a clever answer and stop. Lower-priced reasoning matters in that design because an extended task can involve many rounds of planning, tool use, testing and revision.
The panel expected Grok 4.7 to introduce another ingredient: SpaceX physics and engineering knowledge. That was an expectation about a future model, as was the claim that it would surpass Opus—not a result demonstrated by Grok 4.6.
For the model already released, the company’s stated direction is concrete: take a product idea through implementation, test the work, and keep going across the steps needed to produce a usable first version. The panel’s question was whether the next gains would come from better examples, larger training runs—or making that whole sequence work reliably at the price on offer.