14 September 2026
Heard In AI

What Gemini 3.7 Flash's analyst benchmark win actually measures

On Moonshots with Peter Diamandis, the panel read out a new leaderboard result: Google's Gemini 3.7 Flash on top of the AA-AnalystAgent benchmark with 60%, ahead of Claude Opus 5 at 54%. Diamandis called it proof that Google is back; Alex argued the score measures repeated reliability on spreadsheet analysis rather than frontier capability, and blamed Google Search for pushing Gemini toward speed and determinism. Emad Mostaque agreed the model was decent but said Google's problem is institutional, not a shortage of chips.

A briefing reports one development when it happens. We correct or clarify it later; a new development gets a new briefing. How our formats work

Peter Diamandis had a line ready before the slide even came up. Google, he said, is back: Gemini 3.7 Flash had just taken the top spot on the AA-AnalystAgent benchmark, and the people who had written Google's epitaph in the frontier model race could now read it back.

The numbers he read out: 80 tasks across 14 business and scientific domains, the highest overall accuracy in the field, with tasks completed up to 90% faster than other leading models and 2.4 times faster than the GPT-5.6 variant he named. On the headline metric, Gemini 3.7 Flash scored 60%, ahead of Claude Opus 5 at 54% and a third model the panel called Fable 5 at 49%.

Diamandis drew a pattern from it. Meta was declared dead, he said, and then shipped models; xAI was written out of the race and came back. "To me, it seems like none of these players are out of the race. They're maybe in stealth mode," he said — holding back, then returning "with a fast, furious punch to try and take the top position."

Then he turned to Alex, who has posted hard against Google, and asked him to defend himself. Alex offered a choice: reassurance, or the facts unvarnished. He took the second option. "Google's still out of the running for the capability frontier."

What a 60% score counts

Alex said the result had made him scratch his head, because Gemini 3.7 Flash is "nowhere near the top of the capability frontier." His answer was to look at how the benchmark scores.

AA-AnalystAgent, an independently maintained evaluation, gives an AI agent folders of real spreadsheets and documents and asks quantitative questions about them: financial models, government appropriations, hydrology, energy-cost calculations. The agent works inside a Python sandbox containing the task's files, can fetch URLs and, for models that support it, view images. The published methodology caps a run at 100 turns and asks the agent to submit an answer value rather than an explanation. Grading compares that value with a held-out, human-written reference answer using a language-model judge, with a deterministic numerical check that guarantees a pass when the number matches in the required units and precision. The question set is kept private to limit the chance that models have seen it during training.

The part Alex seized on is the scoring rule. The headline figure counts a question as solved only if all five independent attempts got it right. So 60% means 48 of the 80 questions answered correctly five times out of five. That is a different quantity from average accuracy per attempt, and a much harder one than solving a question at least once.

"This is a benchmark that is fine-tuned for reliability," Alex said. "It rewards agents that give the same answer, basically the same answer every time... but it penalizes stochasticity. It penalizes in some sense creativity. Maybe we don't want creativity out of our analysts, I don't know, but it promotes reliability and determinism."

What the rule demands is the same correct value, not the same words or the same route to it. An agent that reasons differently on each of its five runs and lands on the right number each time passes exactly like one that repeats itself. Alex also said he had checked whether the benchmark imposed a time constraint and found none; the documented limit is on turns rather than on the clock, and the speed figures Diamandis cited are reported alongside the score rather than folded into it.

Having narrowed what the win shows, Alex offered what he called his conspiracy theory for why this particular model wins this particular test.

A year ago, he said, people were wringing their hands about whether Google's ten blue links could survive chatbots and reasoning agents that remove the need to search at all. Google's response was to disrupt itself, building Gemini — or at least some Flash variants of it — directly into the boxes at the top of its results pages. But users expect search results to be fast and to be stable: not "wildly different or unpredictable answers each time."

Those two pressures, he argued, have shaped the models, and they compete internally for scarce computing power. He pointed to the naming as evidence: "note that there's no Gemini 3.7 Pro anywhere. It's just Flash. It's small, it's fast, and it's reliable." His conclusion was that the model is over-optimized for wall-clock speed and determinism, "and as a result, it does well on the one benchmark that rewards highly reliable answers and underperforms" elsewhere. He presented this as his theory of the case, not as anything Google had said.

Compute shortage, or something else

Emad Mostaque half-agreed. Gemini models, as they are used now, are for organizing data across any modality, he said, and Flash is "a perfectly decent model" — but not as good as the Chinese models, singling out a new GLM Flash release he said delivers the same performance ten times cheaper. Google had previewed a Gemini 3.5 Pro, he said, and "it just couldn't keep up."

Where he broke with Alex was the explanation. "You can't excuse them of not having enough compute or it being a scarce resource. They literally have millions of chips," Mostaque said. He put the problem down to turnover and "some institutional malaise." Google has the data, the record of what people search for and a captive audience for the Gemini app, he said, and yet in his view the app has barely advanced; he credited AI Studio as decent and NotebookLM as continuing to be excellent, and saw the company falling behind elsewhere apart from multimodal work and video.

Alex pushed back that Google is more compute-starved internally than people realize. On his account, the company holds regular meetings to apportion scarce processing capacity among three constituencies fighting for the flops: Google Cloud Platform, arguing on behalf of external customers; Google DeepMind, arguing for training and inference; and Google Search, whose own needs grow as search gets more intelligent.

Diamandis objected that every company is compute-starved, and that Google has more than anybody — it is simply spread across all its products and services. Alex's distinction was that Google has large non-AI consumers competing for the same capacity, where OpenAI and Anthropic do not.

Mostaque then made the argument that most narrowed the disagreement. Google is landing around three million TPUs this year, he said, while training a model at or near the frontier takes roughly two to four thousand of them — the evidence being that Chinese labs did it and open-sourced the results, so the recipes are known. Apply the same architecture as GLM or Kimi to the data Google already feeds into Gemini Flash, he argued, "you should have a better outcome, but they're not doing that for some reason." Alex put the same figure as 2,000 to 4,000 chips for 60 to 90 days: "That is a microscopic investment by Google standards... it has nothing to do with compute dominance and everything to do with talent attrition."

Mostaque added a human reason the obvious shortcut does not get taken. Imagine the career position, he said: the best-funded AI engineer in the world walking into a boss's office to say the Chinese labs won, so let's download Kimi and tune it. "You can't say that... because you look like an idiot." Instead, he said, people leave in droves for a clean start. These accounts of Google's internal compute meetings, staff departures and training costs were the panel's own analysis rather than figures either speaker sourced on air.

One number, one axis

Later in the same conversation Alex made the point that complicates Diamandis's comeback reading and his own dismissal alike. On the hardest tier of a research-level mathematics benchmark, he said, the model he calls Fable 5 outperforms OpenAI's latest — ironically, since OpenAI was the primary sponsor behind that tier's development. "The frontier isn't zero-dimensional," he said. "It's a one-plus-dimensional frontier where if you're willing to pay a lot and wait a long time" for one model, the results are impressive; if you are short of money, time or compute, a cheaper model can sit at a better point on the cost curve.

That is the shape of the AnalystAgent result too. It says that on 48 of 80 private spreadsheet-and-document questions, an agent built on a small, fast Google model produced the right value five times running — and the leaderboard it sits on keeps adding later releases. "I just want to point out to everybody listening, it's not obvious, right?" Diamandis said. "There is a lot going on."

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02

Connected ideas and articles

From the conversation

Podcast episodes

Article history

Updates to this article

Tags

Parallel sold patient search agents before it could afford a web index

On Training Data, Parag Agrawal explains how his company Parallel entered web search without first building a giant index: it launched a search agent that crawled after a request arrived, replaced outsourced human data collection for insurance, sales and finance customers, and treated the index as a latency optimization to be grown later. He describes the agent-specific architecture behind it, the 200-millisecond Turbo mode Parallel announced in July, and a Google Cloud deal that puts Parallel Search beside Google Search as a grounding option.

7 min read

Grok 4.6 closes the gap—and the panel asks what would take it ahead

xAI’s August 12 release puts Grok 4.6 alongside GPT-5.6 Sol Max in its launch benchmark table, with pricing aimed at sustained agent work. The Moonshots panel’s debate was about the next step: whether training on other models’ reasoning can only help a challenger catch up, and what computing infrastructure it takes to move beyond that.

6 min read

NVIDIA's open-model push is about GPU demand, the Moonshots panel says

On Moonshots EP #283, Peter Diamandis introduced a reported $6 billion NVIDIA arrangement with the coding startup Poolside as America's answer to Chinese open models. Emad Mostaque argued the real driver is selling more GPUs, while Alex and Dave disagreed about whether licensing-and-hiring deals exist to dodge antitrust review or simply to hire fast — and what happens to the half of Poolside that stays behind.

7 min read

How Waymo's own chip changes the robotaxi cost math

On Moonshots with Peter Diamandis, the panel walked through Waymo's sixth-generation driver: a purpose-built 5-nanometer chip, fewer but sharper cameras, and an autonomous-hardware estimate falling from $115,000 to $20,000. Peter had ridden in the new vehicle; Alex objected that the West is now white-labeling Chinese hardware, and argued Waymo is heading for full vertical integration.

5 min read