15 September 2026
Heard In AI

Anthropic's cheaper cached reads make business context the prize

Anthropic's Fable 5.1 charges $0.25 per million tokens for cached reads, a quarter of the previous rate, which one Moonshots panelist read as an invitation to load an entire company's context into the model and keep it there. The panel connected that price to a wider scramble: with model leads lasting about a month, the labs are racing to convert them into customer workflows, partnerships and proprietary design data that a rival cannot copy.

A briefing reports one development when it happens. We correct or clarify it later; a new development gets a new briefing. How our formats work

A week after Anthropic released Fable 5.1, one panelist on Moonshots with Peter Diamandis skipped past the benchmark charts and went to a line in the pricing table. The cache reads, he said, are "75% cheaper than Fable 5" — and in his account that is where the interesting competition now sits.

A cached read is what a model charges to re-use context it has already processed. When a company first loads its documents, records and instructions into a model, the system works out, as the panelist put it, "a rapid map to get to where it needs to go." Storing that map and reading it back costs a fraction of processing the same material again from scratch on every request. Anthropic's documentation prices Fable 5.1 at $10 per million input tokens and $50 per million output tokens, with cached reads at $0.25 per million — one-quarter of Fable 5's rate. Tokens are the chunks of text models read and write; a million of them is roughly a small shelf of documents.

The saving is not unconditional. Creating the cache is billed separately: $12.50 per million tokens for five-minute caching and $20 per million for one-hour caching. Loading a company's context once and returning to it repeatedly therefore costs something very different from arriving with fresh material on every request. The economics reward organizations that keep coming back to the same body of knowledge.

What Anthropic released

Fable 5.1 and Mythos 5.1 arrived together, described on the show as the world's most advanced models for coding and knowledge work. The host described them as essentially the same underlying intelligence with different safety envelopes: Fable 5.1 broadly available, Mythos 5.1 reserved for tightly controlled cybersecurity and life science programs because Anthropic believes those capabilities require stronger safeguards. Anthropic's documentation dates the release to 1 September 2026, gives both models a one-million-token context window and a maximum output of 128,000 tokens, and offers Mythos 5.1 by invitation through a program it calls Project Glasswing.

The context window is the amount of material a model can hold in view at once. Asked when the industry gets an infinite one, a panelist said a million tokens is now the standard across both OpenAI and Anthropic, with an "effective context" he qualified heavily — considerably larger, he said, if agents are allowed to pass messages to each other rather than stuffing everything into a single window.

Prices at that scale are not abstract for heavy users. One panelist reckoned that he and fellow panelist Alex were "at least $100,000 into it already" after a week, with thousands of pages of output to show for it.

Quality the panel could feel

The cheaper cached read was not presented only as a cost line. The same panelist argued that reading from the cache is also much faster, which is part of why Fable 5.1 responds to its environment more quickly and, in his words, more pleasantly.

The panel's impressions of the earlier model were unkind. One said Fable 5 was "really terrible to talk to. I hated it," and that 5.1 is pleasant. Another called Fable 5 "so geeky. It was almost torture," adding that 5.1 fixed it — and that OpenAI's models have always been friendlier and more concise. Out of the box, he said, the models now have noticeably different personalities, though users can tune them to be wordier or simpler.

One panelist offered a sharper test from his own field, mathematical physics: does the model confuse a constructive method with an axiomatic one on certain physics problems? Fable 5 did, he said; Fable 5.1 does not. He treated that as evidence of improved understanding of context rather than only better answers — knowing, as he put it, that "this is math, not physics." The same discrimination, he expected, would eventually be optimized for business material.

A lead that lasts about a month

Why spend so much effort on context? Because, on the panel's reading, nobody keeps a capability lead for long. Fable 5.1 sits a notch above OpenAI's Astra, one panelist said, but the two releases are about 30 days apart, and Chinese models roughly 60 days behind that. "We're ahead for a minute. So what?"

His answer was that the lead has to be converted into something stickier. Both companies, he expected, would spend it locking up business partnerships, real estate, generators, chips, entire states, countries and governments — "otherwise, what's the point? All you're doing is declaring victory for 30 days."

Another panelist described what that looks like from the customer's side. Labs are signing partnerships vertical by vertical, and a large company in one of those verticals faces what he called a Hobson's choice: partner and risk handing over the keys to the kingdom, or hold off and watch the lab partner with a competitor anyway. He read Salesforce's partnership with Claude as clever on exactly these terms — giving the model its capability in order to stay wired into the loop. "The huge tension they've got is not which is the best model, but how quickly can you convert that model into customer learning fastest?" As the models themselves demonetize, he argued, the value moves to the application layer on top.

That is why, in his account, companies sitting on chip design data or mechanical design data are being approached now. Those archives plug knowledge gaps that abundant public text cannot: a model could consume them in about a week, he said, and come out the best mechanical designer or chip designer around. "That's turf you can defend."

The test a buyer can run

The panel's earlier diagnosis of corporate adoption sets the price cut against what most companies are actually doing. Asked whether the chief executives he meets on the road grasp the speed of change, Salim Ismail said they are "woefully behind" and mostly dabbling. His thought experiment: if you removed AI from your organization today, would any workflows change? For most, he said, the answer is no — which tells you they are tinkering rather than making structural change. The advantage, in his view, lies with the organizations rewriting their workflows and organizational design.

A quarter-price cached read matters only to a buyer who has something worth caching: records, processes and instructions that a model reads again and again. Building that is the structural change Ismail says most companies have not made — and, on the panel's account, the same asset the labs are competing to sit inside.

Share this article

Go to the original

Sources & further reading

  1. 01

Connected ideas and articles

From the conversation

Podcast episodes

Article history

Updates to this article

Tags

Would she pay the real price? Zitron's test for AI adoption

On The Diary of a CEO, critic Ed Zitron praises a chatbot for reading a troubleshooting log and for helping fix his son's Minecraft mod, then argues that neither is worth a trillion dollars. The host counters with his fiancée's one-woman business and his chief of staff's inbox. The argument turns on tokens, subscription rate limits and who is paying the real bill.

7 min read

A 100x claim lands, and the panel hits a harder question: coordinating 10,000 agents

On Moonshots, the panel revisits Elon Musk's January prediction that models were "off by two orders of magnitude" in intelligence per gigabyte, after Tim Sweeney tweeted that it had come true and Musk replied that specialist AIs add another 100x. Dave calls 100x a lower bound and asks what anyone would actually do with 10,000 brilliant agents; Emad Mostaque describes running specialized agent teams, while Alex argues Musk's "specialist models" are really sparsification inside generalist models.

6 min read

Why a Moonshots panel thinks China's AI tokens go to video and America's to code

Alibaba's Wan 3.0 and a relayed claim that 70% of Chinese AI token use goes to video sent the Moonshots panel into an argument about money: one guest said American labs chase revenue per token while Chinese labs give their weights away, another said video is the only market that will trust a Chinese model. They ended up disagreeing about whether world models or text models reach self-improving AI first.

6 min read

Altman says economic inertia slowed AI's impact. His podcast panel disputes the cause

On Moonshots with Peter Diamandis, the panel watched Sam Altman explain that he expected GPT-4 to put software businesses up for grabs far sooner than it did, and that the economy's inertia has made the transition "smoother and slower." Salim Ismail blamed institutions that move at a different speed from the technology, Alex pointed instead at the abstraction layers of the economy and prescribed vertical integration, and Emad Mostaque objected that the models simply were not good enough until recently.

6 min read