15 September 2026
Heard In AI

Box's Aaron Levie expects open-weight tokens and closed-model revenue to grow together

On Training Data, Box CEO Aaron Levie describes how his customers actually pick models: a default for asking questions of their files, and hard-nosed accuracy evaluations for the high-volume extraction work where most tokens are spent. He endorses Decagon founder Jesse Zhang's argument that mature workflows migrate to open-weight models, and explains why the big labs' revenue and open-weight token volume can climb at the same time.

A briefing reports one development when it happens. We correct or clarify it later; a new development gets a new briefing. How our formats work

A Box customer who wants thousands of contracts or invoices read by an AI agent does not usually start by naming a model. According to Box chief executive Aaron Levie, speaking on Sequoia's Training Data podcast, the customer starts with a target: "I want to really make sure that at this cost profile, I can get 98% accuracy on data extraction." Box then sends an engineer into the customer's data environment to test against five different models, and, as Levie puts it, "then you're just basically at the mercy of the eval."

Box began as a way to store and share files securely in the cloud, and now sells AI agents that work over those files. An "eval" is a fixed set of test tasks with known answers, used to compare one model against another on the work a customer actually cares about. The unit of cost in this business is the token — the chunk of text a model reads or writes, and the thing providers bill for.

A default for questions, an eval for volume

Levie draws a line between two kinds of demand inside the same product. The way most end users meet the agent is by searching and asking questions of their own documents, and there, he says, customers mostly accept Box's default model "because it's just easy and it works extremely well." Box sets those defaults from its own testing against different cost and accuracy thresholds: a document-centric Complex Work Eval covering domains such as life sciences, financial services, the public sector and technology, plus a second held-back test on Box's own instance and how Box employees use their data. Customers can still choose any model from what he calls Box's model garden.

But the tokens are not where the users are. By volume, Levie says, most of them flow through workflow agents and data extraction — the repetitive, high-throughput jobs. That is exactly where customers stop accepting the default and run their own evaluations at a chosen cost profile.

"Higher than people think, lower than what enterprises actually want"

Asked how far open-weight models have spread among Box's customers, Levie gives a three-part answer: "probably higher than people think, lower than what enterprises actually want. And much, much, much, much, much lower than what it'll be in five years." Open-weight models are ones whose parameters can be downloaded, run on your own infrastructure and further trained; the leading closed models can only be reached through their owners' interfaces.

He does not think cost is currently the main driver. More than 30% of the interest, he estimates, comes from "the sexiness of, like, I want to try GLM." He describes hearing chief information officers at Fortune 500 companies say "we're playing with open source here," when he believes a closed model such as Gemini — or whichever one had just run steep discounting — "would have been just fine at that particular cost profile that you're trying to do." The motive he ascribes to them is hedging: not wanting to be locked to one supplier. "Like, we're at that phase still."

The rough edges are practical. Open-weight models can be less token-efficient, meaning they consume more tokens for the same job. And Levie relays stories of a model that "randomly" switches into Chinese mid-chain — partway through its own reasoning. "Okay, well, that'll be weird for a bank."

The maturity argument Levie subscribes to

For the longer run, Levie points to an essay he says he "totally" subscribes to: Jesse Zhang of Decagon, the customer-service AI company, on open source in the enterprise. In that July 2026 post, Zhang reported that roughly 90% of Decagon's workloads ran on open-source models, and attributed that mainly to latency and task-specific fine-tuning rather than to cost savings: a live customer-service conversation needs a small, fast model that still meets a demanding quality bar. Zhang's broader claim is about maturity. New workflows, where you do not yet know the inputs, the desired behavior or the ways things fail, favor the most capable general-purpose models. Once a workflow is understood and stable, it can be moved to a smaller specialized model. He expects that migration to take years, because fine-tuning requires enough data, expertise, scale and return to justify it.

Levie states the condition for peeling a workload off in two forms: either the alternative is "literally cheaper," or "having some post-training gets you X percent more performance" — post-training being additional training on top of a released model to specialize it.

Why token share and spending share can diverge

This is the part Levie expects to read strangely: the revenue of Anthropic and OpenAI going "off the charts" while open-weight usage also grows exponentially. "The pie is growing so fast," he says, but he describes something more specific than a rising tide. He sketches two routing patterns inside a single agent. In one, an expensive frontier model does the orchestration — deciding what needs doing and in what order — and farms long-tail subtasks out to a cheaper model. In the other, a cheap model handles work by default and escalates a case that looks too hard to a heavier model.

Either way, the arithmetic splits. "You might have blended 50% spend on each, but 10 times the amount of tokens on the OpenWeights model." Both suppliers can report growth from the same workload, for different reasons.

Who gets to choose the model

Behind the routing is a question about who is trusted to pick. The host raised it as not wanting "the fox" guarding "the henhouse": the seller of the token also metering and gating which token is best for each use case. Levie takes the same view, and says it may be "singularly the biggest" reason the application layer matters. Even assuming complete benevolence, he argues, when he hands a task to an agentic system he wants it cost-optimized with accuracy held constant — and the party that can do that is the one with no preference among ten models.

The force working against that, he acknowledges, is that model companies can subsidize their models or offer them at different rates. He expects that to be temporary, and says he means it "entirely" as a gross-margin argument rather than an antitrust one: once these companies are public they will face the same laws of capitalism as everyone else, subsidy works only up to a certain threshold of spend that tens of billions of dollars of capital may already exceed, and the training runs still have to be paid for. Pressed on why inference needs subsidizing at all when margins on it are high, the exchange turned to the gap between what a lab charges outside developers through its interface and what it charges for its own consumer products — and to the equilibrium problem that follows: if the developer interface is the high-margin business paying for the subsidy, moving customers onto the lab's own applications removes that revenue.

Levie then adds a set of players whose margin needs differ: Meta, SpaceX, China and even Nvidia, cohorts he thinks could accept something like 10% margin on inference because what they want is to fill massive compute clusters. If that holds, and there are no closely held proprietary secrets, cost per token falls on a like-for-like basis and more value lands on the application layer. His conclusion is a distribution rather than a winner: "there's just value in kind of everybody in the stack," and the one outcome he would not bet on is one or two labs capturing 95% of the value created. He suspects the biggest labs should prefer that too — "at some point you'll just be nationalized if you're the only thing that sort of exists as intelligence."

Box is positioning for the same duality. Levie says the company has a Box Labs effort, "a little bit secret," paying close attention to models. Its current focus is not to build one: "let's make the agent really, really good on any model." Over time, he says, certain use cases get peeled off — on a per-customer basis or across the whole data set — which is the same maturity logic applied one workflow at a time.

Share this article

Go to the original

Sources & further reading

  1. 01

From the conversation

Podcast episodes

Article history

Updates to this article

Tags

Would she pay the real price? Zitron's test for AI adoption

On The Diary of a CEO, critic Ed Zitron praises a chatbot for reading a troubleshooting log and for helping fix his son's Minecraft mod, then argues that neither is worth a trillion dollars. The host counters with his fiancée's one-woman business and his chief of staff's inbox. The argument turns on tokens, subscription rate limits and who is paying the real bill.

7 min read

Why a Moonshots panel thinks China's AI tokens go to video and America's to code

Alibaba's Wan 3.0 and a relayed claim that 70% of Chinese AI token use goes to video sent the Moonshots panel into an argument about money: one guest said American labs chase revenue per token while Chinese labs give their weights away, another said video is the only market that will trust a Chinese model. They ended up disagreeing about whether world models or text models reach self-improving AI first.

6 min read

Box's two rules for software in the agent era: beat the generic agent, then let it in

On Sequoia's Training Data podcast, Box CEO Aaron Levie said any company sitting on customers' data now has two obligations: build an agent measurably better than an off-the-shelf one at its own workflows, and expose the same capabilities to outside assistants like Claude and ChatGPT. He described the tuned search-and-retrieval harness behind Box's agent, the evaluations that track model progress, and his bet that within five years roughly 90% of enterprise tokens will be spent on work nobody asked for directly.

8 min read

Anthropic's cheaper cached reads make business context the prize

Anthropic's Fable 5.1 charges $0.25 per million tokens for cached reads, a quarter of the previous rate, which one Moonshots panelist read as an invitation to load an entire company's context into the model and keep it there. The panel connected that price to a wider scramble: with model leads lasting about a month, the labs are racing to convert them into customer workflows, partnerships and proprietary design data that a rival cannot copy.

5 min read