15 September 2026
Heard In AI

Idea

Paying only when the work is done: the case for outcome-based AI pricing

On Moonshots EP #285, the panel worked through a proposed shift in how AI is sold: not by tokens consumed but by results delivered. This explainer sets out what outcome pricing means, the fixed-price and CRM precedents the panel cites, Alex's advertising-market analogy, Dave's regional-bank argument that it could preserve jobs, and the objections about failed delivery and reward hacking.

An idea page explains a concept, theory or proposal: where it comes from, what supports it, the objections and the open questions. We update it when new material changes the explanation. How our formats work

Peter Diamandis announced the turn himself. "I'm going to move us away from tech innovation to business model innovation," he said on Moonshots EP #285, and then described what he called one of his favorite stories: AI companies starting to charge for finished work instead of for the machinery that produces it.

In his account, Salesforce was first, pricing its Agentforce product on customer revenue generated rather than tokens consumed, and OpenAI had that week done something similar, letting some of its largest customers pay only when its AI actually completes the job. "You don't pay for tokens. You don't pay for compute time. You don't pay for API calls. You pay when the work is done." Those pricing changes are as the panel described them on air; the discussion did not present contract documents.

The distinction matters because of how AI is normally billed. A token is the unit these services meter — roughly a fragment of a word — and customers are charged per million of them, for what they send in and for what the model produces, including the text it generates while working out its answer. Under that arrangement the buyer pays for the attempt. Under outcome pricing the seller only collects if the attempt succeeds. Peter put the difference in one line: the company selling you tokens is selling you compute, and the company selling you results is selling you labor. He compared it to the old contracting choice between fixed-price and time-and-materials work — effectively, he said, a performance guarantee.

Where the idea comes from

Dave reached further back than Salesforce. He credited Tom Siebel at Siebel Systems with the move, before Marc Benioff. Before those companies, he said, a CRM system — the software a sales team uses to track its customers — might cost around fifty dollars a year for a license, "but it wouldn't work particularly well." The change was to wrap the license in something larger: if my salespeople are twice as effective, what is that worth? At that point a customer would pay twenty or thirty thousand dollars a year. The price point went up by a factor of a thousand, in his telling, "but the customer was happier because they got the total solution."

That is the shape of the proposal. The vendor stops selling a component and starts selling a business result, and takes on the work of making the result happen.

Dave's reason for liking it is that token prices cannot express how much a task is worth. Earlier in the same episode the panel had discussed an AI research campaign that spent on the order of ten billion tokens designing a mission to Alpha Centauri. "Okay, what's the pricing model for that?" he asked. It depends entirely on the use case, and could "easily could vary a million to one." Some work — he pointed to a colleague's projects on global peace and global governance — you want the tokens spent on regardless; other work, like drug discovery, should not absorb everything. If a system discovers a drug worth hundreds of billions, he argued, the developer can ask for a share of that instead of a metered fee. He expects outcome pricing to unlock uses that "might otherwise not fit the price model."

He attached an immediate condition: "Delivering it's not so easy, though. You need specialists in every market." His closest existing analogy was enterprise coding companies where the customer receives the finished answer at an attractive price and does not track the tokens burned along the way.

A market that looks like digital advertising

Alex pre-registered a prediction about where this settles: it ends the way quantitative digital advertising is monetized.

Advertisers can buy on three bases. CPM is cost per thousand impressions — paying for the ad being shown. CPC is cost per click. CPA is cost per action, or conversion — paying only when someone buys. In a working market these prices sit in a rough relationship to each other, because for a given product there is an expected rate at which impressions become clicks and clicks become sales.

Alex maps AI onto the same three rungs. The equivalent of CPM is raw compute: flops, or hours of GPU time, as when someone rents hardware to run an open-weight model and prices how many GPU hours a dollar buys. The equivalent of CPC is tokens, which is how many teams already budget their projects. The equivalent of CPA is outcome pricing — decide what counts as success and pay for that. In equilibrium, he said, "to the extent there can ever be an equilibrium in the middle of a singularity," buyers would choose between them from a picker, the way they choose a bidding basis on Google or Facebook ads, "except it'll actually be useful."

The analogy also supplies a release valve. Peter asked where the risk goes on contracts OpenAI takes. Alex answered with the advertising auction: a buyer who bids two cents a click may find that no such price clears, and Google simply auto-pauses the campaign after a few minutes. The same idea applies here. If the money offered is too far from the compute and difficulty the task actually requires, the campaign pauses rather than running at a loss.

And the bids carry information. Alex's second argument is that outcome prices tell the seller what the world values. The panel has argued before that OpenAI missed the enterprise market while concentrating on consumers and that Anthropic went past it; in Alex's view an outcome-priced tier is a way to catch up, because it works as a price discovery mechanism. If pharmaceutical companies and consulting firms are each saying this task is worth ten thousand dollars and that one is worth a million, the seller learns which industries produce the most value per token — a landscape, he said, on which it can "revenue per token max."

The regional bank, and the jobs argument

Dave's case for the model is a story about a large regional bank. Its board knows AI is coming. It has no AI group, no foundation model, no idea. So it buys API access from Anthropic and OpenAI at around two dollars per million tokens, "which is so cheap, it's ludicrous," and proceeds to dork around with it without doing much.

Meanwhile the CEO believes the bank should be able to serve three times as many customers at half the price, and Dave thinks a supplier looking at the business would agree it is doable. So why is it not happening? On his account both ends are stuck. At two dollars per million tokens the account is too small to reach a vendor's priority list, so nobody at the vendor is motivated to go in and do the work; and the bank cannot hire the talent to implement AI properly.

Outcome pricing, in his argument, unsticks both ends. The vendor says: I will get you to exactly that target, three times the customers at half the price. "I will deliver that to you, but I want half the gain." Now the deal is worth thousands of times more than the token bill, so the supplier cares, the implementation actually arrives, and — this is the part he presses — the bank survives and its people keep their jobs. He frames it against the earlier expectation that AI automates everyone's job and leaves a universal basic income in its place, an outcome he says Sam Altman, Dario Amodei and Elon Musk all dislike. This is his prediction about incentives rather than an observed result; nothing in the discussion measured employment at a bank under such a contract.

The objections

Peter's objection is about failure. Take the interstellar mission as the example: suppose a customer offers a fixed sum for an astrodynamics solution meeting stated parameters — minimum energy, minimum time, whatever the constraints are — and the vendor burns all those tokens and misses the parameters. Then it does not get paid. So, he said, there will have to be some mechanism for judging how solvable a task is and how far the vendor can believe it will hit the customer's objective before agreeing to the price. Guaranteeing the result moves the cost of failure onto the seller, and the seller has to be able to estimate it.

Alex's objection is about success of the wrong kind. "Strong optimizers are incredible reward hackers," he said. Put a reward in front of a capable system and "it will find some outstandingly devilishly clever way to meet your criteria while not giving you what you want." If no such loophole exists, he added, it will find a way to make one exist. Outcome pricing requires writing down a measurable definition of the outcome, which is exactly the kind of target that invites this behavior.

The reply offered on the panel was that this problem is a long way off in practice. Most real-world business is, by AI standards, trivially simple — orders of magnitude easier than designing a route to Alpha Centauri in a week — so there is low-hanging fruit everywhere long before anyone reaches a contract the vendor cannot deliver. The supporting observation was about deployment rather than capability: walk into any company's customer service center and ask whether they are using AI yet, and by that estimate the answer is no "99.999%" of the time.

Open questions

The discussion left several things unresolved. Who defines and measures the outcome, and who audits it, when the vendor's fee depends on the number. What the auto-pause equivalent looks like in practice, given that an ad campaign can be stopped after five minutes while a delivery contract cannot. Whether the specialists Dave says are needed in every market actually exist, since the shortage of implementation talent is the reason his bank was stuck in the first place. And whether a share-of-the-gain contract makes a buyer more willing to proceed or more exposed, since half the lift is a much larger number than the token bill it replaces.

Peter closed the segment without settling it. "I still think it's going to be dependent on the bets that they take."

Share this idea

Connected ideas and articles

From the conversation

Podcast episodes

Page history

How this idea page has changed

Tags

The weak point in Friedberg's AI jobs optimism: workers have to choose to move

On The Diary of a CEO, investor David Friedberg argued that AI grows companies rather than shrinking payrolls, using a painter commanding five robots as his image. Pressed on entry-level hiring, Klarna's staffing numbers and a Stanford payroll study, he named the assumption he thinks could be wrong: that displaced people will see an opportunity and take it.

10 min read

Why a Moonshots panel thinks China's AI tokens go to video and America's to code

Alibaba's Wan 3.0 and a relayed claim that 70% of Chinese AI token use goes to video sent the Moonshots panel into an argument about money: one guest said American labs chase revenue per token while Chinese labs give their weights away, another said video is the only market that will trust a Chinese model. They ended up disagreeing about whether world models or text models reach self-improving AI first.

6 min read

Why more agent output left the Moonshots panel working harder

On the Moonshots podcast, Salim Ismail, Alex and Emad Mostaque describe the same problem from different desks: agents now produce more work than a person can review. Their answers range from designing escalation thresholds inside companies to Mostaque's decision to read his research agents' output only once a week.

6 min read

Parallel sold patient search agents before it could afford a web index

On Training Data, Parag Agrawal explains how his company Parallel entered web search without first building a giant index: it launched a search agent that crawled after a request arrived, replaced outsourced human data collection for insurance, sales and finance customers, and treated the index as a latency optimization to be grown later. He describes the agent-specific architecture behind it, the 200-millisecond Turbo mode Parallel announced in July, and a Google Cloud deal that puts Parallel Search beside Google Search as a grounding option.

7 min read