Outcome-Based AI Pricing Didn't Fix the Metering Problem. It Just Put the Vendor in Charge of the Metric.
Salesforce and Intercom both sold pay-per-outcome as fairer than pay-per-seat. Both ended up in public disputes over what their own billing unit actually means.
When a vendor pitches paying per outcome instead of per seat, the pitch always sounds the same. Stop paying for the twelve people on the license who barely log in. Start paying only for the work that actually gets done. It sounds like a strictly better deal for the buyer, and in theory it can be. For a growing share of AI agent products, it hasn't played out that cleanly in practice. The reason is specific, and it has nothing to do with whether usage pricing is a good idea in general. The deal quietly moves who gets to define the thing you're paying for.
A seat was annoying to buy, but nobody had to trust the vendor to count it
Per-seat pricing has a real flaw. It charges for provisioned access, not usage, so a company with 40 unused licenses still pays for 40 unused licenses. That flaw is well understood and it's the entire reason usage pricing exists. But a seat has one property that got lost in the rush toward usage-based and outcome-based AI pricing: it's externally verifiable. A login event is a login event, recorded on both sides. If a vendor's invoice says 40 seats, a buyer's own admin console can confirm or dispute that number without asking the vendor for anything or trusting a single log line the buyer never sees.
"Conversation," "resolution," and "outcome" don't have that property, and this is easy to miss because the words sound concrete. There is no external system of record for what counts as one AI-handled resolution. No independent ledger, no shared API, nothing the buyer's own tools can query. The only system that logs it is the vendor's own backend, running the vendor's own definition of the word, which the vendor is free to change, and in the two cases below, did.
Salesforce redefined "conversation" three times in about a year
Agentforce launched with a flat rate of $2 per conversation. It was a simple number to put in a slide and a hard one to budget against, because a single customer query could trigger several backend actions: pulling a record, checking an order status, drafting a reply. Buyers had no clean way to know in advance which of those steps would get bundled into one billable "conversation" and which would count as separate ones, which meant the same support volume could land at wildly different price points depending on how the backend happened to chunk the work that month.
Salesforce's own pricing history bears this out. In May 2025, it introduced Flex Credits, a consumption model priced around $500 per 100,000 credits that lets buyers pay per action instead of per conversation. That didn't replace the conversation-based rate. It sat alongside it. Per-user licensing starting near $125 a month came next, on top of both. By 2026, Agentforce had three live pricing models running at the same time, and Salesforce admins were still naming pricing friction as a top complaint alongside reliability concerns, well after the "fix" had shipped twice.
Intercom's "resolution by silence" is the same problem from the other direction
Intercom prices its Fin AI agent at $0.99 per resolution, with a $49-per-month base plan that includes an initial allotment. On paper that reads cleaner than Salesforce's conversation model. A resolution is either resolved or it isn't, and there's no bundling question to argue about. In practice, Fin counts a ticket as resolved through two different paths. One is the customer explicitly confirming the answer worked. The other is the customer simply going quiet after Fin's reply and never following up.
That second path is an inference, not a confirmation, and it's where the disputes concentrate. A customer who gives up on a bad answer and quietly emails support through a different channel looks identical, in Fin's logs, to a customer who got helped and moved on with their day. Once ticket volume crosses a few thousand resolutions a month, that inference gets expensive in aggregate. Support teams have reported monthly Fin bills running past $5,000 with no cap in sight, and a meaningful share of that total traces back to resolutions nobody on the support side ever agreed were actually resolutions.
“A per-resolution price is not the same thing as a predictable support budget. The gap between the two is the set of resolutions nobody but the vendor signed off on.”
Seat vs. outcome: what's actually being measured
| Dimension | Per-seat | Per-outcome |
|---|---|---|
| What's counted | A login or provisioned license | A "conversation," "resolution," or "action" |
| Who defines the unit | Fixed by contract at signing | Set by the vendor's product logic, can change |
| Can the buyer verify it independently | Yes, the buyer's own admin console shows seat counts | No, only the vendor's backend logs the event |
| Where disputes happen | Rare, since seat counts are unambiguous | Common, concentrated in inferred or bundled edge cases |
| Failure mode | Paying for unused capacity | Paying for events the buyer never agreed were billable |
Hybrid pricing doesn't solve this. It doubles the surface area
Seat-based pricing has genuinely been shrinking. One 2026 pricing survey found the share of SaaS companies pricing purely per seat fell from 21% to 15% in twelve months, while hybrid models, meaning a base seat fee plus a usage or outcome component, grew from 27% to 41% over the same period and are on track toward roughly six in ten vendors by year-end. Analysts at IDC have projected that most software vendors will have moved off pure per-seat pricing altogether by 2028, largely because agents do work independently of any one logged-in user, which breaks the basic logic of charging by seat in the first place.
Hybrid sounds like the reasonable middle ground, and for a vendor managing revenue predictability, it probably is. For a buyer, it isn't a compromise between an auditable metric and an opaque one. It's both running at the same time. The buyer is now trusting the vendor's seat count and the vendor's outcome count, and when a bill comes in higher than expected, either one could be the cause, with no independent way to tell which before asking the vendor to explain its own math.
Why this is a new problem, not an old one wearing a new label
It's tempting to file this under "read your contracts," which is true but misses what changed. Traditional usage metering, like API calls or gigabytes stored, counts something mechanical and low-level enough that vendor and buyer rarely disagree about what happened. A gigabyte stored is a gigabyte stored regardless of who's asking. An AI agent's "resolution" or "conversation" is a judgment call baked into a model's output, made by software the buyer doesn't control, describing an interaction the buyer often can't fully replay. The metering unit itself now contains an opinion, and that opinion belongs entirely to the party sending the invoice.
What to put in the contract instead of the deck
None of this is an argument for refusing usage-based or outcome-based pricing outright. Agents that do real work independently of a logged-in human are a legitimate reason to stop billing by headcount, and plenty of buyers are genuinely better off on a usage model once it's priced sanely. The argument is that the definitional risk needs to move from the sales deck into the master service agreement before signature, in three specific places.
First, get the billable event defined in writing, in enough detail that a third party could apply the definition to a transcript and land on the same answer the vendor got internally. "Resolved" needs a test, not an adjective. Second, get a dispute mechanism with an actual credit remedy: a defined window to flag events the buyer disputes as billable, and a contractual commitment to credit them, not a vague promise to "reach out to support." Third, get a hard monthly cap or a true-up clause written into the agreement, so a spike in inferred outcomes can't quietly turn into an uncapped invoice at renewal.
A vendor that has already had to redefine its own metric in public, as both Salesforce and Intercom have, in different ways, is not a vendor whose pricing deck should stand as the final word on what a buyer will actually owe. Read the definition the vendor is willing to put in the contract, not the one it put in the pitch.
Frequently asked questions
Related reading
Vendors' AI Costs Fell 90%. Your Renewal Bill Rose 20–37% Anyway.
Vendor AI infrastructure costs fell 40–90% since 2024. Renewal asks still rose 20–37%. Here's where the gap actually goes, and the playbook that claws most of it back.
AI Agent Benchmarks Got Gamed to Near-Perfect Scores Without Solving a Single Task
Eight major AI agent benchmarks hit 73-100% scores without an agent solving the underlying task. A second 2026 study found the same gap honestly: a 37% lab-to-production drop and a 50x cost swing.
SAFE Note Dilution Is Decided Before Your Series A Price Is Set — Most Founders Model It After
Three SAFEs at three caps can look like low-friction fundraising. By the mechanics Y Combinator built into the post-money SAFE, they're also a fixed claim on founder equity that almost nobody totals up.