Intercom charges $0.99 per resolution. Chargeflow takes 25% of recovered chargebacks. Sierra AI crossed $150M in ARR selling outcomes to enterprise brands.
Outcome-based pricing has become the default model for AI agents in customer support. The pitch is simple: the vendor gets paid when the buyer gets value.
Whether that alignment survives contact with an invoice comes down to one test. Outcome pricing works when the outcome is binary, meaning it happened or it didn't, and both sides can verify it independently.
The companies getting it right picked outcomes that pass: a chargeback won or lost, a ticket resolved or escalated. The companies struggling picked outcomes that are subjective, opaque, or defined by the vendor who profits from counting more of them.
This stopped being a niche question fast. Intercom, Zendesk, HubSpot, and Salesforce all sell support AI priced by resolution or conversation now, which means buyers are signing contracts denominated in a word that has no standard definition.
The Binary Test
Chargeflow passes cleanly. A chargeback is either won or lost. The credit card network decides, not Chargeflow, and the merchant can verify the result in their own payment processor dashboard. Chargeflow takes 25% of recovered funds; if the chargeback is lost, the merchant pays nothing. The incentives hold because the outcome is a fact, not an interpretation.
Sierra negotiates outcomes per contract. A "resolved support conversation" is defined in writing before the deal closes. A retention outcome means a customer who came to cancel stayed. A revenue event means money moved. At $150K+ annual commitments, both sides invest in getting the definition right, so the outcome is negotiated but binary once defined.
Intercom's Fin is where the line starts to blur. A "resolution" can be hard, meaning the customer confirms, or soft, meaning no follow-up within 24 hours. The hard resolution is binary. The soft one is an assumption: silence inside an arbitrary window gets treated as success.
That assumption carries a cost.
When Silence Passes for Success
One support team followed up with customers who received soft resolutions from Fin. Only 62% said their issue was actually resolved. A 38% false-positive rate, with every false positive billed at the full $0.99.
Silence can mean the answer worked. It can also mean the customer gave up, called phone support instead, or tried the solution 25 hours later and found it still didn't work. All four count as billable resolutions under a 24-hour window.
Community posts have flagged another edge case. Fin counts a resolution even when a human agent takes over the conversation, as long as the AI responded first. The AI's contribution might have been a generic greeting, and the resolution fee still applies.
At that point the outcome is no longer binary. It's probabilistic, and the entity assigning the probability is the same entity collecting the fee. With token billing you can count tokens yourself, and with seats you know how many you bought. With resolutions, the open question is what an independent audit would even look like.
Five Vendors, Five Definitions
Every vendor in this market bills for something called a resolution or a completed conversation. No two of them mean the same thing by it.
| Provider | Price | What Counts | Verification | Risk |
|---|---|---|---|---|
| Intercom (Fin) | $0.99/resolution | Customer confirms OR 24h silence | 24-hour window | Pays for abandoned conversations |
| Zendesk | $1.50-$2.00/resolution | AI analyzes conversation for relevance | 72-hour AI evaluation | Black-box classification |
| HubSpot (Breeze) | $0.50/resolved conversation | "Conversation successfully completed" | Not disclosed | Lowest price, least transparency |
| Salesforce (Agentforce) | $2/conversation or $0.10/action | Completed conversation or discrete action | 24h inactivity | Simple query costs the same as complex case |
| Sierra | Custom | Contractually defined per deal | Agreed per contract | Requires $150K+ commitment |
Price points vary 4x for superficially similar products. HubSpot charges $0.50 and Zendesk charges up to $2.00, both resolving support tickets with AI. The difference isn't quality. It's definition breadth and verification rigor.
Zendesk's 72-hour window with AI-based verification is more conservative than Intercom's 24-hour silence rule. But the AI that decides whether a resolution was "satisfactory" is a black box. You can't audit the classification logic. You can only compare your internal satisfaction scores against the invoice and hope they correlate.
Salesforce has already moved once. The original $2-per-conversation model drew criticism because a simple order status check cost the same as a complex multi-step case, so Salesforce introduced $0.10 per action pricing. The new unit arrived less than 18 months after the old one launched.
What a Loose Definition Costs in Practice
The definitions stop being abstract the moment you sit in a specific chair.
Start with a buyer comparing quotes. A head of support with 5,000 AI-eligible tickets a month would pay HubSpot about $2,500 and Zendesk up to $10,000 for the same ticket set. The spread says nothing about which agent answers better. The useful move before comparing prices: get each vendor's definition in writing, and ask how you would verify a billed resolution yourself. A low rate on a loose definition can cost more than a high rate on a strict one.
Now the founder pricing her own agent. Say it resolves 65 of every 100 attempts, and inference on an attempt runs about $0.20. She pays for all 100 attempts and bills for 65, so each billed resolution carries roughly $0.31 of model cost. That number moves when the upstream provider reprices, when ticket complexity shifts, when the knowledge base regresses. None of those require her to change a line of code.
The finance lead has the hardest seat. Traditional SaaS is arithmetic: 50 seats at $100 a month is $5,000, every month. One founder's Intercom bill went from $200 to $1,400 during a product launch, a 7x move from a single external event. "Somewhere between $3,000 and $12,000 depending on how good our knowledge base gets" is not a number a CFO can put in a quarterly plan. Zendesk's committed rate ($1.50 instead of $2.00) helps only if you can forecast volume, which is the thing that couldn't be forecast in the first place.
Who Pays for the Failed Attempts?
For vendors, the model has a structural consequence worth sitting with. Outcome pricing shifts risk from the buyer to the vendor. The buyer pays only when value is delivered, which is the selling point and also the exposure.
The vendor absorbs the cost of every failed attempt. If 30% of AI resolutions require human escalation, the vendor eats the inference cost on that 30% with zero revenue: the model inputs, the compute, the API calls to the underlying LLM, all consumed and none recovered. Every failed resolution is a cost that never appears on the invoice but absolutely appears on the margin report.
This is, frankly, how outcome pricing should work. The product mistake belongs to the builder, and the bill shouldn't. Absorbing variance deliberately is very different from absorbing it without measuring it.
And the margin dynamics compound quietly. Better AI resolves more tickets, which is good. Better AI also attempts harder tickets, which raises the failure rate at the tail. The average resolution rate improves while the cost per failed attempt grows.
One support team improved their help center content and made Fin more effective. Resolution rate jumped from 40% to 65%, and costs rose over 60%. The AI got better and the invoice got bigger, the same quality-cost inversion that hit GitHub Copilot.
Where Outcome Pricing Holds
Outcome pricing is, we'd argue, the future for AI products that deliver measurable results. No other model ties payment so directly to impact.
The vendors making it work, Sierra and Chargeflow among them, share two traits. Their outcomes are binary and independently verifiable, and they know what each outcome costs them to deliver. Sierra negotiates the definition; Chargeflow lets the card network decide. Neither leaves the outcome definition to a probabilistic model controlled by the party collecting the fee.
The vendors struggling, or repricing in Salesforce's case, share the opposite trait. The outcome is fuzzy, the measurement is opaque, and nobody on the vendor side can answer the question the whole model rests on: what does it cost us to deliver one successful outcome?
Without that number, outcome pricing is a guarantee made against an estimate. The gap between resolutions billed and resolutions that cost less than $0.99 to deliver decides whether the model compounds or erodes, quarter by quarter.
If you're building toward outcome pricing, Bear Lumen connects inference costs to individual outcomes, so you see your true cost per resolution continuously instead of reading it off the margin report. The pricing framework covers where that number fits in the rest of the pricing decision.