Skip to main content
start
8 min read

The Usage-Based Pricing Trap: Why AI Companies Are Moving Back to Plans

Pure usage-based pricing is the default recommendation for AI products. The default is wrong. Hybrid models outperform pure usage on retention, predictability, and growth.

BLT

Bear Lumen Team

Research

Usage Based PricingPricing StrategyAI PricingPlan Based PricingCustomer Retention

A five-person engineering team budgeted around $100 a month for their AI code editor. Six weeks later they had spent $4,600.

That was Cursor, which switched its Pro plan to token-based billing in June 2025 and apologized within three weeks. Users who planned around $20 a month opened invoices for $60 to $100. Snowflake customers know the pattern at data-warehouse scale: cross-region data sharing quietly outruns the forecast. Notion AI went the other direction: a flat $10/member/month add-on that hit 65% workspace adoption within a year, with churn reduction worth an estimated $30M annually.

Two of those pricing structures produced apology posts and surprise invoices. The flat one kept its customers.

Matt Green, who studies pricing across 500+ SaaS companies at Growth Unhinged, compresses the whole problem into one line: usage-based pricing is easiest to close and hardest to renew.

Pure usage is still the default recommendation for AI products. The evidence below says the default is wrong.


The Year the Whole Budget Went Variable

Three years ago a typical software budget had 14 fixed line items and one variable cloud bill. A CFO could forecast Q3 spend in an afternoon. Today half those tools have moved to consumption pricing.

Budget Line2023 (Fixed)2026 (Variable)
Code editor$20/seat/moPer completion (Cursor)
Support tool$50/seat/mo$0.99/resolution (Intercom Fin)
Analytics$500/mo flatPer query
Data warehouse$2,000/mo flatPer compute credit (Snowflake)
AI assistant$30/seat/moPer message

One variable line item is manageable. Five compounding against each other is a different job.

Picture the CFO closing the quarter. The code editor came in at double plan. The warehouse ran over because one team turned on cross-region sharing in March. The support tool under-ran, which sounds like good news until you ask why. No single miss is large. Together they turn the software forecast into a range instead of a number.

78% of IT leaders report unexpected charges from consumption-based AI pricing, and 90% of CIOs name cost forecasting as their top challenge in AI deployment. Those aren't complaints about one vendor. They describe what happens when an entire stack goes variable at once.

A forecastable invoice has become a renewal argument all by itself.


Two Ways Renewals Break

Green's line is worth unpacking, because the renewal fails in two opposite directions.

The sales motion is genuinely easy. "Only pay for what you use" removes objections, carries no shelfware risk, and asks for almost no commitment. First signature, no problem.

Bill shock is the failure mode where the product worked. The customer adopted enthusiastically, usage spiked, and the invoice landed at 3x expectations. This is the Cursor pattern. Users didn't blame themselves for using the product more. They blamed Cursor for the surprise.

Perceived waste is the failure mode where it didn't. The team used the product lightly and the invoice stayed low. That sounds fine until budget review, when the low invoice tells the CFO the product isn't essential and it becomes the first line item cut. Low usage reads as low value, even when the product did exactly what was needed.

Anyone who has run a renewal meeting knows both versions. In one, procurement opens with the invoice history and asks about the 3x March bill. In the other, your own usage stats argue for the cancellation.

Green reports an executive revenue leader whose gross revenue retention improved after switching from consumption back to per-seat pricing. That's an observed result, not a theoretical argument, and it's the single most inconvenient data point for anyone insisting seats are dead.


Your Best Customers Get the Biggest Invoices

The customer who uses the product most gets the highest bill. The power user who builds workflows around your tool pays more than the person who logs in once a month.

Enterprise buyers notice, because they can't forecast costs for a product they plan to use more over time. A successful deployment means a bigger bill. The incentive structure quietly pushes customers to limit adoption rather than expand it.

Companies moving from pure usage to hybrid aren't retreating from value-based pricing. They're giving customers what they ask for at the negotiating table: a known number that still reflects what they got.


Per-Token Pricing Frames You as a Commodity

Vin Vashishta makes the point most pricing discussions skip. Price per token, per API call, or per credit, and you have framed your product as infrastructure. Infrastructure gets compared on price per unit. The customer shops for cheaper tokens the way they shop for cheaper compute, and when a cheaper alternative appears, switching is simple arithmetic.

An outcome gets compared on results. The customer asks whether the problem got solved, not how many tokens it took, and switching means retraining, reintegrating, and accepting the risk of worse results.

Pricing FrameCustomer EvaluatesSwitching Trigger
Per token / per creditCost per unitCheaper alternative appears
Per resolution / per outcomeProblem solved or notResults decline
Plan with guardrailsTotal value vs. total costBudget review, competitive feature gap

Model economics sharpen this. Inference costs drop roughly 10x every 18 months. Price per token and every cost reduction presses directly on your revenue. Price by plan and the same reduction improves your margin: the customer sees the same bill, and you keep the efficiency gain.


Where the Market Actually Landed

Hybrid pricing adoption grew from 27% to 41% in 2025 and is projected to reach 61% by end of 2026, while pure seat-based pricing fell from 21% to 15% over the same period. The market isn't choosing between seats and usage. It's choosing both.

You could read that adoption curve as vendors hedging rather than customers voting. Some of it probably is. But the retention data points the same direction, and retention is the harder number to argue with.

HubSpot moved its Breeze AI agents from $1.00 per conversation to $0.50 per resolved conversation in April 2026, shifting from activity to outcomes. Intercom charges $0.99 per resolution but bundles it with per-seat helpdesk pricing: a predictable base with a variable outcome layer on top. Notion folded AI into its Business tier entirely in May 2025, absorbing AI cost into a single forecastable line item.

Each move went toward more predictability, not less.


Plans With Usage Guardrails

The companies reporting the best retention numbers share a structure.

ComponentPurpose
Base plan ($X/month)Predictable cost the CFO can budget
Included usage allowanceCovers 80-90% of customers without overage
Published overage rateClear per-unit cost above the allowance
Hard spending capMaximum monthly bill, no surprises

Notion AI at $10/user/month with included responses reached 90% renewal rates and 91% monthly active retention. Jasper and Fireflies.ai use flat-rate add-ons or seat-based AI plans to avoid billing anxiety. Anthropic sells team seats from $25/seat/month with usage limits per tier, keeping API-rate billing separate from the workspace product.

It's worth being precise about what Cursor's backlash was actually about. The anger came from removing the plan, not from adding usage. Users accepted limits. They rejected bills they couldn't predict a week in advance.


The Number Under Every Pricing Debate

Every argument about seats versus usage versus outcomes traces back to the same gap: most companies don't know their per-customer cost-to-serve.

A founder choosing between $49 flat and $0.02 per request is comparing two guesses without that number. You can't set a confident flat price without knowing your cost floor. You can't set a confident usage rate without knowing how much cost varies across customers. And you can't price outcomes without knowing what an outcome costs to produce.

OpenAI, Anthropic, and Vercel all surface granular usage dashboards so teams can model ROI, because cost visibility is the prerequisite to any pricing structure. The 78% of IT leaders reporting unexpected AI charges have a visibility problem more than a pricing model problem. The same holds on the vendor side: every model swap and prompt change shifts the cost underneath the price, whether or not anyone measures it.

Some customer segments are comfortably profitable under flat pricing. Others need guardrails. A few might justify outcome-based pricing where the outcome is measurable and consistent. Which is which only shows up in the numbers, customer by customer.


Pricing Teaches Your Customers How to Behave

The standard framing treats pricing as revenue capture. It's also a behavior decision. Usage pricing teaches customers to limit consumption. Plan pricing teaches them to maximize adoption within their tier. Outcome pricing teaches them to demand measurable results.

The Copilot story showed what happens when a company can't see its own cost distribution. The Cursor story showed what happens when customers can't see theirs. Same root in both cases: not enough visibility into what each customer costs and what each customer gets.

If you're designing a hybrid, the allowances, overage rates, and caps should come from your own cost distribution rather than from a competitor's pricing page. Bear Lumen gives you that per-customer view, so the guardrails you publish match the costs you actually carry.

Share this article