Your AI capability works, early users can see the value, and the launch date is getting close. Now the team wants a price. That is usually one decision too soon.
If you choose a number before deciding what the customer is paying for, you can end up with a commercially elegant model that feels arbitrary on the invoice. A better sequence is to define the value, choose the billable unit, design the package, validate willingness to pay, and then test whether the economics survive contact with real customer behavior.
Key takeaways
- Choose the pricing model and pricing metric before debating the price point.
- Use a billable unit that customers can understand, forecast, influence, and verify.
- Keep pricing and packaging separate: one determines how you charge, while the other determines what the customer receives.
- Run qualitative buyer research before quantitative willingness-to-pay work, so you measure a concrete offer rather than an ambiguous idea.
- Model realized price, adoption, usage, serving cost, discounting, and margin by customer segment before seeking approval.
- Treat billing, metering, sales enablement, customer communication, and dispute handling as part of the product launch.
Decide what the customer is paying for before choosing the amount
Four decisions tend to get collapsed into the word “pricing.” Separating them makes the work much easier:
| Decision | Question it answers | Example output |
|---|---|---|
| Pricing model | What economic structure will govern the purchase? | Fixed fee, access-based, usage-based, outcome-based, or hybrid |
| Pricing metric | Which unit causes the bill to change? | Seat, document, workflow run, minute, or successful resolution |
| Price point | How much will you charge for that unit or entitlement? | A subscription amount, unit price, commitment, or overage rate |
| Package | Which capabilities, allowances, controls, and service levels are included? | An entry plan, advanced plan, enterprise plan, or add-on |
A sound pricing process treats the model and metric as distinct decisions that come before the price point. Otherwise, a discussion about value quickly deteriorates into negotiation over an unsupported number.
Start by completing this sentence: “The customer should pay more when _____.” The blank should describe an increase in value, not merely an increase in your compute bill.
For a customer-support agent, the answer might be a successfully handled query. One concrete definition is a query resolved without further human help, which turns the general idea of an outcome into a countable event. For a copilot used across many loosely defined tasks, access-based pricing may be easier to defend because the value is dispersed and difficult to attribute to a single event. For document processing, a processed document is easy to count, but a verified output may be closer to the value the customer actually receives.
This is why infrastructure cost should inform your margin model without automatically becoming the customer-facing metric. Tokens, model calls, and compute time may be measurable, but measurability alone does not make them meaningful to a buyer. Charging for failed attempts can also create the wrong incentive: the customer pays more when the product performs more work, even if that work does not produce more value.
Test every candidate metric against these questions:
- Value alignment: When the count rises, does the customer’s expected benefit usually rise with it?
- Auditability: Can the customer verify why an event was billed without relying on your interpretation alone?
- Controllability: Can the customer influence consumption, or can your product create an unexpectedly large bill on its own?
- Predictability: Can a buyer make a credible budget before using the product at scale?
- Operational clarity: Can product telemetry, billing systems, contracts, sales materials, and support teams use the same definition?
- Economic durability: Does revenue continue to make sense when model costs, usage patterns, or product performance change?
No metric will be perfect on every dimension. The useful output is an explicit tradeoff, not a claim that the tradeoff does not exist. My rule is simple: if the buyer and the billing system would count the event differently, the metric is not ready.
Turn the value metric into a package customers can understand
The pricing model determines how money moves. Packaging determines what the customer can buy. A good metric can still fail inside a confusing package, especially when several AI capabilities each introduce their own unit, allowance, and exception.
Choose the model by examining the nature of the value and the quality of the measurement:
- Fixed-fee or access-based pricing fits an entitlement whose value is broad, recurring, or hard to attribute to individual events. It improves budget predictability, but light users may feel they are subsidizing heavy users.
- Usage-based pricing fits a consumption unit customers can observe and control. It scales with activity, but it can make experimentation feel expensive and can disconnect the bill from successful results.
- Outcome-based pricing fits a result that is valuable, attributable, and unambiguous. It creates strong value alignment, but disputes appear quickly when success has exceptions or depends on factors outside the product’s control.
- Hybrid pricing combines predictable access with an allowance, overage, capability tier, or outcome charge. It can balance predictability and expansion, but every additional mechanism raises the explanation and implementation burden.
Outcome-based pricing deserves particular care in AI. Calling something an outcome does not make it one. The metric needs an operational contract that answers:
- What exact event starts and completes the outcome?
- Which failures, retries, duplicates, test events, and customer-abandoned interactions are excluded?
- What happens when a person later reverses or corrects the AI’s work?
- Which system is the record of truth when product analytics and billing disagree?
- What evidence will the customer see on the invoice or usage screen?
- How are credits, disputes, and measurement errors handled?
- How will customers be notified if the definition changes?
These details are not billing trivia. A small change in the definition of a metric can materially change how customers experience the value. If an AI support agent closes a conversation, for example, the team still has to decide whether a reopened conversation, a delayed escalation, or a duplicate contact changes the count.
Once the metric is defensible, design the package around buyer progression rather than an inventory of features. For each package, write down:
- The initial job: the smallest complete problem a customer can solve after purchasing.
- The expansion reason: the meaningful change that should cause an upgrade, such as greater volume, broader automation, advanced control, or stronger governance.
- The allowance: what is included, what is metered, and what happens at the boundary.
- The predictability controls: usage visibility, alerts, caps, approval rules, commitments, or negotiated ceilings.
- The enterprise requirements: security, auditability, permissions, support, data controls, and service terms that affect adoption in larger organizations.
Then ask a person unfamiliar with the design to explain the offer back to you. They should be able to say what they get, what causes the bill to increase, and which package they would need. If that requires a long verbal rescue from the product manager, the package is carrying too much complexity.
Also test the catalog as a system, not just one SKU at a time. A separate metric can look rational for every AI capability while leaving the customer with several unrelated meters on one platform. As an AI portfolio expands, the problem shifts from optimizing individual offers to creating a coherent pricing system across products and outcomes. Sometimes the right answer is a common allowance or platform commitment. Sometimes it is separate add-ons. The important point is to make that choice deliberately.
Validate willingness to pay, then model the real business
Do not begin buyer research by asking, “How much would you pay for AI?” The buyer has not been given a stable thing to evaluate, and “AI” is not a billable unit.
Start with qualitative conversations about the buyer’s operating reality. Useful prompts include:
- What changes in the business when this problem is solved well?
- How is that value measured or defended internally?
- Which team owns the budget, and which team experiences the benefit?
- How is the problem funded today: software, labor, outsourcing, or absorbed inefficiency?
- Which unit would feel intuitive on an invoice?
- Which charging behavior would feel unfair even if the total amount were acceptable?
- How much predictability does procurement require?
- Which failure cases would make the buyer challenge a billed event?
The buyer’s job is to explain goals, constraints, expectations, and purchase behavior. Your job is to translate those signals into a model and metric. Buyers should not be expected to design the pricing architecture for you.
Only then should you quantify willingness to pay. The offer in the questionnaire must specify the capability, model, metric, success definition, package, and relevant limits. Methods such as Gabor-Granger and Van Westendorp can help explore how purchase willingness changes across possible prices, but a precise-looking output cannot compensate for an ambiguous offer.
Willingness-to-pay work gives you evidence for a pricing range and exposes price sensitivity. It does not produce a perfect price. It also measures stated intent rather than completed purchases, so the result must be combined with operating data and commercial judgment.
Build the financial model in this order:
- Move from list price to realized price. Include expected discounting, credits, commitments, channel effects, and contract terms. A healthy-looking list price is irrelevant if the average transaction lands somewhere else.
- Estimate billable units by segment. Use product or beta data where available. Do not rely only on a blended average that hides the difference between small, mid-market, and large customers.
- Translate units into customer revenue. Realized unit price multiplied by expected billable units, plus any fixed component, gives you the commercial shape of an account.
- Subtract the costs that move with usage. Model inference, third-party services, human review, support burden, and any other serving cost that increases as customers consume more.
- Apply adoption and attach assumptions. A high unit price does not help if few eligible customers buy. A low price does not guarantee adoption if the package, procurement path, or value story is the real barrier.
- Run conservative, expected, and upside cases. Vary adoption, usage, discounting, serving cost, and performance rather than changing only the price.
This model should connect the recommendation to adoption, revenue, and margin. The relevant inputs include segment differences, discounting, purchase friction, sales capacity, competitor pricing, usage, attach rates, and margin expectations. The goal is not to make uncertainty disappear. It is to show which assumptions determine the result.
Before sign-off, document what would make the recommendation wrong. Examples include usage concentrating far above the expected range, serving costs declining more slowly than assumed, heavy discounting, weak attach in a priority segment, or frequent disagreement over billed outcomes. Assign an owner and a measurable signal to each assumption. That turns the pricing decision into something the company can inspect after launch.
The approval request should cover the whole commercial design: model, metric, success definition, packages, price points, discount rules, forecast assumptions, migration approach, launch dependencies, and conditions for review. Approving only the headline number leaves the consequential decisions unresolved.
Launch the billing experience and keep the system under review
A pricing decision is not finished when an executive approves it. It becomes real when a customer can predict the charge, a salesperson can explain it, the product can meter it, the billing system can invoice it, and support can resolve a dispute.
Use this launch checklist:
- Event specification: Give every billable event a precise definition, inclusion and exclusion rules, required fields, and system of record.
- Metering quality: Reconcile product events with billing records and define what happens when data is delayed, duplicated, missing, or corrected.
- Customer visibility: Show usage, allowances, forecasted charges, and relevant alerts before the invoice arrives.
- Sales enablement: Give the go-to-market team examples that explain the model, the value metric, package boundaries, and common objections in the same language customers will see.
- Value demonstration: Connect the billed unit to a customer-visible result so account reviews do not depend on abstract AI activity.
- Support and dispute handling: Define who investigates contested events, what evidence they use, who can issue a credit, and how recurring measurement problems reach the product team.
- Analytics: Instrument attach, activation, consumption, realized price, discounts, serving cost, margin, expansion, contraction, disputes, and churn by segment.
- Ownership: Name the person responsible for monitoring the assumptions and convening a review when the evidence changes.
Sales education, customer rollout, ROI tooling, and interpretation of results are substantial parts of operationalizing a pricing decision. If those workstreams start after the pricing meeting, the launch plan is already incomplete.
For an existing product, treat migration as a commercial and contractual change, not a website edit. Before announcing anything, have finance and legal confirm contract terms, notice requirements, renewal treatment, credits, tax behavior, and billing-system readiness. Decide how current customers move, what happens to prior commitments, and how exceptions will be governed. An improvised exception process can quietly replace the published pricing model with a collection of one-off deals.
After launch, diagnose the layer that is failing before changing the price:
- Frequent disputes may indicate a weak metric definition or unreliable measurement.
- Strong adoption with poor realized margin may point to discounting, serving cost, allowances, or an underpriced unit.
- Low adoption despite a recognized need may indicate packaging, procurement friction, trust, or positioning rather than a price that is simply too high.
- Customers unable to forecast their bills may need better controls, commitments, alerts, or a more predictable model.
- A valuable new capability that does not map to the existing unit may signal that the metric or model has stopped representing the product.
- Individually sensible offers that create a confusing catalog may require a platform-level packaging redesign.
AI pricing is a living product system because the product, its cost structure, and the way customers obtain value can all change. Pricing therefore needs periodic review as capabilities and product breadth evolve. That does not mean changing the price whenever a metric moves. It means preserving the ability to distinguish a price-point problem from a metric, package, measurement, or value problem.
Before your next pricing meeting, write down the value event, the billable unit, the package boundary, the customer’s method for verifying a charge, and the assumptions that make the economics work. If the team cannot agree on those points, do not debate the number yet. Resolve the product decision underneath it.
References








