tutorialsAugust 8, 2026

Revenue Recognition for AI and Token-Based Pricing Under ASC 606

A technical walkthrough of how the ASC 606 five-step model applies to token and credit-based AI pricing: variable consideration, breakage, output-method progress measurement, true-ups, and contract modifications.

R

RevExOS

Q2C Consulting

Revenue Recognition for AI and Token-Based Pricing Under ASC 606

A customer buys 10 million tokens for $500, prepaid, expiring in 12 months. They use 6.2 million tokens. They get billed an overage true-up in month 4 when a batch job blows through their monthly allocation. Their plan changes in month 7 from a fixed token block to a committed-use tier with a different rate card. None of this is a subscription in any sense ASC 606's drafters had in mind when the standard was framing "distinct services transferred over time." It's metered consumption of a fungible unit, sold in prepaid blocks, with volatile usage and a real chance that a meaningful chunk of what was paid for is never consumed.

This isn't a novel accounting problem: prepaid minutes, gift cards, and telecom data plans have used variants of this model for years. What's new is the scale of the estimation uncertainty. A telecom data plan has usage bounded by human behavior; an AI agent workload can consume 40x its typical token volume in a single session because of a retry loop, a longer-than-usual output, or a multi-step agentic chain nobody scoped for at signing. That volatility is what makes token-based pricing a genuinely hard ASC 606 application, not a relabeling of "usage-based SaaS."

This post assumes you already know the five-step model; for the primer, see our ASC 606 five-step explainer. This is the token-specific application: what changes, where the judgment calls live, and where controllers get it wrong. It builds on the broader usage-based SaaS RevRec gap mapped in our post on SaaS pricing complexity; AI token billing is the sharpest edge case inside that category, since usage variance is structural rather than incidental. For the metering mechanics behind these numbers, see how AI companies bill for token usage.


The Five Steps, Applied to Token Consumption

Step 1: Identify the Contract

The question that actually matters: do you have one contract, or a series of implicit micro-contracts every time a customer tops up a credit balance? A customer with a signed master services agreement and an evergreen self-serve top-up mechanism typically still has one contract: the MSA establishes the enforceable rights, and each top-up fulfills an existing promise to stand ready to deliver capacity. If self-serve customers buy credit packs with no umbrella agreement, each purchase is its own contract with its own collectibility assessment. This matters downstream because it determines whether usage across purchases can be pooled for breakage estimation (below) or must be tracked purchase by purchase.

Step 2: Identify the Performance Obligation

This is where companies get the token application wrong by analogy to software licensing. The performance obligation is not "access to the AI model." It's the obligation to process the customer's input and deliver output on demand, consuming their token balance as you do it. "Access to a platform" is a stand-ready obligation recognized ratably over a term, the way seat-based SaaS works. "Processing tokens on demand" is a series of distinct, substantially similar services delivered over time as consumed, recognized under the usage allocation exception in Step 5.

Most token arrangements have a single performance obligation: processing capacity, metered in tokens. It becomes multi-obligation with bundled offerings, common in this market: a monthly platform fee including a fixed token allocation, plus fine-tuning or dedicated-capacity add-ons, plus premium support. Each is likely distinct: the platform fee is stand-ready and ratable; the included allocation is usage-based; fine-tuning sold as a discrete project is likely point-in-time. Bundling all of this under one blended price without decomposing it overstates revenue at signing if any part of the fee is recognized upfront against work not yet done, or understates it later if the token allocation is buried inside a ratable platform fee it doesn't belong in.

Step 3: Determine the Transaction Price - Variable Consideration

Token-based pricing is close to a textbook example of variable consideration under ASC 606-10-32-5 through 32-14. The total transaction price is not known at contract inception because it depends on future, unknown usage volume, which triggers the estimation and constraint requirements.

Two estimation methods are available: expected value (a probability-weighted sum across possible outcomes) and most likely amount (the single most probable outcome). Narrow-band committed-use contracts can often defend most likely amount. Wide, unpredictable variance calls for expected value, and this is where AI workloads diverge from prior usage-based categories: a customer running deterministic batch jobs has stable consumption, while one running agentic workflows, where the model's own output determines how many subsequent calls, retries, or tool-invocation loops run, can vary by an order of magnitude month to month with no seasonal anchor. Document that difference explicitly rather than reusing a generic usage-based policy inherited from a prior SaaS engagement.

The constraint is the operative control: include variable consideration only to the extent it's probable a significant reversal won't occur when uncertainty resolves. Apply it to unbilled, in-period usage not yet metered, not to usage already metered, which is a fact pattern, not an estimate. It bites hardest on end-of-period accruals for unreconciled usage, and on rebate tiers tied to a future threshold ("rate drops to $0.008 per 1,000 tokens above 50 million a month"). If a customer is trending toward that threshold without reliably crossing it, hold the price at the higher, unrebated rate until crossing is highly probable.

A common failure mode: companies that reconcile metering data 5-10 days after period end default to a straight-line run-rate off the first three weeks of the month. For agentic or bursty workloads that assumption is often wrong by a wide margin, and foreseeably so.

Step 4: Allocate the Transaction Price

Where multiple performance obligations exist (platform fee, included allocation, overage tier, professional services), allocate based on relative standalone selling price. Your published rate card, say $0.015 per 1,000 output tokens, is usually a reasonable SSP anchor for the token component, since it's directly observable and applied consistently. It's harder to establish SSP for a bundled included allocation ("$2,000 a month includes 5 million tokens"), since the implied per-token rate is typically discounted relative to the published overage rate, and you need a defensible basis for how much of the $2,000 is platform access versus prepaid capacity. The residual approach can apply here.

Step 5: Recognize Revenue - Output Method vs. Input Method

ASC 606-10-25-31 through 25-32 gives two families of methods for measuring progress: output methods, which measure value transferred directly (units delivered, milestones reached), and input methods, which measure effort or inputs as a proxy (costs incurred, resources consumed).

For token-based pricing, the output method is almost always correct, and easy to defend: tokens processed and delivered is a direct, objectively measurable output that corresponds precisely to value transferred. Recognized revenue in a period equals tokens consumed, priced per your rate card, at the applicable tier.

The input method fails for a specific reason: your compute cost to serve a token (GPU-seconds, model-call cost, retries) doesn't correspond to value delivered or price charged. Two customers running an identical prompt through an identical model consume the same billable tokens and generate the same revenue, regardless of whether one request hit a cache or required an infrastructure-side retry. Using "compute cost incurred" as the progress measure would make recognized revenue move with your cost structure instead of what the customer received. The one exception is a bundled fixed fee with no natural output measure, such as a flat fee for effectively unlimited usage within soft limits.


Breakage: The Underdiscussed Problem in Prepaid Token Packs

Prepaid, non-refundable credit or token packs create a contract liability (deferred revenue) at purchase, because you've received consideration for an obligation not yet satisfied. As the customer draws down the balance, you recognize revenue proportional to consumption under the output method above. The complication is what happens to the portion never consumed before expiration, forfeiture, or account closure: the breakage.

ASC 606-10-55-46 through 55-49 governs breakage for prepaid stored-value instruments, with two paths depending on whether you expect to be entitled to it. If breakage is probable and estimable, recognize it in proportion to the pattern of rights the customer has exercised, pulling a pro-rata slice of the expected unused balance into revenue as tokens are used, not just at expiration: if historical data shows customers typically use 80% of a purchased pack before expiration, recognize the full purchase price ratably as tokens are consumed, scaled so that consuming the expected-to-be-used 80% yields 100% of revenue recognized, then true up against actual expiration data as it comes in. If you don't expect to be entitled to breakage (a refund obligation, or insufficient historical data), recognize it only when the likelihood of the customer exercising remaining rights becomes remote, typically at contractual expiration.

Three things make this worth a dedicated line in your revenue policy:

  • You almost certainly have breakage, and probably more than you think. Prepaid packs with 6-12 month expiration windows routinely see 15-35% non-utilization, especially packs bought speculatively ahead of a project that gets delayed. Without a measured rate you have no basis for the probable-and-estimable path, and default to expiration-triggered recognition, deferring revenue unnecessarily.
  • Estimation requires a large, stable population. ASC 606-10-55-48 requires a portfolio large enough to reasonably predict the pattern. Forty enterprise customers on customized packs is a weaker basis than thousands of self-serve customers on standardized sizes; a small or heterogeneous base often means expiration-triggered recognition is the more defensible answer.
  • Unclaimed property law can override the conclusion operationally. Some jurisdictions treat unused prepaid balances as escheatable after a dormancy period. That doesn't change the ASC 606 answer, but breakage assumptions need a legal and tax cross-check too.

Track breakage separately from ordinary consumption-based recognition. Auditors will ask for the basis of your rate, the population it's drawn from, and a rollforward of estimated versus actual breakage by cohort. Without that on demand, you have a placeholder, not a policy.


True-Ups and Overage: Timing the "Extra" Revenue

Committed-use contracts are common in enterprise AI pricing: the customer commits to a minimum monthly spend, say $50,000, entitling them to a corresponding token allocation at a committed rate, and usage above that bills at an overage rate, often reconciled monthly via a true-up.

In most structures the committed minimum and the overage are the same performance obligation, processing capacity metered in tokens, with a tiered price: the committed tranche at the discounted rate, consumption above it at the overage rate. Because it's a single obligation, the output method applies uniformly: recognize revenue as tokens are consumed, at the rate the applicable tranche falls into. A customer who commits to 10 million tokens a month at $0.01 and consumes 14 million recognizes revenue at $0.01 for the first 10 million and at the overage rate, say $0.014, for the remaining 4 million, in the period usage occurs, not deferred to a quarterly true-up.

The timing question that trips people up: invoicing cadence and recognition cadence are not the same thing. If a contract bills committed minimums monthly but reconciles overage quarterly, the overage revenue still needs recognizing in the month it was consumed, accrued as an unbilled receivable (contract asset) until the quarterly invoice catches up. Waiting for the true-up invoice because "that's when we know the final number" is a metering-timeliness problem, not a reason to defer recognition.

One more distinction worth making explicit: an unused committed minimum, a take-or-pay arrangement where the customer commits to $50,000 a month, uses only $30,000, but owes the full $50,000, is not a breakage question. Breakage applies to consideration already paid for a right the customer may not exercise. Take-or-pay minimums are owed regardless of exercise, a standing-ready obligation recognized as earned through the passage of the period rather than requiring usage to occur.


Contract Modifications: Plan and Allocation Changes Mid-Term

Token allocation and plan changes mid-contract are routine: a customer upgrades from 5 million tokens a month to 20 million in month 4, or moves from pay-as-you-go to a committed-use tier with a new rate. Under ASC 606-10-25-10 through 25-13, every modification requires classification into one of three treatments.

Separate contract: the modification adds distinct goods or services at a price reflecting their standalone selling price. A customer adding a genuinely separate service, like a fine-tuning engagement priced at its SSP, alongside an existing token plan, is a separate contract, accounted for independently with no reallocation of the original arrangement.

Termination and creation of a new contract, applied prospectively: the remaining goods or services are distinct from those already provided, even if pricing isn't purely at standalone value. A plan upgrade from 5 million to 20 million tokens a month typically qualifies here. You don't restate revenue already recognized under the old plan; you apply the new rate to consumption from the modification date forward. This is the most common treatment for plan-tier changes and the easiest to implement, since it never touches prior-period numbers.

Cumulative catch-up, treated as part of the original contract: the remaining goods or services are not distinct from those already delivered, most relevant when a modification changes the price for an obligation already partially satisfied and not separable, such as a mid-quarter renegotiation that resets pricing for the whole quarter retroactively. Here you adjust revenue in the period of modification to reflect the cumulative effect on obligations already partially satisfied.

The practical test for a controller: does the new rate apply only to tokens consumed going forward (prospective, no restatement), or does it retroactively reprice consumption that already happened (cumulative catch-up, restating the current period)? Sales teams negotiating retroactive discounts to save a renewal frequently don't realize they've created a catch-up adjustment instead of a clean prospective change, exactly the detail that needs to route through finance before it's promised to the customer.


What This Requires Operationally

None of the five steps above are exotic in isolation. What makes token-based pricing hard is that all five interact continuously, every billing period, across potentially thousands of customers with unpredictable usage: variable consideration estimates need updating as metering data lands, breakage rates need cohort-level tracking against actuals, overage needs accruing in the period of consumption regardless of invoice timing, and every plan change needs correct classification before it hits the ledger. This works in spreadsheets at a few dozen customers. It breaks down well before a few hundred, and fastest where the stakes are highest: fast-scaling AI companies whose usage volatility is the whole reason this is hard.

If your token billing and revenue recognition process is still reconciled manually at month-end, or your breakage assumptions have never been tested against actual cohort data, that's a fixable Q2C infrastructure problem, not an argument for simplifying your pricing. Reach out for a Q2C audit and we'll map exactly where your process needs to close the gap.

Stop losing revenue between quote and cash.

Get a free Q2C audit and see exactly where your AR, invoicing, and collections process is leaking money.

Get a free audit