Token-Based Pricing in Legal AI: What It Gets Right, and Where It Gets Risky
We’ve been hearing from a lot of firms this year: the AI tool they signed up for is sending invoices three times the size of the last one, without a corresponding increase in legal work to justify the cost. What has changed is the pricing and the plumbing underneath it. Many vendors are moving contracts off flat, per-seat subscriptions and onto usage pricing billed per token, so usage that used to be buried in a flat fee now shows up as its own line item. The models doing the work have also been upgraded to newer, pricier versions and that cost gets passed through. "Agentic" workflows now turn a single task into a chain of calls instead of one: an agent drafting a clause checks its own output, revises it, verifies a cross-reference, loops through that a few times before it hands back an answer that used to take one pass. Any one of those changes would nudge a bill upward on its own, and stacked together, they explain why bills are increasing rapidly.
One well-known provider has said token usage across its platform grew more than tenfold in about six months, a number that looks great on a board deck and a lot harder to explain to a client when the invoice jumps from five figures to six with no real change in headcount, matter volume, or billable hours.
Is consumption pricing the villain here? Not really, and that's what makes it interesting.
The case for paying by token
Consumption pricing didn't appear out of nowhere. Flat, unlimited-seat pricing forces a vendor to guess at average usage and price around that guess, which means light users subsidize heavy ones, and many vendors quietly cap what the product can do so nobody blows past those assumptions. This is similar to the business model of gyms. Frequent users are subsidized by optimistic users who expect to go to the gym, but never quite get out of bed. Charge for what actually gets used and that ceiling disappears. A firm running a large due diligence sweep, or pointing an agent at a thousand-document review, can just do it, without the vendor throttling the product to protect its margins or the firm being given a surprise bill.
It also ties cost to value more honestly. If an AI workflow saves a lawyer twenty hours on a matter, the firm and the client can see roughly what that work cost to produce, instead of spreading a flat subscription fee across every matter whether the tool was used or not. And it rewards good habits: a firm running tight, well-scoped workflows pays less than one that lets an agent wander.
The case against it
Most firms aren't set up to manage a variable token cost the way they manage other variable costs. Legal budgeting, timekeeping, and client billing all assume a fairly stable cost base. A pricing model that can move 3x in a quarter, tied to a metric like token count that almost nobody outside engineering fully understands, is hard to plan around. Getting a bigger bill because you did more work is one thing. Getting a bigger bill because a newer model needs more tokens to produce the same answer, or because an agent needed six attempts instead of two and nobody could see it happening, is another.
There's also a conflict of interest baked into pure token metering: the vendor makes more money when the software is less efficient, not more. An agent that takes the long way to an answer costs the client more and earns the vendor more, and nothing in the pricing corrects for that on its own. Most vendors probably aren't engineering waste on purpose, but fixing it is entirely the buyer's job, and most buyers don't have the tools to see where it's happening. That lack of transparency into planned agent workload per task is a problem all industries are dealing with when building new processes around the foundational model providers.
The price of a token isn't really the point
The more useful question isn't whether tokens are expensive, but whether the work is being routed sensibly in the first place. A lot of the "AI got expensive" story is actually a workflow design story. A trivial fix, like updating a party name or a date, gets thrown at the same general-purpose agent as a genuinely hard drafting problem, and both burn tokens at similar rates even though only one of them needed real reasoning. Firms that sort work by actual complexity, routing simple changes somewhere lighter and saving expensive, iterative agent work for matters that need it, end up with far more predictable spend than firms letting one tool touch everything. A single tool determining on its own the work path, cost, iteration needs, and validation requirements becomes the cost black box that firms used to having precise control over their profit margins can’t afford to become reliant on.
There's a structural version of that same fix, too. A lot of the token burn described earlier comes from an agent reasoning its way to an answer from scratch, drafting, checking, redrafting, because it has no fixed map of how a clause relates to the rest of the contract or to the firm's own precedent. A system built around that map, a legal ontology that already encodes how terms, clauses, and past deals connect, can often look an answer up instead of arguing its way toward one. That's the real antidote to token maxxing: not a cheaper way to run an undisciplined process, but a process that never needed five rounds of self-correction to get the answer right the first time.
Whether it's routing simple work away from expensive models or building on a fixed map of the firm's own contracts, the pattern is the same: systems that don't have to reason from scratch every time end up cheaper, no matter how the vendor bills for them. That matters more than which pricing model a vendor happens to be running this year. A firm that understands its own token consumption can negotiate metered pricing from a position of strength: usage caps, budget alerts, and a real sense of what a task should cost. A firm that doesn't understand its consumption will struggle under any pricing model, flat or metered, because the actual problem was never visible to begin with.
Before the next renewal
There are a few things worth asking any AI vendor before signing the next contract, regardless of what pricing model they're pitching. Can the tool send simple, high-volume tasks down a cheaper path than complex ones, or does everything run through the same expensive model by default? Does it already know how this contract type and this firm's precedent fit together, or does every request start the reasoning process from scratch? Can the firm see usage close to real time, or does the bill just show up at the month's end? And when usage spikes, is that because people are getting more value from the tool, or because the software has become less efficient at the same work it did last quarter?
Those questions decide whether a 3X bill is a good sign or a bad one. The pricing model printed on the contract matters a lot less than the answer to that.


