AI agents grounded in structured knowledge of your business are cheaper to run for three plain reasons. They send the model a small amount of the right information instead of a large amount of possibly relevant text. They get the answer right first time instead of searching again and again. And they hand repeatable steps to ordinary code rather than paying a model to reason through them on every run. With AI now billed by use rather than by seat, the cost of an agent is driven by how much it reads and how often it retries. That makes your AI bill a knowledge problem as much as a pricing one.
Seats became meters
For most of the last few years, business AI was priced like software: a fee per user per month, used as much or as little as you liked. Through 2026 that model gave way to metering, and the pace picked up over the summer.
Salesforce published a rate card billing agent actions in credits. GitHub moved Copilot to AI credits. Microsoft priced Business Central agents at roughly 50 credits per invoice processed plus 5 per line, and announced that usage billing would be switched on by default for new Microsoft 365 Copilot Business licences from 2 November 2026. HubSpot began charging per resolved customer conversation.
The details differ, but the direction is the same, and it is the shift I flagged in headless is the new mobile-first. You are no longer paying for access to AI. You are paying for every unit of work it does. Under per-seat pricing, an inefficient agent was a performance problem. Under metering, it is a line on the invoice, and it grows with every task.
Cheaper tokens, bigger bills
The strange thing about 2026 is that the unit price of AI fell while budgets overran.
Ramp's index of effective token prices paid by businesses fell to $0.68 per million tokens in September, from $1.15 in March (Ramp, September 2026), and at least one frontier provider launched its newest model 20% cheaper than the last. Yet Futurum's survey of 1,636 organisations found 46.9% over their AI budget (Futurum, 10 September 2026), and Gartner has forecast that inference cost per agentic workflow will rise more than fivefold by 2028 (Gartner, 17 August 2026).
The explanation is volume. A chatbot answers a question with one call. An agent completing a task might plan, search, read, check, search again, draft, verify and summarise, each step a call, each call carrying context. Lower prices do not help much if each task uses far more tokens than it needs.
Where the tokens actually go
When you trace an expensive agent, the tokens usually go to the same five places.
Oversized context: every call carries pages of loosely relevant text, because nobody was sure which passage mattered. Caching softens this for text that never changes, but not for context rebuilt from fresh searches on every call. Repeated searching: the agent is unsure, so it rephrases and searches again, and again. Long conversation histories: the whole transcript is resent with each step. Facts recomputed on every run: the agent works out which customer an email is about, or which contract applies, from scratch each time. And verification loops that re-check work the agent could have got right first time from a reliable source.
Every one of these is a symptom of the same underlying problem. The agent does not know your business, so it pays to rediscover it on every task. I describe how that shows up as wrong answers in why your RAG system keeps missing what is in your documents. It shows up on the invoice too.
How grounding cuts the bill
A grounded agent works from a structured, current picture of the facts it needs: resolved customers, suppliers and contracts, their relationships, and where each fact came from. That changes the economics of most of the five.
Context shrinks, because the agent can look up the three facts it needs rather than reading thirty pages that might contain them. Searching drops, because a lookup against a resolved record either finds the answer or clearly does not, instead of returning a different handful of passages each time. Facts stop being recomputed, because which customer an email concerns was resolved once, when the email arrived, and recorded in an entity registry. And verification gets cheaper, because an answer built from sourced facts can be checked against its sources rather than re-derived.
That work is not free. Resolving records costs tokens when data arrives, and the layer needs maintaining. It pays back because each fact is resolved once and then read on every task.
The research points the same way. For broad questions across a collection of documents, Microsoft found that answering from graph summaries used between 26% and 97% fewer tokens than map-reduce summarisation of all the source text, though not compared with vector search (Microsoft Research, 2024). Long contexts also degrade recall even on simple tasks (Chroma, "Context Rot", July 2025), so the smaller context is usually the more accurate one as well. I compare the approaches in GraphRAG vs vector RAG.
Build the tool, not the process
The biggest saving is often the least glamorous: stop paying a model to do the same thing the same way on every run.
Where the path through a task is known, a workflow beats an agent, a point I make in token budgets and AI FinOps. The new twist is who builds the workflow. Among finance and operations practitioners, a clear rule has emerged over the last year: use AI to build the tool, not to run the process. If a step is repeatable and rule-based, such as matching an invoice to a purchase order, checking a quote against a price list, or routing a request, write ordinary code for it. Use AI to design and write that code quickly, and to handle the exceptions and judgement calls the code cannot. MXD Process, a US maker of industrial mixing equipment, reported building a quoting tool this way in eight weeks, cutting quote turnaround from five days to a day and a half (Eric Seiberling, LinkedIn, August 2026).
Deterministic code costs almost nothing per run, gives the same answer every time, and is easy to test. A model should be doing the parts that need judgement, not the parts that need arithmetic.
Cheaper and more accurate are the same fix
It is tempting to treat AI cost and AI quality as a trade-off: spend more for better answers, or economise and accept worse ones. For most business agents, they move together.
The things that make an agent expensive (oversized context, repeated searching and recomputed facts) are the same things that make it unreliable. Fix the knowledge it works from and both improve together. A grounded agent also needs less raw model power for routine work, so simpler tasks can be routed to smaller, cheaper models, keeping the frontier models for the genuinely hard cases. That flexibility has a second benefit: it reduces your dependence on any single model provider, a risk I wrote about in when your best AI model can vanish overnight.
This is the economic half of the argument I make in your AI does not need a bigger model, it needs to know your business. Better knowledge is not only more accurate. It is cheaper.
How to find your savings
Start by measuring the right thing. Cost per call tells you almost nothing. Cost per completed task tells you what a unit of work actually costs, and lets you compare an agent with the person or process it supports.
Then trace a sample of expensive tasks end to end. How much of each call's context was actually used? How many searches did it take to find an answer? Which facts were recomputed that could have been looked up? Which steps are repeatable enough to become code? The answers usually point to a handful of changes worth far more than any discount you could negotiate.
Finally, put budgets and enforced spending caps in place before the meters run, as I set out in token budgets and AI FinOps.
If your AI costs are climbing faster than the value you can point to, I am happy to help you find where the tokens are going, as a standalone review or as part of a Knowledge Audit. Let's talk.
Related: Uber burned through its token budget by April. Your business will be next · Your AI does not need a bigger model. It needs to know your business · GraphRAG vs vector RAG. Which one does your business need?