The Price Is Real. The Bill Is Somewhere Else
Two of the cheapest numbers in artificial intelligence expire on the thirty first of December. The bill that replaces them is already published, and almost nobody has read it.
By Katie Delaney · 2026-09-07 · 11 min read
What ai consulting services should price, and what they ignore#
the increase in Gemini 3.8 Flash input and output pricing scheduled for 1 January 2027
Every buyer of ai consulting services has been handed the same seductive sentence this year: the models got cheaper. It is true, and it is nowhere near the whole truth. A per-token rate is a sticker on a windscreen. The bill that arrives at the end of the month is assembled from meters the sticker never mentions.
Start with the number that has a date attached. Google's own announcement of Gemini 3.8 Flash on 2 September put the model at seventy five cents per million input tokens and $3.75 per million output tokens. The Gemini API pricing page spells out what the announcement soft-pedals: those rates run through December 31, 2026, and on January 1, 2027 they become $1.50 and $7.50.
That is a doubling, printed in advance, in the vendor's own documentation. Any organisation building a 2027 budget line off today's rate is budgeting at half the real cost, and will discover it in the first invoice of the new year. Good ai consulting services diary that date before they quote a figure.
There is a second trap in the same document, and it only bites at scale. Flash pricing is flat whatever the prompt length, but the Pro tier is not: Gemini 2.5 Pro charges $1.25 per million for prompts up to 200,000 tokens and $2.50 above it. Feed a long document into the wrong tier and the rate doubles mid-workload, silently, with no line on the invoice explaining why.
The quiet meter running under the floorboards#
The same pricing page carries a charge Anthropic does not levy at all. Cached input on Gemini 3.8 Flash is $0.075 per million tokens, rising to $0.15 in January, but caching also bills storage at $0.50 per million tokens per hour, doubling to $1.00. Cache something large and leave it warm, and the meter runs whether or not a single request touches it.
This is the practical difference between a demo and a deployment, and it is precisely the gap a boutique agency is hired to close. The demo runs for ten minutes. The deployment runs for a year. Most ai consulting services quote the demo and invoice the deployment, which is how a sensible budget turns feral.
Nothing here requires a view on which laboratory is ahead. It requires a patient prowl through four pricing pages and the discipline to write down what they say. That is unfashionable work, and it is the work that decides whether an artificial intelligence programme survives its second year.
Where the money goes to ground in the undergrowth#
Prompt caching is the single largest lever on a real bill, and the least understood. Anthropic's prompt caching documentation sets out the multipliers plainly: five minute cache writes cost 1.25 times base input, one hour writes cost 2 times, and cache reads cost 0.1 times. Then it names an exception that changes the arithmetic entirely.
Cache hits on Claude Fable 5.1 and Mythos 5.1 are priced at 0.025 times the base input price. Against the $10 per million base input rate on the published pricing page, that lands cache reads at $0.25 per million. VentureBeat reported the change as a seventy five per cent cut, from a previous dollar, and attributes that comparison to its own reporting rather than to a live vendor page.
Read that ranking slowly, because it reframes every cost estimate built on a single number. On one model, on one day, the span between the cheapest and dearest token is two hundredfold. Which of those meters your workload spends its life on is an architecture decision, and architecture is what generative ai consulting is actually for.
This is the thicket where most estimates get lost. A team compares two headline input rates, picks the lower, and never notices that its workload lives almost entirely on cache reads and reasoning tokens, where the ranking reverses. Careful ai consulting services model the mix before they model the price.
The silent failure with no scent to follow#
The caching documentation also carries the trap that costs the most and announces itself the least. There is a minimum cacheable prefix: 512 tokens on the newest models, rising to 4,096 on some older ones. Below it, Anthropic states, requests are processed without caching, and no error is returned.
A team can therefore ship a caching strategy, watch it silently do nothing, and pay full input rate for months while congratulating itself on the saving. The clock is unforgiving too: cache lifetime is measured from the start of the request that writes it, so a long streamed response eats its own cache window.
Gemini 3.8 Flash pricing table: "Output price (including thinking tokens)" $3.75 / 1M. thinking defaults to medium. you never see the full thoughts. you still pay output rate for them. I log thoughtsTokenCount now.
That practitioner has found the third meter. Reasoning tokens are invisible in the response and billed at the output rate, so a model that thinks harder costs more without producing a longer answer. Nobody reads that on a pricing page. Somebody finds it in a bill.
It is worth saying plainly that none of this is hidden. Every figure in this piece sits on a public vendor page. It is simply spread across four documents nobody reads end to end, which is a fair working definition of what ai consulting services are paid to do.
Why the league table cannot settle the scent#
Faced with this, the instinct is to reach for a benchmark and let the leaderboard choose. That instinct is where a good deal of money goes quietly to ground, and there is now peer-reviewed reason to distrust it.
The Leaderboard Illusion, by Singh and colleagues at Cohere Labs and collaborating universities, examined arena-style evaluation and found the distortion is structural rather than statistical noise. Undisclosed private testing, the authors write, lets a handful of providers test multiple variants before release and retract scores at will, producing biased ratings through selective disclosure. At the extreme they identified 27 private variants tested by one provider before a single model launch.
The access gap compounds it. The paper estimates two providers received roughly 19.2 and 20.4 per cent of all arena data, while 83 open-weight models shared 29.7 per cent between them, worth relative gains of up to 112 per cent on that distribution.
None of this argues that benchmarks are worthless. It argues that a table is evidence about a table. The only evaluation that settles a purchase is the one run against your own workload, on your own data, with your own latency and cost budget attached, which is unglamorous, unpublishable and the thing that actually works.
There is a second reason to run your own trial, and in regulated categories it outranks cost. Where the data sits is now a purchasable property: Anthropic says activity data for its enterprise safeguards can be stored in the customer's own cloud account rather than the vendor's, and that it charges nothing for the capability. For a bank or a clinic, that sentence settles more procurement arguments than any benchmark ever will, and ai consulting services working in those categories should lead with it.
What to put in the budget on Monday#
Strip out the noise and the week hands a buyer four instructions, none of which require a view on which laboratory is winning. This is the part where ai consulting services earn the fee, or fail to.
First, date every price in the model. An undated rate in a spreadsheet is a wrong rate with a delay on it. Second, forecast on running cost: cache reads, cache writes, storage hours and reasoning tokens, not the advertised input figure.
Third, check the cacheable prefix before claiming a saving. Fourth, treat vendor tables as vendor marketing and run your own evaluation. Small teams commissioning ai consulting for small businesses should insist on all four in writing, because the four together are the difference between a forecast and a guess.
The trail from here is unglamorous and short. Take the four pricing pages, put every rate and every expiry date in one sheet, and mark which meter each workload actually runs on. A den built on that sheet holds. A den built on a headline rate does not, and ai consulting services that cannot produce the sheet are guessing in good clothes.

Buyers hunting boutique ai consulting firms tend to ask which model is best. It is the wrong quarry, and chasing it wastes a season. The better question, and the one a careful guide answers, is which meter your workload will live on for the next eighteen months, and what that meter costs on the first of January.
A per-token rate is a sticker on a windscreen. The bill is assembled from meters the sticker never mentions.
Vendors are, to be fair, publishing more of this than they used to. Anthropic's work on enterprise safeguards and its own disclosure of alignment and security failures both put awkward operational detail on the record. The documentation is there. Reading it is the job, and ai automation consulting that skips it is selling a sticker.
The same discipline runs through every lane folkfox reports on, whether the buyer is a bank weighing fintech marketing or a security vendor reading cybersecurity marketing. Published rates sit on the pricing page, the method sits on the AI consultancy page, and the search side of the same problem sits with SEO and GEO. None of it works without a number somebody has checked.
Frequently asked questions#
How much does an AI consultant cost?
Fees vary widely, but the larger variable is usually the running cost of the system rather than the day rate. A model that looks cheap per token can cost several times more in production once cache writes, storage hours and reasoning tokens are counted, so ask any adviser to forecast the monthly bill, not the rate.
What is ai consulting services in practice?
Good ai consulting services turn a capability into a costed, governed system: choosing models against your own workload, designing the caching and context architecture that determines the bill, evaluating output quality independently of vendor benchmarks, and writing the whole thing down so a finance team can forecast it.
Why do model prices have an expiry date?
Introductory pricing is a launch tactic. Google publishes Gemini 3.8 Flash at $0.75 and $3.75 per million input and output tokens through 31 December 2026, then $1.50 and $7.50 from 1 January 2027. The rate is real, and so is the date, so any multi-year forecast has to carry both.
Does prompt caching always save money?
No. Caching bills a write premium and, on some platforms, an hourly storage charge, so a poorly sized cache can cost more than none. There is also a minimum cacheable prefix, below which requests are processed without caching and no error is returned, which means a broken caching strategy can look like a working one.
Can I trust published AI benchmarks?
Treat them as evidence about the benchmark. Peer-reviewed work has found that selective disclosure and unequal data access bias arena-style leaderboards, including one case of 27 private variants tested before a launch. Attribute every score to whoever published it and run your own evaluation on your own data.
What should a small business ask before buying AI work?
Ask for a dated price list, a forecast of monthly running cost broken down by meter, the cacheable prefix assumption, and an evaluation plan that does not rely on a vendor table. If an adviser cannot produce those four things, they are quoting a sticker price rather than a bill. Honest ai consulting services will offer all four before being asked.
Read more on this topic#
Nobody Had Edited It in a Decade. Then Came Four Thousand Pages.
What happens when autonomous agents meet a system nobody was watching, and why containment is a governance question.
Read the pieceThe Score Everyone Is Quoting Belongs to a Different Benchmark.
The companion argument on capability and access being separate purchases, told through a gated security model.
Read the pieceThe definitive european ai model: Multiverse Computing launched Quasar 438B
A closer look at how a model launch is marketed, and what the published numbers do and do not support.
Read the pieceThree Studies Measured the Same Ad Load and Disagreed by Thirteen Times.
When measurements disagree by an order of magnitude, the denominator is usually the culprit.
Read the pieceCosting an AI programme that has to survive January?
Our ai consulting services price the meter, not the sticker. We build the forecast, run the evaluation on your own data, and write it down so your finance team can hold it.
Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.