Ben Rees

The fix for a spiralling AI bill is rarely a cheaper model

Most AI cost overruns come from asking the same easy, judgement-free question over and over through an expensive interface, not from asking anything genuinely hard.

Ben Rees - 9 October 2026

Today, working through a backlog of bugs in my own knowledge base, I stopped asking my AI agent three separate questions and had it write three short scripts instead.

One replaced a document-counting check I'd had it hand-write from scratch three times in one session. One replaced a two-step login-and-query ritual I'd had it retype every time I wanted to ask my own database something. One replaced a three-command ingest sequence I'd had it re-assemble by hand every time a source file changed.

None of those were hard questions. They were the same easy question, asked repeatedly, through an interface built for judgement calls.

Cheaper tokens and a rising bill are not in tension

The standard advice for AI costs is to use a cheaper model. Epoch AI's own data shows frontier inference prices falling roughly 50x a year between 2020 and early 2025. Tokens have never been cheaper.

Plenty of people's bills have gone up anyway. That's not a contradiction. A cheap thing done constantly still adds up to real money, and a model will do the same mechanical thing constantly if nobody tells it not to, because answering is what it's built to do. It doesn't know the question was pointless. It just knows the question arrived.

Someone did this at a much bigger scale than I did today

A Netflix storage architect, Tejas Chopra, found himself burning around $200 a day on agent tool calls stuffed with redundant, machine-generated boilerplate. He built Headroom, an open-source tool that compresses that boilerplate before it ever reaches the model, cutting token use by 60 to 95 percent with no measurable drop in accuracy on real benchmarks.

Same realisation I had today, just further along and at far greater scale: most of what gets sent to a model was never the hard part. It's the filler around the hard part, sent again and again because nobody separated it out.

The actual test isn't "is this expensive"

It's "does this need judgement, or does it have the same shape every time."

A question that needs the model to weigh something, decide something, or notice something new belongs with the model. A question that's identical in structure to the last ten times you asked it doesn't. That one belongs in a script, and a script costs nothing to run a thousand times.

This is a different discipline from model-tiering, which is the other common answer (cheap model for easy tasks, expensive model for hard ones). Tiering still assumes every request should go to a model at all. Half the fix is recognising which requests shouldn't.

Moving the work doesn't remove the risk, it relocates it

The honest complication: converting something into fast, automated work doesn't make the risk disappear. It changes what kind of mistake becomes possible.

While I was cleaning up a throwaway test from earlier in that same session, a single tidy-up command also deleted the actual fix I'd just written, because it was still sitting uncommitted when I ran it. Nothing was lost permanently. It just meant redoing thirty seconds of work I'd already finished once.

That's a trivial example, and it stayed trivial because someone was watching it happen in real time. The same mistake at a higher level of unsupervised automation, running on something that mattered more and with nobody watching, is a different story. Cheap and fast doesn't mean safe. It means the consequences of not paying attention arrive faster too.

The bill isn't the real problem

A model that's never allowed to make a judgement call doesn't need to be a model. The actual saving isn't a smaller number on an invoice. It's noticing, honestly, which of your own requests were never questions at all. They were chores wearing a question's clothes.

Related reading