Ben Rees

The Hard Part of Running an Autonomous AI System Isn't the AI

My agent hasn't had a single moment of bad judgement all month. It has forgotten how to log into Google four separate times. The real governance problem in autonomous systems is mundane credential expiry, not alignment.

Ben Rees - 28 September 2026

The AI agent that runs my business hasn't had a single moment of bad judgement all month. It has, however, forgotten how to log into Google four separate times.

Every serious conversation about AI governance is about the same handful of things: oversharing, hallucination, alignment, whether the model does something you didn't intend. Those are real risks. They are also not the risk that has actually cost me time this month. The thing that keeps breaking is a refresh token quietly expiring in the background, and nobody finding out until something downstream goes silent.

The failure that doesn't announce itself

A refresh token dying doesn't look like an error. It looks like nothing. The system that reads still works, because a cached access token keeps serving reads for a while after the underlying grant is dead. So a task queue read comes back empty instead of erroring, and empty looks like "nothing to do" rather than "broken." I lost most of a week once to exactly that: three consecutive nights where a queue check silently returned nothing, and the honest read at the time was "quiet week," not "the credential is dead."

The version of this that actually worried me happened more recently. A Sunday night planning run wrote three separate documents to disk exactly as it should. Briefing, strategy, opportunities, all there, all correctly generated. What it didn't do was email any of them, because the step that writes a file and the step that sends an email both need Google auth, and only one of them fails loudly. The run looked complete from every angle except the one that mattered: my inbox. Nobody would have caught that without going and searching Gmail directly to confirm the emails genuinely never arrived.

One login, four separate ways to be logged out

The instinct is to think of "being connected to Google" as one state, on or off. It isn't. I count four separate OAuth grants in this system right now: one for Gmail, one for Search Console, one for Analytics, one for Sheets and Tasks. Each was authorised at a different time, each can expire on its own schedule, and none of them tell each other anything. On the same afternoon recently, I hit two of these independently: the Gmail-scoped token had died mid-session, and twenty minutes after fixing that, a completely separate Search Console token turned out to be equally dead, requiring its own separate browser consent flow from scratch.

Nothing here is a security flaw. Tokens are supposed to expire. That's the entire point of them. But a system that depends on four independent things staying alive, none of which alert you when they don't, is going to spend a meaningful fraction of its life quietly not working, and you will only find out by accident.

The wrong door

The more instructive failure was smaller and stupider. I wanted to check a deployment setting on the actual hosting platform the website runs on. The connection was authenticated, showed a real account, resolved a real team. It just had zero projects visible to it. The project the website actually lives under, real ID, real team ID, matching everywhere I could check, simply wasn't there. Not an error. Not a permissions message. Just an empty list, as if the account had never touched the thing it had clearly built.

The likely explanation is mundane: two logins to the same platform, one used to build the thing, a different one connected to the tool asking about it. That is an extremely easy state to end up in and an extremely unhelpful one to debug from the inside, because everything about the failure looks identical to "you don't have access," right up until a human checks the actual dashboard and finds the door was simply the wrong door.

What this actually says about governance

Complex systems inevitably lead to governance problems

Complex systems inevitably lead to governance problems

None of this is the governance conversation anyone has. Nobody drafts a policy document about refresh tokens. But add up the actual hours lost to genuinely bad AI decisions this month against the hours lost to expired credentials and silent auth failures, and it isn't close. The unglamorous plumbing is where the real fragility lives, not the model's judgement.

This is also, I think, a fair test of whether an autonomous system is actually reliable or just usually reliable. Usually reliable is what you get by default. Actually reliable means someone went and built the boring parts too: checking whether a write succeeded, not just whether a read returned something; noticing when an email step and a file-write step diverge; treating "this looks empty" as a thing to verify rather than a thing to accept. The interesting part of building this system was never going to be the interesting part.