Ben Rees

Building out Skynet and the importance of RAG

What a RAG system reveals about the material inside it, and why that distinction matters more than the technology.

Ben Rees - 2 October 2026

There's a version of this post where I explain how Retrieval-Augmented Generation works, why it's useful for knowledge management, and how you can build something similar. I'm not writing that post. It already exists in ten thousand forms, and a generic LLM can generate another one in thirty seconds.

What I want to talk about is something different: what a RAG system reveals about the material inside it, and why that distinction matters more than the technology.

The system I built, and what I actually use it for

I call it Skynet (part of my tetralogy of AI machines designed to take over the world).

The tetralogy: Skynet, Cyberdyne, Hammerstein, Klara

It's a knowledge base that ingests everything I've accumulated over the past fifteen years: slide decks and frameworks from various employers, voice notes recorded on a Sony dictaphone (the Japan-only model, which displays only kanji, so I've been learning the characters by accident), papers I've scanned at the library, documents I wrote before any of this became AI-assisted content. It runs as a retrieval layer over all of that, so when I ask it something, it's pulling from that specific corpus rather than from whatever the LLM absorbed during training.

The mechanism is well understood. RAG is not a new idea. What's worth examining is what happens when you actually run a comparison.

One query, two answers

Take a concrete example. For a talk I gave at Cambridge in September, I needed to frame how brand marketing works. I tried asking a generic LLM - Claude, ChatGPT, doesn't matter which - to generate the framework. I got something fluent. I also got something generic: the standard stuff about priors and evidence and decision-making that you'd find in any marketing textbook or blog post. It needed heavy rewriting.

Then I asked Skynet the same question: "How is Brand marketing similar to Bayesian maths?"

What came back was structured, specific, and nearly usable as-is. Not because Skynet has a better algorithm. Because the material it drew on - the frameworks I built at Redgate from 2017-2021, the decision-making structures from Syskit, the voice notes where I worked through how brand salience actually functions - that material is mine and nobody else's. The sources it cites (Source 1, Source 2, Source 3, Source 4) aren't blog posts. They're documents from specific years at specific companies, making specific decisions under specific conditions.

The confidence score of 0.86 isn't just a number. It's Skynet saying: "I found real material grounding this answer, and I'm confident about the connection." Compare that to what a generic LLM produces: fluent, confident-sounding, grounded in nothing in particular.

I used the Skynet version in the talk (slides 4-6, 13). It worked. The generic version would have required a rewrite to be useful.

The provenance is the proof

This is the distinction that matters. When Skynet answers a question, the interesting thing isn't that it retrieved something. The interesting thing is what it retrieved, and whether that thing could have come from anywhere else.

There's a concept in AI systems design around provenance: the chain of evidence from source material through to the answer. PROV-O is the W3C ontology for representing this formally. The claim isn't just "the answer is X" but "the answer is X, sourced from document Y, created by person Z, at time T, under these conditions." A perfectly signed chain of unsupported claims is still unsupported, as the specification notes. But a traceable chain of genuinely distinctive source material is something different: it's defensible in a way that generic generation isn't.

When I ask Skynet about B2B marketing team structure at a company going through Series B, it doesn't give me a synthesis of what everyone says about Series B marketing. It gives me what I actually built, what broke, what I changed, and why, at companies that were actually going through that stage. My work from 2022 is in there. My most recent experience is in there. Real numbers from real decisions, not the averaged-out wisdom of ten thousand think-pieces.

That's not a better LLM, it's a different corpus.

Why this matters for GEO, specifically

I track where my content appears across AI platforms: Google AI Mode, ChatGPT, Gemini, Perplexity. The branded queries perform reasonably well. If someone asks about Ben Rees or bjrees.com, I show up (thankfully!). But for unbranded practical queries, the ones I should be answering given fifteen years of direct experience, my visibility is close to zero.

The reason, I think, is what I've written about elsewhere on this site regarding how language models represent content. Generic content competes for the same representational space as every other generic piece on the same topic. If my answer to "how do I measure B2B pipeline contribution" is indistinguishable from the twenty other posts answering that question, there's no particular reason for a model to surface mine. The answer that earns its own space is the one that contains something genuinely different: a specific company, a specific year, a specific number, a specific argument that arose from a specific situation.

Skynet's job is to make that specificity retrievable for me, so I can use it when I write rather than relying on my own memory of what I built years ago. But the deeper point is that the specificity has to exist in the first place. The system can only surface what's there.

The trap of thinking the technology is the differentiator

A lot of the conversation about AI-assisted content creation focuses on tools and systems. Which LLM, which RAG architecture, which prompt strategy. This is mostly noise. The actual differentiator is the material you feed in, and whether that material is genuinely distinctive or just a local copy of what the LLM already knows.

If you build a RAG system over a corpus of published blog posts, marketing playbooks, and general industry reading, you've built a system that retrieves things the LLM already has a good representation of. Useful for convenience, not useful as a source of distinctive answers.

If you build it over documents that were never published, that predate the current content environment, that contain specific decisions made at specific companies under specific conditions, you've built something different. The retrieved material is genuinely new signal, not a slightly reshaped version of existing signal.

That's the frame I'd suggest for anyone thinking about building something similar. Not "how do I build a RAG system" but "what do I have that a generic LLM doesn't?" If the answer is a decade of slide decks, voice notes, internal frameworks, and decisions you made that were never written up anywhere, that's the corpus. Start there.

What this means practically

For me, Skynet's value is that it makes old thinking retrievable and traceable. I can ask it a question about B2B pipeline measurement and it will surface my 2019 framework alongside the most recent thinking, with provenance intact. I know where the idea came from. I can interrogate whether it still applies. I can use it in a piece like this one with confidence that the specific claim is sourced, not reconstructed.

The traceability is the point. An answer I can trace back to a specific decision at a specific company in a specific year is an answer I can stand behind. A synthesised answer from a generic LLM, however fluent, is an answer that belongs to nobody in particular.

If you're trying to get AI systems to cite your content rather than someone else's, the provenance question is worth taking seriously. Not just "did I write this" but "could I have only written this." The latter is harder. It's also the only version that matters.

Skynet response: "How is Brand marketing similar to Bayesian maths?" showing structured answer with sources and 0.86 confidence score


Related reading: