On this page
Last week Chuck laid out four rungs. Here's the first one, the one everybody treats as safe, and the two times our own systems gave us a clean, cited, wrong answer.
When a company decides how worried to be about an AI tool, it asks what the tool can do. Can it send an email, change a record, move money, talk to a customer. If the answer is nothing, if it only looks things up and answers questions, the worry goes away. Reading feels like the one place where nothing can go wrong, because nothing leaves the building and nothing changes.
That reasoning puts reads at the bottom of the ladder, and it's right about the bottom of the ladder. It's wrong about the risk. Chuck laid out four rungs last week, and this piece is about the first one, because it's where we've been burned most often while believing we had nothing to be burned by.
An answer is an action for whoever relies on it. When someone at InTech asks what we agreed with a person two months ago and gets a clean answer back, they walk into the next conversation on that answer. The system changed nothing. The person changed what they said, what they promised, what they skipped. If the answer was built on half the record, the damage looks like any other bad action, with one difference: nobody reviewed it, because from the outside there was nothing to review.
An answer is an action for whoever relies on it.
What reads looks like here
Our company has a memory. Meetings get captured, decisions get written down, and what we've said and done accumulates in one place I can ask questions of. When I ask what we agreed to, the answer comes back with where it came from. I use it every day and I'd be slower without it.
You've met the project-sized version of this. In April, Chuck and I wrote about the Context Graph, the living document we load into Claude at the start of serious work so it stops recommending things we already rejected. Chuck put the risk in one sentence in that piece: "A stale graph is worse than no graph, because it gives Claude confident wrong context instead of uncertain right context."
A company memory is that same problem at the size of the whole company. It's fed by meetings, mail, calendars, and whatever people remembered to write down, and every one of those feeds has an edge. The answers it gives you are only as complete as the feeds, and none of them announce where they stop.
The answers it gives you are only as complete as the feeds, and none of them announce where they stop.
The date it got wrong
The clearest example is one we caught inside our own publishing. The first draft of our Sep 15 piece, the one that lined CRAFT up against Anthropic's AI-native playbook, put the start of our methodology in June. The earliest reference to it that our memory could find was from June, and the draft treated that as the beginning.
I knew it was wrong. We started in January. The piece went out titled "We Wrote Ours in January" because I caught the date in review.
Look at what the memory did and didn't do. It returned a consistent set of records. Nothing in them contradicted the date. Nothing said the story might begin earlier than the record did. The source was real and it was cited, and it started in June, so everything built from it started in June too. Coverage that stops partway through a story reads exactly like coverage of the whole story, because the system has no way to see what it never captured.
Coverage that stops partway through a story reads exactly like coverage of the whole story.
The reason that one got caught is that I happened to know the answer. That part should bother you. The questions people put to a company memory are mostly the ones where they don't already know the answer, which is why they're asking. A wrong answer gets caught at the rate people happen to know the truth in advance, and that rate is lowest on the questions that matter most.
The same mistake from the mailbox
The mailbox has the same shape. It's complete about email. In August I asked for a follow-up to someone I'd been talking to, and the draft came back as a reply into a thread that had ended weeks earlier, opening with an apology for the delay.
There was no delay. We'd spoken by phone the night before and traded texts that afternoon, and the texts had already set the agenda. None of that was in the mailbox. The system read the last email as the last contact and wrote an apology for something that hadn't happened.
That draft belongs to the next rung, and I'll get to it. The failure started at reads. "This thread went cold" is an answer to a question about the state of a relationship, and the mailbox can only answer questions about email. Conversations move to the phone, to text, to a hallway. A source that's complete about one channel looks complete about the person.
A source that's complete about one channel looks complete about the person.
The question under the question
In August, Chuck ran the exit test on our own stack, and the finding that mattered most was that the test checks whether a record exists and says nothing about whether it's current. In his words, "The test measures possession. It does not measure freshness." He told you to ask a second question of any record: when was it last written to, and by what.
Reads is where that question earns its keep. A system can hold everything you ever gave it and still answer from the one stretch of the story it happened to catch. Possession gets you access. Freshness and coverage decide whether the answer is any good, and neither shows up in the answer.
Possession gets you access. Freshness and coverage decide whether the answer is any good, and neither shows up in the answer.
What we do at this rung
After each of these we wrote a rule, and they're short.
Before using anything the memory returns about a matter that's still in motion, we look at the dates of what came back. If everything it pulled sits in one narrow window, and the subject is something that would have kept generating events, the record has an edge and the answer is about the edge. The check takes a few seconds, and it catches the case where the picture is dense, consistent, and a month old.
Facts about what happened hold up well. Facts about where something stands go stale. So a status, a date, or a "where things stand" line gets checked against the live source, or against the person who was there, before it goes into anything that matters. The memory is where we start. It isn't where we finish.
And the system doesn't get to characterize a relationship from a single channel. If it can't see the last interaction, it writes neutrally or it asks.
The standard I want from this rung is an answer that says where it came from, when that source was last written to, and what the system couldn't see. We're closer on the first than on the other two.
This is also why reads is a rung and not a default. Chuck's promotion rule says a task moves up when its failure mode at the current rung is known and rare. You can't know that about reads unless someone checks answers against the source, and checking is exactly what almost nobody does at a rung that looks harmless. Every rung above this one runs on what reads returns. A draft written from half the story is fluent and wrong. An action taken on it is wrong and takes effect. A task that sits on reads with an unchecked failure rate carries that rate up the ladder with it.
A task that sits on reads with an unchecked failure rate carries that rate up the ladder with it.
Where a Blueprint starts
This is why a Business OS Blueprint begins with the memory and not with an agent. Before anything gets an agent, we map what the organization knows that an agent would need, where it lives, and what shape it's in. In practice that comes down to three questions.
What gets written down at all? A decision made on a phone call exists nowhere a system can read it. Who or what writes to each record, and when was it last written to? And what can the system not see? That one needs a stated answer, because the alternative is finding out by being wrong in front of someone.
The client keeps that map whether or not they build with us. It's useful to a company that never buys a single agent, because the same gaps trip up the people who work there.
What comes next
Next week is drafts: the rung where the facts usually come out right and the draft still doesn't sound like the person whose name is on it.
In the meantime, pick the place your business keeps what it knows. It might be a shared drive, a CRM, a notes tool, or one person's head. Then ask Chuck's question of it: when was this last written to, and by what? If you can't answer, that's your first finding. Leave a comment with where yours turns out to stop. I'll read every one.
