Build to Thrive | The AI Blueprint | The week of August 10th, 2026
Building a Foundational Brain
Editorial
I have spent the last week reading the research on where AI actually breaks, rather than where people say it breaks. Three studies, measuring the same thing from different angles, and they agree in a way I did not expect. The failure is not capability. Models are not getting worse at the work. What collapses is consistency, and it collapses in the exact conditions we all work in: long conversations, a lot of context, and a session that has to remember what we agreed twenty minutes ago. I wrote in June about why the tools kept making me busier, and last month about what happens when you stop collecting tools, both pieces This issue hands you three fixes, and none of them is a better prompt.
If your AI output feels inconsistent rather than bad, this issue gives you the one page that makes it steady. If you deliver work anyone pays for, it gives you the sentence to have ready before a client asks whether AI touched it. And if your sessions have been getting longer because everything needs to stay in one place, it gives you the two-minute habit that replaces them.
Leak One: reliability, not ability. Aptitude barely moves as a conversation runs on. Consistency falls off a cliff, and once a session takes a wrong turn it does not recover. The fix is a page, not a prompt, and it takes an hour once.
Leak Two: the window on the label is not the window you can trust. Accuracy drops well before the advertised limit, and whatever sits in the middle of a long brief is the first thing lost. The fix is where you put things, and it costs nothing.
Leak Three: your buyer already suspects. Seven in ten say they would not trust a report made with AI, and the ones who have used the tools are the more sceptical group. The fix is one sentence, written before the question arrives rather than during it.
All three run on the same idea. Nothing inside a conversation is authoritative, so nothing in there can correct a wrong turn once it happens. Reliability has to come from outside the chat, from something you wrote down while you were thinking clearly. It is not a property of the model. It is a property of what the model can check itself against.
Do one of them this week. What you are after is not smarter answers. It is the same answer twice, which is the thing you have actually been missing.
Juan
16 and 112. In over two hundred thousand simulated multi-turn conversations, leading models lost 16 percent of their aptitude and gained 112 percent more unreliability, an average 39 percent drop across six tasks. Capability held. Consistency did not. (Microsoft Research and Salesforce, ICLR 2026)
More than 30 percent. The performance drop when the thing that matters moves from the start or end of your input into the middle of it. Models have primacy and recency bias, exactly like people. (Stanford, TACL 2024)
30 to 50 percent. Accuracy loss across eighteen frontier models measured well before the stated context limit. A 200,000-token window can be unreliable at 50,000. (Chroma, named lab, not peer reviewed)
70 percent. Buyers who say they would not trust a consultancy report prepared using AI, from a survey of 3,887 companies. Among those who have actually used consultants’ AI tools, 77 percent call it a bubble, against 55 percent who have not.
20 to 30 percent. The widely quoted share of the workday knowledge workers supposedly lose searching for information. I went looking for the source and found four different figures from four vendors and one analysis calling it a myth. Directionally true, badly sourced, and I am not using it.
What is happening. Researchers at Microsoft and Salesforce took standard single-turn benchmarks and converted them into multi-turn conversations, which is how people actually work. Every leading model got worse. The average drop was 39 percent across six tasks.
Then they decomposed it, and that is the finding. Aptitude fell 16 percent. Unreliability rose 112 percent. Their conclusion: when a model takes a wrong turn early in a conversation, it does not recover, because nothing in a chat is authoritative enough to correct it.
What it means for you. If your AI output feels inconsistent rather than bad, that is measured behaviour and not your imagination. It also means the common fix is the wrong one, because sharper prompts address ability and ability is not what broke.
Sit with their explanation, because the fix is inside it. Nothing in a conversation is authoritative. Every statement in a chat carries the same weight as every other statement, including the wrong turn taken at minute four. There is nothing for the model to check itself against, so a small early error compounds instead of getting corrected.
Which tells you exactly where the fix has to live. Not in the conversation. Outside it, in something written down when you were thinking clearly, that does not drift as the session gets long. Reliability is not a property of the model. It is a property of what the model can check against.
What I would do this week. Open one document and write four things: what you sell and to whom, how you decide pricing, what makes a client a good fit, and three decisions you have already made and do not want reopened. Paste it at the top of your next session. That document is the authority the conversation does not have. One hour, and it is the highest-return hour in this issue.
What is happening. Two separate findings, same direction. Stanford measured a U-shaped curve: models use what sits at the beginning and end of what you give them and lose what sits in the middle, with more than 30 percent degradation as the important thing moves inward. And the lab Chroma tested eighteen current models and found accuracy dropping 30 to 50 percent well before the advertised context limit.
What it means for you. Two habits are quietly costing you. Running one enormous session all day instead of shorter ones with the context written down. And burying the constraint that matters in the middle of a long brief, where it is the first thing to fall out. The number on the box is a capacity figure, not a reliability figure.
What I would do this week. Put the things that must always be true at the top of the conversation, not in the middle of paragraph nine. And when a session gets long, stop and ask it to write down what has been decided so far, then start fresh with that summary. Two minutes, and it is the difference between a session that holds and one that drifts.
What is happening. Source Global Research surveyed 3,887 companies. Seventy percent said they would not trust a consultancy report prepared using AI. Almost a third said learning a firm had used AI would shake their confidence in that firm.
The part that inverts the expectation: the clients who have actually watched these tools work are the more sceptical group. Seventy-seven percent of those who have used consultants’ AI tools think the technology is a bubble, against fifty-five percent of those who have not. Familiarity lowered trust rather than raising it.
What it means for you. If you deliver work anyone pays for, the question is no longer whether to use AI. It is what you say about it, to someone better informed and less impressed than they were a year ago. Hiding it is not a strategy and confessing everything is not either. What the buyer is actually paying for is a named person who is answerable when the work is wrong.
What I would do this week. Write one sentence you would be comfortable saying out loud if a client asked directly. Not a policy. One sentence. If you cannot write it, that is the thing to fix before the question arrives.
Move 1: Write the four-answer memory page. What you sell and to whom, how you price, what makes a good client, and three settled decisions. One page, one hour, pasted into every session from now on. This is the whole fix for Leak One: the conversation has no authority in it, so you supply one.
Move 2: Move the constraint to the top. Whatever must always be true goes first, never buried mid-brief. This is free and it recovers most of that 30 percent.
Move 3: Break the marathon session. When it gets long, have it write down what has been decided, then start clean with that summary. Long sessions feel efficient and measure badly.
Move 4: Write your one sentence about AI use. Before a client asks, not during. Check it against the six things that should stay human.
None of these needs a developer and none needs a new subscription.
PROMPTS
The Four-Answer Memory Page. Walks away with a one-page description of the business that makes every future session specific instead of generic. Asks the four questions, produces the document, says where to keep it.
The Session Handoff. Walks away with a clean summary of everything decided in a long session, formatted so it can start the next one. Built for people whose conversations run all day and drift by afternoon.
The One Sentence. Walks away with a single line they can say to a client about how AI is used in their work, drafted from their own delivery process rather than a template.
Run in order they cover the same three leaks: what the business knows, what a session remembers, and what the buyer is told.
Operator Second Brain Starter Pack (Free Subscribers)
8 plain text files and a morning routine. That’s it. No software to install, no tool to buy, no architecture to design.
Drop these files somewhere your AI assistant can read them (ChatGPT projects, Claude projects, a local folder your AI tool has access to). Fill them in once. Update them as you go. Read them every morning with AI before you start working.
That’s the whole system. Operator Second Brain
These are notion text files. For our paid subscribers, you can access my Second Brain Prompts at the end of this edition.
See how your whole business is organized. There are nine moving parts to a healthy business. One is under strain. Most operators fix symptoms instead of finding the real constraint. The Thrive Business System Assessment shows you all nine and helps you find which one matters first. Numbers & Decisions might be it. Or it might be something else. The map shows you. Start at learn.buildtothrive.co/build
Run the AI Leverage Assessment. If you already know the strain is operational drag, time leaks, or workflow bottlenecks, this is the faster path. It shows you the cost of friction and where AI actually helps. Start at learn.buildtothrive.co/mytimeback
The Thrive Business System Assessment. Run this diagnostic to see all nine moving parts of your business and surface which one is under strain. Most operators fix symptoms instead of finding the real constraint. This assessment walks you through each lever—Clarity, Opportunity, Follow-Through—and names where the bottleneck sits. Numbers & Decisions might be it. The Thrive Business System












