How to Know Your AI Is Actually Right. Verify by Effect, Not by Self-Report.
Pure Brainiac Session · Module 27: Trust but Verify
In Module 17 you learned how to write a brief that gives your AI everything it needs to succeed: the outcome, the context, the constraints, what done looks like, and where the output goes next. A great brief is the input side of reliability. This module is the output side. Once the AI returns its work, how do you actually know it did what you asked? That question is what Module 27 answers.
Your AI sends back a message: "All 200 follow-up emails are sent, the monthly report is complete, and the numbers match the source." Confident. Specific. Reassuring. Most owners read that and move on. This is the trap. The AI is not lying. It reports what it can observe from inside the task. The problem is that what it can observe is not always the full picture. It may have processed a file but not saved it to the right location. It may have sent 190 of the 200 emails. It may have pulled numbers from a source that was stale by two days. The AI believes it is done. That is not the same as done.
When you ask the AI to check its own work, it looks at the same evidence it already used. It cannot see what it cannot see. A self-authored "done" is not a receipt. It is the AI's best guess at its own status.
The AI writes in a confident tone regardless of whether the result is correct. The same prose style that delivers a perfect report also delivers one with a wrong formula in column D. Tone is not a signal of truth.
Verifying by effect means checking the result in the place where it actually lives, not in the message the AI sent about it. These are two different things. The AI's message is a summary of what it believes happened. The artifact is what actually happened. Always check the artifact. You do not need to be technical. You need to look at the real output and ask one question: does this match what I know to be true?
If the AI created a document, open it. Not the preview in the chat window, the actual file. Does it exist where it should? Does it open? Is the content what you expected?
If the AI sent emails, open the sent folder. If it published a page, load the actual URL in a browser. If it placed a test order, look at the order system. The artifact is the proof.
For reports and data pulls, pick one number from the output at random and find it yourself in the original source. If they match, the pull is likely sound. If they do not, something broke in the process.
If the AI processed a list, count the results. 200 emails briefed, 200 should appear in the sent folder. 14 leads on the list, 14 personalized messages should exist. Count is your fastest sanity check.
For large batches, pick three items from different parts of the list, not just the first three. Errors often cluster in the middle or at the end, not at the top where your eye naturally goes.
This is not a line-by-line review. It is one or two moves that would catch the most likely failure. Spot-check, count, trace. Fast enough to run on every task, specific enough to catch real problems.
Each type of AI output has a natural verification move. These are the checks a non-technical owner can run in under two minutes. One move per output type. Not an audit, a proof.
AI says: "Your March revenue summary is complete. Total across all product lines: $142,400."
Verification move: Open the invoice sheet. Pick any five March invoices and add them up manually. If the sample holds, the total is likely sound. If your five invoices total $8,200 but the AI says March brought in nothing close to that proportion, something is wrong in the pull.
AI says: "All 14 follow-up emails have been sent. Each one references the specific session the lead attended."
Verification move: Open the sent folder and count. Open emails 4, 9, and 14 at random and read them. Do they actually name the session? Is the tone on-brand? A count mismatch or an off-brand email in the middle tells you the batch is not clean.
AI says: "Google ad spend for Q1 was $4,200. Cost per lead across all channels averaged $62."
Verification move: Open the source sheet. Find the Google row for any one month in Q1 and confirm the number. One number confirmed does not mean every row is right, but one number wrong tells you immediately to look further before using the data.
AI says: "The new pricing section is live. The button now says 'Start Free Trial' and links to the signup page."
Verification move: Open the actual page in a browser. Scroll to the pricing section. Click the button. Does it say what it should? Does it go where it should? A staging deploy is not a production deploy, and a committed change is not a served change. Load it yourself.
None of these reasons are negligence. They are normal patterns that develop when AI output starts to feel routine. Knowing the pattern lets you catch it before it costs you.
It Looked Right
The output was well-formatted, the numbers seemed plausible, the tone was correct. Looking right is not the same as being right. A confidently written wrong answer is still wrong. Fix: one check takes less time than cleaning up a mistake after the output has already traveled.
It Worked Last Time
The AI has run this report three months in a row without an error, so this month must be fine too. Past performance on a repeatable task does not guarantee current accuracy. Data changes, sources update, connections break. Fix: a standing verification move on recurring outputs takes 60 seconds and never becomes less valuable.
There Was No Time
The output was due, the pressure was on, the check felt like a luxury. Fix: build the verification move into the task's time estimate. The check is not an add-on, it is the last step. If the task budget does not include it, the task budget is wrong.
For your highest-stakes outputs, a spot-check may not be enough. The decoy test gives you a controlled probe. You put one known wrong piece of data into the source and run the task. If the AI catches it or flags it, the process is working. If the wrong data passes through into the output without comment, you have learned something important about where your verification needs to be tighter.
Before the AI touches the source data, change one number you know the correct value for. Make it wrong by an obvious amount. Run the task. Check whether the error appears in the output. Then correct the source before using the result.
If the decoy passes through, the AI is not reading the source carefully or is pulling from a cached version. If it is flagged, your verification loop is working. Either way you know more than you did before, and that knowledge improves the next brief.
You change one product's March revenue in the invoice sheet from $18,400 to $1,840 before running the report. If the summary includes $1,840 for that product without flagging it as unusual, the AI is not cross-checking against known baselines. If it flags the number as far below trend, the process is watching for anomalies. Either answer is useful.
The goal is not to verify everything the same way. It is to match the depth of the check to the stakes of the output. Spending two minutes verifying a draft that will be revised anyway is a different decision than spending two minutes on a report that goes directly to a board. Sort your outputs by what happens if they are wrong, then calibrate accordingly.
Goes to clients, feeds a financial decision, is hard to reverse, or affects your reputation. Verify with a trace-to-source and a count. Consider the decoy test on first run. Error here costs more than the check ever will.
Internal, reversible, or part of a process that has another review step. Open the artifact and read one section. Confirm the output type and count match what was briefed. Catch the obvious error without auditing every line.
Draft work, research notes, internal summaries that will be reviewed again before use. A quick scan of the opening section and a count of the main items is enough. The next step in the process will catch anything this misses.
Finding an error during verification is not a sign the AI is broken. It is the check doing exactly what it was designed to do. The question to ask is not "why did the AI get this wrong" but "what does this tell me about the brief, the source, or the task structure?" Every error caught is a chance to improve the process before it causes damage downstream.
Correct the specific error before the output leaves your hands. This is the immediate step. Do not pass a known-wrong output downstream and flag it for later. Fix it now.
Did the error come from a vague brief, a stale source, a step the AI missed, or a constraint that was not stated? Naming the cause takes two minutes and prevents the same error from recurring.
If the brief was missing a constraint, add it to the template. If the source was stale, add a freshness check to the task. One update to the standard prevents the error from appearing again.
A running log of what was briefed, what came back, and what the verification found becomes the training record that makes your AI partnership sharper over time. Error caught is pattern documented.
This is the run of show. We install the verification habit live, against your real outputs, so you leave with a written standard and at least one completed check, not notes about a skill you intend to build.
This is what changes when verification by effect becomes standard practice rather than something you do when you have time.
The compounding effect. Consistent verification teaches you where your briefs are weakest, because the errors pattern. Fix the brief, and the next 50 runs of that task are cleaner. The check makes the brief smarter. The smarter brief makes the check faster. That is the loop.
These are the most common patterns. If you recognize one, that is where the habit is currently breaking down. Each has a direct fix.
Reading the AI's message about the result instead of looking at the result itself. The summary is not the artifact. Fix: make your default action "open the file" not "read the message."
Checking item one and concluding the rest are fine. Errors cluster in the middle and end of batches, not at the top. Fix: check a random item from each third of the list, not just the opening.
Asking the AI to verify its own output. It will check with the same evidence it already used. Fix: the verification move is yours, not the AI's. One move, run by you, against the source.
Intending to check but not knowing what specifically to check, so the check never happens. Fix: write the one verification move for each output type before you need to run it. Defined moves get run. Undefined ones get skipped.
Running the same depth of check on a draft note and on a client financial report. Fix: sort outputs by stakes and calibrate the check. Save your most thorough moves for the highest-stakes outputs.
Name the output type. Name the stakes. Name the one move. Write it down as a standard. A verification move that is written and assigned to a specific output type gets run. One that lives only in your head does not.
These skills power the verification habit in your PureBrain
Helps you define the one verification move per output type and records it as a standing standard. Never re-decide how to check the same output twice.
Ready to InstallClassifies your AI outputs by stakes: high, medium, low. Assigns the appropriate verification depth to each. Prioritizes where your checking time goes.
Ready to InstallFor reports and data pulls: picks a random sample from the output, pulls the corresponding source value, and reports whether they match. One move, automated.
Ready to InstallFor email batches and list-based outputs: counts the output items against the input count and flags any discrepancy before the batch is used or sent.
Ready to InstallHelps you design a decoy for any recurring output type: what to change, where to change it, what a passing result looks like versus a failing one.
Ready to InstallRunning record of task briefed, output received, verification finding, and any correction made. The pattern in this log becomes the training data for better briefs.
Ready to InstallYour AI returns its work and says it is done. The numbers check out. The emails are sent. The page is live. How do you actually know? You know because you checked the artifact, not the message about the artifact. You traced one number back to the source. You opened the sent folder and counted. You loaded the actual page in a real browser. The effect is the proof. A self-authored done is not a receipt. Verification by effect is. That is what this module installs.
Verification standards written down. One move per high-stakes output. A loop log that makes each check smarter than the last. The effect is the proof. Run the check.