Building the Machine, Day 5: I Made My AI Grade My Own Homework.
It Was Half Genius, Half Lost.
Last night I did something most people building with AI never do.
I wrote the full spec for the product I’m launching. The whole thing. Architecture, pricing, the client flow, the security model, every piece. Hours of work. And then, instead of handing it to my AI and saying “build this,” I handed it over and said something else.
“Tear it apart. Tell me where it breaks. Don’t tell me it’s good. Assume it has problems and find them.”
What came back is the reason I’m writing this.
The genius half
It was sharp. Scary sharp.
It caught a contradiction I’d been walking past all week. I kept saying “build it right,” “hit the deadline,” and “don’t cut corners” in the same breath. Those three can’t all be true at once. It made me pick.
It caught a money problem. One piece of my flow would’ve had the system charging a client automatically at a moment nobody actually approved. Small detail. Would’ve turned into a nightmare and a pile of chargebacks. It flagged it against my own rule, and it was dead right.
It even redesigned a piece of my product better than I’d built it. I described something complicated. It handed back something simpler that does the same job. I took it.
If I’d stopped there, I’d have called it a win and gone to bed.
The lost half
But something felt off. And this is the part that matters.
The entire assessment was built on an assumption I never made. It treated my new product as if it were bolted onto the messy infrastructure I already run. Every concern it raised, every “this is exposed,” every limitation - all of it traced back to that one wrong assumption. It wasn’t answering my question. It was answering a question it invented.
And here’s the kicker. It said it with total confidence. Exact same tone as the genius parts. No flag. No “I’m assuming here.” Just stated it like fact.
The lesson nobody sells you
The value tonight wasn’t the AI being right. It wasn’t even the AI being wrong.
It was that I could tell the difference.
That’s the whole game. That’s the part nobody mentions when they’re selling you the magic robot that runs your business while you sleep. The machine is only as good as your ability to catch it when it’s confidently full of shit. You have to know your own operation well enough that when the AI hands you a beautiful, articulate, wrong answer, an alarm goes off in your gut.
I’ve watched people take the confident wrong answer and run. Build on top of it. Ship it. Because it sounded smart, it came fast, and questioning it felt like more work than trusting it.
The judgment doesn’t get automated. You are still the last line. The AI is the fastest, sharpest, most tireless junior partner you’ll ever have, and it will walk you off a cliff with a smile if you’re not paying attention.
So I fixed the assumption. Sent it back. Told it to reassess against reality. And the spec is genuinely better tonight than it was this morning because I used the machine as a sparring partner rather than an oracle.
That’s what building the machine actually looks like. Not magic. A fight. A good one.
For the paid crew: the actual tool
Everything above is the story. Below is the system - the thing you can steal and use tomorrow.
Because “be able to tell the difference” is useless advice on its own. Nobody can act on that. So here’s what I actually did, in a form you can copy:
The exact “tear it apart” prompt I used - word for word, the one that forces your AI to red-team instead of cheerlead. Most people’s prompts beg for validation without realizing it. This one makes that impossible.
The 3-question gut check I run on every AI answer before I trust it - the specific things I look for that separate a confident-right answer from a confident-wrong one.
The one structural move that made my AI’s wrong assumption obvious instead of invisible - so you catch the cliff before you walk off it, not after.
This is the difference between using AI like a tool and getting used by it. Paid members get the whole kit below…




