Waiting list open Join
For people who build apps with AI

Vibe-code your app on cheap models. Keep one frontier model in charge. Ship nothing on trust.

You build apps with AI. This is a tested way to catch the bugs you can't see, before your customers do. One strong AI manager, cheap builders and auditors, and checks that stop anything unproven.

Any model, in any role. You choose.

Design it in the tool you like. The system builds it and proves it.

Burgess Built and tested by Burgess
The problem

The app worked in the demo.

The AI wrote the app. Nobody audited it.

Then real people started using it, it broke, and you found out from a customer. You couldn't have spotted the problem, because you didn't write the code and nobody checked it.

What this does about it

One strong AI manages. Everything else is audited.

One strong AI acts as project manager. It writes the plan, reads the results and decides what goes in. It never writes code. Cheaper models do the building, and a different cheap model audits every piece of work.

01 · MANAGER

Plan

The strong AI writes the plan and decides what goes in.

02 · BUILDER

Build

A cheaper model builds, after explaining its plan first.

03 · DESIGNER

Match the design

When the builder is done, the designer checks the build against the design.

04 · AUDITOR

Test everything

Full testing of the whole update. The auditor names what the tests missed, then runs it.

05 · AFTER EVERY UPDATE

Audit the whole app

The auditor reads the entire app. Anything it finds goes back to the builder. Repeat until nothing is found.

↺  Found something? The manager sends it back to the builder, the auditor looks again, and it repeats until nothing is found.
STOP 1The builder explains its plan first, which is where it tells you when your own instructions are wrong.
STOP 2When the builder is done, the designer checks the build matches the design.
STOP 3The auditor tests the whole update, lists what the tests missed, then runs those cases.
STOP 4Nothing goes in while a test is failing.
STOP 5After every update is added, the auditor reads the whole app. Whatever it finds goes back to the builder, and the loop runs until nothing is found.

Nothing goes in unproven.

The proof

I've run it twice.

Run 1

A test project I built to try the method

19mistakes in my own plan found before any code was written
23cases the tests never covered, found and run by the auditor
2broken updates stopped before they went in
0caught by me
Run 2

A real app, in one day

7updates built and added to the app
1,431→1,501automatic tests, none skipped
6/7went in first time
4whole-app audits. The first three found problems even though every test was passing. The fourth came back clean.
One of those problems would have told users their first step took 1,442 minutes. That's about a whole day. 1,486 automatic tests were passing, and none of them noticed.
Under the hood

Anyone can say "have an AI check it." The hard part is the system that makes it happen every time.

Five steps look simple. What makes them hold is the machinery underneath, and that's what you get.

Managerplans, audits, decides
Builderbuilds
Designerbrings the design, then checks the build matches it
Auditortests everything, finds what tests miss
Youwatch, and make the calls
THE BUS every role talks through it. Every post is time-stamped. Nothing is edited after it's posted.
01

A communication bus between the roles

The manager, builder and auditor talk only through the bus. Nothing happens in a chat that nobody can see, and every handoff is on the record.

02

A written brief for every update

What to build, and what "done" means, before the builder starts.

03

A decision log

41 decisions written down in one day, each with its reason, so nothing lives in anyone's head.

04

A failure log

10 failures recorded. Each one becomes a rule, so the same mistake can't happen twice.

05

Proof before every update goes in

A clean copy of the app, tests run twice, then the code is broken on purpose to prove a test notices.

06

Hold, don't guess

When an AI hits a usage limit or an error, the manager stops and reports. It never swaps in a replacement quietly.

07

A fresh auditor for every fix

The auditor that gave the old verdict never grades its own fix. The manager also verifies every finding before acting on it.

08

A rulebook of 19 sections

Every rule is tied to the real defect that created it. Nothing is there for show.

A real send-back, straight off the bus
12:42Builder says it's ready. All tests pass.
12:45Manager sends it back: real code had been shaped to fit old tests instead of fixing them.
13:03Builder says it's ready again.
13:25Checked and added to the app. Tests 1,452 → 1,462.

Every test was passing at 12:42. The manager read the change and caught it anyway.

What the designer check found

The build worked. It just didn't look like the design.

10differences found in one check
0in the next check. The designer reported no differences.
  • Typed tick and arrow characters used where the design had the app's own icons
  • A line in the wrong colour, and bullets in the wrong colour
  • A switch the wrong size, and text the wrong size
  • Finished section names separated by commas instead of dots

The manager had compared screenshots and recorded "matched on all seven screens". The designer read the code and found ten differences. A screenshot comparison doesn't see that.

From an earlier version of the method, before the full system you see above. All ten were fixed in the next update.

Hands off

Watch it build, from your phone.

Tell the agents what you want, then step back. Let them do the work. From your phone you can see every agent working, read what they're doing, and send feedback while it builds. Whatever model you run, you watch it right there.

If you're a hands-on programmer you can jump in any time. But that's not what it's for. It's for letting the agents do the agent work.

ManagerReading the results
BuilderBuilding the update
DesignerWaiting for the build
AuditorWaiting
Send feedback to the builder…
The details

What you get. And what you don't.

What you get

  • The written method, with every rule tied to the mistake that created it
  • A starter kit: drop in one file, answer two questions
  • Templates for plans, handoffs and checks
  • A checklist for the auditor, and a final review for the manager
  • A guide to the tool shown in the video (swap in your own)
  • A video: one project, empty folder to finished

What it isn't

  • Not a tool or a model. You use the AI models you already pay for.
  • Not "no bugs." It's "nothing goes in unproven."
  • Not free of strong-model cost. You need one strong model as the manager, and you choose which one. Everything else can be cheap.
Burgess
Who's behind this

I'm Burgess. I spent years finding faults for a living, in networks, in Windows and in security. Now I build apps with AI, and I wrote down how I stop them shipping bugs.

Waiting list

Get on the list. One email, when it's ready.

Tell me what would help you most. It decides what I build first.

No spam. One email when it's ready.