My agents write code faster than I can review it.

That is the actual problem. I would sit down to review, work through the queue, and the queue was longer when I finished than when I started.

So I did the obvious thing first. I tried to review faster.

That held for about a week. Then it stopped holding, and it stopped for a reason I should have seen coming. AI scales output. It does not scale my judgment. I am the same person, reading the same way I read last year. The only thing that changed is how much arrives.

So speed was never the fix. And when I finally mapped my own lifecycle end to end, the real problem was sitting somewhere I had not been looking: nothing in the system decided what was allowed to reach me in the first place.

Nothing was guarding the door.

Why the popular fix fails

Most of the AI policies being written right now ask one question. Was this made by AI?

You cannot answer that. You cannot detect it and you cannot prove it. So the control fails at the exact moment it has to hold.

And honestly, even if you could detect it, you would be solving the wrong problem. Where the work came from does not matter. What matters is that the work shows up unproven, and a human has to do all of the proving.

A gate asks something different. It asks what the work can prove before anyone looks at it.

Every hour your experts spend confirming what a machine could have confirmed is an hour they are not spending on judgment.

Human review should not be your first verification step. It should be your most expensive escalation layer.

That single reordering is the whole design.

The five questions

Five gates decide whether work is allowed to reach a person.

  • Ownership. A human is accountable for this output. Not the agent.
  • Scope. The work is bounded enough that someone can actually review it.
  • Evidence. Everything a machine can check has already been checked, and it proves it.
  • Explanation. The person who submitted it can say what it does, why it exists, and defend the decisions.
  • Risk. Not every output deserves the same depth of human attention.

The first four ask whether the work is ready to consume expert attention at all. The fifth asks a different question: whose attention, and how much. Four gates admit. One routes.

None of the five ask where the work came from.

That is the concept. Here is what happens when you make it machinery instead of a policy sitting on a wiki page.

I recorded the whole walkthrough, gate by gate, against the system actually running it. It is here: the long-form version. You do not need it to follow this. Everything below stands on its own.

Gate by gate, in my own lifecycle

Ownership. Every piece of work in my system starts as a PRD, and every PRD carries a name and a date at the top. When something breaks three weeks later, I do not interrogate an agent about why the thing exists. I look at the document. I know who decided, when, and what they were going after. If no human will put their name on the work, the work does not start.

Scope. This is where most AI workflows come apart. An agent can generate an enormous change in one shot, and an enormous generated change is not a review request. It is a research project. So every story in my system becomes a bounded work contract, capped at what can be done in about a hundred to a hundred and fifty thousand tokens.

That number is not theory and it is not from a paper. It came out of doing it. That much work produces an amount of code, at a level of complexity, that I can read without going numb.

Most teams bound their work units around what the model can produce in one pass.

I did not size the work to what the AI can produce. I sized it to what I can judge.

Those are different numbers, and only one of them is the constraint.

Because every unit is small, I take a two minute break between reviews and my attention resets. The last review of the day gets roughly the same brain as the first one.

Evidence. Every test node in the pipeline carries validators. Before work is allowed to reach me, those validators run everything a machine can check. Tests, policies, checks. And they do not pass silently. They emit evidence. So when work arrives in front of me it is not asking me to trust it. It is carrying proof.

Explanation. Work arrives with a review document that says what changed, why it exists, and which decisions went which way. If the work cannot be explained, it has not earned my review. I am not reverse engineering intent out of a diff. The explanation shows up with the work.

Risk. Same document, and what matters is what it leaves out. Not every line is in there. Only the code carrying real risk, where my judgment actually changes the outcome. A schema migration and a copy tweak are not the same decision, so they do not get the same attention.

My own throughput went up roughly five times after this was running. That is my own number, self reported, not a measured study.

The same five gates, moving legal cases instead of code

Now the part I actually want you to take away, because none of this is about software.

The gates are not a coding practice. They are the answer to one economic question: what does machine output have to prove before it is allowed to spend an expert's attention?

I designed a system for a lemon law firm. Paralegals use AI to write case summaries and rank how viable each case is, before any of it reaches an attorney. The attorney is the scarce expert there, exactly the way my review was the scarce thing in my own pipeline. The same five gates protect them.

Ownership is single sign on, federated through AWS Cognito. Every case review carries the paralegal who worked it. You cannot touch a case without your name attaching to it. Same gate as my PRD, completely different mechanism.

Scope is a queue the attorney never sees. They hold five cases at a time. Paralegals drop finished work in behind them, and a new case only surfaces when the visible pile drops below the count. That is my token cap translated into a different domain. Bound the work to what the expert can judge without fatigue.

Evidence is statute excerpts. The AI does not hand over a verdict on viability and ask to be believed. It pulls the actual statutory language it used to get there. It shows the law.

Explanation is a per case summary with protocol checkboxes and the paralegal's notes. And the detail I like most: they speak it in through speech to text. That is deliberate, because a gate that is expensive to comply with is a gate people route around.

Risk is a confidence ranking on every case. Anything below the threshold gets flagged, and it goes to a senior paralegal before it ever reaches an attorney. The expensive person is the last stop, not the first.

That firm went from thirty or forty cases a month to around a hundred and fifty. Same attorneys. Nobody new hired. Those are their numbers and mine, not an audited result.

One system moves code to an engineer. One moves legal cases to an attorney. Different technology, different industry, different everything, except the five questions: who owns it, is it bounded, did it prove itself, can a human explain it, and how much attention does it deserve.

What this costs you

This is real work and it is not free. Here is the accounting.

Somebody writes the validators, and that somebody is human. Evidence does not generate itself. Every check a machine runs at the gate had to be specified by a person at plan time. That is upfront fatigue you pay before you get any relief.

Bounding scope multiplies your units. More work contracts, more documents, more ceremony per unit shipped. If your overhead per unit is high, cutting the units smaller makes it worse before it makes it better.

A gate is only as good as the evidence behind it. A validator that passes everything is worse than having none, because now you trust it. Weak checks at the door do not reduce your review burden, they hide it.

People route around gates that are expensive. The speech to text in the legal system exists for that reason, and it is the kind of thing that is easy to skip and fatal to skip. If your gate adds fifteen minutes of typing, it will get filled in badly or not at all.

This is the wrong tool for genuinely exploratory work. You cannot bound a research spike to a work contract with a clean definition of done, because you do not know what done is yet. Do not try. Gate the work that has a shape, and leave the rest ungated on purpose.

The obvious objection is that this is just more process, and more process is what everybody is trying to escape by using AI. Fair. The difference is where the process sits. Traditional process taxes the work on its way out, after a human already built it. These gates tax the work on its way in, before anybody expensive has touched it, and everything they reject costs nobody anything, right?

What to actually take from this

Do not copy my implementation. The pipeline, the validators, Cognito, the queue counts. None of that is the point, and none of it will fit your situation anyway.

Copy the thinking.

  • Find your scarce expert. The person whose attention is the actual bottleneck. Attorney, architect, clinician, approver, security lead.
  • Find what is flooding them. Be specific about the volume and where it enters.
  • Make the work prove itself at the door. Pick one workflow and ask what its output must prove before your expert ever sees it.

Start with ownership and evidence if you want somewhere to begin. Those two are the ones you can put in place without asking anyone to change how they build.

AI made production cheap in every field at the same time, which means every field now has the same problem and most have not named it yet.

When production gets cheap, the gate moves to admission.

Expert attention is the bottleneck. Something has to guard the door.

The walkthrough

I take the five gates through the live system and show what each one actually rejects, including the two that reject almost nothing and are still worth running.

Watch it on YouTube