Somewhere in your company right now, somebody is about to approve an agentic AI programme. Not because the case is strong. Because saying no looks like standing still.
I want to show you what standing still is actually worth. It starts in a goalmouth.
The best move is to stand still. Almost nobody does.
In 2007, Michael Bar-Eli and four colleagues analysed 286 penalty kicks from top leagues and championships. They mapped where the ball went, where the keeper went, and what got stopped.
The kicks spread out. 32.2% went to the keeper's left, 28.7% down the middle, 39.2% right. The keepers didn't spread. They dove in 93.7% of kicks. Across 286 penalties, they stayed in the centre 18 times.
Now the part that matters. Dive left and you stop 14.2% of kicks. Dive right, 12.6%. Stay in the centre and the keepers in this sample stopped 6 of those 18. That's one in three, versus roughly one in seven for diving. More than twice as effective.
So the measured best move is the one keepers chose 6.3% of the time.
More than a quarter of penalty kicks go down the middle.
28.7% of kicks go down the middle. Keepers stay there 6.3% of the time.
The best move in the data is the one almost never chosen.

So why do keepers dive?
Not because they can't do the maths. The authors' explanation is norm theory. Diving is what a keeper is supposed to do. So the paper found a goal "yields worse feelings for the goalkeeper following inaction (staying in the center) than following action (jumping), leading to a bias for action."
A goal after a dive is bad luck. A goal while standing still is a failure to try. Same goal, different shame.
And look at who this is happening to. Elite professionals. Enormous incentives. A decision they have faced their whole careers. Expertise didn't fix it. Where acting is the norm, not acting registers as risk, and the numbers stop mattering.
That's the mechanism, right? Now carry it out of the goalmouth, because it's running in your building on three floors at once.
Floor 1: the board
Nobody gets congratulated for the agentic AI programme they declined to start. Building is observable. It's a slide, a headline, a line in your performance review. Waiting leaves no artefact at all.
So the decision to build gets made on visibility, and the business case gets written afterwards to match. The board is the keeper, and the crowd is watching.
Floor 2: the system
Here's the thing about the agents themselves. Give one of these systems a task and it will call a tool, because "take no action" was never defined as a legal outcome with value attached. If doing nothing isn't in the action set, the system cannot select it, no matter how good the model gets. There are now benchmarks, BFCL among them, that score whether a model can decline to act. They exist because models are bad at exactly that.
Why did we build them this way? Here's what I think happened, and I'm flagging it as my read of it: the programme upstairs had to show value. A system that quietly decides nothing needs doing produces no demo, no token count, no slide. So we shaped systems to produce visible activity. We built keepers that aren't allowed to stand still.
Floor 3: 2 a.m.
You're on call. The quality signal wobbles. Four bad outputs in an hour, on an agent whose metrics are noisy and lagging on a good day. You know four outputs prove nothing. You also know that if you wait and it gets worse, tomorrow's question is "you saw it and did nothing?"
So you restart the service, and when that changes nothing you swap the model, and when that changes nothing you edit the prompt, and by morning you have made three production changes on evidence that could not support one. Waiting for real evidence would have been right. Waiting also looks like negligence.
And honestly, Floor 3 is only this violent because of Floor 1. The thing shipped before the data was under control and before the workflows had settled. The 2 a.m. decision is the boardroom decision, arriving late, holding a pager.
Which raises the question the whole industry is dodging: what does "showing value" even mean?
Can anyone prove any of this is working?
Two credible 2025 reports point in opposite directions. MIT's NANDA initiative found 95% of organisations getting zero return on $30 to $40 billion of enterprise GenAI investment. Wharton, same year, found three in four leaders reporting positive returns, with 72% formally measuring ROI.
Both are real. The gap is the definition. MIT counted success as deployment past pilot with measurable KPIs, checked six months out. Wharton asked leaders whether they saw positive returns. That's the entire disagreement, and the gap itself is the finding.

MIT says 95% get nothing. Wharton says three in four see returns. Same year, same subject.
An AI value claim with no measurement definition attached carries no information.
If nobody defines "working" up front, the only available proof is visible activity.
And that loops straight back to Floor 2. When there's no agreed number, tokens burned becomes the evidence of value. The most valuable thing an agent can do, correctly deciding that nothing needs doing, is invisible. Remember the keeper: the save you make by standing still never shows up as a save.
One more MIT number worth keeping. Bought beat built, roughly 67% deployment versus 33%. The report's own caveat is that this may reflect organisational capability more than approach. Directional, but the burden of proof sits on building.
Why this won't feel like a bias
Standing still will feel wrong to you even after you've seen the numbers. It's worth knowing why, because Schopenhauer described your next steering committee in 1851.
First, on how agentic AI got onto your agenda at all:
"A man never feels the loss of things which it never occurs to him to ask for; he is just as happy without them... every man has an horizon of his own, and he will expect as much as he thinks it is possible for him to get."
Nobody suffered from not having an agentic AI programme in 2022. It wasn't in anyone's horizon. Then it entered every executive's horizon in the same few quarters, and from that moment not having one became a felt loss. The business case didn't change. The horizon moved.
Second, the sharper one:
"Fame is something which must be won; honor, only something which must not be lost. The absence of fame is obscurity, which is only a negative; but loss of honor is shame, which is a positive quality."
That is the goalkeeper, exactly. Honour is meeting the expectation. Diving is the expectation. The keeper who dives and concedes loses nothing. The keeper who stands and concedes has proved false to the norm, and that lands as shame, which is felt. The executive who correctly declines isn't risking obscurity. They're risking shame: being the one who was slow, who didn't see it. Shame arrives and gets experienced. The disaster you avoided never arrives, so it never gets credited.
He wrote that 135 years before Kahneman and Miller formalised it as norm theory, and 156 years before Bar-Eli measured it in a goalmouth. To be straight with you about what that convergence is worth: three matching descriptions are not three measurements. Philosophy explains the trap. It doesn't prove you're in it.
But it does tell you the "wait" option will feel like cowardice in the room. Expect that feeling. It's the bias arriving on schedule, dressed as judgment. You don't beat it by arguing with it live. You beat it by writing the rules down before you're in the room.
Five gates, in the order they kill you
A failure at Gate 1 cannot be repaired at Gate 5. Answer honestly. "We're working on that in parallel" is a failed gate wearing a disguise.
Gate 1. Define "working" as a number before anything is built. Which number moves, by how much, measured how, by when, and who agreed that's the number. If the answer is "efficiency" or "we'll know it when we see it," stop. This gate is the whole MIT versus Wharton gap, settled in advance.
Gate 2. Make "the agent did nothing" a legal, valuable outcome. Can the agent close a task by taking no action, and is that logged as a success? If your metric scores it as a failure, stop. Prevented work is real value. Make it countable or it gets optimised away.
Gate 3. Foundations finished, not maturing. Is the data this agent reads under change control, with known freshness and a named owner? If the honest answer is "we're cleaning that up alongside the build," stop. Everything built on a moving foundation is an incident you have scheduled in advance.
Gate 4. A real reason to build instead of buy. Name the specific capability you cannot buy. "Control" and "we're an engineering company" don't count. In MIT's sample, bought deployed twice as often as built.
Gate 5. Write the waiting rule before launch. At 2 a.m., with a degraded signal: how many bad outputs, over what window, before anyone touches production? Who can authorise waiting? If the answer is "use judgment," stop.
Judgment at 2 a.m. under pressure is action bias with a lanyard on.
Set the threshold while you're calm, and make waiting an authorised act instead of a personal risk.
The rule takes the blame so the engineer doesn't have to.
What the gates cost you
The gates are not free, and pretending they are would be its own kind of slide.
Gate 1 costs a political fight before there's anything to show. Agreeing whose number counts is real work, and it lands when enthusiasm is highest and patience is lowest.
Gate 2 costs metric complexity. "Did nothing, correctly" needs its own definition and someone auditing it, or it becomes cover for a system that's simply broken. You are signing up to tell those two apart forever.
Gate 3 costs the calendar. Data under change control before you build can mean quarters. Your competitor's demo does not wait for your data owner.
Gate 5 will burn you at least once. A threshold strict enough to stop 2 a.m. poking will eventually hold someone still through a real failure. The rule takes the blame instead of the engineer. That's the point, and it will still hurt that week.
And the whole list costs visibility. Run every gate seriously and sometimes the outcome is that you don't build. The disciplined no is invisible. Nobody will ever know what it saved. That's the exact asymmetry this piece is about, now pointed at you.
When are the gates the wrong tool? A contained sandbox pilot: cheap, reversible, no production data, no customer in the blast radius. There you can buy learning with action, and moving fast is the cheaper way to learn. The gates protect production launches. Don't use them to strangle an experiment.
What to actually take from this
Standing still is a decision. Score it like one, against the same number the build would be scored against.
Define "working" as a number before you build. The MIT and Wharton reports disagree because nobody does this.
Put "do nothing" in the agent's action set with value attached. If it can't be selected, it won't be.
Write the waiting rule while you're calm. 2 a.m. is too late.
The keeper who dives and concedes loses nothing. The keeper who stands and concedes proved false to the norm. Until your organisation can credit a correct no, everyone in it keeps diving.
Before you quote any of this
The centre save figure rests on 18 keepers staying put, 6 saves against 12 goals. The paper's statistical test holds, but say "more than twice as effective," never a precise multiple. It's one sport, observed in 2007, no experiment. The claim that agents lack a null action is checkable and true. The claim that it happened because leaders needed visible value is my argument, and I've marked it as such. Schopenhauer explains the mechanism; he is not a third piece of evidence. And the 95% figure is stricter than the headlines made it sound: zero measurable return under a demanding definition, from 52 interviewed organisations and 153 surveyed leaders. Skip the Fortune write-up and go to the report.
Sources
Bar-Eli, Azar, Ritov, Keidar-Levin & Schein, "Action bias among elite soccer goalkeepers: The case of penalty kicks," Journal of Economic Psychology 28(5), 2007. Figures from Tables 1 and 2 of the MPRA working paper (No. 4477).
MIT NANDA, The GenAI Divide: State of AI in Business 2025, July 2025. 300+ initiatives, 52 organisation interviews, 153 senior-leader surveys.
Wharton Human-AI Research, Accountable Acceleration: GenAI Fast-Tracks Into the Enterprise, 2025.
Schopenhauer, Aphorisms on the Wisdom of Life (1851), chs. 3 and 4, T. Bailey Saunders translation, Project Gutenberg ebook 10741.
Kahneman & Miller, "Norm theory: Comparing reality to its alternatives," Psychological Review 93(2), 1986. Cited via Bar-Eli et al.