Last night the data layer under a system I'm building fell over. Disk filled on the EC2 box, and roughly 1.2 million queued files had backed up behind a misbehaving flow.
We cleared the queues and expanded the volume. Everything came back.
Except the agents.
What I'm actually building
The client is a third-party logistics operator running a fleet of around a thousand trucks. Their system of record is an AS/400.
I want to be careful here, because this usually gets written as a joke about old technology. It isn't one. That platform has run a large, complicated, physical business reliably for decades, which is more than most of what I've shipped can claim. The constraint it creates is real engineering, not a punchline.
The constraint: you do not point an AI agent at it directly. Nobody sane would, and their security team would not clear it anyway.
So the agent reaches data through Apache NiFi. Speaking precisely, the AS/400 sits on DB2, and what we actually connect to is DB2. A flow in NiFi talks to that database, and the flow is exposed as an API endpoint the agent calls. NiFi runs in the cloud, tunnelled out, with the whole security posture their IT department signed off on.
That mediation layer is the entire reason this project is possible. It is also, as of last night, the thing I know the most about.
Worth knowing for what follows: that same NiFi cluster carries other heavy workloads. It is not dedicated to the agent.
What broke
NiFi moves work as FlowFiles, and FlowFiles live on disk. Something in one of the tool paths went wrong and they stopped draining. About 1.2 million of them accumulated, the volume filled, and the cluster started hiccuping.
The honest version of causation: we were load-testing the agent at the time. I cannot prove the agent test caused the pileup, and I'm not going to claim I can. The timing points at it hard enough that I'd bet on it.
Good news first. This was staging. No customer touched it, nothing shipped, nobody got charged twice. That matters for the rest of this article and I'll come back to it.
We cleared the backed-up files and gave the volume more room. Standard recovery, nothing clever.
The part that stopped me
The non-agentic consumers of that same NiFi cluster came back on their own. Queues drained, throughput returned, and they behaved exactly as they had before. That's what you expect. That's what recovery means.
The agentic system came back up too. And it was not the same system.
Cost per task was up. The same work, more expensive.
We saw duplicate actions. Things done twice that should have been done once.
Task success sat below where it had been. Not zero. Lower.
Retries were up, and token usage with them.
Nothing was down. Every dashboard was green. The infrastructure had fully recovered and the agent had not.
The same outage. The same recovery. The same cluster.
The normal systems came back. The agentic one carried the outage forward.
So I went looking for why, and it turns out this has a name.
The failure class this belongs to
Let me get the prior art out of the way, because most of this is old and pretending otherwise would be the fastest way to lose you.
What I hit is a metastable failure. A temporary trigger pushes a system into a degraded mode that a feedback loop keeps it in, so it stays degraded after the trigger is gone. Bronson and colleagues named it at HotOS in 2021. Huang and colleagues studied real incidents at OSDI in 2022 and found retry policy was the sustaining effect in more than half of them, with outages running from 90 minutes to three days.
Even the word hysteresis is borrowed. Performance Evaluation was publishing on multi-server threshold queues with hysteresis in 1994, and every autoscaler you've used implements it, with a scale-out threshold deliberately different from the scale-in threshold.
So none of the shape is new. What's different is why the loop stayed closed for the agent and opened for everything else on the same cluster.
Your fan-out factor is not in your code
In an ordinary service, the amplification factor is fixed at deploy time. One request means four queries, because somebody wrote four queries. You can read it in the source.
In an agentic system, that number is an output of the model.
An agent takes a job. It calls a model, which decides to spawn subagents, which call tools, which hit APIs and databases, which sometimes fail, which triggers retries, which are more model calls. Two hundred incoming jobs are not two hundred units of load. The real number depends on what the model decided in the moment.
Now put that under contention. My tool endpoint started timing out, and those timeouts came back as tool results, which means the error message landed in the context window. The model read it and reasoned about it.
And a reasoning model responding to a failure does not simply retry. It might decompose the task further. It might select a different tool. It might spawn another subagent to investigate. Every one of those responses increases fan-out at exactly the moment the system is already saturated, right? Which produces more timeouts, which produces more reasoning about failure.
In a normal system the loop that sustains a metastable failure closes through retry policy, which is a config value you can read.
In an agentic system it closes through reasoning.
Stochastic, content-dependent, and in no config file anywhere.
That's why the other consumers on that cluster recovered and mine didn't. Their amplification was fixed. Mine was a decision the model kept making, and it kept making it based on errors that were no longer happening.
An overloaded agent doesn't return a 500
An overloaded web service errors. You see it.
An overloaded agent has options. It can retry the tool, reason around the failure, pick a different tool, or run an operation it already ran. None of those produce an error your dashboard recognises. That's exactly what I was looking at: green infrastructure, duplicate actions, task success quietly below baseline.
This is the failure mode I keep coming back to. The API returned 200. The answer was wrong. Nobody knew.
And notice this is where the metastable failure literature stops being enough. That work is about goodput collapse, meaning throughput falls and stays fallen. Here throughput was fine. Degraded steady state, expressed behaviorally instead of numerically.
Which is the whole reason the ordinary test would never have caught it.
Why the standard advice would have missed this
Grafana's k6 testing guide is specific about load patterns you should not run.
"Avoid rollercoaster series where load increases and decreases multiple times. These will waste resources and make it hard to isolate issues."
That is good advice, and for a conventional service it's correct. If throughput collapses you want a clean ramp so you can name the level where it broke.
But a clean ramp gives you one measurement per load level. If the question is whether behavior depends on load history, one measurement is exactly the wrong number, because the entire question is whether the same level measures differently at different points in the run.
I found this by accident, through an outage. The test that finds it deliberately is a sweep, not a spike.
10, 50, 100, 200, then back down through 100, 50, 10.
Measure task success at every level in both directions. Plot ascending against descending. If the two paths don't lie on top of each other, the enclosed area is your number.
That loop area is what I'd track release over release. Single scalar, comparable over time, and as far as I can find nobody reports it for agent systems.
One correction to my own first instinct here. Measuring only after the spike does not demonstrate hysteresis, it demonstrates an unrecovered transient. Which is all I actually have from last night. The loop is the evidence, and I don't have the loop yet.
Three ways to fool yourself
I want to be blunt about how a naive version of this test produces a confident wrong answer, because I nearly drew one of these conclusions myself.
Post-spike is often better, not worse. Caches are warm, connection pools are sized, and your autoscaler hasn't scaled back in. Measure 20 agents right after a 200 spike and you may be measuring your autoscaler's cooldown policy. Pin replica count, or record it as a covariate and stop pretending it isn't there.
One cycle is one sample. Agent behavior is stochastic. One sweep gives exactly one measurement per load level, which cannot separate a real loop from noise. You need repeats, distributions instead of point values, and a flat-load control arm running at 20 for the same wall-clock duration.
Order is confounded with height. In a single ascending-then-descending run, time since start correlates perfectly with peak height. Degradation at the final 10 might be the 200 spike or might be cumulative wear. Randomise peak ordering across repeats, or run a spike arm against a no-spike arm.
Skip those three and you'll find hysteresis whether or not it's there, which is worse than not testing.
Make the behavior layer measurable, or drop it
The appealing version of this test measures task success, wrong actions, and policy violations at 200 concurrent agents. Two of those three aren't measurable and I'd rather say so.
"Wrong action" needs a ground-truth oracle. On open-ended tasks at scale there isn't one. Put an LLM judge on it and you've created a second workload with its own cost, latency, and error rate, which you now have to validate under load, which is the same problem again. Put a human on it and the test cannot run at 200 concurrent agents, which was the point.
The version that works: seeded fixtures. Don't run open-ended tasks. Run a fixed catalog of tasks whose correct action set you already know, against an instrumented tool layer. Correctness becomes a set difference between expected actions and observed actions. That computes in milliseconds, at any concurrency, with no judge involved.
Evaluation replaces assertion.
A behavioral metric you cannot compute at load is an assertion wearing a metric's clothes.
Keep in the behavior layer: duplicate side-effecting calls, task success against fixtures, abandoned tasks, turns per task, tokens per task, cost per task. Cut or heavily qualify wrong actions and policy violations.
One honest catch, because a careful reader gets there before I do. The instrumentation that detects duplicate actions, meaning idempotency keys on side-effecting operations, is substantially the instrumentation that prevents them. Fair. But idempotency gives you at-most-once at the boundary, not semantic uniqueness. Two different bookings for the same trip pass every idempotency check you have. That's the failure worth hunting.
Do not run this against live customers
I got lucky on this one. We were in staging, so 1.2 million stuck files and a batch of duplicate actions cost us a night and nothing else.
The mechanism of this test is the harm. You're inducing the conditions that cause duplicate actions in order to observe duplicate actions. Chaos engineering handles that with blast-radius control, but ordinary chaos injects infrastructure faults with known blast radii. Here the blast radius is whatever your agents' tool grants allow, which is the thing under test.
So the design I'd actually run: production infrastructure, production models, production data reads, and a shimmed write layer that records intent and returns realistic responses without executing anything. That preserves the mechanisms that produce the effect, which are rate limits, connection pools, disk, lease timeouts, and real third-party latency, and it removes the part that charges a customer twice.
Two more that are easy to miss. Use a separate provider organisation with its own rate-limit bucket, or you'll degrade your own production traffic. And take the shared-cluster point from my night seriously: your load test's blast radius includes every other workload on that infrastructure. Mine took down a data layer that other jobs depended on.
What this costs you
The token bill is the part everyone assumes is prohibitive, and it's the cheapest thing here.
Roughly a few hundred dollars per full sweep. One cycle is a couple of thousand model calls. On a mid-tier model with prompt caching on, that's tens of dollars. Long-horizon agents on a frontier model push it into four figures. Less than half a day of senior engineer time, so cost isn't your blocker.
The metered tool calls probably cost more than the tokens. Thousands of calls against search, enrichment, and data providers add up faster than inference does, and nobody budgets for it.
The instrumentation is the real price. Distributed tracing across model calls, subagents, and tools. A canonical action log keyed by logical operation. A fixture catalog with expected action sets. A diff harness. Two to six engineer-weeks if you already have tracing, closer to a quarter if you don't. That number, not the tokens, is what you're actually deciding on.
You will break shared infrastructure. If your mediation layer serves anything else, and mine does, a serious agent load test is an outage for those consumers too. Either isolate it or schedule it like the disruptive thing it is.
Some of what you measure is not yours to fix. Provider-side rate limiting, capacity, and backoff produce hysteresis you cannot control or reproduce reliably week to week. Split findings into the loop you own, meaning leases, pools, queues, idempotency, and context accumulation, and the loop your provider owns.
When to skip it entirely. Agents with no side effects, or a fan-out factor near one, or anything where a human approves every action before it lands. The test earns its cost when amplification is high and actions are irreversible.
Where I actually am
I have one incident, not a study. I watched an agentic system fail to recover from an outage that everything around it recovered from, and I have the four symptoms it carried forward. I don't yet have the loop measured in both directions, which is the evidence that would turn this from an observation into a number.
That's the next thing I'm building, and I'd rather tell you that than dress up one bad night as a methodology.
Don't ask whether your agents survive 200 concurrent jobs.
Ask whether 20 behaves the same after 200 as it did before.
Because silent failure is the AI failure mode, and a system whose correctness depends on the load it experienced an hour ago will pass every steady-state evaluation you own. Mine did.
If you run the sweep before I do, I want to hear what broke.
The sources, if you want to go deeper. Metastable Failures in Distributed Systems (HotOS 2021) names the failure class. Metastable Failures in the Wild (OSDI 2022) has the incident data. The AWS Well-Architected Agentic AI Lens is where the concurrency and contention guidance lives, and k6's test types guide is the advice I'm arguing with.
Nothing above depends on reading any of them.
Chris