In November 1910, six men travelled separately to a rail terminal in New Jersey. They boarded a private car and rode to a hunting club on an island off the coast of Georgia. They stayed ten days. They told their families it was a duck hunt.
They denied the meeting had happened for the next twenty years.
Senator Nelson Aldrich, who chaired the National Monetary Commission. His private secretary, Arthur Shelton. Henry Davison, a partner at J.P. Morgan. A. Piatt Andrew, Assistant Secretary of the Treasury. Frank Vanderlip, president of National City Bank. Paul Warburg, a partner at Kuhn, Loeb.
On the train they used first names only. Nelson, Harry, Frank, Paul, Piatt, Arthur. That way the staff could not identify them. For decades afterward they called themselves the First Name Club.
They were there to design a central bank for the United States. Three of the six ran, or were partners in, the largest institutions the thing would govern.
This is not a conspiracy theory. It is on the Federal Reserve's own history site, and the Fed states the reason for the secrecy plainly:
Aldrich and Davison chose the attendees for their expertise, but Aldrich knew their ties to Wall Street could arouse suspicion about their motives and threaten the bill's political passage.
So the popular claim has real ground under it. The meeting happened. The bankers were in the room. The secrecy was deliberate.
But it’s still incomplete. It stops one chapter early, and the chapter it skips is the only one that matters.
What they actually wrote
What came off the island was the Aldrich Plan. A Reserve Association of America, one central bank with fifteen branches. Here is how it would be governed, verbatim from the Fed:
Each branch would be governed by boards of directors elected by the member banks in each district, with larger banks getting more votes.
There it is. Not a hidden agenda. A governance clause, written down openly, by people who sincerely believed the thing was good for the country. The Commission's own covering letter called the institution "scientific in its method, and democratic in its control."
Capture does not always arrive hidden.
"With larger banks getting more votes." They wrote it down. They thought it was good policy.
Which is what makes it useful, and what makes the villain version of this story worthless.
Then the plan died, and it died on that clause
Democrats objected to exactly that sentence. The Fed's words: they "objected to the version of democracy it presented, which could have allowed the largest banks to exert outsized influence on the central bank's leadership." Repudiating the Aldrich Plan went into the party platform. Wilson won. The plan was shelved.
What passed in December 1913 was the Federal Reserve Act. And here is the sentence that should reorganize how you think about this:
The technical details of the final bill closely resembled those of the Aldrich Plan. The major differences were the political and decision-making structures.
Look at what that means in the actual governance.
The Aldrich Plan, rejected. A forty-six-member board with only six appointed by the government. And the head of the organization was to be selected from a list of three names supplied by the association itself. Branch boards elected by member banks, weighted by size.
The Federal Reserve Act, enacted. A Federal Reserve Board with supervisory authority over the banks, made up entirely of presidential appointees. Staggered terms, extended to ten years so that no president could appoint all of them across two terms.
Same machine. The elastic currency, the pooled reserves, the regional branches, the discounting of commercial paper, all of it survived. What changed was who counts the votes.
And they kept Warburg. Carter Glass, Robert Owen, their staffs, Wilson's adviser Colonel House, and two future cabinet officers all consulted him directly while drafting the bill that replaced his own plan. They did not purge the expertise. They moved the authority and kept the knowledge.
The regulated wrote the first draft. It was rejected because it let them count the votes.
The final law kept their architecture and replaced their control structure.
Same engineering. Different apex. That is the entire fix.
Capture is structural, not moral
The six men were not villains, and the story gets much more useful once you stop trying to make them into any. They were the most competent people available. They diagnosed the problem correctly. American panics really did arrive every fifteen years, reserves really were immobilized, the currency really was inelastic. Their technical work was good enough to survive their own political defeat almost intact.
The flaw was not in their motives, right? A room of insiders, working in good faith, in seclusion, for ten days, produced a design with a hole in it that was obvious to the first outsider who looked. They could not see it. The thing they could not imagine was themselves being the problem.
So the point is not "keep interested parties out of the room." That version leads to bad engineering. The point is this:
You do not fix capture by finding a checker with no conflicts. You fix it by making sure the checker cannot fail the same way as the checked.
Three jobs, and only one of them is safe
Which brings us to the thing you were probably going to do this quarter. Have a model help build its own evaluation. It is tempting, it is cheap, and it gets described as one activity, honestly. It is three, and they carry completely different risk.
Generating inputs. Proposing test cases, adversarial prompts, edge cases, weird formats. It does not require being right about the answer, because a human still labels it. Defensible.
Defining correct. Authoring the oracle. What the right answer actually is. No.
Deciding it passed. Adjudicating the run. Holding the vote. No.
The middle one is where people go wrong, and it is not a bias problem. It is capability circularity. If a model could reliably produce the correct expected output for a hard case, it could reliably produce the correct output in production. So a self-authored oracle is most reliable exactly where you need it least, and silently wrong exactly where you need it most.
And the first one is not clean either. A model proposes the failure cases it can conceive of, and its conception of the failure space comes from the same weights that produce the failures. The cases it cannot imagine are disproportionately the cases it will fail. That is a coverage problem, not a bias problem, and validating the generated tests does not touch it, because validation only inspects the tests that exist.
Six competent men. Ten days. A hole they could not see.
A model authoring its own oracle is not being dishonest.
The competence needed to write a valid oracle is the competence under test.
Which is why it works on the easy cases and fails quietly on the ones you built it for.
So build a checker that cannot fail your way
The instinct most teams have here is the right one. If your model has blind spots, bring in models that do not share them. Run five from different providers and let them propose the scenarios yours would never think of.
The instinct is right, and the usual wiring quietly reverses it. Here is the wiring that holds. Every piece of it exists because of a specific way this turns back into an echo.
Measure the swarm's independence before you trust it. Different models are not independent failure sources. They share overlapping pretraining corpora, near-identical architectures, similar post-training methods, convergent safety tuning, and they are trained on each other's output. Vendor diversity is a procurement fact, not a statistical one. So a swarm reduces variance and does not touch bias. If they share a blind spot, five models miss it five times and hand you high confidence in a wrong answer. That is worse than one model missing it once, because now there is a consensus artifact to point at in the review. Independence is measurable, so measure it. Take cases you already know are hard, with known ground truth, and see how much the models disagree. Low disagreement on genuinely hard cases means your swarm is one model wearing five hats.
Give them the scope as information, not as a fence. Outside models do not know your project, so they will propose work that is out of bounds. The obvious move is to tighten the requirements until they stop. Do not. Every tightening round converges the swarm onto the requirement author's model of the world. And that author is the same party whose model of the world produced the system under test. The variance you are squeezing out is the blind-spot coverage you paid for. Run it long enough and you have built an expensive echo.
Classify, never filter. "Out of bounds" is doing far too much work. It collapses three different things. Genuinely out of scope, which is noise. Inside scope but nobody wrote the requirement down, which is a requirements gap. And outside the declared scope because the scope is wrong, which is the most valuable thing you will get all quarter. Filter before you classify and the last two die silently. You brought in outsiders to see what insiders cannot, so do not build a step that throws away the outsider's objection for being outside. Label every scenario and keep the reasoning. The boundary judgments become the artifact instead of a deletion nobody witnessed.
Name the adjudicator, and put them outside the team that owns the scope. Somebody has to rule on whether a scope error is real. If that is the team that wrote the scope, you have rebuilt the Aldrich Plan exactly. The regulated adjudicating objections to their own design, in good faith, with excellent technical credentials. This is the presidentially appointed board. It is the only step here that actually relocates an authority, and it is the one most likely to get dropped for being inconvenient.
Run traceability as bookkeeping, not as a gate. Bidirectional traceability, every requirement reaching a test and every test tracing back to a requirement, is the right instrument. Putting it at the end as the gate inverts what it can do. It verifies coverage of stated requirements and is structurally blind to unstated ones. So a clean audit proves internal consistency. It says nothing at all about the requirement nobody wrote. That is the exact failure class the swarm is there to find. It will pass, and it will pass for the wrong reason. Run it continuously instead, and treat an orphan test, one that traces to no requirement, as the prize. It is usually an unwritten requirement in disguise.
Seed defects you already know about. Without this the swarm cannot be wrong. If it comes back with nothing, there is no way to tell "no problems exist" from "the swarm is blind." And an empty result will feel like good news. Plant a held-out set and measure whether it gets found. Until you have done that, the process has never been tested. It has only been run.
Notice what none of that does. It does not remove the models, or the requirements, or the audit. Every component is still in the design. What changed is the order, and which chair the deciding vote sits in.
Which is exactly what happened in 1913.
What this costs you
Relocating an authority is not free, and the version of this argument that skips the bill is not worth much.
You are adding a person who can say no and does not own the outcome. That is the entire mechanism, and it is also the thing your delivery schedule will hate first.
The adjudicator has to be senior enough to overrule the scope owner and available enough to actually do it. On most teams that person does not exist, and inventing a junior one gives you the ceremony without the independence.
Classifying instead of filtering means you keep the noise. Somebody reads all of it. Filtering was not a mistake anyone made for fun. It was just cheaper, and reading everything is kind of the price of the fix.
Seeded defects are real work and they go stale. A held-out set the team has seen is no longer held out.
Measuring inter-model disagreement costs a run and a labelled hard set, and it can tell you the swarm you already bought is one model wearing hats. That is an expensive thing to learn out loud.
None of this is free of the same trap. Whoever designs the adjudication sits above it. You have moved the problem up one floor, which is progress and not escape.
What you can lean on, and what you cannot
Every quote above comes from the Federal Reserve's own account of itself. That is worth something. They are the ones documenting the secrecy, the Wall Street ties, and the capture clause.
The evaluation half is different. The three jobs and the six rules are an argument, not a finding. Nothing here was measured, no study is being cited, and sitting next to well-sourced history does not upgrade it.
The analogy has a limit too. Warburg was safe to consult because of what surrounded him. Legible incentives, published reasoning, and opponents who fought him in public for years. A model contributing to its own evaluation has none of that. So "they kept Warburg" is a principle about relocating authority. It is not permission to let the system under test write the test.
They moved the votes
A system cannot certify itself. Not because it is dishonest, but because its blind spots and its failures come from the same place.
Six of the most capable financial minds in America spent ten days in seclusion and produced a design with a hole in it that the first outsider spotted immediately. They were not corrupt. They were inside.
They did not purge the bankers. They kept the architecture and kept the man who wrote it.
Do not look for a checker with no conflicts. Find one that cannot fail the same way as the checked.
Then go and find out who, on your team, is currently holding all three jobs.
