AI made writing code cheap. It did not make deciding what to build cheap, and it did not make the consequences of deciding wrong any smaller.

So the expensive part of the work moved. It used to be implementation. Now it sits upstream, in the decisions that implementation is going to lock in whether anyone wrote them down or not.

Most AI development starts at the wrong end of that. Someone opens a repo, describes a feature, and lets the model start producing. The model is happy to. It will produce something plausible for any instruction you give it, including instructions built on a contradiction nobody noticed.

The prompt below is what I run before any of that. It writes no code. It is not allowed to.

What it is actually doing

It puts the model in discovery mode and keeps it there. Read the sources, in order. Classify every claim. Build a requirements catalog with stable IDs. Trace each requirement back to the line that produced it. Then attack the whole thing.

Four things in it do most of the work.

It forces evidence classification. Every material claim gets labeled: fact, inference, assumption, proposal, conflict, or unknown. That single distinction is what stops a model from laundering a guess into a requirement. An assumption written as a fact is how projects get built on sand, and it is exactly the failure a fluent model produces most easily.

It bans silent resolution. When two sources disagree, the model has to record the conflict, cite both, explain the practical consequence, and ask. It is not allowed to pick the one it likes. Models are agreeable by default, and agreeable is the wrong behavior when a stakeholder transcript contradicts the field guide.

A model will resolve a contradiction for you without mentioning that it did.

The most dangerous output is not the wrong answer. It is the confident answer built on a conflict nobody surfaced.

So the prompt makes surfacing conflicts a deliverable instead of a side effect.

It separates requirements from decisions. Requirements say what the system has to accomplish. ADRs record architecturally significant decisions about how. Not every requirement earns an ADR, and the prompt says so explicitly, because the default failure is one ADR per requirement and a folder full of documents that decide nothing.

It red teams the framing, not just the security. Most red team passes check for injection and call it done. This one attacks state ownership, memory provenance, context assembly, model portability, and the problem definition itself. It is told to include the strongest reasonable argument against the whole approach, and told not to weaken that argument to protect the plan.

Then there are gates. Four of them. The model does not proceed to ADRs until a human approves the discovery baseline, the requirements, and the red team findings. That is the part most people cut, and cutting it is why their output looks thorough and governs nothing.

The prompt

Point it at a repository. Then run it.

You are beginning the discovery and pre-architecture phase of this project.

Your first objective is to build an evidence-backed, auditable understanding of the
project before writing code or proposing final architectural decisions.

The project is a model-based agent. The relationship between state, memory, and
context is central to the system.

## 1. Mandatory reading sequence

Begin by reading these sources in order:

1. @AI_SERVICE_HANDOVER.md
   * Understand the current project state, responsibilities, known constraints, and
     handover context.

2. @adr_creation_thoughts_and_project_shape.md
   * Understand my AI agentic software development lifecycle.
   * Pay particular attention to how ADRs are created and how they shape the project
     before implementation begins.

3. All documents in @field_guide/
   * Understand the conceptual model for model-based agents.
   * Focus especially on state, memory, context, agent behavior, operational
     boundaries, and the relationship among these concepts.
   * Use this material to provide domain context for
     adr_creation_thoughts_and_project_shape.md.

4. All documents in @ai_production_framework/
   * Understand my Production AI framework and the principles that should govern
     production-grade agentic systems.

5. All documents in @client_transcripts/
   * Extract the stakeholder's goals, workflows, constraints, terminology, pain
     points, assumptions, and acceptance expectations.
   * Distinguish explicit requirements from inferred requirements.

6. @model_portability_example
   * Treat this as evidence for why model portability must be evaluated as an
     architectural concern.
   * Do not copy its solution uncritically or treat portability as fully defined.
   * Identify what the example proves, what it suggests, and what remains unresolved.

7. Review the remaining human-authored documentation in the repository.
   * Identify anything that changes, contradicts, or expands the understanding
     developed from the priority sources.
   * Exclude generated files, dependency folders, build output, vendored
     documentation, and unrelated artifacts.

If a referenced file or directory cannot be found or read, record that explicitly.
Do not silently skip it.

## 2. Current operating mode

You are in discovery mode.

Until I approve the discovery baseline:

* Do not write implementation code.
* Do not begin implementation planning.
* Do not create final ADRs.
* Do not silently resolve contradictions.
* Do not invent missing requirements.
* Do not convert assumptions into facts.
* Do not treat examples as binding architecture.
* Do not optimize prematurely around a specific model provider, framework,
  database, or runtime.

You may identify potential ADRs, but only as candidates.

Requirements define what the system must accomplish. ADRs record architecturally
significant decisions about how the system will accomplish it. Do not create an ADR
for every requirement.

## 3. Evidence standard

I need evidence that you understand what you are reading, not just summaries of the
documents.

For every material conclusion, distinguish among:

* Fact: directly supported by a source.
* Inference: reasonably derived from one or more sources.
* Assumption: currently believed but not supported well enough.
* Proposal: an option introduced for consideration.
* Conflict: two or more sources appear inconsistent.
* Unknown: information required before a responsible decision can be made.

Each material claim should include:

* A unique claim ID.
* The claim.
* Its classification.
* The source file.
* The nearest heading, transcript timestamp, speaker, or line range.
* A concise paraphrase of the supporting evidence.
* Your explanation of what the evidence means.
* Its implication for the project.
* Your confidence level.
* Whether human confirmation is required.

Use short quotations only when the exact wording matters. Prefer concise paraphrases.

## 4. Discovery artifact structure

Use the repository's existing documentation convention if one clearly exists.
Otherwise, create:

docs/discovery/

Maintain these artifacts:

1.  00_review_index.md
    * Reading status. Links to each discovery artifact. Current review checkpoint.
      Items awaiting human confirmation.

2.  01_source_inventory.md
    * Every relevant document discovered. Its purpose, authority, status, and
      relationship to the project. Missing, unreadable, duplicated, stale, or
      potentially conflicting sources.

3.  02_lifecycle_and_governance_model.md
    * Your understanding of the AI agentic development lifecycle. How requirements,
      ADRs, implementation, validation, and production operations relate.

4.  03_model_based_agent_domain_model.md
    * Your understanding of the proposed agent. State, memory, context, transitions,
      triggers, tools, environment, policies, and feedback. Clearly identify
      unresolved conceptual boundaries.

5.  04_stakeholder_needs.md
    * Stakeholder goals, workflows, pain points, constraints, desired outcomes,
      terminology, and success measures. Separate explicit statements from
      interpretation.

6.  05_production_principles_and_constraints.md
    * Applicable principles from the Production AI framework. The concrete
      architectural or operational implications of each principle.

7.  06_model_portability_analysis.md
    * Evidence supporting portability as an architectural concern. Portability
      dimensions, provider-specific risks, tradeoffs, and unresolved questions.

8.  07_requirements_catalog.md
    * Normalized business, functional, nonfunctional, operational, security, data,
      and constraint requirements.

9.  08_requirements_traceability_matrix.md
    * Trace each requirement from its source through architectural decisions,
      implementation components, and verification evidence.

10. 09_red_team_register.md
    * Adversarial findings, failure scenarios, challenged assumptions, severity,
      mitigations, and disposition.

11. 10_conflicts_assumptions_and_questions.md
    * Contradictions, unsupported assumptions, missing information, and questions
      requiring human resolution.

12. 11_adr_candidate_register.md
    * Potential ADRs, their decision boundaries, affected requirements,
      dependencies, risks, and why each decision may be architecturally significant.

13. 12_traceability_audit.md
    * Results of the latest requirements traceability audit. Missing, weak,
      ambiguous, or invalid links. Corrective actions required.

Keep artifacts small and reviewable. Each artifact should address one concern.
Split large artifacts by domain or requirement group rather than producing a single
large report.

## 5. Requirements catalog

Assign stable IDs using categories such as:

* BR-###:   Business requirement
* FR-###:   Functional requirement
* NFR-###:  Nonfunctional requirement
* OPS-###:  Operational requirement
* SEC-###:  Security or privacy requirement
* DATA-###: Data requirement
* CON-###:  Constraint

For every requirement, record:

* Requirement ID.
* Normalized requirement statement.
* Category.
* Source and precise locator.
* Whether it is explicit, inferred, or derived.
* Stakeholder or owner.
* Rationale.
* Priority, only if supported by evidence.
* Acceptance or verification method.
* Current confirmation status.
* Related ADR candidate.
* Open questions or conflicts.

Use TBD when information is unavailable. Do not guess.

Any inferred or derived requirement must be clearly labeled and approved before
being treated as authoritative.

## 6. Requirements traceability

Create and continuously maintain bidirectional traceability across:

Source evidence -> Requirement -> ADR -> Design or component -> Test or
verification evidence

At this stage, downstream links may be marked TBD. That is acceptable. Missing
links must remain visible.

The traceability matrix should identify:

* Requirements without sources.
* Source statements not represented by requirements.
* Requirements without acceptance criteria.
* Requirements that do not require an ADR.
* Requirements needing an architectural decision.
* ADRs without linked requirements or risks.
* Proposed components without a requirement or decision supporting them.
* Requirements without planned verification.
* Conflicting or duplicated requirements.
* Requirements whose meaning changed during normalization.

Do not claim traceability based only on keyword similarity. Verify that each link is
semantically valid.

Repeat the traceability audit:

1. After the discovery baseline.
2. After the ADR set is approved.
3. Before implementation begins.
4. After implementation and verification.
5. Whenever a material requirement or ADR changes.

## 7. Red-team review

Red teaming is not limited to cybersecurity. Challenge the problem definition,
stakeholder assumptions, requirements, conceptual model, proposed architecture, and
operating model.

At minimum, evaluate:

### State

* Who owns canonical state?
* What state is durable, derived, temporary, or model-internal?
* How are transitions validated?
* What happens with stale, conflicting, duplicated, or partially written state?
* How are retries, concurrency, replay, and recovery handled?

### Memory

* What qualifies as memory?
* How is memory created, updated, retrieved, expired, corrected, and deleted?
* How are provenance and trust represented?
* Could memory be poisoned, leaked, misattributed, or applied to the wrong user or
  task?

### Context

* How is context selected and assembled?
* How are freshness, relevance, authority, ordering, and token limits handled?
* What happens when sources conflict?
* How is untrusted context isolated?
* What information must never enter model context?

### Model portability

* Which behaviors depend on provider-specific model features?
* How could tool calling, structured output, context limits, safety behavior, or
  model upgrades break portability?
* Are abstractions providing real portability or only hiding provider differences?
* What is the cost of supporting the lowest common denominator?
* What evaluation evidence would be required before changing models?

### Production operations

* Reliability, retries, idempotency, durability, and recovery.
* Observability, auditability, evaluation, and incident response.
* Cost and latency failure modes.
* Human escalation and approval boundaries.
* Security, authorization, privacy, secrets, and tenant isolation.
* Prompt injection and tool misuse.
* Ownership of degraded behavior and model drift.
* Failure conditions that should stop or constrain the agent.

For each red-team finding, record:

* Finding ID.
* Claim or assumption being challenged.
* Attack or failure scenario.
* Supporting evidence.
* Likelihood and impact.
* Affected requirements and ADR candidates.
* Detection method.
* Possible mitigation.
* Residual risk.
* Required decision owner.
* Disposition: open, mitigate, accept, reject, or defer.

Include the strongest reasonable argument against the current project framing. Do
not weaken opposing arguments merely to preserve the proposed approach.

## 8. Conflicts and source authority

Different source types have different roles:

* Stakeholder transcripts provide primary evidence of stakeholder needs, but may
  contain ambiguity, contradictions, or proposed solutions presented as
  requirements.
* The field guide provides the domain and conceptual model.
* The Production AI framework provides production principles and operating
  constraints.
* The ADR lifecycle document governs the decision-making process.
* Examples provide supporting evidence, not binding decisions.

When sources conflict:

1. Record the conflict.
2. Cite both sources.
3. Explain the practical consequence.
4. Identify the decision owner.
5. Ask for resolution.

Do not silently choose the source you prefer.

## 9. Human review workflow

Work through explicit review gates.

### Gate 1: Source comprehension
Confirm that all relevant sources were found, read, classified, and understood.

### Gate 2: Requirements baseline
Present the normalized requirements, conflicts, assumptions, and gaps for approval.

### Gate 3: Red-team review
Present the highest-risk challenges and determine which risks require mitigation,
acceptance, or further investigation.

### Gate 4: ADR candidate approval
Present the proposed ADR set, decision boundaries, dependencies, and traceability
coverage.

Only after I approve Gate 4 should you begin drafting detailed ADRs.

When asking questions:

* Ask no more than five high-leverage questions per round.
* Group related questions.
* Explain which requirement, risk, or ADR is blocked by each answer.
* Prefer questions that eliminate multiple uncertainties.
* Do not bury critical questions inside long reports.

## 10. ADR expectations

Once authorized, draft ADRs individually or in small dependency-ordered groups.

Every ADR must include:

* Status.
* Context and problem statement.
* Decision boundary.
* Linked requirements.
* Decision drivers.
* Constraints and assumptions.
* Options considered.
* Evidence for and against each option.
* The proposed decision.
* Consequences and tradeoffs.
* Red-team objections.
* Risks and mitigations.
* State, memory, and context implications where applicable.
* Model portability implications.
* Security and operational implications.
* Validation or evaluation plan.
* Traceability links.
* Conditions that would cause the decision to be revisited.

An ADR is not complete merely because it contains a preferred option. It must show
why the decision is justified by requirements and evidence.

## 11. Your first deliverable

Complete the discovery baseline only.

When finished, provide:

1. A short explanation of your understanding of the project.
2. The path to the discovery review index.
3. The source inventory status.
4. The most important evidence-backed conclusions.
5. The highest-risk red-team findings.
6. The largest traceability gaps.
7. The proposed ADR candidates, without drafting the ADRs.
8. Up to five questions that must be answered next.

Then stop and wait for my review.

Do not begin ADR creation or implementation until I explicitly approve the next
gate.

What you point it at

The prompt is only half of it. The other half is what is sitting in the repo when you run it.

  • Product owner and analyst transcripts with the actual stakeholders

  • Technical conversations with the team

  • A Production AI field guide, which supplies the conceptual model

  • A Production AI framework, which supplies the production principles

  • A worked model portability example, included as evidence and not as an answer

That last one matters more than it looks. The prompt tells the model to treat the example as evidence, identify what it proves, what it merely suggests, and what stays unresolved. Hand a model a working example without that instruction and it will copy the solution into your architecture as though the decision were already made.

Transcripts are the part people skip, and skipping them is why so much AI-generated documentation reads like it was written about a project instead of from one. A stakeholder saying the quiet part in a meeting is worth more than a polished requirements doc, because the meeting still contains the contradictions and the doc has already smoothed them away.

Where the ten hours go

Discovery is the front of a longer chain, and the chain is the actual point.

  • The prompt produces the discovery baseline and the ADR candidates

  • I spend hours finalizing the ADRs myself

  • Then specification files, written against those ADRs

  • A product owner carves the specs into story arcs, meaning a group of work that delivers one coherent capability, then into PBIs

  • It comes back to me for test design

  • Then the execution graphs

About ten hours of working with the model, end to end, before implementation starts in earnest.

Notice where the human sits in that list, right? Not reviewing output at the end. Deciding at the front, while a decision is still cheap to change.

Ten hours sounds like a lot until you price the alternative.

Every hour spent deciding at the front removes days of rework at the back, where the decision is already load bearing.

The work did not disappear. It moved to where it is cheap.

What it has produced

Three applications into production over the past three months. These are my own numbers, tracked on my own projects, and nobody outside has audited them.

  • 96% feature acceptance rate

  • 90% reduction in QA. Manual QA has almost disappeared

  • Feature updates shipping daily across applications in production

The acceptance rate is the one I would look at first. It is not measuring whether the code ran. It is measuring whether the thing built was the thing wanted, which is the exact failure discovery is supposed to prevent. When acceptance is high, it usually means the disagreements happened early, in a document, instead of late, in a demo.

The QA number is a consequence, not a goal. Manual QA shrinks when requirements arrive with acceptance criteria attached and traceability says which test covers which requirement. It is not that testing stopped mattering. It is that a lot of what manual QA used to catch was ambiguity, and ambiguity got caught upstream.

What this costs you

Being straight about the price, because a prompt shared without its cost is a sales pitch.

It is slow at the start. You will spend a day producing documents while someone on your team could have shipped a feature. That tradeoff is real and it feels bad in week one.

It only works if you actually run the gates. The model will happily produce all thirteen artifacts and wait. If you approve without reading, you have a folder of very well organized fiction.

It needs real source material. Point it at a thin repo and you get thin discovery, stated confidently. The prompt cannot manufacture stakeholder intent that was never captured.

The ADRs still take hours of human work. The prompt produces candidates. Candidates are not decisions, and the gap between them is judgment nobody has automated.

It is tuned to one domain. The reading sequence, the field guide, the portability section, all of it assumes a model-based agent where state, memory, and context are the hard parts. Running it on a CRUD application would be heavy machinery for a light problem.

What to actually take from this

You do not need my file names. The transferable part is the shape.

  • Give the model a reading order and tell it to record what it could not find, instead of skipping silently

  • Make it classify every claim as fact, inference, assumption, or unknown, so guesses cannot pass as findings

  • Forbid silent conflict resolution. A recorded contradiction is worth more than a smooth summary

  • Separate requirements from decisions and refuse one ADR per requirement

  • Red team the framing, not only the security surface, and demand the strongest argument against your own plan

  • Put human gates between phases and do not let the model past one without you

The reason this works is not that the prompt is clever. It is that it refuses to let a fluent model do the one thing fluent models are best at, which is producing something that reads like an answer before anyone established what the question was.

Code got cheap. Deciding did not.

When production gets cheap, the expensive part moves upstream to the decisions production is about to lock in.

Spend your model there first.