Testing Exceptions and Failure Cases
Agent evaluation uses representative tasks, missing information, conflicting evidence, unavailable tools, unsafe requests, and recovery cases rather than one prepared demonstration. Build an evaluation set from real work and known exceptions, recording expected behaviour, evidence, escalation, and unacceptable outcomes.
What You Will Be Able to Decide
- Explain testing exceptions and failure cases in product and business terms.
- Apply this decision: Build an evaluation set from real work and known exceptions, recording expected behaviour, evidence, escalation, and unacceptable outcomes.
- Recognise this material risk: a convincing happy-path demonstration is mistaken for reliable behaviour under ordinary operational pressure.
- Ask a consultant for evidence rather than reassurance.
A founder or operator is deciding how an AI agent should participate in a real workflow without inheriting undefined authority.
Agent evaluation uses representative tasks, missing information, conflicting evidence, unavailable tools, unsafe requests, and recovery cases rather than one prepared demonstration.
A consultant can recommend and implement the technical approach. The founder still needs to decide which outcome matters, which risk is acceptable, and what evidence is sufficient.
The Founder Situation
A founder or operator is deciding how an AI agent should participate in a real workflow without inheriting undefined authority.
The immediate question is testing exceptions and failure cases. The technical label matters only because it changes a product decision, a responsibility, or the evidence required before launch.
Technical term
Testing Exceptions and Failure Cases
Agent evaluation uses representative tasks, missing information, conflicting evidence, unavailable tools, unsafe requests, and recovery cases rather than one prepared demonstration.
Treat it like a clause in a commercial agreement: its value comes from making expectations and consequences clear, not from sounding formal.
What Matters in Practice
Start with the product consequence, then choose the simplest technical treatment that protects it. A longer tool list is not a stronger plan.
For this decision, the useful standard is that the agent behaves predictably across representative work, respects its boundaries, and produces evidence a responsible person can review.
- Make the decision explicit: Build an evaluation set from real work and known exceptions, recording expected behaviour, evidence, escalation, and unacceptable outcomes.
- Ask what evidence would show that the chosen approach works.
- Name the person or provider responsible when the approach fails.
- Record the result in the agent workflow specification, evaluation set, and operating record.
A Proportionate Decision
Build an evaluation set from real work and known exceptions, recording expected behaviour, evidence, escalation, and unacceptable outcomes.
The principal risk is that a convincing happy-path demonstration is mistaken for reliable behaviour under ordinary operational pressure. This does not require the most expensive possible solution. It requires the consequence to be understood and the control to match it.
- Describe the user or business outcome that must be protected.
- Identify the most credible failure and its consequence.
- Compare the simplest adequate approach with one realistic alternative.
- Set a review point for when the decision may need to change.
Strong Evidence and Weak Reassurance
Warning Signs
- Nobody can explain how testing exceptions and failure cases changes a user or business outcome.
- The proposal does not address this risk: a convincing happy-path demonstration is mistaken for reliable behaviour under ordinary operational pressure.
- The only evidence is a successful demonstration of the easiest path.
- The decision has no named owner, boundary, or review point.
- A provider-specific feature is being mistaken for a permanent product requirement.
Questions to Ask a Consultant
- What decision are we making about testing exceptions and failure cases?
- Which user or business outcome does the recommendation protect?
- How have we reduced or accepted this risk: a convincing happy-path demonstration is mistaken for reliable behaviour under ordinary operational pressure.
- What evidence can I review without relying on the original implementer?
- What is deliberately deferred, and when will it be reconsidered?
- Who owns the accounts, data, documentation, and recovery process?
Key takeaway
Key Takeaway
Agent evaluation uses representative tasks, missing information, conflicting evidence, unavailable tools, unsafe requests, and recovery cases rather than one prepared demonstration. The founder's job is to make the consequence explicit; the consultant's job is to recommend and demonstrate a proportionate implementation.
