Resources/Prepare
Case design
Select up to 10 representative cases
Cover ordinary work, important edges, and known failure risks without copying unnecessary production data.
- Audience
- Product owners, domain reviewers, and AI practitioners choosing the assessment inputs.
- Question
- Which small set of cases best represents the decision and its material risks?
Build from the workload, not from a benchmark
Start with the tasks customers or staff actually need the system to perform. A public benchmark may add context, but it cannot replace product-specific cases and expected behaviour.
- 01
Choose three to five ordinary cases that represent frequent work.
- 02
Add two or three difficult or ambiguous cases.
- 03
Add one or two high-consequence failures that could block the change.
- 04
Include a known regression when the team has one.
Keep every case reviewable
For each case, write the input purpose, the facts available to the model, required elements, unacceptable behaviour, and the reviewer qualified to resolve ambiguity.
Use the least sensitive useful version
Prefer synthetic, public, redacted, or specifically approved examples. Remove names, account numbers, credentials, confidential instructions, and unrelated personal information before assessment work begins.
- 01
Replace identifiers with stable fictional values.
- 02
Preserve the task difficulty without preserving the person behind it.
- 03
Record any feature lost during redaction because it may limit the conclusion.
Worked example
Example: support-response case mix
Four common order questions, two policy edges, two critical billing or warranty cases, one known malformed-output regression, and one safe prompt-injection case form a more useful set than 10 nearly identical happy paths.