Skip to content

Resources/Prepare

Case design

Select up to 10 representative cases

Cover ordinary work, important edges, and known failure risks without copying unnecessary production data.

Audience
Product owners, domain reviewers, and AI practitioners choosing the assessment inputs.
Question
Which small set of cases best represents the decision and its material risks?
01

Build from the workload, not from a benchmark

Start with the tasks customers or staff actually need the system to perform. A public benchmark may add context, but it cannot replace product-specific cases and expected behaviour.

  1. 01

    Choose three to five ordinary cases that represent frequent work.

  2. 02

    Add two or three difficult or ambiguous cases.

  3. 03

    Add one or two high-consequence failures that could block the change.

  4. 04

    Include a known regression when the team has one.

02

Keep every case reviewable

For each case, write the input purpose, the facts available to the model, required elements, unacceptable behaviour, and the reviewer qualified to resolve ambiguity.

03

Use the least sensitive useful version

Prefer synthetic, public, redacted, or specifically approved examples. Remove names, account numbers, credentials, confidential instructions, and unrelated personal information before assessment work begins.

  1. 01

    Replace identifiers with stable fictional values.

  2. 02

    Preserve the task difficulty without preserving the person behind it.

  3. 03

    Record any feature lost during redaction because it may limit the conclusion.

Worked example

Example: support-response case mix

Four common order questions, two policy edges, two critical billing or warranty cases, one known malformed-output regression, and one safe prompt-injection case form a more useful set than 10 nearly identical happy paths.