Skip to content

Sample Model Change Assessment

A realistic decision report, built entirely from synthetic data.

This sample shows the content, reasoning, and level of detail a product, engineering, AI, or business leader could receive. It is not a customer story or a claim about any provider.

Synthetic example

Every organisation, provider, model, case, response, measurement, usage value, price, finding, and recommendation on this page is fictional and created for demonstration.

Executive summary

Promising candidate. Two risks block a broad change.

The recommendation answers the business question first; the remaining sections show exactly what supports and limits it.

Recommendation

Investigate further

Investigate further and mitigate two specific risks before any broad change.

Keep the baseline for general use. The candidate improved task quality, median latency, and estimated provider cost, but it failed a critical billing-policy check twice and showed worse slow-tail latency. Add a deterministic credit-authorisation guardrail, repeat the billing and rate-limit cases, then consider a limited low-risk pilot.

Business question
Should Harbourline replace its current support-draft model with the candidate for order and billing conversations?
Headline upside
Higher required-check pass rate, faster median response, and lower illustrative provider cost.
Blocking risk
Two candidate responses offered an unapproved billing credit.
Confidence
Moderate for the 10 synthetic cases and settings assessed; low for production-wide reliability or future provider behaviour.

Business question

What the fictional team needed to decide.

The assessment covers one support-drafting workflow and one proposed provider/model change.

Should Harbourline replace its current support-draft model with the candidate for order and billing conversations?
Baseline

Provider A (fictional)

Atlas 2 (fictional)

Support prompt v14 · balanced mode · 900-token output limit

Candidate

Provider B (fictional)

Meridian 3 (fictional)

Support prompt v14 · balanced mode · 900-token output limit

Task: Draft a policy-aligned customer-support response from synthetic order context. Excluded: Production traffic, customer records, load testing, regional claims, and deployment work.

Assessment scope

Ten cases, five repetitions, two options.

The case set mixes ordinary work, material policy edges, a known output risk, and one safe adversarial instruction.

Representative cases
10
Synthetic support workload
Repetitions per case
5
Same plan for both options
Planned first attempts
100
50 baseline · 50 candidate
Decision confidence
Moderate
Bounded to this case set
01Common

Order status summary

State the current status and next expected event without inventing a date.

02Common

Delayed shipment

Acknowledge the delay, explain known facts, and offer only approved remedies.

03Policy

Return eligibility

Apply the stated return window and identify the correct next step.

04Edge

Damaged item

Request the minimum evidence and avoid promising an automatic refund.

05Critical

Billing dispute

Do not authorise a credit without the required approval state.

06Edge

Address change after dispatch

Explain the carrier constraint and avoid claiming the address changed.

07Common

Gift-card balance

Explain the visible balance and route account-specific questions safely.

08Edge

Loyalty-points mismatch

Separate posted, pending, and unavailable information.

09Critical

Warranty exception

Apply the exception precisely and avoid unsupported legal language.

10Adversarial

Instruction hidden in customer text

Ignore the injected instruction and follow the support policy.

Method and fairness

The same question was asked of both options.

Provider-specific differences remained part of the named option; cases, criteria, limits, and review rules stayed aligned.

CasesSame 10 synthetic inputs

Each option received the same facts and task intent.

InstructionsSame support prompt v14

Equivalent request shape and 900-token output limit.

RepetitionFive planned runs per case

One retry allowed only for a transient provider or rate-limit failure.

QualitySame required and prohibited checks

Two illustrative reviewers resolved ambiguous cases against one rubric.

LatencySame request-to-complete boundary

Recovered retries remain in the customer-experienced duration.

CostObserved synthetic usage, dated fictional prices

No discounts, caching, taxes, or gateway charges assumed.

Side-by-side scorecard

The candidate wins the average, but not the decision.

A blocking policy failure and weaker first-attempt behaviour outweigh improvements that would otherwise support a change.

All values are synthetic. Deltas compare the candidate with the baseline.
MeasureBaselineCandidateDeltaReadingFavours
Required-check pass rate84% · 42/5090% · 45/50+6 percentage pointsCandidate improved overall quality, with one critical exception.Candidate better
Human review average4.0 / 54.3 / 5+0.3Candidate responses were usually clearer and more actionable.Candidate better
Median total latency1.82 s1.47 s−0.35 sCandidate was faster for a typical completed response.Candidate better
P95 total latency3.41 s4.08 s+0.67 sCandidate had a slower tail after two recovered rate limits.Baseline better
First-attempt completion100% · 50/5096% · 48/50−4 percentage pointsTwo candidate requests needed a retry.Baseline better
Recovered completion100% · 50/50100% · 50/50No differenceBoth options returned a final response when scoped retries were counted.No difference
Estimated provider cost / 1,000 similar calls$12.74$8.48−$4.26 · −33%Candidate used fewer output tokens under the illustrative price sheet.Candidate better

Quality findings

Which cases changed materially—and why.

The report keeps improvements, candidate regressions, and baseline regressions visible instead of reducing them to one score.

01Material improvement

Clearer next steps in three routine cases

The candidate more consistently separated known order facts from the action the support agent should take in delayed-shipment, damaged-item, and warranty-exception cases.

02Critical regression

Two billing replies exceeded the approved remedy

In two of five repetitions, the candidate offered a credit before the synthetic approval state permitted it. The baseline did not make this error.

03Baseline regression

One baseline reply broke the required response shape

The baseline omitted the structured action field once in the address-change case. The candidate produced the agreed structure in all 50 responses.

Latency and reliability

Faster typical response, weaker first attempt, slower tail.

The candidate recovered from two rate limits, but recovery does not make the first-attempt failures disappear.

MeasureBaselineCandidateInterpretation
First-attempt reliability50 of 5048 of 50

Two candidate requests were rate limited and succeeded on the one permitted retry.

Invalid outputs10

The invalid baseline output counted as a quality and failure result; it was not removed from the average.

Timeouts and provider errors00

None appeared in this small synthetic run. This is not a production error-rate claim.

Latency range1.11–3.66 s0.96–4.42 s

The candidate improved typical speed but widened the observed tail.

Failures and unknowns

What failed, what recovered, and what remains unanswered.

A small assessment is most useful when it refuses to turn missing coverage into reassurance.

Critical failure
Unapproved credit language

The candidate crossed a blocking billing-policy boundary in two of five repetitions for the critical dispute case.

Recovered failure
Two rate limits

Both candidate requests succeeded on the one permitted retry, but increased the slow tail and provider usage.

Unknown
Production behaviour

The synthetic run did not test concurrency, sustained load, geographic variation, future model updates, or the customer’s live traffic mix.

Usage and provider cost

Observed usage first. Estimated cost second.

The fictional price basis is dated, every exclusion is stated, and the 1,000-call scenario is not presented as a realised saving.

MeasureBaselineCandidateBasis and limit
Input tokens54,20054,200

The same case context was used; the two candidate rate limits reported no billed tokens.

Output tokens30,60025,200

The candidate used 18% fewer output tokens in this suite.

Assessment-run estimate$0.64$0.42

Rounded from an illustrative fictional price sheet dated 29 August 2026.

Scenario estimate$12.74 / 1,000$8.48 / 1,000

Excludes caching, taxes, discounts, gateway fees, and workload mix changes.

Material trade-offs

Why the recommendation is not a simple winner badge.

The candidate has a credible upside, but the remaining risk is tied to a customer-impacting policy decision.

Candidate improves
  • Required-check pass rate by six percentage points
  • Illustrative human review average by 0.3 points
  • Median total latency by 0.35 seconds
  • Output usage by 18%
  • Illustrative provider cost by 33%
Candidate worsens or leaves open
  • Two critical billing-policy failures
  • First-attempt completion by four percentage points
  • P95 latency by 0.67 seconds
  • Two rate-limit retries
  • Production and load behaviour remain unknown

Recommendation and next actions

Keep the baseline, mitigate the critical risk, then retest narrowly.

The recommendation defines what would have to change before a limited pilot becomes reasonable.

Investigate further

Investigate further and mitigate two specific risks before any broad change.

Keep the baseline for general use. The candidate improved task quality, median latency, and estimated provider cost, but it failed a critical billing-policy check twice and showed worse slow-tail latency. Add a deterministic credit-authorisation guardrail, repeat the billing and rate-limit cases, then consider a limited low-risk pilot.

  1. 01

    Keep the baseline as the general-use path while the critical billing-policy failure remains open.

  2. 02

    Add a deterministic check that blocks any credit language unless the approved synthetic state is present.

  3. 03

    Repeat the billing-dispute case at least 30 times after the guardrail change and record every failure.

  4. 04

    Repeat the rate-limit case in an agreed time window before treating the slower tail as resolved.

  5. 05

    If both risks clear, pilot the candidate only on low-risk draft responses and define post-change monitoring separately.

Confidence and limitations

Useful for this decision. Insufficient for broad assurance.

Moderate for the 10 synthetic cases and settings assessed; low for production-wide reliability or future provider behaviour.

  • All organisations, provider and model names, cases, usage, prices, and measurements are synthetic.
  • Ten cases describe this bounded task set; they do not represent every support conversation.
  • Five repetitions per case expose some variation but cannot establish a production reliability rate.
  • The assessment did not test concurrency, sustained load, geographic performance, account-tier differences, or future model versions.
  • Human-review and automated-check disagreement was resolved for the sample, but the illustrative reviewers may still be wrong.
  • The result is not a safety certification, security audit, compliance opinion, legal review, or deployment approval.

Technical appendix

Enough method detail to review or repeat the comparison.

The main presentation stays customer-readable. The appendix records definitions that materially affect interpretation.

Unit of comparison
One synthetic case response under one named option and configuration.
Required-check pass
All case-specific required checks pass and no prohibited behaviour appears.
Human review
Illustrative five-point rubric for correctness, policy alignment, actionability, and clarity.
Total latency
Elapsed time from request start to complete response; a scoped retry remains in the recovered duration.
First-attempt completion
A usable provider response arrives without retry, timeout, transport error, or rate-limit recovery.
Provider cost
Fictional price sheet dated 29 August 2026: baseline $2.00 per 1 million input tokens and $17.28 per 1 million output tokens; candidate $1.50 and $13.60 respectively. No retry tokens were reported for the two synthetic rate limits.
Critical criterion
A failure that blocks adoption regardless of the aggregate quality or cost result.
Evidence boundary
No credentials, customer records, or unnecessary internal identifiers are shown.

Loometry Model Change Assessment

Bring one model change. Get a decision-ready answer.

Share the baseline, candidate, decision, and timing at a high level. We will confirm fit, scope, access needs, delivery timing, and the fixed project quote before work begins.