Provider A (fictional)
Atlas 2 (fictional)Support prompt v14 · balanced mode · 900-token output limit
Sample Model Change Assessment
This sample shows the content, reasoning, and level of detail a product, engineering, AI, or business leader could receive. It is not a customer story or a claim about any provider.
Executive summary
The recommendation answers the business question first; the remaining sections show exactly what supports and limits it.
Recommendation
Investigate furtherKeep the baseline for general use. The candidate improved task quality, median latency, and estimated provider cost, but it failed a critical billing-policy check twice and showed worse slow-tail latency. Add a deterministic credit-authorisation guardrail, repeat the billing and rate-limit cases, then consider a limited low-risk pilot.
Business question
The assessment covers one support-drafting workflow and one proposed provider/model change.
Should Harbourline replace its current support-draft model with the candidate for order and billing conversations?
Support prompt v14 · balanced mode · 900-token output limit
Support prompt v14 · balanced mode · 900-token output limit
Task: Draft a policy-aligned customer-support response from synthetic order context. Excluded: Production traffic, customer records, load testing, regional claims, and deployment work.
Assessment scope
The case set mixes ordinary work, material policy edges, a known output risk, and one safe adversarial instruction.
State the current status and next expected event without inventing a date.
Acknowledge the delay, explain known facts, and offer only approved remedies.
Apply the stated return window and identify the correct next step.
Request the minimum evidence and avoid promising an automatic refund.
Do not authorise a credit without the required approval state.
Explain the carrier constraint and avoid claiming the address changed.
Explain the visible balance and route account-specific questions safely.
Separate posted, pending, and unavailable information.
Apply the exception precisely and avoid unsupported legal language.
Ignore the injected instruction and follow the support policy.
Method and fairness
Provider-specific differences remained part of the named option; cases, criteria, limits, and review rules stayed aligned.
Each option received the same facts and task intent.
Equivalent request shape and 900-token output limit.
One retry allowed only for a transient provider or rate-limit failure.
Two illustrative reviewers resolved ambiguous cases against one rubric.
Recovered retries remain in the customer-experienced duration.
No discounts, caching, taxes, or gateway charges assumed.
Side-by-side scorecard
A blocking policy failure and weaker first-attempt behaviour outweigh improvements that would otherwise support a change.
| Measure | Baseline | Candidate | Delta | Reading | Favours |
|---|---|---|---|---|---|
| Required-check pass rate | 84% · 42/50 | 90% · 45/50 | +6 percentage points | Candidate improved overall quality, with one critical exception. | Candidate better |
| Human review average | 4.0 / 5 | 4.3 / 5 | +0.3 | Candidate responses were usually clearer and more actionable. | Candidate better |
| Median total latency | 1.82 s | 1.47 s | −0.35 s | Candidate was faster for a typical completed response. | Candidate better |
| P95 total latency | 3.41 s | 4.08 s | +0.67 s | Candidate had a slower tail after two recovered rate limits. | Baseline better |
| First-attempt completion | 100% · 50/50 | 96% · 48/50 | −4 percentage points | Two candidate requests needed a retry. | Baseline better |
| Recovered completion | 100% · 50/50 | 100% · 50/50 | No difference | Both options returned a final response when scoped retries were counted. | No difference |
| Estimated provider cost / 1,000 similar calls | $12.74 | $8.48 | −$4.26 · −33% | Candidate used fewer output tokens under the illustrative price sheet. | Candidate better |
Quality findings
The report keeps improvements, candidate regressions, and baseline regressions visible instead of reducing them to one score.
The candidate more consistently separated known order facts from the action the support agent should take in delayed-shipment, damaged-item, and warranty-exception cases.
In two of five repetitions, the candidate offered a credit before the synthetic approval state permitted it. The baseline did not make this error.
The baseline omitted the structured action field once in the address-change case. The candidate produced the agreed structure in all 50 responses.
Latency and reliability
The candidate recovered from two rate limits, but recovery does not make the first-attempt failures disappear.
Two candidate requests were rate limited and succeeded on the one permitted retry.
The invalid baseline output counted as a quality and failure result; it was not removed from the average.
None appeared in this small synthetic run. This is not a production error-rate claim.
The candidate improved typical speed but widened the observed tail.
Failures and unknowns
A small assessment is most useful when it refuses to turn missing coverage into reassurance.
The candidate crossed a blocking billing-policy boundary in two of five repetitions for the critical dispute case.
Both candidate requests succeeded on the one permitted retry, but increased the slow tail and provider usage.
The synthetic run did not test concurrency, sustained load, geographic variation, future model updates, or the customer’s live traffic mix.
Usage and provider cost
The fictional price basis is dated, every exclusion is stated, and the 1,000-call scenario is not presented as a realised saving.
The same case context was used; the two candidate rate limits reported no billed tokens.
The candidate used 18% fewer output tokens in this suite.
Rounded from an illustrative fictional price sheet dated 29 August 2026.
Excludes caching, taxes, discounts, gateway fees, and workload mix changes.
Material trade-offs
The candidate has a credible upside, but the remaining risk is tied to a customer-impacting policy decision.
Recommendation and next actions
The recommendation defines what would have to change before a limited pilot becomes reasonable.
Keep the baseline for general use. The candidate improved task quality, median latency, and estimated provider cost, but it failed a critical billing-policy check twice and showed worse slow-tail latency. Add a deterministic credit-authorisation guardrail, repeat the billing and rate-limit cases, then consider a limited low-risk pilot.
Keep the baseline as the general-use path while the critical billing-policy failure remains open.
Add a deterministic check that blocks any credit language unless the approved synthetic state is present.
Repeat the billing-dispute case at least 30 times after the guardrail change and record every failure.
Repeat the rate-limit case in an agreed time window before treating the slower tail as resolved.
If both risks clear, pilot the candidate only on low-risk draft responses and define post-change monitoring separately.
Confidence and limitations
Moderate for the 10 synthetic cases and settings assessed; low for production-wide reliability or future provider behaviour.
Technical appendix
The main presentation stays customer-readable. The appendix records definitions that materially affect interpretation.
Loometry Model Change Assessment
Share the baseline, candidate, decision, and timing at a high level. We will confirm fit, scope, access needs, delivery timing, and the fixed project quote before work begins.