Skip to content

How it works

Hold the question steady. Make the trade-offs visible.

The assessment is a controlled comparison before a decision. It keeps the baseline, candidate, cases, measures, failures, and limitations connected from scope to recommendation.

Inputs

Your team supplies five kinds of decision context.

The public contact form asks only for high-level facts. Cases, access, and sensitive details are reviewed after fit is confirmed.

01Decision

What are you considering, who owns the choice, and when does the answer matter?

02Baseline

Which provider, model, endpoint, prompt, settings, gateway, and recovery path represent today?

03Candidate

Which exact change is proposed, and which differences are intentional?

04Cases and checks

Which safe examples represent the work, and what must or must not happen?

05Constraints

What data, provider access, budget, timing, review, and retention boundaries apply?

Method

Five steps from question to action.

Each step produces a reviewable output and narrows what the final recommendation is allowed to claim.

  1. 01

    Frame the decision

    Name the current path, proposed change, decision owner, deadline, required outcomes, and blocking risks.

  2. 02

    Approve safe cases

    Select up to 10 representative examples with expected behaviour and the least sensitive useful input.

  3. 03

    Lock the method

    Agree what stays constant, what intentionally changes, how results are repeated, and how each measure is defined.

  4. 04

    Run and inspect

    Execute both options, preserve failures and retries, and review aggregate and case-level differences together.

  5. 05

    Recommend an action

    Separate observations from interpretation, state confidence and limits, then recommend change, stay, investigate, or mitigate.

Fairness controls

Comparable does not mean pretending the providers are identical.

The method holds decision inputs steady and records any difference that could affect interpretation.

What stays aligned and how unavoidable differences are handled.
ControlHeld steadyDocumented exception
CasesThe same reviewed case set and input factsA case is excluded only for a documented incompatibility.
InstructionsEquivalent system instructions, context, tools, and response contractA provider-specific format change is recorded.
ConfigurationComparable sampling, output, timeout, and retry settingsA non-equivalent provider setting is treated as part of the candidate.
RepetitionThe same planned repetitions and stop conditionsUnexpected provider limits remain visible as results.
QualityThe same required, prohibited, and reviewer criteriaA check changes only through an approved method revision.
Cost and timingThe same measurement boundaries and stated price basisCustomer-supplied commercial terms are labelled separately.

Evaluation versus monitoring

This assessment answers a change question. Monitoring answers what happens next.

The two practices use related measures but support different claims and timelines.

Model Change Assessment

Controlled and time-bounded.

  • Runs before a model decision
  • Uses approved, representative cases
  • Compares a named baseline and candidate
  • Ends with a recommendation and next action
Production monitoring

Continuous and environment-specific.

  • Observes live behaviour after deployment
  • Uses production traffic and operational signals
  • Detects drift, incidents, and new failure patterns
  • Requires its own coverage, alerts, owners, and response plan
Boundary
An assessment does not prove future production behaviour

Up to 10 cases can support a bounded decision. They do not replace production monitoring, load testing, incident response, or a regional reliability study.

From evidence to recommendation

Four layers prevent a metric table from becoming the conclusion.

A decision owner should be able to distinguish what happened from what the team should do about it.

01Observations

Responses, timing, errors, retries, usage, and case-level outcomes.

02Measures

Defined checks and calculations applied to the preserved observations.

03Interpretation

Why an improvement, regression, failure, or unknown matters to the product decision.

04Recommendation

Change, stay, investigate, or mitigate—with confidence, conditions, owner, and next step.

Provider and data boundary

Agree access and data handling before execution.

Credentials and sensitive case material never belong in the public contact form, browser storage, report body, or source code.

You describeProvider, endpoint, models, gateway, cases, and restrictions

High-level fit information comes first; no secrets are needed for an initial discussion.

We approveMinimum access, data scope, budget, retention, and removal

Any provider access is scoped to the agreed assessment and stays server-side.

The report includesEnough configuration and method context to interpret the result

It excludes credentials, customer records, and unnecessary internal identifiers.

Loometry Model Change Assessment

Bring one model change. Get a decision-ready answer.

Share the baseline, candidate, decision, and timing at a high level. We will confirm fit, scope, access needs, delivery timing, and the fixed project quote before work begins.