BDI

Defense technology.
Buyers, markets, opportunities.

How to evaluate a commercial change-detection service

A change-detection trial should explain missed changes, false alerts and unusable observations separately. Otherwise its accuracy score tells a buyer little about the service it will receive.

In this article
  1. Define the change before defining success
  2. Keep observation quality visible
  3. Read errors in both directions
  4. Make the reference evidence inspectable
  5. Distinguish software performance from assisted service performance
  6. Preserve the trial as a baseline for renewal
  7. Sources & evidence

A commercial change-detection service promises to turn repeated observations into a manageable account of what has changed. The buyer's problem is that several different failures can produce the same empty dashboard: nothing changed, no usable observation was available, the software missed a change, or an alert failed to reach the user. An evaluation that treats all four as equivalent cannot explain the service's value.

For commercial teams following industrial development, infrastructure or environmental conditions, the purchase decision should start with a precise description of the change being reported. The evaluation then needs to connect the underlying observations, the analytical result and the work required to review that result. This guide addresses product assessment and reporting, rather than operational surveillance or targeting.

Define the change before defining success

A supplier and customer can agree that an image looks different while disagreeing about whether the difference matters. A reporting service may be intended to identify a completed civil construction stage, a change in land cover or an alteration to a documented industrial site. Those are different analytical claims and should have separate definitions.

Specify the object of the report, the relevant time interval and the evidence required to label a case as a change. Include a way to record an unresolved case. If the reference reviewer cannot establish what happened, the evaluation should not silently turn that uncertainty into a correct or incorrect software decision.

The NIST AI RMF 1.0 core calls for documented evaluation data, attention to representativeness and performance assessment in conditions similar to the intended setting. It also asks users to document limits on generalising results. For a buyer, this supports a simple principle: the evaluation must describe the service the business intends to use, not just a convenient collection of examples.

Keep observation quality visible

A change-detection product depends on the observations it receives. The acceptance report should distinguish the period the customer wanted to assess from the period actually represented in usable data. A result based on two clear observations is a different evidential situation from one in which an intermediate delivery was missing or rejected.

This matters when comparing services with different source-data packages. One may include additional imagery acquisition; another may process only the archive already available to the customer. If both quote a monthly analytics fee, the cheaper offer may simply assume that someone else supplies the observations needed to make its result useful.

Ask for a coverage account alongside the alert account. It should show which requested comparisons were completed, which were deferred and why. A lack of usable input should remain visible rather than appearing as a confident no-change result. The distinction between acquisition opportunities and delivered data is explored further in imagery revisit and delivery latency.

Read errors in both directions

An illustrative evaluation contains 100 independently reviewed cases, of which 20 contain the defined change. A service identifies 16 of those changes and also flags 8 unchanged cases. It therefore produces 24 alerts: 16 useful detections and 8 false alerts. It has also missed 4 changes. These counts describe different consequences for the buyer.

The team reviewing alerts experiences a different workload from the team relying on the service to find changes. Reporting only the proportion of correct alerts would hide the missed cases. Reporting only the fraction of changes found would hide the unnecessary reviews. Neither should be replaced by a general accuracy claim without showing the underlying counts.

The prevalence of the event also matters to interpretation. A trial deliberately containing many obvious changes may be helpful for examining the product, but its alert workload need not resemble a routine reporting period in which change is uncommon. Preserve the selection method so that the customer understands what the trial does and does not estimate.

Make the reference evidence inspectable

The CEOS Land Product Validation subgroup distinguishes limited comparisons with reference data from broader assessment across representative locations and periods. Its higher validation stages also address systematic updates as products and time series evolve. These are scientific validation principles, not a certification of a particular commercial change-detection service.

A buyer can apply the underlying logic without commissioning a global research programme. Choose reference cases that cover the intended commercial scope and document how each label was established. If the service will support several types of industrial reporting, preserve separate results for those types rather than merging them into one attractive average.

Reference review should also have a correction process. If a disputed case is relabelled after discussion, retain the original judgement, the reason for the change and its effect on the reported result. Otherwise a trial can become progressively easier to pass without either party noticing how much its definition of success has moved.

The reference account should also preserve timing. Evidence that a construction milestone existed at the end of a quarter does not establish the exact day on which it occurred. A service reporting a first observation should use that description, and a customer comparing it with a completion record should allow for the difference. This prevents an evaluation from treating two different dates as contradictory when they describe different events.

Distinguish software performance from assisted service performance

Some suppliers provide an automated output. Others provide a report that a specialist has checked or edited. Both can be valuable, but a trial needs to describe which one is being measured. An analyst-assisted result cannot establish the unattended performance of the software alone.

The quotation should make the human work visible. Specify whether review covers every result or only selected exceptions, how much contextual research is included and how unresolved cases are returned. A customer buying a finished analytical report may reasonably pay for that expertise; the problem arises when it is present in the demonstration and absent from the recurring service.

Measure the receiving team's work as well. A technically correct alert may require substantial effort to interpret if it lacks the underlying comparison, timestamps or an explanation of its scope. A slightly smaller set of well-documented results can be more useful to a commercial team than a larger stream that creates a new research backlog.

Preserve the trial as a baseline for renewal

The final evaluation should identify the data package, analytical version, reference cases and service arrangements used. These establish the baseline against which later changes can be understood. A supplier may improve its model, switch an input source or revise the report format; the customer needs to know which changes affect its existing workflow.

Agree how a material update will be assessed. Repeating a bounded subset of reference cases can expose changes in output meaning or review burden without pretending that the original trial proved performance forever. Retain examples of difficult and inconclusive cases alongside successful ones so that later reviews do not focus only on the easiest demonstrations.

For businesses comparing providers, BlackSky's company profile offers a route into one commercial imagery and analytics business. A company description helps identify the offer; the acceptance evidence still needs to match the exact service and customer use under consideration.

A useful change-detection contract ultimately describes an accountable reporting chain. It explains what was observed, what the service concluded, where the evidence was insufficient and what work remains for the customer. That gives both parties a stronger basis for pricing, renewal and expansion than an isolated model score.

Sources & evidence

  1. AI RMF CoreNIST
  2. CEOS Land Product Validation SubgroupCEOS / NASA

NIST AI RMF 1.0 and CEOS land-product validation principles inform the evaluation approach. Numerical examples are illustrative and do not describe any named supplier's performance.

Suggest a correction