BDI

Defense technology.
Buyers, markets, opportunities.

Reading false-alarm claims in counter-drone evaluations

A false-alarm claim is only comparable when the buyer knows what was counted, over what period and against which reference record. Percentages and alerts per hour describe different things.

In this article
  1. Establish what the evaluation calls a false alarm
  2. Keep the denominator attached to the result
  3. Read the reference record alongside the product output
  4. Low alert volume is not sufficient evidence of quality
  5. Measure the review burden separately
  6. Preserve comparability when products change
  7. Sources & evidence

A supplier's claim of a low false-alarm rate can be meaningful, but the number needs a definition. Ten false alerts in a month, one percent of all alerts and one percent of non-drone evaluation cases describe different measures. They cannot be ranked as if they were three versions of the same result.

For a counter-drone buyer, the commercial consequence is the work required to review the system's output and the confidence that the evidence supports. For a supplier, a well-defined claim is easier to defend and more useful in a procurement evaluation than an impressive percentage with no visible denominator.

Establish what the evaluation calls a false alarm

The JRC's counter-drone technology report distinguishes false positives from false negatives. A system can report a drone when the reference record says there was none, or fail to report a drone that was present. The report also explains why alert quality affects the human part of the overall system.

That distinction should be preserved in an evaluation report. A false positive concerns an incorrect positive result. A missed detection concerns a different failure and needs different evidence to identify. A report containing only the alerts produced by the product cannot by itself establish how many relevant cases the product missed.

The buyer should also understand the unit being counted. Is an alarm a single notification, a tracked event or every repeated message associated with the same event? If one occurrence generates several notifications, a raw notification count may describe interface workload more accurately than the number of independent mistakes.

Neither measure is inherently wrong. The problem arises when a supplier reports one and a buyer interprets the other.

Keep the denominator attached to the result

A percentage needs a population. “False alerts divided by all alerts” describes the share of reported alerts that were incorrect. “False positive cases divided by all negative cases” describes a different relationship. An alert count over time describes yet another aspect of the service.

Consider an illustrative evaluation with 100 positive alerts, of which ten are judged incorrect. Ten percent of the alerts are false in that dataset. Without the number of evaluated negative cases or the observation period, the result does not establish a false-positive probability per negative case or an average number of false alerts per day.

The distinction matters because procurement requirements can concern different outcomes. A customer may care about the time staff spend reviewing notifications, while a technical evaluation may compare classification decisions on a labelled dataset. Both are relevant, but the evidence should connect the technical result to the customer requirement.

The accessible primary text of NPSA's DTI requirements guide asks buyers to define an acceptable false-alarm rate in relation to their circumstances and expresses categories over a 24-hour period. That provides a concrete example of a time-based requirement, distinct from a percentage of test cases.

Read the reference record alongside the product output

An alert-quality result needs an agreed basis for deciding whether each relevant case was correct. The evaluation should identify how that reference was established and which cases remain uncertain. Otherwise, disagreements about the reference can be mistaken for product-performance differences.

The reference record should also be independent of the output being assessed. If an event is considered real only because the evaluated product reported it, the process cannot establish whether the product was correct. If the record contains only planned positive examples, it may say little about the false-alert behavior that the buyer wants to understand.

Unresolved cases should remain visible. Excluding them may be appropriate for a clearly stated analysis, but the report should explain the exclusion and its effect on interpretation. Treating every uncertain case as a success creates confidence that the evidence does not support.

This is an evidence-governance question rather than a prescription for a live trial. Qualified evaluators should determine the appropriate method and safeguards; commercial reviewers need enough documentation to understand the result and its limits.

Low alert volume is not sufficient evidence of quality

A system that reports fewer alerts may reduce review workload, but a buyer also needs evidence about what it failed to report. The purchasing decision should therefore consider false alerts and missed cases together within the same stated evaluation scope.

This does not imply a universal balance that every buyer should choose. Different products and customer requirements can justify different acceptance criteria. It does mean that a supplier should not present a reduction in alert volume as proof of improved overall performance without explaining the accompanying evidence.

The same caution applies to a claim of zero false alarms. That can accurately describe an observed evaluation period. It does not establish that the product will never produce a false alarm in future use. The size and character of the observed sample remain material to the claim.

A concise, bounded statement can be more persuasive than an absolute one. It identifies the evaluated configuration, the observation basis, the result and the cases to which the claim applies.

Measure the review burden separately

Two products with the same number of false events can create different amounts of staff work. One may present a single coherent record, while another produces repeated notifications that users must reconcile. The customer should identify the review task it is purchasing the system to support.

A useful commercial evaluation can distinguish the number of events, the number of notifications and the effort needed to resolve them. It should also identify whether the displayed information is sufficient for the responsible staff to understand what the system is claiming.

A staffing estimate should use measured review effort where available and show assumptions where it is not. Multiplying every notification by the same assumed handling time may overstate the burden of duplicates or understate the work associated with ambiguous records. Keeping those categories separate makes the cost estimate easier to update after a customer evaluation.

This is where interface and integration quality can affect the business case. If staff must move between several applications or manually connect records to understand an alert, the cost may not appear in the sensor's headline specification. The supplier can create value by making the result easier to review, even without changing the underlying sensor.

BDI's coverage of NATO's TIE26 interoperability work examines the wider need for systems to work together. An exercise participation record supplies context, while the alert-quality claim for an individual offer still needs its own evidence.

Preserve comparability when products change

Alert processing can change with a software version, sensor configuration or integration update. A comparison should identify the version used for the reported result and explain whether later changes remain within that evidence.

The customer should be able to distinguish an observed improvement from a change in counting rules. If a new release groups repeated notifications differently, the number of alerts may fall without a corresponding change in the number of incorrect events. That may still improve usability, but the claim should describe the actual improvement.

Supplier context can help readers identify the relevant product family. BDI's DroneShield profile provides that market view without serving as a substitute for an offer-specific evaluation report.

A credible false-alarm claim therefore keeps four things together: what was counted, the denominator, the reference record and the evaluated configuration. It allows buyers to compare evidence that answers the same question and gives suppliers a clear basis for demonstrating the commercial value of their product.

Sources & evidence

  1. Technical developments in counter-drone technologyEuropean Commission Joint Research Centre · 27 February 2025
  2. Developing operational requirements for detect, track and identify technologyNational Protective Security Authority

JRC's public technology report and indexed text of NPSA's public requirements guide inform the terminology. The NPSA PDF could not be retrieved directly in this review; only its accessible primary-source text was used. Numerical examples are illustrative.

Suggest a correction