BDI

Defense technology.
Buyers, markets, opportunities.

The human-review cost inside an AI product's business case

AI can move work from drafting to verification or from generalists to scarce specialists. Measure accepted outcomes, review effort and elapsed time before estimating savings.

In this article
  1. Measure the complete task before pricing the saving
  2. Productivity evidence is specific to its setting
  3. Follow review effort to the receiving team
  4. Exceptions determine who must remain available
  5. Compare the costs without hiding quality
  6. Establish what will be reviewed after rollout
  7. Sources & evidence

The cost of an AI-generated document is not the price of generating it. Someone may still need to check its sources, resolve an exception, correct its structure and decide whether it is ready for another team to use. For a defence technology business buying an engineering or commercial assistant, those activities belong inside the business case from the start.

Human review can also be the product's strongest contribution. A system that organises evidence and makes a difficult decision easier to inspect may save more useful time than one that produces a longer first draft. The evaluation should identify where employee effort moves, whose capacity it consumes and whether the final work meets the same acceptance standard.

Measure the complete task before pricing the saving

Begin with the task as it exists today. A technical response may involve finding product records, reconciling differences, drafting an answer and obtaining approval. Measure enough of that process to know which part creates the delay. An assistant can accelerate drafting while leaving the approval queue unchanged, producing a real local saving without the promised improvement in delivery time.

Use separate measures for active employee effort and elapsed time. An engineer can work on another task while an automated process runs, but a customer may still be waiting for the result. Treating all waiting time as labour exaggerates cost; ignoring waiting time conceals a service limitation. Both measures can matter to the purchasing decision.

Define acceptable completion before comparing workflows. If AI-assisted answers receive more extensive checks, record why. If reviewers accept a lower standard because the output is intended only for exploration, compare it with the equivalent exploratory baseline. A faster draft and an approved technical response should not share a denominator without explanation.

Productivity evidence is specific to its setting

METR's July 2025 study involved 16 experienced open-source developers and 246 issues in familiar repositories. With the early-2025 tools studied, AI-allowed tasks took 19% longer. Participants nevertheless believed AI had accelerated them. The authors explicitly limited the finding to the studied setting; it is not a current estimate for all developers or evidence that review work alone caused the result.

The follow-up matters. In February 2026, METR said its newer experiment gave an unreliable estimate of current productivity effects, including because developers and tasks were increasingly selected out when participants did not want to work without AI. Concurrent agent use also complicated time measurement. The researchers believed acceleration had probably improved but treated the data as weak evidence about its magnitude.

For a buyer, these studies support measurement discipline rather than a universal productivity assumption. Employees' reported experience is valuable, but it should sit alongside task records and accepted outcomes. A result from a different model generation, user population or work setting needs a clear explanation before it becomes a forecast for the organisation.

Follow review effort to the receiving team

A hypothetical supplier adopts AI to prepare product compliance matrices. Sales engineers create the first version faster, but senior product specialists spend longer checking whether each answer applies to the quoted configuration. Counting only the sales engineer's time could show an attractive saving while consuming the capacity of the scarcer specialist.

Record effort by role and activity. Useful categories include evidence retrieval, drafting, factual verification, correction and approval. These categories should reflect the actual workflow, not a universal template imposed on every product. They make it possible to distinguish work removed from work transferred, and routine effort from interruptions that disrupt specialist schedules.

Some review is already present in the baseline. Charge the AI workflow for the additional review it creates and credit it for review it removes. Requiring a human approval does not mean every minute of approval is a new AI cost. Equally, an existing approval process should not be assumed capable of absorbing an unlimited increase in generated material.

Examine the amount of material arriving at that stage. If a product allows employees to produce twice as many drafts, total review demand can rise even when each draft is easier to inspect. The business may welcome that extra output, but it needs to fund and organise the resulting work rather than presenting all increased volume as labour saved.

Exceptions determine who must remain available

Average handling time hides the work that requires scarce judgment. A routine answer may take little review, while an unusual product configuration requires a specialist to reconstruct the source record. Keep the frequency, effort and disposition of those exceptions visible. Report unresolved cases as unresolved, rather than excluding them from the productivity calculation.

The interface can materially affect that effort. Reviewers benefit when the product shows what changed, where a claim came from and what remains uncertain. Evaluating citations in enterprise AI search explains why a working reference can still leave substantial verification work. A citation that opens the relevant version and passage may be commercially valuable even if generation speed is unchanged.

Decide who can handle each exception and what happens when that person is unavailable. A system that needs frequent vendor assistance may perform well in a supported pilot but impose a different cost in normal operation. Include the supplier's intervention in the record, and distinguish a product function from a service delivered by its staff.

Compare the costs without hiding quality

A practical calculation begins with accepted tasks over a stated period. Add employee effort, service charges and recurring administration associated with those tasks. Keep initial integration and training visible as separate investments, then explain the period over which the business expects to recover them. This prevents one unusually expensive setup week from being confused with steady operation.

Use the same quality conditions for both workflows. If the AI-assisted process produces more complete documentation, record that improvement alongside time. The business may rationally accept equal or greater effort for better evidence or wider coverage. A time-saving claim would still need its own support; improved quality should not be silently converted into an invented percentage reduction.

Likewise, avoid converting every minute saved into an immediate payroll saving. Small fragments of time may improve responsiveness or allow staff to handle a backlog without reducing headcount. The credible commercial claim depends on how that released capacity can be used. Buyers can assess additional throughput and avoided delay even when the finance case contains no staff reduction.

Establish what will be reviewed after rollout

A pilot is an initial estimate, not the end of measurement. Record the product version, task mix and review rules that produced it. When the system or workflow changes, revisit the activities most likely to affect the conclusion. A new source connector can increase usefulness while introducing a new type of checking work.

The CAE engineering AI account is a useful reminder to preserve the denominator of a reported outcome: its issue-intake result should not become a claim about all engineering productivity. Suppliers strengthen their own evidence by naming the activity improved and the part of the organisation actually measured.

The purchasing decision becomes clearer when the proposal states which work is removed, which work remains and which roles supply it. A product that reduces total accepted-task effort while keeping review manageable has a defensible business case. A product that improves quality or expands capacity can also earn its place, provided that benefit is described on its own terms.

Sources & evidence

  1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer ProductivityMETR · 10 July 2025
  2. We are Changing our Developer Productivity Experiment DesignMETR · 24 February 2026

METR findings are dated and limited to the studied populations; the 2026 follow-up’s selection limitations are retained. Cost examples are BDI analysis.

Suggest a correction