BDI

Defense technology.
Buyers, markets, opportunities.

Evaluating citations in an enterprise AI search product

Working links do not prove an AI answer is supported. Separate retrieval quality, claim support and evidence handover before measuring the product’s value.

In this article
  1. Three scores, three different promises
  2. Build questions around actual evidence relationships
  3. Inspect claims at the size a reader uses them
  4. Automated judging needs its own evidence
  5. Preserve the evidence after the demonstration
  6. Price supported work, not attractive answers
  7. Sources & evidence

An enterprise search answer can contain working links and still misrepresent the documents it cites. For a defence supplier using AI to navigate engineering records, supplier correspondence or programme material, the relevant purchase question is whether an employee can establish why an answer deserves to be used. Counting citation markers answers very little of that question.

A product evaluation should separate three stages: finding relevant material, constructing an answer supported by that material, and giving the reader a faithful route back to the evidence. Each can succeed while another fails. The commercial consequence is practical: a system that retrieves well but explains poorly may remain useful as search, whereas an apparently polished answer service can create substantial verification work.

Three scores, three different promises

The December 2023 ALCE research paper evaluates fluency, correctness and citation quality separately. Its citation measures distinguish whether statements have supporting citations from whether the citations supplied are relevant support. These are useful distinctions for a customer evaluation; the paper's results for the models and datasets tested in 2023 are not a performance estimate for a product offered today.

Consider a hypothetical maintenance knowledge base containing a released procedure and an earlier draft. The search engine might retrieve both, the answer might reproduce the draft's obsolete requirement, and its citation might open the released document. The link works, and the final document is relevant to the subject, but the answer is not supported by that citation. A simple link checker would miss the problem entirely.

Conversely, an answer may accurately reproduce a passage from an obsolete draft. That is faithful citation with an unsuitable source selection. The customer needs to know whether release status, document authority and validity dates are represented in retrieval and visible to the reader. Grounding an answer in a document does not establish that the document governs the present decision.

Build questions around actual evidence relationships

A useful evaluation set includes more than factual lookup questions with an obvious single sentence answer. Engineering and commercial employees also compare revisions, reconcile differences between documents, identify missing information and distinguish proposals from approved decisions. These question types expose different product capabilities and should retain separate results.

For a supplier questionnaire, one question might ask which approved product variants support a named interface. Another might ask whether a proposed variant has completed the same verification. A third might ask what evidence is absent. The correct response to the third could be a statement of uncertainty with relevant records, rather than a confidently completed table.

Prepare the expected evidence before running the demonstration. Record the document version, relevant passage, permissible conclusion and material qualification. This is not necessarily a single preferred wording: several answers can be valid. It is a boundary around what the available evidence supports, which prevents evaluators from rewarding fluent but stronger conclusions after seeing them.

The set should also include questions that the authorised corpus cannot answer. Report how frequently the product declines, answers partially or supplies unsupported detail. A refusal is not automatically a failure when the requested fact is unavailable. Equally, a system that refuses almost everything offers little business value, so coverage and unsupported-answer rates need to be read together.

Inspect claims at the size a reader uses them

A paragraph can combine a product capability, a date and a commercial implication while attaching one citation at the end. Evaluate those claims separately. The source may support the capability but contain neither the date nor the inference about market access. The reader should be able to recognise which parts are reported facts and which are the system's interpretation.

This matters particularly with tables. A citation in a row heading can look as though it supports every cell, although the underlying passage addresses only one attribute. Ask evaluators to follow the evidence for selected cells, including empty values and qualifications. Preserve the original units and status labels when judging correctness; a maximum specification and a measured result are different claims even if the number matches.

The interface also affects review effort. A citation that opens a long document at its first page imposes more work than a stable passage reference with context. Record whether the reviewer can inspect the surrounding text, identify the version and return to the answer without losing their place. These observations measure usability of evidence rather than merely its existence.

Automated judging needs its own evidence

A 2025 study recorded by NIST compared relevance assessments in the TREC 2024 retrieval-augmented generation track. Across 77 runs from 19 teams, automatically generated assessments produced system rankings highly correlated with fully manual judgments. That finding concerns retrieval evaluation in the studied setting; it does not establish that an automated judge will reliably approve every factual claim in a customer's engineering answer.

For a commercial pilot, automated scoring can help identify recurring failure categories and expand coverage between manual reviews. Its acceptance depends on checking the kinds of disagreement that matter to the customer. A judge that treats an obsolete procedure as relevant may rank search results reasonably while missing the very distinction that an engineering team needs.

Keep a manually reviewed sample that includes ambiguous questions, conflicting records and missing answers. Record disagreement between reviewers as well as disagreement with the automatic score. Some questions reveal a poorly specified acceptance rule or an inconsistent knowledge base, rather than a uniquely defective model. Resolving those cases improves both the evaluation and the customer's information management.

Preserve the evidence after the demonstration

An answer displayed on a screen is difficult to investigate later if the corpus, model configuration or retrieved passages have changed. Retain enough information to associate the answer with the input question, source versions and product configuration used. The purpose is to reproduce the basis of the decision, not to promise that a probabilistic system will always generate identical prose.

The same concern appears in maintaining source version history for research products. A current link may point to a newer page than the one originally used. Enterprise records can change just as external sources do. The buyer therefore needs an agreed approach to expired records, updated procedures and answers saved outside the search application.

Exported answers deserve a separate check. If an employee pastes a response into a report, do the references preserve document identity and access conditions? A useful in-app citation may become an unexplained number in an exported file. Assess that handover using the actual document workflow, including what a colleague without identical permissions can inspect.

Price supported work, not attractive answers

A pilot should report the number of useful tasks completed with acceptable evidence, the time spent verifying them and the main reasons for escalation. These measures help a buyer decide whether the product belongs in exploratory search, drafting assistance or a more controlled review process. They also give the supplier a defensible account of where its product saves effort.

The TKMS and Cohere enterprise AI agreement illustrates why company knowledge and workflows are a substantial commercial setting for this technology. An announced agreement, however, does not supply an independent accuracy benchmark for another customer's records. The evaluation must remain tied to the intended corpus and decision.

A strong proposal makes its answer boundaries inspectable. It shows how source authority enters retrieval, how unsupported claims are recognised, and what survives when an answer leaves the application. That evidence gives commercial teams something more durable to sell than the presence of citations: a documented reduction in the work required to reach a supported conclusion.

Sources & evidence

  1. Enabling Large Language Models to Generate Text with CitationsAssociation for Computational Linguistics
  2. A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial LookNIST · 18 July 2025

Research findings are dated to their original studies. The customer evaluation examples and commercial interpretation are BDI analysis.

Suggest a correction