BDI

Defense technology.
Buyers, markets, opportunities.

Evaluating entity resolution in an OSINT platform

Company research depends on knowing which records describe the same business and which describe related but separate entities. A platform trial should examine both kinds of error.

In this article
  1. Decide what the entity represents
  2. False merges and missed links have different costs
  3. Build a reference set around real research questions
  4. Inspect the company cluster, not only individual pairs
  5. Keep confidence distinct from a research conclusion
  6. Evaluate corrections as part of the product
  7. Price the usable research workflow
  8. Sources & evidence

Entity resolution is the work of deciding which records refer to the same entity. In public company research, that can mean connecting a contract announcement, a company filing and an older trading name to the correct business. It can also mean keeping two similarly named companies separate. A platform that gets those decisions wrong can distort a market map before an analyst begins interpreting it.

The commercial value is therefore broader than finding more matches. A useful OSINT platform should preserve distinctions between identity, ownership and other relationships, explain the evidence behind a proposed match and allow mistakes to be corrected without losing the research history. This guide focuses on public companies and institutions, rather than identifying or tracking private individuals.

Decide what the entity represents

A market research team may want to count legal companies, commercial groups, product brands or manufacturing sites. Those are different units. If the platform and the analyst use different definitions, a result can look wrong even when the underlying records have been linked consistently.

Begin the evaluation by defining the unit used in the intended report. A group-level account may legitimately collect several subsidiaries, while a contract analysis may need to preserve the specific legal entity named in each award. A product brand may remain relevant after its owner changes. The platform should support those distinctions explicitly rather than forcing every item into one flat company record.

GLEIF's Level 2 documentation offers a concrete example. It distinguishes identity reference data from relationships to direct and ultimate accounting-consolidating parents. That relationship has a defined scope; it should not be silently relabelled as every possible form of ownership or control.

False merges and missed links have different costs

Consider an illustrative industrial research dataset containing a parent business, a subsidiary with a similar name and an unrelated company using the same abbreviation. Combining all three into one record can attribute contracts or products to the wrong organisation. Keeping every spelling variation separate can instead fragment the history of a single business and overstate the number of competitors.

The first error creates a false merge. The second creates an unnecessary split. A supplier may reduce one while increasing the other, so a broad claim about matching accuracy is insufficient. The evaluation needs examples and counts that make both consequences visible.

The UK Government Data Quality Framework treats completeness, uniqueness and accuracy as separate dimensions. It also relates quality to the intended use. A database with fewer duplicate records is not automatically more accurate if the deduplication process has merged distinct entities.

Build a reference set around real research questions

A reference set should include cases for which the intended identity and relationship are established from suitable public evidence. Record why each case was labelled and preserve unresolved examples separately. A name resemblance alone is a weak basis for declaring that two records concern the same legal entity.

Include the types of difficulty the buyer actually encounters: abbreviated names, historical names, group structures, records from different jurisdictions and announcements that identify a brand rather than a company. These are categories for evaluating a legitimate commercial research product, not instructions for collecting additional personal data.

Avoid selecting only obvious matches. A trial based entirely on records with a common identifier will not reveal how the platform behaves when the most useful identifier is absent. Conversely, a collection made entirely of unusual edge cases cannot estimate routine review workload. Explain the selection so that the customer can interpret both the strengths and the limits of the result.

Inspect the company cluster, not only individual pairs

A 2024 research preprint by Binette and co-authors proposes an entity-centred approach to reference labelling and evaluation. It discusses both pairwise and cluster measures, with an application to inventor-name disambiguation. The paper is a methodological reference, rather than a validation of any OSINT supplier.

The distinction is useful for company research. A platform may make many correct pairwise links while still contaminating an important company record with one unrelated entity. That single connection can affect several downstream summaries if products, contracts and relationships are inherited across the combined record.

Review a sample of complete company records as well as individual match suggestions. Look at whether the assembled history is coherent and whether every important connection has an inspectable basis. A cluster view can reveal a mistake that a series of isolated pair decisions makes difficult to notice. It also better represents the object that an analyst will actually use in a report.

Keep confidence distinct from a research conclusion

A match score should identify what it means and how it is intended to be used. It might rank candidate links for human review, or it might support an automatic merge within a bounded workflow. The same number should not be presented as an unexplained probability that every statement in the resulting company record is true.

Ask how the product handles ambiguous cases. A useful review queue can present the competing records, the evidence supporting each interpretation and the consequences of accepting a match. If the analyst rejects a suggestion, the product should preserve that decision so that the same unresolved issue does not repeatedly consume time.

The customer should also know whether its corrections affect only its workspace or a shared dataset. Those are different service models. An internal analyst may be able to resolve a local research problem quickly while a centrally maintained provider needs a broader review process before changing the record used by many customers.

Evaluate corrections as part of the product

A trial should include a deliberately disputed company match with known reference evidence. The purpose is to observe the review and correction workflow, not to trick the supplier. Determine whether the analyst can separate records, preserve the source material and identify reports that relied on the earlier grouping.

The cost of a mistake depends partly on how widely it has travelled. A corrected company page is useful, but a commercial team may also have exported a spreadsheet or prepared a customer briefing using the previous version. The platform should make material corrections discoverable and provide enough information for the team to assess affected work.

Time matters as well. A record can accurately describe a past relationship while being unsuitable as a current statement. Preserve the date or period represented by the source alongside the date it entered the platform. The principles discussed in OSINT provenance and analytical standards help explain why a current interface should retain the history behind its conclusions.

Price the usable research workflow

Compare the software fee with the human review and correction work required to maintain a reliable company dataset. A service that supplies more suggested connections may be valuable, but those suggestions also need evaluation. The quotation should identify included reference data, update frequency, analyst support and export rights for corrected records.

For readers examining the supplier landscape, Blackdot Solutions' company profile provides context on one OSINT workflow business. The relevant procurement evidence remains the performance of the exact product, data package and review process the customer intends to use.

A strong entity-resolution offer makes uncertainty and correction manageable. It helps the customer build a coherent account of a business without erasing the differences between a legal entity, a group relationship and a commercial brand. That is the foundation for useful industry research and defensible competitive reporting.

Sources & evidence

  1. The Government Data Quality FrameworkUK Government Data Quality Hub · 3 December 2020
  2. How to Evaluate Entity Resolution SystemsBinette and co-authors / arXiv · 8 April 2024
  3. Level 2 Data: Who Owns WhomGLEIF

This guide concerns public company and institutional research. It uses government data-quality guidance, GLEIF's documented relationship scope and a research preprint as methodological references.

Suggest a correction