Summary

An entity registry is a single, authoritative list of the real-world things your business deals with: customers, suppliers, people, products, sites and contracts. Each entity gets one ID, every name and code it goes by across your systems, its key facts, its status, and a record of where each fact came from. It is what lets an AI recognise that "Smith Manufacturing" in the CRM and "Smith Mfg Ltd" on an invoice are the same company, and that "Dave at Smith's" in an email works there. People reconcile these differences from memory. AI only sees the records it is given, so without a registry it treats one customer as three and gives answers that are confidently incomplete. Building one no longer needs a heavyweight master data project: rules and matching tools handle the clear cases, language models judge the ambiguous ones, and people review what is still uncertain, one process at a time.

An entity registry is a single list of the real-world things your business deals with (customers, suppliers, people, products, sites and contracts), where each one has one ID, every name and code it goes by across your systems, its key facts, and a record of where each fact came from. It is what lets an AI recognise that "Smith Manufacturing" in the CRM and "Smith Mfg Ltd" on an invoice are the same company, and that "Dave at Smith's" in an email works there. Your people make those connections without thinking. Your AI can only connect what it has been given, and without a registry it will treat one customer as three. It is the least glamorous part of making AI work, and the part everything else rests on.

The problem it solves

Every growing business accumulates the same mess. A key customer is spelled four ways across the CRM, the finance system and the support desk. A supplier was acquired and renamed, so half its history sits under the old name. A contact moved from one client to another and now appears in both. Account codes changed in a reorganisation, and nobody went back to update the old records.

People cope with this remarkably well. They carry the connections in their heads: they know that the invoice from the renamed supplier belongs with the old contract, and that the email from Dave is about the Smith account. Much of what we call experience is exactly this kind of quiet reconciliation.

AI takes records at face value. Ask it for everything that happened with a customer and it returns a third of the history, because the other two thirds are filed under names it did not connect. The AI is not wrong about what it found. It is wrong about what it missed, and it has no way to know the difference.

What goes into one

An entry in an entity registry is less exotic than the name suggests. For a single customer it might look like this.

FieldExample
Registry IDCUST-00417
Canonical nameSmith Manufacturing Ltd
Also known asSmith Mfg Ltd, Smith Manufacturing, Smiths
IDs in other systemsCRM account 88213, finance account SMITH01, support org 5521
Strong identifiersCompany no. 01234567, VAT GB123456789
StatusActive (previously traded as Smith Engineering until 2023)
Key relationshipsContracts, open issues, main contacts, parent company
ProvenanceEach fact linked to the system or document it came from, and when

Two fields do most of the work. "Also known as" and "IDs in other systems" are what let records from different places be joined together. Provenance is what lets anyone, human or AI, check where a fact came from before relying on it.

How entity resolution works

Filling the registry is called entity resolution: deciding which records refer to the same real thing.

Start with strong identifiers. If two records share a company registration number, they are almost certainly the same legal entity. A VAT number or a corporate email domain is strong supporting evidence, not proof, because groups share them and subsidiaries borrow them. Where those identifiers are filled in, they settle a good share of the duplicates.

Then use weaker evidence together: similar names, the same address, the same people, the same products or contracts. No single clue is decisive, but several agreeing clues are persuasive. Rules and established matching tools handle the clear cases cheaply. Language models earn their place on the ambiguous pairs: they are good at judging whether "Smith Mfg" at one address and "Smith Manufacturing Ltd" at a nearby one are the same firm, and at setting out the evidence for a reviewer.

Finally, people review the uncertain middle. The confident matches are merged automatically, the confident non-matches are left alone, and a person decides the genuinely ambiguous cases. Merges link the source records rather than overwriting them, and every decision is recorded, so a wrong merge can be undone, though not any action already taken on it.

Isn't this just master data management, or our CRM?

It has the same lineage as master data management (MDM), and the goal is identical: one trusted version of each important thing. The difference is in how it is built and what it covers.

Traditional MDM programmes were often large, slow and expensive, focused on structured systems, and run as multi-year projects. Many stalled under their own weight. A modern entity registry is usually scoped to one process at a time, uses language models to do most of the matching, and deliberately includes unstructured sources: documents, contracts and email, the most underrated data source in your business.

Your CRM, meanwhile, knows its own records well and everyone else's not at all. It may have excellent duplicate detection within itself. It usually has no idea that its account 88213 is the finance system's SMITH01. An entity registry is the one place that knows both.

Why AI makes it urgent

The mess described above has existed for decades. It was tolerable because humans were doing the reconciling. AI changes that in two ways.

First, questions. Useful questions span systems. Which customers with an open complaint are due for renewal? The answer depends on joining records correctly. Without a registry, the answer is incomplete and nobody can see that it is. I explain how this shows up in practice in why your RAG system keeps missing what is in your documents.

Second, actions. An agent that updates a record, sends an email or raises an invoice needs to act on the right entity. A wrong merge or a missed match is no longer a reporting error. It is a wrong action, taken quickly and at scale.

This is why the registry sits at the centre of a knowledge graph, as I describe in GraphRAG vs vector RAG, and at the centre of the broader argument in your AI does not need a bigger model, it needs to know your business.

An old problem in new clothes

I first met this problem long before anyone called it entity resolution for AI. I built a maritime surveillance system, now used in more than 30 countries, that fused radar, vessel transponders and other sensors into one live picture. A core problem was deciding, continuously, which radar returns and which transponder signals described the same ship. It also had to stay up: one client later called it "the first and only system they had ever used that never crashed".

A related lesson carried through to the agentic AI we built for a European insurance brokerage, where a unified schema across legacy systems had to work before any models were built. Different decade, different data, same foundation.

How to start

You do not need a company-wide programme. You need one process and a few weeks.

Pick a process where getting the entity wrong is expensive: renewals, collections, supplier risk, claims. List the three to five entity types it depends on, usually customers, contacts, contracts and one or two others. Pull those records from the systems involved and resolve them, strong identifiers first, with uncertain matches reviewed by someone who knows the business.

Then connect the documents and emails that mention those entities, so the registry knows not just who someone is but what has been said and agreed with them. Keep it current by resolving new records as they arrive, rather than in an annual clean-up. Measure two things: how many duplicates you found, and which questions you can now answer that you could not before. Both numbers are usually the best argument for the next process. Finding where those records live in the first place is what I cover in what I look at in a knowledge audit.


If your systems disagree about who your customers are and you want to fix that before you build AI on top, I am happy to help. Let's talk.

Related: Your AI does not need a bigger model. It needs to know your business · GraphRAG vs vector RAG. Which one does your business need? · Email is your most underrated data source

Frequently asked questions

What is an entity registry?
An entity registry is a single list of the real-world things a business deals with, such as customers, suppliers, people, products and contracts. Each entity has one identifier, all the names and codes it is known by in different systems, its key attributes and status, and the source of each fact. It gives people and AI systems one consistent answer to the question of who or what a record refers to.
What is entity resolution?
Entity resolution is the process of working out which records, across different systems and documents, refer to the same real-world thing. It matches on strong identifiers such as company registration numbers first, then on supporting evidence such as VAT numbers, email domains, names, addresses and context, with uncertain matches reviewed by a person. The result is recorded in the entity registry.
Is an entity registry the same as master data management?
They share the same goal: one trusted version of each important thing. Traditional master data management projects were often large, slow and focused on structured systems. A modern entity registry is usually scoped to one process at a time, includes unstructured sources such as email and documents, and uses language models to judge the ambiguous matches, which makes it far lighter to build.
Why does AI need an entity registry?
Because AI takes records at face value. A person knows that two spellings of a customer are the same company; an AI treats them as two, so it misses half the history, double-counts, or acts on the wrong record. For any question or action that spans systems, the AI needs to know which records describe the same thing, and that is exactly what a registry provides.
How do you start building an entity registry?
Pick one process that matters, list the three to five entity types it depends on, pull those records from the systems involved, and resolve them, strong identifiers first and uncertain matches reviewed by a person. Then connect documents and email to the resolved entities and keep the registry current as new records arrive. Expand to the next process once the first is working.
Stay ahead

AI & tech are moving fast.
Get the signal, not the noise

Ready to make AI actually work?

Tell me what you're working on. I'll respond personally. If there's a fit, we'll take it from there.

Limited Fractional CTO capacity · Knowledge Audits start within two weeks