The stages
1. Normalization
Records are brought to one form before they are compared: names split into parts, case and accents folded, transliterations considered, organization names stripped of legal suffixes, URLs canonicalized, dates parsed.
2. Blocking
Comparing every record with every other grows with the square of the dataset. Blocking groups records by cheap keys (a surname, a name's phonetic code, an employer, a shared identifier) and compares only records within the same block.
3. Pairwise matching
Each candidate pair is scored field by field. The classic probabilistic model, described in "A Theory for Record Linkage" (1969), weighs each field by how likely it is to agree for a true match versus a non-match: agreement on a rare surname counts for more than agreement on a common one. Modern systems learn these weights, or use trained classifiers and text embeddings, but the principle is the same.
// A pairwise comparison of two person records, as weights per field.
const weights = { orcid: 12, email_domain: 3, employer: 4, city: 2, name: 1.5 };
function score(a, b) {
if (a.orcid && b.orcid) return a.orcid === b.orcid ? Infinity : -Infinity; // identifiers decide
let s = 0;
for (const field of ["employer", "city", "name"]) {
if (a[field] && b[field]) s += a[field] === b[field] ? weights[field] : -weights[field];
}
return s; // above a threshold: same person; below another: different; between: review
}4. Clustering
Pairwise decisions are combined into clusters, one per entity. Transitivity is the risk: if A matches B and B matches C, A and C end up together even if they conflict. Good systems check clusters for internal contradictions, such as two different ORCID iDs or two incompatible birth years, and split them.
5. Merging and canonical records
A cluster becomes one canonical entity with one identifier. Its values are chosen by source authority and recency, with provenance kept so a wrong merge can be undone.
Identifiers short-circuit the process
Shared persistent identifiers turn matching into a join: two records with the same ORCID iD or Wikidata QID are the same person, and two with different ones are not.sameAs links play a similar role on the open web. This is why registries and stable identifiers matter more than any matching algorithm. SeePerson identifiers compared.
Entity resolution in knowledge graphs
- Wikidata merges duplicate items by hand or with tools, leaving a redirect from the merged QID, and records external identifiers that let other datasets link to it.
- Reconciliation services match a name and a few properties against a dataset and return ranked candidates. OpenRefine popularized the protocol; the W3C Entity Reconciliation Community Group documents it as the Reconciliation Service API.
- Search engines resolve the entities in pages they crawl against their own knowledge graphs, using names, context, structured data and links. Their methods are not public, but the signals they document are the ones above.
POST https://reconcile.example/api
Content-Type: application/x-www-form-urlencoded
queries={"q0": {"query": "Lena Marlowe", "type": "Q5",
"properties": [{"pid": "P108", "v": "Brightfield Analytics"}]}}
# A reconciliation service answers with ranked candidates:
{"q0": {"result": [
{"id": "Q00000000", "name": "Lena Marlowe", "score": 92.4, "match": false,
"type": [{"id": "Q5", "name": "human"}]}
]}}Typical errors
- False merges of namesakes, the more damaging error for people, because facts from one person's life are attached to another.
- Missed matches when a person changes name, employer or city, or when their name is transliterated differently.
- Error propagation: a wrong
sameAsor a wrong merge in one dataset is copied by every dataset that trusts it.
How SelfBadge helps resolution
A SelfBadge profile gives a resolver what it needs in one place: a stable identifier (https://selfbadge.com/<handle>#person), registry identifiers, accounts proven to be the person's in sameAs, a disambiguating description, and the verification level and source of every fact in the.json version, so a resolver can weigh a verified employer above a self-declared one. When SelfBadge creates profiles from Wikidata, it matches them to people already on SelfBadge by identifiers only, never by name, and a claim that merges a seeded profile into a person's own is always reviewed by a person.
Specifications and sources
Related reference
- Disambiguation: how machines tell people with the same name apart.
- Person identifiers compared: Wikidata QID, ORCID iD, ISNI, VIAF, Library of Congress, Crunchbase and others.
- sameAs: how sameAs links one person across sites, best practices and common mistakes.
- How AI assistants identify people: the signals assistants use, and why namesakes get mixed up.
- Open data sources about people: Wikidata, DBpedia, OpenAlex and ORCID public data, and their licenses.
- All reference pages