Two sources of knowledge
A language model answers from two places. The first is its training data: text collected up to a cutoff date and compressed into the model's weights. It cannot be corrected after training, and it explains why old roles linger. The second is retrieval: a live web search run while answering, whose results are added to the model's context. ChatGPT, Claude, Gemini and Perplexity all offer web search, and their APIs expose it as a tool (OpenAI's web search tool, Anthropic's web search tool, Gemini's grounding with Google Search; Perplexity's Sonar models always search). Retrieved pages are where new, clear sources change an answer quickly.
The signals that tie pages to one person
- The name and its variants: full name, short form, earlier names, transliterations.
- Context that travels with the name: employer, job title, city, field and dates. "Lena Marlowe, head of product in London" and "Lena Marlowe, novelist in Leeds" are separable by any one of them.
- Links between accounts: a website that links to a profile that links back, or links marked
rel="me". Two-way links are strong evidence of one owner. - Structured data: schema.org Person markup states name, role, employer and accounts (
sameAs) without guessing, and an@idgives the person a stable identifier across pages. - Knowledge bases: Wikidata and search engines' knowledge graphs hold identifiers for well-known people. Most professionals are not in them.
- Source authority: an organization's own team page or a verified profile counts for more than a scraped directory.
Why namesakes get merged
Retrieval matches on words, and a name is the strongest word in a question about a person. If the pages about one person rarely state their employer, city or field, the search results mix in pages about namesakes, and the model may blend them into one biography. The better-documented namesake usually dominates, because more retrieved pages agree about them. The remedy is the same as in any entity resolution: identifiers, linked accounts and context stated together. See Disambiguation.
Why some people are missing
Many professionals have most of their public record on sites that require a sign-in or limit crawlers. To a retrieval system such a person can look almost absent, and a model may answer that it has no information, or worse, fill the gap with a namesake. An open page that states the same facts, and links to those accounts, fills the gap.
Failure modes, in short
- Merged namesakes: another person's job, age or history attached to the wrong person.
- Stale facts: a role the person left years ago, because older pages outnumber newer ones.
- Missing person: no open source to retrieve.
- Wrong accounts: an account that is not the person's presented as theirs.
- Invented details: plausible facts with no source, where retrieval returned too little.
What a machine-readable profile provides
A SelfBadge profile states each signal in the form machines read best:
- A plain-language summary sentence, and the facts again as short sentences, in the page and in the
.mdversion. - schema.org JSON-LD: a ProfilePage whose main entity is the Person, with a stable
@idsuch ashttps://selfbadge.com/<handle>#person. - Only accounts proven to be the person's in
sameAs, and registry identifiers inidentifier. - A disambiguating description that says how to tell the person apart from namesakes.
- A lookup API and an MCP server, so an agent can look the person up directly instead of searching. See Reading a SelfBadge profile.
Specifications and sources
Related reference
- Disambiguation: how machines tell people with the same name apart.
- Entity resolution: how knowledge graphs merge records about the same person.
- sameAs: how sameAs links one person across sites, best practices and common mistakes.
- The AI check method: how the AI check asks each engine, reads its sources and analyzes the answers.
- Reading a SelfBadge profile: for bots and agents: JSON-LD, .json, .md, llms.txt, the API and the MCP server.
- All reference pages