Skip to main content

Document Duplicity Check

Document duplicity check finds duplicates by the data on the document — name, date of birth, personal number, document number — rather than by face. It catches people re-registering under the same identity details and provides the text signal that complements face duplicity check.

Document duplicity check requires the Stored data tier. It does not require the biometric 1:N add-on — it works on its own, and when face 1:N is also enabled the two signals are merged into one combined duplicity outcome.

1:1 Document Compare​

In order to decide, whether 2 documents are of a same person, rules for comparison need to be defined. Different countries may have different conventions for giving a name and may have different personal identifiers. Thus, each integration may have different comparison rules.

What is compared​

The integrator needs to define, which fields are compared, and how the different comparison scores are weighted to create one matching score that is compared against a threshold.

StrengthFields
Strong (identity anchors)Date of birth, personal number
SupportingFull name, given names, surname, document number

How matching works​

  1. Text cleaning — names are normalized (titles stripped, lower-cased, punctuation removed) so cosmetic differences do not cause misses.
  2. Similarity scoring — names are compared with a fuzzy string-similarity measure that tolerates typos and minor spelling variation; exact fields like date of birth and personal number act as strong anchors.
  3. Weighted score — the platform combines the field scores into a single weighted score and compares it against a match and a similar threshold. The comparison fields, their weights, and the thresholds are configurable per tenant in IDV Configuration → Document dedup, as a guided form or as raw JSON.

The document duplicity check searches all existing records of the tenant, not only a single known Customer. When a workflow reaches the dedup_document step, the document data of the new identity is searched against the watchlists of the tenant:

  1. Name vector — the name fields of the identity are combined and transformed into a numeric name vector. The vector tolerates spelling variants, transliteration and typos, and it is non-reversible: the name cannot be reconstructed from it, so the search index holds no readable personal data.
  2. Vector search — the name vector is searched across the Customer, Blocklist and In-progress watchlists — the same three galleries used by face 1:N. Records whose personal number or document number matches exactly are added to the candidates regardless of the name.
  3. Re-ranking — the candidates are re-scored with the 1:1 comparison rules of the tenant and cut at the match / similar thresholds. The search is designed to answer in about a second even for databases of 100 million records.
  4. Face cross-check — every text hit is compared 1:1 by face against the new identity. This is what lets the combined outcome tell a returning customer from someone using the identity data of another person.

Outcomes​

dedup_document.result is derived from the watchlist of the top hit, using the same vocabulary as face duplicity check:

ResultMeaning
no_matchNo existing record shares the document data.
customerThe data matches an existing Customer — the attempt can be merged or sent to review.
blocklistThe data matches a blocked person — the attempt is rejected.
concurrentThe data matches another verification that is running right now.
reviewThe data matches an identity that is waiting for manual review.

Hits from the face and the document search are collected into one hitlist (up to 20 candidates), each carrying both a face and a text score.

Combining document match with a face match​

The real value comes from cross-checking the document signal against the face signal: a full text match with a non-matching face and a matching face with different text point to different kinds of misuse, while a match on both is a returning customer. The best_duplicate step performs this combination and produces a single outcome for the decision rules — see Customer Uniqueness → Combined duplicity outcome.

See also​