Customer Uniqueness
On the Stored data tier, an accepted onboarding promotes a Digital Identity to a Customer — either a new Customer record is created, or the identity is attached to an existing one. A Customer is meant to represent one real person. Whether that holds is decided by the duplicity checks in the workflow and by the rules that act on their result. This page explains how the checks work together, how much of the decision can be automated, and how to design a workflow that does not let duplicates through.
Once two Customer records exist for the same person, every system that consumed them inherits the split: assignments are divided between the records, authentication has two candidates, reports count the person twice, and an erasure request has to find every copy. Merging afterwards is operator work in the back office plus a correction of the identifiers your own systems already stored. Treat uniqueness as a workflow design decision, not as database clean-up.
Two signals, one decision
Two independent checks establish whether the applicant already exists:
| Check | What it compares | Blind spots when used alone |
|---|---|---|
| Face duplicity check (1:N) | The selfie against the faces in the watchlists. | Twins and look-alikes; strong changes of appearance over time; a genuine person re-registering with the same face but a different, forged identity is a face hit with different data. |
| Document duplicity check (1:N) | Name, date of birth, personal and document numbers against the records in the watchlists. | Identity data reused by another person; namesakes; name changes and re-issued documents; spelling variants across documents. |
Each check covers the other's blind spots. Combined, they answer the two questions that matter — is this the same face? and is this the same identity data? — and the pattern of the two answers tells a returning customer apart from the different flavours of identity misuse. The best_duplicate workflow step performs that combination.
The goal is to automate as much of the decision as the evidence allows: clear duplicates are merged and clear newcomers are created without an operator. Where the two signals disagree, the platform provides the evidence and the routing, but the decision stays with a person.
Combined duplicity outcome
After dedup_face and dedup_document have run, the best_duplicate step synthesizes both signals into one result that decision rules can branch on:
best_duplicate.outcome— how the best candidate relates to the new identity (table below);best_duplicate.type— the watchlist of that candidate:customer,blocklist,concurrentorreview;best_duplicate.hit— the candidate record itself, used by the outcome actions (merge, block, review).
| Face comparison | Document comparison | outcome | Interpretation |
|---|---|---|---|
| Match | Match | full_hit | The same person — a returning customer or a repeated attempt. |
| Similar (review band) | Match | similar_face | Same identity data, face changed significantly — ageing, or an attempt with the data of someone else. |
| Match | Similar | similar_name | Same face, slightly different data — twins, a name change, or a re-issued document. |
| Match | No match | face_missmatch | Same face under different identity data — a look-alike or a fraud attempt. |
| No match | Match | doc_missmatch | Same identity data on a different face — the data may be reused by another person. |
| — | — | no_match | No candidate reached any threshold. |
Several candidates are full_hit | multiple_full_hits | More than one record fully matches — the thresholds are too permissive or the database already holds duplicates. |
When several candidates exist, the outcome is that of the highest-ranked candidate. Ranking goes full_hit, then similar_name, face_missmatch, similar_face, doc_missmatch; within a group the higher face score wins, because the face signal is the stronger evidence. Hits from both searches are collected into one hitlist (up to 20 candidates), each carrying a face score and a text score.
In the back office, the Duplicates tab of the Digital Identity detail lists every hit with both scores side by side and an outcome badge, so an operator sees at a glance why a case was flagged.
What can be decided automatically
| Outcome | Automatic decision | Why |
|---|---|---|
no_match | Create a new Customer. | Neither signal found the person. |
full_hit, type customer | Merge — attach the identity to the existing Customer. | Both signals agree it is the same person. |
Any outcome, type blocklist | Reject. | The person was blocked before; the reason for blocking does not expire. |
Any outcome, type concurrent | Hold for review until the other session finishes. | Two sessions of one person are running at once; only one may create the Customer. |
similar_name | Review. | Twins and name changes are genuine; the same face with altered data is not. The evidence alone cannot tell them apart. |
similar_face | Review. | A person who aged, changed hairstyle or lost weight looks like this — so does someone presenting the data of another person. |
face_missmatch | Review, or reject by policy. | A look-alike is rare but genuine; a face reused under different data is not. |
doc_missmatch | Review, or reject by policy. | Identity data appearing with a different face is a strong fraud indicator, but data-entry errors and shared family documents exist. |
multiple_full_hits | Review. | The database already contains duplicates or the thresholds are too loose; an operator has to pick the record. |
The clear cases — no_match, full_hit, anything on the Blocklist — can and should be decided automatically. The remaining outcomes describe situations where the same evidence is produced by a rare genuine event and by a fraud attempt. A rule can only choose between review (an operator decides, with the cost of a queue) and reject (the genuine rare cases are turned away, with the cost of lost customers). Which trade-off is right depends on the risk of the use case and is a business decision — the platform makes both configurable, but cannot make the ambiguity disappear.
The distinction between fraud flavours is exactly what the two-signal combination buys: a stolen identity shows as doc_missmatch, a forged document on a known face as face_missmatch, a repeated onboarding as full_hit. The business reading of each combination is tabulated in Identity Verification → Duplicity check during onboarding.
Designing a workflow that cannot create duplicates
Every path that ends in create_new_customer must be gated by the duplicity checks. Check a custom workflow against this list before publishing it:
- Run both checks before deciding.
dedup_face,dedup_documentandbest_duplicatecome before thedecisionstep of every outcome that may create a Customer. - Create only on
no_match. The rule that leads tocreate_new_customertestsbest_duplicate.outcome == 'no_match'. Every other outcome value has an explicit rule — merge, review or reject. Do not let adefaultrule create a Customer; a signal you forgot to handle would then silently become a new record. - Handle concurrent sessions. Route
type == 'concurrent'to review with1N_move_to_review, and resolve the held sessions in the outcome of the winning one withreject_concurrentsormerge_rejected_concurrents. Without this, two simultaneous sessions of one person can each create a Customer. - Never create on Incomplete or Rejected. Terminal states other than accept must have no Customer-creating action. The predefined workflows follow this; keep it when forking.
- Know what your tier runs. Without the 1:N add-on the face check is skipped at run time and uniqueness rests on the document check alone. Keep
dedup_documentin the workflow and consider the add-on if face-based misuse is a concern for your use case. - Do not bypass the workflow. Customers created through the Customer API — for example during a migration — are not checked for duplicates; only the uniqueness of the
external_idis enforced. De-duplicate the source data before importing it. - Calibrate and watch. Tune the match and similar thresholds in IDV Configuration on your own population. A rising share of
multiple_full_hitsmeans duplicates are already in the database or the thresholds are too loose; a rising review rate means they are too tight. - Staff the review queue. The ambiguous outcomes only work as a safety net if someone decides them within your service-level target — see Manual Review.
When a duplicate slips through anyway
Operators resolve duplicates in the Duplicates tab of the Digital Identity detail: Merge with Customer attaches the identity to the right record, Resolve as Unique confirms the person is genuinely new, and Put as new on Blocklist stops a repeat attempt. Every resolution moves the corresponding watchlist members so that the next search sees the corrected state.
What the platform cannot do is repair the systems that already consumed the duplicate — an account, a SIM, an insurance policy issued against the second record. Those identifiers have to be re-pointed on your side, which is why preventing the duplicate in the workflow is the cheaper path.
See also
- Face Duplicity check (1:N) and Document Duplicity check — how each signal is computed
- Watchlists — the galleries the checks search
- Workflow Manifest Reference — the steps, signals and outcome actions used above
- Digital Identity Lifecycle — how identities become Customers
- Manual Review