Thinklytics

CRM Data Quality · 10 min read · October 2026

Native CRM deduplication or cross-system customer matching? The count that decides

By Sean Majidi, Founder, Thinklytics

One count decides this, and it is not the duplicate rate. It is how many systems in your estate can create a customer record. If the answer is one, the native tooling is right and cheaper. If it is more than one, native dedupe will keep reporting a clean CRM while the duplicates live between systems.

Two proposals. One configures the deduplication features in the CRM you already pay for. The other builds an identity layer across every system that touches a customer. The second costs considerably more.

One count decides which is right, and it is not the duplicate rate.

Count the systems that can create a customer

Native CRM deduplication against cross-system matching

These solve different problems. Choosing on price rather than on where the duplicates are created is the common mistake.

Native CRM deduplicationCross-system matching
What it comparesRecords inside one CRMRecords across every system that creates a customer
Where it blocks a duplicateOn create and on import, in that CRMAt the identity layer, so every system resolves to the same entity
What it cannot seeThe same customer in billing, support, the web store or an acquired entityNothing, if every creating system is in scope
Cost and timeDays to weeks, often a configuration jobWeeks to months, and it needs a data platform under it
Right whenOne CRM creates the records and the duplicates are typing variantsSeveral systems create customers, or entities were acquired
Wrong whenBilling and the web store also create customersThere really is one source and the problem is data entry

Start by counting the systems that can create a customer record. If the answer is one, the native tooling is the right answer and the cheaper one. If it is more than one, native deduplication will keep reporting a clean CRM while the duplicates live between systems.

Source: Thinklytics engagement pattern across the identity resolution and master data engagements in the case library.

Not the systems that store customers. The systems that can create one.

Usually that list is longer than expected: the CRM, the billing platform, the web store, the support desk, the events or webinar tool, the marketing platform if forms write directly, and anything that arrived with an acquisition.

If the answer is one, native deduplication is the right answer and the cheaper one. It compares records inside that CRM on create and on import, catches typing variants and pasted imports, and is often a configuration job measured in days.

If the answer is more than one, native deduplication will keep reporting a clean CRM while the duplicates accumulate between systems. That report is true and useless, because it describes a subset of the records your customers actually exist in.

The reason the count matters more than the rate is that the rate tells you how bad the stock is and the count tells you where the flow comes from. The six flows are covered in why duplicate CRM records keep returning.

The matching itself

Deterministic and probabilistic matching, and where each belongs

Most estates need both. Running only one is how a project either misses most duplicates or merges two real customers.

MethodHow it decidesWhere it fits and what it costs you
DeterministicExact or normalised match on a key: tax ID, policy number, email, account numberSafe and auditable. Misses everything with a typo, a trading name or a changed email
ProbabilisticA score across several weak fields: name, address, phone, purchase historyCatches what deterministic misses. Produces false matches, so it needs a threshold and a review queue
Both, in orderDeterministic first, probabilistic on the remainder, humans on the uncertain bandWhat we run. An insurer normalised policy numbers, matched insured names and aligned effective dates across four systems
Survivorship rulesWhich field wins when two records merge, per fieldAgreed with the data owners before any merge runs, or the merge creates its own re-entry problem
Audit trail on every mergeWhat merged, on what rule, and how to unpick itThe thing that makes the merge reversible. Non-negotiable

A retailer ran deterministic and probabilistic matching across eight systems to resolve 2.3 million duplicate records into 1.4 million unique profiles. An SAP manufacturer agreed matching and survivorship rules with the data owners first, then deduplicated customers, vendors and materials with audit trails.

Source: Thinklytics case library, published approaches and outcomes per engagement.

Both approaches have to decide when two records are the same thing, and there are two ways to decide.

Deterministic. Exact or normalised match on a key: tax ID, policy number, email, account number. Safe, auditable, explicable to an auditor. It misses anything with a typo, a trading name, a changed email or a different address format.

Probabilistic. A score across several weak fields: name, address, phone, purchase history. It catches what deterministic misses, and it produces false matches, which is why it needs a stated threshold and a review queue rather than being left on a default.

Most estates need both, in that order. Deterministic first, probabilistic on what remains, humans on the uncertain band. A national retailer ran exactly that across eight systems to resolve 2.3 million duplicate records into 1.4 million unique profiles. A regional insurer normalised policy numbers, matched insured names and aligned effective dates across four claims systems before 1.2 million policy records were usable.

Two things that are not optional.

The threshold is yours to set. A missed duplicate repeats information. A false merge destroys it, combining two real customers into one record whose history is now wrong, and the customer usually finds out before you do. Those errors are not equally bad, so set the threshold conservatively and staff the review queue with people who know the accounts.

Survivorship rules, field by field, agreed with the data owners before any merge runs. The default is usually most-recent-wins, which loses the right value on any field where the newer record was created in a hurry. Then whoever needed that value re-creates the record they had, and the merge has manufactured a duplicate. An SAP manufacturer agreed matching and survivorship rules with the data owners first, then deduplicated customers, vendors and materials with audit trails, and turned every cleansing decision into a documented transformation rule so the work was repeatable.

What the engagements measured

Eight identity engagements, what was measured

Published outcomes. The ones with a before and after rate are the useful ones, because a count on its own does not say whether it worked.

EstateWhat was resolvedDelivery and effect
Venue group, 8 ticketing systemsDuplicate fan records from 184,000 to 11,200 within 800,00014 weeks. $940K annual revenue leakage stopped, campaign failures from 144 to 24 a month
National retailer, 8 channels2.3M duplicate records to 1.4M unique profiles14 weeks. $4.2M incremental revenue in 90 days, email open rate from 18 to 31 of 100
Pharmacy benefit managerMember match accuracy from 75 to 94 of every 10014 weeks. $4.8M a year in misrouted claims recovered, 3 stalled ML pilots restarted
Regional P&C insurer, 4 claims systems1.2M policy records unified13 weeks. Loss ratio reporting accuracy from 71 to 97 of 100
Workers comp carrierPolicy data quality score from 61 to 94 on the vendor's own 10-category rubric16 weeks at $340K against a $2.8M vendor quote. $8.4M annual underwriting value unblocked
Casino group, 5 property systems14,200 players active at more than one property identified18 weeks. $2.4M marketing opportunity, $680K new revenue in 90 days
SAP manufacturerDuplicate vendor and material records from 41,000 to 12,0009 weeks. $640K annual reconciliation and rework labour avoided, go-live date held
SAP distributorMaster data quality score from 58 to 89, 31,000 obsolete records flagged7 weeks, before a go-live date was set

Eight engagements, 7 to 18 weeks. Duration tracked the number of creating systems and whether survivorship had to be negotiated, not the record count.

Source: Thinklytics case library, published delivery durations and outcomes per engagement.

Eight identity engagements in our case library published a measured outcome. Seven to 18 weeks, and the duration tracked the number of creating systems and whether survivorship had to be negotiated, not the record count.

The ones with a before and after rate are the useful ones. A venue group went from 184,000 duplicate fan records inside 800,000 to 11,200, so 23% to 1.4%. A pharmacy benefit manager lifted member match accuracy from 75 to 94 of every 100 records. A regional insurer took loss ratio reporting accuracy from 71 to 97 of 100. An SAP manufacturer cut duplicate vendor and material records from 41,000 to 12,000.

A count of merges tells you nothing on its own, which is why it is the number most proposals offer.

One of these is worth reading as a procurement lesson rather than a technical one. A workers compensation carrier had spent $1.2M on an AI underwriting platform that required a data quality score of 80 across 10 categories. Its policy data scored 61. The platform vendor quoted $2.8M to fix it. Reading the vendor's own rubric showed eight failing dimensions, six of which were addressable with automated pipelines standardising address formats, resolving duplicate policy IDs, filling missing coverage start dates and normalising coverage code mappings. The remaining two were legacy records needing manual review. The score reached 94 in 16 weeks for $340K. See the policy data quality engagement.

What cross-system matching does not require

It does not require replacing the CRM, and proposals sometimes imply otherwise.

In all eight of those engagements the source systems were kept and the resolved identity was built above them, feeding back into the operational tools. A retailer loaded resolved profiles into a warehouse and connected them to the marketing automation platform. A casino group built a player identity above five property management systems and surfaced it to the marketing team. The CRM keeps doing what it does. It stops being asked to answer a question about records it cannot see.

What has to be true either way

Four things, and they are independent of which approach you choose. A project that skips them refills the stock regardless of how good the matching was.

A block at the point of entry, live in every system that can create a record. A no-match queue a human clears, because a silent insert on no-match is the largest single source of re-accumulation. A reversible merge with an audit trail, tested before the first production run. A monitored duplicate rate with a stated threshold and a named owner after the project ends.

The venue group did the match, then the block, then the real-time monitoring. Three steps, and the third is the one usually cut for budget. It is the reason the rate held.

What we would do first

Write the list of systems that can create a customer record, and put a name next to each one for who owns it. That list is the scope boundary and it takes an afternoon.

Then measure the duplicate rate on a stated denominator, inside the CRM and across the estate separately. If the two numbers are close, native deduplication is your answer. If the estate rate is far worse than the CRM rate, you have found where the duplicates live and the native tooling was never going to reach them.

What to put in the contract, and what to ask before signing, is in what to ask before a CRM data cleanup project, and the CRM data access and rollback checklist is the worksheet we hand to whoever has to approve the access these projects need.

Delivery sits in master data management for the identity layer, data cleaning and preparation for the remediation, data 360 consultant where the customer view spans channels, vendor-neutral system integration for the no-match queues, and Salesforce Data Cloud consulting on Salesforce estates. The full set of work in this area sits under our systems need people in between them.

Frequently asked questions

How do I choose between native CRM deduplication and cross-system matching?

Count the systems in your estate that can create a customer record. If the answer is one, native deduplication is the right answer and the cheaper one, usually days to weeks of configuration. If the answer is more than one, which it is for most companies above a few hundred staff, native tooling will keep reporting a clean CRM while the duplicates accumulate between systems where it cannot see them. The count is the decision. The duplicate rate is not.

What does native CRM deduplication actually cover?

Records inside that one CRM, compared on create and on import, which is exactly what it claims. It catches typing variants, pasted imports and the same company entered twice by two reps. It cannot see the same customer in billing, the support desk, the web store or an entity that arrived with an acquisition, because those records are not in the CRM. A clean report from it is a true statement about an incomplete picture.

What is the difference between deterministic and probabilistic matching?

Deterministic matches on an exact or normalised key: tax ID, policy number, email, account number. It is safe and auditable and it misses anything with a typo, a trading name or a changed email. Probabilistic scores across several weak fields such as name, address, phone and purchase history, catches what deterministic misses, and produces false matches, so it needs a threshold and a human review queue. Most estates need both, deterministic first and probabilistic on the remainder.

Who should set the match threshold?

You, not the tool and not the implementer. The threshold is a business trade between missed duplicates and wrongly merged customers, and those two errors are not equally bad. A missed duplicate repeats information. A false merge destroys it, combining two real customers into one record whose history is now wrong. Set the threshold conservatively, put the uncertain band into a review queue, and have the reviewers be people who know the accounts.

What are survivorship rules and why do they matter?

They decide which value wins, field by field, when two records merge. Without them the default is usually most-recent-wins, which loses the correct value on any field where the newer record was created in a hurry. Then the person who needed that value re-creates the record they had, and the merge has manufactured a duplicate. Agree the rules with the data owners before any merge runs. An SAP manufacturer did exactly that before deduplicating customers, vendors and materials with audit trails.

How long does each approach take?

Native deduplication is typically days to weeks and often a configuration job. Cross-system matching in our case library ran 7 to 18 weeks across eight engagements, and the duration tracked the number of creating systems and whether survivorship had to be negotiated rather than the record count. Eight ticketing systems took 14 weeks. Eight retail channels took 14. Four claims systems took 13. Duplicate vendors and materials in one SAP estate took 9.

Does cross-system matching mean replacing the CRM?

No. In every one of the eight identity engagements in our case library the source systems were kept and the resolved identity was built above them, feeding back into the operational tools. A retailer loaded resolved profiles into a warehouse and linked them to the marketing automation platform. A casino group built a player identity above five property management systems. The CRM keeps doing what it does; it stops being asked to answer a question about records it cannot see.

What has to be true for either approach to hold?

A block at the point of entry in every system that can create a record, a no-match queue a human clears rather than a silent insert, a reversible merge with an audit trail, and a monitored duplicate rate with a threshold and a named owner. None of those are specific to the approach, and a project without them refills the stock regardless of how good the matching was. A venue group did the match, then the block, then the monitoring, which is why its rate held.

The work behind this

Eight identity resolution engagements in the case library ran 7 to 18 weeks with a published measured outcome, and in all eight the source systems were kept and the resolved identity was built above them.

Data cleaning and labeling, 22 engagements.

Topics covered

  • native CRM deduplication
  • cross system customer matching
  • identity resolution
  • deterministic versus probabilistic matching
  • survivorship rules
  • master data management versus CRM dedupe
  • customer 360

Frequently asked questions

How do I choose between native CRM deduplication and cross-system matching?

Count the systems in your estate that can create a customer record. If the answer is one, native deduplication is the right answer and the cheaper one, usually days to weeks of configuration. If the answer is more than one, which it is for most companies above a few hundred staff, native tooling will keep reporting a clean CRM while the duplicates accumulate between systems where it cannot see them. The count is the decision. The duplicate rate is not.

What does native CRM deduplication actually cover?

Records inside that one CRM, compared on create and on import, which is exactly what it claims. It catches typing variants, pasted imports and the same company entered twice by two reps. It cannot see the same customer in billing, the support desk, the web store or an entity that arrived with an acquisition, because those records are not in the CRM. A clean report from it is a true statement about an incomplete picture.

What is the difference between deterministic and probabilistic matching?

Deterministic matches on an exact or normalised key: tax ID, policy number, email, account number. It is safe and auditable and it misses anything with a typo, a trading name or a changed email. Probabilistic scores across several weak fields such as name, address, phone and purchase history, catches what deterministic misses, and produces false matches, so it needs a threshold and a human review queue. Most estates need both, deterministic first and probabilistic on the remainder.

Who should set the match threshold?

You, not the tool and not the implementer. The threshold is a business trade between missed duplicates and wrongly merged customers, and those two errors are not equally bad. A missed duplicate repeats information. A false merge destroys it, combining two real customers into one record whose history is now wrong. Set the threshold conservatively, put the uncertain band into a review queue, and have the reviewers be people who know the accounts.

What are survivorship rules and why do they matter?

They decide which value wins, field by field, when two records merge. Without them the default is usually most-recent-wins, which loses the correct value on any field where the newer record was created in a hurry. Then the person who needed that value re-creates the record they had, and the merge has manufactured a duplicate. Agree the rules with the data owners before any merge runs. An SAP manufacturer did exactly that before deduplicating customers, vendors and materials with audit trails.

How long does each approach take?

Native deduplication is typically days to weeks and often a configuration job. Cross-system matching in our case library ran 7 to 18 weeks across eight engagements, and the duration tracked the number of creating systems and whether survivorship had to be negotiated rather than the record count. Eight ticketing systems took 14 weeks. Eight retail channels took 14. Four claims systems took 13. Duplicate vendors and materials in one SAP estate took 9.

Does cross-system matching mean replacing the CRM?

No. In every one of the eight identity engagements in our case library the source systems were kept and the resolved identity was built above them, feeding back into the operational tools. A retailer loaded resolved profiles into a warehouse and linked them to the marketing automation platform. A casino group built a player identity above five property management systems. The CRM keeps doing what it does; it stops being asked to answer a question about records it cannot see.

What has to be true for either approach to hold?

A block at the point of entry in every system that can create a record, a no-match queue a human clears rather than a silent insert, a reversible merge with an audit trail, and a monitored duplicate rate with a threshold and a named owner. None of those are specific to the approach, and a project without them refills the stock regardless of how good the matching was. A venue group did the match, then the block, then the monitoring, which is why its rate held.

Related reading

If this is the problem you have