Thinklytics

CRM Data Quality · 10 min read · October 2026

Why duplicate CRM records keep returning after a cleanup

By Sean Majidi, Founder, Thinklytics

A cleanup is a one-off fix to a stock, and duplicates are a flow. Five of the six causes are design choices you can close, and the sixth is measured at 2.1% of contact fields changing every week. Which is why the rate comes back unless something blocks it at the point of entry.

You ran a deduplication project. The database was clean. Eighteen months later sales is complaining about duplicates again, and nobody can say what went wrong.

Nothing went wrong with the cleanup. It fixed a stock, and duplicates are a flow.

The six causes, and the one you cannot prevent

Five reasons the duplicates come back

A cleanup is a one-off fix to a stock. Every one of these is a flow, and a flow refills the stock.

CauseWhat it looks likeWhat stops it
No system owns the recordCRM, marketing platform and billing all create customersOne authoritative system per record type, written down
No match rule at the point of entryA rep types a company name that already exists in another formA match check on create, in the form, before save
Integrations create on no-matchA sync finds no match, so it inserts. Every run adds moreA no-match queue a human clears, never a silent insert
No survivorship ruleTwo records merge and the wrong field wins, so someone re-creates the right oneA field-level survivorship rule agreed with the data owners
Nobody watches the rateThe duplicate rate is measured once, at cleanup, and never againA monitored duplicate rate with a threshold and an owner
Natural decay of the underlying factsPeople change jobs and companies. The record is not duplicated, it is wrongScheduled re-verification, because this one cannot be prevented

Only the last row is unavoidable. The first five are design choices, which is why a cleanup without them is overtaken within a few quarters.

Source: Thinklytics engagement pattern across the identity resolution and master data engagements in the case library, including a venue group that cut duplicate fan records from 184,000 to 11,200 and then blocked new ones at the point of entry.

No system owns the record. The CRM creates customers. So does the marketing platform, the billing system and the web store. Four creators, no authority, and the same company arrives four times in four spellings.

No match rule at the point of entry. A rep types a company name that already exists in another form. Nothing checks before save.

Integrations create on no-match. A sync looks for a match, does not find one, and inserts. That is correct behaviour for a sync and catastrophic as a default, because it adds records on every run and nobody sees it happen.

No survivorship rule. Two records merge, the wrong value wins on some field, and the person who needed that value re-creates the record they had. The merge caused the duplicate.

Nobody watches the rate. The duplicate rate is measured once, during the cleanup, and then never again. By the time anyone notices, the argument is about whether the first project worked.

The underlying facts decay. People change jobs and companies. This one cannot be prevented, only re-verified on a schedule.

Five of six are design choices. That is the useful finding, because it means the recurrence is fixable rather than inevitable.

How fast the sixth one moves

How fast CRM contact fields go stale, measured weekly

Annualised from continuous weekly re-verification rather than from an annual snapshot. This is the inflow a one-time cleanup is competing against.

  • Overall record decay, compounded from 2.1% per week
  • Job title
  • Company affiliation
  • Email validity

Source: Cleanlist AI Waterfall Decay Study, Q1 2026: 5,000 anonymised CRM contacts re-verified weekly for field-level changes over 13 consecutive weeks, 7 January to 7 April 2026. Vendor-run, and it is a measured study rather than a survey. The study notes the familiar 30% a year figure traces to a 2017 HBR article and a 2018 Gartner report, both using annual snapshots.

The Cleanlist AI Waterfall Decay Study took 5,000 anonymised CRM contacts and re-verified them weekly for 13 consecutive weeks, between 7 January and 7 April 2026. It observed 2.1% of fields changing per week, compounding to roughly 67% a year. By field: job title 1.1% weekly, which annualises to 43.6%; company affiliation 0.7%, so 30.6%; email validity 0.4%, so 18.8%.

Two caveats worth stating. It is vendor-run, by a company that sells data cleaning. And it is a measured study rather than a survey, which in this case cuts in its favour: the study points out that the familiar 30% a year figure traces back to a 2017 HBR article and a 2018 Gartner report, both of which used annual snapshots rather than continuous measurement. An annual snapshot misses a contact that changed twice.

Take the direction rather than the decimal. The direction is that a meaningful share of your contact fields are wrong within a quarter, which is the inflow your one-time cleanup is competing against.

Decay is not duplication, and the distinction decides the fix. Decay means the record is wrong. Duplication means the entity exists twice. They connect because a rep who cannot find a contact under its stale details creates a new one, so decay feeds duplication from below. Re-verification addresses one. A match rule at entry addresses the other.

What it costs the people doing the selling

What sales teams say the data is costing them

Reported by sales professionals about their own working week. The top two data issues are both duplicate-shaped.

FindingFigureScope
Reps' time spent meeting with customers40% of an average workweekMore than half of their time goes to nonselling work including data entry and prospecting
Top data issues where agents are in use1 manual errors, 2 duplicate dataThen security concerns, incomplete data, corrupt data
Sales pros with agents saying data quality issues hurt their sales46%Of those whose teams have deployed agents
Sales leaders with AI saying tech silos delay or limit those initiatives51%Of sales leaders whose teams use AI
Lack of a unified customer view36% severe impact, 51% some impactFrom Salesforce State of Data and Analytics, 2025
Sales reps overwhelmed by too many tools42%And 84% of teams without an all-in-one platform plan to consolidate

Duplicate data ranks second on a list the respondents wrote themselves, above incomplete and corrupt data. It is not a hygiene footnote, it is the second thing they name.

Source: Salesforce State of Sales, 7th edition: anonymous survey of 4,050 sales professionals across 22 countries, conducted August through September 2025, all respondents third-party panelists. Vendor-run. Note the fieldwork predates the report year.

Salesforce's State of Sales, seventh edition, is an anonymous survey of 4,050 sales professionals across 22 countries, conducted August through September 2025 with third-party panelists. The fieldwork predates the report year, which is worth knowing when quoting it.

Reps reported spending 40% of an average workweek meeting with customers, and more than half their time on nonselling work including data entry and prospecting.

The interesting finding is a ranking the respondents produced themselves. Among teams with agents deployed, the top data issues were, in order: manual errors, duplicate data, security concerns, incomplete data, corrupt data. Duplicate data sits second, above incomplete and corrupt. It is not a hygiene footnote in that list, it is the second thing they name.

Alongside it, 46% of sales pros with agents said data quality issues hurt their sales, 51% of sales leaders with AI said tech silos delay or limit those initiatives, and 42% of reps said they are overwhelmed by too many tools. On a unified customer view specifically, Salesforce's State of Data and Analytics 2025 found 36% reporting a severe impact from the lack of one and 51% some impact.

Why the CRM says the database is clean

Because native CRM deduplication compares records inside that CRM, which is exactly what it claims to do.

If billing, the web store, the support desk or an acquired entity can also create a customer, the duplicates live between systems, and no amount of in-CRM scanning will see them. A clean CRM report is then a true statement about an incomplete picture, which is the most expensive kind of reassurance.

So the first count is not the duplicate rate. It is the number of systems in your estate that can create a customer record. Which approach follows from that count is covered in native CRM deduplication vs cross-system matching.

What it looks like when it holds

A regional venue division ran eight venues on eight separate ticketing systems and fan databases. Duplicate fan records across those systems were misattributing ticket sales, issuing loyalty points twice, and sending upsell campaigns to people who had already bought. The revenue team put the cost at close to $1M a year.

The work had three parts, and the third is the one usually missing. A single fan identity built by matching email, phone and purchase data across the eight databases. Then rules to block new duplicates from entering. Then a dashboard monitoring the duplicate rate in real time.

Duplicate records fell from 184,000 to 11,200 inside 800,000 records, so 23% to 1.4%. Misattributed sales and double loyalty points stopped, worth $940K a year. Upsell campaign failures fell from 144 a month to 24. See the venue group fan identity engagement.

A national retailer did the equivalent across loyalty, point of sale, e-commerce, mobile and email: eight systems, deterministic and probabilistic matching, 2.3 million duplicate records resolved into 1.4 million unique profiles in 14 weeks. A targeted campaign on the resolved data produced $4.2M of incremental revenue in three months and email open rates moved from 18 to 31 of every 100.

The AI version of the problem

An agent reads whatever customer records it can reach. Two records for one account produce two answers, two next-best-actions, and two emails, delivered with more fluency than the dashboards managed.

That is why duplicate data ranks second specifically among teams that have deployed agents, rather than among all teams. A pharmacy benefit manager had three machine learning pilots stalled on this: member match accuracy was running at 75 of every 100 records. Lifting it to 94 restarted all three and recovered $4.8M a year in misrouted claims. See the member matching engagement.

What we would do first

Two counts, and no procurement conversation until both exist.

List every system that can create a customer record. All of them, including the web store, the support desk, the billing platform, the events tool and anything that arrived with an acquisition. Most teams find the list is longer than they expected, and the length of it decides the approach.

Then measure your duplicate rate on a stated denominator. 184,000 in 800,000, not "about 60,000 merges". A rate on a denominator is checkable by someone else and comparable afterwards. A count of merges is neither.

With those two numbers you know which route you are on, and you have the before figure you will need to show the work landed. The routes are in native CRM deduplication vs cross-system matching, and the SAP-specific version is in master data deduplication in SAP.

Delivery sits in master data management for the identity layer, data cleaning and preparation for the remediation, data 360 consultant where the customer view spans channels, and vendor-neutral system integration for the no-match queues between systems. The full set of work in this area sits under our systems need people in between them.

Frequently asked questions

Why do duplicate CRM records keep coming back after a cleanup?

Because a cleanup fixes the stock and the duplicates are a flow. Six things create them: no system owns the record type, there is no match check at the point of entry, integrations insert when they find no match, there is no field-level survivorship rule so merges lose data and someone re-creates the right record, nobody monitors the duplicate rate after the project ends, and the underlying facts decay on their own. Five of those six are design choices. Close them and the rate holds. Skip them and you buy the same cleanup again.

How fast does CRM contact data actually go stale?

The Cleanlist AI Waterfall Decay Study re-verified 5,000 anonymised CRM contacts weekly for 13 consecutive weeks between 7 January and 7 April 2026 and observed 2.1% of fields changing per week, which compounds to roughly 67% a year. By field: job title 1.1% weekly, company affiliation 0.7%, email validity 0.4%. It is vendor-run and it is a measured study rather than a survey, which matters because the familiar 30% a year figure traces to a 2017 HBR article and a 2018 Gartner report that used annual snapshots.

Is decay the same thing as duplication?

No, and conflating them leads to the wrong fix. Decay means the record is wrong: the person changed jobs, the email bounces. Duplication means the same entity exists twice. Decay drives duplication indirectly, because a rep who cannot find a contact under its stale details creates a new one. Decay needs scheduled re-verification. Duplication needs a match rule at the point of entry. Buying one when you needed the other is common.

How much does this cost a sales team?

Salesforce's State of Sales, seventh edition, surveyed 4,050 sales professionals across 22 countries between August and September 2025. Reps reported spending 40% of an average workweek meeting customers, with more than half of their time on nonselling work including data entry. Among teams using agents, the top two data issues named were manual errors and duplicate data, ahead of security, incomplete and corrupt data. 46% of sales pros with agents said data quality issues hurt their sales, and 51% of sales leaders with AI said tech silos delay or limit those initiatives.

What does a duplicate rate look like when it is fixed properly?

A venue group running eight separate ticketing systems went from 184,000 duplicate fan records inside 800,000 to 11,200, so 23% down to 1.4%. The order of work is the point: match and merge across the eight databases, then rules that block new duplicates at the point of entry, then a dashboard monitoring the duplicate rate in real time. All three. That stopped $940K a year of revenue leakage and cut upsell campaign failures from 144 a month to 24.

Why does my CRM report a clean database when sales says otherwise?

Because native CRM deduplication only compares records inside that CRM. If billing, the web store, support or an acquired entity can also create a customer, the duplicates live between systems where the CRM cannot see them. Count the systems that can create a customer record. If the answer is more than one, a clean CRM report is a true statement about an incomplete picture.

Does AI make this better or worse?

Worse, faster, if the records are not resolved first. An agent reads whatever customer records it can reach, so two records for one account produce two answers and two actions. In the Salesforce data, duplicate data is the second-ranked issue specifically among teams that have deployed agents, and 36% report a lack of unified customer view as a severe impact with 51% reporting some impact. A pharmacy benefit manager had three ML pilots stalled until member match accuracy went from 75 to 94 of every 100.

Where should we start?

Two counts, no procurement. List every system that can create a customer record, including the web store, the support desk, the billing platform and anything acquired. Then measure your duplicate rate on a stated denominator, so 184,000 in 800,000 rather than a count of merges. Those two numbers tell you whether you need native deduplication or cross-system matching, and they give you the before figure you will need to prove the work landed.

The work behind this

Eight identity resolution engagements in the case library published a measured outcome, from a duplicate fan record rate cut from 23% to 1.4% across eight ticketing systems to member match accuracy lifted from 75 to 94 of every 100.

Data cleaning and labeling, 22 engagements.

Topics covered

  • duplicate CRM records
  • CRM data quality
  • duplicate customer records across systems
  • CRM data decay
  • golden record
  • identity resolution
  • RevOps data hygiene

Frequently asked questions

Why do duplicate CRM records keep coming back after a cleanup?

Because a cleanup fixes the stock and the duplicates are a flow. Six things create them: no system owns the record type, there is no match check at the point of entry, integrations insert when they find no match, there is no field-level survivorship rule so merges lose data and someone re-creates the right record, nobody monitors the duplicate rate after the project ends, and the underlying facts decay on their own. Five of those six are design choices. Close them and the rate holds. Skip them and you buy the same cleanup again.

How fast does CRM contact data actually go stale?

The Cleanlist AI Waterfall Decay Study re-verified 5,000 anonymised CRM contacts weekly for 13 consecutive weeks between 7 January and 7 April 2026 and observed 2.1% of fields changing per week, which compounds to roughly 67% a year. By field: job title 1.1% weekly, company affiliation 0.7%, email validity 0.4%. It is vendor-run and it is a measured study rather than a survey, which matters because the familiar 30% a year figure traces to a 2017 HBR article and a 2018 Gartner report that used annual snapshots.

Is decay the same thing as duplication?

No, and conflating them leads to the wrong fix. Decay means the record is wrong: the person changed jobs, the email bounces. Duplication means the same entity exists twice. Decay drives duplication indirectly, because a rep who cannot find a contact under its stale details creates a new one. Decay needs scheduled re-verification. Duplication needs a match rule at the point of entry. Buying one when you needed the other is common.

How much does this cost a sales team?

Salesforce's State of Sales, seventh edition, surveyed 4,050 sales professionals across 22 countries between August and September 2025. Reps reported spending 40% of an average workweek meeting customers, with more than half of their time on nonselling work including data entry. Among teams using agents, the top two data issues named were manual errors and duplicate data, ahead of security, incomplete and corrupt data. 46% of sales pros with agents said data quality issues hurt their sales, and 51% of sales leaders with AI said tech silos delay or limit those initiatives.

What does a duplicate rate look like when it is fixed properly?

A venue group running eight separate ticketing systems went from 184,000 duplicate fan records inside 800,000 to 11,200, so 23% down to 1.4%. The order of work is the point: match and merge across the eight databases, then rules that block new duplicates at the point of entry, then a dashboard monitoring the duplicate rate in real time. All three. That stopped $940K a year of revenue leakage and cut upsell campaign failures from 144 a month to 24.

Why does my CRM report a clean database when sales says otherwise?

Because native CRM deduplication only compares records inside that CRM. If billing, the web store, support or an acquired entity can also create a customer, the duplicates live between systems where the CRM cannot see them. Count the systems that can create a customer record. If the answer is more than one, a clean CRM report is a true statement about an incomplete picture.

Does AI make this better or worse?

Worse, faster, if the records are not resolved first. An agent reads whatever customer records it can reach, so two records for one account produce two answers and two actions. In the Salesforce data, duplicate data is the second-ranked issue specifically among teams that have deployed agents, and 36% report a lack of unified customer view as a severe impact with 51% reporting some impact. A pharmacy benefit manager had three ML pilots stalled until member match accuracy went from 75 to 94 of every 100.

Where should we start?

Two counts, no procurement. List every system that can create a customer record, including the web store, the support desk, the billing platform and anything acquired. Then measure your duplicate rate on a stated denominator, so 184,000 in 800,000 rather than a count of merges. Those two numbers tell you whether you need native deduplication or cross-system matching, and they give you the before figure you will need to prove the work landed.

Related reading

If this is the problem you have