CRM Data Quality · 10 min read · October 2026
What to ask a consultant before a CRM data cleanup project
By Sean Majidi, Founder, Thinklytics
Eight questions, four of them about reversibility, because a merge is the hardest data operation to undo. Plus the eight acceptance criteria written as rates rather than counts, and the access question that should be in the contract rather than in an email.
A proposal to clean up customer data is in front of you. The technical approach probably looks fine. What decides whether this works is narrower than the methodology, and most of it is about what happens when a merge is wrong.
Eight questions
Questions to ask before you sign
The answers tell you more than the methodology section. Four of these are about reversibility, because merges are the hard part to undo.
- Which systems in my estate can create a customer record, and are all of them in scope?. Ask them to list the systems back to you. If the list is shorter than yours, the scope is wrong.
- What duplicate rate will you report, on what denominator, before and after?. A rate on a stated denominator is checkable. A count of merges is not.
- What is the match threshold, and who decides it?. It should be you, because the threshold is a trade between missed duplicates and wrongly merged customers.
- Show me the field-level survivorship rules from a previous engagement. If they cannot produce one, survivorship will default to most-recent-wins and you will find out which fields that ruins.
- How do I undo a merge, and have you tested the rollback on my data?. Before the first production run, not after. This is the question that separates a data project from an incident.
- What blocks the next duplicate at the point of entry?. If there is no answer, you are buying a cleanup and not a fix, and the decay rate will refill it.
- Who owns the duplicate rate after you leave, and at what threshold do they act?. A named person and a number. Not a committee and not a dashboard nobody opens.
- What access do your people need, for how long, and how is it revoked?. Cleanup work needs broad read and write access to customer data. Scope it, time-box it, and write the revocation date down.
A carrier compared a vendor's $2.8M remediation quote against the vendor's own 10-category data quality rubric, found eight failing dimensions, and had six fixed by automated pipelines for $340K. Asking which dimensions actually fail is how a quote gets right-sized.
Source: Thinklytics case library, published engagement scopes, approaches and outcomes.
Which systems in my estate can create a customer record, and are all of them in scope? Ask them to list the systems back to you. If their list is shorter than yours, the scope is wrong, and the duplicates will survive in whatever was left off. Which approach that count implies is in native CRM deduplication vs cross-system matching.
What duplicate rate will you report, on what denominator, before and after? A rate on a stated denominator is checkable by someone who was not in the project. A count of merges is not.
What is the match threshold, and who decides it? It should be you. The threshold trades missed duplicates against wrongly merged customers, and those two errors are not equally bad.
Show me field-level survivorship rules from a previous engagement. If they cannot produce one, survivorship will default to most-recent-wins and you will discover which fields that ruins.
How do I undo a merge, and have you tested the rollback on my data? Before the first production run. This question separates a data project from an incident.
What blocks the next duplicate at the point of entry? If there is no answer, you are buying a cleanup and not a fix.
Who owns the duplicate rate after you leave, and at what threshold do they act? A named person and a number.
What access do your people need, for how long, and how is it revoked? Covered below, because it is the question most often settled in an email.
Why reversibility gets four of the eight
A merge is the hardest common data operation to undo.
Most data work is recoverable. A bad transformation is re-run. A bad load is rolled back. A merge combines two records into one, and if the match was wrong you have destroyed information rather than duplicated it: two real customers now share one history, one of them has the other's orders or claims attached, and the customer usually notices before you do.
Across integrated systems it is worse, because the merge propagates. The CRM merge fires a sync, the marketing platform merges its own records, the support desk reassigns tickets, and by the time someone reports it the undo path crosses four systems.
So the rollback is designed, documented and tested on production-shaped data before the first merge runs. Demonstrated, not described. An SAP manufacturer deduplicated customers, vendors and materials with audit trails and turned every cleansing decision into a documented transformation rule, which is what made the work repeatable and the trial conversion pass on the next attempt.
What belongs in the scope
What belongs in a cleanup scope, and what gets left out
The right-hand column is where these projects fail. Each omission is invisible at signature and obvious two quarters later.
| In scope | Why | Left out of most proposals |
|---|---|---|
| Every system that can create a customer record, listed | The scope boundary. If billing creates customers and is not listed, the duplicates survive | A scope defined as the CRM only |
| A measured duplicate rate before and after, on the same denominator | 184,000 of 800,000 to 11,200 of 800,000 is checkable. A count merged is not | A count of merges, with no rate and no denominator |
| Match rules, deterministic then probabilistic, with the threshold stated | The threshold is a business decision about false merges, not a technical default | The threshold, left to the tool |
| Field-level survivorship rules, agreed with the data owners | Decides which value wins on merge. Get it wrong and someone re-creates the right record | Survivorship, assumed to be most-recent-wins |
| A block at the point of entry, plus a no-match queue for integrations | This is what stops the stock refilling. Without it you buy a cleanup again next year | Prevention, which is the whole difference |
| A monitored duplicate rate with a threshold and a named owner | Turns a project into a controlled rate | Monitoring, so nobody notices the drift |
| A reversible merge with an audit trail, and a tested rollback | Merges touch records the business runs on. You need the undo before the first run | Rollback, until the day it is needed |
A venue group did the cleanup, then blocked new duplicates at entry, then built a real-time duplicate rate dashboard. All three, which is why the rate held.
Source: Thinklytics case library, published engagement scopes and outcomes.
Seven items. The last three are what make the result last, and they are the three most often absent.
Every system that can create a customer record, listed by name. A measured duplicate rate before and after on the same denominator. Match rules, deterministic then probabilistic, with the threshold stated rather than left to the tool. Field-level survivorship rules agreed with the data owners. A block at the point of entry plus a no-match queue for every integration. A monitored duplicate rate with a threshold and a named owner. And a reversible merge with an audit trail and a tested rollback.
The venue group that cut duplicate fan records from 184,000 to 11,200 inside 800,000 did three things in order: matched and merged across eight ticketing databases, then set rules to block new duplicates from entering, then built a dashboard monitoring the duplicate rate in real time. The third step is the one cut for budget most often, and it is why that rate held.
Acceptance criteria, as rates
Acceptance criteria for a cleanup, written as numbers
Tested on the last day by someone who was not in the project. Most of these are rates rather than counts, deliberately.
- Duplicate rate at or below a stated threshold, on the agreed denominator. A venue group went from 184,000 duplicates in 800,000 records to 11,200, so 23% to 1.4%. State the target before the work.
- Match accuracy at or above a stated figure, measured on a held-back sample. Member match accuracy from 75 to 94 of every 100 at a pharmacy benefit manager. Hold the sample back so the figure means something.
- Zero wrongly merged records in the reviewed sample, with the review documented. A false merge costs more than a missed duplicate, because it destroys information rather than repeating it.
- A blocked-on-create rule live in every system that can create a record. Named systems in the statement of work. This is the criterion that makes the result last.
- A no-match queue for every integration, with a named person clearing it. Silent inserts on no-match are the single largest source of re-accumulation.
- A tested rollback, exercised on production-shaped data before go-live. Demonstrated, not described.
- A monitored duplicate rate with an alert threshold and a named owner. The project ends. The rate does not.
- A downstream effect measured, where one is measurable. Campaign failures from 144 to 24 a month. Email open rate from 18 to 31 of 100. $4.8M a year in misrouted claims. These are what the cleanup was for.
The last one is the criterion that justifies the spend. A clean database with no measured downstream effect is a hygiene project; a 23% to 1.4% duplicate rate that stopped $940K of revenue leakage is an investment.
Source: Thinklytics case library, published delivery outcomes and acceptance measures per engagement.
Eight, written before the work starts and tested on the last day by someone who was not in the project. Mostly rates rather than counts, deliberately.
Two are worth expanding.
Match accuracy measured on a held-back sample. Hold the sample back before the matching runs, or the figure is a measurement of the thing that produced it. A pharmacy benefit manager moved member match accuracy from 75 to 94 of every 100 records, which is the shape of claim you want: a rate, on a sample, before and after.
A measured downstream effect. This is the criterion that justifies the spend. Campaign failures from 144 a month to 24. Email open rates from 18 to 31 of every 100. $4.8M a year of misrouted claims recovered. $940K a year of revenue leakage stopped. A clean database with no measured downstream effect is a hygiene project. A duplicate rate that moved from 23% to 1.4% and stopped $940K of leakage is an investment, and the difference is entirely in whether anyone agreed to measure it.
Right-sizing a quote
Worth a section, because the largest number in this category is often the least examined.
A workers compensation carrier had spent $1.2M on an AI underwriting platform that required a data quality score of 80 across 10 categories. Its policy data scored 61, and the platform vendor quoted $2.8M to remediate.
Reading the vendor's own rubric showed eight failing dimensions. Six were addressable with automated pipelines: standardising address formats, resolving duplicate policy IDs, filling missing coverage start dates, normalising coverage code mappings. The remaining two were legacy records that needed manual review workflows. The score reached 94 by week 16, the platform went live three weeks later against a projected $8.4M of annual underwriting value, and the total cost was $340K. See the policy data quality engagement.
The transferable part is not the saving. It is that the quote was sized against a rubric nobody had read, and reading it changed the scope from ten categories to six pipelines and two review queues. Ask which dimensions actually fail, and ask to see the rubric.
The access question
Cleanup work needs broad read and write access to customer data across several systems, which is more than most data projects require.
Four things belong in the contract rather than in an email. The named individuals. The systems and objects in scope. The time box. And the revocation date.
Then sandbox-first for the match runs, production access only for the merge window, and an audit log of what each account touched. The CRM data access and rollback checklist is the worksheet we hand to whoever has to approve this, because it is usually a different person from the one who approved the budget and they have different questions.
The two warning signs
No prevention step. If the scope covers matching and merging but nothing blocks the next duplicate at the point of entry and no integration gets a no-match queue, you are buying a cleanup rather than a fix. The measured decay rate will refill it: 2.1% of contact fields change per week in the Cleanlist study of 5,000 contacts re-verified weekly over 13 weeks to April 2026, and a rep who cannot find a contact under stale details creates a new one.
A count of merges instead of a rate on a denominator. That is the metric you choose when you do not intend to be measured a year later.
Why integrators under-price this work is covered in why system integrators underinvest in data cleansing.
What we would do first
Before reading the proposal again, produce two artifacts yourself. The list of systems that can create a customer record, with an owner next to each. And your current duplicate rate on a stated denominator.
Then check the proposal against the seven scope items and the eight acceptance criteria. Most proposals cover the matching and omit the prevention and the rollback, and those are the two that decide whether you are doing this once.
Delivery sits in master data management for the identity layer, data cleaning and preparation for the remediation, data 360 consultant where the customer view spans channels, vendor-neutral system integration for the no-match queues, sales and CRM AI automation where agents will read the resolved records, and SAP data quality and governance on SAP estates. The full set of work in this area sits under our systems need people in between them.
Frequently asked questions
What should I ask a consultant before a CRM data cleanup?
Eight questions. Which systems in my estate can create a customer record, and are all of them in scope. What duplicate rate will you report, on what denominator, before and after. What is the match threshold and who decides it. Show me field-level survivorship rules from a previous engagement. How do I undo a merge, and have you tested the rollback on my data. What blocks the next duplicate at the point of entry. Who owns the duplicate rate after you leave, and at what threshold do they act. And what access do your people need, for how long, and how is it revoked.
Why do four of the questions concern reversibility?
Because a merge is the hardest common data operation to undo. It combines two records into one, and if the match was wrong you have destroyed information rather than duplicated it: two real customers now share one history, and the customer usually notices before you do. So you want the rollback path designed, documented and tested on production-shaped data before the first merge runs, not improvised afterwards. Everything else in a cleanup is recoverable from a backup. Merges across integrated systems frequently are not.
What is the right way to state the duplicate rate?
As a rate on a named denominator, measured the same way before and after. A venue group went from 184,000 duplicate fan records inside 800,000 to 11,200, so 23% to 1.4%. That is checkable by someone who was not in the project and comparable a year later. A count of merges performed is neither, which is why it is the number most proposals offer. Ask for the denominator in the statement of work.
What belongs in the scope?
Every system that can create a customer record, listed by name. A measured duplicate rate before and after on the same denominator. Match rules, deterministic then probabilistic, with the threshold stated. Field-level survivorship rules agreed with the data owners. A block at the point of entry plus a no-match queue for every integration. A monitored duplicate rate with a threshold and a named owner. And a reversible merge with an audit trail and a tested rollback. Seven items, and the last three are the ones that make the result last.
What acceptance criteria should go in the contract?
Eight, mostly stated as rates. Duplicate rate at or below a target on the agreed denominator. Match accuracy at or above a stated figure measured on a held-back sample. Zero wrongly merged records in the reviewed sample, with the review documented. A blocked-on-create rule live in every named system. A no-match queue per integration with a named clearer. A tested rollback exercised on production-shaped data. A monitored rate with an alert threshold and an owner. And a measured downstream effect, which is the one that justifies the spend.
How do I right-size a quote that looks too high?
Find the rubric the number was built on and check which dimensions actually fail. A workers compensation carrier was quoted $2.8M by its AI underwriting vendor to lift a policy data quality score from 61 to the required 80 across 10 categories. Reading the vendor's own rubric showed eight failing dimensions, six of which automated pipelines could address: address standardisation, duplicate policy IDs, missing coverage start dates, coverage code normalisation. The score reached 94 in 16 weeks for $340K.
What access will the team need, and how should it be controlled?
Broad read and write access to customer data across several systems, which is more than most data projects need and worth treating accordingly. Put four things in the contract rather than an email: the named individuals, the systems and objects in scope, the time box, and the revocation date. Sandbox-first for the match runs, production access only for the merge window, and an audit log of what each account touched. This is the conversation the CRM data access and rollback checklist exists for.
What is the warning sign that a proposal will not hold?
No prevention step. If the scope covers matching and merging but nothing blocks the next duplicate at the point of entry and no integration gets a no-match queue, you are buying a cleanup rather than a fix, and the measured decay rate will refill it. The second warning sign is a count of merges instead of a rate on a denominator, because that is the metric you choose when you do not intend to be measured a year later.
The work behind this
Twenty-two engagements in the case library carry data cleaning and labeling. Eight of them published a measured identity outcome with a before and after figure, including a duplicate rate cut from 23% to 1.4% and a policy data quality score lifted from 61 to 94 for $340K against a $2.8M vendor quote.
Data cleaning and labeling, 22 engagements.
Topics covered
- CRM data cleanup project
- deduplication project scope
- CRM cleanup acceptance criteria
- merge rollback
- data cleanup consultant questions
- duplicate rate threshold
- CRM data access
Frequently asked questions
What should I ask a consultant before a CRM data cleanup?
Eight questions. Which systems in my estate can create a customer record, and are all of them in scope. What duplicate rate will you report, on what denominator, before and after. What is the match threshold and who decides it. Show me field-level survivorship rules from a previous engagement. How do I undo a merge, and have you tested the rollback on my data. What blocks the next duplicate at the point of entry. Who owns the duplicate rate after you leave, and at what threshold do they act. And what access do your people need, for how long, and how is it revoked.
Why do four of the questions concern reversibility?
Because a merge is the hardest common data operation to undo. It combines two records into one, and if the match was wrong you have destroyed information rather than duplicated it: two real customers now share one history, and the customer usually notices before you do. So you want the rollback path designed, documented and tested on production-shaped data before the first merge runs, not improvised afterwards. Everything else in a cleanup is recoverable from a backup. Merges across integrated systems frequently are not.
What is the right way to state the duplicate rate?
As a rate on a named denominator, measured the same way before and after. A venue group went from 184,000 duplicate fan records inside 800,000 to 11,200, so 23% to 1.4%. That is checkable by someone who was not in the project and comparable a year later. A count of merges performed is neither, which is why it is the number most proposals offer. Ask for the denominator in the statement of work.
What belongs in the scope?
Every system that can create a customer record, listed by name. A measured duplicate rate before and after on the same denominator. Match rules, deterministic then probabilistic, with the threshold stated. Field-level survivorship rules agreed with the data owners. A block at the point of entry plus a no-match queue for every integration. A monitored duplicate rate with a threshold and a named owner. And a reversible merge with an audit trail and a tested rollback. Seven items, and the last three are the ones that make the result last.
What acceptance criteria should go in the contract?
Eight, mostly stated as rates. Duplicate rate at or below a target on the agreed denominator. Match accuracy at or above a stated figure measured on a held-back sample. Zero wrongly merged records in the reviewed sample, with the review documented. A blocked-on-create rule live in every named system. A no-match queue per integration with a named clearer. A tested rollback exercised on production-shaped data. A monitored rate with an alert threshold and an owner. And a measured downstream effect, which is the one that justifies the spend.
How do I right-size a quote that looks too high?
Find the rubric the number was built on and check which dimensions actually fail. A workers compensation carrier was quoted $2.8M by its AI underwriting vendor to lift a policy data quality score from 61 to the required 80 across 10 categories. Reading the vendor's own rubric showed eight failing dimensions, six of which automated pipelines could address: address standardisation, duplicate policy IDs, missing coverage start dates, coverage code normalisation. The score reached 94 in 16 weeks for $340K.
What access will the team need, and how should it be controlled?
Broad read and write access to customer data across several systems, which is more than most data projects need and worth treating accordingly. Put four things in the contract rather than an email: the named individuals, the systems and objects in scope, the time box, and the revocation date. Sandbox-first for the match runs, production access only for the merge window, and an audit log of what each account touched. This is the conversation the CRM data access and rollback checklist exists for.
What is the warning sign that a proposal will not hold?
No prevention step. If the scope covers matching and merging but nothing blocks the next duplicate at the point of entry and no integration gets a no-match queue, you are buying a cleanup rather than a fix, and the measured decay rate will refill it. The second warning sign is a count of merges instead of a rate on a denominator, because that is the metric you choose when you do not intend to be measured a year later.
Related reading
If this is the problem you have
- Our systems need people in between them, resolved by 13 services.
- CRM Data Access and Rollback Checklist, the worksheet for whoever has to approve the spend.
- The 30 day Corporate Drag and Risk Diagnostic, findings yours either way.