Thinklytics

Express Scripts · Healthcare · St. Louis, MO · 14 weeks

3 stalled ML pilots restarted

Three machine learning pilots had stalled for more than a year because member identity data was inconsistent. We implemented unified entity resolution using Databricks and rebuilt the feature engineering pipeline from the ground up. This cleared the data issues and allowed all three pilots to move into production.

Challenge

Express Scripts had three different member ID systems that never aligned. Their ML models only matched 75 of every 100 records correctly, causing about $4.8M a year in misrouted claims and extra manual work. Three AI projects were stuck for over a year because internal teams couldn’t prioritize fixing the data issues.

Approach

We created an entity resolution layer on Databricks to match member identities from three source systems. We used probabilistic matching first, then applied deterministic logic when needed. After that, we rebuilt the feature engineering pipelines for each ML model and validated them against a six-month holdout dataset. Finally, we delivered the updated pipelines to the internal ML team.

Outcome

We improved member match accuracy to 94 of every 100 records on the validation set. Within four weeks of delivery, the client restarted all three ML pilots. Reducing misrouting is expected to save $4.8 million annually. The internal ML team now fully owns the pipeline, supported by complete documentation.

Customer IDs were a total mess, and it completely threw off our data quality.

We ran into a major headache with the ML pilots because the training data was all over the place. Duplicate member records were everywhere, but each came with a different ID and no way to connect them. So, the models got fed mixed signals and, unsurprisingly, their accuracy took a nosedive.

We started out with probabilistic matching, but when we needed tighter accuracy, we switched over to deterministic. Simple as that.

Here’s the deal. We started with a two-step matching process. First, we threw in probabilistic matching using name, birth date, and address to catch the fuzzy matches. Then, we got stricter with a deterministic check using SSN and member ID to establish exact matches. When we tested it, we hit 94 of every 100 records. Not too shabby.

Handing it off to the in-house ML team so they can take it from here

We jotted down all the data tweaks, matching rules, and those tricky edge cases in a runbook. Nothing left to guess. Then, just two weeks after the handoff, the internal ML team was already running their first production inference by week 16. It was great to see things move that fast.

Results

  • $4.8M Annual claims recovery
  • 75 to 94 of 100 Member match accuracy
  • 3 ML pilots restarted
  • 14 wks Delivery timeline

We had $2 million stuck in three machine learning pilots that weren’t moving forward because our member data was all over the place. Thinklytics cleaned up the data in 14 weeks, and now all three pilots are up and running.

VP of Data Science, Express Scripts

Thinklytics

Data and AI consulting for Fortune 500s, health systems, and growth-stage companies. Clean data, governed metrics, analytics ready for AI.

Austin, TX · United States

[email protected]