Key takeaways
  • Test against your own historical book, where you already know the answers
  • Measure time to availability from date of death, because that is the metric that predicts prevented loss
  • Run a known deceased control set so you can measure recall, not just volume
  • Decide the go or no go thresholds before you see any results

Set the question before you set the schedule

A pilot that asks whether the data is good will produce an argument. A pilot that asks four specific numerical questions will produce a decision.

Write these down first, agree them with whoever signs the contract, and set the pass thresholds while you are still neutral.

  • What proportion of our known deceased population does the file find? That is recall.
  • Of the matches it returns, what proportion survive manual review? That is precision.
  • How many days after date of death did each record become available to us?
  • How many matches would have been actionable at the decision point we care about, rather than after it?

Week one, build the sample

Pull a historical population where the outcome is already known. Twelve to twenty four months of applications or accounts is usually enough, and history is essential because you need to be able to check the answer.

Then build a control set. Take accounts you already know were deceased, ideally confirmed through an estate process or a family notification, and hold them aside. Without a control set you can measure how much the file returns but not how much it misses, and the miss rate is the number that matters.

Agree the file format and the minimum identity elements at the same time. Name and date of birth is workable. Adding last known address materially improves match confidence and reduces the review burden.

Week two, run it and instrument it

Submit the sample, capture the raw results and keep everything. You want the confidence score on every match, the matched elements, the date of death and the last known address.

Then compute the availability metric properly. For each match, take the date of death and compare it to the date the record became queryable. Report the median and the ninetieth percentile rather than an average, because averages hide the tail and the tail is what determines whether the file prevents anything.

If you are evaluating two providers, run the identical sample through both and compare on the same four questions. Different samples produce different numbers and no useful comparison.

Week three, adjudicate

Take a random sample of the matches and review them properly against whatever internal evidence you have. This is where precision gets measured, and it is worth doing by hand rather than by rule.

While you are in there, split the results by confidence band. You are looking for the band where precision is high enough to act automatically and the band below it where a human should look. Those two numbers become your production thresholds, so the pilot pays for itself twice.

Also check the control set. Every known deceased account that the file did not return is worth understanding. Sometimes it is an identity data quality problem on your side, which is useful to know regardless.

Week four, write the decision

The output is a short document, not a slide deck. State the four numbers, compare them to the thresholds set in week one, and recommend go or no go.

Then quantify. Take the matches that were available before your decision point, multiply by the average exposure on that product, and you have the avoided loss figure. Take the matches on accounts already in outreach and you have the complaint and reputational figure. Those two lines are what a budget holder needs.

If it is a go, the production design is mostly already written. You have thresholds from week three, a file format from week one, and a measured latency profile that tells you what cadence to run at.

  • Recall, precision, median and ninetieth percentile availability, actionable match rate
  • Comparison to the thresholds agreed before results were seen
  • Avoided loss estimate and outreach complaint estimate
  • Recommended production thresholds and run cadence

Keeping the pilot cheap

The reason this fits in a month is that a per record model with no charge for non matches makes the sample size question disappear. You are not paying to test a hundred thousand records against a file, you are paying only for the small number that come back confirmed.

That removes the usual procurement bottleneck, which is committing to a minimum before you have any evidence. Test at real volume, pay for what is actually found, and let the numbers decide.

Common questions

How large should the pilot sample be?

Large enough that the expected number of deceased matches is comfortably into the hundreds, so precision and recall are measurable rather than anecdotal. For most consumer books that means tens of thousands of records over a twelve to twenty four month window. Use history rather than live traffic so you can verify the answers.

What is a good deceased match rate?

There is no single benchmark, because it depends on the age profile of your book and the age of the records in it. That is exactly why the control set matters. Recall against known deceased accounts is comparable across providers and across time. Raw match volume is not.

Can we pilot without sending real customer data?

You can start with a synthetic or masked sample to validate the integration and the file format, but the four metrics that decide the question need real identities matched against real records. Run the security review first, then use real data under the agreed controls.

See what this looks like in your portfolio

Upload a sample file or call the API and get confirmed deceased matches with date of death, age and last known address. You are only charged for records we confirm.