I will match, merge and deduplicate your messy datasets
PhD Candidate in Astrophysics, Statistical Consulting
About this Gig
Two lists of the same people, products or places, no shared ID: names spelled differently, typos, missing fields. Excel's VLOOKUP gives up immediately.
I do probabilistic record linkage: instead of "match / no match", every candidate pair gets a score for how likely it is to be the same entity, based on how much agreement on each field is actually worth. Agreeing on a rare surname is strong evidence; agreeing on a common one is weak. The method knows the difference.
Why this matters to you: you get to choose the threshold. High confidence for automatic merging, and a separate list of ambiguous pairs for a human to look at, instead of silently merging two different customers, or keeping the same one twice.
I use this on astronomical catalogs, where merging the wrong two objects invalidates the science. Your data deserves the same care.
What I deliver: your merged dataset with a confidence score on every match; a separate "needs review" list for ambiguous cases; a quality report on how many matched, how many didn't, and why; and duplicates found within each file, not just across them.
Send a small sample first if you'd like to see it work before ordering.
Technology:
Python
FAQ
What file formats do you accept?
CSV or Excel. If your data is sensitive, send a sample with fake or anonymized values first, the method works the same on real data later.
I don't know which fields to match on, can you help?
Yes, tell me what fields you have (name, address, phone, etc.) and I'll tell you which combination will work best for your data.
Should I prioritize precision or recall?
Whichever matters more to you: never merging two different records, or never missing a true match. You can't fully have both, so tell me which matters more and I'll tune the threshold for that.

