I will match, merge and deduplicate your messy datasets

Mexico

I speak Spanish, English, French

PhD Candidate in Astrophysics, Statistical Consulting

PhD candidate in astrophysics. I spend my days doing statistics on messy, incomplete, correlated data — galaxy catalogs where nothing is clean and every error bar has to be defended. I bring that sta...
About this Gig

Two lists of the same people, products or places, no shared ID: names spelled differently, typos, missing fields. Excel's VLOOKUP gives up immediately.


I do probabilistic record linkage: instead of "match / no match", every candidate pair gets a score for how likely it is to be the same entity, based on how much agreement on each field is actually worth. Agreeing on a rare surname is strong evidence; agreeing on a common one is weak. The method knows the difference.


Why this matters to you: you get to choose the threshold. High confidence for automatic merging, and a separate list of ambiguous pairs for a human to look at, instead of silently merging two different customers, or keeping the same one twice.


I use this on astronomical catalogs, where merging the wrong two objects invalidates the science. Your data deserves the same care.


What I deliver: your merged dataset with a confidence score on every match; a separate "needs review" list for ambiguous cases; a quality report on how many matched, how many didn't, and why; and duplicates found within each file, not just across them.


Send a small sample first if you'd like to see it work before ordering.

Technology:

Python

Expertise:

Data manipulation

Data validation

ETL

Normalization