I will clean, deduplicate and normalize your messy dataset in excel

Venezuela

I speak Spanish, English

Backend Engineer for Payments, Ledgers and Verified Data

I build the part of the software where money cannot be wrong. Since 2019 I have been the backend engineer behind a cross-border payments platform: remittances, wallets, multi-provider payouts and a do...
About this Gig

The dangerous part of cleaning a dataset isn't the mess you can see. It's the 412 near-duplicates you can't.


Two rows for the same customer with the name spelled differently. Three product entries that are one product. Dates in six formats where two of them are ambiguous, so half your timeline is silently wrong.


Send me the file. You get it back clean, plus a change report: rows in, rows out, what merged with what, and every row I touched. Nothing disappears without appearing in that report.


On near-duplicates: I flag them for your review by default instead of merging them automatically. You know your data. Silently merging two rows that looked similar is how a clean dataset becomes a wrong one, and you'd never find out.


What I fix: exact and fuzzy duplicates, inconsistent dates and number formats, mixed currencies, capitalization and spacing, split or merged columns, broken encodings, malformed URLs and codes, missing values, and merging several files into one consistent dataset.


CSV, XLSX, JSON, TSV. No scraping involved - this is your data, processed.