I will build a data cleaning and validation pipeline in python


About this gig
Most data problems are not analysis problems. They are cleaning problems nobody wants to look at.
A pipeline that silently "fixes" bad data is worse than one that refuses it. It turns 15/01/2026 into the wrong month, drops 3% of rows, and produces a report that looks clean and is wrong.
WHAT YOU GET
- A Python pipeline that reads your files, cleans what it can prove, and quarantines the rest with a written reason.
- Every rejected row goes to its own file with its source and line number. Nothing is dropped quietly.
- Validation rules you can read and change: types, ranges, required fields, date formats, duplicates.
- An auditable report plus exit codes, so it fails loudly instead of writing a wrong file.
- Tests, so you can change the rules without breaking them.
- A handover doc: what was built, how to run it, what was verified, and the limitations.
HOW I WORK
I look at your data before quoting. If a column is ambiguous I will ask rather than guess.
STACK
Python, standard library only. Nothing to install.
Tell me what "clean" means for your data and I will say if it is a fit. Samples on GitHub - message for the link.
Get to know natsuki
all
- FromChina
- Member sinceSep 2026
- Avg. response time1 hour
Languages
Chinese, English
FAQ
What file formats can you handle?
CSV by default. Excel, TSV and fixed-width text are fine too. Your source file is never modified - the pipeline reads it and writes new files.
What happens to rows you cannot clean?
They go to a separate rejected file with the source row number and the reason. Nothing is dropped silently, and the report counts them so you can see exactly what was excluded.
Will you change my original data?
No. The pipeline only reads your input. Every output is a new file, so you can always diff the result against the source.
Can it run on a schedule?
Yes. It exits with a non-zero code on a fatal problem, so cron or Task Scheduler can tell success from failure without anyone reading the log.
My columns mean something specific to my business.
That is exactly why I ask before quoting. Send a sample and the validation rules get written to your definitions rather than guessed from the column names.

