I will clean, transform, and build etl pipelines
Data Whisperer
About this Gig
I help businesses and researchers turn messy, inconsistent, or raw data into clean, reliable, and analysis-ready datasets with full pipelines built in Python or R. I also offer consulting for teams that need advice or a second opinion rather than (or in addition to) hands-on build work.
Hands-on Engineering Work
- Clean and validate messy datasets (missing values, corrupted entries, inconsistent formats, duplicate records)
- Build ETL pipelines to automate data extraction, transformation, and loading from CSV, Excel, APIs, or databases
- Normalize inconsistent column names, data types, and structures from multiple sources (data integration)
- Design data models and schemas for clean, scalable storage
- Set up and manage data warehouses (Snowflake, BigQuery, Redshift)
- Build transformation pipelines with dbt and orchestrate recurring runs with Apache Airflow
- Store and manage data in cloud storage (Amazon S3)
- Data governance: validation rules, quality checks, and documentation standards
- Data migration between systems, formats, or platforms
- Handle time-series data: gaps, sensor dropout, seasonal patterns, interpolation
- Basic forecasting and trend analysis on cleaned time-series data
FAQ
What file formats do you accept?
CSV, Excel (.xlsx), JSON, and most SQL database exports. If your data comes from an API, share the docs and I'll confirm feasibility before you order.
Do you write the code in Python or R?
Both - I'll recommend whichever fits your existing tools and team, or use whichever you specify.
Can you handle recurring/automated data feeds, not just a one-time file?
Yes, this is what the ETL pipeline packages are for — the script is built to run again on new data, not just clean what you send once.
What if my dataset is larger than the package limits?
No problem, message me before ordering with your row count and I'll send a custom quote.

