I will build your etl data pipeline using python sql airflow and AWS
Data Engineer
About this Gig
Struggling with messy, scattered data that's hard to trust? I build reliable ETL pipelines that turn raw data into clean, query-ready information the way production data teams do it.
I work with Python, SQL, Apache Airflow, PySpark, and AWS to move your data from source to a structured warehouse, with validation along the way so bad data gets caught early.
What I can help with:
ETL/ELT pipeline design and development
Apache Airflow orchestration (scheduled DAGs)
Data ingestion from APIs, CSVs, and databases
Data transformation with Python/PySpark
Data warehouse modeling (fact/dimension tables, PostgreSQL)
Data quality validation and automated checks
I've built end-to-end pipelines handling 100,000+ records across multiple sources, structured into clean, analytics-ready warehouses.
I care about clean code and clear communication you'll know what's happening at every stage, not just get a black box at the end.
Got a different data source or need? Message me first happy to check the fit before you order.
Destination Platform:
Amazon Redshift
•
PostgreSQL
•
MySQL
•
Amazon S3
Tools & Platforms:
AWS Glue DataBrew
•
Other
