I will build etl pipelines for data engineering using apache airflow
Junior Data Scientist, ML Engineer, Python, ETL Pipelines
About this Gig
Is your data sitting in APIs or files with no automated way to collect and store it? I build production-ready ETL pipelines and data engineering workflows using Apache Airflow 3, Python, and SQL that extract, transform, and load your data on a schedule, fully automated, with zero manual work.
What I deliver:
- Automated pipeline with independent extract, transform, and load tasks
- Apache Airflow 3 with TaskFlow API and daily scheduling
- Multi-container Docker stack for clean, reproducible deployment
- PostgreSQL or MySQL database with structured, queryable records
- Full source code delivered via GitHub
Why me? I have a peer-reviewed IEEE conference publication, dual DataCamp certifications (Certified Data Scientist and Certified Associate Data Scientist), and a research internship with a UK-based AI lab. My ETL data pipeline runs live in production, accumulating 365+ structured records per year with zero manual intervention.
I work with: REST APIs and file-based data sources, loading data into SQL databases such as PostgreSQL or MySQL.
Note: Please message me before ordering to discuss your data source and requirements.
Destination Platform:
PostgreSQL
•
MySQL
Tools & Platforms:
Other
My Portfolio
FAQ
What data sources can you connect to?
Currently REST APIs. If you have a different source such as CSV files or a database, message me first and we can discuss feasibility.
Do you handle full data engineering, or just basic ETL scripts?
Full data engineering: pipeline architecture, Apache Airflow orchestration, Docker containerization, and production-ready code, not a one-off script.
Do I need Apache Airflow already installed?
No. I will set up the pipeline environment for you, including Docker configuration if needed.
Will the pipeline run automatically without me doing anything?
Yes. The Standard and Premium packages include fully scheduled automation using Apache Airflow that runs on your defined schedule without any manual triggering.
Will I receive the source code?
Yes, all packages include the full Python source code and DAG files.
Can you work with my existing database?
Yes, as long as you can provide connection credentials securely. I recommend discussing this before ordering.
Does the pipeline check for bad or invalid data?
Yes. Incoming data is validated before it loads, checking types, required fields, and clearly invalid values, so a malformed record gets caught rather than silently landing in your database.
What happens if a step in the pipeline fails?
Each task retries a few times automatically before failing, and you can be alerted by email if a run ultimately fails, so problems get caught immediately rather than discovered days later.

