I will build full stack open source production data pipeline with cicd automation

Pakistan

I speak Urdu, English

Data Engineer

Data Engineer specializing in building end-to-end, self-hosted data lakehouses on Kubernetes using open-source technologies. I designed and operated a full production-grade data platform implementing ...
About this Gig

Are you looking for flexible data engineering support whether it's a 

quick fix or a full pipeline build? I offer three engagement levels:


Hourly Support Quick debugging, small Spark/ETL tasks, or consultation

Part-Time Engagement Focused 20-hour pipeline or dbt model development  

Full Project Build Complete end-to-end lakehouse setup (50+ hours)


I am a Data Engineer with hands-on experience building a full self-hosted 

data lakehouse on Kubernetes processing ~5 million rows using Apache Spark, 

Apache Iceberg, Trino, dbt-Trino, and Google BigQuery.


What I deliver:

Apache Spark ETL pipelines ingestion, cleansing, transformation

Data quality checks using PyDeequ

dbt-Trino Gold layer models Star Schema, Data Vault, OBT, Marts

Dagster orchestration and CI/CD automation

Full documentation (Standard & Premium tiers)


Message me before ordering so we can discuss your specific requirements.

Warehouse Platform:

Snowflake

BigQuery

Databricks

Project Type:

New Build

My Portfolio