I will build AWS etl pipelines with airflow, spark and python
Senior Data Engineer
About this Gig
Senior Data Engineer 10+ years building production ETL for Fortune 500 clients (Expedia, Ford, HP).
I build reliable, production-ready AWS data pipelines with Python, PySpark, Airflow, and Spark not throwaway scripts. Real engineering: partitioning, broadcast joins, orchestration, and testing.
Track record:
Cut a 64-hour HiveQL job to under 3 hours (95.6% faster)
Migrated 500 TB from Hadoop to AWS, fully validated
Built and run my own multi-tenant SaaS platform end-to-end
What you get: clean, documented, deployable code pipelines your team can maintain, not a black box.
Based in Mexico (US timezone overlap) real-time collaboration during your business hours. Fluent English.
Not sure which package fits? Message me first and I'll scope it with you.
Tools & Platforms:
AWS Glue DataBrew
•
Other
My Portfolio
FAQ
Q: Do you work with data sources outside AWS?
A: Yes — I ingest from databases (PostgreSQL, MySQL), APIs, files, and streaming sources like Kafka, landing data into S3, Redshift, or your warehouse of choice.
Q: Will I get the source code and documentation?
A: Always. You receive clean, version-controlled code plus documentation so your team can run and maintain everything independently.
Q: I'm not sure which package I need. Can you help?
A: Absolutely. Message me with your sources, destination, and volume, and I'll recommend the right package or build a custom offer.
Q: Can you fix or optimize an existing pipeline instead of building new?
A: Yes. I audit broken or slow pipelines and optimize them — my HiveQL work cut one job from 64 hours to under 3.

