I will build and optimize large scale apache spark data pipelines

India

I speak Telugu, English, Hindi

Senior Data Engineer

Senior Data Engineer with 8+ years of experience building and optimizing large-scale distributed data pipelines. Experienced in Apache Spark, Scala, PySpark, Hive, Hadoop, SQL, and AWS, with hands-on...
About this Gig

Expert Apache Spark Optimization & Data Pipeline Architecture

Are your Spark jobs failing with Out-of-Memory (OOM) errors, timing out on broadcasts, or costing too much in cloud compute?

I am a Senior Data Engineer with over 8 years of enterprise experience building, troubleshooting, and optimizing large-scale distributed data pipelines. I specialize in processing multi-terabyte workloads and resolving complex Big Data bottlenecks using Apache Spark, Scala, PySpark, Hadoop, and Hive.

What I Can Do For You:

  • Spark Performance Tuning: Resolve data skew, optimize expensive joins, and fine-tune memory allocation to stop job failures and drastically reduce execution time.
  • ETL Pipeline Development: Design and build robust, scalable, and fault-tolerant data ingestion pipelines from scratch using best practices for partitioning and storage.
  • Production Debugging: Deep-dive into lineage, shuffles, and cluster resource utilization to fix silent performance killers in your existing code.
  • Core Tech Stack: Apache Spark, Scala, Python/PySpark, Hadoop, Hive, SQL, and cloud platforms (AWS).

Whether you need a quick diagnostic fix for a failing query or an end-to-end enterprise data architecture, I d