p
piyushkumar_03

Piyush K

@piyushkumar_03

Data EngineerII

India
English, Hindi
About me
I am a Data Engineer II with extensive experience in building and optimizing large-scale data platforms. I specialize in managing production Trino clusters, implementing Apache Iceberg Lakehouse architectures, and automating ETL pipelines using tools like dbt, Airflow, and Spark. I have a proven track record of driving significant cost reductions and improving data accessibility through self-serve tooling and robust governance.... Read more

Skills

p
piyushkumar_03
Piyush K
Offline • 

See my services

Data ETLs
I will build scalable etl pipelines and spark lakehouses

Work experience

PW

Data Engineer

PW • Full-time

Jun 2023 - Present • 3 yrs 4 mos

Trino Infrastructure & Cost Optimization: Managed 4 specialized production Trino clusters processing 450–500 TB and 3,000–4,000 CPU-hours daily; drove a 30–35% cloud cost reduction using KEDA autoscaling while sustaining tight query latencies (p90 0.8–1.1s, p99 17–30s). Trino Ops & Resource Guardrails: Built Trino operational tooling for automated bad-query termination, query-level performance profiling, and proactive alerts to prevent rogue workloads from degrading shared cluster stability. Self-Serve Transformation Platform: Engineered an internal dbt-style UI on Trino and Apache Iceberg enabling analysts to author, validate, schedule, and monitor gold-layer models independently, slashing pipeline turnaround time from days to minutes. Secondary Transformation & CI/CD Governance: Owned end-to-end dbt secondary transformation delivery (Trino/Iceberg) across 2,850 models, serving as primary MR reviewer/approver to maintain strict production standards and deployment safety. Developer Platform Engineering (Phoenix): Co-developed and operated Phoenix (internal Virtual Data Engineer), a self-serve portal for table diagnostics, Airflow pipeline insights, and query debugging that eliminated operational bottlenecks across the DE team. Framework Automation & Scale: Standardized repetitive ETL workflows into automated frameworks, including an automated sync for 1,000+ Bronze-to-Silver tables, consolidation of 100+ BigQuery event streams into a unified Gold table, and automated SLA/freshness reporting. Lakehouse Governance & Security: Architected Lakehouse PII security using Apache Ranger for fine-grained column masking and access controls; integrated an approval-driven, time-bound access grant workflow via Phoenix to prevent PII exposure without breaking analytics workflows. Real-Time CDC Analytics: Integrated Kafka CDC event streams into ClickHouse for PW’s Vishwas Diwas event, powering real-time dashboards for sub-second tracking of live signups, order volume, and re