I will build data pipelines that

India

I speak English

AI and ML Engineer for LLM Fine Tuning RAG and Data Science

AI/ML developer with hands-on production experience in LLM fine-tuning (QLoRA/PEFT), RAG systems (ChromaDB, MongoDB Atlas vector search, agentic memory), and data engineering (PySpark, Airflow, Delta ...
About this Gig

I build production-style ETL pipelines using PySpark for transformation, Airflow for orchestration, and Delta Lake for storage the same stack I use in my day-to-day data engineering work.

What's covered:

  • Bronze-to-Silver (raw-to-cleaned) transformation logic in PySpark
  • Airflow DAG design with proper scheduling and dependency handling
  • Delta Lake table setup with versioning
  • Optional: Trino/Metabase dashboard on top of the cleaned data
  • Fully in-house, self-hosted pipeline architecture no dependency on third-party managed ETL platforms (Fivetran, Airbyte Cloud, etc.) if you'd rather own the full stack

I'll ask about your data volume and source format before starting pipeline design changes a lot between a 10k-row CSV and a streaming source, and I'll scope honestly rather than reuse a generic template.

Destination Platform:

Amazon Redshift

Databricks Lakehouse

Tools & Platforms:

AWS Glue DataBrew

Google Cloud Dataflow