v
vanshbhardwa261

Vansh Bhardwaj

@vanshbhardwa261

DevOps Consultant : Cloud Infrastructure, CI CD, Kubernetes, Terraform, AWS

India
English, Hindi
About me
DevOps Consultant DevOps Engineer with 4+ years of experience helping startups and businesses build reliable DevOps workflows and scalable cloud infrastructure. Skilled in AWS, GCP, Terraform, Ansible, CI/CD, Jenkins, GitHub Actions, Argo CD, Docker, Kubernetes, Linux, Bash, ELK, Datadog, Prometheus, and Grafana. I deliver clean, production-ready solutions with clear communication, quick response times, reliable support, and a strong focus on long-term scalability and reliability, ensuring your systems stay healthy, secure, and cost-efficient.... Read more

Skills

v
vanshbhardwa261
Vansh Bhardwaj
Offline • 

Portfolio

Work experience

Hitachi

Senior Software Engineer

Hitachi • Full-time

Jan 2025 - Present1 yr 7 mos

• Owned the availability, reliability, and scalability of production systems across AWS and GCP, supporting large-scale microservices. • Led production releases using Helm and modernized CI/CD with GitHub Actions and GitOps (Argo CD), reducing deployment failures by 35% and improving release reliability. • Managed production Kubernetes platforms across AWS EKS and Google Kubernetes Engine (GKE), ensuring high availability of containerized workloads. • Troubleshooted Kubernetes production issues including CrashLoopBackOff, image pull failures, pod/node failures, networking, DNS, storage, and resource constraints, while implementing Horizontal Pod Autoscaling (HPA). • Provisioned and scaled cloud infrastructure using Terraform, improving consistency and reducing manual provisioning effort by 60%. • Automated infrastructure and operational workflows using Ansible, Rundeck, Bash, and Python, reducing operational toil. • Built observability solutions with Prometheus and Grafana to monitor infrastructure and application health. • Defined and monitored SLIs, SLOs, and SLAs using the SRE Golden Signals to improve service reliability and proactive detection of performance issues. • Participated in 24×7 on-call, responding to production alerts including pipeline delays, Kafka consumer lag, MongoDB, GCP VM, quota, load balancer, disk, and container failures. • Led incident response for critical outages, driving RCA, cross-functional recovery, and preventive improvements. • Delivered cloud cost optimization initiatives by decommissioning unused environments/VMs and implementing resource lifecycle policies, saving over USD $1 Million / annually. • Led AI-driven operations by developing Atlassian Rovo AI Agents and an AI-powered Rovo On-Call Assistant to streamline operations and accelerate incident response.

Taboola

Software Engineer

Taboola • Full-time

Jan 2021 - Dec 20243 yrs 11 mos

DevOps / SRE