
Santhosh S
Certified DevSecOps, GitOps Engineer, Cloud Computing Expert!
Skills

See my services


Work experience
Rapyder Cloud Solutions
Full-time • 1 yr 4 mos
Senior DevOps Engineer
Jun 2026 - Present • 3 mos
Lead the design and implementation of the Safexpress disaster recovery environment on AWS — multi-region failover, Velero-based Kubernetes backup and restore, database replication and documented recovery runbooks — achieving an RTO under 30 minutes with zero data loss across quarterly DR drills. • Own the production Kubernetes (EKS) platform lifecycle: HA cluster design, control-plane and worker-node patching, version upgrades, and root-cause resolution of pod, node, networking, storage and ingress incidents. • Serve as the primary technical point of contact for the client’s cloud infrastructure, running architecture reviews, DR readiness assessments and technical workshops on DevOps and Kubernetes best practices. • Standardize delivery through centralized Helm charts and reusable Terraform, Terragrunt and CloudFormation modules across 20+ microservices, backed by Python and Bash automation, cutting environment provisioning from days to under 2 hours. • Maintain and extend end-to-end CI/CD pipelines (GitHub Actions, AWS CodePipeline, Jenkins, ArgoCD) with automated quality gates, sustaining an 80% reduction in deployment errors and release cycle time under 20 minutes. • Lead the AWS Control Tower governance rollout across 8+ AWS accounts and drive enterprise observability with Prometheus, Grafana, ELK, SigNoz and New Relic.
AWS DevOps Engineer
May 2025 - Jun 2026 • 1 yr 1 mo
Migrated 30+ on-premises and legacy workloads to AWS for banking, fintech and AI clients, holding cutover downtime to under 1% of the migration window. • Managed and modernized production Kubernetes (EKS) clusters across multiple client accounts — HA setup, node patching, version upgrades, and troubleshooting of pod, node, networking, storage and ingress issues. • Supported regulated enterprise application environments for BFSI clients, resolving operational tickets within contractual SLAs and cutting mean time to detect incidents by around 40% through improved monitoring and alerting. • Built end-to-end CI/CD pipelines from the ground up (GitHub Actions, AWS CodePipeline, Jenkins, ArgoCD) with Git-based version control, reducing deployment errors by 80% and release cycle time to under 20 minutes. • Implemented Velero-based disaster recovery with multi-region failover, achieving RTO under 30 minutes with zero data loss across quarterly DR drills. • Configured enterprise observability (Prometheus, Grafana, ELK, SigNoz, New Relic) and contributed to the AWS Control Tower governance rollout across 8+ accounts. • Delivered customer-facing proof-of-concept en
Cloud Platform Engineer
E2E Networks Limited • Full-time
May 2024 - Apr 2025 • 11 mos
Migrated 20+ client on-premises and multi-cloud workloads to E2E Cloud with minimal downtime, managing the Kubernetes platform lifecycle end to end. • Operated and troubleshot NVIDIA GPU server fleets and enterprise application environments across on-premises and cloud, sustaining 95%+ utilization for AI/ML workloads. • Ran Kubernetes platform operations — patching, golden image updates, CNI and ingress troubleshooting, and etcd backup and recovery for kubeadm on-premises clusters. • Built CI/CD pipelines with Jenkins, Git and Ansible, cutting manual deployment effort by roughly 60% across compute and storage environments