I will setup prometheus grafana monitoring and alerts for your servers
AWS DevOps and SRE, 7 yrs, Kubernetes CKA, Terraform, CICD
About this Gig
Finding out your server was down from a customer is the wrong way to find out. Let me set up monitoring that tells you first.
Senior DevOps / SRE, 7+ years at Zoom, Airtel, Infosys and Mindtree. I run on-call for 150+ Linux hosts and multi-cluster EKS at 99%+ uptime, using Prometheus, Grafana, Datadog and PagerDuty every day. Certified: CKA, AWS Solutions Architect, AWS SysOps, Terraform.
WHAT I SET UP
Prometheus with node exporter, cAdvisor and service exporters
Grafana dashboards that show what actually matters, not 40 useless panels
Alertmanager rules routed to Slack, email or PagerDuty
Loki for centralised logs, searchable next to your metrics
Kubernetes monitoring: pods, nodes, deployments, resource pressure
AWS CloudWatch alarms where they fit better than Prometheus
WHY IT MATTERS
Alerts that fire constantly get ignored. I tune thresholds so every alert means something is genuinely wrong and needs a human.
HOW I WORK
Step one: tell me what you run and what has hurt you before.
Step two: I confirm scope and what to alert on. Consultation is free.
Step three: working stack, tuned alerts and a runbook per alert.
Message me with your setup for a free consultation.
Tools:
Docker
•
Kubernetes
•
Amazon EKS
•
Other
Frameworks:
Terraform
•
Ansible
Cloud Provider:
Amazon Web Services
Programming language:
Bash
•
Python
Expertise:
Installation
•
Debugging
•
Configuration

