I will set up observability and alerting
SRE and DevOps Specialist in Kubernetes, Terraform and Cloud
About this Gig
Need reliable visibility into your systems, not just another dashboard? I will set up, improve, or troubleshoot a bounded observability solution for an agreed application, service, workload, namespace, or telemetry pipeline.
A full Kubernetes cluster is included only when its size, workloads, telemetry sources, and required integrations fit the selected package and are agreed before ordering.
Depending on the package and agreed scope, I can help with:
- Metrics, logs, and traces collection
- OpenTelemetry, OTLP, and Grafana Alloy configuration
- Grafana, Prometheus, Loki, and Alertmanager integration
- Dashboards and actionable alerts
- SLI and SLO definition
- Missing telemetry, noisy alerts, and pipeline troubleshooting
- Validation, documentation, and handover
I am a Staff Site Reliability Engineer with 12+ years of DevOps and SRE experience. My work includes production observability, application instrumentation, telemetry pipelines, incident response, and reliability reviews.
Each package has defined limits for environments, telemetry sources, dashboards, alerts, and SLOs. Multi-environment, multi-cluster, production-critical, or large existing platforms require discussion before ordering.
Tools:
Docker
•
GitLab
•
Jenkins
•
GitHub
•
BitBucket
•
Kubernetes
•
Amazon EKS
Frameworks:
Terraform
•
Ansible
Programming language:
Bash
•
Python
Expertise:
Installation
•
Development
•
Configuration
FAQ
What observability tools do you support?
My primary experience includes Grafana, Grafana Cloud, Grafana Alloy, Fluentbit/D, Promtail Prometheus, Loki,Mimir, Elasticsearch, Alertmanager, OpenTelemetry, OTLP pipelines, and Kubernetes observability. Share your current stack before ordering so I can confirm compatibility and scope.
Can you improve an existing observability setup?
Yes. I can review, troubleshoot, and improve existing telemetry collection, dashboards, alert rules, OpenTelemetry pipelines, and Grafana configurations. The scope depends on the number of applications, workloads, telemetry sources, environments, dashboards, alerts, and integrations involved.
What counts as one telemetry source?
One telemetry source means one agreed application, service, workload, collector pipeline, or compatible system sending a defined set of metrics, logs, or traces. Multiple applications, clusters, environments, accounts, or independent pipelines require additional scope.
Does a package include an entire Kubernetes cluster?
Not automatically. Packages are limited by the number and complexity of workloads, telemetry sources, dashboards, alerts, integrations, and SLOs. A small cluster may fit after review, but multi-namespace, multi-team, or large clusters require a custom offer.
What does the Premium package cover?
Premium covers one agreed system scope within the listed limits for telemetry collection, dashboards, alert rules, and SLIs or SLOs. It does not automatically cover every application, workload, namespace, cluster, account, or environment in an organization.
Can you work with production systems?
Yes, after reviewing the environment and agreeing on access, backups, validation, rollback, and the change window. Production-critical work may require a custom offer and must be explicitly authorized before any change is made.
Will you need access to my systems?
It depends on the work. I may need repository access, config files, architecture details, samples, or time-limited access to a staging/observability platform. Access should be individual, least-privilege, limited to the agreed scope. Don't send passwords, private keys, or secret in ordinary messages
Are application-code changes included?
Small, bounded instrumentation changes may be included when agreed before ordering. Application features, unrelated defects, extensive refactoring, and unsupported application stacks are outside scope and require separate assessment.
What information should I provide before ordering?
Please provide your application or platform type, current stack, num of workloads/services, environment, telemetry already available, desired dashboards or alerts, known problems, access constraints, and expected outcome. Contact me before ordering for production,multi-cluster or complex environment
Do you provide ongoing monitoring or on-call support?
No. This Gig covers the agreed setup, validation, documentation, and limited post-delivery support. It does not include 24/7 monitoring, emergency incident response, or ongoing operational ownership.

