remotely.living

Senior Site Reliability Engineer (SRE)

EPAM Systems · Remote - Kazakhstan / Kyrgyzstan / Uzbekistan · 2026-09-29

Apply for this job

Job description

We are looking for a Senior Site Reliability Engineer to work hands-on with a Grafana-based observability stack and AWS/Kubernetes (EKS), defining SLIs/SLOs, reducing alert noise, building actionable dashboards, strengthening incident response, and improving release safety through progressive delivery and automated deployment analysis.

Responsibilities

- Own the observability charter for the platform: build monitoring, alerting, synthetic checks, dashboards, and runbooks

- Define meaningful SLIs/SLOs and reduce alert noise to improve signal quality

- Design and optimize release pipelines with progressive delivery, health gates, and automated rollback mechanisms

- Apply a performance engineering mindset through load testing, capacity analysis, and latency profiling

- Automate operational toil through scripting and infrastructure-as-code

- Accelerate SRE maturity by applying AIOps capabilities to improve detection, diagnosis, and reduce manual effort

- Lead incident response practices including on-call readiness and blameless post-mortems

- Collaborate with DevOps, Cloud teams, product engineering teams, and Tech Leads to drive reliability improvements

Requirements

- 3+ years of experience in Site Reliability Engineering, DevOps, or platform/production engineering supporting customer-facing systems

- Expertise in observability tools such as Grafana, Prometheus, and log/trace aggregation (Loki, Tempo, OpenTelemetry) covering metrics, logs, traces, and events

- Knowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviews

- Experience operating workloads on Kubernetes (ideally EKS) and AWS, with the ability to debug issues across application, container, and infrastructure layers

- Proficiency in Python, Bash, or Go, with exposure to infrastructure-as-code (Terraform) and CI/CD pipelines

- Background in incident management including triage, escalation, communication, post-mortems, and on-call processes/rotations

- A proactive ownership mindset with the ability to identify problems from telemetry before they're reported and follow through with engineering teams

- Strong communication skills to turn noisy signals into crisp findings, runbooks, and recommendations

- English Level: B2+ (Upper-Intermediate) or higher

Nice to have

- Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactively

- Skills in performance engineering, including load/stress testing (k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecks

- Familiarity with leveraging AIOps capabilities to advance SRE maturity and drive innovation

- Experience with Datadog or similar enterprise observability platforms

- Background in evangelizing best practices and setting standards across engineering teams

- Exposure to programmatic advertising or adtech platforms