Jobiglo

No results.

Senior Site Reliability Engineer (SRE)

EPAM Systems

Senior 🇬🇧 English
Grafana Prometheus Loki Tempo OpenTelemetry AWS Kubernetes (EKS) Python Bash Go Terraform CI/CD pipelines k6 JMeter Locust Datadog

Job description

About the role

We are seeking a Senior Site Reliability Engineer to lead the observability and reliability initiatives for our platform. The role focuses on building a Grafana‑based monitoring stack, defining SLIs/SLOs, reducing alert noise, and improving release safety on AWS/EKS.

Key responsibilities

  • Own the observability charter: build monitoring, alerting, synthetic checks, dashboards, and runbooks.
  • Define meaningful SLIs/SLOs and reduce alert noise to improve signal quality.
  • Design and optimise release pipelines with progressive delivery, health gates, and automated rollback mechanisms.
  • Apply performance‑engineering practices such as load testing, capacity analysis, and latency profiling.
  • Automate operational toil through scripting and infrastructure‑as‑code.
  • Accelerate SRE maturity by applying AIOps capabilities for detection and diagnosis.
  • Lead incident‑response practices, including on‑call readiness and blameless post‑mortems.
  • Collaborate with DevOps, Cloud, product engineering teams and Tech Leads to drive reliability improvements.

Required profile

  • 3+ years of experience in Site Reliability Engineering, DevOps, or platform engineering supporting customer‑facing systems.
  • Strong knowledge of observability tools (Grafana, Prometheus, Loki, Tempo, OpenTelemetry) covering metrics, logs, traces, and events.
  • Deep understanding of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning, and blameless post‑incident reviews.
  • Experience operating workloads on Kubernetes (preferably EKS) and AWS, with ability to debug across application, container, and infrastructure layers.
  • Proficiency in Python, Bash or Go, and exposure to Terraform and CI/CD pipelines.
  • Background in incident management, including triage, escalation, communication, post‑mortems, and on‑call rotations.
  • Proactive ownership mindset and ability to identify problems from telemetry before they are reported.
  • Strong English communication skills (B2+).

Required skills

  • Grafana
  • Prometheus
  • Loki
  • Tempo
  • OpenTelemetry
  • AWS
  • Kubernetes (EKS)
  • Python
  • Bash
  • Go
  • Terraform
  • CI/CD pipelines
  • k6
  • JMeter
  • Locust
  • AIOps concepts
  • Datadog (nice‑to‑have)

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec EPAM Systems.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches for Кыргызстан.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 2 апта мурун

Expires 1 ай ичинде

20 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

EPAM Systems