remotely.living

Cloud Engineering Manager

EPAM Systems · Remote - Ukraine · 2026-09-29

Apply for this job

Job description

We are looking for a skilled Cloud Engineering Manager to drive the strategic direction of cloud infrastructure and DevOps practices within our organization.

This position requires a blend of leadership, technical expertise, and operational excellence to oversee high-performing teams and deliver scalable, reliable, and resilient systems for mission-critical applications.

Responsibilities

- Oversee the design and implementation of infrastructure and product monitoring systems

- Define and utilize SLI/SLO standards to strengthen reliability and system performance

- Conduct root cause analysis to identify and address system inefficiencies

- Facilitate postmortem analyses and drills to optimize incident response strategies

- Evaluate product performance, scalability, and reliability to ensure operational excellence

- Automate operational tasks to optimize workflows and improve productivity

- Deploy CI/CD pipelines and champion modern DevOps practices

- Manage cloud infrastructure and configuration through Infrastructure-as-Code tools

- Collaborate with cross-functional teams to align cloud strategies with business goals

- Support the growth of engineers through mentoring and professional development initiatives

- Plan and execute staffing strategies to build a scalable and efficient engineering team

Requirements

- Minimum of 7 years of experience in cloud engineering, DevOps, or SRE

- Expertise in scripting languages such as Python, Go, Bash, or PowerShell

- Proficiency in observability tools including Prometheus, Grafana, DataDog, and ELK

- Background in cloud infrastructure management tools like Terraform and cloud-specific CLI tools (gcloud, az, aws)

- Skills in configuration management tools such as Ansible

- Knowledge of CI/CD platforms including Jenkins (Groovy SDK, Jenkinsfile), GitLab-CI, or Azure DevOps

- Expertise in containerization technologies such as Docker and Kubernetes

- Strong ability to analyze incidents and develop strategies for improvement

- Capability to design scalable, cloud-native solutions in alignment with business objectives

Nice to have

- Familiarity with hybrid cloud environments and multi-cloud architectures

- Background in integrating machine learning workloads into cloud ecosystems

- Showcase of working with observability stacks tailored for microservices architecture

- Understanding of serverless cloud platforms and associated workflows