remotely.living

Lead Operational Intelligence Engineer

EPAM Systems · Remote - Ukraine · 2026-09-29

Apply for this job

Job description

Our team is seeking a dynamic, highly experienced professional to fill the role of Lead Operational Intelligence Engineer.

This role involves taking charge of developing, maintaining, and enhancing our cloud-based Elastic & Observability Platform. The successful candidate will spearhead strategic initiatives, mentor a top-performing technical team, and maintain platform reliability while promoting innovation and self-service capabilities for platform users. On-call rotation duties for monitoring platform health and functionality are also part of this position.

Responsibilities

- Ensure observability and search platforms exceed business SLAs in terms of availability, functionality, performance, and security

- Deliver technical leadership when complex incidents arise and ensure prompt escalation of resolutions during on-call shifts

- Create and maintain thorough platform documentation, standard operating procedures, and knowledge-sharing materials

- Work with cross-functional teams, stakeholders, and vendors to manage operational needs, advance strategic initiatives, and handle installations, troubleshooting, and upgrades

- Drive improvements to platform features and self-service tools, including advanced Elastic Synthetics and automated chargeback processes

- Design and build proofs-of-concept to advance platform innovation, such as AI-driven observability, sophisticated data processing models, or migration to Kubernetes-based platforms

- Guide the construction, deployment, and upkeep of Elastic clusters using Infrastructure-as-Code tools such as Terraform and Ansible, and coach team members on best practices

- Manage platform lifecycle tasks, such as component upgrades, capacity planning, cost optimization, and adapting to new compliance requirements

- Regularly evaluate and optimize ELK stack performance, covering ingestion, indexing, and query tuning for large-scale environments

- Build and improve alerting and incident management processes by integrating advanced monitoring tools like Kibana Rules, Watchers, and PagerDuty

- Manage the ingestion, enrichment, backup, and restoration of large-scale platform data, optimizing data workflows along the way

- Direct and plan major operational events, including SSL certificate rotations, cluster migrations, and scalability optimization efforts

Requirements

- At least 5 years of experience in Operational Intelligence, demonstrating leadership and technical skill in managing large-scale observability platforms

- Proven ability to design and oversee Elastic clusters within complex, multi-cloud environments

- Comprehensive knowledge of Elastic Stack components, including advanced setups of Elasticsearch, Kibana, and Logstash

- High-level skills in Infrastructure-as-Code tools such as Terraform and Ansible, with flexibility to work with tools like Jenkins CI or GitOps frameworks

- Strong Python scripting abilities for automating processes, handling data, and expanding platform interoperability

- Solid grasp of incident management frameworks and workflows using tools such as PagerDuty, Uptrends, and other enterprise monitoring platforms

- Demonstrated success in diagnosing and resolving intricate platform issues within strict SLA timeframes

- Strong skills in managing and scaling fault-tolerant platforms, ensuring performance, security, and compliance across large distributed systems

- Proven track record of mentoring team members, managing priorities, and serving as a liaison between technical and non-technical groups

- Strong English communication skills (B2+ level), both written and verbal, with an emphasis on technical communication

Nice to have

- Skills in Groovy scripting or advanced Linux administration experience to streamline platform operations

- History of enhancing observability workflows through custom integrations in tools such as Uptrends, PagerDuty, or Elastic

- Practical experience configuring advanced Elastic Synthetics for reliable monitoring and custom synthetic testing

- Background in leading strategic initiatives like AI-driven modernization, cloud-native migrations, or cost-saving observability improvements