remotely.living

Senior Site Reliability Engineer (Linux, Python, Prometeus)

Link Group · Remote · 2026-09-25

Apply for this job

Job description

About the Role

We are seeking a Senior Site Reliability Engineer (SRE) to drive reliability, performance, and operational excellence across our enterprise virtualization platforms and high-performance AI compute infrastructure. In this role, you will act as a technical leader and mentor, automating away operational toil, elevating system observability, and designing resilient, self-healing infrastructure at scale.

This position is ideal for an experienced Linux Systems Specialist / DevOps Engineer who thrives on architecting robust systems, optimizing containerized workloads, and guiding engineering teams toward modern SRE best practices.

Key Responsibilities

- Infrastructure Automation & Toil Reduction: Design, build, and maintain automation tooling and scripts (using Python, Golang, and IaC ) to streamline software deployments, safeguard release processes, and minimize manual intervention.

- Observability & Performance Optimization: Define Service Level Objectives (SLOs) and enhance system telemetry using Prometheus, Grafana, and distributed tracing to accelerate error detection and improve the reliability of our core virtualization platform.

- AI Infrastructure & Capacity Management: Drive capacity planning, workload scheduling, and auto-scaling strategies tailored for high-demand AI compute infrastructure.

- Incident Management & Reliability: Participate in on-call rotations, serving as a technical guide during critical service-impacting incidents to speed up recovery and root-cause resolution.

- Technical Leadership & Mentorship: Elevate the engineering team by coaching developers on SRE principles, promoting a culture of reliability, and fostering continuous learning.

Required Qualifications & Technical Expertise

- Core Experience: Deep, expert-level background in Linux/Unix Systems Administration, SRE, or DevOps supporting large-scale, distributed production environments.

- Containerization & Orchestration: Proven expertise in Kubernetes and orchestrating large-scale containerized systems.

- Software Engineering & IaC: Advanced proficiency in at least one modern programming language ( Python or Golang ) alongside configuration and IaC tools ( Terraform, SaltStack, or Ansible ).

- Observability & SLOs: Hands-on experience establishing SLOs/SLIs and implementing enterprise observability stacks ( Prometheus, Grafana, tracing frameworks).

- System Architecture & Design: Practical experience architecting software systems and cloud/bare-metal infrastructure at scale.

- SRE Evangelism: Strong accountability for system health, with a collaborative mindset to introduce SRE practices to teams new to the discipline.