remotely.living

Lead Data & AI Platform Engineer

EPAM Systems · Remote - Australia · 2026-09-29

Apply for this job

Job description

We are seeking a Lead Data & AI Platform Engineer to design, build, and operate the platforms that support data engineering and AI delivery across a range of client engagements. You will lead platform automation, deployment, governance, observability, and reliability for data pipelines, machine learning systems, LLM applications, and AI agents.

This is a hands-on technical leadership role. You will guide architectural decisions, establish reusable engineering practices, and work closely with data and AI teams to bring solutions into production.

Responsibilities

- Lead the design and implementation of secure, scalable cloud infrastructure for data and AI workloads

- Build reusable infrastructure and CI/CD patterns for data platforms, applications, models, and AI agents

- Establish operational practices for AI solutions, including deployment, versioning, evaluation, monitoring, and rollback

- Implement observability across pipelines and AI applications, including logs, metrics, traces, alerts, quality measures, and cost monitoring

- Improve platform reliability through automation, incident response, performance tuning, and capacity planning

- Define standards for cloud security, identity and access management, secrets, networking, and governance

- Support teams deploying LLM applications and agentic systems, including their model integrations, tool calls, and external APIs

- Use AI-assisted tools to improve infrastructure development, troubleshooting, and operational workflows, while validating their output

- Partner with architects, engineers, and stakeholders to make platform decisions and explain technical trade-offs

- Mentor engineers and contribute to technical standards, reusable templates, and platform roadmaps

Requirements

- Strong hands-on experience in platform engineering, DevOps, site reliability engineering, or MLOps, with experience leading technical delivery

- Experience with at least one major cloud platform: Azure, AWS, or Google Cloud

- Strong infrastructure as code skills, such as Terraform

- Experience designing CI/CD pipelines and automated deployment processes

- Scripting or programming experience, preferably in Python

- Experience operating data platforms, distributed workloads, or cloud-native applications

- Strong understanding of monitoring, logging, alerting, incident response, and production troubleshooting

- Knowledge of cloud networking, identity and access management, secrets management, and security practices

- Practical understanding of AI/ML platform operations, including model deployment, versioning, monitoring, and lifecycle management

- Understanding of the operational needs of LLM applications and AI agents, including tracing, evaluation, latency, reliability, and cost

- Ability to lead technical decisions, mentor engineers, and collaborate directly with client and delivery teams

Nice to have

- Experience with Databricks, Snowflake, Microsoft Fabric, or similar data and AI platforms

- Experience with MLflow, model serving platforms, or model registries

- Experience deploying and observing RAG applications or agentic systems

- Familiarity with LLM evaluation, guardrails, and monitoring for response quality

- Experience with Kubernetes, containers, API gateways, or service orchestration

- Experience applying AI to operations, such as alert correlation, incident triage, or root cause analysis

- Experience with GitHub Actions, Azure DevOps, GitLab CI, or similar tools

- Experience with Databricks Asset Bundles, Unity Catalog, or equivalent capabilities

- FinOps experience, including workload optimisation and cloud cost attribution

- Consulting experience, including platform assessments, solution design, and technical estimation