remotely.living

Principal AI Infrastructure & Accelerator Architect

EPAM Systems · Remote - United States · 2026-09-29

Apply for this job

Job description

Are you a visionary technical leader passionate about defining the future of AI infrastructure and squeezing every drop of performance out of advanced hardware accelerators at scale? We are seeking a Principal Architect to lead, shape, and execute our technical strategy for AI performance, optimization, and hardware-software co-design.

In this elite, highly visible role, you will define the architectural vision for both AI training and serving infrastructure, delivering massive industry-wide impact. You will spearhead our Center of Excellence (CoE), scaling our practice and guiding the technical roadmap across next-generation Tensor Processing Units (TPUs), Graphics Processing Unit (GPU) fleets, state-of-the-art ML models, and advanced compiler toolchains.

Your architectural decisions will directly enable cutting-edge AI research and large-scale production deployments across Google Cloud, major enterprise customers, and the broader open-source ecosystem. If you thrive on solving intractable performance bottlenecks and redefining what is physically and computationally possible in AI infrastructure, this is your platform.

Req# 1067369299

Responsibilities

- Define and drive the multi-year technical roadmap for high-performance AI kernels, custom operations, and hardware-software co-design targeting TPU and GPU architectures

- Scale and mentor a world-class technical practice, establishing architectural governance, engineering standards, and best practices across the organization

- Act as the principal technical liaison partnering with ML researchers, core framework architects (JAX, PyTorch), and compiler engineering teams (XLA, MLIR) to eliminate systemic bottlenecks and shape future hardware/software requirements

- Architect foundational infrastructure—including enterprise-grade benchmarking suites, automated autotuning frameworks, regression analysis pipelines, and comprehensive documentation—empowering the global developer community

- Anticipate industry shifts by tracking advancements in hardware architectures, emerging model topologies, and compiler innovations to unlock step-changes in AI training and inference efficiency

Requirements

- Bachelor's degree in Computer Science, Electrical Engineering, or equivalent practical experience (Master's or Ph. D. preferred)

- 15+ years of software engineering experience, with 8+ years focused on distributed systems, AI infrastructure, or high-performance computing (HPC) architecture

- 7+ years of experience designing and developing complex software systems in C++ or Python

- 5+ years of experience leading the architecture, design, and delivery of large-scale software products, frameworks, or developer ecosystems from inception to production

- Proven track record of architecting performance-critical systems at the kernel level, bridging hardware accelerators and high-level software frameworks

Nice to have

- Deep expertise in optimizing TPU/GPU execution, leveraging low-level kernel languages/abstractions such as Pallas, Mosaic, Triton, or CUDA

- Comprehensive knowledge of modern ML frameworks (JAX, PyTorch), attention mechanisms, Mixture of Experts (MoEs), model quantization, and low-precision arithmetic

- Advanced understanding of modern accelerator architectures, including heterogeneous compute, complex memory hierarchies, data movement optimization, and multi-node scale-out fabrics

- Deep familiarity with compiler principles, code generation, and modern toolchains such as MLIR, OpenXLA, and LLVM

- Demonstrated leadership in building and scaling developer infrastructure, widely adopted Open-Source Software (OSS) libraries, and extensible high-performance APIs

- Exceptional strategic communication and stakeholder management skills, with a history of influencing cross-functional engineering teams, researchers, and executive leadership