remotely.living

Senior AI Engineer with Spark, AWS Services

EPAM Systems · Remote - Georgia / Armenia / Kazakhstan / Kyrgyzstan / Uzbekistan · 2026-09-29

Apply for this job

Job description

We are seeking a Senior AI Engineer with Spark and AWS Services expertise to join the RBQM Production Pod within the program. You will build and maintain data pipelines that power AI/GenAI applications for Risk-Based Quality Management in clinical trials. This role focuses on RAG document ingestion, vector indexing, and building data APIs for AI applications.

Responsibilities

- Design and build RAG document ingestion pipelines (chunking, embedding, vector indexing) for clinical trial quality data

- Build and manage vector databases (AWS OpenSearch) for RAG-powered AI workflows

- Develop batch and streaming ETL/ELT pipelines from scratch for unstructured clinical data (PDF, DOCX, clinical reports)

- Build and expose data APIs for AI application consumption

- Optimize chunking strategies, embedding generation, and retrieval performance for RAG architectures

- Manage data quality, lineage, and governance for AI/ML data pipelines

- Deploy and maintain AWS data infrastructure (S3, Lambda, Glue, Athena, Step Functions, DynamoDB)

- Collaborate with Data Scientists and Backend Developers in an integrated pod team

Requirements

- 5+ years of hands-on data engineering experience at scale

- Expertise in RAG document ingestion pipelines (chunking, embedding, vector indexing)

- Proficiency in AWS OpenSearch as a vector database for RAG workflows

- Advanced proficiency in Python, including SQL and Spark SQL

- Skills in unstructured data transformation (PDF, DOCX) for RAG/LLM applications

- Familiarity with AWS Services: S3, Lambda, Glue, Athena, Bedrock, Step Functions, API Gateway, CloudWatch, DynamoDB

- Knowledge of containerization with Docker

- Capability to build custom pipelines from scratch, beyond configuring out-of-the-box services

- Proficiency in English at a B2+ level

Nice to have

- Background in pharmaceutical or life sciences domain

- Familiarity with Snowflake, Pinecone (vector DB alternative)

- Knowledge of SageMaker processing jobs

- Skills in CI/CD tools (Jenkins, Git/Bitbucket) and infrastructure tools (CDK or Terraform)

- Understanding of clinical data standards (CDISC, ADaM, SDTM)