remotely.living

Agent Evaluation Engineer — Build-Time Framework & Deployment Gates

EPAM Systems · Remote - Portugal · 2026-09-29

Apply for this job

Job description

We're looking for an Agent Evaluation Engineer — Build-Time Framework & Deployment Gates to join our team in Portugal in a fully remote working mode. In this role, you will design and maintain an evaluation framework for AI agents, ensuring quality and compliance through automated tests and CI/CD deployment gates. You will develop multi-layer evaluation suites that blend deterministic checks with LLM-powered graders, simulate multi-turn conversations, and define reliability metrics. The position also involves implementing staging validations, shadow-mode traffic analysis, and A/B rollout strategies, with feedback loops from production environments to enhance overall system robustness.

Responsibilities

- Design and implement build-time evaluation frameworks for agentic workflows using LangGraph or comparable orchestration frameworks

- Create deterministic and LLM-as-judge grading pipelines covering reasoning, trajectory accuracy, and output quality

- Develop test harnesses for multi-turn conversational simulations and context-retention scoring

- Define reliability assessment methods including multi-trial metrics (pass@k, pass^k)

- Implement CI/CD deployment gates that enforce quality thresholds and block releases not meeting standards

- Integrate staging validation, shadow-mode traffic comparison, and A/B rollout control in deployment pipelines

- Leverage AWS AgentCore Evaluations for on-demand and online scoring components connected to production feedback

- Convert production incidents into reusable regression cases for continuous quality improvement

- Collaborate with engineering and DevOps teams to embed evaluation gates into automated workflows

Requirements

- 4+ years of experience in building automated testing or evaluation frameworks for ML, LLM, or agentic systems

- Proven hands-on experience designing multi-layer evaluation suites with deterministic and LLM-based graders

- Expertise with CI/CD pipelines and implementing metric-based quality gates for automated deployments

- Practical knowledge of LangGraph or similar agent orchestration frameworks

- Strong background in designing simulation-based evaluation strategies and conversation-level tests

Nice to have

- Experience with AWS AgentCore Evaluations API (CreateEvaluation, custom evaluators)

- Familiarity with shadow-mode, canary, or A/B deployment practices for ML-based platforms

- Background in transforming production failures into build-time regression tests for agent workflows