remotely.living

AI Safety Practitioner

Mercor · Remote - Albania / Austria / Bosnia & Herzegovina / Belgium / Bulgaria / Switzerland / Czechia / Germany / Denmark / Estonia / Spain / Finland / France / United Kingdom / Greece / Croatia / Hungary / Ireland / Iceland / Italy / Liechtenstein / Lithuania / Luxembourg / Latvia / Monaco / Moldova / North Macedonia / Malta / Netherlands / Norway / Poland / Portugal / Romania / Serbia / Sweden / Slovenia / Slovakia / San Marino / United States · Contract · 2026-07-16

Apply for this job

Job description

We are seeking experienced **AI Safety Practitioners** to evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and feedback.

## Responsibilities

- Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality.

- Review content involving misinformation, political persuasion, self-harm, violence, cyber, biosecurity, and other sensitive domains.

- Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking.

- Identify unsafe outputs, hallucinations, reasoning failures, and policy violations.

- Provide structured feedback to improve model alignment and safety performance.

- Collaborate with AI researchers and safety teams on ongoing evaluation initiatives.

## Required Qualifications

- Bachelor's degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline.

- 5+ years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a related field.

- Excellent written English, critical thinking, and analytical reasoning skills.

- Ability to consistently evaluate nuanced and policy-sensitive scenarios.

## Preferred Qualifications

- Experience with AI Safety, RLHF, SFT, Trust & Safety, or AI evaluation.

- Familiarity with safety policies, content moderation, or evaluation rubric development.

- Experience reviewing complex, high-risk, or ambiguous content.

## Why Join?

- Shape the safety and behaviour of frontier AI models used by millions worldwide.

- Work on challenging, real-world safety evaluations across nuanced and high-impact domains.

- Collaborate with leading AI researchers, engineers, and safety teams.