AI Safety Practitioner – Frontier Model Evaluation & Alignment
Paid remote expert bounty in Expert network (Global). Work in the software you already know; rates and ladder published up front.
- $30-$250
- per hr, published up front
- 1
- slots left of 1
- Remote
- no AI experience required
What you'll be paid to show
About the role
We're looking for experienced **AI Safety Practitioners** to help evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics. This is a critical trust-and-safety function: you'll be reviewing AI-generated responses against evolving safety policies, flagging risks, and providing structured feedback that directly shapes how these models behave in production.
This role suits detail-oriented thinkers who are comfortable navigating nuanced ethical, legal, and policy questions, and who can apply rigorous judgment even when the "right answer" isn't obvious.
What you'll do
- Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality
- Review content spanning sensitive domains including misinformation, political persuasion, self-harm, violence, cybersecurity, and biosecurity
- Apply and help refine evaluation rubrics and safety policies based on observed model behavior
- Identify edge cases, policy gaps, and emerging risk patterns in model outputs
- Provide structured, well-documented feedback that supports model improvement and alignment efforts
- Collaborate with policy and research teams to calibrate judgments on ambiguous or contested cases
- Maintain consistency and rigor across large volumes of nuanced evaluation tasks
Requirements
- Strong ability to reason through ambiguous, policy-sensitive, and ethically complex content
- Excellent written communication skills for documenting evaluation rationale clearly
- Familiarity with AI safety concepts, content moderation, trust & safety, or policy evaluation work
- Comfort engaging with sensitive subject matter (e.g., violence, self-harm, misinformation) in a professional, objective manner
- High attention to detail and consistency when applying structured rubrics
- Self-directed and reliable in a fully remote, asynchronous work environment
- Prior experience in AI red-teaming, content policy, research annotation, or related evaluation work is a plus
Compensation
Hourly rate: **$60–$70/hour**, based on experience and evaluation throughput. Flexible, remote, contract-based engagement with opportunities for ongoing work.
What we're looking for
- AI safety evaluation
- Content policy review
- Risk assessment
- Model alignment
- Trust & safety
- Structured feedback and annotation
- Ethical reasoning
- Policy interpretation