Open bounties

AI Safety Practitioner – Frontier Model Evaluation & Alignment

Paid remote expert bounty in Expert network (Global). Work in the software you already know; rates and ladder published up front.

$30-$250
per hr, published up front
1
slots left of 1
Remote
no AI experience required
Frontier AI/LLM models Evaluation and annotation platforms

What you'll be paid to show

About the role

We're looking for experienced **AI Safety Practitioners** to help evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics. This is a critical trust-and-safety function: you'll be reviewing AI-generated responses against evolving safety policies, flagging risks, and providing structured feedback that directly shapes how these models behave in production.

This role suits detail-oriented thinkers who are comfortable navigating nuanced ethical, legal, and policy questions, and who can apply rigorous judgment even when the "right answer" isn't obvious.

What you'll do

  • Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality
  • Review content spanning sensitive domains including misinformation, political persuasion, self-harm, violence, cybersecurity, and biosecurity
  • Apply and help refine evaluation rubrics and safety policies based on observed model behavior
  • Identify edge cases, policy gaps, and emerging risk patterns in model outputs
  • Provide structured, well-documented feedback that supports model improvement and alignment efforts
  • Collaborate with policy and research teams to calibrate judgments on ambiguous or contested cases
  • Maintain consistency and rigor across large volumes of nuanced evaluation tasks

Requirements

  • Strong ability to reason through ambiguous, policy-sensitive, and ethically complex content
  • Excellent written communication skills for documenting evaluation rationale clearly
  • Familiarity with AI safety concepts, content moderation, trust & safety, or policy evaluation work
  • Comfort engaging with sensitive subject matter (e.g., violence, self-harm, misinformation) in a professional, objective manner
  • High attention to detail and consistency when applying structured rubrics
  • Self-directed and reliable in a fully remote, asynchronous work environment
  • Prior experience in AI red-teaming, content policy, research annotation, or related evaluation work is a plus

Compensation

Hourly rate: **$60–$70/hour**, based on experience and evaluation throughput. Flexible, remote, contract-based engagement with opportunities for ongoing work.

What we're looking for

  • AI safety evaluation
  • Content policy review
  • Risk assessment
  • Model alignment
  • Trust & safety
  • Structured feedback and annotation
  • Ethical reasoning
  • Policy interpretation