Open bounties

AI Safety Red Teamer — English & Indonesian Bilingual Specialist

Paid remote expert bounty in Expert network (Global). Work in the software you already know; rates and ladder published up front.

$30-$250
per hr, published up front
1
slots left of 1
Remote
no AI experience required
Large language models AI evaluation platforms Annotation tooling

What you'll be paid to show

About the role

We're building a global red team of human data experts dedicated to making AI systems safer before they reach the public. This is not a typical QA or content moderation job — you'll be actively probing AI models with adversarial prompts, surfacing hidden vulnerabilities, and generating the structured red-team data that helps our AI systems detect and resist harmful, biased, or misleading outputs.

We're specifically looking for bilingual specialists fluent in **English and Indonesian** who can evaluate model behavior across both languages and cultural contexts, ensuring safety coverage isn't limited to English-only use cases.

What you'll do

  • Design and execute adversarial prompts intended to expose weaknesses in AI model outputs
  • Review and annotate AI responses for bias, misinformation, harmful behavior, and policy violations
  • Evaluate model outputs in both English and Indonesian, flagging language- or culture-specific failure modes
  • Document vulnerabilities with clear, structured write-ups that feed directly into model safety improvements
  • Collaborate with a distributed team of reviewers to calibrate standards and maintain consistent evaluation quality
  • Provide feedback on edge cases involving sensitive topics such as misinformation, discrimination, and unsafe instructions
  • Iterate on red-teaming strategies as models evolve and new failure patterns emerge

Requirements

  • Native or near-native fluency in **both English and Indonesian** (written and spoken)
  • Strong analytical and critical thinking skills, with attention to nuance and context
  • Comfort engaging with sensitive, controversial, or potentially disturbing content in a professional, objective manner
  • Excellent written communication skills for documenting findings clearly
  • Reliable access to a computer and stable internet connection for remote work
  • Prior experience in content moderation, linguistics, trust & safety, AI training data, or related fields is a plus
  • Ability to work independently while following detailed evaluation guidelines

Compensation

This role is paid hourly, in the range of $17–$25/hour depending on experience and evaluation throughput. Work is fully remote and flexible, with opportunities for ongoing engagement as the project scales.

What we're looking for

  • AI safety evaluation
  • Red teaming
  • Bilingual content review (English/Indonesian)
  • Bias and misinformation detection
  • Adversarial prompt design
  • Data annotation
  • Trust and safety analysis