AI Safety Red Teamer — English & Indonesian Bilingual Specialist
Paid remote expert bounty in Expert network (Global). Work in the software you already know; rates and ladder published up front.
- $30-$250
- per hr, published up front
- 1
- slots left of 1
- Remote
- no AI experience required
What you'll be paid to show
About the role
We're building a global red team of human data experts dedicated to making AI systems safer before they reach the public. This is not a typical QA or content moderation job — you'll be actively probing AI models with adversarial prompts, surfacing hidden vulnerabilities, and generating the structured red-team data that helps our AI systems detect and resist harmful, biased, or misleading outputs.
We're specifically looking for bilingual specialists fluent in **English and Indonesian** who can evaluate model behavior across both languages and cultural contexts, ensuring safety coverage isn't limited to English-only use cases.
What you'll do
- Design and execute adversarial prompts intended to expose weaknesses in AI model outputs
- Review and annotate AI responses for bias, misinformation, harmful behavior, and policy violations
- Evaluate model outputs in both English and Indonesian, flagging language- or culture-specific failure modes
- Document vulnerabilities with clear, structured write-ups that feed directly into model safety improvements
- Collaborate with a distributed team of reviewers to calibrate standards and maintain consistent evaluation quality
- Provide feedback on edge cases involving sensitive topics such as misinformation, discrimination, and unsafe instructions
- Iterate on red-teaming strategies as models evolve and new failure patterns emerge
Requirements
- Native or near-native fluency in **both English and Indonesian** (written and spoken)
- Strong analytical and critical thinking skills, with attention to nuance and context
- Comfort engaging with sensitive, controversial, or potentially disturbing content in a professional, objective manner
- Excellent written communication skills for documenting findings clearly
- Reliable access to a computer and stable internet connection for remote work
- Prior experience in content moderation, linguistics, trust & safety, AI training data, or related fields is a plus
- Ability to work independently while following detailed evaluation guidelines
Compensation
This role is paid hourly, in the range of $17–$25/hour depending on experience and evaluation throughput. Work is fully remote and flexible, with opportunities for ongoing engagement as the project scales.
What we're looking for
- AI safety evaluation
- Red teaming
- Bilingual content review (English/Indonesian)
- Bias and misinformation detection
- Adversarial prompt design
- Data annotation
- Trust and safety analysis