AI Safety Red Teamer — English & Finnish Bilingual Expert
Paid remote expert bounty in Expert network (Global). Work in the software you already know; rates and ladder published up front.
- $30-$250
- per hr, published up front
- 1
- slots left of 1
- Remote
- no AI experience required
What you'll be paid to show
About the role
We're building a global red-teaming initiative dedicated to making AI systems safer, more reliable, and less prone to harmful or biased outputs. We're looking for bilingual English & Finnish speakers who can bring native-level linguistic and cultural fluency to the evaluation of AI model behavior. This is a remote, flexible engagement for detail-oriented individuals who enjoy probing systems for weaknesses and contributing to the frontier of responsible AI development.
As an AI Safety Expert on this project, you'll act as an adversarial tester — crafting prompts designed to surface edge cases, biases, misinformation risks, and other harmful behaviors in AI model outputs. Your findings will directly shape how AI systems are trained to behave more safely and responsibly across languages and cultural contexts.
What you'll do
- Design and execute adversarial prompts ("red-teaming") to expose vulnerabilities in AI model responses
- Evaluate AI-generated content for bias, misinformation, harmful behavior, or policy violations
- Document and categorize findings with clear, structured feedback for model improvement teams
- Review sensitive and nuanced content in both English and Finnish, ensuring cultural and linguistic accuracy
- Collaborate with a distributed team of safety experts to refine testing methodologies
- Provide qualitative and quantitative assessments of model performance against safety benchmarks
- Maintain strict confidentiality and adhere to responsible AI evaluation guidelines throughout the engagement
Requirements
- Native or near-native fluency in both English and Finnish (written and spoken)
- Strong analytical and critical thinking skills, with comfort engaging with sensitive or adversarial content
- Prior exposure to AI, machine learning, content moderation, linguistics, or related fields is a plus
- Excellent written communication skills for documenting findings clearly and objectively
- Ability to work independently in a remote setting with minimal supervision
- High attention to detail and consistency when applying evaluation guidelines
- Comfort discussing topics such as bias, misinformation, and harmful content in a professional, objective manner
Compensation
This role offers an hourly rate of $48–$62, commensurate with experience and evaluation quality. Engagement is remote and flexible, with opportunities for ongoing work based on performance and project needs.
What we're looking for
- AI red-teaming
- Content evaluation
- Bilingual linguistic analysis (English/Finnish)
- Bias and misinformation detection
- Adversarial prompt design
- AI safety and alignment
- Qualitative data annotation