Open bounties

AI Safety Red Teamer — Adversarial Testing for Frontier AI Systems

Paid remote expert bounty in Expert network (Global). Work in the software you already know; rates and ladder published up front.

$30-$250
per hr, published up front
1
slots left of 1
Remote
no AI experience required
Large Language Models Frontier AI Systems Prompt Engineering Tools

What you'll be paid to show

About the role

We're building a distributed team of skilled adversarial testers to help identify vulnerabilities in frontier AI systems before they reach the public. As an AI Safety Red Teamer, you'll be on the front lines of AI safety — probing large language models for weaknesses, uncovering unsafe behaviors, and helping make these systems more robust against misuse.

This is intellectually demanding work that sits at the intersection of security research, policy, and applied AI. You'll need to think like an attacker, reason across ambiguous and high-stakes topics, and communicate findings with precision.

What you'll do

  • Design and execute adversarial prompts intended to stress-test frontier AI models under realistic and creative attack scenarios.
  • Identify jailbreaks, prompt injection vulnerabilities, unsafe completions, hallucinations, and policy violations.
  • Systematically evaluate model behavior across sensitive and high-risk domains, including misinformation, cybersecurity, biosecurity, fraud, and political content.
  • Probe model responses in ambiguous "grey-area" scenarios where safety boundaries are unclear or contested.
  • Document discovered vulnerabilities with clear reproduction steps, severity assessments, and supporting evidence.
  • Collaborate with safety and policy teams to refine attack taxonomies and improve evaluation rubrics.
  • Iterate on testing methodologies as models and mitigations evolve.

Requirements

  • Strong understanding of large language model behavior, failure modes, and common jailbreak/adversarial techniques.
  • Demonstrated interest or background in AI safety, security research, red teaming, or adversarial ML.
  • Excellent written communication skills for documenting vulnerabilities clearly and objectively.
  • Comfort engaging with sensitive, high-risk, or ethically ambiguous content in a professional, safety-focused manner.
  • Sound judgment and discretion when handling potentially harmful information.
  • Ability to work independently and reliably in a remote, asynchronous environment.
  • Nice to have: prior experience with AI red teaming programs, bug bounty work, cybersecurity, misinformation research, or trust & safety roles.

Compensation

This role pays **$70–$84/hour**, commensurate with experience and demonstrated expertise. Engagement is remote and flexible, with multiple openings available for qualified specialists.

What we're looking for

  • AI Red Teaming
  • Adversarial Prompt Design
  • AI Safety Evaluation
  • Jailbreak & Exploit Identification
  • Risk & Policy Analysis
  • Cybersecurity Awareness
  • Misinformation Analysis
  • Technical Documentation