Open bounties

AI Evaluation Analyst — LLM Output Quality & Spec Compliance

Paid remote expert bounty in Legal (Global). Work in the software you already know; rates and ladder published up front.

$30-$250
per hr, published up front
1
slots left of 1
Remote
no AI experience required
Large Language Models (LLMs) Annotation Platforms Evaluation/Rubric Tools

What you'll be paid to show

About the role

We're looking for a detail-obsessed AI Evaluation Analyst to help us assess and improve the quality of frontier large language model outputs. In this role, you'll work closely with detailed evaluation specs to judge, annotate, and document model behavior at scale. This is a remote, flexible-hours opportunity ideal for someone who thrives on precision, clear writing, and independent judgment.

You'll be part of a distributed team helping shape how AI systems are measured against real-world quality standards — your work directly informs how these models are trained and refined.

What you'll do

  • Evaluate frontier LLM outputs against detailed, evolving rubrics and specifications
  • Annotate and label data with a high degree of accuracy and consistency
  • Write clear, well-structured English assessments, notes, and justifications for your evaluations
  • Maintain spec fidelity even at high volume, without sacrificing quality or attention to detail
  • Work independently to interpret and apply detailed guidelines with minimal oversight
  • Flag edge cases, ambiguities, or inconsistencies in specs back to the team
  • Continuously calibrate your judgments against evolving quality standards and feedback

Requirements

  • Working knowledge of frontier LLM behavior and common failure modes
  • Strong experience with data annotation or evaluation workflows
  • Excellent written English with clarity, structure, and precision
  • Demonstrated ability to maintain fidelity to detailed specs across high volumes of work
  • Strong self-direction and ability to work independently against written guidelines
  • Comfortable with remote, asynchronous collaboration across time zones

Compensation

This role pays $20–$30 per hour depending on experience and evaluation throughput/quality. Engagement is remote and open to candidates globally.

What we're looking for

  • LLM Behavior Analysis
  • Data Annotation
  • Written Communication
  • Spec Compliance
  • Independent Judgment
  • Quality Evaluation