AI Session Annotator: Voice & Vision Assistant Reviewer (English)
Paid remote expert bounty in Creative & Content (Global). Work in the software you already know; rates and ladder published up front.
- $30-$250
- per hr, published up front
- 1
- slots left of 1
- Remote
- no AI experience required
What you'll be paid to show
About the role
We're building a team of sharp, detail-oriented English speakers to evaluate recorded sessions between users and a voice-and-camera AI assistant. In each session, a person speaks to the assistant while their camera streams a live view of whatever is in front of them — a recipe, a broken appliance, a math problem, a piece of furniture being assembled. Your job is to watch, listen, and judge: did the assistant actually understand what it saw and heard, and did it respond accurately and helpfully?
This is not passive video-watching. Your written notes become the primary signal used to prioritize what gets fixed next in the underlying model. You're the human judgment layer between raw AI output and product quality — precision and honesty in your scoring matter far more than speed.
What you'll do
- Review recorded sessions where a user interacts with a voice-and-camera AI assistant in real time.
- Cross-check the assistant's verbal responses against what is actually visible in the accompanying video feed.
- Score each session against a structured evaluation rubric covering accuracy, relevance, and helpfulness.
- Write clear, specific comments explaining *why* a response succeeded or failed — vague feedback isn't useful, so we expect concrete detail.
- Flag hallucinations, missed visual details, mismatched context, or unsafe/incorrect guidance.
- Maintain consistent judgment across sessions so that scoring stays comparable and defensible over time.
- Occasionally participate in calibration discussions to align on edge cases and rubric interpretation.
Requirements
- Native or near-native fluency in written and spoken English, with strong attention to nuance and detail.
- Comfortable watching video content closely and cross-referencing it against audio/text in real time.
- Strong written communication skills — you'll be producing rubric-based评价 notes that others rely on.
- Sound judgment and ability to apply a rubric consistently, even in ambiguous cases.
- Reliable internet connection and a quiet environment suitable for focused audio/video review.
- Comfortable working independently and asynchronously with minimal supervision.
**Nice to have:**
- Prior experience in content moderation, QA, transcription review, or AI/ML data annotation.
- Familiarity with how voice assistants or multimodal AI systems work.
- Experience writing structured feedback or evaluation reports.
Compensation
This role pays **$24–$30/hour**, fully remote, with flexible scheduling. Compensation is based on demonstrated accuracy and consistency during onboarding and calibration rounds.
What we're looking for
- Content review and annotation
- Multimodal (audio/video) evaluation
- Rubric-based scoring
- Written feedback and reporting
- Attention to detail
- Quality assurance judgment