AI Developer Trace Task Auditor — Software Development Quality Review
Paid remote expert bounty in Finance Ops (Remote — United States). Work in the software you already know; rates and ladder published up front.
- $30-$250
- per hr, published up front
- 1
- slots left of 1
- Remote
- no AI experience required
What you'll be paid to show
About the role
We're looking for experienced software developers to serve as Trace Task Auditors, evaluating the quality and correctness of AI-assisted coding sessions used to train and benchmark frontier AI models. In this role, you'll review end-to-end developer workflows — from prompt to final code — produced using modern AI-assisted development tools, and provide precise, rubric-based feedback that directly shapes model quality.
This is a remote, contract-based opportunity ideal for senior engineers who enjoy deep code review, technical writing, and working at the intersection of software engineering and AI evaluation.
What you'll do
- Review full coding session traces generated with AI-assisted developer tools, including agentic and spec-driven workflows
- Judge the correctness, efficiency, and soundness of code changes, debugging steps, and architectural decisions
- Evaluate the reasoning quality behind AI-generated suggestions and developer responses to them
- Apply structured rubrics consistently to score traces across multiple quality dimensions
- Write clear, actionable, and well-organized feedback explaining scoring decisions
- Flag edge cases, ambiguous rubric scenarios, or systemic issues observed across sessions
- Collaborate with a distributed team of auditors to maintain scoring consistency and calibration
Requirements
- 3+ years of professional software development experience
- Hands-on experience with AI-assisted coding tools and agentic or spec-driven development workflows
- Strong code-reading and debugging skills across full-stack or backend systems
- Excellent written communication skills, with the ability to produce clear, structured feedback
- Comfort working independently against detailed rubrics and evaluation guidelines
- Ability to work remotely within the United States
**Nice to have:**
- Prior experience with AI model evaluation, data labeling, or RLHF-style annotation work
- Exposure to multiple programming languages and frameworks
- Experience mentoring or reviewing code for other engineers
Compensation
This role pays **$70–$90 per hour**, based on experience and demonstrated evaluation quality. Engagement is remote, flexible, and contract-based, with opportunities for ongoing work based on performance.
What we're looking for
- Software Development
- Code Review
- Debugging
- AI-Assisted Development Tools
- Technical Writing
- Quality Auditing
- Rubric-Based Evaluation