Coding & Rubric QC Expert (Pilot)
Engagement type: Independent contractor, paid per completed task
Location: Remote (must be based in North America)
Duration: Paid pilot, ~1 week, with potential for continued or expanded work for strong performers
Time commitment: ~10+ hours minimum during the pilot window, flexible scheduling
About the Role
We're looking for detail-oriented reviewers to help evaluate the quality of AI benchmark tasks. Each task gives you a question, a proposed "gold" answer, and a rubric of claims describing what a correct answer should contain. Your job is to judge whether that rubric actually holds up — not to write new content, but to critically audit work that's already been built.
Day to day, this looks like:
- Reviewing each rubric claim and judging whether it's clear, specific, and actually verifiable — or vague, redundant, and unusable.
- Marking each claim as pass, fail, or flagged for follow-up, with a short written reason.
- Spotting requirements that don't actually connect to the question or the answer.
- Double-checking that the "gold" answer is actually correct before treating it as ground truth.
- Giving an overall accept or reject call on the task, backed by your specific findings.
This is evaluative, judgment-driven work rather than open-ended coding or writing. If you enjoy code review, QA, or grading/auditing other people's work, this will feel familiar.
Average task time is expected to be around 20 minutes.
What We're Looking For
Required:
- Based in North America
- Bachelor's degree in Computer Science (or equivalent completed degree)
- At least 3 months of prior experience doing human-data / annotation / AI-training work of any kind
- At least 1 year of software engineering experience outside of human-data work (any language or stack is fine — no specific tech stack required)
Strongly preferred:
- Direct experience grading or auditing rubrics, benchmarks, or "golden set" answers against a reference or acceptance criteria
- Experience critically reviewing someone else's work product — flagging missing requirements, vague criteria, or incorrect reference answers
- A track record of writing clear, evidence-backed feedback rather than simple pass/fail calls with no explanation
- Prior experience as a task reviewer or QC lead on another data-annotation project
- Prior experience specifically on coding-related annotation projects
Compensation & Pay
- Paid per completed, accepted task (not hourly)
Requirements
CS degree, 1+ yr software engineering experience, 3+ months human-data/annotation experience, North America-based