← All jobs
Verita AI Verified
remote ·

Task & Evaluation Specialist

$60/hr
Share
Software Engineering remote
Posted Sep 2, 2026

What this role does Design realistic, difficult shopping-related tasks (e.g. meal planning, restocking, gifting, travel prep) that stress-test AI models, then build the rubrics used to grade how well a model performs them.

Who we're looking for

  • Subject-matter expertise in retail, e-commerce, consumer shopping behavior, or a related domain (nutrition, logistics, home/office management, etc.)
  • Strong analytical and writing skills — able to define precise, low-variance grading criteria
  • Comfortable working with AI outputs and evaluating them critically and objectively
  • Prior experience in eval design, QA, tutoring, or curriculum/rubric design is a plus

Core responsibilities

  • Author realistic shopping tasks with a clear, concrete end goal
  • Build a grading rubric/verifier for each task that two independent reviewers would score consistently
  • Review an AI agent's attempt at the task and grade it against the rubric
  • Document where the AI succeeded or failed, with clear reasoning

Work style

  • Independent, deadline-driven work
  • High attention to detail and consistency
  • Comfortable with ambiguity — tasks are self-directed with light guidelines
Pay range
$60/hr
Share
Similar roles

You might also like

micro1 Verified New
remote · hourly
Software Engineer
Software Engineering
Posted Sep 11, 2026
$90–$140/hr
micro1 Verified New
remote · hourly
Puzzle Solver (Coding)
Software Engineering
Posted Sep 11, 2026
$50–$80/hr
Turing Verified New
remote
Software Engineer Pod Lead
Software Engineering
Posted Sep 11, 2026
Pay on listing
Turing Verified New
remote
Software Engineer Mining
Software Engineering
Posted Sep 11, 2026
Pay on listing