What this role does
Design realistic, difficult shopping-related tasks (e.g. meal planning, restocking, gifting, travel prep) that stress-test AI models, then build the rubrics used to grade how well a model performs them.
Who we're looking for
- Subject-matter expertise in retail, e-commerce, consumer shopping behavior, or a related domain (nutrition, logistics, home/office management, etc.)
- Strong analytical and writing skills — able to define precise, low-variance grading criteria
- Comfortable working with AI outputs and evaluating them critically and objectively
- Prior experience in eval design, QA, tutoring, or curriculum/rubric design is a plus
Core responsibilities
- Author realistic shopping tasks with a clear, concrete end goal
- Build a grading rubric/verifier for each task that two independent reviewers would score consistently
- Review an AI agent's attempt at the task and grade it against the rubric
- Document where the AI succeeded or failed, with clear reasoning
Work style
- Independent, deadline-driven work
- High attention to detail and consistency
- Comfortable with ambiguity — tasks are self-directed with light guidelines