In this hourly, remote contractor role, you will work as a STEM & Technical Subject Matter Expert (SME) to create challenging expert-level questions, evaluate AI-generated technical content, and design detailed rubrics that define what a correct, rigorous, and expert-quality answer should contain.
You will develop problems within your area of expertise that test advanced reasoning rather than simple factual recall. You will identify weaknesses in frontier AI models, including incorrect assumptions, incomplete reasoning, mathematical or scientific errors, missed edge cases, and answers that appear plausible but fail under expert scrutiny.
A major focus of this role is rubric design and writing. You will translate complex technical judgment into clear, specific, and measurable evaluation criteria, including required concepts, reasoning steps, acceptable alternative approaches, critical errors, and partial-credit considerations. You may also critique and improve rubrics developed by other experts.
This role is with SME Careers, a fast-growing AI Data Services company and subsidiary of SuperAnnotate, delivering training data for many of the world’s largest AI companies and foundation-model labs. Your technical expertise directly helps improve the world’s premier AI models by making their reasoning more accurate, rigorous, and reliable.
Requirements
- Bachelor’s degree or higher in Mathematics, Physics, Chemistry, Engineering, Computer Science, Statistics, Applied Science, or another technical discipline.
- Strong professional proficiency in English, minimum C1, with the ability to write precise technical explanations and evaluation criteria.
- 3+ years of professional, academic, research, or industry experience in your stated technical domain.
- Demonstrated advanced expertise in a clearly defined technical field or specialization.
- Ability to create challenging technical questions that require multi-step reasoning, domain expertise, or professional judgment.
- Strong ability to design detailed evaluation rubrics defining required reasoning, correct methodology, acceptable alternatives, partial-credit criteria, and critical errors.
- Ability to distinguish between a correct final answer and an answer supported by valid versus flawed reasoning.
- Comfortable reviewing and critiquing other experts’ rubrics for ambiguity, missing criteria, redundancy, technical inaccuracies, or poor scoring design.
- High attention to detail when evaluating formulas, assumptions, units, methodology, edge cases, logical consistency, and technical terminology.
- Experience with exam writing, academic grading, peer review, research review, technical QA, standards development, or assessment design is strongly preferred.
- Prior experience with AI evaluation, RLHF, SFT, benchmarking, data annotation, prompt design, or LLM evaluation is preferred.
- Reliable, self-directed, and able to deliver consistent quality in an hourly, remote contractor workflow.