We are hiring senior full-stack engineers to help build and refine next-generation coding environments that evaluate AI agents. This work directly supports the development of robust grading systems that accurately measure how effectively an AI agent solves complex coding tasks. The goal is to create secure evaluation suites that cannot be bypassed or easily cheated by the models.
Throughout this long-term engagement, you will work remotely with real website clones to identify bugs and write technical specifications. You will build comprehensive automated test suites designed to grade AI-generated code. This involves analyzing how an agent interacts with a given task and ensuring the grading logic is watertight. The process requires deep technical scrutiny and creative problem-solving to account for unpredictable AI behavior.
We welcome senior full-stack software engineers with extensive hands-on experience in React. You should be highly comfortable writing technical specs and building automated testing frameworks for complex web environments. Candidates with a strong background in web security, QA automation, or AI evaluation are highly encouraged to apply.
New
New
New