We're running a paid study on the effectiveness of coding evaluation harnesses used to test AI agents. Our team is building a comprehensive suite of programming environments designed to measure agent performance accurately. This work ensures that AI coding assistants are evaluated against realistic, high-quality development scenarios.
You will walk through a series of proposed coding tasks and environments during an AI-moderated session. We will ask you to review the structure, difficulty, and realism of these programming challenges. You will provide technical feedback on the evaluation harnesses and suggest improvements to the test suites. The conversation will focus on ensuring these tasks accurately reflect real-world engineering requirements.
We are hiring experienced software engineers based in South Asia who have a strong background in building and reviewing complex codebases. We welcome backend developers, test automation engineers, and full-stack engineers who understand evaluation harnesses. Ideal candidates have hands-on experience verifying realistic programming tasks in professional environments.