We are running a paid trial to build a bench of people who are exceptionally good at designing hard tasks that expose where AI models fall short. This work helps us understand the boundaries of current model capabilities and evaluate their performance on complex, multi-step requests. High performers in this trial may be considered for ongoing paid work of the same kind.
You will spend about an hour translating a real workflow you know deeply into a demanding prompt for an AI model. Using your own ChatGPT account, you will test your prompt, identify where the AI fails, and refine it to ensure it requires real reasoning and produces a concrete file. You will then write a detailed grading rubric that a stranger could use to evaluate the AI's attempt. Throughout this process, your screen, camera, and microphone will be recorded so we can assess your critical thinking.
We are looking for professionals and subject matter experts who possess deep knowledge of a specific, complex workflow in their job or personal life. You must be capable of quickly judging the quality of an output in your chosen domain. We welcome applicants from all professional backgrounds who have a computer and a basic understanding of how to interact with large language models.