We are running a paid study on advanced AI capabilities to find people who excel at designing tasks that expose where models fall short. This work feeds into a larger effort to evaluate and improve how artificial intelligence handles complex, multi-step requests. High performers in this trial assessment will be considered for ongoing paid research projects.
You will spend about an hour creating a demanding prompt based on a real workflow you know deeply. During a screen-recorded session, you will test your prompt in ChatGPT, identify where the model fails to deliver a complete file, and refine the request to make it harder. Finally, you will write a comprehensive grading rubric that someone else could use to evaluate any AI's attempt at your task. You will hand back your prompt, failure notes, rubric, and the generated file for review.
We welcome professionals, domain experts, and knowledge workers from any field who deeply understand complex, specialized workflows. You must be able to recognize a high-quality output in your area of expertise within seconds. We are looking for individuals who can creatively translate their daily tasks into challenging scenarios that require real reasoning and calculation.