A leading AI research organization is seeking advanced LLM power users to evaluate how well AI systems handle personalized, real-world life tasks.
This role is for people who use AI tools heavily in their personal lives and can clearly judge whether an AI response is useful, personalized, realistic, and successful.
Strong candidates have:
Heavy personal usage of LLM products
Experience using AI for multi-step tasks, planning, research, decision-making, or personal workflows
Familiarity with tools such as ChatGPT, Claude, Gemini, Perplexity, Cursor, Windsurf, Codex, or other AI agents
Ability to explain what makes an AI output good, bad, incomplete, unsafe, or unrealistic
Strong written judgment and attention to detail
LLMs are quickly becoming personal assistants for everyday decisions, but truly useful AI needs to do more than produce generic advice. It needs to understand context, preferences, constraints, tradeoffs, and what success looks like in real life. Your evaluations will help improve how AI systems support people with practical, high-context tasks across food, health, productivity, careers, and learning. This work directly contributes to making AI assistants more personalized, trustworthy, and useful for real-world personal workflows.