Mercor is seeking experienced Python engineers to build the dataset that teaches AI agents to reason about real codebases.
Codebase Q&A is a reinforcement-learning environment for training AI agents to explore and reason about software repositories they have never seen before. You will build difficult codebase-exploration tasks: high-level engineering questions about production Python repositories that cannot be answered by reading a single file, and that require runtime evidence to resolve — the kind of question a senior engineer answers by actually investigating the system.
This track works on widely deployed Python codebases including Django, Twisted, Scrapy, SymPy, mypy, pip, coverage.py, mitmproxy, Certbot, Tornado, Werkzeug, sqlfluff and Meson.
Contract, task-based work paid at $130 per approved task, performed remotely in Mercor Studio.
Select tasks from a pre-validated pool of Python repositories, each pinned to a specific git commit and arriving with an engineering question, positive and negative rubrics, foils, and a golden solution
Explore the repository until you genuinely understand the subsystem in question, including the relevant PR or issue history
Rewrite the task's question so it defeats frontier coding agents while remaining well-defined and fairly answerable
Adjust rubrics and foils so they reward correct reasoning and reject plausible-but-wrong answers
Run the full validation loop in Studio — Check, Golden, Run, Analyze — and iterate until every stage passes
Document your work in the task README and submit for expert review, addressing reviewer feedback on returned tasks
3+ years of professional software engineering experience, with substantial production Python
Demonstrated ability to navigate a large, unfamiliar Python codebase and explain how it behaves at runtime — not just what the source says
Fluency in Python internals that matter in real systems: import machinery, descriptors and metaclasses, async and event loops, packaging, and C-extension boundaries
Experience with web frameworks, developer tooling, type checkers, or scientific/numeric libraries is highly relevant — the task mix is weighted toward architecture and system design (41%) and code onboarding (31%)
Precise written English: the questions and rubrics you write are the product
Comfortable with git at commit level, containerized environments, and command-line tooling
Fully remote and asynchronous, with expert team leads across US, India, and Nigeria time zones
Work is performed in Mercor Studio; time is tracked with Insightful, and platform access is provisioned through Okta
Onboarding requires identity verification, a background check, signed Terms of Work and CIIAA, a tax form, and passing a short calibration quiz before task access is granted
Payment is per approved task; every task is reviewed by a second engineer and a super reviewer before acceptance
Submit your resume or relevant technical background to get started
Qualified applicants may be asked to complete a brief technical assessment or provide additional information