Mercor is assembling a panel of energetic materials and propulsion experts to red-team frontier AI models. The goal is to test whether a model can correctly judge the misuse potential of a technical request — answering legitimate questions fully while refusing genuinely dangerous ones.
You will:
This domain lives on the dual-use line more than most. The same propulsion analysis that qualifies a motor for flight describes a capability; the same safe-and-arm logic that keeps a system inert is the part someone would want to defeat. A model that refuses a routine performance or qualification question is failing as badly as one that answers a question it shouldn't. Drawing that line reliably is why we need practitioners rather than generalists.
This is the engineering half of the domain: we're looking for people who have designed, qualified or tested propulsion and initiation systems, with real hardware and real test data behind them. What matters most is being able to tell a routine professional question from one that's fishing for something else: the same effects model that sizes a safety standoff also sizes an intent. You will be writing prompts that sit exactly on that line, then judging whether the model held it.
Examples of relevant backgrounds (ideally with red-teaming experience):
This is writing-intensive work. Every judgment you make needs a written rationale that a non-specialist can follow. Prior technical writing, published research, or expert witness experience is a strong signal; please include a sample or link. You will also be reading and writing about misuse scenarios in your field for sustained periods. We brief experts on this in advance, and you can pause or step away at any point without penalty.
Your work here will not involve, and must not draw on, classified or export-controlled information, or anything covered by an NDA or prepublication review obligation. If you hold such obligations you may still be a good fit; tell us in your application and we will scope the work accordingly.
New
New
New