About the role
A frontier engineering-reasoning evaluation run in collaboration with a leading AI research lab. The work measures whether state-of-the-art models can reason from first principles in your engineering domain rather than retrieve facts from training data — and you get visibility into the model's internal reasoning on your own tasks.
What you'll do
Domain experts write simulations and specification sheets; the model attempts to design an artifact — controller gains, circuit parameters, geometry — that satisfies every spec.
You define the simulation, a set of specs with pass/fail thresholds, and the prompt. The model probes your simulation with a limited number of calls, then submits a final design. An agentic grader runs the simulation and scores spec satisfaction.
Who we're looking for
Support
Onboarding calls run daily, alongside internal tooling built to help you work faster. Prior model-evaluation experience is not required.