About the role
We build materials-science and engineering tasks that test how well an AI model does real expert work. You write a realistic task — a prompt, a data room, and a way to grade it — run it against the model, and tighten it until the model can no longer reason through it cleanly.
What you'll do
- Author domain-authentic tasks drawn from work you've actually done.
- Build the data room and grading criteria that make the task objectively scorable.
- Run the task against the model, analyse where it succeeds or fails, and iterate until it genuinely discriminates.
- Work with the project lead and other domain experts to settle what "good" looks like in your domain.
Who we're looking for
- Senior expertise in materials science — metals, ceramics, composites, characterization, additive manufacturing — or in electrical engineering (analog, RF, power) or mechanical engineering (design, CFD and fluids, heat transfer, pressure vessels, NDE, CAE).
- Comfortable navigating ambiguity — much of the work is still being defined, and shaping it is part of the role.
- Able to explain not just the right answer, but why a wrong answer is wrong.
Support
Daily onboarding sessions and multiple office hours every day.