Before foundation models are deployed to hundreds of millions of users, they must be stress-tested by humans who are actively trying to make them fail. Sovrano AI runs red-teaming projects for the world's leading AI labs - and we need diverse, intelligent adversarial testers from across Europe to do it.
Red-Teaming Fellows work directly on safety projects commissioned by the world's top LLM companies. The labs behind the most powerful AI systems on the planet rely on human red-teamers to find the failures their automated tests miss. You will be one of those people - and the work you do will influence how safe these models are before they reach the public.
Red-teaming requires you to think like an adversary - and then document your findings with academic rigour. It is intellectually demanding and emotionally draining at times. You will encounter model outputs that are wrong in surprising, disturbing, or deeply frustrating ways. That is the job.
KEY RESPONSIBILITIES Design and execute adversarial prompts intended to surface model failure modes - hallucinations, jailbreaks, logical inconsistencies, and harmful outputs Document failure cases with structured reports: prompt used, model response, severity classification, and recommended mitigation Test model behaviour across edge cases in your domain of expertise Contribute to structured red-teaming campaigns coordinated by Sovrano AI's safety leads IDEAL QUALIFICATIONS Bachelor's or Master's students across all disciplines at any European Sovrano AI partner university Genuinely adversarial thinkers who enjoy finding holes in systems Strong written English and structured documentation skills Comfortable sitting with ambiguity and making judgment calls without clear rules NICE TO HAVE Interest in AI ethics, safety, or alignment is a strong plus WHAT SUCCESS LOOKS LIKE Direct involvement in AI safety work at the world's leading LLM companies - one of the most consequential and sought-after areas in all of AI €16–26/hr, paid weekly - rate scales with output quality and domain expertise Verified Sovrano AI Safety Fellow credential Rare hands-on experience in AI safety and alignment that most researchers only read about CONTRACT & PAYMENT TERMS
Project-based. Red-teaming campaigns run in focused sprints of 2–6 weeks. Fellows are invited onto campaigns based on domain fit and quality scores from previous work. €16–26/hr paid weekly.