We're hiring MBA interns to oversee and quality-check the AI evaluation work produced by our team. This is a 12-week, full-immersion internship. You sit above the evaluators. They produce the data. You make sure it's right. Around 20% of your time goes toward structured AI training: how large language models work, how RLHF fits into the development pipeline, what evaluation quality standards look like, and how quality assurance works in AI data production. No prior AI knowledge is required. You'll learn everything on the job. The other 80% is QA, oversight, and product testing.
KEY RESPONSIBILITIES 1. Reviewing and auditing evaluation output. Our evaluation team produces thousands of RLHF preference rankings, fact-checks, and QA datasets per week across multiple domains and European languages. Each evaluation includes a rating, a rubric score, and a written rationale. Your job is to audit that output. You'll review samples, check whether evaluators are applying rubrics correctly, whether their rationales hold up under scrutiny, and whether quality is consistent. When you find issues, you document them, identify patterns, and feed that back to improve the process. This is the quality layer between our evaluators and the AI labs buying the data. 2. Calibration and consistency. AI evaluation data is only useful if it's consistent. If two evaluators score the same AI output differently, that's a signal. You'll run inter-annotator agreement checks, identify where evaluators are diverging, and work out whether it's a rubric problem, a training gap, or a genuine edge case that needs resolution. You'll help calibrate new evaluators during onboarding and flag quality trends over time. 3. Testing Sovrano AI's products and tools. We're building internal tooling and AI-powered products (interview preparation agents, evaluation platforms, workflow automation). You'll be one of the first people to use them. That means running through workflows end to end, finding bugs, edge cases, and UX issues, documenting everything clearly, and working directly with the product team to get things fixed. 4. Process improvement. With your MBA background and 3-5 years of professional experience, you've seen how operations work (and don't work) in real companies. You'll bring that perspective to how we run things: evaluator onboarding, quality feedback loops, client reporting, internal documentation. If something could work better, you're expected to say so and build the fix. 5. Contributing to industry benchmarks. The evaluation work you do here feeds directly into the benchmarks top-tier AI labs use to assess and publicly report the factual capabilities of their models. When OpenAI, Google DeepMind, or Anthropic measure how well their models reason, datasets built by evaluators like you are part of what they test against. IDEAL QUALIFICATIONS Currently enrolled in an MBA program (or recently completed) 3-5 years of professional experience before your MBA, in any industry Fluent in English plus at least one other European language Comfortable working independently in a fully remote setup Reliable internet connection and a quiet workspace NICE TO HAVE Prior experience in QA, operations, consulting, project management, or any role where you were responsible for output quality Experience working across cultures or in international teams Comfort with structured data, spreadsheets, and basic analytics CONTRACT & PAYMENT TERMS
€250 per week