We're hiring business students to evaluate AI-generated content across strategy, consulting, and general business domains. This is a 12-week, full-immersion internship. The work is technical, hands-on, and paid. Around 20% of your time goes toward structured AI training: how large language models work, how RLHF fits into the development pipeline, what good evaluation looks like and why it matters. No prior AI knowledge is required. You'll learn everything on the job. The other 80% is live evaluation work on production AI models. This is the core of the internship. Your business education and your languages are the reason you're here.
KEY RESPONSIBILITIES 1. Teaching AI models what good business reasoning looks like. You'll get two AI-generated responses to the same business prompt and decide which one is better. That might mean comparing two AI-written strategy memos on a market entry question, checking whether the reasoning is sound, whether the frameworks are applied correctly, and whether a consultant or manager would consider the output useful. This is called RLHF (reinforcement learning from human feedback), and it's the most important step in making AI models useful for business applications. Your judgment directly shapes how the model behaves for every future user. 2. Fact-checking AI in your native language. AI models hallucinate. They make up market data, cite reports that don't exist, and confidently state things that are wrong. You'll catch them. If you speak Spanish and English, you might review business AI outputs in both languages, flagging incorrect claims, made-up sources, and translations that sound technically correct but no native speaker in a professional setting would write. Most AI evaluation today only covers English and Chinese. European languages are massively underserved. 3. Evaluating AI on cross-lingual business content. AI labs are building multilingual business tools, and they need evaluators who can judge quality in their native language, not just translate from English. You might receive a business prompt in English and evaluate an AI-generated response written natively in German, French, or Spanish. The question isn't whether the grammar is correct. It's whether the content reads like it was written by someone who thinks in that language and understands the local business context. 4. Red-teaming AI for safety. The EU AI Act now requires companies deploying AI to test their models for harmful outputs. You'll try to break them. That means finding prompts that cause the model to produce biased, misleading, or inappropriate business advice, documenting exactly how it fails, and scoring the severity. 5. Contributing to industry benchmarks. The evaluation work you do here feeds directly into the benchmarks top-tier AI labs use to assess and publicly report the factual capabilities of their models. When OpenAI, Google DeepMind, or Anthropic measure how well their models reason, datasets built by evaluators like you are part of what they test against. IDEAL QUALIFICATIONS Currently enrolled in a Bachelor's or Master's in Business Administration, Management, International Business, or a closely related program Fluent in English plus at least one other European language (German, French, Spanish, Italian, and Polish are in highest demand) Comfortable working independently in a fully remote setup Reliable internet connection and a quiet workspace NICE TO HAVE Prior experience in consulting, startups, project management, business development, or operations Exposure to strategy work, market analysis, or client-facing roles Experience working across cultures or in international teams CONTRACT & PAYMENT TERMS
Depending on experience (paid)