We're hiring finance master's students to evaluate AI-generated financial content for the AI labs training the next generation of models. This is a 12-week, full-immersion internship. The work is technical, hands-on, and paid. Around 20% of your time goes toward structured AI training: how large language models work, how RLHF fits into the development pipeline, what good evaluation looks like and why it matters. No prior AI knowledge is required. You'll learn everything on the job. The other 80% is live evaluation work on production AI models. This is the core of the internship. Your finance knowledge is the reason you're here.
KEY RESPONSIBILITIES 1. Teaching AI models what good financial reasoning looks like. You'll get two AI-generated responses to the same finance prompt and decide which one is better. That might mean comparing two AI-written analyses of a company's debt-to-equity ratio, checking whether the reasoning holds up under scrutiny, whether the terminology is used correctly, and whether a professional analyst would trust the output. You might evaluate AI responses on corporate valuation, macroeconomics, portfolio theory, or financial statement analysis. Each evaluation gets scored on a rubric and includes a written rationale. This is called RLHF (reinforcement learning from human feedback), and your judgment directly shapes how these models behave for every future user. 2. Fact-checking AI in your native language. AI models hallucinate. They fabricate numbers, cite reports that don't exist, and confidently state things that are wrong. You'll catch them. If you speak German and French, you might review financial AI outputs in both languages, flagging incorrect figures, made-up sources, and translations that are technically passable but that no native-speaking finance professional would write. Most AI evaluation today only covers English. European languages are massively underserved, which is exactly why your bilingual profile matters. 3. Evaluating AI on ESG and sustainability reporting. European business schools teach ESG and sustainability as core curriculum. AI labs need people who understand CSRD, EU Taxonomy, and GRI frameworks to evaluate whether AI-generated ESG summaries are accurate, complete, and compliant. US-based competitors don't have evaluators with embedded EU regulatory knowledge. You do. 4. Building financial QA datasets. AI labs training financial assistants need question-answer pairs built by people who actually read financial statements. You might write and evaluate QA pairs ranging from basic extraction ("What was EBITDA in FY2024?") to analytical reasoning ("How does the debt-to-equity trend compare to industry benchmarks?"). Each answer includes a step-by-step reasoning chain. This structured data trains the next generation of enterprise financial AI tools. 5. Contributing to industry benchmarks. The evaluation work you do here feeds directly into the benchmarks top-tier AI labs use to assess and publicly report the factual capabilities of their models. When OpenAI, Google DeepMind, or Anthropic measure how well their models reason about finance, datasets built by evaluators like you are part of what they test against. IDEAL QUALIFICATIONS Currently enrolled in a Master's in Finance, Financial Economics, or a closely related program Fluent in English plus at least one other European language (German, French, Spanish, Italian, or Portuguese are in highest demand) Comfortable working independently in a fully remote setup Reliable internet connection and a quiet workspace NICE TO HAVE Prior experience in investment banking, asset management, corporate finance, audit, or financial consulting Familiarity with IFRS standards and European financial reporting Experience working across cultures or in international teams CONTRACT & PAYMENT TERMS
Depending on experience (paid)