← All jobs
Sovrano AI Verified
remote ·

AI Evaluation Intern - STEM & Engineering

$500/mo
Share
Other Engineering remote

We're hiring engineering and STEM master's students to evaluate AI-generated technical and mathematical content. This is a 12-week, full-immersion internship. The work is technical, hands-on, and paid. Around 20% of your time goes toward structured AI training: how large language models work, how RLHF fits into the development pipeline, what good evaluation looks like and why it matters. No prior AI knowledge is required. You'll learn everything on the job. The other 80% is live evaluation work on production AI models. This is the core of the internship. Your quantitative skills and technical reasoning are the reason you're here.

KEY RESPONSIBILITIES 1. Evaluating AI on mathematical and quantitative reasoning. You'll get two AI-generated solutions to the same math or engineering problem and decide which one is better. These cover calculus, linear algebra, probability, statistics, and optimization at the undergraduate and early graduate level. You're not just checking whether the final answer is right. You're tracing the reasoning chain step by step, identifying where errors creep in, and ranking explanatory clarity. AI models are getting better at math, but they still make subtle mistakes that only someone with genuine quantitative training can spot. This is called RLHF (reinforcement learning from human feedback), and it's the single highest-demand category in AI evaluation right now. 2. Evaluating AI-generated code. AI coding assistants are used by over 50% of professional developers. AI labs need evaluators who can review AI-generated code across languages (Python, JavaScript, SQL, C++, and others) and assess it for correctness, efficiency, readability, and whether it actually solves the stated problem. You'll compare multiple AI-generated solutions and rank them. Competitive programming benchmarks and real-world software engineering tasks are both in scope. 3. Fact-checking AI in your native language. AI models hallucinate. They produce plausible-looking but wrong calculations, cite formulas incorrectly, and confidently misapply technical concepts. You'll catch them. If you speak German or French alongside English, you'll review technical AI outputs in both languages and flag errors. Most AI evaluation today only covers English. Technical content in European languages is almost entirely unevaluated. 4. Red-teaming AI for safety. The EU AI Act now requires companies deploying AI to test their models for harmful outputs. You'll try to break them in your domain. That means finding prompts that lead to dangerously incorrect technical calculations, flawed engineering recommendations, or outputs that could cause real-world harm if someone trusted them. Documenting failure modes in technical domains is one of the highest-value evaluation tasks. 5. Contributing to industry benchmarks. The evaluation work you do here feeds directly into the benchmarks top-tier AI labs use to assess and publicly report the factual capabilities of their models. When OpenAI, Google DeepMind, or Anthropic measure how well their models reason about STEM subjects, datasets built by evaluators like you are part of what they test against. IDEAL QUALIFICATIONS Currently enrolled in a Master's in Engineering (any field), Computer Science, Mathematics, Physics, or a closely related STEM program Fluent in English plus at least one other European language Strong quantitative skills (you should be comfortable with graduate-level math) Comfortable working independently in a fully remote setup Reliable internet connection and a quiet workspace NICE TO HAVE Prior experience in engineering, software development, data science, or technical R&D Competitive programming, research publications, or applied math experience Experience working across cultures or in international teams CONTRACT & PAYMENT TERMS

Depending on experience (paid)

Pay range
$500/mo
Share
Similar roles

You might also like

Turing Verified
United States · remote
Senior Infrastructure / Backend Engineer — GCP Production Systems
Other Engineering
Posted Aug 28, 2026
Pay on listing
Turing Verified
remote
Domain Experts - Aerospace / Flight-Dynamics Engineer
Other Engineering
Posted Aug 27, 2026
Pay on listing
Mercor Verified
United States · remote · hourly
Mechanical Engineers
Other Engineering
Posted Aug 21, 2026
$70–$110/hr
Mercor Verified
United States · hybrid · full-time
Materials Science Domain Expert
Other Engineering
Posted Aug 19, 2026
$70–$110/hr 40h/wk