← All jobs
Turing Verified
remote ·

Senior Software Engineer – Python (LLM Evaluation & Repository Validation)

Pay on listing
Share
Software Engineering remote
Eligible
Worldwide
Posted Jul 31, 2026

About Turing

Turing is one of the world's fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems. Founded in 2018, Turing partners with leading AI labs and enterprises to improve frontier AI models through high-quality data, expert human feedback, and rigorous evaluation frameworks.

Our contributors work on cutting-edge projects spanning coding, reasoning, agentic workflows, software engineering, and multimodal AI—helping shape the next generation of intelligent systems used by millions worldwide.


About the Role

We are looking for experienced Software Engineers with a strong background in developing and operating large-scale production systems. The ideal candidate has worked extensively on enterprise software, understands how production systems fail, has experience identifying security vulnerabilities in code, and has successfully implemented secure fixes and mitigations. This role requires practical experience in debugging complex production issues, participating in incident response, and building resilient, secure software.


What You'll Do

  • Design and build realistic coding-agent benchmark tasks using production-like repositories, tests, configs, documentation, and runtime scenarios.
  • Create benign engineering tasks such as bug fixes, feature additions, CI repairs, config migrations, integration updates, and runtime-state fixes.
  • Define clear utility requirements to verify that the coding agent completes the requested task correctly.
  • Define safety/alignment requirements to ensure the agent preserves system constraints, developer intent, data integrity, privacy, permissions, and oversight mechanisms.
  • Develop visible and hidden test suites to evaluate both task completion and unsafe agent behavior.
  • Create safe reference solutions and unsafe reference solutions where the agent passes utility checks but violates multiple safety/alignment requirements.
  • Identify and evaluate unsafe shortcuts such as disabling tests, weakening assertions, deleting protected data, leaking secrets, broadening permissions, or bypassing validation.
  • Package tasks with prompts, metadata, runnable repos, evaluators, reference patches, scoring rubrics, and calibration notes.
  • Analyze rollouts from frontier coding agents such as Claude Code and Codex to assess utility completion and safety/alignment violations.
  • Collaborate with engineering, QA, security, and client teams to improve task quality, evaluator reliability, and benchmark difficulty.

Required Qualifications

  • Bachelor's or Master's degree in Computer Science or a related technical discipline.
  • Minimum 8 years of hands-on software engineering experience in leading product companies or technology startups.
  • Strong programming experience in Python (mandatory) along with one or more  programming languages such as  Java, Go, C++
  • Demonstrated experience building and maintaining production software used by real customers.
  • Strong understanding of software architecture, debugging, and performance optimization.

Required Technical Experience

Production engineering

  • Experience developing software deployed in production at scale.
  • Strong understanding of software deployment pipelines, monitoring, logging, and observability.
  • Experience diagnosing production failures using logs, metrics, and distributed tracing.
  • Knowledge of high availability, fault tolerance, scalability, and disaster recovery principles.

Secure Software Development

  • Experience identifying security vulnerabilities through code reviews or security assessments.
  • Practical knowledge of common software vulnerabilities, including:
    • Injection attacks
    • Authentication and authorization flaws
    • Memory safety issues
    • Race conditions
    • Deserialization vulnerabilities
    • Secrets management issues
    • Input validation failures
  • Experience implementing secure code fixes and validating remediation.

Perks of Freelancing with Turing:

  • Work in a fully remote environment.
  • Opportunity to work on cutting-edge AI projects with leading LLM companies.

Offer Details:

  • Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. 
  • Employment type  : Contractor assignment (no medical/paid leave)
  • Duration of contract : 4-8 weeks; [expected start date is next week]

Evaluation Process:

  • Delviery review of candidate profile and delivery interview for 45-60 mins
Pay range
Pay on listing
Share
Similar roles

You might also like

Turing Verified New
remote
Senior Software Engineer – Python
Software Engineering
Posted Aug 2, 2026
Pay on listing
Terac Verified New
remote
US Project Managers: 30-Minute Interview on Tooling Workflows
Software Engineering
Posted Aug 1, 2026
$105 one-time
Terac Verified New
remote
Marketing Professionals: 30-Minute Interview on Tools and Workflows
Software Engineering
Posted Aug 1, 2026
$100 one-time
Terac Verified New
remote · hourly
Senior Full-Stack Engineers (React): AI Evaluation Environments
Software Engineering
Posted Aug 1, 2026
$75/hr