Turing is one of the world’s leading AGI infrastructure companies, working with frontier AI labs to accelerate model development through high-quality training data, evaluations, and engineering talent.
Compensation: Market, please provide a specific hourly rate expectation
Availability: 40 hours/week, with at least 6 hours of IST/day
Type: Independent contractor
Duration: Approximately 2 months
Start: 01-Nov-2026
Bottleneck profiling and self-healing: Enable the agent to inspect execution DAGs and query profiles when jobs stall or degrade, isolate root causes, and autonomously create isolated branches that synthesize fixes for syntax errors, upstream schema changes, and column renames. Each fix is verified through dry-run assertions and submitted as a review-ready change request with full test results for human approval.
Statistical and semantic data quality forensics: Replace static thresholds with rolling time-series baselines that account for seasonality, day-of-week swings, and business close cycles, and profile categorical entropy, column distributions, and foreign key orphan rates to catch subtle corruption that row counts and schemas miss.
Lineage reconciliation and quarantine: Trace datasets from raw ingestion through marts and reporting views, run cross-tier checksums and consistency assertions before reporting cycles begin, and automatically quarantine partitions with severe violations so bad metrics never reach executive dashboards.
Conversational copilot and incident management: Generate root-cause incident briefs covering failure cause, affected downstream assets, and recovery steps to cut alert noise, and build a natural language interface for checking pipeline health, triggering selective partition re-runs, and analyzing data distributions.
Python (core): strong production-grade Python for automation, agent logic, testing, and data processing.
SQL (core): advanced SQL for profiling, validating, and troubleshooting data in warehouses such as Snowflake, BigQuery, or Databricks.
ETL/ELT and orchestration (core): hands-on experience building and operating pipelines with Airflow (or Dagster/Prefect) and dbt, and understanding of how failures originate and propagate.
Data quality, observability, and statistics: experience with tools like Great Expectations, Soda, or Monte Carlo, plus a solid grasp of anomaly detection, time-series baselines, and forecasting.
LLM/agent engineering: experience with tool calling, multi-step agents, and frameworks such as the Claude Agent SDK or LangGraph, with a focus on safety, evaluation, and reliability.
AI interview (~25 minutes)
Practical code/AI evaluation exercise (~30 minutes)
Hiring manager interview (~20 minutes)
The practical exercise focuses on your ability to review and evaluate AI-generated code, not competitive programming or algorithm puzzles.