About the Role
Join our global AI Community as a Bioinformatics AI Expert specializing in advanced benchmark dataset creation and AI model evaluation.
In this role, you will design complex, reasoning-intensive, "information-reversal" evaluation tasks to measure the autonomous tool-use and data analysis capabilities of Agent-style Large Language Models (LLMs) on real "wet lab / dry lab" biological data (e.g., RNA-seq, FASTQ, BAM, mzML).
Key Responsibilities
- Information-Reversal Problem Creation: Construct complex scientific scenario questions where AI models receive anonymized raw/processed biological datasets and must deduce hidden identities, experimental conditions, variant locations, or functional pathways.
- Data Workspace Packaging: Build clean, self-contained zip workspace archives containing real or synthesized biological datasets (e.g., expression tables, sequence files, metadata).
- Multi-Step Tool Orchestration: Design tasks that strictly require multi-step external tool execution (e.g., BLAST, genomic interval tools, variant annotators, phylogenetic software, differential expression pipelines) rather than simple code execution or direct text lookups.
- Chain of Thought (CoT) & Rubric Engineering: Author step-by-step resolution paths and rigid all-or-nothing grading rubrics featuring canonical answers and explicitly accepted variants for automated LLM evaluation.
- Quality Assurance & Verification: Ensure all benchmark deliverables pass automated validation checks, eliminating answer leakage in metadata or filenames, and guaranteeing single, deterministic correct answers.