We're assembling a group of senior software and ML engineers to work on one of the hardest open problems in AI today: how well can frontier models reason about LLM inference systems, GPU-level performance optimization, and model-serving architecture? You'll help design evaluation scenarios, write reference solutions, and grade model outputs on real systems-engineering problems — the same class of problems you likely work on day-to-day.
This is a research-and-evaluation role, not a traditional engineering job — you won't be shipping production code for us, you'll be defining what "correct" and "excellent" look like for AI models tackling the problems you already know deeply.
Responsibilities
- What You'll Do Design realistic technical scenarios and problem sets in LLM inference optimization, GPU kernel design, and systems performance engineering
Requirements
- 5–10 years of professional software engineering or ML systems experience
- Bachelor's degree or higher in Computer Science, Computer Engineering, Electrical Engineering, or a related field
- Strong, hands-on background in one or more of: GPU inference optimization, custom CUDA kernel development, model-serving systems (e.g., vLLM, TensorRT-LLM, SGLang), quantization, or optimizer design
- Experience at a recognized technology or AI company, in a role with real systems-engineering ownership
- Strong written communication — you'll be authoring technical explanations and feedback, not just code
Preferred Qualifications
- Direct experience with SGLang, Mamba/Mamba2 architectures, or IBM Granite-family models
- Open-source contributions to inference/serving frameworks (vLLM, SGLang, TensorRT-LLM, etc.)
- Experience with distributed inference (tensor/expert parallelism, continuous batching, KV-cache optimization)
Why Apply
- Remote, contract, part-time
- Flexible — work on your own schedule