About the company
Turing’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks that improve model capabilities in software engineering, enterprise knowledge work, and advanced STEM reasoning. In software engineering, Turing is the largest and longest-running data provider in the category. Turing also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems inside mission-critical workflows. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment, turning real-world deployment signals into better data, evaluations, and more capable models.
Responsibilities
- Own Cloud Reliability at Scale: Operate and continuously improve production infrastructure across GCP, AWS, and Azure. Participate in a global on-call rotation, lead incident response, troubleshoot complex issues, drive root-cause remediation, and improve monitoring, logging, alerting, performance, capacity, and cloud cost efficiency. Use AI-assisted engineering tools to support log analysis, debugging, troubleshooting, and optimization.
- Build the Developer Platform: Build and support engineering platforms, release pipelines, and deployment automation using GitHub Actions, Google Cloud Build, Jenkins, and related tooling. Design and maintain secure, scalable Terraform-based IaC, develop reusable modules, automate workflows with Python, Go, Shell, or similar technologies, and operate production Kubernetes/GKE environments across deployment, scaling, networking, security, observability, and troubleshooting.
- Secure and Shape the Future of Infrastructure: Design and troubleshoot cloud networking, including VPCs, firewalls, load balancers, routing, DNS, and WAFs. Partner with security and engineering teams to implement secure-by-default infrastructure and meet requirements across identity and access management, network security, secrets management, and infrastructure security. Collaborate with engineering teams on scalable architectures, deployments, production issues, and operational best practices; mentor engineers and lead initiatives that improve reliability, security, and developer productivity.
Requirements
- 5+ years of experience in cloud, infrastructure, DevOps, SRE, or platform engineering, including hands-on work in GCP production environments.
- Independent experience with Terraform/IaC, CI/CD, using tools such as GitHub Actions, Jenkins, Cloud Deploy, or equivalent, and Kubernetes in production.
- Strong understanding of cloud architecture, distributed systems, and networking, including VPCs, firewalls, load balancers, and WAFs, along with experience securing cloud environments, preferably in GCP.
- Strong scripting or programming skills in Python, Go, Shell, or equivalent.
- Experience responding to production incidents, troubleshooting complex issues, and conducting root-cause analyses.
- Strong communication, ownership, and problem-solving skills.
- Willingness to participate in an on-call rotation alongside a globally distributed infrastructure team.
Conditions
- Work at the frontier of AI, helping the world’s leading AI labs improve their most advanced models by building expert datasets, RL environments, and first-of-a-kind benchmarks.