About the company
Turing builds large-scale datasets and reinforcement learning (RL) environments that power post-training for the world’s leading AI labs and enterprises. We create RL environments to evaluate and improve our customers' models on complex, long-range, multi-step workflows across high-GDP-value domains such as Finance, Sales, Retail, Developer Tools, Collaboration, Customer Experience.
Responsibilities
- End-to-End Ownership: Data Quality, Process Design, and Team Building: Lead the creation of datasets, RL environments, and evals focused on Coding Agents / Software Engineering for one or more AI lab customers. Ensure that everything you ship to clients meets frontier standards for realism, correctness, diversity, and difficulty. Set up quality rubrics, automated validation scripts, and human review processes for every stage of data generation. Build and lead cross-functional teams of software engineers, researchers, QAs, and data creators drawn from Turing’s 4M+ developer network. Interview, onboard, train, and mentor team members to ensure consistent output quality and technical excellence.
- Collaborate with Researchers at Frontier Labs: Act as the primary technical point of contact for your customer projects, interfacing directly with researchers and engineers at frontier AI labs to understand their coding agent roadmap and model data needs, to gather feedback, and to co-define success criteria for your projects. Provide regular progress updates, surface insights from model evaluations, and incorporate client feedback to improve future iterations.
- Drive Research, Sales Enablement, and Industry Thought Leadership: Fine-tune models in-house on Turing-generated datasets or Turing-rl-environment generated trajectories to determine model improvement as a proof of data quality. Proactively build benchmarks and run evals on frontier models and coding agents to identify strengths and weaknesses on SWE tasks, and leverage these insights to inform product roadmap. Equip customer-facing teams with the Evaluation reports, sample datasets, and trainings to enable them to communicate your data offerings to customers most effectively. Publish research papers and technical posts on Turing’s data products, innovations in our synthetic data generation / automation pipelines, evaluations of frontier agents and models, and Turing’s model fine-tuning results on our datasets.
- Build Tools and Infrastructure: Oversee development of internal tools that accelerate data generation and verification (e.g., automated data scraping pipelines, unit test generators, repo sandboxing). Design dashboards and APIs for customers to run model evals, view performance reports, and integrate Turing data directly into their post-training pipelines.
Requirements
- Post-training experience on SWE tasks or experience building coding agents: We expect that you have a deep understanding of data ingredients and design principles that lead to measurable coding model improvements, either from fine-tuning models to improve SWE capabilities or from building coding agents.
Conditions
- Work at the frontier of AI, helping the world’s leading AI labs improve their most advanced models by building expert datasets, RL environments, and first-of-a-kind benchmarks.
- Contribute to leading-edge AI research and showcase your work at top conferences such as ICLR, ICML, and NeurIPS.
- Bring frontier AI innovation to the enterprise, applying lessons learned from leading AI labs to solve real-world business challenges.
- Collaborate with and learn from exceptional colleagues with deep AI experience from Google, Meta, Amazon, and other leading technology companies.
- Move at the pace of AI innovation, with the speed, ownership, and impact of a startup.