About the company
At Toloka AI we create data that powers leading GenAI models and innovations. We work with frontier labs, big tech, renowned AI startups, enterprises and non-profit research organizations worldwide. We use a combination of Experts + Crowd + Tech Platform to teach AI models to reason and evaluate their efficacy and safety. We have experts in more than 50 different domains—from doctors and lawyers to physicists and engineers and boast one of the most diverse global crowds, representing over 100 countries and speaking 40+ languages. We are a well-funded startup with an enviable portfolio of clients including Anthropic, Amazon, Microsoft, Poolside, Recraft, and Shopify.
Responsibilities
- Create tasks and the criteria to evaluate them. Each task consists of:
- An isolated environment - an emulation of a developer's workstation: a Linux machine with dev tools (terminal, CLI), MCP servers (repository, task tracker, messenger, documentation, etc.), and a real web application codebase;
- A task - a prompt, just like in real work (a ticket link, a messenger conversation, anything). The agent figures it out on its own - reads the ticket, checks discussions, searches docs, explores the code. All of that context is also part of the task and needs to be created and thought through;
- Tests and evaluation criteria - how to verify the agent solved the task correctly.
- The work breaks down into two phases:
- Building the virtual company - following a high-level plan, you create a company: codebase, infrastructure, context (conversations, documentation, tickets). The result is a realistic environment with development history.
- Assembling and calibrating tasks - from intermediate states of the company, task drafts are assembled (based on a taxonomy). You turn a draft into a full task: write the prompt, define evaluation criteria, ensure the task is solvable and the evaluation is fair.
- Then an AI agent receives the task and tries to solve it the way a developer would - and you evaluate the result.
- Iterating with an AI agent on tests - they should catch real problems, not miss bad solutions, not break on good ones;
- Reviewing code written by agents;
- Analyzing why an agent failed or succeeded;
- Designing edge cases and adversarial scenarios.
Requirements
- You don't need to be an expert in every item, but you should be comfortable reading and reasoning about code across the stack - that's what lets you design realistic tasks, write meaningful tests, and evaluate results.
- Core tech stack: Backend: Python, FastAPI; Frontend: JavaScript/TypeScript, React; Infrastructure: Docker, Postgres, Kafka, Redis.
Conditions
- You start with task and dataset work, and if it goes well, we expect you to move into a project team - with a broader scope and deeper.