About the company
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
About the product
In a rapidly evolving world, trust in AI depends on AI agents being grounded in fresh, verified real-world data. Search is the foundation that makes this possible. We are building an agent-native search platform designed specifically for AI systems rather than human users. Our product provides programmatic, low-latency, and observable search APIs that AI agents use to retrieve, filter, and reason over real-world information at scale. Behind every search request is a rich stream of signals — query patterns, retrieval decisions, crawling outcomes, ranking quality, usage, and revenue events. Turning that stream into a trustworthy, queryable data platform is what makes the product improvable, the business measurable, and the models trainable.
Responsibilities
- Contribute to the design, development, and operation of Tavily's data platform — from real-time ingestion through data warehouse medallion layers to consumer-facing datasets and dashboards.
- Build and maintain reliable batch and streaming pipelines that ingest data from production services and external systems.
- Design and evolve scalable, analytics-ready data models in the data warehouse.
- Work closely with engineers across the company to ensure data produced by production systems is reliable, well-structured, and usable downstream.
- Improve observability across the data platform, including data quality checks, freshness monitoring, lineage, schema evolution, and cost controls.
- Partner with researchers, engineers, analysts, finance, and product managers to deliver trustworthy datasets for product, search quality, ML, and GTM analytics.
- Contribute to defining the objects, entities, and relationships that represent Tavily's search domain — including agent inputs, URLs, chunks, agent sessions, crawls, and the connections between them — and translate them into clean, queryable data models.
- Improve engineering practices around testing, documentation, deployment, and incident response.
- Investigate and resolve production data issues, including broken pipelines, corrupted datasets, schema changes, and large-scale backfills.
- Contribute to technical standards and best practices for data engineering across the company.
- Help maintain high standards of data quality, integrity, security, and governance across environments.
Requirements
- Have 5+ years of Data Engineering experience, with strong experience designing and implementing scalable, analytics-ready data models and cloud data warehouses such as Snowflake or BigQuery.
- Have hands-on experience with Snowflake, or a comparable cloud data warehouse, and a strong understanding of modern data warehouse architecture, preferably including medallion-style modeling.
- Have deep knowledge of databases, including schema design, query optimization, and familiarity with NoSQL use cases.
- Have strong experience with modern data orchestration and transformation frameworks such as Airflow and dbt.
- Understand cloud data services on AWS or GCP and have experience with streaming platforms such as Kafka or Pub/Sub.
- Have hands-on experience with Spark, MapReduce, or similar distributed processing systems, and understand when distributed processing is the right tool.
- Are fluent in Python and SQL for production data work.
- Have operated data systems in production: debugged them under pressure, recovered from data incidents, handled schema changes, and backfilled corrupted or incomplete datasets.
- Care deeply about data quality and about making datasets understandable and trustworthy for the people using them.
- Are comfortable working on ambiguous, cross-functional data problems and collaborating closely with both technical and non-technical stakeholders.
Conditions
- Competitive compensation.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.