About the company
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
Responsibilities
- Owning delivery of region and GPU cluster deployment projects from kickoff through production handover
- Coordinating the capacity acceptance procedure: validating infrastructure availability, quota allocations, monitoring/alerting, and documented sign-off before a region goes to customers
- Coordinating acceptance testing across engineering teams (compute, storage, networking, GPU) and clarifying test ownership between them
- Acting as the real-time escalation point during live deployments — resolving blockers such as hardware failures, provisioning issues, or capacity conflicts by engaging the right on-call teams
- Defining and maintaining clear contracts and responsibilities between the teams involved in delivery to reduce cross-team ambiguity and friction
- Driving continuous improvement and automation of repetitive deployment and acceptance-testing work
- Running structured, high-signal status communications so stakeholders always know where a deployment stands
Requirements
- Proven experience leading technical projects or engineering teams, ideally in infrastructure, cloud, or datacenter environments
- Demonstrated ability to coordinate across many engineering teams simultaneously and drive alignment without formal authority
- Calm, decisive communication under pressure, including during live incidents
- Strong bias toward clear ownership and documentation over ad hoc firefighting
- Fluent English
Conditions
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams