About the company
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
Responsibilities
- Create technical content demonstrating how to effectively use computing workloads with VMs, GPU clusters, k8s, SLURM, Soperator, etc.
- Develop sample code, tutorials, and reference architectures showcasing best practices for cloud computing and ML infrastructure
- Create video tutorials and live coding sessions demonstrating effective use of Nebius cloud infrastructure
- Collaborate with academic partners (universities, e.g., Stevens, MIT) to understand their requirements and develop solution architectures that align with their needs: design and document Infrastructure as Code solutions, documentation, and technical how-to guides in collaboration with the Nebius Solutions Architect Team
- Act as a trusted advisor to our academic partners, providing technical expertise on GPU cloud technologies and best practices
Requirements
- Strong understanding of cloud infrastructure and distributed computing principles
- Experience with virtual machines, containerization, and managing compute resources
- Experience building with IaC solutions, preferably Terraform
- Knowledge of GPU clusters and techniques for optimizing ML workloads
- Working knowledge of container orchestration systems like Kubernetes and job schedulers like SLURM
- Familiarity with infrastructure components including networking, storage optimization, and resource management
- Experience optimizing performance of diverse workloads in cloud environments
- Strong programming skills, particularly in Python, and familiarity with the PyTorch ecosystem
- Understanding of cloud infrastructure concepts and deployment patterns
- Excellent written communication skills and ability to clearly express technical ideas in text
- 2+ years of experience in software development, cloud engineering, DevOps, or a similar technical role
- Demonstrated experience with cloud technologies and infrastructure
- Previous work with infrastructure-as-code, containerization, and cloud environments
Conditions
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams