About the Company
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
Responsibilities
- Acting as the main technical interface for customers running workloads on Nebius GPU infrastructure.
- Supporting customers in deploying, configuring, and tuning GPU-based environments for performance and reliability.
- Investigating and resolving complex issues spanning hardware, networking, operating systems, and cluster-level behavior.
- Partnering with internal teams (data center operations, networking, platform engineering) to coordinate and drive resolution of customer-impacting issues.
- Converting customer requirements into practical architectures, configurations, and execution plans.
- Identifying opportunities to improve system performance, stability, and overall customer experience.
- Developing and maintaining technical documentation, including solution patterns, troubleshooting guides, and operational best practices.
- Contributing to continuous improvement by surfacing recurring issues, gaps, and optimization opportunities to internal teams.
Requirements
- Experience in a customer-facing technical role (e.g., solutions engineer, support engineer, technical account manager, or similar).
- Strong understanding of GPU infrastructure, including NVIDIA-based systems, multi-node environments, and performance considerations.
- Hands-on experience with Linux systems and system-level troubleshooting.
- Familiarity with large-scale compute environments such as GPU clusters, AI infrastructure, or supercomputing systems.
- Ability to diagnose issues across hardware, networking, and software layers.
- Strong analytical and problem-solving skills, with the ability to operate effectively in high-pressure situations.
- Excellent communication skills, with the ability to explain complex technical concepts clearly to customers.
- A proactive, ownership-driven approach with a strong focus on customer success.
Conditions
- Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan: up to 4% company match with immediate vesting.
- Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.
- Remote work reimbursement: up to $85/month for mobile and internet.
- Disability & life insurance: company-paid short-term, long-term and life insurance coverage.
- Competitive salaries ranging from $180K to $220K OTE, which includes base salary and performance bonus. Equity in the form of RSUs may be available at certain salary grades.