About the Company
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
Responsibilities
- Support and troubleshoot production bare-metal GPU infrastructure across Nebius data center environments.
- Diagnose high-priority issues across Linux, hardware, firmware, BIOS, drivers, networking, storage, optics, cabling, and physical infrastructure.
- Perform hardware diagnostics, component replacement, firmware updates, BIOS configuration, break/fix, and infrastructure maintenance.
- Use Linux command-line tools, system logs, BMC data, and hardware health signals to identify root cause.
- Partner with data center technicians and engineering teams to coordinate remote troubleshooting, incident response, RCA, and operational recovery.
- Troubleshoot network connectivity issues involving TCP/IP, VLANs, DNS, DHCP, switch connectivity, optics, cabling, and link-level failures.
- Create troubleshooting guides, runbooks, escalation notes, and knowledge base documentation.
- Use scripting or automation to improve support workflows and reduce manual intervention.
- Participate in on-call or after-hours support for production infrastructure as needed.
Requirements
- 5+ years of experience in data center operations, infrastructure support, systems administration, hardware support, cloud infrastructure, or similar technical environments.
- Hands-on experience troubleshooting bare metal servers, Linux systems, hardware components, and networking issues in production environments.
- Intermediate Linux command-line proficiency, including service validation, log review, process inspection, network configuration checks, and system diagnostics.
- Experience with server hardware diagnostics, component replacement, firmware updates, BIOS configuration, driver troubleshooting, or BMC tools such as iDRAC, iLO, IPMI, Redfish, or similar.
- Strong networking fundamentals, including TCP/IP, VLANs, DNS, DHCP, routing/switching concepts, optics, cabling, and link-level troubleshooting.
- Experience participating in incident response, escalation handling, root cause analysis, or operational recovery in high-availability environments.
- Strong communication and documentation skills, with the ability to work cross-functionally with data center, infrastructure, network, systems, and engineering teams.
Conditions
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams