About the company
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
Responsibilities
- Perform advanced firmware and hardware diagnostics on enterprise server platforms, including CPU, memory, PCIe devices, GPUs, storage subsystems, and power components
- Troubleshoot complex hardware failures using system logs, BMC/IPMI interfaces, BIOS diagnostics, and vendor-specific tooling
- Act as the primary escalation point for L1 and L2 technicians on high-impact hardware incidents
- Conduct structured root cause analysis and document findings to prevent repeat failures
- Own the full RMA lifecycle, including validation of failed components, warranty claim creation, vendor coordination, tracking, and resolution
- Interface directly with OEM vendors to escalate recurring hardware defects and drive corrective action
- Analyze hardware failure trends and report metrics such as repeat RMA rates and component reliability
- Develop and standardize diagnostic playbooks, troubleshooting workflows, and hardware validation procedures
- Validate replacement components prior to redeployment into production environments
- Collaborate cross-functionally with data center operations, procurement, and engineering teams to improve hardware lifecycle processes
- Contribute to reducing MTTR and improving fleet-wide reliability through process improvements and knowledge sharing
Requirements
- 5+ years of hands-on experience working with enterprise server hardware in a production data center environment
- Deep understanding of x86 server architecture, including CPUs, memory, PCIe devices, storage controllers, GPUs, and power subsystems
- Strong experience performing firmware and BIOS/BMC diagnostics and upgrades
- Advanced Linux command-line troubleshooting skills, including log analysis and hardware-level diagnostics
- Experience working with remote management interfaces such as IPMI, iDRAC, iLO, or equivalent
- Proven experience managing hardware RMA processes and working directly with OEM vendors
- Ability to conduct structured root cause analysis and document technical findings clearly
- Familiarity with hardware monitoring systems and failure trend analysis
- Strong ownership mindset and ability to operate independently in mission-critical environments
- High proficiency in spoken and written English
Nice to have:
- Experience performing board-level diagnostics and component-level repair (SMD rework)
- Familiarity with data center networking equipment and basic network troubleshooting
- Experience supporting GPU-dense or high-performance compute environments
- Valid driver’s license
Conditions
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
Compensation: $112,700.00 - $140,800.00 OTE (On Target Earnings)