About the company
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.
Responsibilities
- Lead the investigation and resolution of complex infrastructure, networking, and platform-related incidents.
- Act as a senior escalation point for operational teams during critical service-impacting events.
- Support large-scale NVIDIA GPU infrastructure and high-performance networking environments.
- Troubleshoot complex Linux, Kubernetes, networking, storage, and hardware-related issues.
- Analyze platform performance, capacity, stability, and reliability trends to proactively identify risks.
- Lead root cause analysis activities and drive long-term corrective actions.
- Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve complex technical challenges.
- Participate in major incident management and service restoration activities.
Requirements
- 7+ years of experience
- Deep expertise in Linux, Kubernetes, networking, storage, and hardware
- Experience with NVIDIA GPU infrastructure
- Strong troubleshooting and incident management skills
- Technical leadership experience