About the company
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025.
Responsibilities
- Lead a team of engineers responsible for Kubernetes infrastructure running on bare metal
- Set clear goals, priorities, and execution plans for the team, and ensure reliable delivery against them
- Partner with senior ICs and adjacent teams on the roadmap for cluster lifecycle management, upgrades, reliability, observability, and infrastructure automation
- Improve the team's operational excellence across incident response, on-call health, root-cause analysis, and service ownership
- Drive engineering best practices for safe change management, testing, rollout quality, and production readiness
- Support the design and operation of platform capabilities for provisioning, patching, upgrades, scaling, and troubleshooting of Kubernetes clusters
- Build strong cross-functional relationships with compute, networking, storage, security, and product stakeholders
- Hire, coach, and develop engineers while creating a high-accountability, high-trust team culture
- Establish and improve mechanisms for planning, prioritization, execution tracking, and continuous improvement
- Help translate complex platform and infrastructure work into clear business and customer value
Requirements
- Experience managing an infrastructure, platform, or SRE-oriented engineering team
- Strong technical depth in Kubernetes, distributed systems, and production infrastructure
- Experience operating Kubernetes in complex environments, ideally including bare metal, hybrid, or highly performance-sensitive systems
- Familiarity with cluster lifecycle management, including provisioning, upgrades, node operations, observability, and reliability engineering
- Track record of improving team execution, engineering quality, and operational maturity
- Experience leading incident response cultures and driving follow-through on reliability improvements
- Strong partnership skills across engineering, product, and operations functions
- Ability to coach engineers at different levels and create clarity in ambiguous or fast-scaling environments
- Strong written and verbal communication, including the ability to explain technical trade-offs and priorities clearly
Conditions
- Competitive compensation and benefits package