About the Role
CoreWeave is seeking a highly skilled and motivated HPC Performance Engineer to join our HAVOCK Team, reporting into the Manager of Systems Engineering. In this role, you will play a crucial part in the design, development, and optimization of our bare-metal systems from POST through joining a Kubernetes cluster. The team’s primary responsibilities include maintaining a custom Linux kernel, various OS images (Ubuntu-based), the virtualization stack (kubevirt/qemu/vfio), and the container/pod runtime stack (containerd/nydus/kubelet). You will collaborate closely with cross-functional teams, up stack engineering teams, and stakeholders to ensure our low-level software stack is performant in the context of hardware updates; and providing data, metrics, dashboards, and analysis to substantiate performance assertions.
Responsibilities
- Develop and maintain tools for establishing systems performance baselines
- Develop and maintain performance regression analysis testing automation
- Design and maintain performance regression test pipelines for HPC workloads
- Debug and Tune fabric-level performance to ensure low-latency high throughput configurations
- Development of telemetry for performance analysis across distributed clusters of servers
- Triage and fix performance issues in Linux
- Collect data, produce metrics and visualizations that communicate performance information compared to benchmarks; this data should lead to appropriate business decisions and toward greater automation that improves customer experience in relation to performance
- Define Linux and OS requirements, specifications, and system architecture in relation to systems performance, in collaboration with cross-functional teams. Along with these responsibilities there will also be cross team collaboration to triage and resolve bottlenecks
Requirements
- 5+ years of professional experience in Systems/HPC Performance Engineering, Benchmarking, and/or Validation.
- Bachelor’s degree in Computer Engineering, Electrical Engineering, Computer Science, or a related field.
- Strong experience with MPI workloads and distributed system performance analysis
- Familiarity with RoCE, InfiniBand, and GPUDirect/Data Direct I/O, NUMA, etc in HPC workloads
- Hands-on use of public HPC benchmarks (HPCC, HPL, OSU, MLPerf-HPC, STREAM, IO500)
- Extensive, deep experience in Linux internals
- Fluency with a programming language geared toward automation (Python preferred, but others possible)
- Experience writing robust, testable code
- Experience diagnosing and fixing systems performance issues
- Experiencing with implementing automation testing
- Ability to effectively prioritize and communicate proposed features and fixes in a remote-employee environment
- Strong passion for automation, with a commitment to automating processes comprehensively
- Excellent documentation skills and attention to detail
- Strong analytical and problem-solving abilities
Conditions
- The base salary range for this role is $165,000 to $242,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location.
- In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).