← Все вакансии/Senior/CoreWeave
SeniorOfficeNew York, NY / Sunnyvale, CA / Bellevue, WA

HPC Performance Engineer

C
CoreWeave
Зарплата
$13,800–$20,200
Уровень
Senior
Формат
Office
О роли

Описание вакансии

About the Role

CoreWeave is seeking a highly skilled and motivated HPC Performance Engineer to join our HAVOCK Team, reporting into the Manager of Systems Engineering. In this role, you will play a crucial part in the design, development, and optimization of our bare-metal systems from POST through joining a Kubernetes cluster. The team’s primary responsibilities include maintaining a custom Linux kernel, various OS images (Ubuntu-based), the virtualization stack (kubevirt/qemu/vfio), and the container/pod runtime stack (containerd/nydus/kubelet). You will collaborate closely with cross-functional teams, up stack engineering teams, and stakeholders to ensure our low-level software stack is performant in the context of hardware updates; and providing data, metrics, dashboards, and analysis to substantiate performance assertions.

Responsibilities
  • Develop and maintain tools for establishing systems performance baselines
  • Develop and maintain performance regression analysis testing automation
  • Design and maintain performance regression test pipelines for HPC workloads
  • Debug and Tune fabric-level performance to ensure low-latency high throughput configurations
  • Development of telemetry for performance analysis across distributed clusters of servers
  • Triage and fix performance issues in Linux
  • Collect data, produce metrics and visualizations that communicate performance information compared to benchmarks; this data should lead to appropriate business decisions and toward greater automation that improves customer experience in relation to performance
  • Define Linux and OS requirements, specifications, and system architecture in relation to systems performance, in collaboration with cross-functional teams. Along with these responsibilities there will also be cross team collaboration to triage and resolve bottlenecks
Requirements
  • 5+ years of professional experience in Systems/HPC Performance Engineering, Benchmarking, and/or Validation.
  • Bachelor’s degree in Computer Engineering, Electrical Engineering, Computer Science, or a related field.
  • Strong experience with MPI workloads and distributed system performance analysis
  • Familiarity with RoCE, InfiniBand, and GPUDirect/Data Direct I/O, NUMA, etc in HPC workloads
  • Hands-on use of public HPC benchmarks (HPCC, HPL, OSU, MLPerf-HPC, STREAM, IO500)
  • Extensive, deep experience in Linux internals
  • Fluency with a programming language geared toward automation (Python preferred, but others possible)
  • Experience writing robust, testable code
  • Experience diagnosing and fixing systems performance issues
  • Experiencing with implementing automation testing
  • Ability to effectively prioritize and communicate proposed features and fixes in a remote-employee environment
  • Strong passion for automation, with a commitment to automating processes comprehensively
  • Excellent documentation skills and attention to detail
  • Strong analytical and problem-solving abilities
Conditions
  • The base salary range for this role is $165,000 to $242,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location.
  • In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).
Стек и навыки

С чем работаем