← Все вакансии/Senior/Essnova Solutions, Inc.
SeniorOfficeBerkeley

Site Reliability Engineer

ES
Essnova Solutions, Inc.
Зарплата
$100
Уровень
Senior
Формат
Office
О роли

Описание вакансии

Position Overview

Essnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy Research Scientific Computing Center (NERSC), a mission-critical high-performance computing (HPC) and data environment supporting scientific research for the U.S. Department of Energy (DOE) Office of Science.

The Site Reliability Engineer will work as part of a 24/7 operations environment responsible for maintaining the accessibility, reliability, security, and operational health of large-scale computing and data systems.

This is a highly hands-on position combining Linux systems administration, infrastructure monitoring, incident response, programming and scripting, automation, networking, ServiceNow, and physical data center operations.

Responsibilities
  • Monitor high-performance computing systems, storage infrastructure, networks, and other data center and facility-related systems.
  • Review and respond to infrastructure and system alerts, perform initial triage, and engage appropriate on-call personnel when escalation is required.
  • Respond to alerts across multiple systems to help ensure monitoring and data collection remain operational 24/7.
  • Troubleshoot system, application, network, monitoring, and infrastructure issues affecting system reliability.
  • Develop solutions that improve operational processes, prevent recurring issues, and automate responses to routine service conditions.
  • Identify opportunities to improve monitoring capabilities, alerting, incident triage, and operational automation.
  • Develop and maintain tools within the monitoring pipeline in collaboration with operations personnel.
  • Develop software and integrations capable of generating alerts and notifications from HPC system APIs into monitoring pipelines.
  • Build and maintain application and tool configurations to ensure reliable operation as data volumes and user demands increase.
  • Utilize ServiceNow to support incident management, trouble-ticketing, operational workflows, and service management activities.
  • Collaborate across technical teams to identify and resolve operational bottlenecks and maintain system reliability.
  • Coordinate with technical groups during center-wide maintenance activities.
  • Manage diagnostic, monitoring, and notification software during planned maintenance periods.
  • Perform regular physical and logical walkthroughs of the data center floor.
  • Monitor environmental conditions, power distribution units (PDUs), cooling infrastructure, and other facility systems supporting reliable data center operations.
  • Maintain accurate trouble-ticket documentation for outages, incidents, maintenance activities, troubleshooting actions, and operational updates.
  • Analyze problems of varying complexity and evaluate technical data to determine appropriate troubleshooting and remediation methods.
  • Exercise independent technical judgment when selecting methods and approaches for resolving operational issues.
Requirements
  • 5+ years of relevant professional experience in Site Reliability Engineering, systems/infrastructure engineering, DevOps, data center operations, HPC operations, network/system operations, or a closely related technical environment.
  • Strong hands-on experience working with Linux, including Linux shell and command-line environments such as SSH.
  • Programming and/or scripting experience using one or more languages such as Python, C, C++, Perl, Java, or comparable scripting or programming languages.
  • Knowledge of standard software development practices.
  • Experience supporting large-scale IT infrastructure, highly available systems, data centers, critical installations, or comparable technical environments.
  • Knowledge of large data communications networks and common network protocols.
  • Network security experience, including knowledge of firewalls and access control lists (ACLs).
  • Experience troubleshooting infrastructure, application, system, network, or operational issues.
  • Experience responding to monitoring alerts and performing technical incident triage.
  • Ability to analyze operational and system data to identify problems and determine appropriate solutions.
  • Experience collaborating across multiple technical teams to resolve operational issues and maintain system reliability.
  • Strong written and verbal communication skills.
  • Ability to independently learn and apply new technologies in a complex technical environment.
Conditions
  • Work Location: Onsite – California
  • Schedule: Full-Time | 5 Days Per Week | Midnight–8:00 AM (Owl Shift)
  • Compensation: $80.00 per hour
  • This position does not offer sponsorship. Must be authorized to work in the United States.
Стек и навыки

С чем работаем