Position Overview
Essnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy Research Scientific Computing Center (NERSC), a mission-critical high-performance computing (HPC) and data environment supporting scientific research for the U.S. Department of Energy (DOE) Office of Science.
The Site Reliability Engineer will work as part of a 24/7 operations environment responsible for maintaining the accessibility, reliability, security, and operational health of large-scale computing and data systems.
This is a highly hands-on position combining Linux systems administration, infrastructure monitoring, incident response, programming and scripting, automation, networking, ServiceNow, and physical data center operations.
Responsibilities
- Monitor high-performance computing systems, storage infrastructure, networks, and other data center and facility-related systems.
- Review and respond to infrastructure and system alerts, perform initial triage, and engage appropriate on-call personnel when escalation is required.
- Respond to alerts across multiple systems to help ensure monitoring and data collection remain operational 24/7.
- Troubleshoot system, application, network, monitoring, and infrastructure issues affecting system reliability.
- Develop solutions that improve operational processes, prevent recurring issues, and automate responses to routine service conditions.
- Identify opportunities to improve monitoring capabilities, alerting, incident triage, and operational automation.
- Develop and maintain tools within the monitoring pipeline in collaboration with operations personnel.
- Develop software and integrations capable of generating alerts and notifications from HPC system APIs into monitoring pipelines.
- Build and maintain application and tool configurations to ensure reliable operation as data volumes and user demands increase.
- Utilize ServiceNow to support incident management, trouble-ticketing, operational workflows, and service management activities.
- Collaborate across technical teams to identify and resolve operational bottlenecks and maintain system reliability.
- Coordinate with technical groups during center-wide maintenance activities.
- Manage diagnostic, monitoring, and notification software during planned maintenance periods.
- Perform regular physical and logical walkthroughs of the data center floor.
- Monitor environmental conditions, power distribution units (PDUs), cooling infrastructure, and other facility systems supporting reliable data center operations.
- Maintain accurate trouble-ticket documentation for outages, incidents, maintenance activities, troubleshooting actions, and operational updates.
- Analyze problems of varying complexity and evaluate technical data to determine appropriate troubleshooting and remediation methods.
- Exercise independent technical judgment when selecting methods and approaches for resolving operational issues.
Requirements
- 5+ years of relevant professional experience in Site Reliability Engineering, systems/infrastructure engineering, DevOps, data center operations, HPC operations, network/system operations, or a closely related technical environment.
- Strong hands-on experience working with Linux, including Linux shell and command-line environments such as SSH.
- Programming and/or scripting experience using one or more languages such as Python, C, C++, Perl, Java, or comparable scripting or programming languages.
- Knowledge of standard software development practices.
- Experience supporting large-scale IT infrastructure, highly available systems, data centers, critical installations, or comparable technical environments.
- Knowledge of large data communications networks and common network protocols.
- Network security experience, including knowledge of firewalls and access control lists (ACLs).
- Experience troubleshooting infrastructure, application, system, network, or operational issues.
- Experience responding to monitoring alerts and performing technical incident triage.
- Ability to analyze operational and system data to identify problems and determine appropriate solutions.
- Experience collaborating across multiple technical teams to resolve operational issues and maintain system reliability.
- Strong written and verbal communication skills.
- Ability to independently learn and apply new technologies in a complex technical environment.
Conditions
- Work Location: Onsite – California
- Schedule: Full-Time | 5 Days Per Week | Midnight–8:00 AM (Owl Shift)
- Compensation: $80.00 per hour
- This position does not offer sponsorship. Must be authorized to work in the United States.