← Все вакансии/Senior/ServiceNow
SeniorFlexibleSanta Clara

Staff Reliability Engineer

S
ServiceNow
Уровень
Senior
Формат
Flexible
О роли

Описание вакансии

About the company

ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow—helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together.

Responsibilities
  • Design, build, and operate cloud-native engineering platforms for software validation, release validation, and production readiness
  • Design and maintain production-like release and test ServiceNow environments that improve release confidence and deployment readiness
  • Build and integrate automated test pipelines, observability, reliability signals, deployment intelligence, and quality gates into CI/CD workflows
  • Develop automation solutions that improve engineering productivity, streamline operations, and reduce manual toil through shift-left engineering practices
  • Build reusable frameworks, self-service engineering environments, test data management, mock services, and developer productivity tooling
  • Design and enhance Kubernetes-based platforms supporting scalable test infrastructure, release automation, cloud-native workloads, and developer self-service
  • Implement automated validation for failure detection, deployment verification, policy enforcement, security checks, resilience testing, and operational health assessments
  • Resolve complex platforms, infrastructure, and networking challenges through software engineering, systems design, and automation
  • Partner closely with engineering teams to improve platform reliability, release quality, cloud-native adoption, and engineering best practices
  • Participate in architecture reviews, technical design discussions, and implementation of scalable, automation-first engineering solutions
  • Influence technical decisions through strong engineering execution, collaboration, and delivery of high-quality platform capabilities
  • Mentor engineers through technical guidance, code reviews, knowledge sharing, and engineering best practices
  • Foster a culture of reliability, automation, operational excellence, continuous improvement, and customer-focused engineering
Requirements
  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving
  • 8+ years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, Software Engineering, or Infrastructure Engineering with a Bachelor's degree; or 6 years and a Master's degree; or a PhD with 3 years experience; or equivalent experience
  • Hands-on experience with Kubernetes across cluster operations, networking, storage, security, autoscaling, and multi-cluster environments
  • Experience building and operating cloud-native platforms supporting scalable, highly available services
  • Experience integrating Kubernetes with CI/CD, GitOps, automated test pipelines, deployment validation, and cloud-native deployment workflows
  • Experience designing and implementing automation to improve developer productivity
Стек и навыки

С чем работаем