← Все вакансии/Senior/StradIT
SeniorHybridDallas

Site Reliability Engineer

S
StradIT
Уровень
Senior
Формат
Hybrid
О роли

Описание вакансии

About the company

(No company description provided)

Responsibilities
  • Participate in design reviews, sprint zero, and delivery planning to define and validate reliability requirements, including resiliency, observability, fault tolerance, performance, scalability, holiday and special-day processing, and disaster recovery.
  • Collaborate with Major Release Management to ensure each release meets SRE standards for observability, resiliency, and reliability requirements, support readiness, and knowledge base coverage.
  • Define and improve monitoring, observability, dashboards, telemetry coverage, and alert strategy to strengthen outage detection, reduce noise, improve signal quality, and accelerate incident response.
  • Assist in major incident response and root cause analysis by identifying observability gaps, improving telemetry and knowledge articles, and driving actions that reduce repeat incidents.
  • Drive automation, intelligent tooling, and AI-assisted remediation to reduce manual toil, improve consistency, accelerate recovery, and scale operational support.
  • Serve as the operational readiness authority before production releases by validating reliability requirements, assessing support readiness, surfacing production risks, and confirming release supportability.
  • Lead capacity, performance, workload trend, and resiliency analysis to ensure applications scale reliably under normal, peak, and stress conditions.
  • Establish and track reliability metrics such as availability, incident volume, MTTx, alert quality, automation coverage, reliability requirement compliance, change failure rate, and repeat incident reduction.
  • Participate in application reliability governance and service reviews by presenting incident trends, compliance metrics, operational risks, improvement actions, and readiness gaps.
  • Prepare executive reporting on reliability posture, release readiness, observability maturity, alert quality, incident trends, automation progress, risks, and improvement outcomes.
  • Promote SRE practices through mentoring, standards adoption, best-practice sharing, and approved AI tools that improve knowledge, observability, performance, security, and maintainability.
Requirements
  • Minimum of 10+ years of related technical experience across application support engineering, software engineering, site reliability engineering, production support, or application operations.
  • Bachelor’s degree preferred or equivalent practical experience.
  • Experience supporting business-critical applications in production environments.
  • SRE, observability, automation, or ITIL certifications are a plus.
  • Proven experience in one or more in-scope roles including Application Support Engineer, SDET, Software Engineer, or SRE, with responsibility for improving reliability practices, application validation, observability coverage, automation frameworks, and reliability standards.
  • Strong understanding of monitoring and observability platforms, including dashboard design, alert tuning, telemetry coverage, log analysis, metrics, traces, and event correlation.
  • Programming or scripting proficiency in one or more languages such as Python, Java, Go, PowerShell, or similar for automation, tooling, and operational efficiency.
  • Familiarity with distributed applications, middleware, messaging, batch processing, real-time processing, and production application behavior in high-availability environments.
  • Experience in financial services, capital markets, regulated environments, or other high-availability operational settings.
  • Demonstrated participation in disaster recovery, performance testing, resiliency testing, release readiness, incident response, and root cause analysis.
  • Knowledge of AI concepts, data platforms, anomaly detection, incident correlation, and intelligent automation use cases.
  • Strong collaboration skills across application support, application development, release management, risk, security, business, and vendor stakeholders.
  • Ability to translate production support insights into actionable engineering improvements that reduce risk, improve stability, and enhance customer experience.
Стек и навыки

С чем работаем