About the company
(No company description provided)
Responsibilities
- Participate in design reviews, sprint zero, and delivery planning to define and validate reliability requirements, including resiliency, observability, fault tolerance, performance, scalability, holiday and special-day processing, and disaster recovery.
- Collaborate with Major Release Management to ensure each release meets SRE standards for observability, resiliency, and reliability requirements, support readiness, and knowledge base coverage.
- Define and improve monitoring, observability, dashboards, telemetry coverage, and alert strategy to strengthen outage detection, reduce noise, improve signal quality, and accelerate incident response.
- Assist in major incident response and root cause analysis by identifying observability gaps, improving telemetry and knowledge articles, and driving actions that reduce repeat incidents.
- Drive automation, intelligent tooling, and AI-assisted remediation to reduce manual toil, improve consistency, accelerate recovery, and scale operational support.
- Serve as the operational readiness authority before production releases by validating reliability requirements, assessing support readiness, surfacing production risks, and confirming release supportability.
- Lead capacity, performance, workload trend, and resiliency analysis to ensure applications scale reliably under normal, peak, and stress conditions.
- Establish and track reliability metrics such as availability, incident volume, MTTx, alert quality, automation coverage, reliability requirement compliance, change failure rate, and repeat incident reduction.
- Participate in application reliability governance and service reviews by presenting incident trends, compliance metrics, operational risks, improvement actions, and readiness gaps.
- Prepare executive reporting on reliability posture, release readiness, observability maturity, alert quality, incident trends, automation progress, risks, and improvement outcomes.
- Promote SRE practices through mentoring, standards adoption, best-practice sharing, and approved AI tools that improve knowledge, observability, performance, security, and maintainability.
Requirements
- Minimum of 10+ years of related technical experience across application support engineering, software engineering, site reliability engineering, production support, or application operations.
- Bachelor’s degree preferred or equivalent practical experience.
- Experience supporting business-critical applications in production environments.
- SRE, observability, automation, or ITIL certifications are a plus.
- Proven experience in one or more in-scope roles including Application Support Engineer, SDET, Software Engineer, or SRE, with responsibility for improving reliability practices, application validation, observability coverage, automation frameworks, and reliability standards.
- Strong understanding of monitoring and observability platforms, including dashboard design, alert tuning, telemetry coverage, log analysis, metrics, traces, and event correlation.
- Programming or scripting proficiency in one or more languages such as Python, Java, Go, PowerShell, or similar for automation, tooling, and operational efficiency.
- Familiarity with distributed applications, middleware, messaging, batch processing, real-time processing, and production application behavior in high-availability environments.
- Experience in financial services, capital markets, regulated environments, or other high-availability operational settings.
- Demonstrated participation in disaster recovery, performance testing, resiliency testing, release readiness, incident response, and root cause analysis.
- Knowledge of AI concepts, data platforms, anomaly detection, incident correlation, and intelligent automation use cases.
- Strong collaboration skills across application support, application development, release management, risk, security, business, and vendor stakeholders.
- Ability to translate production support insights into actionable engineering improvements that reduce risk, improve stability, and enhance customer experience.