About the company
Riot Games was established in 2006 by entrepreneurial gamers who believe that player-focused game development can result in great games. In 2009, Riot released its debut title League of Legends to critical and player acclaim. As the most played PC game in the world, over 100 million play every month.
Responsibilities
- Own and continuously improve the reliability, availability, scalability, performance, and operational health of the Efficiency team’s (web) platform and the tools deployed within it
- Design, build, and maintain the infrastructure, deployment systems, and operational foundations required to support a growing portfolio of production AI services and internal tools
- Improve CI/CD pipelines, release engineering practices, environment management, and deployment automation so software can be shipped safely, quickly, and consistently
- Establish production-readiness standards and ensure new utilities have appropriate monitoring, alerting, ownership, documentation, rollback strategies, and support plans before launch
- Define and operationalize service health indicators, SLIs, SLOs, error budgets, and reliability metrics that guide engineering priorities and tradeoffs between reliability, velocity, cost, and complexity
- Build comprehensive observability across applications, infrastructure, service dependencies, and user workflows using metrics, logs, traces, dashboards, synthetic monitoring, and actionable alerts
- Establish sustainable incident-management and on-call practices, including escalation paths, runbooks, severity definitions, communication protocols, and clear service ownership
- Lead or contribute to the diagnosis and resolution of production incidents, coordinating across teams and driving blameless post-incident reviews and durable corrective actions
- Build automation that reduces operational toil, improves mean time to detect and recover, and eliminates recurring sources of failure or manual intervention
- Implement safe deployment patterns such as automated validation, progressive delivery, canary releases, feature flags, health checks, rollback mechanisms, and controlled environment promotion
- Perform capacity planning, load testing, performance analysis, and resource forecasting to ensure the Toolkit can support increasing adoption and usage across Riot
- Design and validate resilience, backup, recovery, failover, and disaster-recovery strategies for critical services, data, configurations, and infrastructure
- Identify single points of failure and systemic risks across applications, cloud infrastructure, networking, databases, queues, caches, third-party dependencies, and operational workflows
- Improve developer experience by building self-service workflows, reusable infrastructure components, local development environments, test environments, deployment tooling, and clear operational documentation
- Establish and maintain infrastructure-as-code, configuration-management, secrets-management, and environment-governance practices that make infrastructure changes safe, repeatable, and auditable
- Partner with engineers throughout the software development lifecycle to embed operational excellence