About the company
Roblox is building the tools and platform that empower a global community of developers and creators. The Infrastructure Compute Site Reliability Engineering team owns and manages the underlying cell infrastructure system.
Responsibilities
- Design and develop systems & libraries that promote fault-tolerance and resilience, automate cluster management, and ensure observability.
- Promote and institute reliability best practices across the Infra Compute group.
- Build, automate and standardize process automation to create a "golden path" of tooling and platform support.
- Create tooling that provides production guardrails, evaluating release candidate capacity with load testing.
- Create performance monitoring services and observability for capacity issues and platform degradations.
- Analyze systems and system designs for production readiness.
Requirements
- Bachelor's degree in Computer Science or related engineering field with at least 6 years as an SRE or Software Engineer.
- 3+ years of professional experience with high-level programming languages like Go, Java, C#.
- Experience with Kubernetes or similar orchestration systems. Experience in Nomad, Vault, and Consul is strongly desired.
- Experience building software and tools and getting them adopted.
- Strong problem-solving, planning, and collaboration skills.
Conditions
- Based in San Mateo, CA Headquarters.
- Onsite Tuesday, Wednesday, and Thursday, with optional presence on Monday and Friday.
- Annual Salary Range: $196,750 - $243,290 USD.
- Eligible for equity compensation and benefits.