About the company
Xsolla is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with Xsolla to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, Xsolla is resolute in the mission to bring opportunities together, and continually make new resources available to creators. Headquartered and incorporated in Los Angeles, California, Xsolla operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world.
Responsibilities
- Own the application-level infrastructure: Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, and service-level networking and integrations
- Own the domain's observability: design and implement SLOs/SLIs, monitors, alerts, and dashboards for critical services on Datadog and OpenTelemetry-based tooling
- Help to set up and evolve CI/CD pipelines for domain services (GitLab CI, GitHub Actions), including deploy and rollback automation
- Perform capacity planning and performance tuning ahead of expected load - product launches, sales events, and regional rollouts - including load testing and performance regression investigation
- Run Production Readiness Reviews for new services and major changes; define and enforce what "production-ready" means for the domain
- Support domain incident response: assist with deep investigation of complex incidents, contribute to post-mortems, drive follow-up reliability improvements, and maintain runbooks
- Build domain-specific automation that reduces operational toil: runbook automation, deploy helpers, recurring operational scripts
- Maintain and drive a forward-looking reliability roadmap for the domain together with product engineering leads
- Participate in product team planning, refinements, and architecture reviews, bringing the reliability perspective before design decisions become expensive to change
- Co-author company-wide SLO/SLI, capacity, and operational standards together with the broader SRE team; contribute improvements directly to shared SRE-operated subsystems
- Participate in the SRE duty rotation, supporting developers across the company
Requirements
- 3+ years of proven SRE, DevOps, or platform engineering experience: on-call or incident response duty, SLO/monitoring ownership, deploy pipeline and infrastructure work for production services
- Software development background: you have built and shipped backend services, not only operated them - comfortable reading application code during an investigation and writing production-quality automation in at least one language (e.g., Go, PHP)
- Hands-on Kubernetes experience: Helm, manifests, deploy strategies, debugging application-level performance and networking issues (GKE or another managed Kubernetes)
- Solid observability practice: building monitors, dashboards, and SLOs/SLIs on a modern platform (Datadog preferred; Prometheus/Grafana experience also relevant), familiarity with OpenTelemetry
- Infrastructure as Code exposure (Terraform/Terragrunt) for collaboration with platform teams
- GCP experience (IAM, networking, managed services)
- Experience building and maintaining CI/CD pipelines (GitLab CI and/or GitHub Actions)
- Programming/scripting proficiency sufficient to build automation and tooling (e.g., Python, Go, or Bash)
- Practical experience with incident response, post-mortems, and driving reliability improvements from incidents
- Strong collaboration and communication skills — this role works embedded with product development teams daily
- Experience in payments, fintech, e-commerce, or gaming — high-traffic transactional systems
Conditions
- Nice to Have: Kubernetes certifications, Google Cloud Platform certifications, HashiCorp certifications