About the company
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster.
The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems.
Responsibilities
- Keep user-facing services and production systems reliable, scalable, and efficient
- Build automation and tooling that reduces toil and replaces manual work with repeatable, infrastructure-as-code-driven workflows
- Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling
- Write and maintain infrastructure as code, and ship changes safely through CI/CD and GitOps
- Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately
- Contribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early rather than just outages
- Take part in incident response and post-incident reviews, turning learnings into changes in automation and process
- Document runbooks, architecture decisions, and reviews so your findings become repeatable practices
Requirements
- Experience keeping production systems reliable, combining an operations mindset with real software engineering practice
- Experience building net-new infrastructure tooling and automation, not just configuring existing tools. For example, Terraform modules, Kubernetes operators or controllers, or production automation and services written from scratch
- The ability to read, debug, and reason about code. Most of our teams work in Go; some work in Ruby. You can discuss a piece of code's behavior, performance, and failure modes
- Experience with infrastructure as code, and with Kubernetes and its ecosystem, at a depth appropriate for the level you're applying for
Conditions
- Remote work opportunities
- Flexible work hours
- Professional growth and development opportunities
- Inclusive and diverse work environment