About the Company
At Databricks, we are passionate about empowering data teams to tackle the world’s most challenging problems. We build and operate the world’s best data and AI infrastructure platform, enabling our customers to leverage deep data insights and enhance their business.
Responsibilities
- Lead platform incident investigation, coordinating cross-functional teams through rapid detection, mitigation, and resolution
- Conduct thorough post-incident root cause analysis across infrastructure, services, and cloud providers
- Design and implement customer-focused alerting pipelines and end-to-end observability workflows
- Build automation tools, establish reusable monitoring patterns, and resolve reliability gaps
- Provide mentorship to junior engineers on observability patterns, alert design, and service health metrics
- Participate in on-call rotation
Requirements
- Minimum of 6 years of experience as an SRE, DevOps Engineer, Production Engineer, or similar role
- Production-level experience with at least one major cloud provider (AWS, Azure, GCP) and proficiency in Docker, Kubernetes
- Hands-on experience with monitoring, logging, and alerting tools such as ELK, Prometheus, Grafana, PagerDuty
- Strong proficiency in Python or similar languages
- Experience owning critical phases of the incident lifecycle from detection through resolution and post-mortem analysis
- BS, Master's, or PhD in Computer Science, Computer Engineering, or related field