About the company
Databricks is the Data and AI company. More than 20,000 organizations worldwide — including adidas, AT&T, Bayer, Block, Mastercard, Rivian, Unilever, and 70% of the Fortune 500 — rely on the Databricks Data + AI Platform to build and scale data and AI apps, analytics and agents.
Responsibilities
- Architect and Automate: Design and deploy production-grade infrastructure on AWS using Terraform or Pulumi.
- Orchestration: Manage and scale containerized workloads using AKS (Azure Kubernetes Service) or EKS, focusing on cluster security and resource efficiency.
- CI/CD Excellence: Architect robust deployment pipelines using GitHub Actions, managing both GitHub-hosted and self-hosted runners for specialized build requirements.
- Drive "Observable by Default" Frameworks: Create underlying infrastructure to ensure new internal applications are secure and have logging and metrics enabled by default.
- Tooling, Scripting & AI: Build internal CLI tools, AI plugins and automation scripts to streamline developer workflows and enhance operational efficiency.
- Partner Cross-Functionally: Collaborate with stakeholders across Security, Engineering, Infrastructure, and Support to deliver impactful projects with real business outcomes.
- Mentor and Document: Participate in code reviews, document solutions and failure triage playbooks, and mentor junior engineers on the platforms you own.
Requirements
- Software Engineering Expertise: 5+ years of production-level experience with a strong proficiency in Python (non-negotiable).
- IaC: Expert-level proficiency in Terraform (modules, state management) or Pulumi (Preferred).
- Cloud & Infrastructure Breadth: Hands-on experience with AWS (or Azure/GCP), Kubernetes, Docker and containerization concepts.
- Automation & Integration Mindset: Experience building and troubleshooting integrations between infrastructure, data pipelines, and observability platforms.
- CI/CD: Advanced knowledge of Github Actions, Github Runners.
- Strong Observability Mindset: Understanding of observability pillars (logging, metrics, tracing) and hands-on experience with tools like Datadog, Prometheus, or ELK.
- Distributed Systems: Proficiency in running systems through concepts like Kafka or messaging queues.
- Independent Execution: Ability to operate with minimal guidance, take ownership of ambiguous projects, and follow a vision set by tech leads to execute independently.
Conditions
- Pay range transparency: Zone 1 Pay Range: $159,900—$219,900 USD.
- Comprehensive benefits and perks.
- Commitment to diversity and inclusion.