← Все вакансии/Senior/Databricks
SeniorOfficeCosta Rica

Site Reliability Engineer

D
Databricks
Уровень
Senior
Формат
Office
О роли

Описание вакансии

About the company

Databricks is the Data and AI company. More than 20,000 organizations worldwide — including adidas, AT&T, Bayer, Block, Mastercard, Rivian, Unilever, and 70% of the Fortune 500 — rely on the Databricks Data + AI Platform to build and scale data and AI apps, analytics and agents.

Responsibilities
  • Architect and Automate: Design and deploy production-grade infrastructure on cloud platforms (AWS/Azure) using Infrastructure as Code (IaC) tools like Terraform or Pulumi.
  • Reliability and Performance Engineering: Optimize system performance, architecture, and scaling to ensure maximum uptime and minimal latency for critical IT services.
  • CI/CD Excellence: Architect robust deployment pipelines (e.g., GitHub Actions), managing both hosted and self-hosted runners for specialized build requirements.
  • Observable by Default: Create underlying infrastructure to ensure new internal applications are secure and have logging, metrics and alerts enabled by default.
  • Agentic Tooling: Build internal AI plugins, and automation scripts to streamline developer workflows and enhance operational efficiency.
  • Incident Response: Focus on subsequent data usage, incident management workflows, and creating necessary dashboards to maintain service health. Participate in a shared on-call rotation, leading rapid incident response and technical troubleshooting for production outages. Facilitate blameless post-mortems to identify root causes and implement permanent preventive engineering solutions.
  • Partner Cross-Functionally: Collaborate with Security, Engineering, and Support teams to deliver real business outcomes.
Requirements
  • Software Engineering Expertise: 5+ years of production-level experience with strong proficiency in Python (non-negotiable).
  • Infrastructure as Code (IaC): Expert-level proficiency in Terraform (modules, state management) or Pulumi.
  • Cloud & Containers: Hands-on experience with AWS, Azure, or GCP, along with Kubernetes, Docker, and containerization concepts.
  • Observability Mindset: Deep understanding of observability pillars (logging, metrics, tracing) and experience with tools such as Datadog, Prometheus, or ELK.
  • Distributed Systems: Proficiency in running systems using concepts like Kafka or messaging queues.
  • CI/CD Proficiency: Advanced knowledge of GitHub Actions and GitHub Runners.
  • Independent Execution: Ability to take ownership of ambiguous projects, follow a vision set by tech leads, and execute independently with minimal guidance.
Conditions
  • Comprehensive benefits and perks.
  • Commitment to diversity and inclusion.
Стек и навыки

С чем работаем