About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Sr. Staff Platform/Data Reliability Engineer, Databricks based in United States.
This is a senior technical leadership opportunity focused on making Databricks a reliable, secure, scalable, and operationally mature enterprise platform.
Responsibilities
- Own operational excellence for the Databricks platform, including monitoring, alerting, observability, incident response support, and production runbooks for data jobs and platform services.
- Define and maintain CI/CD, version control, deployment, and environment-promotion standards for Databricks workflows, jobs, notebooks, code packages, and infrastructure configuration.
- Establish reliable promotion practices across development, testing, and production environments.
- Design and maintain standards for job orchestration, cluster and compute policies, service-principal usage, workload isolation, and secure production execution.
- Develop reusable operational templates and onboarding patterns for new data domains, including logging conventions, job tagging, metadata capture, and support handoff processes.
- Partner with senior data engineering stakeholders to ensure ingestion and medallion architectures are observable, recoverable, cost-efficient, and secure in production.
- Collaborate with cloud and infrastructure teams to align Databricks configurations and usage with broader enterprise cloud standards across commercial and future government-hosted environments.
- Help implement technical controls for data segregation, access boundaries, auditability, and operational compliance within regulated environments.
- Monitor and improve platform health indicators such as job success rates, pipeline reliability, incident trends, cost efficiency, and environment drift.
- Establish clear platform standards, operational expectations, documentation, and support models that enable sustainable scaling.
- Mentor engineers developing platform expertise and help expand operational knowledge across the broader engineering organization.
Requirements
- 12+ years of relevant experience in data platform engineering, platform operations, site reliability engineering, or modern cloud data infrastructure.
- Hands-on experience operating Databricks or a comparable cloud data platform in production environments.
- Strong experience designing or managing CI/CD pipelines, environment promotion, version control, and deployment automation for data platforms and pipelines.
- Deep understanding of observability, monitoring, alerting, incident management, production support, and reliability engineering principles.
- Experience designing compute policies, workload isolation, service-principal models, and secure production execution patterns on cloud data platforms.
- Strong understanding of access control, auditability, security, and operational discipline in regulated or security-sensitive environments.
- Ability to collaborate effectively with cloud/infrastructure, security, data engineering, analytics, and other technical stakeholders.
- Databricks certification and/or demonstrated expertise with Delta Lake, Unity Catalog, Workflows, and Databricks Asset Bundles is preferred.
- Experience with infrastructure-as-code and platform automation in enterprise environments is a plus.
- Experience supporting commercial and government environments, or other segregated environments with different compliance and access requirements, is preferred.
- Experience in defense, aerospace, federal, or another highly regulated industry is advantageous.
- Strong communication, mentoring, documentation, and technical leadership skills, with the ability to establish standards and influence engineering practices.
Conditions
- Base salary range of $180,000–$270,000 per year, with actual compensation determined by factors including skills, experience, certifications, and work location.
- Full-time regular employees receive base pay within the listed range plus bonus, benefits, and equity.
- Temporary employees receive compensation within the listed range plus an applicable temporary benefits package after 60 days of employment.
- Opportunity to work on advanced, mission-critical technology in a highly technical and security-sensitive environment.
- Exposure to complex cloud data platforms and production systems spanning commercial and government-oriented requirements.
- Opportunities for technical leadership, mentoring, and long-term platform strategy ownership.
- Equal opportunity workplace committed to an inclusive environment and supporting reasonable accommodations.
- Employment offers are contingent on a cleared background and, where applicable, reference checks.