About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Database Reliability Engineer based in the United States.
This is a high-impact engineering role responsible for the reliability, performance, availability, and cost efficiency of production databases powering a modern healthcare technology platform.
Responsibilities
- Own the reliability, performance, availability, and operational health of production databases running on AWS Aurora MySQL across EHR, Data, and AI workloads.
- Manage database observability end to end, maintaining the metrics and alerting pipeline into Datadog and integrating after-hours database alerts into the DevOps on-call rotation.
- Establish automated safeguards to identify and terminate long-running or runaway queries and provide immediate visibility into database activity across instances.
- Investigate database performance issues by analyzing execution plans, optimizing queries, re-indexing where appropriate, and reducing unnecessary database load and latency.
- Review application and EHR queries before production release, serving as a performance gate to prevent inefficient workloads from reaching production.
- Educate engineering teams on database hygiene, query optimization, and practices that improve reliability and performance.
- Create and maintain comprehensive database runbooks so first responders can resolve incidents quickly and consistently.
- Architect and continuously optimize Aurora reader topologies and read-routing strategies based on actual workload requirements.
- Own database cost efficiency through instance right-sizing, reserved-capacity and Savings Plan strategies, and appropriate storage tiering.
- Manage replication health, backups, restore testing, failover procedures, and disaster recovery capabilities.
- Establish safe, repeatable standards for database schema changes and migrations across engineering teams.
- Partner with platform architecture stakeholders to evaluate and determine the best technical approach for new and existing database workloads.
- Handle protected health information responsibly while maintaining database operations within a HIPAA-compliant environment.
Requirements
- 6+ years of experience in database reliability engineering, database administration, database engineering, or a closely related discipline, including ownership of production systems at scale.
- Deep expertise in MySQL, including query optimization, execution-plan analysis, indexing strategies, and replication; hands-on Aurora MySQL experience is strongly preferred.
- Experience operating large, multi-reader Aurora or RDS clusters at terabyte scale, including read-routing and connection-management strategies.
- Strong knowledge of database observability and monitoring technologies such as Datadog, Percona Monitoring and Management (PMM), Performance Insights, Prometheus, or Grafana.
- Proficiency with Python, Bash, or a comparable scripting language for automation, along with practical experience using infrastructure-as-code tools such as Terraform.
- Strong AWS operational knowledge, particularly RDS/Aurora, database sizing, reserved capacity, storage options, and cloud cost optimization.
- Experience with production on-call rotations, incident response, troubleshooting, and creation of operational runbooks.
- Strong understanding of database reliability, security, backup, recovery, failover, and disaster-recovery practices.
- Excellent communication and collaboration skills, with the ability to educate and influence engineers across multiple teams.
- Experience with cloud data warehouses such as Snowflake, Databricks, or Redshift is a plus, particularly where analytical workloads can be moved away from OLTP databases.
- Experience working with HIPAA-regulated or other compliance-heavy environments involving sensitive data is advantageous.
- Familiarity with automated query-remediation approaches such as Percona Toolkit, pt-kill, statement timeouts, or custom query termination systems is beneficial.
- Application-side experience, particularly with PHP, is a plus for effective collaboration on query and application performance.
- Knowledge of database and cloud security best practices, agile methodologies, and fast-paced engineering environments is valued.
Conditions
- Competitive salary and compensation package.
- Remote/hybrid working environment.
- Potential equity compensation based on outstanding performance.
- Flexible PTO.
- Company-sponsored lunches.
- Company-paid disability and life insurance.
- Company-paid family and medical leave.
- Medical, dental, and vision insurance.
- Discounted pet insurance.
- FSA/DCA and commuter benefits.
- 401(k) retirement plan.
- Credits toward online fitness classes and gym memberships.
- Access to an HQ recovery suite featuring a cold plunge, sauna, and shower.
- Opportunity to work on meaningful healthcare technology and solve complex infrastructure challenges.
- Collaborative environment focused on smart, sustainable work.