About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer, Ads based in United States.
As a Staff Site Reliability Engineer, you will provide technical leadership for reliability across a large-scale advertising technology ecosystem.
Responsibilities
- Lead reliability initiatives across multiple advertising technology domains, including ad serving, auctions, targeting, reporting, measurement, and billing.
- Partner with engineering leadership to define and execute a roadmap focused on reliability, scalability, operational excellence, and developer productivity.
- Design and build platforms, tooling, automation, and engineering solutions that improve system resilience and developer efficiency at scale.
- Lead architecture reviews and influence technical decisions affecting critical, revenue-generating distributed systems.
- Participate in on-call rotations, lead complex production investigations, and coordinate cross-functional responses to major incidents.
- Identify systemic reliability risks and drive durable engineering solutions that strengthen platform resilience.
- Establish and monitor reliability metrics for critical customer and advertiser journeys, including campaign creation, ad delivery, auction participation, reporting, attribution, and billing.
- Improve operational maturity through SLOs, automation, incident management, observability, and performance optimization.
- Troubleshoot complex issues across modern distributed system stacks and drive root-cause analysis toward long-term improvements.
- Mentor engineers and provide technical leadership across multiple teams and initiatives.
- Influence roadmap and investment decisions by ensuring reliability requirements are incorporated into product and infrastructure planning.
Requirements
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or a related discipline, with experience operating large-scale distributed systems.
- Strong experience evolving and supporting high-traffic, user-facing production environments.
- Deep expertise in distributed systems, scale engineering, cloud-native architectures, and highly available system design.
- Strong software engineering capabilities in a general-purpose backend language such as Go.
- Solid understanding of observability practices and technologies, including metrics, logging, tracing, and alerting.
- Proven experience improving reliability through SLOs, automation, incident management, performance optimization, and operational best practices.
- Demonstrated ability to diagnose and resolve complex issues across modern distributed technology stacks.
- Strong cross-functional leadership skills, with the ability to influence technical direction and drive operational improvements across multiple teams.
- Excellent written and verbal communication skills and the ability to collaborate effectively with engineering leadership and technical stakeholders.
- Experience with Kubernetes, cloud infrastructure, and large-scale distributed systems is highly valued.
- Experience with advertising technology or other revenue-critical platforms is a plus, particularly ad serving, real-time auctions, budget pacing, campaign delivery, measurement, attribution, or billing systems.
- Experience operating high-QPS, low-latency services where performance directly affects business outcomes is advantageous.
- Familiarity with large-scale data technologies such as Kafka, ClickHouse, Spark, Flink, or BigQuery is a plus.
- Experience partnering with Product, Data Science, and advertising engineering teams is beneficial.
- Exposure to machine learning inference or recommendation systems operating at scale is also valued.
Conditions
- Comprehensive health benefits, including medical, dental, and vision coverage.
- 401(k) program with employer matching.
- Equity compensation in the form of restricted stock units, subject to the position offered.
- Flexible vacation policy and global company days off.
- 4+ months of paid parental leave.
- Family planning support.
- Workspace benefits and support for a home office.
- Personal and professional development funds.
- Paid volunteer time off.
- Flexible-first workforce approach with remote-friendly work arrangements.
- Reasonable accommodations available for qualified candidates with disabilities and disabled veterans.
- Base salary range of $217,000–$303,900 USD, with final compensation determined by factors such as skills, depth of experience, relevant credentials, level, and location.
- Certain positions may also include additional variable compensation depending on the role.