About Us
At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company.
At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in.
About the Role
Emerging Technologies & Incubation (ETI) builds and launches new products on Cloudflare’s global network. Within ETI, you’ll join the Storage Infrastructure team to build and operate a shared storage platform for the teams behind stateful products such as R2, Workers KV, and Durable Objects.
We manage the underlying storage hardware, distributed databases, and object storage clusters, including fleet lifecycle automation, capacity, hardware validation, failure-domain planning, observability, and production operations. In this role, you will build platform software and automation that makes these fleet operations safe, repeatable, and scalable.
Responsibilities
- Design and build automation and operator tooling for provisioning, configuring, expanding, upgrading, and decommissioning storage hardware, distributed databases, and object storage clusters.
- Engineer globally distributed storage fleets to tolerate hardware failures, network disruption, capacity pressure, and recovery events across multiple failure domains.
- Develop observability, alerting, and safety controls that help engineers understand fleet health and make auditable production changes.
- Diagnose incidents across Linux, storage, networking, and distributed systems; participate in on-call and implement durable fixes that reduce operational toil.
- Use AI tools throughout development and operations to accelerate debugging, incident triage, root-cause analysis, and toil reduction, while applying sound engineering judgment and verifying outputs before taking production action.
- Validate new storage hardware and characterize its performance during normal operation, failures, rebuilds, and other recovery conditions.
- Partner with R2, Workers KV, Durable Objects, Network, Capacity Planning, Performance Engineering, and Infrastructure Operations to turn service requirements into infrastructure capabilities.
Requirements
- Experience designing, building, and operating infrastructure platforms or large-scale production fleets, including automation of lifecycle operations.
- Strong Linux systems knowledge and experience troubleshooting across compute, storage, and networking layers.
- Experience operating distributed systems in production, including observability, incident response, reliability improvements, and safe change management.
- Ability to write maintainable software in at least one programming language, such as Go, Rust, or Python.
- Strong written and verbal communication skills, with experience collaborating across engineering teams and technical stakeholders.
Conditions
- Compensation may be adjusted depending on work location.
- For New York City, New Jersey, Washington, Washington DC, and California (excluding Bay Area) based hires: Estimated annual salary of $185,000 - $254,000
- This role is eligible to participate in Cloudflare’s equity plan.
- Cloudflare offers a complete package of benefits and programs to support you and your family. Our benefits programs can help you pay health care expenses, support caregiving, build capital for the future and make life a little easier and fun! The below is a description of our benefits for employees in the United States, and benefits may vary for employees based outside the U.S.
Health & Welfare Benefits
- Medical/Rx Insurance
- Dental Insurance
- Vision Insurance
- Flexible Spending Accounts
- Commuter Spending Accounts
- Fertility & Family Forming Benefits
- On-demand mental health support and Employee Assistance Program
- Global Travel Medical Insurance
Financial Benefits
- Short and Long Term Disability Insurance
- Life & Accident Insurance