About the company
At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year.
The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us.
Responsibilities
- Designing, building and operating reliable, scalable infrastructure in AWS and other cloud platforms.
- Evolving our container orchestration (Kubernetes), Helm charts, deployment systems (Argo) and the guardrails around them to ensure engineers can safely deploy, run, understand and debug their services.
- Evolving our observability stack (Openmetrics, Datadog) to make insights into our systems as good and easy to understand as possible.
- Setting standards for our Infrastructure-as-Code codebases (Terraform).
- Creating and supporting developers with CI/CD pipelines (Buildkit, Github Actions), build tooling (Gradle, Bazel, npm/pnpm/bun, Go and Cargo) and artifact repositories (Artifactory, ECR, github).
- Improving our incident response tooling to ensure we can efficiently handle incidents when they occur.
- Crossing organizational boundaries to align with product teams and ship changes that impact the whole engineering org.
Requirements
- 10+ years of software engineering experience, site reliability engineering or infrastructure engineering experience, with significant time spent on production system infrastructure.
- Deep knowledge of cloud infrastructure, with hand-on AWS experience.
- Strong Infrastructure-as-Code skills, preferably with Terraform and a track record of creating safe and reusable infrastructure patterns. Clickops should be your tool of last resort.
- Solid software engineering skills and the ability to build automation and production tooling in a general-purpose programming language.
- Experience leading incident response, conducting blameless post-mortems or COE processes and executing on the follow ups.
- An independent mindset: you thrive in ambiguity, find the highest-leverage problems yourself, and don't wait to be told what to do.
- Sound judgment around infrastructure: you will make the right decisions around security, IAM, networking and change management.
- Excellent communication: you can write a clear design or incident document, explain operational tradeoffs, and align stakeholders across teams on platform-wide changes.
- A generally helpful attitude. Everyone should know you within no-time as the person who gets stuff unblocked, points them in the right direction and makes their work easier. The 90’s can keep their BOFH.
- A drive for impact: you're not satisfied until our systems are reliable and every engineer is confident operating what they own.
- Nice to have: Experience operating AWS at scale across multiple accounts, regions and availability zones; Experience with GitOps and continuous delivery tooling such as Argo CD; Experience with Identity and Access Management at scale, including OIDC; Experience designing disaster recovery, business continuity, and resilience testing programs; Experience improving cloud cost efficiency and resource utilization.
Conditions
- We come to the office 3 times a week to hang out, whiteboard, and ship together.
- We have the latest hardware and software, including frontier AI models on day one.
- We're agile, but not dogmatic. Teams decide how they work best.
- Massive leverage: every improvement you ship is multiplied across hundreds of engineers and thousands of deploys a day.
- Full ownership: design, implementation, rollout, and impact measurement: it's yours.
- You shape the foundation: the infrastructure patterns and operational standards.