About the company
This position is listed on behalf of a partner company, who manages all applications and next steps.
Responsibilities
- Architect, build, and optimize highly available, multi-region cloud infrastructure
- Drive the transition from manually configured environments toward standardized, globally scalable infrastructure managed entirely through code
- Design and implement automated, self-healing infrastructure using Infrastructure-as-Code, GitOps practices, modern CI/CD pipelines
- Establish and continuously improve platform reliability, resiliency, availability, and operational standards
- Build comprehensive observability and telemetry capabilities using Prometheus, Grafana, and OpenTelemetry
- Own incident management activities, lead blameless post-mortems
- Develop reusable platform capabilities and "Golden Paths"
- Partner with Product, Security, and Core Engineering teams
- Mentor engineers and influence technical standards
- Architect and operate high-performance, low-latency private connectivity
- Apply security as a core engineering consideration
- Take end-to-end ownership of complex distributed-systems challenges
Requirements
- 8+ years of experience owning outcomes within complex, large-scale distributed systems
- Advanced AWS expertise
- Strong Infrastructure-as-Code experience, particularly with Terraform
- Extensive hands-on experience with Kubernetes, particularly EKS, as well as Docker
- Strong knowledge of GitOps workflows and tools such as Flux, Argo, and GitHub Actions
- Strong programming and automation capabilities using Python, Go, Bash, or comparable languages
- Deep experience implementing and operating observability solutions using Prometheus, Grafana, OpenTelemetry
- Solid understanding of cloud security, API gateways, load balancing, network isolation
- Demonstrated ability to lead complex technical initiatives independently
- Strong communication, collaboration, and mentoring skills
- Experience with tokenization, payment processing, cryptography, security products is a plus
- Bachelor’s degree in a relevant discipline is preferred
- Experience with distributed data streaming platforms such as Kafka or Amazon MSK is a plus
- Experience with database performance tuning and query optimization is beneficial
- Familiarity with Java and the Spring Framework is advantageous
- Must be based in an eligible U.S. hiring location; certain roles may also be available in select Canadian provinces
Conditions
- Competitive annual salary ranging from $145,000 to $260,000
- Early equity opportunities
- Medical, dental, and vision insurance
- HSA and FSA options
- Life and disability insurance
- Pet insurance
- Unlimited paid time off
- Paid national holidays
- Global parental leave program
- Internet stipend
- Learning and professional development stipend
- Flexible working hours
- Remote and hybrid work opportunities
- Opportunities for internal promotion and career growth
- Lunch-and-learn sessions, team events, and company summits
- Weekly office lunches and employee discount programs