About the Company
Sezzle is a fintech company that provides interest-free installment plans, revolutionizing the shopping experience. We are building an innovative team to shape the future of fintech and retail.
About the Role
We are seeking a Senior Site Reliability Engineer. This role presents an exciting opportunity to thrive in a dynamic, fast-paced environment. As a Senior SRE, you will have a high degree of autonomy and authority to identify and resolve problems in our infrastructure, deployments, and operational workload.
Responsibilities
- Architect, upgrade, design, and build scalable infrastructure solutions leveraging Kubernetes, AWS, RDS (MySQL/Postgres), and modern distributed patterns.
- Help drive the infrastructure team’s roadmap, leading us to higher levels of reliability, recoverability, and scalability.
- Drive capacity planning, benchmarking, and work with the team to stress test our systems, find bottlenecks, and prepare for further growth.
- Define, maintain and enforce SLAs and alerts across our infrastructure.
- Lead the teams towards stronger signal anomaly detection, better, more flexible alerting.
- Help lead Sezzle’s AI enablement efforts, identifying opportunities to apply AI and automation to enhance infrastructure reliability, developer productivity, and internal tooling.
- Build in consistency and scalability across a distributed microservices architecture while maintaining performance and reliability.
- Establish and evolve engineering best practices for observability, security, and CI/CD across teams.
- Mentor engineers and champion a culture of learning, innovation, and operational excellence.
- Collaborate cross-functionally to translate business goals into technical roadmaps and deliver results that matter.
Requirements
- 6+ years of professional software engineering or infrastructure engineering experience, including significant SRE and backend experience.
- Deployed significant changes to a production application or infrastructure configuration in the past 30 days.
- Strong proficiency in Golang, with experience building and maintaining RESTful APIs.
- Expertise with SQL-based RDBMS (MySQL, PostgreSQL) and experience optimizing schema and queries for performance at scale.
- Proficiency in observability tools (Prometheus, Grafana, Datadog, New Relic).
- Solid understanding of distributed systems design patterns (e.g., transactional outbox, event-driven architecture and stream processing, queues).
- Demonstrated ability to bring new ideas forward, influence decisions, and lead complex technical initiatives.
- Demonstrated experience working with Claude or equivalent large language model tools is required.
- Bachelor's degree in Computer Science, Engineering, or a related technical field.
- Preferred: Experience with AWS cloud infrastructure, mainly AWS Aurora RDS, both MySQL and Postgres.
- Preferred: Experience with data engineering, data pipelines and data warehousing.
- Preferred: Experience with CI/CD pipelines and deploying containerized microservices in Kubernetes.
- Preferred: Familiarity with AI developer tooling like Claude Code, Gemini CLI, Codex, Cursor.
- Preferred: Track record of shipping commercial APIs and data-driven applications in high-growth environments.
- Preferred: Proven leadership in guiding technical direction, improving system reliability, and scaling high-traffic services.