About the company
Beamery's jobs, skills and tasks data platform helps organizations make informed decisions across Talent Lifecycle Management. Solutions power recruitment, mobility, upskilling, diversity, work architecture and workforce planning.
Responsibilities
- Design solutions, produce architectures and write RFCs for platform, reliability and infrastructure initiatives
- Maintain hands-on credibility by helping teams ship scalable, reliable and cost-effective services
- Partner with Product and Engineering Directors on planning large new platform initiatives
- Create and update the architectural vision
- Set and advocate company-wide standards for operational excellence, observability, reliability and incident response
- Take a whole-company view of major incidents, identifying recurring themes
- Own and evolve the platform and infrastructure tech radar
- Set the standard for operational ownership
- Coach and mentor Engineers up to Staff level
- Collaborate on customer solutions with non-technical stakeholders
- Stay connected to production through on-call
Requirements
- Proven track record of designing and delivering scalable, reliable cloud-based infrastructure and platform services
- Extensive experience running and supporting services in production, including on-call leadership and incident command
- Previous experience as an individual contributor in an Engineering leadership position (Staff+, Principal, Architect)
- Deep expertise managing Kubernetes production clusters at scale — cluster lifecycle and upgrades, resource optimization (autoscaling, quotas), security (RBAC, Network Policies), high availability and troubleshooting
- Expertise with Infrastructure as Code (IaC), particularly Terraform, and modern GitOps practices
- Strong software engineering foundations, experience with Go preferable; NodeJS is a plus
- Strong understanding of observability, SLOs, alerting and cost management
- Operational experience with Kafka, MongoDB, PostgreSQL, Elasticsearch, and Istio
- Familiarity with LLMOps is a plus (LiteLLM, MLflow)
- A FinOps mindset