← Все вакансии/Lead/jobgether
LeadRemoteUK

Head of SRE

J
jobgether
Уровень
Lead
Формат
Remote
О роли

Описание вакансии

About the company

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Head of SRE based in United Kingdom.

Responsibilities
  • Own the overall SRE strategy, standards, roadmap, and execution while embedding reliability and operational excellence across engineering teams.
  • Assess existing infrastructure, systems, and operational processes, identifying opportunities to improve, modernize, or rebuild them where necessary.
  • Architect, deploy, and maintain scalable, secure, highly available production environments across cloud infrastructure.
  • Define and implement Service Level Indicators, Service Level Objectives, uptime targets, and reliability standards.
  • Establish robust monitoring, alerting, logging, and observability practices to proactively identify and resolve production issues.
  • Design and improve incident management, root-cause analysis, escalation, and postmortem processes.
  • Build sustainable on-call structures and escalation frameworks that balance operational coverage with team effectiveness.
  • Automate software delivery and infrastructure workflows to improve release safety, consistency, and predictability.
  • Create reproducible environments and Infrastructure-as-Code provisioning frameworks using modern automation practices.
  • Drive improvements in system performance, availability, scalability, and overall production reliability.
  • Support the productionization and operational reliability of data platforms, machine learning workloads, and AI systems.
  • Partner with QA and engineering leadership to improve release quality, deployment processes, and production stability.
  • Ensure infrastructure and operational practices meet enterprise-grade security, compliance, and regulatory requirements.
  • Evaluate infrastructure tooling and vendors where appropriate and help establish scalable technology standards.
  • Hire, manage, mentor, and develop a strong team of SRE engineers while establishing a culture of ownership and continuous improvement.
  • Collaborate with enterprise customers and technical stakeholders when infrastructure must be deployed and operated within customer-controlled environments or VPCs.
Requirements
  • Proven hands-on experience in Site Reliability Engineering, Production Engineering, DevOps, or a closely related engineering discipline.
  • Strong expertise in cloud infrastructure, with AWS or Azure experience preferred.
  • Deep practical knowledge of Infrastructure as Code, particularly Terraform, and container orchestration technologies such as Kubernetes.
  • Experience establishing, scaling, or significantly maturing SRE practices within an engineering organization.
  • Demonstrated success improving system uptime, reliability, performance, and operational processes.
  • Strong understanding of CI/CD, developer experience, infrastructure automation, observability, and modern software delivery practices.
  • Experience designing incident response, on-call, escalation, and operational readiness frameworks.
  • Previous experience managing, mentoring, and developing an SRE engineering team.
  • Experience supporting data platforms and machine learning systems in production.
  • Practical MLOps experience, including model deployment, monitoring, and retraining workflows.
  • Strong understanding of enterprise security, compliance, and regulatory requirements within cloud environments.
  • Experience working with large-scale B2B or B2C products and high-growth technology environments is advantageous.
  • Exposure to AI, ML, LLM, or other advanced technology platforms is a plus.
  • Experience implementing regulatory controls within cloud infrastructure is desirable.
  • Strong communication and interpersonal skills, with the ability to influence engineering teams and senior stakeholders.
  • Practical, hands-on leadership style with a willingness to lead from the front and remain technically engaged.
  • Structured approach to solving ambiguity, building processes, and establishing scalable operational frameworks.
  • High ownership, accountability, and continuous improvement mindset.
Conditions
  • Fully remote working model within the European time zone.
  • Full-time opportunity with significant ownership and influence across engineering and infrastructure.
  • Opportunity to establish and lead an SRE function and shape reliability standards across the organization.
  • Direct exposure to product and engineering leadership and involvement in strategic technology decisions.
  • Opportunity to work on large-scale AI infrastructure and production machine learning systems.
  • Significant autonomy to review, improve, and rebuild infrastructure and operational processes.
  • Opportunity to hire, mentor, and develop a high-performing SRE team.
  • International, high-growth environment focused on advanced AI and agentic technology.
  • Opportunity to work with enterprise-scale systems and demanding security, reliability, and compliance requirements.
  • Strong emphasis on technical ownership, continuous learning, operational excellence, and innovation.
Стек и навыки

С чем работаем