← Все вакансии/Senior/jobgether
SeniorRemoteNetherlands

Staff Site Reliability Engineer

J
jobgether
Уровень
Senior
Формат
Remote
О роли

Описание вакансии

About the company

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer based in Netherlands.

Responsibilities
  • Architect, deploy, operate, and continuously improve scalable, secure production environments, with a strong preference for AWS-based infrastructure.
  • Lead reliability initiatives across multiple engineering streams and establish consistent SRE practices throughout the organization.
  • Design, evolve, migrate, and optimize Kubernetes-based infrastructure, including production hardening and scaling.
  • Establish and enforce robust Infrastructure-as-Code standards using Terraform or equivalent technologies.
  • Define, implement, and operationalize SLIs, SLOs, error budgets, and other reliability practices.
  • Strengthen observability across applications, infrastructure, data pipelines, and ML systems to improve visibility into system health and performance.
  • Collaborate with product and data teams to incorporate product telemetry, model analytics, and operational data into reliability insights.
  • Design and optimize CI/CD pipelines across the complete software lifecycle, from build and testing through deployment and rollback.
  • Improve release safety, deployment frequency, operational predictability, and adherence to service-level objectives.
  • Lead incident response for complex, cross-system failures and drive thorough post-incident reviews and corrective actions.
  • Reduce operational toil through automation, platform engineering, and improved tooling and processes.
  • Design scalable processes and infrastructure for absorbing, standardizing, monitoring, and troubleshooting customer environments.
  • Support and productionize ML workloads by implementing MLOps practices for model deployment, monitoring, and retraining workflows.
  • Ensure infrastructure and operational practices meet enterprise-grade security, compliance, and regulatory requirements.
  • Mentor engineers, share best practices, and raise the overall reliability and engineering standards across teams.
  • Collaborate with Staff Engineers and Architects to influence global product architecture and long-term technology strategy.
Requirements
  • Extensive hands-on experience in Site Reliability Engineering, Production Engineering, or a closely related infrastructure role.
  • Proven experience establishing or scaling SRE practices within high-growth, complex, or highly distributed technical environments.
  • Deep expertise with AWS or Azure cloud infrastructure and modern cloud-native architectures.
  • Strong production experience with Kubernetes, including migration, scaling, optimization, and security hardening.
  • Advanced Infrastructure-as-Code expertise using Terraform or an equivalent technology.
  • Demonstrated experience designing, implementing, and optimizing end-to-end CI/CD pipelines.
  • Strong knowledge of observability practices and tooling across distributed applications and infrastructure.
  • Experience troubleshooting complex multi-tenant, customer-hosted, or enterprise environments.
  • Experience supporting production data platforms and machine learning systems.
  • Practical MLOps experience, including model deployment, monitoring, and operational lifecycle management.
  • Strong understanding of distributed systems, scalability, resilience, fault tolerance, and failure modes.
  • Ability to think across systems and understand the interactions between infrastructure, applications, data, and ML workloads.
  • Strong communication and collaboration skills, with the ability to work effectively across engineering and business functions.
  • Experience with large-scale global B2B or B2C products is desirable.
  • Experience working with AI/ML, NLP, or LLM-based products is a strong advantage.
  • Familiarity with integrating product analytics and model performance metrics into operational monitoring is beneficial.
  • Experience operating in enterprise environments with stringent security, compliance, and regulatory requirements is preferred.
  • Experience implementing regulatory controls within cloud infrastructure is a plus.
  • Experience scaling infrastructure during periods of rapid growth is advantageous.
  • Experience evaluating infrastructure tools, platforms, and vendors is desirable.
  • Experience deploying and operating solutions within large enterprise customer accounts or VPCs is a strong plus.
  • Strong problem-solving skills, high ownership, and accountability.
  • Ability to anticipate failure modes, operate across multiple engineering streams, and influence technical decisions without formal authority.
  • Continuous-learning mindset with a strong commitment to improving systems, processes, and engineering practices.
Conditions
  • Full-time, permanent employment.
  • Fully remote position within European time zones.
  • Opportunity to work on infrastructure supporting advanced AI and agent-based workloads.
  • Significant technical
Стек и навыки

С чем работаем