About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability / Gitops Engineer based in Italy.
Responsibilities
- Drive the development of automation and GitOps practices within the team while acting as an embedded technical lead.
- Collaborate closely with the infrastructure architecture function to align technical solutions with broader architecture objectives.
- Design and architect infrastructure services that can be delivered as reusable products across the organization.
- Develop and strengthen Infrastructure as Code practices by continuously improving automation, processes, consistency, and reusability.
- Automate software operations across private and public clouds while accounting for the complexity and operational requirements of distributed systems.
- Maintain operational responsibility for core services, networks, and infrastructure, ensuring reliability and continuity.
- Troubleshoot complex infrastructure issues, support capacity planning, investigate performance challenges, and improve system resilience.
- Implement and maintain observability, monitoring, and alerting solutions using technologies such as Prometheus, Grafana, and Elasticsearch.
- Collaborate with globally distributed engineering, operations, and support teams to resolve technical challenges and deliver reliable services.
- Dedicate focused development time to larger engineering projects and the automation of repetitive or manual operational tasks.
- Share technical knowledge, experience, and best practices through design sessions, mentorship, collaborative implementation, and team development.
- Take final responsibility for resolving time-critical escalations and ensuring appropriate technical follow-through.
Requirements
- Strong understanding of modern hosting architectures and an automation-first approach based on Infrastructure as Code across private and public cloud environments.
- A product-oriented mindset with an interest in building reusable infrastructure products rather than one-off solutions.
- Professional experience with Python development, including work on large or complex projects.
- Hands-on experience with Kubernetes or other container orchestration technologies.
- Proven experience managing and deploying cloud infrastructure through code and automation.
- Practical knowledge of Linux networking, routing, firewalls, and related infrastructure concepts.
- Familiarity with Linux storage technologies, ranging from distributed storage such as Ceph to database-backed systems.
- Hands-on experience administering enterprise Linux servers.
- Strong knowledge of cloud computing concepts, architectures, and technologies.
- Bachelor’s degree or higher, preferably in Computer Science, Engineering, or a related technical discipline.
- Strong English communication skills across email, chat, video calls, voice communication, and in-person collaboration.
- Ability to troubleshoot problems across the technology stack, from the Linux kernel through application and web layers, while knowing when to seek input from others.
- Flexibility, curiosity, and the ability to learn new technologies and approaches quickly.
- Ability to adapt to fast-changing technical environments and remain focused on delivering reliable outcomes.
- Experience working effectively within globally distributed teams.
- Passion for open-source technologies, with familiarity with Ubuntu or Debian considered valuable.
Conditions
- Compensation shaped according to geographic location, experience, and performance, with regular compensation reviews.
- Performance-driven annual bonus or commission in addition to base compensation.
- Distributed work environment with twice-yearly in-person team sprints.
- USD 2,000 annual personal learning and development budget.
- Recognition and performance rewards.
- Annual holiday leave.
- Maternity and paternity leave.
- Team Member Assistance Program and Wellness Platform.
- Opportunities to travel internationally and meet colleagues in new locations.
- Priority Pass and travel upgrades for long-haul company events.
- Opportunity to work on large-scale cloud infrastructure, SRE, GitOps, Infrastructure as Code, automation, observability, and open-source technologies.
- Dedicated development time for larger technical initiatives and automation projects.
- Opportunities to provide technical mentorship and influence infrastructure engineering practices.