About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SRE Engineer based in Brazil.
This is an opportunity for an SRE professional to strengthen the reliability, resilience, and performance of critical digital environments.
Responsibilities
- Define, implement, and monitor SLIs, SLOs, SLAs, MTTR, and MTTD, establishing measurable reliability objectives.
- Implement and evolve observability solutions, including monitoring, alerting, dashboards, and APM.
- Monitor system latency, traffic, errors, saturation, availability, and overall application and infrastructure performance.
- Prevent, investigate, and resolve incidents, contributing to effective incident response and service restoration.
- Conduct root-cause analyses and establish corrective and preventive actions to avoid recurring incidents.
- Identify infrastructure risks, bottlenecks, single points of failure, and opportunities to strengthen system resilience.
- Support the design and evolution of resilient, scalable, and highly available solutions.
- Automate operational activities and reduce repetitive manual work and operational toil.
- Operate and continuously improve Kubernetes and Docker environments.
- Support capacity planning, business continuity, and disaster-recovery strategies.
- Participate in deployments and help stabilize applications and environments after releases.
- Partner with development, infrastructure, security, and other teams to incorporate reliability practices from the earliest stages of solution design.
- Create and maintain operational dashboards, alerts, procedures, runbooks, and technical documentation.
- Promote a culture centered on reliability, observability, automation, and continuous improvement.
Requirements
- Professional experience working as a Site Reliability Engineer (SRE) or in an equivalent reliability, DevOps, or infrastructure engineering role.
- Practical experience with cloud environments, particularly GCP, AWS, and/or Azure.
- Solid knowledge of Kubernetes and Docker.
- Experience with observability, monitoring, alerting, and APM solutions.
- Strong understanding of SRE concepts and metrics, including SLI, SLO, SLA, MTTR, MTTD, and error budgets.
- Experience managing, investigating, and resolving production incidents.
- Strong troubleshooting skills across applications and infrastructure.
- Experience administering Linux environments.
- Knowledge of networking, security, performance optimization, scalability, and high availability.
- Experience with automation and Infrastructure as Code (IaC).
- Experience working with CI/CD pipelines and modern software delivery practices.
- Strong analytical and problem-solving capabilities, with a proactive approach to preventing issues before they affect production.
- Excellent communication skills and the ability to collaborate effectively with multidisciplinary engineering teams.
- Differentials: experience with GKE, EKS, or AKS; Dynatrace, Datadog, Grafana, Prometheus, ELK, Elasticsearch, or Kibana; Terraform and Ansible; distributed and mission-critical systems; regulated or financial environments; cloud capacity and cost optimization; disaster recovery and business continuity; and cloud, Kubernetes, or SRE certifications.
Conditions
- Meal allowance (Vale Refeição).
- Food allowance (Vale Alimentação).
- Home office allowance.
- Medical insurance.
- Dental insurance.
- Life insurance.
- Birthday day off.
- TotalPass / Wellhub wellness benefit.
- Access to the Boon Saúde health platform.
- Discounts and partnerships with businesses and educational institutions.
- Welcome kit.
- Structured onboarding program.
- Access to continuous learning through Verity Learning.
- Internal initiatives focused on knowledge sharing and professional development.
- Programs and initiatives supporting employee well-being and connection.