← Все вакансии/jobgether
RemoteBrazil

SRE Engineer

J
jobgether
Формат
Remote
О роли

Описание вакансии

About the company

Our partner is looking for a SRE Engineer based in Brazil.

Responsibilities
  • Define, monitor, and continuously improve reliability indicators, including SLIs, SLOs, SLAs, MTTR, MTTD, and error budgets.
  • Implement and evolve observability, monitoring, alerting, and APM solutions across applications and infrastructure.
  • Monitor latency, traffic, errors, saturation, availability, and overall system performance.
  • Prevent, investigate, and resolve incidents, minimizing their impact on users and business operations.
  • Conduct root cause analyses and define corrective and preventive actions to avoid recurring incidents.
  • Identify operational risks, bottlenecks, single points of failure, and opportunities to strengthen system resilience.
  • Support the design and evolution of highly available, scalable, resilient, and fault-tolerant solutions.
  • Automate operational activities and reduce repetitive manual tasks through scripting, automation, and Infrastructure as Code.
  • Operate and continuously improve Kubernetes and Docker environments.
  • Support capacity planning, business continuity, disaster recovery, and cloud cost optimization initiatives.
  • Participate in deployments and contribute to application stabilization following releases.
  • Collaborate with engineering, development, infrastructure, and other technical teams to incorporate reliability from the earliest stages of solution design.
  • Create and maintain operational dashboards, alerts, procedures, runbooks, and technical documentation.
  • Promote a culture centered on reliability, observability, automation, prevention, and continuous improvement.
Requirements
  • Proven professional experience as a Site Reliability Engineer, SRE, or in an equivalent reliability/platform engineering role.
  • Practical experience with cloud environments, using one or more of GCP, AWS, or Azure.
  • Hands-on knowledge of Kubernetes and Docker.
  • Experience implementing and managing observability, monitoring, alerting, and APM solutions.
  • Strong understanding of SRE concepts and metrics, including SLI, SLO, SLA, MTTR, MTTD, and error budgets.
  • Experience managing, investigating, troubleshooting, and resolving production incidents.
  • Knowledge of application and infrastructure troubleshooting in complex environments.
  • Experience administering Linux environments.
  • Understanding of networking, security, performance, scalability, and high availability.
  • Experience with automation and Infrastructure as Code practices.
  • Hands-on experience with CI/CD pipelines and modern software delivery practices.
  • Strong communication and collaboration skills, with the ability to work effectively across multidisciplinary teams.
  • Analytical, proactive, collaborative, and prevention-oriented mindset.
  • Experience with GKE, EKS, or AKS is a plus.
  • Knowledge of Dynatrace, Datadog, Grafana, Prometheus, or comparable observability platforms is a plus.
  • Experience with ELK Stack, Elasticsearch, and Kibana is desirable.
  • Knowledge of Terraform and Ansible is desirable.
  • Experience supporting critical systems and distributed architectures is an advantage.
  • Experience in financial institutions or other regulated environments is a plus.
  • Experience with cloud capacity management and cost optimization is desirable.
  • Knowledge of disaster recovery and business continuity practices is beneficial.
  • Cloud, Kubernetes, or SRE certifications are considered a plus.
Conditions
  • Meal and food allowance.
  • Home office allowance.
  • Medical insurance.
  • Dental insurance.
  • Life insurance.
  • Birthday Day Off.
  • TotalPass / Wellhub access.
  • Health and wellness support through the Boon Saúde app.
  • Discounts and partnerships with a variety of establishments.
  • Partnerships with educational institutions and other services.
  • Welcome kit.
  • Structured onboarding program.
  • Access to continuous learning and professional development initiatives.
  • Dedicated learning and knowledge-sharing programs.
  • Employee support and engagement initiatives.
  • Fully remote work opportunity.
Стек и навыки

С чем работаем