About the company
This is a fully remote opportunity focused on ensuring the stability, scalability, security, and performance of cloud-based SaaS products.
Responsibilities
- Support and continuously improve production applications running in AWS, ensuring high availability, reliability, performance, and operational stability.
- Monitor system health, availability, latency, performance, logs, metrics, traces, and alerts, proactively identifying and addressing reliability risks.
- Participate in incident response, troubleshooting, escalation management, root cause analysis, and post-incident improvement activities.
- Strengthen operational readiness, resiliency, disaster recovery capabilities, and overall production support processes.
- Engineer and optimize AWS environments using services such as EC2, ECS/EKS, Lambda, S3, RDS, CloudWatch, IAM, VPC, Elastic Beanstalk, Load Balancers, and related cloud technologies.
- Apply AWS Well-Architected Framework principles across reliability, security, performance efficiency, cost optimization, and operational excellence.
- Support cloud-native architecture decisions and contribute to infrastructure modernization and optimization initiatives.
- Build and enhance end-to-end observability through dashboards, monitoring, alerting, logging, metrics, and tracing.
- Automate recurring operational processes using Python, Bash, PowerShell, CI/CD pipelines, Terraform, and infrastructure-as-code practices.
- Support applications across both on-premises and AWS environments, contributing to modernization and migration initiatives.
- Partner with application development teams to identify performance bottlenecks, infrastructure constraints, security concerns, and reliability risks.
- Manage and improve containerized environments using Kubernetes, Docker, and Amazon EKS, including GitOps-based deployment approaches.
- Support CI/CD and release management processes using tools and methodologies such as Git, Azure DevOps, ArgoCD, and FluxCD.
- Troubleshoot APIs, microservices, network connectivity, HTTP-based applications, and cloud infrastructure to improve application performance and uptime.
- Manage and optimize database environments, including RDS configuration, schemas, users, performance troubleshooting, data integrity, and storage practices.
- Collaborate across Development, Business, Platform, and Infrastructure teams in Agile environments using Scrum and Kanban practices.
- Serve as a technical point of contact for infrastructure, reliability, and operational requirements within product teams.
- Continuously identify opportunities for reengineering, process improvement, efficiency gains, automation, and cloud optimization.
Requirements
- Bachelor's degree or an equivalent combination of education and professional experience.
- 5+ years of relevant industry or technical experience, with substantial hands-on experience in DevOps, cloud infrastructure, site reliability, or related engineering disciplines.
- Strong DevOps background, ideally with an emphasis on approximately 70% operations and 30% development activities.
- Extensive experience with AWS cloud infrastructure and cloud-based production environments.
- Strong experience with containerization and orchestration technologies, particularly Kubernetes, Docker, and Amazon EKS.
- Hands-on experience building and managing CI/CD pipelines using GitOps methodologies and tools such as ArgoCD or FluxCD.
- Strong knowledge of Git-based version control and experience with Azure DevOps, including tickets, releases, CI/CD pipelines, and infrastructure-as-code workflows.
- Strong Terraform experience and knowledge of infrastructure-as-code best practices.
- Advanced scripting and automation skills using Python, Bash, and/or PowerShell.
- Experience with monitoring and observability tools, particularly AWS CloudWatch, including alert configuration, troubleshooting, and escalation management.
- Strong understanding of AWS services including EC2, AMIs, S3, RDS, Lambda, ECS/EKS, IAM, VPC, Elastic Beanstalk, Load Balancers, and Transfer Family.
- Experience managing large-file transfers and supporting highly available cloud infrastructure.
- Strong knowledge of APIs and microservices, including configuration, performance tuning, security, and reliability best practices.
- Solid understanding of HTTP concepts and protocols, with the ability to analyze and optimize web application performance.
- Experience using diagnostic tools such as curl and wget to troubleshoot connectivity and network performance issues.
- Strong SQL and database administration skills, including RDS configuration, schemas, catalogs, users/logins, synonyms, performance troubleshooting, and database hygiene.
- Strong networking and network management capabilities.
- Experience with service mesh technologies such as Linkerd or Istio is a plus.
- Excellent written and verbal communication skills.