About the Role
We at Innovaccer are looking for a Site Reliability Engineer-III to build secured modern healthcare cloud infrastructure and a massive data stack and aim to write everything as code.
Responsibilities
- Build and automate secure cloud infrastructure (Infrastructure as Code) with various pillars like Cost, Reliability, Scalability, Performance etc.
- Design and architect various domains of SRE.
- Build CI/CD stack collaborating across Dev and QA/Automation team and drive organization to new level of (daily/hourly) continuous delivery and deployment.
- Collaborate with Dev and QA team to bring given initiative to a closer, increased adoption of DevOps practices and tool chain.
- Apply strong analytical skills to understand production system metrics, drive change, optimize system utilization and drive cost efficiency.
- Ensure that the Platform is secured as per guidelines established by CISO. e.g., Secure against DDoS attacks by implementing WAF, Vulnerability and Patch management, install required security agents etc.
- Lead least privilege based RBAC for various production services and tool chains.
- Build and execute Disaster Recovery plan.
- Key stakeholder to participate in case of IR (Incident Response).
Requirements
- 6–9 years in production engineering, site reliability, or related roles.
- Solid hands-on experience with at least one cloud provider (AWS, Azure, GCP) with automation focus (certifications preferred).
- Strong expertise in Kubernetes and Linux.
- Proficiency in scripting/programming (Python required).
- Observability is very critical for the scale of our systems and ability to find insights/behavior, detect problem/failures. Looking for leads to drive this charter spanning across logs, metrics, mesh, tracing etc.
- Knowledge of CI/CD pipelines and toolchains (Jenkins, ArgoCD, GitOps).
- Familiarity with persistence stores (Postgres, MongoDB), data warehousing (Snowflake, Databricks), and messaging (Kafka).
- Exposure to monitoring/observability tools such as ElasticSearch, Prometheus, Jaeger, NewRelic, etc.
- Proven experience in production reliability, scalability, and performance systems.
- Experience in 24x7 production environments with process focus.
- Familiarity with ticketing and incident management systems.
- Security-first mindset with knowledge of vulnerability management and compliance.
- Excellent judgment, analytical thinking, and problem-solving skills.
- Ability to quickly identify and drive optimal solutions within constraints.
- Able to perform with cool head under pressure situations without taking any shortcuts.
- Collaboration with solid verbal and oral communication skills are very critical to this role. Strong cross-functional collaboration skills, relationship building skills, and ability to achieve results without direct reporting relationships.
- Self-motivated individual that possesses excellent time management and organizational skills.
- Strong sense of personal responsibility and accountability for delivering high quality work.