← Все вакансии/Senior/jobgether
SeniorRemoteIndia

Staff DevOps Engineer

J
jobgether
Уровень
Senior
Формат
Remote
О роли

Описание вакансии

About the company

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Staff DevOps Engineer based in India.

This is a high-impact technical leadership role within a global Infrastructure Platform team responsible for operating large-scale, cloud-native SaaS infrastructure.

You will shape the architecture, reliability, security, and scalability of platforms running across AWS, Azure, and GCP.

A major focus will be Kubernetes leadership, including EKS, AKS, service mesh, GitOps, and cloud-native deployment patterns.

You will also help advance AI-assisted and agentic automation across DevOps and SRE workflows.

The role offers significant organizational influence through architecture decisions, technical standards, mentoring, and cross-functional leadership.

You’ll work with globally distributed engineering teams across India, the US, EMEA, and APAC in a fast-paced 24×7 environment.

This is an individual-contributor leadership position with substantial scope to improve reliability, automation, security, and engineering productivity.

Responsibilities
  • Lead cloud infrastructure strategy: Design, build, and evolve resilient, secure, scalable, and cost-efficient multi-account cloud infrastructure, primarily across AWS while supporting Azure and GCP environments.
  • Drive Kubernetes excellence: Serve as a technical authority for production Kubernetes, including EKS and AKS cluster architecture, upgrades, networking, storage, capacity planning, security, multi-tenancy, and troubleshooting.
  • Own service mesh capabilities: Design and operate production service mesh solutions, implementing secure service-to-service communication, mTLS, traffic management, observability, resilience, and progressive delivery.
  • Establish platform standards: Define and promote best practices for Kubernetes, cloud networking, IAM, secrets management, infrastructure governance, workload design, Helm, deployment patterns, and security controls.
  • Advance Infrastructure as Code and GitOps: Automate infrastructure provisioning, deployment, monitoring, incident response, and capacity management using Terraform, CI/CD, GitOps, and related platform engineering practices.
  • Strengthen reliability and operations: Improve SLIs, SLOs, error budgets, observability, runbooks, incident response, on-call practices, post-incident reviews, and systemic remediation across a 24×7 SaaS environment.
  • Support security and compliance: Help implement secure platform controls and maintain infrastructure aligned with requirements such as PCI DSS, including audit readiness, governance, and secure-by-design practices.
  • Champion AI-enabled operations: Introduce LLM-based tooling, AI coding assistants, and agentic workflows to improve infrastructure development, incident triage, root-cause analysis, deployment validation, compliance checks, and operational efficiency.
  • Build safe AI automation: Establish appropriate guardrails, observability, cost controls, and human-in-the-loop practices for production AI and agentic workflows, including solutions built with services such as Amazon Bedrock.
  • Provide technical leadership: Influence engineers across teams and geographies without direct management responsibility, drive consensus on architecture, contribute to critical escalations, and raise engineering standards.
  • Mentor engineering talent: Coach Staff, Senior, and mid-level engineers while documenting reusable patterns, sharing technical knowledge, and encouraging stronger engineering practices.
  • Lead strategic initiatives: Drive cross-functional projects that improve uptime, deployment velocity, cloud consistency, cost efficiency, operational toil, and engineering productivity.
Requirements
  • Experience: 12+ years working in 24×7 production operations and highly available SaaS or cloud environments, with prior experience as a technical lead or Staff+ individual contributor in a global engineering organization.
  • Cloud expertise: 5+ years of hands-on experience with multi-account AWS infrastructure, including AWS Organizations, Account Factory, guardrails, SCPs, landing zones, networking, IAM, and cross-account connectivity.
  • Kubernetes: 5+ years of production Kubernetes experience at scale, with deep expertise in EKS and/or AKS, cluster operations, networking, storage, security, workload management, and performance optimization.
  • Infrastructure as Code: 5+ years of Terraform experience managing infrastructure across multiple AWS accounts and regions.
  • CI/CD & GitOps: 5+ years designing and implementing CI/CD pipelines for Terraform, Kubernetes, and microservices, plus practical experience with GitOps platforms such as ArgoCD, Kargo, or Flux.
  • Programming & systems: Strong Python, Go, or similar programming skills combined with advanced shell scripting and solid knowledge of Linux, networking, distributed systems, and production troubleshooting.
  • Service mesh: Hands-on experience implementing and operating a production service mesh such as Istio, Linkerd, or AWS App Mesh.
  • Observability: Experience with monitoring and logging technologies such as Prometheus, Grafana, OpenSearch, or equivalent platforms.
  • SRE practices: Strong understanding of SLIs, SLOs, error budgets, incident management, observability, reliability engineering, and operational excellence.
  • AI & automation: Experience applying AI tools on AWS or equivalent platforms to improve engineering productivity, automation, or operational efficiency; experience with Amazon Bedrock, LLM agents, or agentic workflows is highly valuable.
  • Security & compliance: Experience designing secure cloud platforms and familiarity with regulated enterprise environments; knowledge of PCI DSS and related compliance practices is preferred.
  • Regional infrastructure: Experience supporting data sovereignty or regional cloud deployments is a plus.
Стек и навыки

С чем работаем