About the company
ZoomInfo is where careers accelerate. We move fast, think boldly, and empower you to do the best work of your life. You’ll be surrounded by teammates who care deeply, challenge each other, and celebrate wins. With tools that amplify your impact and a culture that backs your ambition, you won’t just contribute. You’ll make things happen—fast.
Responsibilities
- Search Stack Orchestration: Architect, scale, and maintain our core search and NoSQL technologies, including Solr (indexing/sharding), HBase (distributed storage), and Redis (high-speed caching).
- Cloud Architecture & IaC: Use Terraform to manage multi-cloud environments (AWS/GCP), ensuring that the search stack and its supporting resources are fully versioned and reproducible.
- Kubernetes Mastery: Oversee the deployment of data services within Kubernetes, focusing on stateful sets, persistent storage performance, and resource isolation for search workloads.
- Evolution toward AI Search: Lead the infrastructure integration of Vector search databases and high performance compute to support AI-driven architectures.
- Observability & Reliability: Implement deep-stack monitoring and alerting using Datadog and Prometheus to ensure proactive issue detection and resolution.
- Automation-First Mindset: Maintain and evolve an active codebase in Python, Go, or Bash to automate repetitive tasks. You will also integrate LLMs into your workflow to accelerate scripting, documentation, and operational efficiency.
- Networking & Traffic: Manage complex cloud networking topologies, including VPCs, Load Balancing, Service Meshes, and Caching layers (e.g., Redis) to minimize latency.
- Incident Response: Lead the debugging of complex, distributed systems issues, performing root cause analysis to prevent recurrence, as part of an on-call rotation.
Requirements
- 7+ years of experience in Infrastructure, DevOps, or SRE, with a significant portion dedicated to managing high-scale distributed data systems.
- Direct experience with Solr, HBase, and Redis is highly preferred. Experience with similar technologies (e.g., Elasticsearch/OpenSearch, Cassandra, or BigTable) is acceptable if you have the senior-level depth to transition quickly.
- Expert-level knowledge of Linux performance tuning, specifically for data-intensive applications (I/O scheduling, memory management, and JVM tuning).
- Proven track record of managing large-scale infrastructure using Terraform or OpenTofu.
- Deep experience running production-grade stateful workloads on Kubernetes.
- Strong proficiency in Python or Go, with the ability to build custom tools and operators to manage data lifecycles.
- Practical experience using LLMs (GitHub Copilot, ChatGPT, Claude) to increase your personal and team productivity.
- Demonstrated ability to learn new technologies quickly and independently.
- Strong technical, organizational and interpersonal skills.
- Strong written and verbal communication skills.
- Must be able to read, understand, and communicate complex problems and solutions in English over a textual medium (such as Slack).