About the company
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
Responsibilities
- Design, build, and evolve CI/CD pipelines that support reliable and efficient build, test, and release workflows across the organization.
- Own and improve artifact lifecycle systems, including versioning, storage, distribution, dependency management, and reproducible builds at scale.
- Partner with development teams to design and improve code review workflows, branching strategies, and automated integration processes.
- Provision, monitor, and optimize cloud infrastructure supporting CI workloads, balancing cost, performance, scalability, and reliability.
- Troubleshoot complex build failures, pipeline bottlenecks, and infrastructure issues, driving root-cause analysis and implementing durable fixes.
- Design and drive improvements to internal build infrastructure, test infrastructure, developer tooling, and automation that increase developer velocity and engineering productivity.
- Identify systemic bottlenecks across developer workflows and lead architectural improvements to developer infrastructure.
- Contribute to the company’s efforts around AI tooling to improve engineering productivity and automate repetitive workflows.
- Provide technical leadership across projects, influence engineering standards, and mentor other engineers within the Developer Productivity organization.
- Participate in on-call and incident response for Developer Productivity systems and services.
Requirements
- 7+ years of professional experience in software engineering, infrastructure engineering, DevOps, developer productivity, or a related area.
- Deep hands-on experience with CI/CD systems and building or evolving automated build, test, and deployment infrastructure.
- Experience with artifact repositories, software packaging, dependency management, and reproducible build concepts.
- Experience with cloud computing platforms, with AWS preferred, and programmatic infrastructure provisioning.
- Strong experience with distributed version control systems, code review workflows, branching strategies, and repository management.
- Strong understanding of Linux/Unix systems, networking fundamentals, and scripting or programming for automation.
- Experience with containerization and container orchestration, with Kubernetes preferred.
- Strong troubleshooting skills and the ability to debug complex distributed systems and infrastructure issues.
- Demonstrated experience leading technical initiatives across multiple teams and driving improvements to developer infrastructure at organizational scale.
- Ability to identify architectural bottlenecks, evaluate tradeoffs, and drive long-term improvements rather than only addressing operational issues.
- Experience operating production systems and participating in on-call, incident response, and postmortem processes.
Preferred qualifications
- Experience with infrastructure-as-code tools and practices.
- Proficiency in Python, Go, Shell, or another language used for infrastructure automation and developer tooling.
- Experience with build systems, build graph optimization, or large-scale build infrastructure.
- Experience with observability practices including monitoring, logging, alerting, and performance analysis.
- Experience building internal developer platforms, self-service tooling, or developer-facing infrastructure.
- Experience applying AI/LLM tooling to engineering workflows or developer productivity.
- BS/MS in Computer Science or a related field, or equivalent practical experience.