About the company
Perplexity's mission is to power curiosity. Perplexity runs one of the highest-throughput inference stacks in the industry, serving Ask, Computer, and API traffic across a large and constantly shifting portfolio of first-party and third-party models.
Responsibilities
- Execute the roadmap for the inference platform — request handling, rate limits and quotas, usage controls, and the reliability and observability surface engineering and product teams depend on.
- Be the connective tissue between model providers and Perplexity's engineering and product teams — coordinating onboarding, launch readiness, and rollout for new models and capacity.
- Drive latency, throughput, uptime, and cost-efficiency as core execution metrics, surfacing tradeoffs between them rather than letting them become side effects.
- Run the operating model for model-release and optimization programs, including day-zero launches, across performance engineering, infrastructure, and product teams.
- Lead cross-functional delivery for inference-stack changes, from planning through launch and post-launch validation.
- Build the mechanisms that make releases predictable — rituals, dashboards, launch checklists — so inference releases stay low-risk at Perplexity's scale.
- Partner with GPU capacity and compute teams to reconcile execution decisions against cost, capacity, and vendor constraints.
Requirements
- Strong experience with technical program management or product management in infrastructure, distributed systems, or ML/model-serving products.
- Direct experience with production LLM or ML inference — understanding what makes serving fast, reliable, and cheap rather than just what a roadmap slide says about it.
- Comfort orchestrating across external partners and internal engineering teams with competing priorities and timelines.
- Experience with data and metrics, and the judgment to surface difficult tradeoffs between latency, throughput, uptime, and cost.
- Thrives in a small, agile team; has initiative and desire for ownership without much precedent to lean on.
- 6+ years of combined technical program management or product management experience.