← Все вакансии/baseten
OfficeSan Francisco

Solutions Architect

B
baseten
Формат
Office
О роли

Описание вакансии

About the company

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital.

Responsibilities
  • Partner with Sales on customer discovery calls (most often second calls, occasionally first calls for large accounts).
  • Lead demos and technical scoping to align on success criteria, architecture, and deployment approach.
  • Own benchmarking and repeatable deployments, including handling standard deployment patterns and configurations across many modalities – LLMs, embeddings, image and video generation, Voice AI, etc.
  • Advise on tradeoffs like H100s vs B200s and latency-optimized vs throughput-optimized setups.
  • Drive consistent "playbook" style deployments for common models and use cases.
  • Become a power user of different runtimes such as vLLM, SGLang, and TRT-LLM and all the common configurations and tradeoffs between them.
  • Drive POC and project execution, including scoping POCs and keeping stakeholders aligned on timeline, deliverables, and next steps.
  • Act as the "ringleader" or project manager for POCs.
  • Pull in Forward Deployed Engineering (FDE) support when deeper or more complex technical work is needed.
Requirements
  • AI/ML background and the ability to credibly discuss AI/ML topics with technical stakeholders.
  • Strong customer-facing communication skills, including the ability to run structured discovery and clarify ambiguous requirements.
  • Technical depth to scope solutions, without needing to write production code.
  • Ability to script and prototype as needed, including comfort "vibe coding" to move quickly in technical workflows.
Nice to have
  • Experience running or supporting benchmarks for ML inference deployments.
  • Familiarity with infrastructure tradeoffs relevant to inference performance and cost (for example GPU selection and latency versus throughput tuning).
  • Experience serving as a cross-functional technical lead for customer POCs, including coordination across Sales and Engineering.
Conditions
  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents.
  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!).
  • Paid parental leave.
  • Fertility and family-building stipend through Carrot.
  • Company-facilitated 401(k).
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Стек и навыки

С чем работаем