About Nebius
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
The role
We are building a high-performance AI inference platform for developer-native teams running latency- and cost-sensitive workloads at scale. We are looking for a Senior Sales Engineer to become a foundational technical partner to our customers and a force multiplier for Sales and Engineering.
Your responsibilities will include:
- Strategic Technical Discovery — Lead deep technical discovery with engineering teams and technical founders; Understand model requirements, traffic expectations, latency constraints, GPU economics, and system dependencies; Translate customer ambition into production-feasible architecture; Identify hidden technical risks early
- Commercial Acceleration — Partner tightly with Sales on strategic deals; Influence deal strategy through architectural clarity; Prevent misaligned commitments before engineering allocation; Increase PoC-to-production conversion by ensuring technical realism
- PoC Architecture & Validation — Define measurable success criteria (latency, TTFT, throughput, cost envelope); Classify workload complexity and required optimization depth; Align appropriate resources (ML Solution Architects, engineering, GPU capacity, etc.); Drive structured Go / No-Go decisions; Prevent uncontrolled customization or hidden R&D
- Pattern Recognition & Platform Leverage — Identify recurring configuration patterns across customers; Quantify demand for advanced optimizations (quantization, speculative decoding, etc.); Surface structured insights to Product and Engineering; Help evolve platform capabilities based on real workload data
We expect you to have:
- Deep understanding of AI inference systems and GPU-backed infrastructure
- Experience with LLM workloads and performance-sensitive environments
- Experience with inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM)
- Ability to reason about latency, throughput, cost, and architecture tradeoffs
- Strong customer presence with engineering-first organizations
- Comfort challenging assumptions and pushing back constructively
- Commercial awareness – you understand that engineering time is a strategic resource
Preferred Technical Stack
- Programming Languages – Python
- Frameworks and Libraries – vLLM, SGLang, TensorRT-LLM, OpenAI/Anthropic SDKs
- Frameworks for Agentic Pipelines : Langchain / Langsmith / smolagents / equivalent
- API and Web Frameworks – FastAPI, Flask
- MLOps and DevOps tools – Kubernetes (K8s), Docker, Git
- Cloud Platforms – AWS (SageMaker, Bedrock), GCP (Vertex AI), Azure (Azure ML)