About the company
Deliveroo's GenAI Platform team sits within Machine Learning Platform and builds shared infrastructure that helps DoorDash, Wolt and Deliveroo teams bring GenAI-powered products to production.
Responsibilities
- Lead the design of infrastructure that moves GenAI ideas from prototype to production
- Own and evolve the open-weights serving stack: real-time GPU endpoints, high-throughput batch inference, and fine-tuning (SFT/DPO/LoRA)
- Architect scalable, high-performance systems for model serving, batch inference, GPU autoscaling and fine-tuning
- Push the cost and latency frontier of GPU inference
- Build platforms supporting rapid experimentation while meeting production standards
- Partner with ML engineers, product engineers, data scientists and platform teams
- Set technical direction for the centralised GenAI platform
Requirements
- BSc, MSc or PhD in Computer Science or equivalent
- 5+ years of industry experience in software engineering
- Deep backend engineering fundamentals, especially in Python and distributed systems
- Track record of designing and owning production services, APIs, data pipelines or ML infrastructure at scale
- Experience operating systems in production: observability, debugging, reliability, incident response, performance/cost optimization
- Deep hands-on experience with LLM inference and/or fine-tuning of open-weight models in production
- Demonstrated technical leadership and mentoring
- Proficiency in using AI coding tools (Claude Code, Codex, Cursor)
Nice to Have
- Experience with LLM inference engines and serving frameworks (vLLM, SGLang, TensorRT-LLM)
- Experience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA)
- GPU performance work: multi-node/distributed inference, KV-cache/memory optimisation, quantisation (FP8/INT8/AWQ/GPTQ)
- Experience with Kubernetes, cloud infrastructure (AWS/GCP), GPUs, serverless/elastic GPU platforms (Modal)
- Experience with LLM gateways, model routing, vendor abstraction, or cost attribution
- Experience building developer platforms or self-serve infrastructure
- Experience building and deploying AI agents or MCP servers in production
- Experience with eval systems, LLM observability, tracing, RAG, search, or vector databases