About the company
Pleo builds spend solutions that make managing money seamless, empowering, and surprisingly effective for finance teams and employees alike. With over 40,000 customers and a decade of unique spend data, we're committed to delivering the future of business spending.
Responsibilities
- Build and ship multiple AI-powered product features, setting the bar for delivery at Pleo, across agentic workflows, spend intelligence, automated actions, and more.
- Bring deep applied AI expertise and strong judgment on trade-offs (quality vs latency/cost, build vs buy, agent patterns vs simpler approaches) and help teams avoid hype-driven decisions.
- Work directly with Product, Design, Engineering, and business stakeholders to advise and prioritise on what's actually worth building. You aren't just implementing specs, you're discovering and defining the product.
- Own the evaluation, monitoring, and operationalisation of AI features: setting up evals, tracking drift and performance, and managing prompt changes safely in production. You will help establish how Pleo does this at scale and own the development in production.
- Act as a design partner to the GenAI Core platform team: challenge decisions with evidence from real feature delivery, bring clear requirements, and validate platform choices in production.
- Establish and enforce practical standards for AI feature delivery in product squads (evaluation strategy, monitoring expectations, safe prompt/versioning practices, privacy & safety guardrails) using GenAI Platform tooling (not building the platform itself).
- Upskill the team through mentorship, reviews, pairing, and lightweight playbooks that make other engineers faster.
Requirements
- Proven experience shipping multiple GenAI features into production, at scale in a customer-facing product. You've moved past prototype phases and ideally have experience shipping multi-step, tool-using agents in a user-facing product.
- The ability to translate complex business challenges and product visions into scalable AI solutions. In other words, you can autonomously scope, design and build.
- Deep applied AI judgment: you can reason about evaluation, retrieval quality, tool use, failure modes, and what is/isn't worth building.
- Experience applying evaluation and observability to LLM systems (tests, golden sets, online metrics, monitoring) and using those signals to iterate.
- Strong understanding of privacy and security concerns when building LLM applications: prompt injection, handling PII and data leakage risk.
- Solid experience with the modern AI stack including Vector DBs, orchestration and durable execution frameworks, and LLM APIs.
- Experience building APIs, services and data retrieval pipelines (RAG, vector search etc) to feed data into LLMs.
- Deep proficiency with Python for both data and ML engineering, SQL and major cloud providers.
- An extensive background in traditional ML engineering and a deep understanding of how to architect and build data systems for reliable and scalable production-use.
- Ability to influence cross-functionally with Product/Design/Engineering and bring teams along on decisions.
Bonus: Experience with GCP, BigQuery, Airflow, Python, SQL on the Data side and AWS, Kotlin, Javascript, Typescript on the Product side; infrastructure containerised with Kubernetes.
Conditions
- Hybrid working options
- 25 days of holiday + public holidays
- Comprehensive private healthcare (Vitality, Alan or Médis)
- Paid parental leave
- Mental health support via MyndUp
- Your own Pleo card