Cortea is a Berlin startup transforming audits with AI
AI-powered software and specialized AI agents remove repetitive work so auditors can focus on judgment
Backed by top-tier VCs with >15m EUR funding, working product and paying customers
Values: first-principles thinking, speed, trust, and kindness
Building side by side in the Berlin office
Responsibilities
Build online and offline evaluation systems for LLM agents, including pipelines using golden datasets, ground-truth data, human review workflows, and experiment results
Create automated quality gates so changes to prompts, context, models, or agent logic can be tested before production
Analyze large volumes of agent traces in columnar and analytical databases such as BigQuery or ClickHouse to identify failure modes, quality regressions, latency issues, reliability gaps, and cost optimization opportunities
Build reliable data retention and replay mechanisms for long-term analysis of production agent behavior
Manage observability tools for tracing, monitoring, debugging, and experiment management
Team up with backend engineers to improve speed and reliability of retrieval and reasoning agents