About the company
Apify is the largest marketplace of tools for AI. 40,000+ Actors helping people and agents get real-time web data, track competitors, generate leads, or integrate their apps. Actors are built by a global creator community that now earns more than $1.2 million every month.
Join us to help people put the web to work. Apify can find missing children, protect consumers from fake discounts across the EU, and feed data to AI chatbots.
Responsibilities
- Operate and improve our monitoring stack (Prometheus, Grafana, OpenTelemetry) - instrument services to expose the right metrics, define what we watch in production, and shape alerting so teams get actionable signals without the noise.
- Help define how we run incidents - clear communication, structured learning afterward, and supporting artifacts (status page, runbooks).
- Work with platform and product engineers to make reliability standards practical - help teams adopt better tooling or practices when things change, and write documentation people actually use.
Requirements
- Hands-on experience choosing what to measure in production - not just reading dashboards, but picking signals that reflect the customer experience.
- Comfortable with incidents and alerts, from early detection through resolution and follow-up so similar issues are less likely to recur.
- Hands-on experience with Prometheus, Grafana, OpenTelemetry, or similar, and with alert-routing tools such as PagerDuty.
- Read and write code: can follow services and pipelines across the stack and collaborate on technical details with the teams building them.
- Know what good post-incident culture looks like in practice - blame-free, learning-focused, and actually used to make things better.
- Can write clear, concise guidance that teams adopt, and work constructively toward sound decisions.
- Driven to automate repetitive tasks and improve developer workflows.
Nice to have:
- Meaningful hands-on experience as an application or backend developer.
- Experience building and maintaining infrastructure on AWS (EC2, EKS, S3, CloudFormation, or similar), and hands-on experience with container technologies.
- Some familiarity with CI/CD pipelines or release practices.
Conditions
- Space, support, and autonomy for personal growth, with a direct impact on Apify's success.
- Full-time position in our amazing offices in Prague (100Yards) or Brno (Titanium).
- Option to work remotely.
- Flexible working hours.