About the company
Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure, data, models, and security into one place, using AI to detect and resolve issues before they impact customers.
Responsibilities
- Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams.
- Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more.
- Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation.
- Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog.
- Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code.
Requirements
- Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments.
- Strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, incident.io, Rootly, Jira Cloud Platform, Cortex, or similar.
- Approximately 2+ years of experience in storytelling and creating compelling content e.g. written content, speaking at conferences like SREcon, DevOpsDays or SREday, leading workshops, live streaming or recording videos, contributing to open source and participating in community events.
- Publicly available writing samples, blog posts, demos, or recordings of presentations on technical topics.
- Comfortable with at least one scripting and programming languages (eg Node.js, Python, Go, bash) and familiar with working with APIs and modern infrastructure such as IaaS cloud services and containers.
- Enjoy self driven exploration and education on new technologies and languages.
Conditions
- Bonus points: hands-on ITSM platform experience (ServiceNow, Jira Service Management); active member of SRE, DevOps, or incident-response communities.