Work Hard Everywhere logo Work Hard Everywhere

Developer Advocate - Service Management

Datadog
📍 Anywhere in the World 💰 🕑 Any timezone
Full-time Mid-level Engineering DevOps And Sysadmin

Job Description

Headquarters: California, USA, Remote; New York, USA, Remote

We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management.

What You'll Do:

- Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams

- Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more

- Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation

- Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog

- Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code

Who You Are:

- Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments

- You have a strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, incident.io, Rootly, Jira Cloud Platform, Cortex, or similar and know...