Datadog's Incidents Product

Datadog

An end-to-end solution to automate and speed up incident response.

Role:

Research, Strategy, UX & UI Design

Three views of the Incidents product: the declaration modal, the overview tab, and the timeline

Outages are a major headache for tech companies. Beyond fixing the problem, engineers need to coordinate the response, document what’s happening, and identify lessons.

Datadog — a cloud monitoring platform used by engineering teams to track the health of their infrastructure — built its Incidents product to automate and speed up several of these tasks.

This product started as an internal tool they’d built for when their own services broke. The goal was to productize it at any scale. This case study covers the three core views that make it possible: declare an incident, describe what’s unfolding, capture the messy middle of triage.

The Problem

Incident management is inherently chaotic. Engineers coordinate across Slack, terminals, and dashboards simultaneously. The product had to meet teams in that chaos.

In addition, every company handles incident management differently. Research surfaced a central tension: encourage structure without decreasing speed. We defined goals across companies:

  • Keep declaration frictionless (early and often is better)
  • Respect existing tools and workflows rather than replacing them
  • Stay opinionated about good incident response while remaining highly customizable
  • Reduce manual work as much as possible

I interviewed over a dozen potential customers and ~7 internal Datadog engineers, and joined live incident sessions. The most important finding: most of the real action happens outside the app, in terminals and Slack threads.

A few constraints shaped every decision:

  • A short design phase meant accumulating debt early; we favored fast iteration over polish
  • The design system constrained form design — this project pushed its boundaries without breaking precedent
  • Customers wanted more structure than expected: more preset fields, more taxonomies
  • Incidents can start anywhere, limiting control over user context

The Users

Three entry points into an incident:

The experienced engineer

The experienced engineer — gets paged, assesses severity quickly, and often stays as Incident Commander start to finish

The browsing engineer

The browsing engineer — discovers anomalies in metrics and decides something’s wrong enough to escalate

The manager

The manager — notified downstream as severity escalates; needs context fast, not detail

Three Views, One Arc

I led strategy and design on three interconnected views — each mapped to a specific moment in the incident lifecycle, each feeding into the final output: the post-mortem.

Diagram of the four incident phases: Detection, Triage, Reparation, Retrospective
Detection → Triage → Reparation → Retrospective. The three views map to the first three phases; the post-mortem closes the loop.

1. Declare — Make it frictionless

The incident creation modal

The creation modal is intentionally minimal. At the moment of declaration, users don’t have much information — so we don’t ask much. Severity and commander have defaults. The name field is auto-focused. One required input, then go.

But context is valuable. The modal captures the originating signal — a graph, a monitor alert, a log error — as a data artifact attached from the start. That signal enriches the timeline and helps correlate the incident with related Datadog resources.

Diagram of the most common workflows leading to incident declaration
Incidents can start from a monitor alert, a log error, a customer complaint, or a Slack message. All paths lead to the same modal.

Entry points exist across the platform:

  • A primary Incident Button handles most cases
  • For denser UIs like Dashboards, minimized buttons and context menus serve as secondary options — centrally implemented, locally decided
  • For teams that live in Slack, a separate flow lets non-Datadog users (like customer support agents) declare incidents without leaving their tool
The final creation modal displayed in the context of a Datadog dashboard
Required fields at the top, smart defaults throughout. The side panel gives teams space to add their own response guidelines.
The creation modal design across multiple iterations
Early versions organized fields around numbered steps. Reorganizing around required vs. optional — and redesigning the components themselves — reduced friction significantly.
The Slack incident declaration modal
For non-Datadog users, the Slack modal adapts to their context. Customer support agents can escalate to engineering without ever opening the platform.

2. Describe — Mirror the post-mortem

The incident overview tab

The overview tab is structured to match the document it’s building toward. Post-mortems need a summary, a cause, a timeline of impact, and lessons learned. The overview mirrors that arc from the start — so when the retrospective comes, the structure is already there.

Exploration of information architecture options for the overview tab
Several IA options explored. Ultimately: three columns, each serving a different moment in the incident lifecycle.

Three columns:

  1. Preview modules — auto-populated from other sections; the condensed timeline with time metrics proved especially useful for newcomers joining mid-incident
  2. Freeform fields — the core storyline: what happened, root cause, impact
  3. Attributes sidebar — categorization fields filled toward resolution, used for aggregated reporting
The final overview tab layout
Custom fields can live in any section — placed where they feel contextually relevant, not stacked at the bottom of a long form.

3. Capture — The timeline as a record, not a log

The incident timeline tab

The timeline is where the real story lives. It aggregates Slack messages alongside automatic updates — status changes, severity, customer impact — plus graphs, code snippets, and signals from tools like PagerDuty and GitHub. Built like an activity feed, designed like an editorial timeline. Time is the primary axis.

Diagram showing how Slack messages flow into the timeline
Messages sent in the incident Slack channel are captured directly in the timeline — without leaving the conversation.

Users with admin access can edit timestamps retroactively. That matters. It means the timeline reflects when things actually happened, not when someone found a moment to log them. From those timestamps come the metrics that matter:

  • Time to detection
  • Time to resolution
  • Customer impact duration

All three are now core analytics measures in the product, queryable across dashboards and notebooks. They also feed directly into Datadog’s DORA Metrics framework, connecting incident data to engineering performance at the org level.

The Post-Mortem

All three paths converge here. A wizard flow lets users select relevant content and export it to Notebooks, Datadog’s built-in editor. Graphs stay interactive. The document exports as PDF for external stakeholders, and integrates with Confluence — with further platform support ongoing.

The post-mortem notebook in Datadog
Graphs copied from the timeline remain interactive in the notebook. The post-mortem is a living document, not a static report.

Impact

The product was pre-monetization and the UI wasn’t yet instrumented for behavioral data. Validation came through a handful of pilot companies we partnered with and also heavy internal usage and evaluation.

Interview feedback was consistent. Customers responded to:

  • Customizable fields and contextual help. Different teams have different criteria for severity levels and incident declaration; the system accommodates that.
  • Explicit time metrics — time to detection and resolution, surfaced in the overview and timeline, made response performance quantifiable.
  • A structure that mirrors the post-mortem. Less cleanup at the end. The documentation builds itself along the way.
  • Timeline flexibility — integrations with Slack, PagerDuty, GitHub, and others met teams where they already worked.

Since then, the product has shipped as a paid seat-based SKU. The time metrics designed into the timeline — time to detection, time to resolution, customer impact duration — are now the core analytics measures of the product, queryable across dashboards and feeding into Datadog’s DORA Metrics framework. The timeline itself is described in the product documentation as “the primary source of information for the work done during an incident.”

On the design side, the redesign introduced patterns that hadn’t existed in the system: incident-specific tagging and toggles, a collapsible modal side panel, signal chips with embedded interactive graphs, and a more editorial typographic scale. Each solved a specific problem. Together they gave incident response the experience it needed, and laid the groundwork for incident management to grow into the rest of the Datadog platform.

More Projects

D

Datadog's Incidents Product

Datadog
An end-to-end solution to automate and speed up incident response.
T

The Digital Magazine

The New York Times Magazine
Digital art direction and strategy to shape and advance the magazine‘s online presence.
O

Olympics 2016: The Fine Line

The New York Times
What makes athletes like Simone Biles the world’s best? This mobile-first project explains it.
A

Ambulante

Freelance
UX and content strategy for the new website of this famous Mexican film festival
I

Innovation & Expansion

Museum of American Jewish History
UX/UI design and development of a site-specific interactive exhibition about the expansion of American Jewish communities in the 19th century.
A

A Murderous War on Drugs

The New York Times
A Pulitzer winning photo story to denounce the murderous Philippine war on drugs
2

25 Songs: The Music Issue

The New York Times Magazine
The top songs of the years in a single playlist to listen… and read
E

Editorial Design

The New York Times
Interactive articles selected from my work as Graphics Editor for The New York Times
B

Beam Signing

9/11 Memorial & Museum & Local Projects
An special interactive guestbook to collect and map messages of hope and tributes to the 9/11 victims
L

Last Column

9/11 Memorial & Museum & Local Projects
Multitouch application to navigate the mementos on the last standing column of the World Trade Center site.
D

Desperate Crossing

The New York Times Magazine
A cinematic photo-essay to portrait the tragedy of migrants crossing the Mediterranean sea.
T

The Ring

Salamanca University / Liceu Barcelona Opera House
A website about Wagner‘s Ring cycle of Operas told through audio, video and custom narratives.