I’m a senior product designer based in New York with more than fifteen years of experience across consumer and B2B products, editorial design, and interactive installations. My work spans DoorDash, Datadog, The New York Times, and institutions like the 9/11 Museum in NYC. I get equally excited talking about conversion flows, user insights, system architecture, and Pulitzer-winning photo stories.
I’m obsessed with making complexity simple by understanding users, strategizing to focus on the right problems, and building things in design or directly in code.
At Datadog I built and led a team of five designers. At DoorDash I drive design strategy and execution as a staff IC: defining scope, shaping roadmaps, and mentoring designers while staying deeply hands-on. I also enjoy teaching, giving talks, and organizing workshops. Mentorship is a core part of how I work.
In my spare time, I love catching up on performing and visual arts, be that an opera, an off-off Broadway show or the latest exhibit at MoMA.
An end-to-end solution to automate and speed up incident response.
Role:
Research, Strategy, UX & UI Design
Outages are a major headache for tech companies. Beyond fixing the problem, engineers need to coordinate the response, document what’s happening, and identify lessons.
Datadog — a cloud monitoring platform used by engineering teams to track the health of their infrastructure — built its Incidents product to automate and speed up several of these tasks.
This product started as an internal tool they’d built for when their own services broke. The goal was to productize it at any scale. This case study covers the three core views that make it possible: declare an incident, describe what’s unfolding, capture the messy middle of triage.
The Problem
Incident management is inherently chaotic. Engineers coordinate across Slack, terminals, and dashboards simultaneously. The product had to meet teams in that chaos.
In addition, every company handles incident management differently. Research surfaced a central tension: encourage structure without decreasing speed. We defined goals across companies:
Keep declaration frictionless (early and often is better)
Respect existing tools and workflows rather than replacing them
Stay opinionated about good incident response while remaining highly customizable
Reduce manual work as much as possible
I interviewed over a dozen potential customers and ~7 internal Datadog engineers, and joined live incident sessions. The most important finding: most of the real action happens outside the app, in terminals and Slack threads.
A few constraints shaped every decision:
A short design phase meant accumulating debt early; we favored fast iteration over polish
The design system constrained form design — this project pushed its boundaries without breaking precedent
Customers wanted more structure than expected: more preset fields, more taxonomies
Incidents can start anywhere, limiting control over user context
The Users
Three entry points into an incident:
The experienced engineer — gets paged, assesses severity quickly, and often stays as Incident Commander start to finish
The browsing engineer — discovers anomalies in metrics and decides something’s wrong enough to escalate
The manager — notified downstream as severity escalates; needs context fast, not detail
Three Views, One Arc
I led strategy and design on three interconnected views — each mapped to a specific moment in the incident lifecycle, each feeding into the final output: the post-mortem.
Detection → Triage → Reparation → Retrospective. The three views map to the first three phases; the post-mortem closes the loop.
1. Declare — Make it frictionless
The creation modal is intentionally minimal. At the moment of declaration, users don’t have much information — so we don’t ask much. Severity and commander have defaults. The name field is auto-focused. One required input, then go.
But context is valuable. The modal captures the originating signal — a graph, a monitor alert, a log error — as a data artifact attached from the start. That signal enriches the timeline and helps correlate the incident with related Datadog resources.
Incidents can start from a monitor alert, a log error, a customer complaint, or a Slack message. All paths lead to the same modal.
Entry points exist across the platform:
A primary Incident Button handles most cases
For denser UIs like Dashboards, minimized buttons and context menus serve as secondary options — centrally implemented, locally decided
For teams that live in Slack, a separate flow lets non-Datadog users (like customer support agents) declare incidents without leaving their tool
Required fields at the top, smart defaults throughout. The side panel gives teams space to add their own response guidelines.Early versions organized fields around numbered steps. Reorganizing around required vs. optional — and redesigning the components themselves — reduced friction significantly.For non-Datadog users, the Slack modal adapts to their context. Customer support agents can escalate to engineering without ever opening the platform.
2. Describe — Mirror the post-mortem
The overview tab is structured to match the document it’s building toward. Post-mortems need a summary, a cause, a timeline of impact, and lessons learned. The overview mirrors that arc from the start — so when the retrospective comes, the structure is already there.
Several IA options explored. Ultimately: three columns, each serving a different moment in the incident lifecycle.
Three columns:
Preview modules — auto-populated from other sections; the condensed timeline with time metrics proved especially useful for newcomers joining mid-incident
Freeform fields — the core storyline: what happened, root cause, impact
Attributes sidebar — categorization fields filled toward resolution, used for aggregated reporting
Custom fields can live in any section — placed where they feel contextually relevant, not stacked at the bottom of a long form.
3. Capture — The timeline as a record, not a log
The timeline is where the real story lives. It aggregates Slack messages alongside automatic updates — status changes, severity, customer impact — plus graphs, code snippets, and signals from tools like PagerDuty and GitHub. Built like an activity feed, designed like an editorial timeline. Time is the primary axis.
Messages sent in the incident Slack channel are captured directly in the timeline — without leaving the conversation.
Users with admin access can edit timestamps retroactively. That matters. It means the timeline reflects when things actually happened, not when someone found a moment to log them. From those timestamps come the metrics that matter:
Time to detection
Time to resolution
Customer impact duration
All three are now core analytics measures in the product, queryable across dashboards and notebooks. They also feed directly into Datadog’s DORA Metrics framework, connecting incident data to engineering performance at the org level.
The Post-Mortem
All three paths converge here. A wizard flow lets users select relevant content and export it to Notebooks, Datadog’s built-in editor. Graphs stay interactive. The document exports as PDF for external stakeholders, and integrates with Confluence — with further platform support ongoing.
Graphs copied from the timeline remain interactive in the notebook. The post-mortem is a living document, not a static report.
Impact
The product was pre-monetization and the UI wasn’t yet instrumented for behavioral data. Validation came through a handful of pilot companies we partnered with and also heavy internal usage and evaluation.
Interview feedback was consistent. Customers responded to:
Customizable fields and contextual help. Different teams have different criteria for severity levels and incident declaration; the system accommodates that.
Explicit time metrics — time to detection and resolution, surfaced in the overview and timeline, made response performance quantifiable.
A structure that mirrors the post-mortem. Less cleanup at the end. The documentation builds itself along the way.
Timeline flexibility — integrations with Slack, PagerDuty, GitHub, and others met teams where they already worked.
Since then, the product has shipped as a paid seat-based SKU. The time metrics designed into the timeline — time to detection, time to resolution, customer impact duration — are now the core analytics measures of the product, queryable across dashboards and feeding into Datadog’s DORA Metrics framework. The timeline itself is described in the product documentation as “the primary source of information for the work done during an incident.”
On the design side, the redesign introduced patterns that hadn’t existed in the system: incident-specific tagging and toggles, a collapsible modal side panel, signal chips with embedded interactive graphs, and a more editorial typographic scale. Each solved a specific problem. Together they gave incident response the experience it needed, and laid the groundwork for incident management to grow into the rest of the Datadog platform.