Datadog favicon

Datadog US1 outage on 2026-09-25

datadoghq.com
September 2026 25 Minor incident

For 27 minutes on 25 September, Datadog's On-Call Pages feature broke for users trying to acknowledge, escalate, or resolve pages. Datadog's status page confirms the fault sat on its side, in the Incident Response component.

Started

15:11 UTC

Duration

Lasted 27m

Source

IsDown

Next time

Datadog is degraded. Is your product working?

An incident on their side does not always mean an outage on yours, and the only way to know is to already be watching your own endpoints. Set that up in under a minute and you will be notified the next time your own site is affected. And there will be a next time.

50 monitors for your own websites. No card.

Affected components: Incident Response

What happened?

Datadog logged a single incident on 25 September, starting at 15:11 UTC and resolved by 15:38 UTC, a 27 minute window. The title on Datadog's status page is specific: users were unable to acknowledge, escalate, or resolve On-Call Pages. This sits inside the Incident Response component, the part of Datadog that handles on-call paging and escalation rather than monitoring data itself. Datadog posted five updates across the window. The first, at 15:11, said they were investigating users unable to acknowledge on-call pages. A second update six minutes later widened the description to cover escalating and resolving pages too, suggesting the initial scope was narrower than what they found. At 15:23 Datadog said they had identified the issue and were applying mitigations. By 15:29 they were monitoring the fix, and at 15:38 they marked it resolved. No root cause is given anywhere in that sequence. Datadog's status page does not say why pages stopped being actionable, only that mitigations were applied and the issue cleared. No user reports were filed during this window on record, and there is no post-mortem published for this incident.

Learning

This incident sat entirely inside Datadog's own Incident Response layer, the mechanism that lets you act on a page once it fires, not inside the monitors or checks that generate alerts in the first place. If your own synthetic checks or health pages were green that day, they would have been telling the truth: the data pipeline and alerting logic were working, but the human-facing step of acknowledging or escalating a page was broken on Datadog's side. That is a layer you cannot see by watching your own endpoints, because the fault was in how Datadog's UI and backend processed your response to an alert, not in anything you control. Over the past 90 days Datadog has logged 12 such incidents with a median duration of 51 minutes, and across its history it averages 4.3 a month, so this is not a one-off. UptimeRobot's own 30 day data shows 6 incidents with a median of 50 minutes, all resolved, which lines up with Datadog's own pattern. Your monitoring covers your half, UptimeRobot now watches the provider's half too, so next time you know which side broke without guessing.

4.7
stars out of 5
284+ reviews on

Start monitoring in 30 seconds.

There's nothing to install. No credit card required. 50 monitors for free.