Waroom

Waroom is an incident management and on-call scheduling platform for engineering teams and SREs.

Incident management that turns crisis into clarity

One product for on-call paging, incident response, and blameless post-mortems. AI drafts the write-up, grades it against your rubric, and answers questions about your incident history in plain language.

Start free trial See all features

14-day trial. No card. Slack or Teams connected in about five minutes.

Incident detail header, Waroom app
INC-2481

Checkout p99 latency above 4s in eu-west-1

Mitigated Declared 19m ago Channel Call
Severity Critical Impact Customer-facing

One product

On-call paging and incident response, not two vendors and two invoices.

Scored, not just drafted

Post-mortems graded against a nine-dimension rubric, with the trend over time.

Agent native

A hosted MCP server on api.waroom.co/mcp, sharing one tool registry with Ask AI.

Reachable

iOS critical alerts and Android high-priority push cut through Do Not Disturb.

The AI layer

AI that works the incident, not a chat box bolted on the side

Three surfaces, one tool registry. The agent that answers your questions in the app is the same one your editor talks to over MCP, and the same context that drafts and grades your post-mortems.

Real product UI, not illustrations

See Ask AI and MCP in action

A post-mortem that gets graded, not just generated

The draft is built from the real timeline: events, status updates, deploys, runbook runs, related incidents. Then a rubric pass scores it dimension by dimension, so quality is measured instead of assumed. Teams see the score move across quarters.

  • Variable-driven templates pre-fill the doc, with a non-AI draft path when you want one
  • Action items sync to Jira and Linear, doc publishes to Confluence or Notion
  • Anthropic, OpenAI, or Google on your own keys, every call traced in Langfuse
Post-mortem quality score card, Waroom app
84 / 100 B Rubric v1
Completeness 100/100 20 pts
Follow-up actions 79/100 15 pts
Timeliness 100/100 10 pts
Timeline depth 84/100 8 pts
Human edit effort 70/100 7 pts
Root cause quality 74/100 15 pts

Names the pool-sizing default but stops short of why the limit was never revisited.

Impact clarity 82/100 10 pts
Action item specificity 65/100 8 pts
Blamelessness 91/100 7 pts
Light edits

Stop re-diagnosing the same outage

Waroom clusters recurring root causes out of your post-mortems and maps relationships between incidents, so the sixth pool starvation reads as a pattern with an owner instead of a fresh mystery.

  • Clusters carry an owning team and a status, with team showback
  • Severity-weighted burden ranks what is actually costing you
  • Unfinished follow-ups surface next to the cause that keeps returning
Recurring root causes analysis, Waroom app
Recurring Root Causes Updated 2h ago

RECURRING ROOT CAUSES

7

2 without an owner

RECURRING INCIDENTS

23

Last 90 days

SEVERITY-WEIGHTED BURDEN

96

sev1x8 . sev2x4 . sev3x2 . sev4x1

OPEN FOLLOW-UPS

11

unfinished follow-up actions

Cluster Incidents First seen Follow-ups

Connection pool exhaustion

Payments . Checkout

6 Mar 4 2 open

Expired third-party credential

Integrations

3 Feb 19 0 open

Cache stampede after deploy

Search

2 Apr 11 1 open

Retry storm on ledger writes

Payments

2 Mar 28 0 open

Your keys, your provider

Anthropic, OpenAI, or Google credentials stay yours. Waroom never marks up usage.

Traced and budgeted

Langfuse tracing on every call, token budgets on post-mortem chat, usage metered per org.

Nothing writes silently

Read tools run free. Declares, acks, and status updates always ask a human first.

Competitive advantages

What you get that alerting tools don't ship

Waroom combines native paging, scored post-mortems, and an agent-native tool surface in one product, so teams improve after every incident instead of just surviving them.

01

AI-drafted, AI-scored post-mortems

Rubric-scored, not just drafted

Nine weighted dimensions with per-dimension rationale and a trend over time, so quality is measured instead of assumed.

02

Native on-call and paging

Built in, not bolted on

Schedules, escalation policies, and paging over SMS, voice, WhatsApp, email, Slack, and Teams inside the core product.

03

Policy-driven nudges

Mute, snooze, and response rates

Automated nudges for stale incidents, overdue updates, and unfinished post-mortems, with owner controls and follow-through analytics.

04

Bring your own AI and paging keys

No vendor markup, no lock-in

Anthropic, OpenAI, or Google keys and your own Twilio or Meta WhatsApp credentials. Usage is never marked up.

05

Recurring root-cause clustering

See the patterns across incidents

Causes cluster with an owning team, severity-weighted burden, and the follow-ups nobody finished.

06

Slack and Microsoft Teams at parity

Two-way mirroring, both platforms

Slash commands, dedicated channels, and mirroring work the same on either platform, not as an afterthought.

07

Native iOS and Android apps

Critical alerts break through silent mode

Critical push on iOS and high-priority alerts on Android cut through Do Not Disturb. One tap to acknowledge, with biometric lock.

08

MCP server and Ask AI agent

Human confirmation on every mutation

One tool registry serves the hosted MCP server and the in-app agent, and every write waits for an explicit approval.

The incident lifecycle

Detect, run, close (formerly Detect, Mobilize, Resolve, Learn)

01

Detect and page (detect, mobilize)

Datadog, New Relic, or any webhook opens the incident, with a PagerDuty payload preset for migrations. Escalation policies walk the ladder over SMS, voice, WhatsApp, email, Slack, and Teams on your own Twilio and Meta credentials.

Escalation policies Calendar export Auto-resolve on clear
02

Run the incident (resolve)

Slack and Microsoft Teams run at parity: slash commands, dedicated channels, two-way mirroring. Bridges on Zoom, Meet, or Teams. Runbook steps execute real HTTP actions, and the nudge engine chases stale incidents and overdue updates.

Interactive runbooks Private channels Stakeholder updates
03

Close the loop (learn)

AI drafts the blameless write-up, the rubric grades it, action items become Jira or Linear issues with aging tracked, and the doc publishes to Confluence or Notion. Recurring root causes cluster themselves.

Rubric scoring Root-cause clustering Published post-mortems

Measure

Response and delivery, computed from the same events

MTTX for how you respond, DORA for how you ship. A change failure links straight to the incident it caused, and the bands for elite, high, medium, and low are configurable per org.

  • MTTA, MTTM, MTTR with mean, p50, and p95
  • Action item completion rate and aging
  • Post-mortem quality by dimension, nudge response rates
  • Contact reachability and on-call readiness scoring
Delivery metric Value Band
Deployment frequency 14 / day Elite
Lead time for changes 6h 20m Elite
Change failure rate 11.4% Medium
Failed deploy recovery 38m High
Deploy 7f21c9 on checkout-api, marked as the cause of INC-2481.

Integrations

Works with the tools already in the incident

Native integrations, plus signed outbound webhooks, a REST API, SAML and OIDC SSO, and SCIM provisioning.

Slack Microsoft Teams Zoom Google Meet Jira Linear Confluence Notion Datadog New Relic GitHub Twilio WhatsApp SendGrid PagerDuty payload preset Custom webhooks

Questions

Before you switch

Do I still need a separate on-call and paging tool?
No. On-call scheduling, escalation policies, and multi-channel paging over SMS, voice, WhatsApp, email, Slack, and Teams are part of the core product, alongside AI post-mortems, action item tracking, and MTTX dashboards.
Whose AI models run, and who pays for the tokens?
Yours. Bring your own Anthropic, OpenAI, or Google keys and Waroom uses them directly with no markup. Every call is traced in Langfuse, post-mortem chat runs under a token budget, and usage is metered per organization.
Can the AI change things without me knowing?
No. Read tools run automatically inside the agent loop, but anything that mutates state, such as declaring an incident, acknowledging an alarm, or posting a status update, surfaces as a pending tool call with its exact arguments and only runs after a human approves it. That rule holds in Ask AI and over MCP.
What does the MCP server actually expose?
Around 35 tools across incidents, alarms, insights, services, teams, and post-mortems, defined once in a shared registry and consumed by both the MCP server at api.waroom.co/mcp and the in-app Ask AI agent. It runs its own OAuth authorization server with a consent screen, so connecting from Claude or another client is a normal authorize flow.
How long does setup take?
Most teams are running in under five minutes: connect Slack or Microsoft Teams, invite the team, create the first on-call schedule. Alert ingestion supports Datadog, New Relic, and generic webhooks, including a PagerDuty inbound payload preset.

Your next incident is going to happen anyway

Set Waroom up before it does. Connect Slack or Teams, invite the team, create the first schedule.

No credit card. Free plan available. Cancel anytime.

Powered by Clickroom Product analytics and feature flags inside Waroom run on Clickroom, our sibling product at Waroom Co. One SDK, zero config, 100k events a month free. clickroom.co