Databricks iconDatabricksSep 19, 2026 ~7 min source read

RADAR: How Databricks catches ‘gray failures’ with anomaly detection

Gray failures are partial, slow-spreading outages that standard health checks miss. Databricks’ RADAR pipeline turns user-facing error signals into fast, metric-agnostic detection, alerting, and root-cause steps to find those failures in minutes.

RADAR: Catch gray failures with anomaly detection

Share this story

Send the public story page.

Useful takeaways from this story.

RADAR is a four-stage pipeline — reliability metrics, anomaly detection, alerting, and root-cause analysis — designed to flag unusual error spikes by user slice and other dimensions.

A practical, high-signal input is spikes in user-facing errors that would normally look like user fault (e.g., INVALID_ARGUMENT) but occur at scale for many users at once.

The useful part

This post is about catching them early with anomaly detection — how we do it at Databricks with a system called RADAR, and how you can build the same thing for whatever metric matters most to your business. Catch gray failures with anomaly detection | Databricks Blog Skip to main content Summary Gray failures slip past green dashboards, quietly costing you customers and revenue before anyone notices. You can build the same system on Databricks for any metric — billing, conversion, or model performance — using native components and an AI-agent scaffold.

How it works

  • These "gray failures" leak users and revenue for hours before anyone connects the dots.
  • It's written for the people who own service reliability: SREs, platform and data engineers, on-call responders, and the engineering leaders they answer to.
  • Keeping a customer-facing service reliable is your job, and every dashboard on your wall is green — CPU healthy, latency fine, servers up, database connected.
  • There's no server crash you'd catch instantly, just a messy middle where most users are fine and one group fails the whole time.
  • Systems calls it differential observability — your failure detectors don't notice a problem even while your users clearly do.

What to take from it

From the outside the house looks fine, but inside the damage is spreading — and the longer you wait, the bigger the blast radius. Get the RADAR scaffold on GitHub Because the best outcome isn't a faster response to angry customers — it's that your customers never have to discover your incidents for you. Some of the most damaging outages are the ones your monitoring never flags: a slice of your customers quietly fails while every health check still reads normal.

Example or evidence

  • Picture a bunch of users in one region who suddenly can't spin up a certain type of cluster.
  • an error that is politely saying, "this one's on you." But when many users hit the same "your fault" error at the same moment, it stops being their fault.
  • At every point in time, record two things: how many errors are happening, and how many separate users hit each one.
  • Build it yourself on Databricks The best news: every piece you need is already on Databricks.

Details worth keeping

RADAR is a four-stage, metric-agnostic pattern — reliability metrics, anomaly detection, alerting, and root-cause analysis — that Databricks runs on itself to catch these failures in minutes, at over 90% precision and 95% faster discovery. When everything is green but nothing is fine Picture a normal Wednesday. By every signal your team watches, the system looks perfect.

Related coverage

  • Qualitymag: Conventional screening asks: does this unit violate a limit we set?
  • Medium: A demo that nails every question in the interview room, then embarrasses itself in production a month later. Here's why that happens — and… Continue reading on Medium »
  • Towards Data Science: How to verify your app aligns with your intent without ever reading a line of generated code.

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app