Snowflake iconSnowflakeSep 25, 2026 ~7 min source read

Customer Journey Observability: Start with Impact, Drill to Root Cause

Capital One’s journey-first approach implemented in Observe by Snowflake shifts incident investigation to begin with customer interactions, then drills down through services, operations, traces, and logs to confirm root cause faster.

Customer Journey Observability: From Impact to Root Cause

Share this story

Send the public story page.

Useful takeaways from this story.

Begin incident investigation at the customer interaction layer so impact is measured first and mitigation becomes the main focus.

Customer journey graphs map interactions (nodes) and user flows (edges) while overlaying RED metrics and session counts to show scope and severity.

From a degraded journey step you can pivot to the responsible service and operation, then confirm the failure path with traces and logs.

# What this story describes

Capital One contrasts the traditional service-first observability model with a journey-first approach implemented in Observe by Snowflake. Instead of starting at infrastructure or service alerts and estimating customer impact afterward, teams begin with a defined customer interaction, measure business impact, then follow relationships in telemetry down to services, operations, traces, and logs to confirm root cause.

# Why this matters

# Snowflake

  • Telemetry ingested into Observe's data lake includes real user monitoring (RUM) logs, clickstream data, and API access logs.
  • After ingestion, interactions and services are modeled and enriched without requiring app-level instrumentation.
  • Customer journey graphs are constructed where nodes represent interactions (for example: Home, Product Details, Add Item, Complete Order) and edges represent the user flow.
  • RED metrics (rate, errors, duration) and unique session counts are aggregated and overlaid on the journey graph so the operational surface shows business-level impact first.

# The practical workflow (e-commerce example)

1) Start at the customer interaction

  • Operators identify a degraded journey step such as Complete Order. The graph immediately shows how many sessions, error rate, and latency are affected so teams know scope and severity before troubleshooting technical layers.

2) Drill down to the failing operation

  • From the failing interaction you pivot to service- and operation-level breakdowns. In the example, the EmptyCart operation in the cart service was causing latency and errors. This narrows ownership and the domain for remediation.
  • Once a suspect operation is identified, engineers pivot into traces to inspect where time is spent and where errors originate in a specific request. They then drill into downstream span logs to confirm the explicit error message and complete the failure narrative: users can't complete checkout because a downstream operation cannot read cached cart data.

# Benefits observed

  • Immediate, business-focused measurement of impact so executives and SREs share the same single source of truth.
  • Implementation without adding application instrumentation because interactions are modeled after telemetry ingestion.

# Operational realities and constraints

  • The structure and relationships are encoded in the data, but investigations remain human-driven. Engineers still decide where to pivot, which signals matter, and how to summarize findings under pressure.
  • Synthesis of findings for the incident response room still requires time and coordination even when impact and root cause are surfaced more quickly.

# Bottom line

More context around this story.

Продуктовая гипотеза или баг фронтенда: как перестать угадывать почему пользователи не дошли до цели
Habr iconHabrSep 2, 2026

Продуктовая гипотеза или баг фронтенда: как перестать угадывать почему пользователи не дошли до цели

Возьмём типичный отчёт воронки интернет‑магазина: «Каталог → Товар → Корзина → Оформление заказа → Оплата». До последнего шага дошло 38% сессий, остальные 62% потерялись между «Оформлением» и «Оплатой». Для продуктовой аналитики эти 62% — повод для гипотезы «форма оформления слишком длинная» и задачи на редизайн. Для к

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app