Amazon iconAmazonSep 24, 2026 ~7 min source read

Delivery Hero moved ad measurement from hourly batches to real time with Amazon Managed Service for Apache Flink

Delivery Hero rebuilt its ad measurement pipeline to process events in real time, cutting the delay between an event and its recording from 61 minutes to 1.2 seconds and shrinking monthly infrastructure costs by about 57%.

How Delivery Hero rebuilt real-time ad measurement with Apache Flink

Share this story

Send the public story page.

Useful takeaways from this story.

Switching to Amazon Managed Service for Apache Flink produced measurable cost savings—about a 57% reduction in monthly operational costs—and improved data quality.

The legacy system’s limits were structural: lack of event-time semantics, synchronous enrichment that blocked scaling, an ill-suited NoSQL design for per-event mutation, and brittle reprocessing.

Incomplete event context was common (30–40% missing session IDs), and the new architecture targeted richer, recoverable enrichment to close those gaps.

Delivery Hero operates an advertising platform across roughly 65 countries, connecting more than 1.5 million restaurant partners and local vendors to millions of consumers. The platform handles tens of thousands of messages per second and processes billions of ad events per day. Advertising contributed nearly EUR 1.5 billion in revenue in 2025. That scale forced a re-evaluation of how ad impressions, clicks, and orders were measured and billed.

The problem with the legacy pipeline

  • No event-time semantics: Events were bucketed by processing time because most lacked usable event timestamps. When ingestion lagged or events arrived out of order, time-based metrics (for example ROAS) skewed.
  • Slow processing: Hourly batching produced an average latency of 61 minutes between when an event occurred and when it was recorded. That delay was too long for budget pacing and real-time ad serving.
  • Synchronous enrichment: Each event triggered blocking external API calls to fetch campaign and product metadata. Spikes exhausted connection pools and caused cascading failures across billing, ad serving, and reporting.

Incomplete event context compounded data quality problems. Enrichment was best-effort and synchronous, so some events were written with blank fields. Missing session IDs affected roughly 30–40% of events, leaving a significant share unattached to user sessions.

Practical takeaways for teams considering a similar move

  • Event-time processing matters: If events arrive out of order or without timestamps, your metrics will drift when you rely on processing time.
  • Avoid synchronous per-event enrichment calls in high-throughput systems: they create back-pressure and hidden coupling between ingestion and external services.
  • Reassess database choices as workloads change: storage and access patterns that worked for occasional document reads don't hold up for continuous per-event mutations and deduplication.
  • Build reprocessing paths into the streaming architecture so recovery doesn't require ad-hoc tooling or manual coordination.

Delivery Hero's migration to a managed Flink service shows how a streaming-first architecture can reduce latency and costs while improving the reliability of billing-grade ad measurement.

More context around this story.

dbt Meets Apache Flink: One Workflow for Data Engineers
Dzone iconDzoneSep 15, 2026

dbt Meets Apache Flink: One Workflow for Data Engineers

Data engineers managing batch SQL pipelines on Snowflake, BigQuery, and increasingly Databricks, and streaming pipelines on Apache Flink face a familiar problem: two toolchains, two skill sets, two CI/CD pipelines.dbt is now extending into stream processing. This post explains what that means in practice, why it matter

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app