High-throughput systems often need to answer a binary membership question on a hot path (for example, "is this card blocked for this merchant?"). A relational lookup can be too slow for strict latency budgets. A standalone Bloom filter is fast and memory-efficient but allows false positives, which some applications cannot tolerate. The design goal is to serve most decisions with sub-millisecond latency while guaranteeing no false-positive business outcomes.
Tier 1 — Bloom filter (ElastiCache for Valkey)
- Role: fast-negative gate. A BF.EXISTS call checks k hash positions in a compact bit array.
- Behavior: if any bit is zero, the item is guaranteed absent (no downstream I/O). If all bits are set, the item is possibly present and must go to Tier 2.
Tier 2 — Exact-match cache (ElastiCache for Valkey)
- Role: deterministic, authoritative answer for keys that were populated there.
Tier 3 — Aurora PostgreSQL (source of truth)
- Role: canonical authoritative store. A primary-key point query returns the definitive answer, populates the cache, and drives updates to upstream tiers.
Why this composition is correctness-safe
Each tier provides a narrow guarantee. The Bloom filter never yields false negatives relative to the data it has been updated with, but it can yield false positives. Those false positives are harmless because Tier 2 and Tier 3 will resolve them: a BF.EXISTS=1 leads to a cache lookup and, if needed, a database query. The database always has the final word, so no false-positive business outcomes occur.
Operational notes and implementation choices
- Co-locate tiers: ElastiCache for Valkey supports both Bloom commands (BF.EXISTS, BF.ADD) and key-value operations (GET, SET) on the same cluster. You can host Tier 1 and Tier 2 on one cluster using distinct key prefixes, reducing the deployment surface to one ElastiCache cluster plus one Aurora cluster.
- Composite keys: when membership is scoped (tenant, merchant, entity), encode a composite key pattern so the Bloom filter and cache index membership consistently and unambiguously.
- False-positive rate (FPR) tuning: choose the filter FPR when creating the Bloom filter. Lower FPR requires more memory but reduces the fraction of requests that advance past Tier 1.
Performance and cost considerations
- Bloom filters provide large memory savings versus set-based indexes, reducing memory footprint for high-cardinality membership sets.
- The Bloom filter reduces I/O by stopping absent-key requests at Tier 1. The remaining requests incur cache and occasional database read costs dictated by the FPR and cache miss rate.
- Hosting both Valkey data structures on a single ElastiCache cluster simplifies operational cost and latency between Tier 1 and Tier 2.