The useful part
If you build on Amazon DynamoDB, the news is a good reason to look at how well the two work together. Analytical queries, like revenue by category or order volume by hour, run better against a separate copy of the data, which keeps your table focused on serving low-latency application traffic. In this post, we show you how to run ad hoc SQL queries on your DynamoDB data with DuckDB.
How it works
- A quick look at Amazon DynamoDB Amazon DynamoDB is a serverless, fully managed, distributed NoSQL database with single-digit millisecond performance at any scale.
- Iceberg metadata carries per-file column statistics, so engines skip data files that can't match a query's predicates and read only what they need.
- The data your DynamoDB table replicates into S3 Tables is immediately readable by Amazon Athena, Amazon Redshift, Amazon EMR, Apache Spark, and other Iceberg-compatible engines, DuckDB included.
- Solution overview The following diagram illustrates the solution architecture.
- Solution architecture for querying Amazon DynamoDB data with DuckDB over a zero-ETL integration to Amazon S3 Tables Here's how the flow works:
What to take from it
A zero-ETL integration replicates your table on a refresh interval into Apache Iceberg tables on Amazon S3 Tables, a capability of Amazon Simple Storage Service (Amazon S3). Introducing Amazon S3 Tables Amazon S3 Tables provide storage that is purpose-built for Apache Iceberg, with compaction, snapshot management, and unreferenced file cleanup handled for you. Node.js 20 or later and Docker, because the Lambda container image builds locally.
Example or evidence
- The standard approach is to keep operational traffic on DynamoDB and run analytics against a replicated copy, which historically meant building and maintaining an ETL pipeline.
- DuckDB and its httpfs, aws, avro, and iceberg extensions installed at build time.
- Python 3.9 or later with Boto3 1.35.74 or later for the helper scripts.
- Check progress with the status script: python3 scripts/status.py.
Details worth keeping
Run DuckDB analytics on your Amazon DynamoDB data with zero-ETL | AWS Database Blog Skip to Main Content AWS Database Blog Run DuckDB analytics on your Amazon DynamoDB data with zero-ETL The team behind DuckDB, the open source analytical database, is joining AWS. You design your table around your application's access patterns, and DynamoDB serves those patterns with consistent low latency whether you're handling 10 requests per second or 10 million. They aggregate across the whole table or ranges of the table, rather than reading individual items, and they change often.