Towards Data Science iconTowards Data ScienceSep 17, 2026 ~7 min source read

Building a Data Lakehouse with DuckDB and DuckLake

A practical walkthrough showing how to start from a local Parquet file, extend the lake to an S3 object, and use DuckDB plus the DuckLake extension to manage metadata, query data, and evolve tables with SQL.

Share this story

Send the public story page.

Useful takeaways from this story.

DuckDB is a fast, embedded analytical database well suited for single-node, selective queries on Parquet files up to a few hundred gigabytes.

You can build a local lakehouse first with Parquet and DuckLake, then join or ingest S3 Parquet data and use DuckLake snapshots to inspect prior table versions.

DuckLake is production-ready (DuckDB team released DuckLake v1.0 recently), but very large multi-user workloads may require stronger compute or distributed engines.

# Overview

# DuckLake are DuckDB is an in-process analytical SQL database designed for fast, memory-resident queries. It performs well on small to medium datasets (the author mentions up to a couple of hundred GBs as a reasonable range for single-node use).

DuckLake is an extension developed by the DuckDB team that changes where table metadata is kept. Instead of metadata files colocated with Parquet data, DuckLake stores metadata in a relational database. That metadata records which Parquet files belong to a table, partitioning, statistics, table versions and change history, enabling SQL-driven inserts, updates, deletes, schema changes and time travel without rewriting Parquet files directly.

# Why this approach Parquet is the de facto columnar file format for analytical lakes because it avoids reading unnecessary columns. Traditional table formats like Delta Lake, Apache Iceberg and Apache Hudi use metadata files alongside Parquet to track table state. DuckLake moves that metadata into a relational store. That makes metadata queries a regular SQL operation and keeps the data files as plain Parquet.

# What you will build The walkthrough covers these concrete steps:

  • Query local and remote Parquet files with DuckDB.
  • Create a DuckLake catalog/table to manage local data.
  • Use SQL to update data and evolve a table schema inside DuckLake.
  • Use DuckLake snapshots to examine prior table versions.
  • Join a DuckLake table to an external S3 Parquet file.

# Practical notes and constraints Prerequisites listed in the article include a modern desktop OS (Windows, macOS or Linux), a terminal or PowerShell, internet access to install DuckDB and the DuckLake extension, and an AWS account and S3 permissions for the S3 portion. The example file sizes are intentionally small — the walkthrough demonstrates technique rather than scale.

# Next steps the article demonstrates The walkthrough gives concrete SQL and command-line steps to:

  • Read Parquet files locally with DuckDB.
  • Initialize DuckLake and create a table pointing at the Parquet files.
  • Execute inserts/updates and show how the metadata tracks changes and versions.
  • Upload a small orders Parquet to S3 and query/join it with the local DuckLake customer table.
  • Create a DuckLake table backed by the S3 data if desired.

This approach is useful when you want a minimal-cost lakehouse that keeps data in Parquet and manages table state with SQL-accessible metadata. It is a practical option for development, small production workloads, and for teams that want a simple, open-source path to a lakehouse without heavy vendor lock-in.

More context around this story.

DuckDB And The Embedded Analytical Database Trend
Javacodegeeks iconJavacodegeeksSep 1, 2026

DuckDB And The Embedded Analytical Database Trend

For most of the history of data warehousing, running a serious analytical query meant talking to a server somewhere: a cluster, a managed warehouse, or at minimum a separate database process listening on a port. DuckDB broke that assumption. It is a full SQL analytical engine that runs inside your own process, whether

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app