Towards Data Science iconTowards Data ScienceSep 29, 2026 ~6 min source read

Compacting 1,000 Apache Iceberg Files Into 6: why the author ran the test and how they did it

A practical briefing of the experiment that measures how rewriting many tiny Iceberg data files into a few larger ones affects SQL query workloads — including the problem, the toolchain, the compaction operation used, and the experimental approach.

Share this story

Send the public story page.

Useful takeaways from this story.

The author built a local, reproducible test: an Iceberg table with 50 million rows spread across 1,000 tiny files, then used Iceberg’s default bin-pack rewrite to produce 6 files and benchmarked three SQL workloads before and after.

The article documents a minimal dev environment and exact software choices so readers can reproduce the experiment locally without cloud services.

The useful part

When you deal with large analytical datasets stored in table formats such as Apache Iceberg, there is a known issue called the small files problem. This happens when data is written in lots of tiny files rather than a smaller number of reasonably sized ones. This increases metadata overhead and can slow down query planning and execution.

How it works

  • This reduces metadata overhead and gives the query engine fewer files to open, scan and manage, which can improve performance significantly.
  • Iceberg supplies the rewrite_data_files procedure and the metadata needed to select and process files for compaction but we as the system admins still have to execute that procedure or configure another...
  • The downside of this setup is that we can't reproduce truly massive datasets that are used in real-world systems but hopefully our results will give us some useful insights.
  • An Iceberg table normally contains Parquet, Avro or ORC data files, plus metadata describing which files currently belong to the table.
  • Its snapshots provide a history of table changes, while manifest files help query engines locate relevant data files without recursively listing every directory.

What to take from it

Iceberg supports atomic changes, schema evolution, partition evolution and time-travel queries without converting the underlying data into a proprietary storage format. Adaptive Query Execution is disabled because Spark might otherwise combine our deliberately small tasks. I'm focussing on Apache Iceberg as it's rapidly growing into one of the leading open table formats for large-scale analytical data, bringing features such as schema evolution, time travel, partition evolution and reliable transactions to data stored in object storage or distributed file systems.

Example or evidence

  • Microsoft is not responsible for, nor does it grant any licenses to, third-party packages.
  • Apache Iceberg is an open table format for accessing huge analytic datasets, bringing database-like features such as schema evolution, partition evolution, time travel, and reliable transactions to data...
  • Python 3.10–3.13 Java 17 or Java 21 PySpark 4.0.3 Apache Iceberg 1.11.0 powershell REM Install JAVA c:\iceberg-compaction> winget install EclipseAdoptium.Temurin.21.JDK Found Eclipse Temurin JDK with...
  • Writing the Python code Create a file named iceberg_compaction_demo.py.

Details worth keeping

I Compacted 1,000 Apache Iceberg Files Into 6. Importantly, although Iceberg supports compaction it doesn't do it for us automatically. But, really, does compaction make that much of a difference?

Related coverage

  • Habr: В феврале этого года проект с красивым именем Polaris незаметно стал полноценным проектом фонда Apache.
  • Amazon: Accelerate slow, repetitive Apache Spark analytical queries on Apache Iceberg tables without rewriting any SQL.
  • Habr: Инженеры данных построили агрегат — маленькую таблицу «продажи по магазинам по дням».

More context around this story.

Apache Iceberg: Индиана Джонс и Каталог судьбы в Lakehouse
Habr iconHabrSep 20, 2026

Apache Iceberg: Индиана Джонс и Каталог судьбы в Lakehouse

В феврале этого года проект с красивым именем Polaris незаметно стал полноценным проектом фонда Apache. Новость прошла как-то мимо, ну стал и стал, мало ли. Между тем это одна из тех новостей, которые лет через пять, возможно, будут называть поворотной точкой. Потому что Polaris — это каталог. А каталог в мире Iceberg

Accelerating Spark queries with Iceberg materialized views
Amazon iconAmazonSep 10, 2026

Accelerating Spark queries with Iceberg materialized views

Accelerate slow, repetitive Apache Spark analytical queries on Apache Iceberg tables without rewriting any SQL. This post shows how automatic query rewrite in Amazon EMR and AWS Glue uses Iceberg materialized views in the AWS Glue Data Catalog to transparently substitute matching query plans, and how to design material

Почему база не видит ваш предагрегат
Habr iconHabrSep 3, 2026

Почему база не видит ваш предагрегат

Инженеры данных построили агрегат — маленькую таблицу «продажи по магазинам по дням». Отчёт из неё собирается за доли секунды. А сводная в Excel всё равно ждёт двадцать секунд и читает миллиард строк. Разбираемся на живом ClickHouse, почему база не видит предагрегат, который для неё построили, какая форма запроса это л

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app