Skip to content
Frank Vitetta

N° 001

Field note / Table formats

Open table formats explained: Iceberg, Delta Lake and Hudi

A plain-language guide to the metadata layer that turns a pile of files into a dependable table. It also covers where the three main formats stand in October 2026.

Author
Published
Last updated
Reading time
6 min

§ 1 The problem a table format solves

Most large analytics platforms now keep their data as files in low-cost cloud storage instead of inside a single database. The 2021 research paper that named the lakehouse describes how data lakes became the usual landing place for raw data. That means cheap storage holding files in open formats such as Apache Parquet.[1] Parquet itself is only a file format. It is a column-oriented way of laying out data for efficient storage and retrieval.[2]

A folder of files is not a table, though. Nothing in the folder says which files make up the current version. Nothing says what should happen when two jobs write at once or how to undo a bad load.

An open table format closes that gap. It is a published specification for a layer of metadata. That metadata records which files belong to the table, in which version and under which schema. So many tools can treat the same files as one consistent table.

§ 2 What "lakehouse" means

The term comes from a paper presented at the CIDR conference in January 2021 by Michael Armbrust, Ali Ghodsi, Reynold Xin and Matei Zaharia. They define a lakehouse by three properties. The first is open data formats that tools can read directly. The second is first-class support for machine learning and data science. The third is performance that competes with a data warehouse.[1] The authors list Databricks as an affiliation. Databricks sells a lakehouse platform. So read the paper as an argument as well as a description.

In plain terms, a lakehouse is warehouse-style tables kept in lake storage and readable by more than one engine. The table format is the part that makes this possible.

§ 3 How the three formats keep track of a table

Apache Iceberg tracks individual data files. It does not track directories. Its specification describes a tree. A table metadata file lists snapshots. Each snapshot points to a manifest list. Manifests list the data files along with statistics about each one. A commit replaces the old metadata file with a new one in a single atomic swap. So a reader always sees a complete snapshot and never a half-finished write.[3]

Delta Lake keeps an ordered transaction log beside the data. Its protocol document says writers first write new data files, then commit by adding a log entry that records which files were logically added and removed. Readers use the log to pick one consistent snapshot.[4]

Apache Hudi records every action on a timeline, which its documentation calls the source of truth for the table's state.[5] Hudi offers two table types. Copy-on-write rewrites files when rows change and favours read speed. Merge-on-read appends changes to log files and merges them later, which favours fast writes.[6]

Plate 001.1Conceptual drawingScroll sideways

How an Apache Iceberg table is layered Query engines ask a catalog for a table. The catalog points to the current table metadata file. The metadata file points to one manifest list per snapshot. The manifest list points to manifest files, which list the data files held in object storage. The metadata file, manifest list and manifests together form the table format. The data files are the storage layer. Query engines Spark, Trino, Flink, DuckDB, warehouses Engines do not list folders. They follow the pointers down. Catalog finds the current metadata file Table metadata file schema, partitions, snapshots Manifest list one per snapshot Manifest files data file paths and statistics Data files (Parquet) Table format the metadata layer Storage layer files in object storage
Figure 1 The layers of an Iceberg table, drawn from the structure described in the Iceberg specification. Delta Lake and Hudi fill the middle layer differently. Delta uses an ordered transaction log. Hudi uses a timeline of actions.
Table 1. The three formats side by side, as of 2 October 2026
AspectApache IcebergDelta LakeApache Hudi
Project homeApache Software Foundation[7]Linux Foundation[8]Apache Software Foundation[9]
How table state is recordedMetadata file, manifest list and manifests, committed by an atomic swap[3]An ordered transaction log[4]A timeline of actions[5]
Row-level changesDelete files in version 2, deletion vectors in version 3[3]Deletion vectors and row tracking are part of the protocol[4]Copy-on-write or merge-on-read table types[6]
Latest release1.12.0, 30 September 2026[10]4.4.0, 20 August 2026[11]1.2.1, 24 September 2026[12]

§ 4 Where the formats stand in October 2026

Checked on 2 October 2026 against the projects' and vendors' own pages.

Iceberg: version 3 is settled, version 4 is being built

The Iceberg specification states that format versions 1, 2 and 3 are complete and adopted. It states that version 4 is under active development and has not been formally adopted. Version 3 added new data types, default column values, row lineage and binary deletion vectors. The new types include a variant type for semi-structured data, plus geometry and geography.[3] The 1.12.0 release of 30 September 2026 includes early building blocks for version 4, such as a reader for its new manifest layout.[10]

Vendor support for version 3 arrived during 2026. Snowflake's release notes record general availability on 7 May 2026.[13] Databricks announced a public preview on 9 April 2026.[14] Its documentation now lists version 3 features on Databricks Runtime 18 LTS and above. A few items, such as nanosecond timestamps, are marked as unsupported.[15] Amazon describes its S3 Tables service as fully managed Iceberg tables.[16]

Delta Lake: catalog-managed tables and a bridge to Iceberg

Delta Lake 4.4.0 was released on 20 August 2026. It added support for Apache Spark 4.2 and extended its integration with the Unity Catalog API for catalog-managed tables.[11] Delta's UniForm feature generates Iceberg and Hudi metadata alongside the Delta log so that other clients can read the same Parquet files. The open source documentation is clear that those outside clients get read-only access.[17]

Hudi: leaning into AI workloads

Hudi 1.2 was announced on 7 June 2026. It added a vector type with built-in vector search, a blob type for unstructured files and support for the Lance file format.[18] Version 1.2.1 followed on 24 September 2026.[12]

Convergence

The formats are growing closer. In June 2024 Databricks, where Delta Lake was created, agreed to buy Tabular. Tabular is a company founded by the original creators of Iceberg. Databricks said it would work towards compatibility between the two formats.[19] Apache XTable, an incubating project, translates metadata between all three. It stresses that it is not a new format.[20] Iceberg also defines a REST API for catalogs. It was created so that engines and catalogs written in different languages can work together.[21] Apache Polaris is an open source catalog that implements it.[22]

My reading, offered as opinion, is that the question has shifted from which format wins to which catalog you trust. Some form of Iceberg compatibility is now on offer from each of the large vendors named above.

§ 5 What this means if you are not an engineer

When a vendor or an internal team proposes a lakehouse, these are the questions I would start with.

  • Which format and which version? Newer features only help if every tool that touches the table can read them. The Iceberg specification itself notes that tables can stay on an older version until processing engines catch up.[3]
  • Which catalog holds the tables? The catalog is where engines go to find a table, so in my view it is where lock-in now tends to live.
  • Who can write? A bridge such as UniForm is read-only for outside clients.[17]

For marketing and analytics teams the benefit is simple to state. If the campaign, customer and revenue tables live once in an open format, every tool can read the same rows. That means the dashboard, the analyst's notebook and any AI system.

You may not need a cluster to read them either. DuckDB, a single-machine engine, can read Iceberg tables and write to them through a catalog.[23] None of this fixes poor inputs.

Sources

Every source listed here was opened and checked on 2 October 2026.

  1. Armbrust, Ghodsi, Xin and Zaharia, "Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics", CIDR 2021 (PDF). cidrdb.org/cidr2021/papers/cidr2021_paper17.pdf
  2. Apache Parquet project home page. parquet.apache.org
  3. Apache Iceberg Table Spec, source file in the project repository (rendered at iceberg.apache.org/spec/). github.com/apache/iceberg/blob/main/format/spec.md
  4. Delta Transaction Log Protocol. github.com/delta-io/delta/blob/master/PROTOCOL.md
  5. Apache Hudi documentation, Timeline. hudi.apache.org/docs/timeline
  6. Apache Hudi documentation, Table types. hudi.apache.org/docs/table_types
  7. Apache Iceberg project home page. iceberg.apache.org
  8. Delta Lake project home page. delta.io
  9. Apache Hudi documentation, Overview. hudi.apache.org/docs/overview
  10. Apache Iceberg 1.12.0 release notes, 30 September 2026. github.com/apache/iceberg/releases/tag/apache-iceberg-1.12.0
  11. Delta Lake 4.4.0 release notes, 20 August 2026. github.com/delta-io/delta/releases/tag/v4.4.0
  12. Apache Hudi releases, including 1.2.1 of 24 September 2026. github.com/apache/hudi/releases
  13. Snowflake release notes, "Support for Apache Iceberg version 3 (General availability)", 7 May 2026. docs.snowflake.com/release-notes/2026/other/2026-05-07-iceberg-v3-ga
  14. Databricks blog, Apache Iceberg v3 public preview announcement, 9 April 2026. databricks.com/blog/next-era-open-lakehouse-apache-icebergtm-v3-public-preview-databricks
  15. Databricks documentation, Apache Iceberg v3 features. docs.databricks.com/aws/en/iceberg/iceberg-v3
  16. Amazon Web Services, Amazon S3 Tables feature page. aws.amazon.com/s3/features/tables/
  17. Delta Lake documentation, Universal Format (UniForm). docs.delta.io/latest/delta-uniform.html
  18. Apache Hudi blog, release 1.2 announcement, 7 June 2026. hudi.apache.org/blog/2026/06/07/apache-hudi-release-1-2-announcement/
  19. Databricks press release on the agreement to acquire Tabular, 4 June 2024. databricks.com/company/newsroom/press-releases/databricks-agrees-acquire-tabular-company-founded-original-creators
  20. Apache XTable (incubating) project home page. xtable.apache.org
  21. Apache Iceberg, REST Catalog Spec. iceberg.apache.org/rest-catalog-spec/
  22. Apache Polaris project home page. polaris.apache.org
  23. DuckDB documentation, Iceberg extension overview. duckdb.org/docs/current/core_extensions/iceberg/overview