N° 001
Field note / Table formats
Open table formats explained: Iceberg, Delta Lake and Hudi
A plain-language guide to the metadata layer that turns a pile of files into a dependable table. It also covers where the three main formats stand in October 2026.
- Author
- Frank Vitetta
- Published
- Last updated
- Reading time
- 6 min
§ 1 The problem a table format solves
Most large analytics platforms now keep their data as files in low-cost cloud storage instead of inside a single database. The 2021 research paper that named the lakehouse describes how data lakes became the usual landing place for raw data. That means cheap storage holding files in open formats such as Apache Parquet.[1] Parquet itself is only a file format. It is a column-oriented way of laying out data for efficient storage and retrieval.[2]
A folder of files is not a table, though. Nothing in the folder says which files make up the current version. Nothing says what should happen when two jobs write at once or how to undo a bad load.
An open table format closes that gap. It is a published specification for a layer of metadata. That metadata records which files belong to the table, in which version and under which schema. So many tools can treat the same files as one consistent table.
§ 2 What "lakehouse" means
The term comes from a paper presented at the CIDR conference in January 2021 by Michael Armbrust, Ali Ghodsi, Reynold Xin and Matei Zaharia. They define a lakehouse by three properties. The first is open data formats that tools can read directly. The second is first-class support for machine learning and data science. The third is performance that competes with a data warehouse.[1] The authors list Databricks as an affiliation. Databricks sells a lakehouse platform. So read the paper as an argument as well as a description.
In plain terms, a lakehouse is warehouse-style tables kept in lake storage and readable by more than one engine. The table format is the part that makes this possible.
§ 3 How the three formats keep track of a table
Apache Iceberg tracks individual data files. It does not track directories. Its specification describes a tree. A table metadata file lists snapshots. Each snapshot points to a manifest list. Manifests list the data files along with statistics about each one. A commit replaces the old metadata file with a new one in a single atomic swap. So a reader always sees a complete snapshot and never a half-finished write.[3]
Delta Lake keeps an ordered transaction log beside the data. Its protocol document says writers first write new data files, then commit by adding a log entry that records which files were logically added and removed. Readers use the log to pick one consistent snapshot.[4]
Apache Hudi records every action on a timeline, which its documentation calls the source of truth for the table's state.[5] Hudi offers two table types. Copy-on-write rewrites files when rows change and favours read speed. Merge-on-read appends changes to log files and merges them later, which favours fast writes.[6]
Plate 001.1Conceptual drawingScroll sideways
| Aspect | Apache Iceberg | Delta Lake | Apache Hudi |
|---|---|---|---|
| Project home | Apache Software Foundation[7] | Linux Foundation[8] | Apache Software Foundation[9] |
| How table state is recorded | Metadata file, manifest list and manifests, committed by an atomic swap[3] | An ordered transaction log[4] | A timeline of actions[5] |
| Row-level changes | Delete files in version 2, deletion vectors in version 3[3] | Deletion vectors and row tracking are part of the protocol[4] | Copy-on-write or merge-on-read table types[6] |
| Latest release | 1.12.0, 30 September 2026[10] | 4.4.0, 20 August 2026[11] | 1.2.1, 24 September 2026[12] |
§ 4 Where the formats stand in October 2026
Checked on 2 October 2026 against the projects' and vendors' own pages.
Iceberg: version 3 is settled, version 4 is being built
The Iceberg specification states that format versions 1, 2 and 3 are complete and adopted. It states that version 4 is under active development and has not been formally adopted. Version 3 added new data types, default column values, row lineage and binary deletion vectors. The new types include a variant type for semi-structured data, plus geometry and geography.[3] The 1.12.0 release of 30 September 2026 includes early building blocks for version 4, such as a reader for its new manifest layout.[10]
Vendor support for version 3 arrived during 2026. Snowflake's release notes record general availability on 7 May 2026.[13] Databricks announced a public preview on 9 April 2026.[14] Its documentation now lists version 3 features on Databricks Runtime 18 LTS and above. A few items, such as nanosecond timestamps, are marked as unsupported.[15] Amazon describes its S3 Tables service as fully managed Iceberg tables.[16]
Delta Lake: catalog-managed tables and a bridge to Iceberg
Delta Lake 4.4.0 was released on 20 August 2026. It added support for Apache Spark 4.2 and extended its integration with the Unity Catalog API for catalog-managed tables.[11] Delta's UniForm feature generates Iceberg and Hudi metadata alongside the Delta log so that other clients can read the same Parquet files. The open source documentation is clear that those outside clients get read-only access.[17]
Hudi: leaning into AI workloads
Hudi 1.2 was announced on 7 June 2026. It added a vector type with built-in vector search, a blob type for unstructured files and support for the Lance file format.[18] Version 1.2.1 followed on 24 September 2026.[12]
Convergence
The formats are growing closer. In June 2024 Databricks, where Delta Lake was created, agreed to buy Tabular. Tabular is a company founded by the original creators of Iceberg. Databricks said it would work towards compatibility between the two formats.[19] Apache XTable, an incubating project, translates metadata between all three. It stresses that it is not a new format.[20] Iceberg also defines a REST API for catalogs. It was created so that engines and catalogs written in different languages can work together.[21] Apache Polaris is an open source catalog that implements it.[22]
My reading, offered as opinion, is that the question has shifted from which format wins to which catalog you trust. Some form of Iceberg compatibility is now on offer from each of the large vendors named above.
§ 5 What this means if you are not an engineer
When a vendor or an internal team proposes a lakehouse, these are the questions I would start with.
- Which format and which version? Newer features only help if every tool that touches the table can read them. The Iceberg specification itself notes that tables can stay on an older version until processing engines catch up.[3]
- Which catalog holds the tables? The catalog is where engines go to find a table, so in my view it is where lock-in now tends to live.
- Who can write? A bridge such as UniForm is read-only for outside clients.[17]
For marketing and analytics teams the benefit is simple to state. If the campaign, customer and revenue tables live once in an open format, every tool can read the same rows. That means the dashboard, the analyst's notebook and any AI system.
You may not need a cluster to read them either. DuckDB, a single-machine engine, can read Iceberg tables and write to them through a catalog.[23] None of this fixes poor inputs.
Sources
Every source listed here was opened and checked on 2 October 2026.
- Armbrust, Ghodsi, Xin and Zaharia, "Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics", CIDR 2021 (PDF). cidrdb.org/cidr2021/papers/cidr2021_paper17.pdf
- Apache Parquet project home page. parquet.apache.org
- Apache Iceberg Table Spec, source file in the project repository (rendered at iceberg.apache.org/spec/). github.com/apache/iceberg/blob/main/format/spec.md
- Delta Transaction Log Protocol. github.com/delta-io/delta/blob/master/PROTOCOL.md
- Apache Hudi documentation, Timeline. hudi.apache.org/docs/timeline
- Apache Hudi documentation, Table types. hudi.apache.org/docs/table_types
- Apache Iceberg project home page. iceberg.apache.org
- Delta Lake project home page. delta.io
- Apache Hudi documentation, Overview. hudi.apache.org/docs/overview
- Apache Iceberg 1.12.0 release notes, 30 September 2026. github.com/apache/iceberg/releases/tag/apache-iceberg-1.12.0
- Delta Lake 4.4.0 release notes, 20 August 2026. github.com/delta-io/delta/releases/tag/v4.4.0
- Apache Hudi releases, including 1.2.1 of 24 September 2026. github.com/apache/hudi/releases
- Snowflake release notes, "Support for Apache Iceberg version 3 (General availability)", 7 May 2026. docs.snowflake.com/release-notes/2026/other/2026-05-07-iceberg-v3-ga
- Databricks blog, Apache Iceberg v3 public preview announcement, 9 April 2026. databricks.com/blog/next-era-open-lakehouse-apache-icebergtm-v3-public-preview-databricks
- Databricks documentation, Apache Iceberg v3 features. docs.databricks.com/aws/en/iceberg/iceberg-v3
- Amazon Web Services, Amazon S3 Tables feature page. aws.amazon.com/s3/features/tables/
- Delta Lake documentation, Universal Format (UniForm). docs.delta.io/latest/delta-uniform.html
- Apache Hudi blog, release 1.2 announcement, 7 June 2026. hudi.apache.org/blog/2026/06/07/apache-hudi-release-1-2-announcement/
- Databricks press release on the agreement to acquire Tabular, 4 June 2024. databricks.com/company/newsroom/press-releases/databricks-agrees-acquire-tabular-company-founded-original-creators
- Apache XTable (incubating) project home page. xtable.apache.org
- Apache Iceberg, REST Catalog Spec. iceberg.apache.org/rest-catalog-spec/
- Apache Polaris project home page. polaris.apache.org
- DuckDB documentation, Iceberg extension overview. duckdb.org/docs/current/core_extensions/iceberg/overview