Skip to content
Frank Vitetta

Reference

Index of terms: big data concepts in plain language

Twenty terms that come up whenever data platforms are discussed. Each definition is short and names its source. Where a field note explains the term in context, the definition links to it.

Author
Last updated
Terms
20

Find a term

Definitions

01Big data#

A loose label for datasets too large or awkward for ordinary tools. There is no agreed threshold. One working definition, used by Jordan Tigani, is whatever does not fit on a single machine. That means the bar rises as machines grow.

Source: Tigani, "Big Data is Dead".

02Data warehouse#

A central database that collects data from operational systems in a structured form. It can then be used for reporting and business intelligence.

Source: Armbrust and others, CIDR 2021. Explained in context in Open table formats explained.

03Data lake#

Low-cost storage that holds raw data as files, usually in open formats such as Parquet. The data does not have to be structured before it is stored.

Source: Armbrust and others, CIDR 2021. Explained in context in Open table formats explained.

04Lakehouse#

An architecture that keeps warehouse-style tables in data lake storage. The paper that named it defines it by three properties. Tools can read its open data formats directly. It has first-class support for machine learning and data science. Its performance competes with a warehouse.

Source: Armbrust and others, CIDR 2021. Explained in context in Open table formats explained.

05Apache Parquet#

An open, column-oriented file format designed for efficient data storage and retrieval. It is a format for individual files. It does not cover whole tables.

Source: Apache Parquet home page. Explained in context in Open table formats explained.

06Open table format#

A published specification for the metadata that lets a collection of data files behave as one table. It gives the table versions, a schema and safe concurrent writes. Apache Iceberg, Delta Lake and Apache Hudi are the main examples.

Source: Apache Iceberg Table Spec. Explained in context in Open table formats explained.

07Catalog#

The service an engine asks in order to find a table. It keeps track of each table's current metadata. Iceberg defines a REST API for catalogs so that engines and catalogs from different suppliers can work together.

Source: Apache Iceberg REST Catalog Spec. Explained in context in Open table formats explained.

08Snapshot#

The state of a table at one point in time. Readers work from a snapshot, so they are not affected by a write that is still in progress.

Source: Apache Iceberg Table Spec. Explained in context in Open table formats explained.

09ACID transaction#

A change to data that is applied completely or not at all. Readers never see it part-way through. The letters stand for atomicity, consistency, isolation and durability. The Delta protocol describes its purpose as bringing these properties to large collections of files in shared storage.

Source: Delta Transaction Log Protocol. Explained in context in Open table formats explained.

10Schema evolution#

Changing a table's columns over time, for example adding, dropping, renaming or reordering them. It is done in a way that stays safe for the data already stored.

Source: Apache Iceberg Table Spec. Explained in context in Open table formats explained.

11Copy-on-write and merge-on-read#

Two ways of handling changed rows. Copy-on-write rewrites the affected files at write time, which keeps reads fast. Merge-on-read appends changes to log files and combines them later, which keeps writes fast.

Source: Apache Hudi documentation, Table types. Explained in context in Open table formats explained.

12Deletion vector#

A compact record of which rows in a data file have been deleted. It is kept separately so that the data file does not have to be rewritten. Deletion vectors are part of Iceberg format version 3 and of the Delta protocol.

Source: Apache Iceberg Table Spec. Explained in context in Open table formats explained.

13OLAP (analytical processing)#

Short for online analytical processing. Workloads made of complex, fairly long-running queries that read large parts of a dataset, such as aggregations over whole tables or joins between large tables.

Source: DuckDB, "Why DuckDB".

14Vectorised execution#

A way of running queries in which the engine processes a large batch of values from a column in one operation. It does not handle one row at a time. It reduces the processor work spent on each value.

Source: DuckDB, "Why DuckDB".

15In-process database#

A database engine that runs embedded inside the application that uses it, with no separate server to install or maintain. DuckDB is an example built for analytics.

Source: DuckDB, "Why DuckDB".

16Distributed cluster#

A group of machines that share one job. In Apache Spark, a driver program coordinates the work, a cluster manager allocates resources and executor processes on worker nodes carry out the tasks.

Source: Apache Spark, Cluster Mode Overview.

17Data quality#

Fitness for purpose. Data is of good quality when it meets the needs of the people using it for a particular task. The level required depends on the use.

Source: UK Government Data Quality Framework.

18Data quality dimension#

A feature of data that can be measured or assessed against defined standards to judge its quality. The six core dimensions named by DAMA UK are completeness, uniqueness, timeliness, validity, accuracy and consistency.

Source: DAMA UK, Defining Data Quality Dimensions (2013).

19Data cascade#

A term from a 2021 Google research study. It describes compounding events in which a data problem causes negative effects further downstream, often surfacing late.

Source: Sambasivan and others, 2021.

20Retrieval-augmented generation#

A design in which a language model combines what it learned in training with documents retrieved from an external index. It retrieves them when it answers. Often shortened to RAG.

Source: Lewis and others, 2020.