Skip to content
Data System School · Papers Distilled

The papers, distilled

Every foundational data-systems paper on the Data System Timeline, read end to end and distilled into one page you can actually finish.

33Papers
5Eras
14Categories
2Languages
Timeline
1960s–1970s · Foundations 3 papers
1980s–1990s · Hardening the core 10 papers
Late 1970s–1980s

The Transaction Concept: Virtues and Limitations

Transactions
Jim Gray (Tandem Computers Incorporated, Cupertino, California)

Names atomicity, consistency and durability as the transaction's defining properties, then shows exactly where nested and long-lived transactions break them.

Proceedings of the Seventh International Conference on Very Large Databases, Sept. 1981; also Tandem Technical Report TR 81.3, June 1981 Read →
Late 1970s–1980s

Eight Transaction Papers by Jim Gray

Transactions
Philip A. Bernstein (Microsoft Research)

A retrospective walk through eight Jim Gray papers that built the transaction abstraction, from two-phase locking to Paxos Commit.

Book chapter, ACM Turing Award winners' series (Curiosity, Clarity, and Caring); arXiv preprint, October 2023 Read →
1986 onward

Looking Back at Postgres

Extensible DBMS
Joseph M. Hellerstein, UC Berkeley

A retrospective on Berkeley Postgres, showing how one extensible object-relational design seeded PostgreSQL and a generation of database systems.

arXiv:1901.01973, January 2019; solicited for Michael Stonebraker's Turing Award book, Making Databases Work (Morgan & Claypool, 2019) Read →
1986–1990s

GAMMA - A High Performance Dataflow Database Machine

Parallel DBMS
David J. DeWitt, Robert H. Gerber, Goetz Graefe, Michael L. Heytens, et al. (Krishna B. Kumar, M. Muralikrishna) - Computer Sciences Department, University of Wisconsin

The first working shared-nothing parallel database: one processor per disk, every relation partitioned, queries run as self-scheduling dataflow.

VLDB 1986 (Proceedings of the Twelfth International Conference on Very Large Data Bases, Kyoto, August 1986), pp. 228-237 Read →
1986–1990s

Parallel Database Systems: The Future of High Performance Database Processing

Parallel DBMS
David J. DeWitt (Computer Sciences Department, University of Wisconsin-Madison) and Jim Gray (San Francisco Systems Center, Digital Equipment Corporation)

Shared-nothing hardware plus partitioned data and a split/merge dataflow runtime give relational queries near-linear speedup and scaleup.

CACM 35(6), June 1992 Read →
1992

ARIES: A Transaction Recovery Method Supporting Fine-Granularity Locking and Partial Rollbacks Using Write-Ahead Logging

Transactions
C. Mohan, Don Haderle, Bruce Lindsay, Hamid Pirahesh, et al. (IBM Almaden Research Center and IBM Santa Teresa Laboratory)

ARIES made crash recovery correct and fast under record-level locking by repeating history, then undoing losers with redo-only compensation records.

ACM Transactions on Database Systems 17(1), March 1992, pages 94-162 Read →
1995–1997

Data Cube: A Relational Aggregation Operator Generalizing Group-By, Cross-Tab, and Sub-Totals

Data warehouse / OLAP
Jim Gray, Surajit Chaudhuri, Adam Bosworth, Andrew Layman et al. (Microsoft Research, Redmond, WA), with Hamid Pirahesh and Frank Pellow (IBM Research, San Jose, CA)

The CUBE operator: one SQL clause that computes every group-by over N dimensions at once, and returns it as a relation.

Data Mining and Knowledge Discovery 1(1): 29-53, 1997; Microsoft Technical Report MSR-TR-97-32 (extended abstract at ICDE 1996) Read →
1996

The Log-Structured Merge-Tree (LSM-Tree)

Storage engine
Patrick O'Neil and Elizabeth O'Neil (Dept. of Math & C.S., UMass/Boston), Edward Cheng (Digital Equipment Corporation), Dieter Gawlick (Oracle Corporation)

A disk index that defers and batches inserts into cascading sorted merges, cutting disk-arm cost by nearly two orders of magnitude.

Acta Informatica 33, 1996 Read →
1989–2001

The Part-Time Parliament

Distributed systems
Leslie Lamport (Systems Research Center, Digital Equipment Corporation), with an editorial submission note by Keith Marzullo (University of California, San Diego)

Paxos: a majority-quorum protocol that keeps replicated logs consistent through crashes and lost messages, and progresses when the network settles.

ACM Transactions on Computer Systems 16(2), May 1998 (received January 1990, accepted March 1998) Read →
1989–2001

Paxos Made Simple

Distributed systems
Leslie Lamport (Microsoft Research)

A plain-English derivation of Paxos, showing fault-tolerant consensus follows almost unavoidably from wanting a majority of acceptors to agree.

ACM SIGACT News 32(4), December 2001; manuscript dated 1 Nov 2001 Read →
2000s · The web-scale break 7 papers
2000–2002

Perspectives on the CAP Theorem

Distributed systems
Seth Gilbert (National University of Singapore) and Nancy A. Lynch (Massachusetts Institute of Technology)

The definitive restatement of CAP: not pick two of three, but the impossibility of safety plus liveness on an unreliable network.

IEEE Computer 45(2), 2012 Read →
2003

The Google File System

Distributed storage
Sanjay Ghemawat, Howard Gobioff, and Shun-Tak Leung (Google)

A cluster file system that makes commodity failure routine, files enormous, and concurrent append a first-class atomic operation.

SOSP 2003 Read →
2004

MapReduce: Simplified Data Processing on Large Clusters

Big data processing
Jeffrey Dean and Sanjay Ghemawat, Google, Inc.

A programming model that hides parallelization, fault tolerance, locality and load balancing behind two user-written functions: map and reduce.

OSDI 2004 Read →
2005

C-Store: A Column-oriented DBMS

Data warehouse / OLAP
Mike Stonebraker, Daniel J. Abadi, Adam Batkin, Xuedong Chen, et al. (MIT CSAIL; Brandeis University; UMass Boston; Brown University)

A read-optimized column store built from overlapping sorted projections, compressed columns, and a hybrid write store with snapshot isolation.

VLDB 2005 (31st VLDB Conference, Trondheim, Norway) Read →
2005

The Vertica Analytic Database: C-Store 7 Years Later

Data warehouse / OLAP
Andrew Lamb, Matt Fuller, Ramakrishna Varadarajan, Nga Tran, et al. (Vertica Systems, an HP Company, Cambridge, MA)

The engineering post-mortem of C-Store: which column-store research ideas survived seven years of paying customers, and which were dropped.

PVLDB 5(12), VLDB 2012 Read →
2008–2010

Cassandra - A Decentralized Structured Storage System

Distributed storage
Avinash Lakshman and Prashant Malik, Facebook

A production store that fused Dynamo's leaderless ring with Bigtable's column families to absorb billions of writes a day.

LADIS 2009 (ACM SIGOPS Workshop on Large Scale Distributed Systems and Middleware); reprinted in ACM SIGOPS Operating Systems Review 44(2), 2010 Read →
2009

Hive - A Warehousing Solution Over a Map-Reduce Framework

Big data processing
Ashish Thusoo, Joydeep Sen Sarma, Namit Jain, Zheng Shao, et al. (Facebook Data Infrastructure Team)

A SQL-like warehouse over Hadoop: HiveQL compiles into map-reduce DAGs, backed by a catalog, partitions, buckets and pluggable SerDes.

VLDB 2009, Lyon, France (demonstration paper) Read →
2010s · Cloud, streams and consensus 8 papers
2010

Spark: Cluster Computing with Working Sets

Big data processing
Matei Zaharia, Mosharaf Chowdhury, Michael J. Franklin, Scott Shenker, et al., University of California, Berkeley

Resilient distributed datasets: cached, lineage-recoverable collections that let clusters reuse a working set in memory across many operations.

HotCloud 2010 (2nd USENIX Workshop on Hot Topics in Cloud Computing) Read →
2011

Kafka: a Distributed Messaging System for Log Processing

Streaming
Jay Kreps, Neha Narkhede, Jun Rao (LinkedIn Corp.)

A distributed commit log for high-volume event data: partitioned append-only segments, pull-based consumers holding their own offsets, zero-copy delivery.

NetDB 2011 (Workshop on Networking Meets Databases), Athens, Greece, June 2011 Read →
2012

Spanner: Google's Globally-Distributed Database

Distributed SQL
James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, et al. (23 authors), Google, Inc.

The first database to give globally distributed transactions external consistency, by exposing clock uncertainty in the time API and waiting it out.

OSDI 2012 (10th USENIX Symposium on Operating Systems Design and Implementation) Read →
2012

Calvin: Fast Distributed Transactions for Partitioned Database Systems

Distributed SQL
Alexander Thomson, Thaddeus Diamond, Shu-Chun Weng, Kun Ren, et al. — Yale University

Order transactions deterministically before executing them, and a partitioned database can drop two-phase commit entirely.

SIGMOD 2012 Read →
2012–2013

Presto: SQL on Everything

Query processing
Raghav Sethi, Martin Traverso, Dain Sundstrom, David Phillips, et al. (Facebook, Inc.)

One adaptive distributed SQL engine that serves sub-second dashboards and multi-hour ETL over dozens of pluggable data sources.

ICDE 2019 Read →
2012–2013

Evolution of Development Priorities in Key-value Stores Serving Large-scale Applications: The RocksDB Experience

Storage engine
Siying Dong, Andrew Kryczka, Yanqin Jin (Facebook Inc.) and Michael Stumm (University of Toronto)

Eight years of running RocksDB at Facebook scale, and how that moved its optimization target from write amplification to space to CPU.

FAST 2021 (19th USENIX Conference on File and Storage Technologies), February 2021 Read →
2014

In Search of an Understandable Consensus Algorithm (Extended Version)

Distributed systems
Diego Ongaro and John Ousterhout (Stanford University)

A leader-based consensus algorithm designed for understandability, equivalent to multi-Paxos in safety and efficiency but teachable and implementable.

Stanford tech report, published May 20, 2014; extended version of the paper in USENIX ATC 2014 Read →
2014–2016

The Snowflake Elastic Data Warehouse

Cloud data systems
Benoit Dageville, Thierry Cruanes, Marcin Zukowski, Vadim Antonov, et al. (Snowflake Computing)

Split a warehouse into blob storage, ephemeral compute clusters and a shared metadata brain, making elasticity an architectural property.

SIGMOD/PODS 2016, San Francisco, CA, USA Read →
2020–2026 · Lakehouse and the AI-native era 5 papers
2016–2020, mainstream in 2020s

Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores

Lakehouse
Michael Armbrust, Tathagata Das, Liwen Sun, Burak Yavuz, et al. (Databricks, with CWI, UC Berkeley and Stanford University)

ACID tables on plain cloud object stores, built from a Parquet-checkpointed write-ahead log that needs no always-on metadata service.

PVLDB 13(12), 2020 (VLDB 2020) Read →
2016–2020, mainstream in 2020s

Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics

Lakehouse
Michael Armbrust, Ali Ghodsi, Reynold Xin, Matei Zaharia (Databricks; UC Berkeley; Stanford University)

A blueprint for running warehouse-grade transactions, indexing and SQL performance directly over open Parquet files in cloud object storage.

CIDR 2021 (11th Annual Conference on Innovative Data Systems Research), Online, January 2021 Read →
2019–2020s

DuckDB: an Embeddable Analytical Database

Query processing
Mark Raasveldt and Hannes Mühleisen (CWI, Amsterdam)

An in-process SQL engine that brings vectorized OLAP execution to the embedded niche SQLite left empty.

SIGMOD 2019 (demonstration paper), Amsterdam, Netherlands Read →
2009; open-sourced 2018

FoundationDB: A Distributed Unbundled Transactional Key Value Store

Distributed SQL
Jingyu Zhou, Meng Xu, Alexander Shraer, Bala Namasivayam, et al. - 21 authors from Apple Inc., Snowflake Inc., and antithesis.com

Serializable ACID transactions at NoSQL scale, built by unbundling the database and proving every feature correct in deterministic simulation.

SIGMOD 2021 (ACM SIGMOD International Conference on Management of Data), June 2021 Read →
2025

Disaggregated State Management in Apache Flink 2.0

Streaming
Yuan Mei, Zhaoqian Lan, Lei Huang, Yanfei Lei, et al. (Alibaba Group; Boston University; KTH Royal Institute of Technology)

Flink 2.0 makes remote storage the primary home of streaming state, hiding its latency with asynchronous, out-of-order record execution.

PVLDB 18(12), 2025 (VLDB 2025) Read →
No papers match those filters.