Data engineering

Modern Data Lakehouse

Open files, open tables, any engine.

From object storage and Parquet up through Delta Lake, Apache Iceberg and Apache Hudi, to catalogs, pipelines, query engines and real platforms on AWS, Google Cloud, Azure and open source. By the end you can reason about how a lakehouse behaves, and design one.

9
chapters
29
modules
~15h
of learning
15
live now
29 modules

Take modules in any order. The route suggests one that builds up naturally.

01

Foundations

Where lakehouses come from, and the storage and file formats everything else stands on.

4 modules · 100 min

02

Open table formats

How Delta Lake, Apache Iceberg and Apache Hudi turn folders of files into real tables.

5 modules · 155 min

03

How tables behave

Transactions, row-level changes and schema changes: what really happens underneath.

3 modules · 85 min

04

Performance & layout

Layout decisions that make queries fast and keep storage costs down.

3 modules · 85 min

05

Catalogs & governance

Who knows where every table lives, and who is allowed to read it.

2 modules · 60 min

06

Building pipelines

Getting data in, and shaping it into tables people can trust.

3 modules · 95 min

07

Querying & serving

How engines read tables, and how BI, ML and AI use them.

3 modules · 95 min

08

Platforms & ecosystem

The same ideas on AWS, Google Cloud, Azure and pure open source.

4 modules · 120 min

09

Capstone

Put it all together: design one lakehouse, then fix a broken one.

2 modules · 75 min