Data engineering
Modern Data Lakehouse
Open files, open tables, any engine.
From object storage and Parquet up through Delta Lake, Apache Iceberg and Apache Hudi, to catalogs, pipelines, query engines and real platforms on AWS, Google Cloud, Azure and open source. By the end you can reason about how a lakehouse behaves, and design one.
9 chapters · 29 modules · ~15 hours
- 1Foundations4
- 2Open table formats5
- 3How tables behave3
- 4Performance & layout3
- 5Catalogs & governance2
- 6Building pipelines3
- 7Querying & serving3
- 8Platforms & ecosystem4
- 9Capstone2
A taste of module 6: time travel on the Delta transaction log. Click any version.
Files in storage · 3 in the table at version 0
- added by this commit
- in the table
- removed by this commit
- still in storage for time travel
Removed files are only unlinked from the log, not deleted. That is why VERSION AS OF works, and why VACUUM exists.
* orders 0