Modern Data Lakehouse
Module 18 · Core
Getting data in
Batch loads, file-arrival triggers and streaming writes, and the tools that do them.
In plain words
Before data can be analysed it has to arrive: in nightly batches, or continuously as a stream. This module follows data from apps and databases into lakehouse tables, and shows the trade-off between freshness and file sizes.
Helps to take first
Words you'll meet
This module is in production
Here's what it will cover.
Centrepiece
Follow one app event from a phone to a lakehouse table, then tune a streaming writer and watch latency fall as small files pile up
You will understand
- Batch vs micro-batch vs continuous ingestion
- Spark Structured Streaming, Flink, Kafka Connect
- Exactly-once sinks and idempotent writes
- Managed ingestion: Auto Loader, Firehose, Datastream, Airbyte, Fivetran
Formats
Scroll storyLive simulationBuild & connectCheckpoints