Modern Data Lakehouse

Module 18 · Core

Getting data in

Batch loads, file-arrival triggers and streaming writes, and the tools that do them.

In plain words

Before data can be analysed it has to arrive: in nightly batches, or continuously as a stream. This module follows data from apps and databases into lakehouse tables, and shows the trade-off between freshness and file sizes.

Helps to take first

Words you'll meet

This module is in production

Here's what it will cover.

Centrepiece

Follow one app event from a phone to a lakehouse table, then tune a streaming writer and watch latency fall as small files pile up

You will understand

  • Batch vs micro-batch vs continuous ingestion
  • Spark Structured Streaming, Flink, Kafka Connect
  • Exactly-once sinks and idempotent writes
  • Managed ingestion: Auto Loader, Firehose, Datastream, Airbyte, Fivetran

Formats

Scroll storyLive simulationBuild & connectCheckpoints