Lake and Lakehouse

Data Engineeringsystem

First met at Engineer4h

In simple words

Table formats put warehouse guarantees on object storage. What the lakehouse actually adds over a folder of files.

The fuller explanation

A lake is object storage holding files, usually Parquet, and that is the whole of it. It is cheap, it takes anything, and it has no idea what a table is. Two writers can produce a half-finished state that a reader sees as real, a schema change is whatever the newest file happens to contain, and the current version of the truth is a convention someone documented once.

A lakehouse is the same files with a metadata layer above them: Iceberg, Delta Lake or Hudi. The metadata records which files make up the table right now, so a commit becomes atomic, readers get a consistent snapshot, schemas can evolve without rewriting history, and you can query the table as it stood last Tuesday. The storage stays open and cheap; the guarantees start to look like a warehouse.

What it does not hand you is warehouse performance by default. The maintenance the warehouse was quietly doing becomes yours: streaming writes produce thousands of small files that must be compacted, metadata and manifests grow and need expiring, and old snapshots keep data alive long after you thought you deleted it. That last one matters when a deletion request arrives. A lakehouse is a good trade, not a free one.

Where does this sit on your route?

The free assessment places you on the same map and names which terms stand between you and the role you want.

Take the free assessment

See it in context

The Atlas shows this term with everything that leads into it and everything that follows, as one picture.

Open the map