← all topics

Data Engineering

Batch and Streaming

Bounded versus unbounded data. Latency you can promise against complexity you have to operate.

Levels foundation / engineer / advanced
Depth 2
Time 4h
Kind concept

Grasp

A batch job is handed a set of data that has an end. Yesterday's orders, this hour's log files, every row in a table: the input is bounded, so the job can read all of it, compute an answer, write it once and exit. If it fails, you run it again. That rerunnability is the quiet reason batch is still how most of the world's data is processed.

A stream has no end. Events arrive continuously and the job never finishes, so every question that batch answered by waiting has to be answered another way. Aggregating over a bounded set becomes aggregating over a window, and windows force you to decide what to do about an event that belongs to a window you already closed, because events arrive late, out of order, or twice. State now lives inside a long-running job rather than in the tables it reads, so it has to be checkpointed, restored after a crash, and migrated when the code changes.

The choice is not a technical preference; it is a latency promise you are willing to operate. Genuinely ask what happens if the number is an hour old. If nothing happens, batch it, because a batch job that fails at three in the morning can wait until morning and a streaming job usually cannot.