Data Engineering
Data Pipelines
Moving data from where it is produced to where it is asked about, reliably and on a schedule.
Grasp
A pipeline reads from somewhere that was not designed for you, reshapes what it finds, and writes it somewhere that will be queried by people who were not in the room. The reading and writing are easy. Everything expensive comes from the word repeatedly.
Run it once and a script is enough. Run it every hour for a year and four questions arrive that a script has no answer to. What happens when the source is late, and the job runs against yesterday's data as though it were today's? What happens when it fails halfway, and half the rows are already written? What happens when someone finds a bug and the last ninety days must be reprocessed? What happens when the source adds a column, or renames one? Each of those is a design decision, and making them explicitly is the difference between a pipeline and a script on a schedule.
The one worth internalising first is idempotency: running the same load twice should leave the same result as running it once. Delete and insert by partition, or merge on a key, never a bare insert. Every pipeline is eventually retried, usually by someone who has been awake too long, and a pipeline that cannot survive being run twice will one day be run twice.