Data Pipelines
In simple words
Moving data from where it is produced to where it is asked about, reliably and on a schedule.
The fuller explanation
A pipeline reads from somewhere that was not designed for you, reshapes what it finds, and writes it somewhere that will be queried by people who were not in the room. The reading and writing are easy. Everything expensive comes from the word repeatedly.
Run it once and a script is enough. Run it every hour for a year and four questions arrive that a script has no answer to. What happens when the source is late, and the job runs against yesterday's data as though it were today's? What happens when it fails halfway, and half the rows are already written? What happens when someone finds a bug and the last ninety days must be reprocessed? What happens when the source adds a column, or renames one? Each of those is a design decision, and making them explicitly is the difference between a pipeline and a script on a schedule.
The one worth internalising first is idempotency: running the same load twice should leave the same result as running it once. Delete and insert by partition, or merge on a key, never a bare insert. Every pipeline is eventually retried, usually by someone who has been awake too long, and a pipeline that cannot survive being run twice will one day be run twice.
Learn these first
Real prerequisites, taken from the map rather than guessed.
Sources
Where this came from, so you can go past us.
Where does this sit on your route?
The free assessment places you on the same map and names which terms stand between you and the role you want.
Take the free assessmentSee it in context
The Atlas shows this term with everything that leads into it and everything that follows, as one picture.
Open the map