Track
Data Engineer
“I want to become a data engineer, and then an AI engineer.”
The least contested entrance to AI work. Ends with someone who can model, move and serve data a team can trust - and who is one short step from RAG, because chunking and ingestion are ETL problems wearing AI costumes.
- Python
The only genuinely non-negotiable prerequisite in the entire Atlas.
- SQL
Where the training data actually lives, and the fastest way to interrogate it.
- Data Modelling
Choosing the grain. The decision that determines every query anyone will ever write against your tables.
- Data Pipelines
Moving data from where it is produced to where it is asked about, reliably and on a schedule.
- Data Warehouse
Columnar storage and separated compute. Why a scan of a billion rows can cost cents or hundreds of dollars.
- Distributed Processing
Spark and its relatives. Partitions, shuffles, and the skew that leaves one task running for six hours.
- Orchestration
Dependencies, retries and backfills. Where a folder of scripts becomes a system you can reason about.
- Data Engineering
Designing the systems that make data trustworthy, timely and cheap enough for analytics and AI to run on.