The tree
Artificial Intelligence at the root, expanding down into domains, topics and techniques. Click any branch to open it.
This is the taxonomy - what each topic is a kind of. It is a tree because every topic has one parent. Learning order is a different structure: a topic can require several others from anywhere in the tree, which is why the Atlas is a map rather than an outline.
Artificial IntelligenceThe root of the Atlas. Everything else in this map is a descendant of this one idea.6
Generative AIModels that produce new artefacts rather than predicting a label. The densest region of the Atlas.5
Large Language ModelsTransformers trained on enough text to become general-purpose reasoning surfaces.9
RAGGive the model the right documents at query time instead of hoping it memorised them.7
- Production RAGIngestion pipelines, freshness, permissions, caching and cost at real volume.
- ChunkingHow you split documents. Where most RAG systems quietly lose their quality.
- RAG EvaluationstubSeparating retrieval failures from generation failures, so you fix the right one.
- Vector SearchstubApproximate nearest neighbours at scale. HNSW, IVF, and the recall/latency trade.
- RerankingstubA second, slower pass over the top results. Usually the largest single quality win available.
- Vector DatabasesstubWhere embeddings live: indexes, filters, and the operational reality of them.
- Hybrid SearchstubCombining keyword and semantic retrieval, because each fails where the other works.
- EmbeddingsMeaning as coordinates. The representation that makes search, clustering and RAG possible.
Fine-TuningAdapting a pretrained model to your task. Often the wrong first answer.2
- LoRA & PEFTstubTrain a few million parameters instead of a few billion, and lose almost nothing.
- RLHF & Preference TuningstubAligning output with human judgement when there is no correct answer to train on.
- Prompt EngineeringSpecifying a task precisely enough that a probabilistic system does it reliably.
- LLM EvaluationMeasuring quality when there is no single correct output. The hardest unsolved problem in shipping.
- Context WindowsstubThe model’s entire working memory, and the constraint that shapes most system design.
- TokenisationstubHow text becomes numbers, and why the model cannot count the letters in a word.
- PretrainingstubThe expensive part: corpus, curriculum, scaling laws, and the compute bill.
- Decoding StrategiesstubTemperature, top-p, beam search. How a distribution becomes the text you see.
- TransformersAttention, feed-forward layers, residuals and normalisation. The block that ate the field.
- AttentionLetting every position look at every other position. The hinge the modern field turns on.
- Multimodal AIstubOne model, several senses. Text, image, audio and video in a shared representation.
- Diffusion ModelsstubLearn to reverse noise, then run it backwards from pure noise to generate images.
Machine LearningSystems whose behaviour is learned from data rather than written as rules.5
Supervised LearningLearning a mapping from labelled examples. The workhorse of applied ML.2
- Trees & EnsemblesstubGradient boosting still wins on tabular data. This is not a historical note.
- Linear & Logistic RegressionstubThe two models you should always try first, and often ship.
Model EvaluationThe skill that separates people who ship models from people who publish notebooks.2
- Overfitting & RegularisationstubWhy a model that memorises the training set is worse than one that does not.
- Cross-ValidationstubGetting an honest performance estimate when you do not have data to spare.
- Gradient DescentThe optimisation procedure underneath essentially every model in this Atlas.
- Feature EngineeringstubShaping raw data into signal. Where domain knowledge earns its keep.
- Unsupervised LearningstubFinding structure in data nobody labelled: clustering, density, dimensionality.
AI AgentsModels that decide what to do next, call tools, and act in a loop rather than answering once.6
- Tool UseFunction calling, schemas, and validating what the model asks you to run.
- Planning & ReasoningstubDecomposition, reflection, and letting the model check its own work.
- Agent MemorystubShort-term scratchpads, long-term stores, and deciding what is worth remembering.
- Agent EvaluationstubJudging a trajectory, not an answer. Compounding error is the thing to measure.
- Model Context ProtocolstubA standard interface between models and the tools and data they need.
- Multi-Agent SystemsstubSeveral specialised agents coordinating. Powerful, and usually premature.
Deep LearningMany-layered neural networks that learn their own features instead of being handed them.3
Neural NetworksLayers, weights and activations. The unit of construction for everything downstream.5
- BackpropagationThe chain rule applied at scale. How a network learns which weights were at fault.
- Convolutional NetworksstubWeight sharing over spatial structure. Still the efficient default for images.
- OptimisersstubAdam, AdamW, schedules and warmup. What you actually tune when training stalls.
- Activation FunctionsstubThe non-linearity without which a deep network collapses into a linear one.
- Recurrent NetworksstubRNNs and LSTMs. Largely superseded, but the reason attention was invented.
- Transfer LearningstubStart from someone else’s trained weights. The economics of modern AI in one idea.
- Training at ScalestubData, tensor and pipeline parallelism, mixed precision, and the memory wall.
AI EngineeringThe bridge from a working notebook to a system real users depend on.7
- AI ArchitectureSystem-level design: model routing, boundaries, failure modes, governance. The summit node.
- Observability & EvalsTracing, logging and regression suites. Knowing you broke it before a user tells you.
- Cost & LatencystubToken economics, model routing and caching. The constraint that decides architecture.
- Guardrails & SafetystubInput validation, output filtering, prompt injection, and where to put the boundary.
- Model ServingstubAPIs, batching, streaming, and the deployment surface between model and product.
- Data StrategystubProvenance, licensing, freshness and permissions. The unglamorous moat.
- Inference OptimisationstubQuantisation, caching, speculative decoding. Buying back latency and margin.
Computer VisionGetting structured meaning out of pixels: classification, detection, segmentation.3
- Vision-Language ModelsstubModels that read an image and talk about it. Where vision rejoined the mainstream.
- Object DetectionstubWhat is in the image, and where. Boxes, anchors and non-maximum suppression.
- Image SegmentationstubPer-pixel classification, for when a bounding box is not precise enough.
- PythonThe only genuinely non-negotiable prerequisite in the entire Atlas.
Data EngineeringDesigning the systems that make data trustworthy, timely and cheap enough for analytics and AI to run on.11
Data PipelinesMoving data from where it is produced to where it is asked about, reliably and on a schedule.3
- ETL and ELTThe transform step moved. That single change is what the warehouse era is actually about.
- Batch and StreamingBounded versus unbounded data. Latency you can promise against complexity you have to operate.
- Change Data CapturestubReading the database log instead of polling the table. How a warehouse stays minutes behind production.
Data WarehouseColumnar storage and separated compute. Why a scan of a billion rows can cost cents or hundreds of dollars.2
- Lake and LakehouseTable formats put warehouse guarantees on object storage. What the lakehouse actually adds over a folder of files.
- Columnar Storage FormatsstubParquet, Iceberg, Delta. Row groups, predicate pushdown, and why layout decides your query bill.
- Data ModellingChoosing the grain. The decision that determines every query anyone will ever write against your tables.
- Cloud Data PlatformsstubStorage, compute, IAM and the bill. Pick one cloud properly rather than three of them badly.
- OrchestrationstubDependencies, retries and backfills. Where a folder of scripts becomes a system you can reason about.
- Distributed ProcessingstubSpark and its relatives. Partitions, shuffles, and the skew that leaves one task running for six hours.
- Data QualitystubTests on data, not just on code. The pipeline that succeeds every night while writing zero rows.
- Pipeline ObservabilitystubFreshness, volume and cost as monitored signals, so a silent failure is loud before a stakeholder finds it.
- Analytics EngineeringstubSoftware practice applied to transformation: version control, tests and review on the models analysts use.
- Data Infrastructure and CI/CDstubContainers, IaC and a deploy that runs your transformations against a staging warehouse before production.
- Governance and LineagestubWho may read this, where it came from, and what breaks if it changes. The questions auditors and outages both ask.
Math for AIThe specific mathematics that appears in practice, and nothing beyond it.3
- Linear AlgebrastubVectors, matrices and the operations every model is secretly made of.
- Probability & StatisticsstubDistributions, expectation, and why every evaluation number needs an error bar.
- Calculus & OptimisationstubDerivatives, the chain rule, and what it means to walk downhill in a loss landscape.
- Data WranglingstubLoading, cleaning and reshaping data. Realistically most of the job.
- SQLstubWhere the training data actually lives, and the fastest way to interrogate it.
- Git & CollaborationstubVersion control, review and reproducibility. Assumed silently by every team.