← all topics

AI Engineering

Inference Optimisation

Quantisation, caching, speculative decoding. Buying back latency and margin.

Levels advanced / architect
Depth 12
Time 7h
Kind technique
On AI Architect
This node is a stub. It exists on the map because the topic is real and its place in the graph is settled, but the writing has not happened yet. The sources on the right are the honest starting point until it does.