Observability & Evals
In simple words
Tracing, logging and regression suites. Knowing you broke it before a user tells you.
The fuller explanation
An ordinary service tells you when it breaks. It throws, the status code changes, the dashboard turns red. A system built on a language model usually does not: it returns a fluent, well-formed, confident answer that happens to be wrong, and every metric you already had stays green. Observability and evals are the two halves of the answer to that.
Observability is what happened in production, per request: the prompt as it was actually assembled, the context that was retrieved, the tool calls and their results, the tokens, the latency and the cost. Without the retrieved context stored alongside the answer, a complaint about a bad response is unreproducible, because you cannot see what the model was looking at. Traces are what turn an anecdote into a bug report.
Evals are the other direction: a fixed set of cases you run before shipping, so a change is judged rather than hoped about. They matter more here than in ordinary software because the system has no compiler and no type checker, and because a prompt edit, a model version, a chunking change or a reranker swap can each move quality in ways nobody predicted. The discipline is unglamorous and it is the difference between improving a system and merely changing it.
Learn these first
Real prerequisites, taken from the map rather than guessed.
Sources
Where this came from, so you can go past us.
Where does this sit on your route?
The free assessment places you on the same map and names which terms stand between you and the role you want.
Take the free assessmentSee it in context
The Atlas shows this term with everything that leads into it and everything that follows, as one picture.
Open the map