← all topics

Large Language Models

RLHF & Preference Tuning

Aligning output with human judgement when there is no correct answer to train on.

Levels advanced / architect
Depth 9
Time 7h
Kind technique
This node is a stub. It exists on the map because the topic is real and its place in the graph is settled, but the writing has not happened yet. The sources on the right are the honest starting point until it does.