Large Language Models
RLHF & Preference Tuning
Aligning output with human judgement when there is no correct answer to train on.
This node is a stub. It exists on the map because the topic is real and its place in the graph is settled, but the writing has not happened yet. The sources on the right are the honest starting point until it does.