In RLHF, humans rate different model responses, and a reward model is learned from these preferences. The language model is then optimized with reinforcement learning to produce preferred responses. This makes models more helpful, polite, and safe in line with human expectations. RLHF was a key step that turned raw language models into useful assistants.
