Reinforcement Learning from Human Feedback
SyllabusScience and Technology: effects in everyday life
Reinforcement learning from human feedback, or RLHF, is a post-training method that teaches a language model which responses people prefer. Instead of learning only to predict text, the model is optimized to produce answers that receive a higher human-derived reward, thereby shaping its tone, helpfulness, refusals and compliance with instructions.
How the feedback becomes a learning signal
A typical RLHF process converts subjective human judgements into a numerical objective that can guide model behaviour.
- Human evaluators compare or rank alternative responses to the same prompts, creating preference data.
- These comparisons are used to train a reward model that predicts which outputs people are likely to prefer.
- The language model is then treated as a policy and updated through reinforcement learning to increase its predicted reward.
- Because optimization generalizes beyond the rated examples, RLHF influences responses to new prompts rather than merely storing approved answers.
Effects on conversational behaviour
RLHF can make a model more responsive to instructions and social expectations, but it optimizes predicted approval rather than truth itself.
- It can encourage helpful and coherent answers, polite language, relevant explanations and appropriate refusal of unsafe requests.
- The model may learn preferred conversational styles, such as acknowledging uncertainty or adapting detail to the user.
- If evaluators favour agreeable answers, the model may display sycophancy, echoing a user's assumptions instead of correcting them.
- Fluent or confident falsehoods can persist because a favourable rating is not equivalent to factual verification.
Limitations and safeguards
The resulting behaviour depends on who supplies feedback, how prompts are sampled and how the reward objective is designed.
- A proxy reward can be exploited, producing responses that appear satisfactory without fulfilling the intended goal.
- Annotator preferences may transmit cultural or demographic biases and may not represent all users.
- Performance can weaken on unfamiliar situations because feedback covers only a limited portion of possible conversations.
- Diverse evaluators, independent testing, red-teaming and continuous human oversight can reveal failures that reward optimization misses.
Keep reading
The news behind topics like this, explained every day
Every day Gyaanam reads The Hindu, the Indian Express and PIB and picks what matters for UPSC. Each story is written up against the syllabus line it belongs to. Your first 7 days or 20 articles are free, whichever ends first.