Support the channel for $2.50:
How is an artificial intelligence “rewarded” during training? Does ChatGPT receive a digital cookie for a correct answer—or punishment when it gets something wrong?
In this video, I explain how large language models are trained using pretraining, loss, backpropagation, supervised fine-tuning and reinforcement learning from human feedback. We’ll also examine hallucinations, reward hacking and why an AI doesn’t feel pride, disappointment or satisfaction when it receives a high score.
The reward is just a number. The AI doesn’t enjoy it. In fact, it doesn’t even know the reward exists.
Source: Ryan McBeth