LOVE IT OR HATE IT. LET'S TALK ABOUT IT.
BREAKING
GENERAL September 5, 2026

How AI is Trained by Rewards: A Deep Dive


Support the channel for $2.50:

How is an artificial intelligence “rewarded” during training? Does ChatGPT receive a digital cookie for a correct answer—or punishment when it gets something wrong?

In this video, I explain how large language models are trained using pretraining, loss, backpropagation, supervised fine-tuning and reinforcement learning from human feedback. We’ll also examine hallucinations, reward hacking and why an AI doesn’t feel pride, disappointment or satisfaction when it receives a high score.

The reward is just a number. The AI doesn’t enjoy it. In fact, it doesn’t even know the reward exists.

Source: Ryan McBeth

DISCUSSION

Leave a Reply

Your email address will not be published. Required fields are marked *

Quantum24 Network