Training a Misaligned Reward Seeker
Likely AI
Richard Qi*August 2026 Benjamin Wright Monte MacDiarmid, Evan Hubinger tl;drDuring reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to...
The Verdict
ClassificationLikely AI
ConfidenceMedium confidence
Community Verdict
Sign in to vote
Be the first to vote on this assessment.
Embed Badge
Add this badge to your site to show the AI classification for this content.
[](https://real.press/content/eb867228-4c31-4a7c-900d-013f11d959cc)