What Humans Want: Models of Preference and Reward in Minds and Machines

AI alignment rests on a simple model of human decision-making: people are noisily rational agents maximizing fixed rewards that exist in the environment. This talk examines that model from both sides: minds and machines. In humans, choices follow resource-rational heuristics rather than optimal policies, and pursuing even an arbitrarily assigned goal devalues the alternatives. In machines, reward models trained on human preferences show striking biases, many inherited from pretraining rather than from the preferences themselves. Finally, LLMs share humans’ inability to “un-know” biasing information, but, unlike humans, can converse with a blind, counterfactual self, substantially mitigating demographic bias and sycophancy.
Speaker: Brian Christian, UC Berkeley
Attend in person or watch online (see weblink)
Wednesday, 10/07/26
Contact:
Website: Click to VisitCost:
FreeSave this Event:
iCalendarGoogle Calendar
Yahoo! Calendar
Windows Live Calendar
