2026/08/05 by Aaditya Mehta, Arya Shah
Computer Science · #cs.AI
12 pages, 6 figures, 3 tables
arxiv created 2026/08/05 · arxiv updated 2026/08/06
Cooperative multi-agent reinforcement learning often adds social terms to individual rewards, yet the scale of those terms is usually chosen by hand. We ask whether a guilt signal can instead be calibrated from human neural and behavioural data and transferred to artificial agents. Using the public SoDec responsibility fMRI dataset (40 participants), we fit a subject-fixed-effects regression of momentary-happiness changes on outcome-type counts and recover a guilt weight as the Partner-negative minus Social-negative contrast (w=1.118, Cohen's d=0.214). We embed this weight in a two-agent Social Lottery environment and train independent Proximal Policy Optimization actor-critics under four shaping regimes: neurally calibrated, uniform constant, zero (selfish), and a unit-coefficient oracle. Across 1,000 evaluation episodes per condition, the calibrated agents track the human Social safe-choice rate most closely (0.459 vs. human 0.484; KL=0.0012), while the other three conditions deviate by one to three orders of magnitude in KL. Human neurobehavioural priors can therefore act as quantitative constraints on prosocial reward shaping.