2026/07/22 by Hongju Pae
#cs.LG #q-bio.NC
A recent line of work measures causal emergence in reinforcement learning agents through Integrated Information Decomposition, reporting that Φr grows with training and tracks reward improvement. For active inference, this raises the question of how reward-free predictive organization relates to such information-theoretic signatures. I test this within an active inference agent whose architecture separates a fast perception latent z from a slow global latent g, where g is driven by prediction error and structurally decoupled from policy gradients. In a reward-free environmental regime-switching protocol, Φr concentrates in g; its aggregate magnitude is largely architectural and decreases with training. The substantive effect of learning becomes legible only at the atom-compositional level: decoupling flips sign from negative to positive and becomes regime-invariant under environmental change, while downward causation carries the regime-dependent adjustment. These results identify g as the architectural locus of Φr-relevant temporal organization in an active inference agent, and argue against reading scalar Φr as a direct index of learned integration.