Decomposing deep neural network-brain representational similarity reveals distinct sources across the human visual hierarchy
Deep neural networks predict neural responses across the visual hierarchy, yet published alignment scores do not reveal whether this correspondence reflects learned representations, architectural inductive biases, low-level image statistics, or categorical structure. We decompose DNN-brain alignment into four sources using variance partitioning on representational similarity analysis, applied to 7T fMRI from eight participants viewing approximately 10,000 natural images. The composition of alignment shifts systematically along the cortical hierarchy: in primary visual cortex, architecture and low-level statistics account for 46% of total explained variance, whereas in high-level visual cortex model-specific learned representations dominate, reaching 68% in the parahippocampal place area.
The same framework applied to language models processing s reveals that alignment is strongly model-dependent: a sentence embedding model contributes 51% unique variance in higher visual cortex, whereas an autoregressive model contributes only 4%, with the remainder attributable to surface text statistics and category structure. Across nine models, alignment scores varied substantially in composition. Decomposing alignment into its sources is necessary before a score can be interpreted as evidence for convergent processing.
The Collection of the Natural Scenes Dataset was supported by NSF IIS-1822683 (K.K.) and NSF IIS-1822929 (T.N.). This research is also by Mid Sweden University for Open Access Publication. Department of Psychology, Lancaster University, Lancaster, UK Faculty of Education, Beijing Normal University, Beijing, China College of Intelligent Science and Engineering, Beijing University of Agriculture, Beijing, China Department of Computer and Electrical Engineering, Mid Sweden University, Sundsvall, Sweden Correspondence to Fangyao Zhang or Yuxuan Zhang.
The authors declare no competing interests. Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Below is the link to the electronic supplementary material.
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder.
To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/. Li, X., Zhang, F. & Zhang, Y. Decomposing deep neural network-brain representational similarity reveals distinct sources across the human visual hierarchy.
Extract — continue reading at the source.