Abstract
People have a remarkable ability to identify the objects that they are looking at, as well as remember the images that they have seen. Researchers know that high-level visual cortex contributes in important ways to supporting both of these functions, but developing models that describe how processing in high-level visual cortex supports these behaviors has been challenging. Recent breakthroughs in this modeling effort have arrived by way of the illustration that deep artificial neural networks trained to categorize objects, developed for computer vision purposes, reflect brainlike patterns of activity. Here we summarize how deep artificial neural networks have been used to gain important insights into the contributions of high-level visual cortex to object identification, as well as one characteristic of visual memory behavior: image memorability, the systematic variation with which some images are remembered better than others.
Understanding the relationship between neural processing in high-level visual cortex and behavior requires quantifications of behavior that can be compared with patterns of neural activity. In the case of object identification, behavioral measures emphasize humans’ robust ability to accurately identify the objects present in an image across variation in the details with which those objects are presented, such as their position, size, pose, and background context (DiCarlo et al., 2012). To quantify systematic patterns of object confusions, researchers use object-identification tasks, which typically require subjects to operate in challenging conditions, for example, with limited time for viewing the images (e.g., Rajalingham et al., 2015; Fig. 1a). Patterns of object confusions are often visualized as a matrix that depicts the rate of confusion between all possible pairwise combinations of objects (Fig. 1b). Although these object-confusion patterns tend to follow reasonable expectations based on the similarity of objects’ shapes—for example, elephants and rhinos are confused more often than elephants and calculators—they provide important constraints for understanding how processing in high-level visual cortex contributes to behavior.

Quantifying object-identification and image-memorability behavior. Each trial of the object-identification task employed by Rajalingham et al. (2018) began with a period of fixation (500 ms), after which subjects viewed a gray-scale image that included an object; across trials, the objects varied in position, size, pose, and background context. This brief (100-ms) image-viewing period was followed by the presentation of two objects, and subjects indicated which of the two objects was present in the first, briefly viewed image (a). Data were pooled across many human subjects to create a matrix depicting confusion rates for all possible pairings of objects (b). The rate of object confusion for a given pair of objects was quantified by d′, computed as the difference between the proportion of trials in which one object was correctly identified (the z-scored hit rate) minus the proportion of trials in which the other object was incorrectly reported as the first object (the z-scored false alarm rate). In the matrix, warmer colors indicate higher levels of confusion; two objects that were often confused (elephant vs. rhino) and two objects that were rarely confused (elephant vs. calculator) are highlighted. The patterns in this matrix provide a quantification of object-identification behavior that can be compared with real and artificial neural responses. Panels (a) and (b) used with permission of the authors. In the image-memorability task used by Khosla et al. (2015), subjects viewed a sequence of images separated by blank interstimulus intervals (ISIs), and they pressed a button during the ISIs to indicate when an image was repeated (c). Image memorability was measured by the ability to remember images that were repeated after approximately 4 min and 100 intervening image presentations. The memorability score for a given image was computed as the rate at which that image was reported to be familiar subsequent to its first presentation (the hit rate), corrected for the rate at which novel images were reported to be familiar (the false alarm rate), averaged across approximately 80 human subjects. The resulting memorability scores were normalized to range from 0 to 1, and they can be interpreted as the fraction of subjects who remembered seeing an image. The top row in (d) shows examples of images ranging from low to high memorability. The histogram in (d) shows the distribution of image-memorability scores in the LaMem data set, which includes 60,000 images with diverse content as well as their memorability scores (Khosla et al., 2015). The arrow indicates the mean memorability score across all images in the set. Image-memorability scores provide a quantification of visual memory behavior that can be compared with real and artificial neural responses.
Tasks measuring visual recognition memory provide a complement to tasks that probe object identification. They require subjects to view the same types of images but answer a different question: whether they have seen those images before (Fig. 1c). Humans are extremely good at remembering the images that they have seen (Standing, 1973). When a visual recognition memory task is configured to be challenging by requiring subjects to remember large numbers of images across timescales longer than minutes, one characteristic behavioral pattern that emerges is variation in image memorability: Some images are systematically remembered better than others (Bainbridge et al., 2013; Goetschalckx & Wagemans, 2019; Isola et al., 2014; Khosla et al., 2015). Memorability for each image is typically quantified with a score that ranges from 0 to 1, and these scores can be interpreted as the proportion of subjects that will remember seeing an image after first seeing it minutes earlier and after seeing many other images since (Fig. 1d). Across a set of images with diverse content, image memorability typically ranges from scores near .5, indicating images that half of subjects will remember, to scores near 1, indicating images that nearly every subject will remember (Fig. 1d; Isola et al., 2014; Khosla et al., 2015). Many different factors combine to determine image memorability (Rust & Mehrpour, 2020). For example, images that contain people tend to be more memorable, whereas images of nature scenes tend to be less memorable (Isola et al., 2014). Similarly, atypical pictures of objects tend to be more memorable than typical ones (Isola et al., 2014). As a complement to object-confusion patterns (Fig. 1b), image-memorability scores (Fig. 1d) provide an important benchmark for understanding the relationship between neural responses in high-level visual cortex and behavior.
Object Identification and Image Memorability: Two Complementary Coding Schemes in High-Level Visual Cortex
Strong analogues of the behavioral patterns that have been documented for object identification and image memorability in humans have been recapitulated in one animal model to date, the rhesus macaque monkey (Jaegle et al., 2019; Rajalingham et al., 2015). As an extension to foundational work in humans using functional MRI (Bainbridge et al., 2017; Kriegeskorte et al., 2008), investigation of these behavioral patterns in monkeys allows for the determination of their neural correlates at the level of detail thought to be most relevant for neural coding: the spatial resolution of individual neurons and the temporal resolution of individual events of neuron activation. In the monkey brain, neural activity patterns that reflect object-identification behavior emerge first in inferotemporal cortex (IT; Majaj et al., 2015). The fact that these patterns are not reflected in the visual brain area that provides input to IT, V4, suggests that these patterns emerge from processing that occurs within IT itself. Variation in image-memorability behavior is also reflected in the neural activity of IT, although by a complementary coding scheme: Whereas object identity is largely reflected as population response patterns, image memorability is reflected by overall magnitude, or vigor, of the IT population response (Jaegle et al., 2019). That is, a strong correlation exists between image-memorability scores and the magnitude of the IT population response.
In the case of object identity, coding of the population response pattern is a consequence of individual IT neurons that are selectively responsive, or tuned, for the high-level image properties that define objects. In the geometric space that is typically used to conceptualize object representations (Fig. 2), this translates into different images being represented in different regions of this space (DiCarlo et al., 2012). Because IT neurons tend to maintain their rank-order object selectivity across different transformations of an object (e.g., changes in position or background), representations of different images containing the same object tend to cluster in this space. In this format, object identity can be easily decoded, for example, by determining the object cluster that a particular IT population response pattern is most similar to. In comparison, within an IT representational cluster of images that all contain the same content—for example, different pictures of the same animal or the same person—the images that evoke the most vigorous responses are the ones that are remembered best. The fact that object identity and image memorability are reflected via different coding schemes in IT means that they impose complementary constraints on models of processing up to and within high-level visual cortex.

Complementary coding of object identity and image memorability in inferotemporal cortex (IT). Image representations in high-level visual cortex are often depicted in a geometric space in which each axis reflects the response of one neuron (DiCarlo et al., 2012). Although the total dimensionality of the space is determined by the total number of neurons in a population (e.g., ~10 million in monkey IT), useful intuitions can be gained by considering a handful of dimensions (e.g., the magnitude of the responses of four neurons, r1–r4, as depicted here by black arrows). In this space, the response of the IT population can be parsed into two independent factors: the direction in which each population response vector points, which is determined by the pattern of responses across the population, and the length of each population response vector, which is determined by the overall vigor, or magnitude, of the IT population response. Shown are the hypothetical responses (white arrows) of the IT population to three images from each of two categories (giraffe, rhino); each image is labeled with its memorability score. Within IT, different objects evoke different IT population response patterns because the individual neurons tend to respond selectively to different object identities. In contrast, different images containing the same object tend to evoke similar IT population response patterns because IT neurons tend to maintain their selectivity for objects across identity-preserving transformations. This translates into population vectors for different images of the same object that tend to cluster in this geometric space. Thus, in this illustration, the vectors for population responses to the giraffe images form one cluster, and the vectors for population responses to the rhino images form another). In other words, object identity is largely reflected as coding by a population response pattern (or, equivalently, the direction in which a vector points). In contrast, image memorability is reflected by magnitude coding in IT: The images that produce the responses of the largest magnitude (reflected by the longest population response vectors) are the ones that are remembered best (as indicated by the memorability scores). This account explains how representations of different images containing the same object can cluster in the representational space, and at the same time, some images within the cluster can be more memorable than others.
Deep Neural Networks Trained to Categorize Objects Reflect Brainlike Object Representations
Pinpointing the brain area or areas in which neural activity aligns with visual behavior is an important first step toward understanding the visual system. The next step involves creating models that describe how those patterns of neural responses arise, or, equivalently, models that predict how a neuron or population of neurons will respond to any arbitrary image. This type of model is often referred to as a “descriptive model” to reflect the fact it is a mathematical description of the mapping from images as input to firing-rate responses as output. Although the components of descriptive models are loosely biophysical, the broader goal of this type of model is accurate prediction of neural activity that generalizes to predict the response to any arbitrary image, as opposed to biological realism (Rust & Movshon, 2005). In addition to playing a role in describing the neural computations performed by the brain, descriptive models are an important step toward defining the specific neural circuits and biophysical mechanisms that give rise to behavior (Carandini, 2012).
Historically, developing accurate descriptive models of neural processing up to and including high-level visual cortex has been difficult. This is because by the time neural signals reach IT, they have been processed by many different brain areas, and consequently, the mapping of images’ pixel patterns to IT neural responses is complex. Major breakthroughs in this modeling effort have arrived by way of advances in computer vision, through the development of new approaches to train deep artificial neural networks to categorize objects (Fig. 3a)—that is, to accept images as input and report the categories of the objects contained in the images as output (Yamins & DiCarlo, 2016). The crucial observation of relevance to visual neuroscience is that, once trained, the higher layers of these deep artificial neural networks reflect patterns of activity that bear some degree of similarity to patterns of neural activity in monkey IT (Khaligh-Razavi & Kriegeskorte, 2014; Yamins et al., 2014; Fig. 3b), as well as indicators of neural activity in human high-level visual cortex, measured by functional MRI (Guclu & van Gerven, 2015; Khaligh-Razavi & Kriegeskorte, 2014). In addition, the alignment of intermediate layers of these networks with intermediate stages of visual processing, including areas V1 (Cadena et al., 2019) and V4 (Yamins et al., 2014), suggests that they do not simply produce the same output as the brain, but also arrive at that output in an analogous way.

Illustration of the use of deep artificial neural networks trained to categorize objects. The diagram in (a) depicts a classic convolutional neural network (CNN) architecture, used both for AlexNet (Krizhevsky et al., 2012) and for the HybridCNN (Zhou et al., 2014). The two networks differ in the images used for their training, but both include five convolutional layers and two fully connected layers. The eighth layer has been omitted as it reflects the output of a 1,000-way object classification, which is irrelevant for the purpose of defining and comparing the population encoding of images within the network. The dimensions of each layer are depicted loosely proportional to their true dimensionality. The matrices in (b) provide a comparison of object representations in monkey inferotemporal cortex (IT) and two layer of AlexNet. Each matrix reflects the dissimilarity between the population activation patterns for all possible pairings of 92 objects, organized by class (animate vs. inanimate) and subclass. H = human; NH = nonhuman. Entries in the dissimilarity matrix for monkey IT were determined by considering the population vectors of firing-rate responses for the images in each pair, computing the Pearson correlation (r) between these population vectors as a measure of response similarity, and subtracting that value from 1 (i.e., dissimilarity = 1 – r). Values for AlexNet network layers were computed in the same manner, but applied to the activation patterns of artificial network units within the same layer. The Kendall rank correlations between the dissimilarity patterns reflected by monkey IT and the different layers of AlexNet are shown. Adapted from “Deep Supervised, but Not Unsupervised, Models May Explain IT Cortical Representation,” by S.-M. Khaligh-Razavi and N. Kriegeskorte, 2014, PLOS Computational Biology, 10(11), Article e1003915, Figs. 5 and 6 (https://doi.org/10.1371/journal.pcbi.1003915). Copyright 2014 by the authors. The original article is available under the Creative Commons CC-BY license. The graphs in (c) show image-memorability representations in monkey IT versus the HybridCNN. In the graph on the left, firing rate (number of neuron spikes per second) across the IT population is plotted as a function of image memorability; each point depicts the response to a different image. The graph on the right shows the mean correlation between image-memorability scores and the magnitude of the population response for different layers of the HybridCNN, for both a trained and a randomly connected version of the network. The shaded bands represent 95% confidence intervals. Adapted from “Population Response Magnitude Variation in Inferotemporal Cortex Predicts Image Memorability,” by A. Jaegle, V. Mehrpour, Y. Mohsenzadeh, T. Meyer, A. Oliva, and N. Rust, 2019, eLife, 8, Article e47596, Fig. 2 Supplement 3b and Fig. 3 (https://doi.org/10.7554/eLife.47596). Copyright 2019 by the authors. The original article is available under the Creative Commons CC-BY license.
Deep artificial neural networks contain many layers that each contain some of the same components incorporated into descriptive models of neurons at earlier stages of visual processing, such as V1 (Yamins & DiCarlo, 2016). “Deep” refers to the fact that these networks have many layers and the input is processed through them sequentially. Most networks include early convolutional layers, in which the same filtering operation is applied at all spatial positions in an image, as well as higher fully connected layers, in which all units are interconnected. Deep networks are typically trained in a supervised fashion in which each image used for training is associated with a categorical label (e.g., “elephant”) and errors (differences between actual labels for images and the network’s predictions) are backpropagated through the network to adjust its parameters. Different networks can differ by their architecture (i.e., the numbers of layers and types of processing within each layer) and their optimization (i.e., details about how they are trained to categorize objects, including the set of images used to train them as well as other training details). For example, popular networks that have been trained for object categorization include AlexNet (Krizhevsky et al., 2012) and the HybridCNN (Zhou et al., 2014), which share the same architecture but differ in the sets of images used to train them, as well as VGG-16 (Simonyan & Zisserman, 2015), which has a different architecture (with more and different layers) than AlexNet but was trained on the same image set, ImageNet (Deng et al., 2009).
What have researchers learned about object identification from deep artificial neural networks? First and foremost, many of these networks include simple models of neurons configured in a purely feedforward hierarchy, and the fact that this configuration generates object representations analogous to brain activity confirms long-held suspicions that stacks of simple model neurons are sufficient to recapitulate the core aspects of object-identification behavior (DiCarlo et al., 2012; Yamins & DiCarlo, 2016). Second, the brain-analogous behaviors that these models fail to recapitulate have provided insights into the contributions of other elements that are known to exist in the brain, such as recurrent connections between units and feedback connections from higher to earlier stages (Kar et al., 2019; Kietzmann et al., 2019). Finally, although these models are complex, and researchers do not yet fully understand how they work, their existence provides an opportunity to perform experiments that can lead to understanding but are extremely difficult to perform in real brains. One example involves the comparison of a trained neural network with one that has randomly connected units to distinguish the contributions of network architecture from those of the specific connectivity patterns that emerge through training.
Deep Neural Networks Trained to Categorize Objects Reflect Brainlike Image-Memorability Representations
Extending findings that IT-like object representations emerge in higher layers of deep artificial neural networks trained to categorize objects, Jaegle et al. (2019) determined that these same networks also reflect a correlate of image-memorability variation. That is, the IT-analogous layers of deep artificial neural networks trained to categorize objects respond more vigorously to some images than others, and the magnitudes of the population responses in these layers are correlated with the image-memorability scores measured in humans (Fig. 3c). Notably, in earlier V1-analogous network layers, correlations between population response magnitude and image memorability are weak, and correlations increase at higher levels of the network hierarchy (Fig. 3c).
These results were surprising, as a deep neural network trained to categorize objects is not optimized to reflect memorability in any way nor, once it is trained and its parameters are fixed, can it remember anything (i.e., it responds exactly the same way the first, second, and 1000th time it sees an image). The precise reasons that deep artificial neural networks trained for object categorization respond more vigorously to some images than others are not yet well understood. However, the fact that this correlate of image-memorability variation emerges from these networks provides important insights into the origins of image memorability, insofar as it suggests that at least some component of this variation can be attributed to a system optimized to see and categorize objects (as opposed to remember the images that it has seen).
Summary
The systematic behavioral patterns that characterize by object identification and image memorability (Fig. 1) impose complementary constraints on descriptions of how processing in high-level visual cortex supports behavior. Any accurate description of processing must be able to account for these two behavioral signatures independently. The recapitulation of humanlike behavioral patterns in rhesus monkeys has allowed for the determination of the neural correlates of these behaviors in high-level visual brain area IT. There, these two behaviors are reflected via complementary coding schemes: Whereas object identity is largely reflected as population response patterns, image memorability is reflected by population response magnitude (Fig. 2). Although it remains unclear why the system operates in this way, one consequence is that the two behaviors place complementary constraints on models intended to describe how behavioral patterns emerge from the many stages of processing up to and including IT.
Historically, developing models that can account for the conversion of images into IT population responses has been difficult. Recent breakthroughs in this modeling effort have emerged from demonstrations that deep artificial neural networks trained to categorize objects bear some degree of similarity to IT neural activity signatures, both for object identification (Fig. 3b) and for image memorability (Fig. 3c). The existence of these models has led to a number of important insights, including the confirmation that stacks of simple model neurons can recapitulate the core aspects of object-identification behavior, and the revelation that at least some component of image-memorability variation emerges from a system optimized for object categorization. Going forward, increasingly refined models will be evaluated by their ability to account for both perceptual and mnemonic behavioral signatures and their neural correlates, thereby providing a natural way for understanding the degree to which the computations that contribute to perception and memory are shared versus distinct. Likewise, the development of these types of models holds tremendous promise for understanding the specific processing that happens in each visual brain area as well as how that processing contributes to visual behavior.
Recommended Reading
Kriegeskorte, N., & Golan, T. (2019). Neural network models and deep learning. Current Biology, 29(7), R231–R236. https://doi.org/10.1016/j.cub.2019.02.034. A primer on deep learning, intended for biologists, that covers feedforward as well as recurrent neural networks.
Rust, N. C., & Mehrpour, V. (2020). (See References). An extensive review of image memorability that synthesizes what is known about image-memorability behavior, its reflection in the brain, and its emergence in deep artificial neural networks trained to categorize objects.
Serre, T. (2019). Deep learning: The good, the bad, and the ugly. Annual Review of Vision Science, 5, 399–426. https://doi.org/10.1146/annurev-vision-091718-014951. A comprehensive review of recent advances in deep learning applied to human visual intelligence.
Yamins, D. L., & DiCarlo, J. J. (2016). (See References). An extensive review of deep artificial neural networks as models for different types of sensory processing, with an emphasis on visual object identification.
