Abstract

Revolutionary advances in our journey towards a better understanding of minds, brains, and the relation between them are rare. One occurred during the 1980s when concepts of parallel distributed computing within networks of neuron-like elements first became highly influential in both psychology (e.g., Rumelhart & McClelland, 1986) and statistical physics (e.g., Hertz, Krogh, & Palmer, 1991). I see Rolls’ latest book as a culmination of the application of those concepts to cerebral cortex. I see it as representative of the limitations of that revolution as well as of its strengths, however. So, after reviewing the many principles discussed by Rolls, I explain why the book shows the need to extend its conceptual foundations by putting greater emphasis on the information processing that occurs within neocortical pyramidal cells (e.g., Häusser & Mel, 2003; Häusser, Spruston, & Stuart, 2000; Jadi, Behabadi, Poleg-Polsky, Schiller, & Mel, 2014; Larkum, 2013; Major, Larkum, & Schiller, 2013; Phillips, 2017; Sjöström, Rancz, Roth, & Häusser, 2008; Williams & Stuart, 2003).
The many principles of cerebral operation explored by Rolls and colleagues
“Cerebral Cortex: Principles of Operation” is an extraordinary book, encyclopaedic, ambitious, challenging, labyrinthine, and long—near to 1,000 pages in all. Its aim is to synthesise the fruits of a lifetime spent studying the firing rates of neurons in the brains of mammals and of primates in particular. Detailed views, often supported by detailed computational modelling, are expressed on a wide array of issues concerning cerebral cortex, defined as the six-layered neocortex plus the three-layered olfactory and hippocampal cortices. The urge to be all encompassing is so great, however, that there is also a chapter on the basal ganglia and one on the cerebellum. Sub-systems and issues are viewed from multiple perspectives. For example, Chapter 12 is on “Memory and the hippocampus,” Chapter 24 is on “The hippocampus and memory,” Chapter 5 is on “The noisy cortex: stochastic dynamics, decisions, and memory,” Chapter 16 is on “Noise in the cortex, stability, psychiatric disease, and aging,” Chapter 14 is on “Invariance learning and vision,” and Chapter 25 is on “Invariant visual object recognition learning.” This all leads to much repetition, so I often had the impression of reading an anthology of papers, in which the background perspective is re-visited in each paper. Indeed, Rolls has an astonishingly large collection of published papers on which to draw, now approaching 600 papers in 50 years, thus approaching about one paper per month sustained over five decades! Even for someone with so many high-quality colleagues, this is a rare feat. The range of issues covered by the book is enormous. Although important issues and areas of relevant research are neglected, that reflects the breadth of the cognitive and neurosciences, not a narrowness of Rolls’ ambition. Thus, an adequate assessment of this book requires assessment of a conceptual framework that has been developed over many years by many researchers and applied to many issues. This cannot be done in a few easy pages, which may explain why I have been unable to find any previous review of the book.
The predominant methodologies that Rolls uses to achieve his aims are extra-cellular single-neuron electrophysiology and computational modelling. The physiology focuses on measurements of neuronal spike-rates in many different brain regions. The modelling consists predominantly of the design and testing of neural architectures that reflect those rates and analysis of their dynamics, capabilities, and limitations using mathematical techniques from statistical physics. Many other methodologies are also drawn upon, such as neuroanatomy, psychophysics, neuroimaging, and pathology.
The whole perspective is based on the rigorously defined and modelled concept of a continuous, single-compartment, point-like, integrate-and-fire neuron. There are two stages of processing within the model of an integrate-and-fire neuron. First, it models the way in which synaptic inputs produce many local excitatory or inhibitory post-synaptic potentials depending on the pre-synaptic input and the local membranes’ conductance and capacitance. Second, it then models the way in which these potentials are integrated at the cell body such that, if a threshold is reached, an all-or-nothing spike is fired and transmitted to other neurons, and the membrane potential then reset. It is assumed that “For most pyramidal cells, the dendrite diameter is sufficiently large that linear summation to produce the net current injected into the cell bodies is a reasonable approximation” (p. 389). In models using this integrate-and-fire assumption, a sigmoidal, logistic, activation function converts the integrated sum into the probability of a spike and thus into spike frequency when conditions are stable. It is made clear that this fundamental simplifying assumption implies that “the post-synaptic membrane is electrically short, and so summates its inputs irrespective of where on the dendrite the input is received” (p. 7).
Many aspects of mental function are viewed in detail from this perspective, including memory and learning, perception, attention, short-term memory, decision-making, emotion, reward, sleep, dreaming, consciousness, syntax and language, and motor control. Several other issues are also discussed at length, including localisation of function, hierarchical organisation, the unavoidability and utility of noise, principles of neural coding, synaptic adaptation and facilitation, evolutionary trends, and genetics of single-layer and multi-layer nets.
I have long been searching for a few, rigorously formulated, unifying principles underlying the distinctive capabilities and limitations of neocortex (Kay & Phillips, 2011; Phillips, 2017; Phillips & Silverstein, 2003; Phillips & Singer, 1997; Phillips, von der Malsburg, & Singer, 2010). That is not what I found in Rolls’ book. Instead, the principles proposed are many, diverse, and often informally expressed. I was unable to find any definitive list. Chapter 26 (Synthesis) does summarise many of them informally under 18 headings such as hierarchical organisation, localisation of function, the noisy cortex, synaptic modification, and memory systems, but the chapters referred to by those summaries often propose far more than one principle. As a consequence, I am still unsure of the criteria used by Rolls to consider something a “principle of operation.” Many of those proposed are simply generalisations concerning the anatomy and physiology of neocortex that are common in textbooks.
Rolls goes to great lengths to make the book reader-friendly. There are many useful figures, highlights at the end of every chapter, ubiquitous cross-referencing, four detailed appendices, and a thorough index. Many of the issues discussed are intellectually challenging, however, so the central messages might have been more effectively conveyed by pruning tangential issues, reducing repetition, and writing shorter, clearer, sentences.
Many of the central messages are clear, nevertheless. One is that cerebral cortex can be understood by assuming it to be built from the repeated use of three idealised architectures. These are (a) autoassociative attractor nets formed by recurrent excitatory collaterals, (b) competitive nets formed by lateral inhibition, and (c) pattern associators such as those underlying classical conditioning. They are often used in combination, and cortex is assumed to be composed of many thousands of such nets. At one point, it is estimated that neocortex is composed of about 27,000 local autoassociative nets, each with a capacity of at least several thousand items. A major focus of the book is on the dynamics of recurrent autoassociative nets and analysis of their dynamics using the mathematics of statistical physics.
Hebbian learning within pattern associators is used to explain how unconditioned responses can be evoked by new stimuli after appropriate conditioning, as when orbitofrontal cortex learns to associate new visual information with taste. Competitive nets are assumed to depend upon the lateral inhibition that is ubiquitous throughout cortex. Their main function is proposed to be pattern separation, which I think of as producing different neural responses to patterns of input that are similar but which need to be distinguished, for example, because they have different statistical relations to on-going activity elsewhere or because they form independent episodic memories.
Autoassociative attractor nets formed by recurrent excitatory collaterals within regions are a central theme throughout. The number of long-term memories that they can store is a paramount issue considered, but as that has not been clearly quantified at either the neural or the behavioural level, I am less sure of the importance of its exact quantification than is Rolls. Autoassociative attractor architectures are defined as those in which all pyramidal neurons in a group have modifiable excitatory connections with all, or many, of the pyramidal neurons in that group, including themselves. This appears to be the case in neocortex where pyramidal neurons within a region are connected to many of those within a vicinity of up to about 2 mm, including themselves. Rolls argues that they produce similar outputs in response to different patterns of input by rapidly “relaxing” towards a stable state in which the synaptic strengths and current activities are in maximal agreement. From a functional point of view, this can produce the “pattern completion” much emphasised by Rolls. As he notes, computation based on positive feedback of this kind is playing with fire, because it is inherently unstable and dangerous. It can lead to gross over-activity as in epilepsy, and perhaps also to hallucinations and delusions as in schizophrenia. Rolls suggests that the taming of this fire may be the key to the evolutionary success of neocortex because autoassociative nets have fast, content addressable, constraint satisfying, pattern completing dynamics. By being able to store many potential attractors in the patterns of synaptic connection strengths, and by retrieving one of them at a time, it is possible that they provide neural bases for LTM, STM, decision-making, and language. Much of the book is devoted to exploring these possibilities. It is concluded that excitatory recurrent collaterals provide the way that the cortex uses to implement short-term memory by continuing firing maintained by the positive feedback, and which provides a computational basis for planning ahead; for completion of a long-term memory from a partial retrieval cue; for decision-making in which the result of the decision can be maintained “on-line” by continuing attractor-related firing, to guide actions performed on the basis of the decision; and for associating together semantic items which provide the basis for semantic structures in language. (p. 415)
It is proposed that the three-layered regions of olfactory cortex and hippocampal CA3 also form autoassociative nets, but with one fundamental difference from the six-layered neocortex. Whereas each region of neocortex is seen as being composed of many local attractor nets, CA3 and olfactory cortex are each seen as being composed of one large attractor because the collaterals within each of those regions are widely spread across the whole region. It is hypothesised that in the case of CA3, this large attractor provides the neuronal basis of episodic memory. There is also much discussion of the relation between function and sparseness of coding within such attractor nets.
Given the centrality of recurrent excitatory connections between nearby pyramidal neurons in neocortex to the arguments Rolls presents, I would have been more persuaded had wider sources of anatomical evidence been considered (e.g., D’Souza & Burkhalter, 2017; Koestinger, Martin, Roth, & Rusch, 2017), along with evidence on the extent to which the relevant anatomy varies with cortical layer, area, age, gender, and species (e.g., DeFelipe, 2015). Koestinger et al. (2017), for example, conclude that the local anatomy is well suited to context-dependent processing, but that is not equivalent to forming an autoassociative attractor net, of which they make no mention. I would have been even more persuaded by Rolls’ arguments on this issue had he discussed the challenges and limitations of using action potentials and somatic recordings as a window on the active mechanisms of dendritic processing (e.g., Chadderton, Schaefer, Williams, & Margrie, 2014).
One particular omission from Rolls discussion of intra-regional recurrent connections, that did puzzle me throughout, concerns their relation to the feedback and long-range horizontal intra-regional collaterals thought to contribute to the surround modulation that is an example of Gestalt organisation. These connections are related to contour integration in vision (e.g., Nurminen & Angelucci, 2014) and are thought to be representative of much of neocortex, but these connections and contour integration are neglected by Rolls. This issue may have been avoided because it complicates the arguments for pattern completion and orthogonalisation that are so central to the book.
Attention is interpreted as involving both biased competition and biased activation. Biasing is conceptualised as a “gentle” interaction, implemented by connections within and between adjacent regions of the cortical hierarchy. Rolls argues that the interaction must be “gentle” because selective attention must not dominate the bottom-up input. If it did, the attractors would lose their local identity, and important information from the environment could be missed. Biased competition is interpreted as reflecting the choice between attractors within a region that are separated by no more than a few millimetres and which compete via inhibitory interneurons. Biased activation is interpreted as reflecting the selective enhancement of processing within whole brain regions. It is not necessarily competitive, and being less local is more visible to macroscopic neuroimaging. Differences between serial and parallel visual search for features or conjunctions are modelled in detail using feedback within a hierarchy of local attractor nets. Increases in reaction time with the number of distractors in search tasks is shown not to necessarily imply serial search because it can also arise in parallel architectures of attractor nets that take longer to settle into an attractor when there are more distractors.
Six principles of attention are proposed, which I summarise as follows. (1) Recurrent collateral connections implement short-term memory attractor states to hold the object of attention active and online. (2) Competitive nets select winning populations of neurons that provide a useful output to the next stage of processing. (3) Cortical resources are allocated so as to respond to current processing needs efficiently, which may, in humans, include multi-step syntactic plans. (4) Whole cortical areas and streams can be biased by top-down attention, as modelled by biased activation. (5) Cognition at the highest level can have top-down effects that bias processing in low-level areas. (6) There is a system for bottom-up attention whereby salient sensory stimuli elicit orienting.
In contrast to principle 5 in that list, however, he also says, It is a general principle of operation of the cortex that decision making areas tend to be higher in the hierarchy than cortical areas that linearly represent reward value or the stimulus . . . . . One reason may be that a nonlinear operation such as a binary choice inevitably throws away information, and it is a goal of earlier cortical stages to build useful representations using all of the information available, and to leave selection to later stages. (p. 292, italics mine)
The contrast between this and principle 5, summarised above, reflects both the crucial role of “usefulness” and Rolls’ uncertainty about the benefits of keeping all the information available.
The long chapter on invariance learning and vision discusses a model of invariant visual object recognition in the ventral visual stream. This is a four-layer unsupervised model of visual object recognition trained by a form of competitive learning that utilises a temporal trace learning rule to implement the learning of invariance using views that occur close together in time. Its success in doing this is in large part due to the fact that non-accidental distinctive features tend to change slowly with changes in viewpoint. As Rolls notes, such a strategy for acquiring visual invariance has also been proposed by several other research groups.
Rolls provides hypotheses on many other issues, including views unlikely to receive widespread assent. For example, the chapter on sleep concludes that it may have evolved to conserve energy and keep us out of harm when it is dark. Consciousness is hypothesised to be the higher order syntactic thought that enables reporting via the sequential syntactic manipulation of symbols. He also suggests that the large amount of feedback to LGN may have little or no functional utility. Although I am not convinced of the validity of all of Rolls’ hypotheses, I learnt many useful things from the book, so my attempt to read it all was well rewarded. Overall, I see it as representative of many of the strengths of the conceptual revolution for which I argued in the 1980s. In the following, I explain why I also see it as showing why we now need to advance beyond the limitations of that revolution. In putting so much emphasis upon network architecture, it underestimates the information processing capabilities of the elements from which the networks are built.
Intimations of revolutionary advances in the conceptual foundations of the cognitive and neurosciences
In the early days of the connectionist, parallel distributed processing, neural net revolution, I wrote a critical notice of the two volumes that expressed its core beliefs (Phillips, 1988). The three network architectures central to the view of cerebral cortex offered by Rolls’ book were all much emphasised in those early days, as were the content-addressing capabilities of autoassociative architectures and the use of the mathematics from statistical physics to analyse their dynamics. Thus, though Rolls distinguishes his work from connectionism, I see it as a neurobiological culmination of that revolution. As I emphasised in 1988, the search for the “microstructure” of cognition in no way implies a lack of concern for the whole organism. Any aspects of microstructure with implications for macroscopic system capabilities and limitations are surely of great importance to psychology. What I did not realise so clearly then is that this search for the psychologically relevant “microstructure” of cognition would take us deep into the subcellular levels of dendritic information processing within neocortical pyramidal cells.
Studies of dendritic computation show that much information processing occurs within neocortical pyramidal cells
A central goal of the cognitive and neurosciences is to understand how our capabilities, actions, and direct phenomenological experience depend upon the properties and dynamics of the elementary information processing components from which our brains are built. As a simplifying first-approximation, it has long been assumed that, from an information-processing point of view, those elementary components can be adequately conceived of as neurons that simply sum their excitatory and inhibitory inputs and use that sum to generate a train of spikes that is communicated to other neurons. This view is made explicit in the widespread use of network diagrams composed of unstructured point-like nodes, with all inputs to a node, or a neuron, being essentially equivalent. Psychological theories, computational models, and machine learning algorithms based on that simplifying assumption are common. Nevertheless, despite its use by Rolls and most other cognitive neuroscientists, there is now ample evidence that this simplifying assumption greatly underestimates the information processing that occurs within each single neocortical pyramidal cell. Research on dendritic computation, for which there has long been a good case, is now growing rapidly (Chadderton et al., 2014; Häusser & Mel, 2003; Häusser et al., 2000; Jadi et al., 2014; London & Häusser, 2005; Major et al., 2013; Petersen & Crochet, 2013; Polsky, Mel, & Schiller, 2004; Sjöström et al., 2008; Williams & Stuart, 2003). Furthermore, the common assumption that all pyramidal cells can be treated as essentially equivalent is clearly an over-simplification (e.g., Luebke, 2017). I frequently refer explicitly to the pyramidal cells, with which I am predominantly concerned, as “neocortical” because there are many differences between hippocampal and neocortical pyramidal cells. Although Rolls notes some of these differences in his text, all of his many diagrams show them as being equivalent.
These advances in our knowledge of dendritic computation in neocortical pyramidal cells now suggest that many psychological phenomena may be most effectively and efficiently understood as the macroscopic consequences of a rich variety of subcellular processes that transcend the summing of excitatory and inhibitory inputs at a single site of integration within the neuron. Even though we are still in the early days of this revolution, functional differences between basal or perisomatic and distal apical sites of integration can already be related to a wide range of psychological phenomena as outlined below. Furthermore, the evolution and development of cognitive capabilities may arise from changes in the information processing that occurs within pyramidal neurons as well as from changes in the macroscopic architectures built from them. In ancient three-layer cortex, the main driving input to pyramidal cells is predominantly via distal apical synapses, whereas in six-layer neocortex, it is via basal or proximal synapses, with distal apical inputs providing top-down connections (Shepherd, 2011). Communication between distal apical and somatic integration zones is influenced by voltage-gated and other non-synaptic ion channels (Biel, Wahl-Schott, Michalakis, & Zong, 2009; He, Chen, Li, & Hu, 2014), some of which have a low density at birth and a long developmental time-course (Atkinson & Williams, 2009). Thus, these and other crucial evolutionary and developmental changes will not be adequately understood if we assume that, except for being either excitatory or inhibitory, all cortical neurons are essentially equivalent, and do no more than sum their excitatory and inhibitory inputs.
Different post-synaptic regions of the dendritic tree have different functions, and distal apical synapses can have some kind of modulatory function
There is clear evidence that the effects of synaptic input on axon-potential generation depend upon synaptic location, with clear differences between basal or perisomatic and distal apical dendrites (D’Souza & Burkhalter, 2017; Major et al., 2013). There is also clear evidence that synaptic learning rules are dependent on the location of synapses on the dendritic tree, again with clear differences between apical and basal/perisomatic inputs (Sjöström et al., 2008). This dependence of both processing and learning on post-synaptic location is closely related to notions of modulation because anatomical, physiological, Event Related Potential (ERP), and much other evidence implicates the distal apical inputs in some form of modulatory function (e.g., Bachmann & Hudetz, 2014; Cauller & Connors, 1992; LaBerge, 2005; Spratling, 2002). Distal apical function may thus be implicated in a wide range of psychological phenomena, including figure-ground segregation, attention, learning, emotional arousal, slow-wave sleep, dreaming, and other variations in conscious state, such as those associated with general anaesthesia (e.g., Bachmann & Hudetz, 2014; Körding & König, 2000; Mather, Clewett, Sakaki, & Harley, 2016; Phillips, 2017; Phillips, Bachmann, & Storm, 2018; Phillips, Larkum, Harley, & Silverstein, 2016; Spratling, 2002). Note that, though the distinct functions of inputs to the distal apical dendrites are emphasised here because of their relation to some form of modulation, there is good evidence for other forms of functional specialisation that are dependent on synaptic location (e.g., Chadderton et al., 2014; Häusser & Mel, 2003; Jadi et al., 2014).
Prima facie, Rolls presents his perspective as being based upon local computing elements that simply sum their excitatory and inhibitory inputs in a way that implies a “dendritic democracy” in which all synapses have equal opportunities. Nevertheless, the necessity of “weak,” “mild,” “gentle,” “potentially supra-linear,” “non-linear,” or “modulatory” forms of interaction is frequently emphasised, and the possibility that this involves synaptic locations in the apical tuft is noted at several points. To understand the operation of cerebral cortex, we therefore need to understand these “modulatory” interactions and their relation to apical function. Referring to them as “weak,” “mild,” or “gentle” may be misleading because amplification can have large effects, and use of such terms does not sit comfortably with describing them as “potentially supralinear.” Thus, the restrictions that Rolls finds it necessary to place on “dendritic democracy” implies the need for functional conceptions of neocortical pyramidal cells that sees them as more than passive integrators of excitatory and inhibitory inputs and distinguishes distal apical from basal/perisomatic sites of integration.
The function of distal apical sites of integration has psychological significance because, as Rolls notes, there are several indications that they mediate top-down and other contextual effects. Ambiguity and its resolution using CONTEXTUAL constraints is so ubiquitous (Merker, 2012) that it usually goes unnoticed. For example, when reading the capitalised word in the previous sentence, the ambiguity of the second symbol in it is not apparent. It is ambiguous, nevertheless, as you can see by noting that when the identical symbol occurs in the context of digits, as in 2O18, for example, it is highly unlikely to be interpreted as a letter. I therefore assume some form or forms of contextual modulation to be common throughout cortex. The notions of “ambiguity” and “context” are themselves highly ambiguous, however, so I must make clear that I intend their more general interpretations. Ambiguities of presence and of relevance to current goals are intended along with ambiguities of interpretation, thus implicating apical sites in attention, which can be seen as resolving ambiguities of relevance. In the limit, the notion of “context” can be generalised to include any information that is used to resolve ambiguities of presence, interpretation, or relevance. Although this generalisation is wide, simple logical constraints can be applied to it. We can see that an O is missing from C NTEXT, which demonstrates that the context is neither necessary nor sufficient to see the symbol, and that the symbol itself is both necessary and sufficient. As much evidence suggests that context-sensitive modulation or gain-control is widely distributed throughout cerebral cortex, Salinas and Sejnowski (2001) argue that gain modulation is a central principle of brain function, and Salinas (2004) argues that it plays a central role in mapping sensory inputs to actions. Thus, although the simple demonstrations given above concern Gestalt perception, I assume that some form of selective context-dependent amplification is relevant to cerebral function in general. This would provide a more general-purpose form of context-sensitivity than the special-purpose mechanisms found throughout the early stages of all sensory systems.
The role of distal apical synapses in mediating context-dependent amplification may therefore have wide-ranging consequences for mental life in general. Direct investigation of the effects of distal apical input by several labs using multi-site patch-clamping provides evidence for a process referred to as back-propagation activated calcium-spike firing (BAC-firing) (Larkum, 2013). BAC-firing was initially interpreted as simply detecting coincidence between top-down and feedforward signals, but it can also be interpreted as using context to amplify outputs that are consistent with the context (Larkum & Phillips, 2016). This emphasises the basic asymmetry between the effects of distal apical and more proximal inputs that is required to meet the logical constraints used to define the context-dependent amplification demonstrated above. Thus, from this perspective, three new primitives can be added to the primitives of excitation, inhibition, and disinhibition that have for so long dominated attempts to interpret mental function in terms of neural activity. These three new primitives are amplification, attenuation, and disattenuation (Phillips, 2017; Phillips et al., 2016). In contrast to the long-established primitives, they imply two distinct inputs, with the effects of the modulatory inputs being conditional on the effects of the driving inputs.
Evidence for BAC-firing is strongest for Layer 5B pyramidal neurons (e.g., Major et al., 2013), but it would have psychological significance, even if limited to them alone, because they have a pivotal role in conveying neocortical signals to sub-cortical centres. The dendritic tree of Layer 5 pyramidal neurons spans most layers of the cortical column and their apical tufts receive input from diverse sources carrying brain-wide contextual information. Thus, they are well suited to the integration of information from across the cortical column in the context of current activity in other cortical and sub-cortical regions. Restriction of BAC-firing, and thus apical amplification, to pyramidal cells of Layer 5B alone is unlikely, however, because pyramidal cells in all layers except Layer 6 have tufts in Layer 1. This includes pyramidal cells in Layers 2/3, and there is direct evidence for some form of apical amplification in them also (e.g., Palmer et al., 2014). There is clear evidence that Layer 1 receives diverse contextual inputs from many sources (e.g., Lawrence, Formisano, Muckli, & de Lange, 2017; Muckli et al., 2015; Petro & Muckli, 2016; Roth et al., 2016; Rubio-Garrido, Perez-de-Manzo, Porrero, Galazo, & Clasca, 2009), so the many network diagrams in Rolls book would be greatly enhanced if they were modified to reliably distinguish between basal/perisomatic and distal apical synapses. Those diagrams do sometimes hint at such a functional distinction between distal apical sites and more proximal perisomatic sites, but not reliably, and not explicitly in relation to the amplifying effects of distal apical synapses.
The tenacious persistence of the assumption of a single site of integration within pyramidal neurons is easily understood. Much of our knowledge of cerebral cortex comes from observations of spiking output or somatic potentials. The technical difficulties that must be overcome to directly observe communication and interaction between distant parts of the same neuron are very great. Even now multi-site patch-clamping of well-separated sites on single neurons is rare, and only a few labs in the world can do it. Furthermore, most of these experiments are in cortical slices, not in whole animals, and multi-site patch-clamping of distant parts of the same neuron in awake behaving animals is even rarer. Thus, in the absence of direct data against it, the simplifying assumption of a single site of integration within the soma of the cell seems justifiable. Theorists have had about a century to formulate explanations of macroscopic phenomena based on that assumption. Such explanations are always possible because networks composed of such simple elements can in principle compute anything computable (McCulloch & Pitts, 1943). As a consequence of this long-term commitment to it, explicitly relaxing the assumption of a single site of integration within pyramidal cells will require a radical change of perspective by cognitive and systems-level neuroscientists. Scientific revolutions are, and should be, rare. Nevertheless, it seems to me that we are now in the early stages of such a revolution.
Even the simple advance of adding amplification, attenuation, and disattenuation to the long-established primitives of excitation, inhibition, and disinhibition has major implications for conceptions of cognition and its neuronal bases. The attractor dynamics of architectures composed of recurrent collaterals that amplify but do not drive will differ from the recurrent architectures emphasised by Rolls. Among many other things, explicit incorporation of recurrent amplifying connections within regions might show how to include Gestalt phenomena, such as contour integration, within the general perspective. More importantly from a clinical perspective, our understanding of many of the pharmacological and pathological phenomena discussed by Rolls may be improved if psychological phenomena are explicitly related to a more differentiated account of information processing at the subcellular level.
Short-term memory may be related to recurrent excitatory drive
When autoassociative nets with recurrent drive reach a state of activity that is an attractor, they can stay in that state until driven away from it by further input from external sources. Therefore, it has often been suggested that some form of recurrent excitation may provide a basis for short-term memory. Rolls suggests that recurrent excitatory drive within local autoassociative nets may provide a neuronal basis for short-term memory. I see several basic differences between that and the visual short-term memory (VSTM) that I explored so long ago (Phillips, 1974), however. First, as he suggests that each cortical region is composed of many such recurrent nets, their information storing capacity and domain specificity have more in common with sensory storage than with VSTM, which has very limited capacity and is highly dependent on prefrontal, domain-general, attentional resources (Morey, 2018). Second, VSTM was concerned with the temporary maintenance of novel visual descriptions, whereas attractor dynamics apply only to familiar states. Third, the VSTM that I explored was clearly restricted to high-level abstractions. It was neither tightly tied to spatial location nor automatically over-written by new retinal input. STM is likely to involve some kind of active post-stimulus persistence, but, as that might involve intrinsic persistence within individual cells, relations between STM, intrinsic persistence, recurrent excitatory drive, and apical function remains an open issue.
Neuromodulators regulate dendritic function and mental state
Neocortical pyramidal cells are far more than passive integrators of excitatory and inhibitory synaptic input. They have dendritic compartments with different functions and with communication between those compartments that is dependent upon ion flow through synaptic, extra-synaptic, and non-synaptic voltage-dependent channels. Being dependent on such active mechanisms, this communication is open to dynamic regulation by the classical neuromodulators. For example, after reviewing many physiological studies of dendritic function and synaptic plasticity, Sjöström et al. (2008) conclude that the electrical properties of different parts of the dendritic tree are not stable but are dynamically altered on a wide range of time scales by the adrenergic, cholinergic, dopaminergic, and serotonergic systems. On reviewing the functional properties of Layer 2/3 pyramidal cells, Petersen and Crochet (2013) conclude that they are under profound regulation by brain state and current behavioural circumstances. Such state-dependent variations of pyramidal cell function have major implications for perception, attention, learning, and higher conscious cognition.
Levels of adrenergic and cholinergic arousal vary on a daily time-scale from being low during slow-wave sleep, intermediate when awake, and very high during stressful emergencies (Atzori et al., 2016; Rho, Kim, & Lee, 2018). Cholinergic, but not adrenergic, levels are high during dreaming and are associated with tonically sustained levels of arousal, whereas adrenergic levels are more associated with fast phasic changes in level of arousal (Reimer et al., 2016). Intermediate adrenergic levels when awake vary phasically from moment-to-moment in a way that is closely linked to alertness and the state of attention (McGinley et al., 2015; Reimer et al., 2016). Their slow tonic and fast phasic changes can be non-invasively monitored in humans and other species through their close correlations with pupillary dilation and contraction (McGinley et al., 2015; Reimer et al., 2016), which provides a convenient way for experimental psychologists to study the dependence of a wide range of cognitive processes on adrenergic and cholinergic arousal. Distal apical tufts are clearly implicated in these basic processes because they are highly sensitive to the state of arousal (Labarrera et al., 2018) and there is a particularly high density of noradrenergic varicosities in Layers 1 and 2 of neocortex (Agster, Mejias-Aponte, Clark, & Waterhouse, 2013; Audet, Doucet, Oleskevich, & Descarries, 1988) and also of cholinergic axons and receptors.
There is already direct evidence of mechanisms by which adrenergic arousal regulates apical function and thus the context-sensitivity of learning and processing. This depends on interactions between the adrenergic system and the Ih current flow through hyperpolarization-activated cyclic nucleotide-gated (HCN) ion channels, which, though known to few psychologists, may have a crucial role in the regulation of cerebral function (Biel et al., 2009; He et al., 2014). Labarrera et al. (2018) show that Ih current flow through HCN channels tends to disconnect the distal apical compartment from the soma. They also show that the activation of adrenergic receptors tends to close those channels, thus increasing the effects of apical amplification. This makes simple intuitive sense in relation to the contrast between sleeping and waking. When we are asleep, there is little or no need to amplify currently relevant activities, so adrenergic levels are low. When adrenergic levels increase, we awake, and the particular activities that are relevant to the current moment are then selectively amplified. The findings of Labarrera et al. (2018) also indicate that a small increase in adrenergic arousal from a very low level can have a strong effect on the excitability of the tuft, and that differences in the level of arousal beyond that are related in a more gradual linear way to tuft excitability. This suggests that there is a discontinuity from no amplification to some amplification on awaking, followed by graded continuous fluctuations in the strength of amplification when awake. This suggestion is further supported by evidence that processes of selective attention have much in common with the change in adrenergic arousal that occurs from sleeping to waking (Harris & Thiel, 2011).
These close relations between state of arousal and apical function provide firm empirical grounds from which to consider the neuronal bases of variations in conscious state. Reviews of the evidence suggest that consciousness is closely related to the selective amplification of those neuronal signals that are relevant to the current circumstances as informed by inputs to the apical tufts in Layer 1 of neocortex (Bachmann & Hudetz, 2014; Phillips et al., 2016). Computational studies show that the functions to which apical amplification could contribute are indeed those that either require consciousness or raise signals into a conscious state (Phillips, 2017). To further explore this possibility, we examined research on the pathways by which general anaesthetics regulate conscious state. We found strong evidence that all or most of them do indeed, as predicted, operate via effects on the selective contribution of apical inputs to the generation of action potentials in neocortex (Phillips et al., 2018). This perspective on consciousness is clearly very different from that proposed by Rolls. Whether the two views can be reconciled, and, if not, which, if either, is valid remains to be seen.
The potential psychological significance of apical function is further increased by evidence that, although apical amplification has been emphasised above, there is also evidence that apical drive can occur. Phillips et al. (2018) hypothesise that in human adults apical drive is more likely than apical amplification when stress is high. This is in part based upon evidence that apical input alone can drive action potential output when HCN ion channels are absent or blocked (Atkinson & Williams, 2009), and that blocking is more likely when stress is high. Although we still have much to learn concerning the conditions under which apical input is amplifying, rather than driving, my reading of the currently available evidence suggests that, in addition to being dependent on the state of arousal, it may become more prominent during the later stages of development and evolution.
Information theory has advanced beyond two-way mutual information
Information theory is central to Rolls’ perspective, and he provides a long and clear introduction to it in one of the four appendices. His view of pyramidal neurons as effectively composed of a single site of integration that transforms an input into an output is complemented by an emphasis upon the mutual information between an input and an output, and Rolls often assumes that more is better.
If all inputs to a neuron are summed to form a single input variable, then a measure of mutual information between the output and that summed input is all that is needed to quantify information transmission. If there is more than one site of integration with different roles in affecting output, however, then the classical Shannon definition of two-way mutual information needs to be complemented by recent advances in multivariate mutual information decomposition.
Although it will come as a surprise to many, mutual information in classical Shannon information theory was formulated only for the case of communication through a channel between a single input vector and a single output vector. Each vector can consist of many variables, but mutual information was defined only for the input and output variables considered as a whole. Multivariate mutual information relating different components of output to different aspects of input is not well defined in Shannon’s theory. Several ways of partitioning multivariate mutual information to do this have now been proposed, however (e.g., Griffith & Koch, 2014; Williams & Beer, 2010), and a Special Issue of the journal Entropy is devoted to this advance in information theory (edited by Lizier, Bertschinger, Jost, and Wibral, 2018). In addition to quantifying the transmission and storage of information, these new techniques clarify and quantify the modification and selective amplification of information.
Consider the case of two inputs whose contributions to a single output are to be distinguished. Multivariate mutual information decomposition decomposes the information transmitted by an output about two inputs into four components: that unique to one of the inputs, that unique to the other input, that shared by the two inputs, and a synergistic component. Synergistic components transmit information that is in neither of the inputs alone, but which depends on both together. The shared component can be further decomposed into that due to correlation between the two inputs and that due to the way in which the output is computed from input. These advances have now been used to provide rigorous conceptual and analytic tools for specifying neural goal functions (Wibral, Priesemann, Kay, Lizier, & Phillips, 2017) and for formally distinguishing context-dependent amplification from other forms of interaction (Kay, Ince, Dering, & Phillips, 2017; Kay & Phillips, 2018). In contrast to the two-way mutual information between input and output emphasised by Rolls and many others, this puts more emphasis upon transmitting information about relations between distinct classes of input (Wibral, Lizier, & Priesemann, 2015) and upon selectively transmitting just that information that is relevant to the use made of it (Kay et al., 2017).
From this perspective, a goal of cerebral cortex can be seen as abstracting, transmitting, storing, and acting upon information that is relevant to its use while suppressing irrelevant information. Different streams of cerebral processing are concerned with different uses, so what is relevant in one stream may be irrelevant in another, as in the division between ventral and dorsal visual streams. In all streams, however, a goal is to abstract and transmit only that information relevant to its use, which includes the discovery of latent statistical structure, so usefulness then depends on the particular combination of inputs within which latent structure is sought. This perspective offers new formal criteria for assessing the effectiveness and efficiency of information processing, thus advancing beyond the assumption that more information is better. These new criteria include the preservation and enhancement of organised complexity (Kay & Phillips, 2011) as well as the discovery of latent structure by Bayesian inference. They also include the suppression of what one of the foremost Bayesian theorists calls “nuisance” variables (Jaynes, 2003). As “nuisance” variables are fundamentally different from the random stochastic noise whose advantages and disadvantages are discussed in depth by Rolls, this suggests that, in addition to the signal-to-noise ratio that has been so much used, a “relevance ratio” may also be useful. Although such a notion may seem novel, essentially that has already been used to study why the back-propagation algorithm of deep learning works so well (Schwartz-Ziv & Tishby, 2017).
Many details of ion-channel dynamics and so on are explicitly incorporated into the models that Rolls advocates, and limitations of the standard integrate-and-fire conception of pyramidal cells are clearly indicated by frequent acknowledgement of the need for non-linear processes. Nevertheless, he explicitly argues for the simplifying conception on the grounds that pyramidal neurons are electrically compact (p. 9). Reviews of dendritic computation such as those cited above greatly weaken that argument. He also argues for linear integration on the grounds that multiplicative synapses are implausible. Multivariate mutual information decomposition shows that linear and multiplicative are not the only options. It shows that there are fundamental contrasts between amplifying interactions and those that are simply additive, multiplicative, or divisive (Kay & Phillips, 2018), all of which are likely to have biological significance.
I am well aware that these current advances in information theory are as yet little known in neuroscience and experimental psychology, where the use of information theory is limited to the classic case of mutual information between one input and one output, as in Rolls’ book, and in Hick’s law (Proctor & Schneider, 2018). I do not expect that the advances for which I argue will be rapidly or easily incorporated into psychology and neuroscience. The mathematics of multivariate mutual information theory is challenging, and many decades of research have been devoted to developing and promoting theories based on Shannon’s bivariate information theory combined with bivariate integrate-and-fire notions of the information processing capabilities of pyramidal cells. I do expect these advances to have a major influence in the long-term, however, because there is far more to the “information processing” with which the cognitive and neural sciences must deal than can be adequately understood using only classical Shannon measures of two-way mutual information and views of neocortical pyramidal neurons that under-estimate their information processing capabilities.
Deep-learning algorithms will advance beyond artificial intelligence to real intelligence by enhancing the information processing capabilities of the local-processing elements of which they are composed
I do not expect decisive advances in the revolution that I advocate to come from arguments presented by theorists such as me. Nor do I expect them to come from empirical discoveries relating cognition and behaviour to dendritic information processing, wonderful though those discoveries have been, and will no doubt continue to be. The assumption that most of the complexities revealed by many decades of cellular neurophysiology and biophysics can be safely ignored by psychologists and cognitive neuroscientists has been too comforting to too many for too long. Instead, and in keeping with Braitenberg’s view that it is easier to synthesise complex systems than to analyse them (Phillips, 1988), I expect the decisive advances to be made within the field of machine learning, and of deep-learning in particular. The great potential of the abstract principles of aerodynamic lift was most emphatically demonstrated by the Wright brothers. My guess is that the great potential of the basic principles underlying human information processing will be most emphatically demonstrated by incorporating them into machine learning algorithms. That will reveal their potential in a concrete, undeniable, form. Behemoths such as Google have the human, technological, and financial resources needed to make this happen at an alarming rate that clearly validates Braitenberg’s view.
Rolls says that the back-propagation rule used in deep learning is biologically implausible. Hinton doubts the validity of each of the arguments usually proposed for that claim, and my view is that something functionally equivalent is not only biologically plausible, but is a major candidate for being a crucial part of the special magic of neocortex. We know that neocortex has abstraction hierarchies with up to about 10 layers. We know that early layers in different streams adapt as a function of their use higher in the hierarchy. We also know that there is massive feedback of information from higher to lower levels, and we know that this is largely to Layer 1. If neocortical pyramidal neurons have two sites of integration, with one amplifying or attenuating response to the other, then all the essentials for implementing back-propagation learning in neocortex are present, and something like this has already been shown to be feasible (Guerguiev, Lillicrap, & Richards, 2017). Information about irrelevant, or nuisance, variables is discarded in ascending a neocortical hierarchy, and different information is discarded in ascending different hierarchies. This is exactly what deep learning by the back-propagation algorithm achieves; it learns to provide compact codes for those variables that predict their deeper uses and to generalise better by learning to suppress irrelevant nuisance variables (Schwartz-Ziv & Tishby, 2017). Furthermore, I predict that, by using the contextual-integration site within the local processors to guide processing as well as to guide learning, the capabilities of deep-learning algorithms will be fundamentally enhanced. By using context to simultaneously guide both learning and processing, this will remove the highly artificial distinction between states in which knowledge is being acquired and that in which it is being used.
“Crystalized” intelligence, seen as the acquisition of a large knowledge base, has long been implemented in silicon, with capabilities that far exceed those of humans. Real intelligence (RI) requires this to be complemented by a “fluid intelligence” that can flexibly select from that knowledge base just the information that is relevant to the ever-changing needs of the current moment and use it to solve novel problems. Could the apical amplification emphasised above provide cellular-level mechanisms for the context-sensitive selectivity that such fluid intelligence requires? Prima facie, the answer may seem to be “No,” simply because all neocortical regions have pyramidal cells with apical tufts in Layer 1, whereas the fluid intelligence that is so developed in primates and particularly in mature humans seems to depend upon specific sub-regions of prefrontal cortex (Dumontheil, 2014). I think that answer is too simplistic, however. What is distinctive about those regions of prefrontal cortex is that their distance from the sensorimotor surface enables them to deal with high-order abstractions. Therefore, their dependence on internal mechanisms for selecting what to amplify is greater, which increases the need for apical amplification. If so, this would help explain why pyramidal neurons in those regions have a higher than average number of synapses (Ramnani & Owen, 2004). This is obviously speculative, but it encourages research on the possibility that there have been evolutionary changes in intracellular processing in the primate line with major consequences for cognition (e.g., Verhoog et al., 2013).
If human information processing is indeed a realisation of the potential of neural networks built from such context-sensitive local processors, and if they are used to engineer RI, then the conceptual framework within which psychologists and neuroscientists work will be transformed.
Such advances cannot resolve major ethical issues, but they may enhance our understanding of them
Unavoidable ethical issues arise, both in relation to the possible impact of further advances in machine learning and in relation to animal rights. I am heavily conflicted on both issues. I do not know whether advances in the cognitive and neurosciences, and in the machine learning algorithms inspired by them, will be to the overall benefit of humans or life on earth. Nor do I know whether they justify the invasive animal experiments upon which they draw. Nevertheless, I am confident that such advances will have huge implications for our conceptions of mental health, and for the treatments and practices designed to improve it. Most importantly, by transforming our conceptions of what it is to be alive, conscious, and human they may offer a more informed conceptual framework within which to develop a better understanding of these major ethical issues.
Footnotes
Acknowledgements
Big thanks to Philip Quinlan for encouraging me to write this, and to Jaan Aru, Talis Bachmann, Benjamin Dering, Jim Kay, Jan Kuipers, Candice Morey, Lucy Petro, Leslie Smith, Jim Stone, Johan Storm, and Michael Wibral for insightful and detailed comments on earlier drafts. Thanks also to Ed Rolls for providing the pdf of the book which was of great use in addition to the book itself.
