Abstract
A compelling demonstration of implicit learning is the human ability to unconsciously detect and internalize statistical patterns of complex environmental input. This ability, called statistical learning, has been investigated in people with dyslexia using various tasks in different orthographies. However, conclusions regarding impaired or intact statistical learning in dyslexia remain mixed. This study conducted a systematic literature search of published and unpublished studies that compared statistical learning between people with and without dyslexia using different learning paradigms in different orthographies. We identified 49 papers consisting of 59 empirical studies, representing the data from 1,259 participants with dyslexia and 1,459 typically developing controls. The results showed that, on average, individuals with dyslexia performed worse in statistical learning than age-matched controls, regardless of the learning paradigm or orthography (average weighted effect size d = 0.47, 95% confidence interval [0.36, 0.59], p < .001). Meta-regression analyses further revealed that the heterogeneity of effect sizes between studies was significantly explained by one reader characteristic (i.e., verbal IQ) but no task characteristics (i.e., task paradigm, task modality, and stimulus type). These findings suggest domain-general statistical learning weakness in dyslexia across languages, and support the need for a new theoretical model of statistical learning and reading, that is, the SLR model, which elucidates how reader and task characteristics are regulated by a multicomponent memory system when establishing statistically optimal representations for deep learning and reading.
Keywords
Humans have a remarkable ability to unconsciously detect and internalize statistical regularities of sensory input from the environment, including frequency, variability, and co-occurrence probability. This ability, called statistical learning (Saffran et al., 1996), has been investigated in people with developmental dyslexia (DD), a persistent difficulty in learning to read despite normal intelligence and schooling (for reviews, see Lum et al., 2013; Schmalz et al., 2017; van Witteloostuijn et al., 2017). However, because of the complexity and difficulty of evaluating statistical learning, a variety of tasks have been used to assess people of different ages with DD in different orthographies. Consequently, mixed and inconclusive findings were reported, with some studies showing impaired statistical learning (e.g., Lukasova et al., 2016; Przekoracka-Krawczyk et al., 2017), while others indicated intact statistical learning (e.g., Inácio et al., 2018; Staels & Van den Broeck, 2017). These mixed results necessitated the present meta-analysis, which examined both the extent to which people with DD differed from their age-matched typically developing (TD) controls in statistical learning and the key participant and task factors responsible for this difference.
Previous Meta-Analytical Reviews
To date, three meta-analytical analyses have compared the differences in statistical learning between people with DD and their TD controls (Lum et al., 2013; Schmalz et al., 2017; van Witteloostuijn et al., 2017). These reviews indicated that individuals with DD had poorer statistical learning than their TD controls, with average effect sizes of d = 0.45 in Lum et al.’s (2013) review, d = 0.47 in Schmalz et al.’s (2017) review, and g = 0.46 in van Witteloostuijn et al.’s (2017) review. However, all of these reviews reveal methodological shortcomings, and none of them have addressed theoretical concerns about the mechanisms underlying the link between statistical learning and reading acquisition.
Specifically, Lum et al. (2013) synthesized the data of 14 published studies comparing individuals with DD and their age-matched TD controls in serial reaction time (SRT) tasks, which assessed the participants’ reaction time difference to randomly ordered and sequentially ordered stimuli (Nissen & Bullemer, 1987). By using a random effects model to compute and test the average weighted effect size of the 14 selected studies, Lum et al. (2013) showed that individuals with DD performed significantly worse than TD controls on sequence learning in an SRT task (d = 0.45, 95% confidence interval [CI; 0.20, 0.69], p < .001). They also found substantial differences in effect sizes of individual studies, some of which could be explained by the interactive influence of the participants’ age with both the type and number of sequence exposures in the SRT task. In particular, their meta-regression analysis revealed a smaller effect size when older participants were tested with a higher order or larger number of sequences. However, their interaction effects should be interpreted with caution since the main effect of participants’ age and each task characteristic was not simultaneously tested. The validity of their results is further questioned by the exclusion of unpublished studies.
Subsequently, Schmalz et al. (2017) evaluated nine published studies investigating statistical learning in individuals with DD using visual artificial grammar learning (AGL), that is, learning the hidden rules of sequences of visual stimuli through exposure. They found that individuals with DD were weaker than their TD peers on the AGL task (d = 0.47, 95% CI [0.04, 0.90]). As in Lum et al.’s (2013) review, only published studies were included. However, in contrast to Lum et al. (2013), publication bias was detected. Thus, the limited number of studies combined with the presence of publication bias indicated insufficient evidence for impaired visual AGL in individuals with DD.
By extending Schmalz et al.’s (2017) study to include unpublished studies, van Witteloostuijn et al. (2017) synthesized the data of 13 studies (11 published and two unpublished) to investigate visual AGL in individuals with DD and TD controls. Using the random effects model, van Witteloostuijn et al. (2017) showed that individuals with DD performed significantly worse than TD controls in visual AGL (g = 0.46, 95% CI [0.14, 0.77], p < .01). Furthermore, their meta-regression analysis indicated that the study-level effect size difference was moderated by participants’ age, with a smaller effect size for adults than for children. However, van Witteloostuijn et al. (2017) claimed that the age effect could be due to a single study (Pothos & Kirk, 2004) reporting better AGL in adults with DD than their TD controls. Despite its inclusion of unpublished studies, van Witteloostuijn et al.’s (2017) review still suffered from methodological limitations. For example, by neglecting the possible correlations of multiple effect sizes from a single study, their use of the random effects model potentially resulted in an inaccurate estimation of the overall effect size (Borenstein et al., 2021).
Together, these previous meta-analytical reviews were each confined to one statistical learning paradigm, with Lum et al.’s (2013) focusing on the SRT paradigm, and Schmalz et al.’s (2017) and van Witteloostuijn et al.’s (2017) focusing on the AGL paradigm. This not only prevented the possibility of evaluating the effect of task paradigm on study-level effect size difference but also partly led to the inclusion of a limited subset of articles on statistical learning in DD, which further contributed to an insufficient power for detecting other important moderators of effect sizes. Only a small number of moderators, especially participant characteristics, were investigated to account for the study-level effect size difference in the previous meta-analyses. However, other participant (e.g., verbal and nonverbal intelligence, native orthography, and reading skill level) and task (e.g., learning paradigm and stimulus type) characteristics crucial to understanding the relations between statistical learning and reading acquisition were excluded.
Given the limited focus and the small subsets of studies included in the previous meta-analyses, we currently risk an incomplete understanding of statistical learning abilities in individuals with DD. Thus, it is necessary to use various paradigms to systematically synthesize data from all studies of statistical learning in DD. To this end, the present study extends previous meta-analyses by more comprehensively evaluating whether people with DD exhibit any statistical learning weaknesses compared to age-matched TD controls and by assessing possible factors responsible for inconsistent findings reported in the literature. Table 1 presents a summary of advancement of our current meta-analytical review in relation to the three mentioned above.
A comparison between the previous and current meta-analyses
Note. Relevant main effects of the interactions were not tested in Lum et al. (2013). SRT = serial reaction time; AGL = artificial grammar learning; RVE = robust variance estimation; WR = word reading; NWR = nonword reading; OD = orthographic depth.
The Present Review
The goal of this study was twofold: First, by incorporating all statistical learning paradigms into the analysis, and thus extending previous meta-analytical reviews, we evaluated whether people with DD across languages exhibit any statistical learning weaknesses compared with TD controls, while controlling for the potential correlations of multiple effect sizes obtained from a single study. Second, to adequately address whether some methodological factors influenced the differences in statistical learning between people with DD and TD controls, we determined which participant/reader characteristics and task characteristics accounted for discrepancies between studies. These participant/reader and task characteristics are described below.
Participant/Reader Characteristics
In addition to age group, we examined cognitive factors (i.e., verbal and nonverbal intelligence, working memory, rapid automatized naming, and phonological processing abilities), which were neglected in the previous meta-analyses. Lum et al. (2013) and van Witteloostuijn et al. (2017) reported the moderating effect of age, which was interpreted as an indirect influence of the involvement of declarative memory. However, given the possible dynamic of competing and complementary roles of implicit procedural memory and explicit declarative memory during the course of statistical learning (Sawi & Rueckl, 2019), the influence of declarative memory on statistical learning cannot be represented by age alone. Thus, in the present study, we proposed that verbal IQ (intelligence quotient) functions as a primary index of the involvement of explicit declarative memory since it places high demands on an individual’s semantic and episodic memory.
Furthermore, our assumption that working memory, rapid automatized naming, and phonological skills might contribute to statistical learning variability in individuals with DD was derived from evidence showing that individuals with DD had deficits in these cognitive skills (e.g., Cancer & Antonietti, 2018). Also, working memory could be a moderator as it may be necessary for the processes of stimulus encoding and retention (Arciuli, 2017) and extraction of statistical regularities during the course of statistical learning (Erickson & Thiessen, 2015).
Another set of reader characteristics was orthographic depth and reading abilities (i.e., word and nonword reading fluency and accuracy). Orthographic depth refers to the reliability of print-to-speech correspondence, and it can be quantified in terms of complexity and unpredictability (Schmalz et al., 2015). For instance, English is regarded as a deep orthography because of its large number of irregular words and position-specific sublexical rules. In contrast, Italian is considered a shallow orthography because of its consistent symbol-to-sound mapping. Although orthographic depth was originally proposed on the basis of print-sound mapping in alphabetic orthographies, an increasing number of studies have extended this concept to nonalphabetic orthographies, such as Chinese (M. Wang & Geva, 2003). In fact, the reliability of print-to-speech correspondence can be applied to Chinese. Given that the majority of Chinese characters (over 80%; Shu et al., 2003) are semantic-phonetical compound characters comprising a phonetic radical (sound clue) and a semantic radical (meaning clue), the phonetic radical-sound mapping is prevalent in the Chinese writing system, and the consistency and reliability of radical-syllable mappings ranged from low to high along a continuum (X. Tong & McBride, 2018). Likewise, a number of studies on Hebrew reading acquisition have clearly indicated the existence of print-to-speech correspondence in the standard, unvowelized orthographic version of Hebrew, where consonants are fully represented and vowels are partially represented by four letters (Schiff, 2012). On the basis of these previous studies, Chinese and Hebrew can be classified as deep orthography when coding the orthographic depth for the included studies.
Among alphabetic orthographies, learning to read a deep orthography (e.g., English) is slower and more difficult than learning to read a shallow orthography (e.g., Italian; Verhoeven & Perfetti, 2021). Higher capacity in statistical learning might be required to learn the symbol-to-sound mapping in a deep orthography than in a shallow orthography (Nigro et al., 2015). Moreover, reading impairments in people with DD, especially reading accuracy, were moderated by orthographic depth (for a review, see Carioti et al., 2021). As orthographic depth may influence the importance of statistical learning to reading acquisition and moderate the key factors differentiating individuals with and without DD, we examined the possible interaction effects of orthographic depth and word reading abilities to account for the variability in statistical learning in people with DD. Specifically, we hypothesized that reading skills for readers of deep orthographies, but not shallow orthographies, might account for effect size differences between the studies since more complex multilevel statistical regularities are embedded in a deep orthography than in a shallow orthography.
Task Characteristics
To extend the preliminary findings of moderating effects of task characteristics in Lum et al.’s (2013) and van Witteloostuijn et al.’s (2017) meta-analytical reviews, we examined whether task paradigm, task modality, and type of stimuli influenced the effect size differences between studies.
The most commonly used statistical learning paradigms for investigating statistical learning in people with DD are SRT (Nissen & Bullemer, 1987) and AGL (Reber, 1967). As described above, the goal of the SRT paradigm is for the participant to become implicitly familiar with a sequence embedded in a continuous stream of stimuli through mere exposure. Although not informed of the sequence, participants respond faster to sequenced trials than to random trials. Such difference in reaction time reflects statistical learning. In contrast, the AGL tasks assess whether participants are able to implicitly learn the rules of the sequence of stimuli. In the AGL tasks, participants are presented with strings of letters, numbers, or nonlinguistic symbols arranged in sequences created by a complex set of rules (also called grammar). During the stimuli presentation, participants may be asked to memorize strings of stimuli (i.e., active exposure) but are not informed of the underlying grammar. In the test phase, participants are asked to distinguish between grammatical and ungrammatical strings. Statistical learning is indexed by whether the participants’ grammar judgement is significantly above chance level. Therefore, the SRT and AGL paradigms differ in both the rule-based structure that governs the statistical predictability in stimuli and the implicitness of instructions.
In addition to SRT and AGL, several other statistical learning paradigms are employed to study statistical learning in people with DD. One such paradigm is the segmentation paradigm (Saffran et al., 1996), which assesses whether participants are able to extract and internalize small groups of stimuli repetitively presented to them in a continuous stimulus stream. Unlike the SRT paradigm, which uses a long sequence of stimuli, the segmentation paradigm usually consists of multiple groups of stimuli, with two or three stimuli in one group (i.e., pairs or triplets). For example, in a visual triplet segmentation task, participants were exposed to a continuous visual string that contained repeatedly occurring triplets composed of nonsense shapes (X. Tong et al., 2019). As in the AGL paradigm, participants are not informed of the input structure, but their knowledge of the input structure is subsequently tested after the exposure phase. Another paradigm is the Hebb learning paradigm (Hebb, 1961), which measures implicit learning of serial order information by contrasting recall for repeating and nonrepeating sequences. Other statistical learning tasks may require participants to associate targets based on an implicit mapping structure (Gabay & Holt, 2015), or to identify a visual target based on implicit contextual cues (Chun & Jiang, 1998), positional cues (Nigro et al., 2016), or predictor-target contingencies (Singh et al., 2018). All these statistical learning paradigms vary in multiple dimensions (i.e., the amount of input structure, input exposure, and explicit instruction) along a continuum (Conway, 2020).
Given that different task paradigms test different types of statistical structures, each task places different demands on cognitive faculties and neuro-circuits (Bogaerts et al., 2021), possibly leading to a domain-specific effect on statistical learning. For example, SRT tasks usually involve a repetitive motor response to a single long sequence, thus activating the cerebellum and basal ganglia responsible for motor and procedural learning more than other statistical learning paradigms (Bogaerts et al., 2021). Moreover, these paradigms differ in response elicitation, which may also place different demands on explicit declarative memory. Specifically, some statistical learning tasks, such as AGL, explicitly instruct participants to select the more probable answers in a test phase following an implicit learning phase. Such explicit, attention-demanding processing is not necessary for SRT tasks though explicit processing can be involved, especially when participants are exposed to a larger number of sequences (Lum et al., 2013).
The difference in task modalities involved in these task paradigms may also contribute to the heterogeneity in effect sizes between studies. For instance, most AGL and SRT tasks are conducted via visual modalities, whereas other task paradigms involve mixed modalities (e.g., auditory, visual, or audiovisual). According to Frost et al.’s (2015) componential model, statistical learning in visual, auditory, and tactile modalities is achieved by a set of domain-general, yet modality-constrained computation principles. Such a modality-constrained account is consistent with Conway’s (2020) principles, which suggest the modality-specific effect of statistical learning results from the neuroplasticity in different cortical regions that are responsible for auditory, visual, or tactile input. Furthermore, the modality-constrained mechanism might result in unequal statistical learning performance in different perceptual modalities. For example, Kemeny and Lukacs (2019) revealed that adults were better at learning the associations auditorily than visually. More important, since individuals with DD could have auditory and/or visual processing impairments (L. C. Wang et al., 2018), as well as weak phonological and/or orthographic skills (Reis et al., 2020), they might perform differently on statistical learning tasks that are conducted via different modalities.
Additionally, we tested stimulus-specificity in terms of linguistic knowledge involvement. According to the linguistic entrenchment hypothesis (Siegelman et al., 2018), statistical learning is influenced by the degree of prior linguistic knowledge and exposure. Given that individuals with DD are disadvantaged in language and literacy abilities, the abstraction of a statistical pattern underlying linguistic stimuli might place extra burden on people with DD. Therefore, we hypothesized that statistical learning with linguistic stimuli would lead to a larger effect size of statistical learning weakness in people with DD.
Finally, the incorporation of a wide range of participant/reader and task characteristics in our study enabled us to not only determine which factors actually influence statistical learning in DD but also to clarify possible cognitive and environmental constraints, thus providing new insights regarding a theoretical model of statistical learning and reading acquisition. The following research questions were addressed:
Method
Literature Search
Search Strategy
We identified studies relevant to our meta-analysis through bibliographic databases, reference lists, and researchers focusing on statistical learning. First, we identified the target articles by systematically searching titles, abstracts, and/or keywords using a broad range of search terms that enabled us to localize all possible statistical learning articles for this current meta-analysis (for a full list of search terms, see the appendix). We retrieved 442 published papers from five bibliographic databases (i.e., MEDLINE, ERIC, PsycINFO, CINAHL, and LLBA) and 39 unpublished studies from Open Access Theses and Dissertations (OATD). An additional nine published articles resulted from a manual search of the reference lists of included papers (e.g., Szmalec et al., 2011) and meta-analyses (Lum et al., 2013; Schmalz et al., 2017; van Witteloostuijn et al., 2017). In total, we identified 490 articles for our meta-analysis.
Study Inclusion Criteria and Selection
In accordance with the protocols used by Lum et al. (2013) and van Witteloostuijn et al. (2017), we included articles that met the following criteria: (1) studies administered at least one valid and clearly described statistical learning task that measured the degree of learning in response to repeated exposure to a statistical probability embedded in the stimuli regardless of domains and modalities; (2) studies compared the statistical learning performance of DD and age-matched TD peers, with the DD group comprising individuals (a) formally diagnosed as DD according to standardized criteria by a clinical psychologist or an educational psychologist (e.g., a reading level of at least one standard deviation below average), (b) having a certificate of DD through an official institution (e.g., a government-approved diagnostic center), or (c) having records documenting reading problems; (3) the studies were written in English; and (4) the studies were accessible before May 2021.
Figure 1 represents the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) flow diagram of this meta-analysis. Of the initial 490 papers, 213 duplicates were removed. Among the nonduplicated papers, 48 met the inclusion criteria and were featured in our meta-analysis, while 229 were excluded in the abstract and full-text screening phases for one or more of the following reasons: (1) not an original paper; (2) no statistical learning tasks; (3) no DD group; (4) no TD group; (5) no data for effect size calculation; (6) not a full-text; or (7) contained the same participants as another study. Finally, one unpublished study was retrieved after contacting the corresponding authors of the studies included herein, resulting in 49 papers for this current meta-analysis.

PRISMA flow diagram showing the study selection processes (N = number of papers).
Study Description
The data structure of this meta-analysis is complex as some studies involved DD and TD participants performing the same statistical learning task, which generated an individual effect size (e.g., Bogaerts et al., 2015; Katan et al., 2017), while other studies involved DD and TD subjects performing multiple statistical learning tasks, sometimes under different conditions, which yielded multiple correlated effect sizes (e.g., Henderson & Warmington, 2017; Pavlidou & Williams, 2014).
Despite variations in characteristics, the statistical learning tasks used in each study can be generally classified under three paradigms: SRT, AGL, and Other (i.e., Hebb learning, segmentation, association, contextual cueing, and target prediction).
SRT tasks measure the motor response to sequences in visual or auditory domains. In our meta-analysis, studies utilized an SRT paradigm to examine (1) whether implicit learning was impaired in adults or children with DD (e.g., Rüsseler et al., 2006; Vicari et al., 2005), (2) whether the type or complexity of a repeated sequence influenced the statistical learning of the study groups (e.g., Deroost et al., 2010; He & Tong, 2017; Howard et al., 2006; Jiménez-Fernández et al., 2011), and (3) the extent that sequential knowledge acquired in the learning and consolidation phases differed between the DD and TD groups (e.g., Hedenius et al., 2013). Additionally, one study implemented the SRT paradigm to investigate brain activation patterns associated with implicit learning deficits (Menghini et al., 2006).
AGL tasks investigated (1) the weakness in statistical learning processes among children or adults with DD (e.g., Kahta & Schiff, 2016; Samara & Caravolas, 2017) and (2) the effects of grammar complexity, chunk strength, and feedback on learning performance in individuals with and without DD (e.g., Katan et al., 2017; Schiff, Katan, et al., 2017; Schiff, Sasson, et al., 2017). Other statistical learning tasks used were Hebb learning (Bogaerts et al., 2015; Henderson & Warmington, 2017; Staels & Van den Broeck, 2015), association (Aravena et al., 2013; Gabay & Holt, 2015), segmentation (e.g., X. Tong et al., 2019), contextual cueing (e.g., Howard et al., 2006; Jiménez-Fernández et al., 2011), and target prediction (e.g., Roodenrys & Dunn, 2008; Singh et al., 2018).
In the included studies, the participants’ native languages categorized as having deep orthographies were Chinese (Cantonese and Mandarin), English, and Hebrew, while those having shallow orthographies were Dutch, Finnish, German, Icelandic, Italian, Polish, Portuguese, Spanish, and Swedish (e.g., Borgwaldt et al., 2005; Borleffs et al., 2017).
Word reading and nonword reading were assessed either by standardized measures, such as TOWRE (i.e., Test of Word Reading Efficiency; Torgesen et al., 1999) in English, or nonstandardized tasks developed by researchers. Verbal and nonverbal IQs were measured using standardized tests such as the Wechsler Intelligence Scale for Children–Third Edition (WISC-III; Kort et al., 2005) and Raven’s Progressive Matrices (Raven et al., 1992). Specifically, verbal IQ measures comprised vocabulary, verbal reasoning, and verbal comprehension. Reliable effect sizes for visual working memory could not be retrieved since most studies used verbal measures only, especially digit span. Rapid automatized naming tests consisted of naming objects, colors, numbers, and letters, while phonological skill measures included phonemic awareness, phonological awareness, and phonological processing tasks.
Data Coding and Reliability
Coding System
Our coding system consisted of 20 variables, both categorical and continuous. These variables included the following reader characteristics: (1) sample size for each group; (2) chronological age of the participants; (3) age group, dummy coded as 0 for adults (18 or above) and 1 for children (below 18 years old); (4 to 13) effect sizes of statistical learning, word reading fluency, word reading accuracy, nonword reading fluency, nonword reading accuracy, verbal intelligence, nonverbal intelligence, working memory, rapid automatized naming, and phonological skills; (14) native language of participants; and (15) orthographic depth, dummy coded as 0 for “deep” orthography and 1 for “shallow” orthography. The statistical learning task characteristics included (16) paradigm of the statistical learning tasks (i.e., SRT, AGL, or Other); (17) task modality (i.e., auditory, visual, or audiovisual); (18) type of stimulus, with nonlinguistic and linguistic coded as 0 and 1, respectively; (19) study quality scores in the adapted Newcastle Ottawa Scale; and (20) source of effect size extraction (i.e., means with SDs, means with pooled SDs, F-value, or t-value; see the section on data extraction).
Reliability and Study Quality
Although the initial 490 papers were identified and the duplicated papers were removed by the first author, both the first and second authors were involved in the abstract and full-text screening of the remaining 277. Fifteen percent of the work performed by the first author was randomly selected and independently assessed by the second author, and vice versa. The interrater reliability was strong (Cohen’s kappa = .82; McHugh, 2012), and a disagreement regarding the eligibility of one article was resolved by discussion.
Furthermore, we assessed the quality of nonrandomized studies in our meta-analysis using the Newcastle Ottawa Scale (Wells et al., n.d.), which evaluated three quality parameters, namely, selection, comparability, and exposure. In our modified Newcastle Ottawa Scale, the maximum score for the parameter of selection was four, with one point given to each of the following items: (1) DD was clearly defined and documented with ADHD ruled out; (2) the sample size for each group was larger than 15; (3) the participants were recruited from more than one school or by way of public announcement; and (4) the characteristics of the control group were explicitly described (e.g., no history of learning difficulties). Regarding the comparability parameter, the maximum score was two, with one point given if both groups were matched on age and another point if they were matched on intelligence. For the exposure parameter, the maximum score was three, with one point given for each of the following items: (1) the experiment contained a secure record of statistical learning performance, (2) the same method of ascertainment was used for both groups, and (3) the nonresponse rate was the same for both groups.
Despite the challenges in using, as well as issues with the reliability of, the Newcastle Ottawa Scale (Hartling et al., 2013; Stang, 2010), we obtained a strong interrater reliability score (Cohen’s kappa = .81; McHugh, 2012) for independent agreement between the first and second coders, which indicates consistency in the coding of the qualities of our selected studies. The mean score of the study quality on the 9-point scale was 7.26, indicating that the overall study quality was satisfactory. Study quality did not significantly influence the estimation of overall statistical learning effect size (k = 58, n = 91, β = −0.05, 95% CI [−0.24, 0.14], p = .556).
Meta-Analytical Approach
Data Extraction
Cohen’s standardized d was adopted as our effect size index, which represents the mean difference between the DD and TD groups divided by their pooled standard deviation (Cohen, 1988). After extracting the means and standard deviations (SDs) of the statistical learning performance, indicated by accuracy or RT, for both groups in each study, we calculated the effect size by applying the R package esc (Lüdecke, 2017). In cases where statistical learning was indicated by the difference between two conditions (e.g., sequenced and random conditions for SRT tasks), we pooled the SDs of both conditions for each group if the SD of the between-condition difference was not reported (Borenstein et al., 2021). If an F-value or a t-value was reported instead of the means and SDs, we used the same R package to transform the F- or t-value into d. For some studies (e.g., Hedenius et al., 2013), we extracted the means and 95% confidence intervals or standard errors from figures using a digitizer software (Digitizelt; e.g., van Witteloostuijn et al., 2017). If the relevant data were not available in the papers themselves, we contacted the corresponding authors to acquire them; however, we received only five out of 20 replies by the end of May 2021. The data from the authors were integrated into our meta-analysis. Effect sizes were extracted and calculated by the first author and assessed by the second author, and vice versa. Interrater reliability for the data extraction was high (Cohen’s kappa = .83). It should be noted that the calculation using means with pooled SDs as the data source may underestimate the true effect size. The source of effect size extraction (i.e., means with SDs, means with pooled SDs, F-value, and t-value) significantly influenced the estimation of overall statistical learning effect size, F(3, 6.78) = 8.87, p = .011, with effect sizes extracted from means with pooled SDs smaller than those extracted from the other three sources. Therefore, in the Results section, we also report the overall effect size estimation without the effect sizes extracted from means with pooled SDs. Importantly, the inclusion or exclusion of effect sizes extracted from means with pooled SDs did not change the conclusions of the moderator analysis.
For SRT tasks, group difference in statistical learning was based on the extent of changes in RTs from the random to sequenced blocks. To calculate a single effect size for each study, the mean difference in RTs between random and sequenced blocks, the corresponding SD, and the sample size were extracted. To estimate the effect size for studies that did not report these data, the means and SDs in random and sequenced blocks for each group were extracted to generate the within-group mean differences and the pooled SDs. If the means and SDs in RTs were not reported, we extracted the F- or t-values, including those testing the interaction between groups (DD vs. TD) and blocks (random vs. sequenced), as well as the main effect of group difference in RTs between random and sequenced blocks. It was noted that some studies compared the RT difference between the last random block and the preceding sequenced block (e.g., He & Tong, 2017), whereas others measured the group difference in RTs between the last random block and the mean median RTs from the final two sequenced blocks (e.g., Jiménez-Fernández et al., 2011).
In AGL tasks, the group difference in statistical learning was measured by the disparity of overall accuracy in grammatical judgement between the DD and TD groups during the test phase. To calculate a single effect size, the mean, SD in accuracy, and sample size were extracted (e.g., Kahta & Schiff, 2016; Schiff, Katan, et al., 2017). However, in the study by Pothos and Kirk (2004), the F-value revealing group effect on statistical learning was extracted because the relevant mean and SD were not accessible. Similarly, in the study by Samara and Caravolas (2017), t-values representing the difference in accuracy between the DD and TD groups during the test phase were extracted.
For the other statistical learning tasks, the measure of statistical learning included the overall accuracy, RT, and slope of the identification test. The means and SDs in accuracy and RTs, and the sample size were extracted to calculate effect sizes (e.g., Aravena et al., 2013). Furthermore, we extracted the F-values that tested the interaction between group and condition, which reflected the different amounts of regularities implicitly acquired by individuals with and without DD (e.g., Howard et al., 2006). The F-value representing the main effect of implicit learning was extracted to generate an effect size (e.g., Bogaerts et al., 2015; Garcia et al., 2019).
Robust Variation Estimation
In the present meta-analysis, 30% of the studies reported multiple results. For instance, in the study by He and Tong (2017), the same group of participants completed the SRT task under two conditions (i.e., large and small numbers of exposures to the implicit sequential regularity) and yielded two distinct but related effect sizes. In other studies, the same TD and DD groups participated on more than one statistical learning task (e.g., Howard et al., 2006; Laasonen et al., 2014; Rüsseler et al., 2006; Staels & Van den Broeck, 2017). Therefore, more than one effect size was generated. When a study generates multiple effect sizes by using the data from the same participants, it is recommended that all of the effect sizes be included in the analysis. Concurrently, the dependent structure within a study ought to be considered as well (Scammacca et al., 2014). Although a multivariate meta-analysis could be an ideal method to synthesize effect sizes that are statistically dependent, it requires the information of the underlying covariance among effect sizes within one study. Yet this information was often neglected in many of the primary studies used in the present meta-analysis. Thus, we applied robust variance estimation (RVE; Hedges et al., 2010) to address the statistically dependent structure within each study, which enabled the inclusion of all valid effect sizes without the need to understand the covariance structure, thereby best utilizing the exacted information.
We used the R package robumeta (RVE; Hedges et al., 2010; Tanner-Smith et al., 2016) to estimate the overall effect size and to perform the meta-regression and subgroup analyses. The correlated effects weighting scheme was adopted given that the predominant dependency was derived from studies providing multiple outcomes based on the same participants. We initially used the default correlation value ρ = .80 to conduct the analyses, and then performed additional sensitivity analyses to examine the robustness of the estimation by adopting various correlation values (i.e., the correlation between the multiple effect sizes within a study for calculating the weights in the RVE correlated effects model). Furthermore, small sample size corrections were adopted if the analysis included less than 40 studies (Tanner-Smith & Tipton, 2014).
Overall Combined Effect Size
To estimate the overall statistical learning effect sizes of the DD and TD groups, we fitted an intercept-only correlated effects model (Coles et al., 2019; Tanner-Smith & Tipton, 2014). The estimate of the intercept can be interpreted as the overall effect size of the difference in statistical learning between the two groups. We used the same approach to assess the overall effect sizes of reading and cognitive moderators, including word and nonword reading fluency and accuracy, verbal and nonverbal intelligence, rapid automatized naming, verbal working memory, and phonological skills, with a positive value indicating that TD outperformed DD.
Influential Analysis
If a large amount of heterogeneity was detected in effect sizes, we examined whether there were extreme effect sizes distorting the overall estimation or influencing the effects of the moderators. We identified the influential cases in the statistical learning effect sizes through a random-effects intercept-only meta-regression using the base R functions influence measures and Baujat plot (Viechtbauer & Cheung, 2010). These R functions identify an influential case using the following diagnostic values: studentized deleted residual, DFFITS value, Cook’s distance, covariance ratios, leave-one-out τ2, Q values, hat matrix, and study weight (Coles et al., 2019). In our meta-analysis, two influential statistical learning effect sizes were detected in studies by Kahta and Schiff (2019) and X. Tong et al. (2019), with Cook’s D = 0.45, 0.10; studentized deleted residual = 7.12, 2.33; and hat matrix = 0.01, 0.01, respectively. After the removal of the two influential cases, the overall effect size changed from d = 0.56 (95% CI [0.40, 0.71]) to d = 0.47 (95% CI [0.36, 0.59]), the RVE-based τ2 decreased from 0.25 to 0.12, and the RVE-based I2 decreased from 72.42 to 56.50. It is worth noting that, since the removal of these two studies led to considerable reductions in the estimated amount of heterogeneity (i.e., decreases in τ2 and I2 by 52% and 22%, respectively), neither study was included in the subsequent effect size estimation and moderator analysis for statistical learning in DD.
In the moderator analysis, we conducted influential analyses by regressing statistical learning effect sizes onto specific individual moderators. Regarding reading skills, influential cases were detected in studies by Nigro et al. (2016), Laasonen et al. (2014), and Vaquero et al. (2021), and were therefore removed from further meta-regression analyses. In terms of cognitive variables, we also removed five influential cases found in the measurements of nonverbal IQ (Samara & Caravolas, 2017), rapid automatized naming (Inácio et al., 2018), verbal working memory (Singh et al., 2018), and phonological skills (Hedenius et al., 2021; Inácio et al., 2018).
Heterogeneity Analysis
We analyzed three types of heterogeneity measures. Q is defined as a weighted sum of squared deviations (for a review, see Hoaglin, 2016), indicating how heterogeneous the separate effect sizes are. Since the Q value was not available in the RVE, we fitted the random effects model with aggregated effect sizes to examine the Q value. We also reported the RVE-based τ2 and I2, which represent an absolute measure of between-studies variability (Schwarzer et al., 2017; von Hippel, 2015) and the ratio of true heterogeneity to total variance (Borenstein et al., 2021), respectively.
Sensitivity Analysis
In order to test the robustness of the overall effect size of statistical learning, we conducted a series of sensitivity analyses to investigate the impact of different decisions during the review and analysis processes (Higgins et al., 2019). First, we estimated the overall effect sizes under a range of possible correlations: ρ = .00, .50, .80, 1.00 in the RVE approach. Second, since an effect size extracted from means with pooled SDs might reduce the true overall effect size, we reexamined the overall effect size by excluding the statistical learning effect sizes extracted from means with pooled SDs. Third, we investigated the overall effect size on the published studies only. Last, we performed the standard meta-analysis in a random effect model to estimate the overall effect size of statistical learning, where the multiple effect sizes were aggregated using the R package MAd (Del Re & Hoyt, 2014) and implementing Gleser and Olkin’s (1994) procedure. We also generated the forest plot based on aggregated effect sizes.
Meta-Regression and Subgroup Analyses
To explore the source of between-studies variation, we evaluated the contributions of participant and methodological factors to the heterogeneity in effect sizes. Specifically, we regressed the statistical learning effect sizes on each moderator to examine their influence on the effect size differences across studies and reported all statistically significant (≤.05) and insignificant (>.05) p values.
For continuous and categorical moderators with two levels, one moderator was entered into a meta-regression equation (Coles et al., 2019). Regarding the categorical moderators that comprised more than two levels (e.g., paradigm and task modality), we performed omnibus tests by conducting Wald-tests using the R package clubSandwhich (Pustejovsky, 2017), where F values were generated for testing any difference across multiple levels. Additionally, we investigated whether interactions between reading effect sizes and orthographic depth accounted for the heterogeneity in effect sizes. The interaction terms were created by multiplying the continuous variable (i.e., fluency and accuracy in both word and nonword reading) and dichotomous variable (i.e., orthographic depth). Each interaction term and the corresponding main effects were entered and evaluated in each of the meta-regression models. To examine the statistical learning effect sizes and the heterogeneity in different subgroups, we carried out subgroup analyses for four categorical moderators (i.e., age group, paradigm, stimulus type, and task modality) by fitting an intercept-only meta-regression model.
Publication Bias
Research has shown that the studies with nonsignificant findings, small effects, or small sample sizes are suppressed and thus less likely to be published (e.g., Borenstein et al., 2021). Therefore, in this meta-analysis, we acquired unpublished articles by searching the OATD database and requesting unpublished papers directly from researchers.
In terms of analyzing the publication bias, we first aggregated the dependent effect sizes (k = 59) and then performed standard publication bias tests (Coles et al., 2019). We used these aggregated effect sizes to examine the funnel plot distribution and conducted regression analyses to confirm our visual inspection (Egger et al., 1997). If there was observed asymmetry distribution of effect sizes, we employed the Trim and Fill method (Duval & Tweedie, 2000) to adjust the effect sizes. We also implemented p-uniform analysis (van Aert et al., 2016) to assess the possibility of publication bias. The p-uniform provides a test for publication bias with statistical values and a more efficient estimator of the overall effect size.
Results
The meta-analysis results were reported in four main sections: (1) the characteristics of readers and statistical learning tasks of the included studies, (2) the overall effect size of the difference between DD and TD groups in statistical learning, (3) whether reader characteristics (i.e., age group, verbal intelligence, nonverbal intelligence, verbal working memory, rapid automatized naming, phonological skills, and the interactions between reading skills and orthographic depth) and task characteristics (i.e., paradigm, task modality, stimulus type, and quality of study) affected the differences in effect sizes between studies, and (4) the publication bias.
Characteristics of Readers and Statistical Learning Tasks
We found 59 studies in 49 papers in which people with DD were compared with TDs in statistical learning. Table 2 contains all 49 papers, which are also marked with an asterisk in the Reference list. The 59 studies that involved 2,718 participants (1,259 DD and 1,459 TD) yielded 92 effect sizes. These studies were conducted between 2003 and 2021, with the majority carried out in the past decade (k = 47). Twenty-four studies focused on adult DD and TD, while 35 assessed 7- to 13-year-old children with DD and their TD peers.
Summary of study characteristics
Note. The number after the study citation (i.e., superscripts 1 and 2) represents an independent study from the same paper. DD = developmental dyslexia; TD = typically developing; AGL = artificial grammar learning; SRT = serial reaction time; Other = other statistical learning tasks (i.e., Hebb learning, segmentation, association, contextual cueing, and target prediction).
The averaged sample size across experiments or conditions.
As shown in Table 2, 24 and 17 studies were from SRT and AGL tasks, respectively, while 23 came from other statistical learning tasks, including Hebb learning (k = 7), association (k = 7), segmentation (k = 3), contextual cueing (k = 3), and target prediction (k = 3). In terms of task modality, 49 studies were visual, eight were auditory, and seven were audiovisual. Regarding the stimulus type, 20 studies adopted linguistic stimuli and 45 employed nonlinguistic stimuli. For orthographic depth, more than half of the studies (k = 31) were conducted in shallow orthographies, while the remaining were carried out in deep orthographies (k = 28).
To examine the characteristics of readers, we compared the effect size of the difference between the DD and TD groups in terms of their cognitive and reading profiles, with a positive value indicating weaker performance in the DD group. Our results showed that the DD group was weaker than the TD group in word reading fluency (k = 34, n = 62, d = 2.21, 95% CI [1.79, 2.62], p < .001), word reading accuracy (k = 29, n = 48, d = 2.31, 95% CI [1.54, 3.08], p < .001), nonword reading fluency (k = 27, n = 48, d = 2.06, 95% CI [1.77, 2.34], p < .001), nonword reading accuracy (k = 19, n = 26, d = 2.80, 95% CI [1.55, 4.04], p < .001), nonverbal intelligence (k = 38, n = 57, d = 0.11, 95% CI [0.03, 0.20], p < .05), rapid automatized naming (k = 10, n = 13, d = 1.34, 95% CI [0.92, 1.76], p < .001), working memory (k = 15, n = 23, d = 1.01, 95% CI [0.79, 1.23], p < .001), and phonological skills (k = 16, n = 28, d = 1.40, 95% CI [1.09, 1.72], p < .001). There was no significant group difference in verbal intelligence (k = 15, n = 23, d = 0.22, 95% CI [−0.03, 0.48], p > .05). These results suggest that individuals with DD performed poorly on all literacy and cognitive measures except verbal intelligence.
Overall Effect Size Estimation
RQ1: Do Individuals With DD Exhibit Weaknesses in Statistical Learning?
Figure 2 shows the forest plot of the statistical learning effect sizes. Using RVE methods with the within-study effect size correlation specified as .80, the overall weighted effect size of statistical learning was positive, d = 0.47 (k = 57, n = 90, 95% CI [0.36, 0.59], p < .001). In terms of the between-studies variation, the true effect sizes appeared to be heterogeneous, with RVE-based τ2 = 0.12 and I2 = 56.50. A sensitivity analysis was conducted to test the robustness of the results in the RVE correlated effect model. By manipulating the ρ values from .00 to 1.00 in the robumeta package, the overall effect size and variance component estimates remained identical (d = 0.47, τ2 = 0.12), which suggested that the results generated under the default ρ value (.80) for computing within-study effect sizes were robust and reliable. Therefore, we reported analyses based only on the default value of ρ = .80. Excluding the effect sizes yielded from means with pooled SDs increased the overall effect size to d = 0.59 (k = 46, n = 67, 95% CI [0.45, 0.73], p <.001, τ2 = 0.15, I2 = 61.47), while excluding the effect sizes yielded from unpublished studies did not change the overall effect size (k = 55, n = 86, d = 0.47, 95% CI [0.35, 0.59], p < .001, τ2 = 0.12, I2 = 56.25). As shown in the forest plot in Figure 2, a standard random-effects model with aggregated effect sizes yielded similar results (k = 57, d = 0.47, 95% CI [0.35, 0.58], p < .001, Q(56) = 129.15, p < .001, τ2 = 0.13, I2 = 56.60).

Forest plot of overall effect size without influential cases.
Meta-Regression and Subgroup Analyses
A series of meta-regression analyses and subgroup analyses were conducted to explore whether reader characteristics (i.e., age group, verbal intelligence, nonverbal intelligence, working memory, rapid automatized naming, phonological skills, and reading ability for a deep or shallow orthography) and task characteristics (i.e., paradigm, task modality, and stimulus type) affected the differences in effect sizes between studies. Results are reported in Table 3.
Meta-regression of effect sizes of statistical learning on reader and task factors
Note. k = number of studies; n = number of effect sizes; CI = confidence interval.
Orthographic depth was dummy coded as 0 for deep orthography and 1 for shallow orthography.
RQ2: Do Different Characteristics of Readers or Tasks Affect the Statistical Learning Difference Between DD and TD Groups?
Age group
No significant difference was found between effect sizes obtained from studies of children and those obtained from studies of adults (k = 57, n = 90, β = 0.05, 95% CI [−0.20, 0.30], p = .704). The effect size in the adult group (k = 23, n = 38, d = 0.46, 95% CI [0.25, 0.66], p < .001, τ2 = 0.17, I2 = 64.76) was comparable to that in the child group (k = 34, n = 52, d = 0.49, 95% CI [0.34, 0.64], p < .001, τ2 = 0.09, I2 = 47.89).
Cognitive variability
No significant moderating effect was found for group difference in nonverbal intelligence (k = 36, n = 55, β = −0.10, 95% CI [−0.70, 0.50], p = .733), working memory (k = 14, n = 22, β = 0.65, 95% CI [−0.08, 1.38], p = .074), rapid automatized naming (k = 9, n = 12, β = 1.14, 95% CI [−0.38, 2.66], p = .110), and phonological skills (k = 15, n = 26, β = −0.31, 95% CI [−0.71, 0.10], p = .115). However, verbal intelligence was significant (k =15, n = 27, β = −0.39, 95% CI [−0.65, −0.14], p = .012), indicating that statistical learning in people with DD was inversely related to their verbal intelligence.
Reading × orthographic depth
To explore the possible moderating effects of reading abilities and orthographic depth on the statistical learning difference between the DD and TD groups, we created four sets of interactions between reading skills and orthographic depth (i.e., word reading fluency × orthographic depth, word reading accuracy × orthographic depth, nonword reading fluency × orthographic depth, and nonword reading accuracy × orthographic depth). Both the main effects of reading skills and orthographic depth, as well as their interaction effect, were evaluated in each model. As shown in Table 3, all main effects of reading skills and orthographic depth were not significant (ps > .05). The interaction effect of word reading fluency × orthographic depth (k = 31, n = 58, β = −0.27, 95% CI [−0.55, 0.01], p = .058) and nonword reading accuracy × orthographic depth (k = 17, n = 23, β = −0.49, 95% CI [−1.07, 0.08], p = .067) approached statistical significance, whereas word reading accuracy × orthographic depth (k = 26, n = 44, β = −0.36, 95% CI [−0.86, 0.14], p = .124) and nonword reading fluency × orthographic depth (k = 26, n = 46, β = 0.38, 95% CI [−0.42, 1.18], p = .275) were not significant.
Types of pradigms
No significant difference in statistical learning effect size was found across different types of paradigms, F(2, 31.8) = 1.88, p = .169. The subgroup analyses further revealed that overall effect sizes in different paradigms were positive and small: SRT (k = 24, n = 33, d = 0.41, 95% CI [0.20, 0.61], p < .001, τ2 = 0.13, I2 = 58.55), AGL (k = 16, n = 19, d = 0.54, 95% CI [0.25, 0.82], p < .001, τ2 = 0.22, I2 = 70.03), and Other (k = 22, n = 38, d = 0.42, 95% CI [0.24, 0.59], p < .001, τ2 = 0.07, I2 = 41.36).
Task domain
Statistical learning tasks were grouped into two types of stimuli (i.e., linguistic vs. nonlinguistic) and three types of modalities (i.e., visual vs. auditory vs. audiovisual). No significant difference in statistical learning effect size was found between linguistic and nonlinguistic stimuli (k = 57, n = 90, β = −0.18, 95% CI [−0.40, 0.04], p = .103), despite a larger effect size in studies with linguistic stimuli (k = 19, n = 27, d = 0.53, 95% CI [0.35, 0.72], p < .001, τ2 = 0.04, I2 = 30.30) than with nonlinguistic stimuli (k = 44, n = 63, d = 0.42, 95% CI [0.28, 0.56], p < .001, τ2 = 0.14, I2 = 60.71). Among all included studies, nonlinguistic stimuli were predominantly visually presented (visual: n = 61; auditory: n = 1; audiovisual: n = 1), while linguistic stimuli were comparatively more evenly presented in different task modalities (visual: n = 13; auditory: n = 8; audiovisual: n = 6). To control the impact of any modality effect, we further examined the stimulus type effect in visual statistical learning tasks. In our analysis, no significant difference in statistical learning effect size was found between visual linguistic and visual nonlinguistic stimuli (k = 48, n = 74, β = −0.08, 95% CI [−0.42, 0.27], p = .638). However, we could not examine the stimulus type effect for the auditory or audiovisual statistical learning tasks because there was only one study utilized nonlinguistic stimuli in their auditory task and another one employed nonlinguistic stimuli in their audiovisual task. Likewise, we investigated the modality effect (i.e., visual vs. auditory vs. audiovisual) only on the studies using linguistic stimuli. Our results showed no significant difference in effect sizes among the three types of modalities, F(2, 9.2) = 2.09, p = .179.
Publication Bias
Figure 3 shows the funnel plot with all aggregated effect sizes (k = 59) plotted against their standard errors. Visual inspection revealed a skewed distribution of effect sizes, suggesting that studies with larger sample variance yielded larger effect sizes. We then used rank and regression analysis (Egger et al., 1997) to confirm our visual inspection of the funnel plot asymmetry (z = 6.88, p < .001; Kendall’s tau = 0.42, p < .001). Given the asymmetric distribution of the effect sizes, we employed the trim and fill method (Duval & Tweedie, 2000) to adjust the asymmetrically distributed effect sizes. No missing studies were estimated on the left side. Last, the p-uniform analysis, which was conducted based on our original set of 59 studies, suggested that 28 studies were significant, and the publication bias was not significant, p = .993. The estimated overall effect size based on the p-uniform analysis was d = 0.53, 95% CI [0.30, 0.75], p < .001, τ2 = 0.18, which was not far different from the overall effect size estimate based on RVE correlated effects model (d = 0.47).

Funnel plot.
Discussion
This meta-analysis evaluated whether individuals with DD exhibited statistical learning weaknesses compared with their TD controls, and further investigated whether participant characteristics and/or task characteristics accounted for inconsistent findings between studies. We found that individuals with DD exhibited a significant moderate weakness in statistical learning compared with the TD controls (d = 0.47). Additionally, the meta-regression analysis revealed that only one reader characteristic, that is, verbal IQ, but no task characteristics, was inversely related to statistical learning effect size. Moreover, statistical learning effect size was marginally significantly correlated with (1) word reading fluency for a deep orthography but not a shallow one and (2) working memory. These findings demonstrated a statistical learning disadvantage in individuals with DD across orthographies and elucidated the computational cognitive processes that regulate statistical learning and reading, which were described in our proposed Statistical Learning and Reading model (SLR; see Figure 4).

A statistical learning and reading (SLR) model.
Two key novel findings provided a more complete picture of statistical learning in people with DD. First, our study extended previous meta-analysis studies, which focused solely on either SRT (Lum et al., 2013) or AGL (Schmalz et al., 2017; van Witteloostuijn et al., 2017), by showing that, compared with their TD peers, individuals with DD across orthographies exhibited poor statistical learning on all statistical learning tasks regardless of the task paradigm, with an effect size of 0.47. This finding is consistent with previous meta-analyses (Lum et al., 2013; Schmalz et al., 2017; van Witteloostuijn et al., 2017), which showed that individuals with DD performed worse than the controls, with average weighted effect sizes (measured in Cohen’s d or Hedges’ g) ranging from 0.45 to 0.47. Together, these findings suggest that individuals with DD across languages demonstrated statistical learning weaknesses on all task paradigms. Such weaknesses can be explained, at least in part, by the procedural deficit hypothesis of DD (Nicolson & Fawcett, 2007), which assumes that abnormal language-procedural learning circuits are responsible for the primary impairment in people with DD.
Another key finding is that some reader characteristics—in particular, verbal IQ, reading in an opaque orthography, and working memory, but not age or other cognitive markers (i.e., nonverbal IQ, rapid automatized naming, and phonological skills)—might account for statistical learning differences between the DD and TD groups. Specifically, verbal IQ effect size was inversely correlated with statistical learning effect size, so that individuals with DD who had higher verbal IQ also had weaker statistical learning. This result elucidates the possible interdependence between implicit procedural memory and explicit declarative memory. Specifically, verbal IQ, as a proxy of an individual’s acquired lexical and world knowledge, entails the explicit declarative memory network that stores semantic and episodic memory. Importantly, the explicit declarative memory network might compete with the implicit procedural memory network during learning (R. M. Brown & Robertson, 2007; Kim, 2020). As demonstrated by Kim (2020), the accuracy of procedural learning decreased with the amount of declarative memory formation during interleaved procedural and declarative learning; conversely, procedural learning impaired declarative learning, suggesting a reciprocal interference between the prefrontal-parietal-cerebellar and the hippocampal-prefrontal network for motor and declarative learning, respectively. Therefore, the inverse relation between verbal IQ and statistical learning, which involve an explicit declarative and implicit procedural memory subsystem, respectively, could result from the reciprocal interference of the two memory subsystems during learning.
Furthermore, our results showed that the statistical learning disadvantage tended to be related to the degree of reading impairment in individuals with DD of a deep orthography, but not a shallow one. This tendency is, in part, consistent with evidence showing that children with DD of a shallow orthography, such as Italian, were sensitive to statistical knowledge of the orthography, such as distributional properties of sound-spelling mappings (Marinelli et al., 2021). As an opaque orthography contains various sets of stable statistical regularities that relate to multilevel word constituents (e.g., in English, the letter “u” is associated with nine phonemes, but when followed by the letter “e,” it is mostly pronounced /juː/, as in “tissue”, or /u/, as in “clue”; in Chinese, the radical 台 is frequently associated with the rimes /iː/ and /ɔːy/, as in “始” /t͡sʰ iː2/ and “抬” /tʰɔːy4/, respectively), good readers can use these correlational or conditional relationships to make inferences about imperfect observations of print-sound input, thereby facilitating their acquisition of reading and spelling. In contrast, as demonstrated by S. X. Tong et al. (2020), children with DD were more distracted by the uncertainty of statistical regularities embedded in print, and exhibited slower and inefficient learning of these conditional relationships, thus impeding their reading acquisition.
It is also worth noting that, though marginally significant, the effect size of group difference between DD and TD in statistical learning appeared to be positively related to working memory, which may indicate the possible involvement of working memory in statistical learning. On a theoretical level, working memory could serve as a workspace for holding, encoding, and extracting information during statistical learning (Arciuli, 2017; Erickson & Thiessen, 2015). However, distinct from the standard definition of working memory, which emphasizes a goal-oriented simultaneous storing and processing of information (Baddeley, 2010; Persuh et al., 2018), we argue that statistical learning may rely more on an implicit working memory that operates unintentionally (Arciuli, 2017; Hassin et al., 2009; Soto et al., 2011). Additionally, the marginally significant correlation between statistical learning and working memory could be due to diverse statistical learning tasks with various statistical complexity. According to Conway (2020), working memory may play a role in statistical learning but only for learning complex, not simple, statistical regularities. Given the controversy surrounding implicit working memory (for a review, see Persuh et al., 2018), and inconsistent findings on the relationship between statistical learning and working memory, future research should endeavor to elucidate further the relation between these two constructs.
Finally, our results demonstrated that the extent of statistical learning differences between people with DD and the TD controls was not influenced by any of these task characteristics, including task paradigm, task modality, and stimulus type, which indicates that statistical learning weaknesses in DD are domain-general. In fact, the absence of paradigm effect is, in part, in line with the overlapping cognitive faculties framework, which suggests that, despite their differences, all existing statistical learning paradigms tap into the core mechanism of implicit procedural learning, that is, extracting rule-based structures or sequences (Bogaerts et al., 2021). Similarly, the lack of modality and domain specific effects in statistical learning reflects an interplay between domain generality in the statistical learning mechanism, in which similar sets of computation principles may be evoked across tasks and modalities (Frost et al., 2015), and the phenotypic heterogeneity of DD. Specifically, the modality constraints (Frost et al., 2015) and linguistic entrenchment (Siegelman et al., 2018) in statistical learning might be masked by the heterogeneous difficulties in DD across multiple cognitive domains, such as auditory and visual processing and phonological and orthographic skills (Reis et al., 2020). In general, these results provide preliminary evidence for the SLR model (see Figure 4), which is discussed in detail below.
Theoretical Implications: A Model of Statistical Learning and Reading
Despite the increasing number of studies on statistical learning in reading development and difficulties, the theoretical models are relatively underdeveloped. In particular, we know remarkably little about how statistical learning occurs across a variety of tasks and stimulus types, or what mechanisms underpin it. According to the preliminary model proposed by Frost et al. (2015), implicit statistical learning involves a set of separate processing principles for detecting statistical properties in visual, auditory, and tactile modalities, which explains modality-specific constraints and domain-general computations in behavioral (e.g., Conway & Christiansen, 2005; Emberson et al., 2011; Siegelman & Frost, 2015) and neuroimaging (e.g., Karuza et al., 2013; Shohamy & Turk-Browne, 2013; Turk-Browne et al., 2009) studies. Despite its theoretical and neurobiological validity, the framework developed by Frost et al. (2015) did not justify the role of statistical learning in a multicomponent memory system, which includes declarative and non-declarative memory.
In contrast, Sawi and Rueckl’s (2019) framework addressed the nature of statistical learning in terms of implicit/procedural memory and explicit/declarative memory. Schapiro et al. (2017), on the other hand, explored the mechanism of statistical learning using a neural network modelling approach and suggested that two separate sets of pathways within the hippocampus might be responsible for statistical and episodic learning. However, neither of these studies explicitly conceptualized statistical learning as a multicomponent ability (Arciuli, 2017) emerging from several independent memory processes (Conway, 2020; Thiessen, 2017; Thiessen & Erickson, 2013). Furthermore, none of these previous models had the core common characteristic of statistical learning and reading—uncertainty or quasi-regularity—been described and quantified. Actually, the sound-form mappings in some orthographies are more quasi-regular than regular so that some words exhibit form-sound correspondence (e.g., in Chinese, the phonetic radical 青 /t͡sʰɪŋ1/ in 請 /t͡sʰɪŋ2/, 情 /t͡sʰɪŋ4/, 清 /t͡sʰɪŋ1 /), while others deviate from these central tendencies in varying degrees (e.g., 精 /t͡sɪŋ1/, 倩/siːn3/, 猜/t͡sʰaːi1/). Thus, these previous models attempted to explain the existence of statistical learning in a variety of tasks and domains, but they did not elucidate the cognitive computational processes that enable our brain to involuntarily cope with quasi-regularity of environmental inputs, such as print, and establish the optimal representation for further learning and reading.
Incorporating our meta-analytical results into the previous statistical learning frameworks (e.g., Arciuli, 2017; Erickson & Thiessen, 2015; Frost et al., 2015; Sawi & Rueckl, 2019), we developed the SLR model (see Figure 4) to explain the mechanisms underlying statistical learning and reading and their relations. As shown in Figure 4, our proposed SLR model comprises an input layer, a multicomponent memory network, and an output layer. The input layer encodes a range of feature vectors related to statistical learning and reading, including but not limited to acoustic-temporal and visual-spatial features. These features are regulated by a system of environmental or contextual structure quantifying the level of uncertainty/quasi-regularity that is indexed by the implicitness, predictability, and complexity of sensory input related to statistical learning or word reading across modalities.
The multicomponent memory network represents an interconnecting domain-general system that consists of a short-term memory subsystem, an explicit declarative memory subsystem, and an implicit procedural memory subsystem. The short-term memory subsystem is a capacity-limited memory store that enables short-term storage of incoming and ongoing information in auditory and visual modalities (e.g., Baddeley, 2010), whereas the explicit declarative and the implicit procedural are long-term memory subsystems that differ in the intention and consciousness of information encoding and storage (e.g., Sawi & Rueckl, 2019). These three subsystems also differ in their attentional demands: the explicit learning that happens in the explicit declarative subsystem necessitates a more controlled attention (i.e., a top-down selective attention to learning stimuli) than automatic attention (i.e., a bottom-up involuntary attention to salient stimuli; Oberauer, 2019), whereas implicit learning performed in the implicit procedural subsystem relies more on automatic attention than controlled attention. However, the level of attention required for storing small amounts of stimulus-specific information in short-term memory depends on the context, with some information storing relying on automatic attention while goal-oriented manipulation of information relies on sustained and controlled attention (e.g., for a review, see Norris, 2017).
The connective strength between the short-term memory and the explicit declarative memory subsystems, and between the short-term memory and the implicit procedural memory subsystems, is modulated by the implicitness, predictability, and complexity of a specific learning task or reading activity. The more implicit, less predictable, or more complex the task is, the more the implicit procedural memory subsystem is involved compared with the explicit declarative memory subsystem. For reading and spelling, the involvement of explicit declarative memory and implicit procedural memory subsystems depends on the degree of quasi-regularity of print-sound mappings that was indexed by the complexity and predictability (i.e., orthography depth) of the given orthography.
Finally, the output layer contains the optimal response, which is a statistically optimal representation of input features as manifested by the neural and behavioral response of any activities involving mere exposure or observation. The optimal processing unit is used for high-level associative learning. Unlike the previous statistical learning frameworks, our SLR model clearly recognizes that statistical learning and reading involve a set of domain-general explicit and implicit memory processes constrained by the representation properties of task input features and individual differences (see Figure 4). Based on our meta-analysis results, the SLR model assumes three core principles and one auxiliary principle.
Principle 1: Reading as Statistical Learning
The SLR model assumes that statistical learning and reading share the same set of cognitive processes responsible for learning regularities. A written language contains various degrees of statistical regularities (such as the consistency of letter-sound mappings) that can be incidentally acquired by human learners through statistical learning (Arciuli, 2018; X. Tong & McBride, 2018). According to the SLR model, the same cognitive faculties underlie learning of regularities in statistical learning and word reading so that when the capacities of these cognitive faculties are reduced or impaired, both statistical learning and reading are affected. This principle is supported by our first key finding that individuals with DD exhibited weaker statistical learning than the TD controls across all statistical learning tasks. Based on this principle, the inefficiency of the implicit memory system may lead to a reduced capacity to recognize statistical regularities of sensory input and inhibit the interference of irrelevant deviations from these regularities when establishing optimal statistical representations for further associative learning. In fact, this assumption is also supported by several neurophysiological studies showing that a faster decay of sounds and words in adults with DD can be attributed to lower neural adaptation compared to the controls (e.g., Jaffe-Dax et al., 2017). Moreover, adults with DD exhibited attenuated prediction error arising from their failure to integrate top-down stimulus repetition into their encoding of unexpected changes in input across face and word stimuli (Beach et al., 2021).
Principle 2: Dynamic Involvement of a Multicomponent Memory System
The SLR model assumes that the short-term memory, explicit declarative memory, and implicit procedural memory subsystems interact in a competitive and complementary way during statistical learning and reading. The activation of these subsystems is indirectly constrained by the implicitness, predictability, and complexity of a specific learning task or reading activity. The more implicit, less predictable (i.e., more quasi-regular), or more complex the task is, the more the implicit procedural memory subsystem is involved compared with the explicit declarative memory subsystem. When the task goals are explicit, controlled attention will be applied to the short-term memory subsystem, turning it into a working memory subsystem for temporarily storing and processing information. Critically, explicit declarative memory may compete with implicit procedural memory during memory formation (e.g., Kim, 2020), whereas working memory may complement implicit procedural memory for holding and processing novel stimulus-specific features of individual items (e.g., Erickson & Thiessen, 2015; Hall et al., 2015). The interactions of sublearning systems within the multicomponent memory system are partially supported by our findings that verbal IQ, which denotes explicit declarative ability, was inversely correlated with implicit statistical learning weakness in people with DD, and that the group differences in working memory and statistical learning tended to be positively related.
Principle 3: The Complementary Domain-General and Context-Constrained Mechanisms Within SLR
The SLR model assumes that the cognitive and neural mechanisms underlying statistical learning and reading support the learning of regularities across various tasks in different modalities. However, this domain-general hypothesis does not deny the constraint of modality-specific mechanisms in fine-tuning or pruning of statistical representation of inputs from a specific modality. This assumption aligns well with the neural network model of the complementary learning systems (CLS) theory, which suggests that, despite the domain-general role of the hippocampus in statistical learning, establishing a robust statistical presentation may also involve modality-specific neural substrates (Schapiro et al., 2017). More important, the domain-general nature of statistical learning implies that the impact of a statistical learning deficit is far-reaching, which helps explain the phenotypes of DD and the heterogeneity of multiple deficits manifested in individuals with DD. Based on the SLR model, the short-term memory and implicit procedural memory subsystems dynamically connect with the input layer, explicit declarative memory subsystem, and output layer. Thus, any problem that occurs within the short-term memory and implicit procedural memory subsystems may lead to multiple difficulties across domains in different modalities. For example, a weakness in statistical learning of phonological and orthographic regularities may increase the risk of developing phonological and orthographic deficits, respectively. Similarly, a weakness in procedural mapping between phonological and orthographic representations may result in rapid naming deficits. All of which heightens the probability of the future development of DD (e.g., Catts et al., 2017). This assumption is supported by our findings that people with DD experienced weaknesses across statistical learning paradigms, modalities, and stimulus types; and that people with DD showed substantial weakness across cognitive domains, including rapid naming and phonological skills.
Auxiliary Principle
As described in Figure 4, the subsystems in the SLR model are interconnected, allowing for bidirectional feed-forward and feed-backward interactions. This reflects the spiral processing and learning of the SLR model, which assumes that optimal, robust statistical representation is established through the accumulation of bottom-up inputs, and, in turn, high-quality representation is fed back to the system to facilitate the encoding of newly input stimuli. More specifically, the explicit declarative memory subsystem operates in a localist connectionist manner where it stores deterministic information in local representations, with each representational unit corresponding to a specific environmental entity. In contrast, the implicit procedural memory subsystem is a distributed connectionist network that operates on probabilistic information in distributed representations, with many units mapping onto one environmental entity. The number of hidden layers within the implicit procedural memory subsystem’s multilayer structure varies across individuals, with less hidden layers in people with DD compared to TD individuals (G. D. A. Brown, 1997; Seidenberg & McClelland, 1989). Importantly, different from previous statistical learning-specific (e.g., L. Wang et al., 2021) and reading-specific (e.g., Seidenberg & McClelland, 1989; Ziegler et al., 2020) connectionist models, our SLR model operates over multifaceted inputs and representations within the same connectionist architecture. However, such a multicomponent memory system for understanding statistical learning and reading in people with and without DD requires further computational modelling and empirical testing.
Limitations and Conclusions
By combining the results of 59 studies conducted with 1,259 individuals with DD and 1,459 TD controls, this study is the first to evaluate (1) the weakness of statistical learning in individuals with DD across task paradigms, modalities, and stimulus types; and (2) the impact on this weakness of reader characteristics and statistical learning task characteristics. Despite the comprehensive nature of this study, it should be noted that the statistical power for detecting the effects of moderators is often low (Borenstein et al., 2021). Therefore, a statistically insignificant effect of a moderator (i.e., reader or task characteristic) shall not be interpreted as evidence that the effect is the same across subgroups or that no relationship exists between the statistical learning effect size and the moderators. On the other hand, a statistically significant effect of a moderator could be explained by a linear (or nonlinear) association between statistical learning and the reader/task characteristic in both the DD and TD groups, or in either. Finally, the relationships obtained in the meta-regression cannot be used to prove causality.
In conclusion, our analysis of studies comparing statistical learning between individuals with DD and the TD controls in different learning paradigms across different orthographies demonstrated a clear weakness of statistical learning in individuals with DD, regardless of task paradigm, task modality, stimulus type, and age group. Furthermore, only reader characteristics (e.g., verbal IQ), but no task characteristics, influenced the statistical learning differences between the two groups. Our findings provide a new overarching theoretical model of SLR that underscores the constraints and the interplay between explicit declarative memory and implicit procedural memory in the process of developing statistically optimal representations for cracking print codes.
Footnotes
Appendix
Note
Authors Stephen Man-Kit Lee and Shelley Xiuli Tong contributed equally to this article. This work was supported by funding from the General Research Fund (17609518 and 17620520) and Research Fellow Scheme (RFS 2021-7H05) of the Hong Kong Government Research Grant Council to Dr. Shelley Xiuli Tong.
Authors
STEPHEN MAN-KIT LEE is a doctoral student at the Speech, Language and Reading Lab, Faculty of Education, The University of Hong Kong, 801, Meng Wah Complex, Hong Kong, China; email:
YANMENGNA CUI is a doctoral student at the Speech, Language and Reading Lab, Faculty of Education, The University of Hong Kong, 801, Meng Wah Complex, Hong Kong, China; email:
SHELLEY XIULI TONG is an associate professor and director of the Speech, Language, and Reading Lab, Faculty of Education, The University of Hong Kong, Room 804C, Meng Wah Complex, Hong Kong, China; email:
