Abstract
Objective
An up-to-date meta-analysis of experimental research on talking and driving is needed to provide a comprehensive, empirical, and credible basis for policy, legislation, countermeasures, and future research.
Background
The effects of cell, mobile, and smart phone use on driving safety continues to be a contentious societal issue.
Method
All available studies that measured the effects of cell phone use on driving were identified through a variety of search methods and databases. A total of 93 studies containing 106 experiments met the inclusion criteria. Coded independent variables included conversation target (handheld, hands-free, and passenger), setting (laboratory, simulation, or on road), and conversation type (natural, cognitive task, and dialing). Coded dependent variables included reaction time, stimulus detection, lane positioning, speed, headway, eye movements, and collisions.
Results
The overall sample had 4,382 participants, with driver ages ranging from 14 to 84 years (M = 25.5, SD = 5.2). Conversation on a handheld or hands-free phone resulted in performance costs when compared with baseline driving for reaction time, stimulus detection, and collisions. Passenger conversation had a similar pattern of effect sizes. Dialing while driving had large performance costs for many variables.
Conclusion
This meta-analysis found that cell phone and passenger conversation produced moderate performance costs. Drivers minimally compensated while conversing on a cell phone by increasing headway or reducing speed. A number of additional meta-analytic questions are discussed.
Application
The results can be used to guide legislation, policy, countermeasures, and future research.
Keywords
Introduction
The effects of cell, mobile, or smart phone use on driving safety continue to be societally contentious, and negative effects have been used as the basis for distracted driving legislation in many countries. Many legislatures have passed or are considering restrictions on cell phone use while driving, including Canada, Europe, the United States, and other parts of the world (Canadian Council on Motor Transport Administrators, 2014; Governors Highway Safety Association [GHSA], 2017; Janitzek, Brenck, Jamson, Carsten, & Eksler, 2009; World Health Organization [WHO], 2011, 2015). These legislative activities have been based—in varying degrees—on politics, public perception of safety, and science. A comprehensive synthesis of experimental evidence on cell phones and driving is needed so that the empirical basis for policy, legislation, countermeasures, and future research is available for knowledge translation.
After decades of research and policy activity, talking on a cell phone while driving continues to be a common behavior identified as a contributor to crashes. During daylight hours, 3.8% of U.S. drivers, in about 542,000 vehicles, were observed holding phones to their ears in 2015 while driving during a given moment (Pickrell, Li, & KC, 2016), and the pattern of observations of drivers holding phones to their ears has declined slightly over the past 5 years (2011–2015). When surveyed, 61.8% of drivers in the United States reported that they make or accept phone calls while driving (Schroeder, Meyers, & Kostyniuk, 2013). In 2013, cell phone use—which includes talking, listening, or dialing—was a contributor to 411 fatal and about 34,000 injury crashes in the United States (National Highway Traffic Safety Administration, 2015). Statistics of crashes, injuries, and fatalities likely underestimate cell phones as a contributor (Farmer, Braitman, & Lund, 2010; GHSA, 2011) because drivers are unlikely to tell a police officer that they were using a cell phone at the time of a crash (National Safety Council, 2013, 2017). Talking on a cell phone while driving is common, contributes to crashes, and continues to be a societal safety issue.
Talking with passengers is frequently compared with talking on a cell phone and is a commonly reported and observed behavior. For example, Schroeder et al. (2013) reported that 94% of sampled U.S. drivers talked to passengers. Naturalistic studies observed that drivers spend about 11% to 15% of their time talking with passengers while vehicles are in motion (Farmer, Klauer, McClafferty, & Guo 2015b; Fitch et al., 2013; Huisingh, Griffin, & McGwin, 2015; Sayer, Devonshire, & Flannagan, 2005; Stutts et al., 2005).Passenger conversation is considered safe and socially acceptable, whereas talking on a cell phone is considered by many to be unsafe. The prevailing wisdom is that passengers regulate their speech with the driver to accommodate the demands imposed by the traffic environment. Although this hypothesis may apply to some passengers, to others it may not. Passengers are a contributor to approximately 11% of distraction-related crashes (Stutts et al., 2001). The presence of teen passengers with teen drivers increases crash risk as the number of passengers increases, and it is a known safety issue (Chen, Baker, Braver, & Li, 2000; Ouimet et al., 2015; Williams, Ali, & Shults, 2010; Williams, West, & Shults, 2012). Observational studies found that younger drivers are more likely to be talking with a passenger or on the phone (Huisingh et al., 2015; Pickrell, Li, & KC, 2016). The mechanism by which crash risk is affected by driver and passenger behavior remains to be fully elaborated but likely involves distraction (including conversation) and risk taking, manifest as speeding and hazardous maneuvers (Simons-Morton, Lerner, & Singer, 2005; Simons-Morton, Guo, Klauer, Ehsani, & Pradhan, 2014; White & Caird, 2010; Williams et al., 2010; Williams et al., 2012).
A number of qualitative literature reviews sought to summarize this vast body of literature on cell phone and passenger conversation. Substantive critiques of survey, observational, experimental, and epidemiological studies summarized portions of this research (e.g., Collett, Guillot, & Petit, 2010a, 2010b; Ferdinand & Menachemi, 2014; GHSA, 2011; Goodman, Tijerina, Bents, & Wierwille, 1999; Kircher, Patten, & Ahlström, 2011; McCartt, Hellinga, & Bratiman, 2006; Ranney, 2008; WHO, 2011). However, literature reviews are not without limitations. Reviewers tend to pick which studies to qualitatively analyze, rely on a limited number of studies, overlook study quality, do not necessarily resolve conflicting results, adopt interpretations that support the views of the authors, and reiterate the prevailing understanding of the literature. After >25 years of research, the complexity of variable relationships and the breadth of cell phone research defy logical grouping of studies, counting or listing of previous results, or repetition of conclusions, which may or may not be meaningful and reliable.
In contrast, meta-analyses require quantitative precision, intimacy with published and unpublished data, comprehensive search for studies, focused hypotheses, and identification of moderators (Rosenthal & DiMatteo, 2001). Thus, meta-analysis provides a quantitative summary of experimental research on cell phones and driving, with a variety of attractive empirical qualities, including confidence in the combined results (Murad et al., 2014).
Due to the continuous intensity of research on cell phones and driving, previous cell phone conversation meta-analyses require updating. Specifically, Caird, Willness, Steel, and Scialfa (2008) and Horrey and Wickens (2006) meta-analyzed 33 and 23 cell phone studies, respectively. Overall, reaction time (RT) to events was slower in the presence of conversation through handheld (HH) and hands-free (HF) phones. In addition, drivers did not appreciably compensate by increasing headway or decreasing speed. However, since these earlier studies, the number of publications in this area of research has at least tripled. Thus, the reliability and validity of our previous meta-analytic results and conclusions are uncertain due to the volume of publication activity. The half-life of a meta-analysis in an active area of research is estimated to be 2 to 5 years (Shojania et al., 2007). An updated meta-analysis is indicated if changes in results affect credibility or if prior meta-analyses are frequently cited (Garner et al., 2016). Thus, the purpose of this study is to update and extend our previous meta-analyses of experimental studies on cell phone or passenger conversation while driving.
This meta-analysis addresses a number of important questions by combining results from experimental studies so that effect size estimates can be calculated for a range of dependent variables. A number of conflicting results have been reported across studies, and these tend to center on the following questions: (1) Does cell phone conversation affect driving performance? (2) Do HH and HF phones have similar driving performance costs? (3) Does passenger conversation have similar performance costs as cell phone conversation? (4) Do drivers engage in compensatory behaviors while conversing? (5) Does dialing affect driving performance? (6) Are the results of studies similar across laboratory, simulator, and test track settings? (7) Do cognitive tasks such as adding two digits have similar or different driving performance effects as naturalistic conversation? (8) Are the results on mobile phone conversation from naturalistic, epidemiological, and driving simulation studies convergent or divergent? Each of these questions is addressed per a meta-analysis of the available studies.
Method
The title, rationale, methods, results, discussion, and funding sources reported here adhere to the PRISMA guidelines and checklist (i.e., Preferred Reporting Items for Systematic Reviews and Meta-Analyses; Moher, Liberati, Tetzlaff, Altman, & PRISMA Group, 2009).
Data Sources and Search Strategy
All studies examined in the work by Caird et al. (2008) and Horrey and Wickens (2006), including those previously excluded, were added to a list of publications that were analyzed per a set of inclusion and exclusion criteria. In addition, a number of databases (Embase, PubMed, MEDLINE, Web of Science, Google Scholar) were queried through keyword variants of convers*, driv*, cell phone, mobile phone, and distract* to identify studies without restriction on year of publication through December 2014. A number of sources were also systematically searched for gray literature (i.e., abstracts, technical reports, and proceedings papers): journals (e.g., Accident Analysis and Prevention, Human Factors, Transportation Research: Part F), conference proceedings (e.g., Association of Computing Machinery Special Interest Group in Computer Human Interaction, International Driving Symposium on Human Factors in Driving Assessment, Training and Vehicle Design, International Conference on Driving Distraction and Inattention, Human Factors and Ergonomics Society, and Transportation Research Board), and government Web sites (e.g., National Highway Traffic Safety Administration, Swedish Road and Traffic Institute). Cross-checking and backtracking of references in each article and consultation with authors were also used as search strategies.
Study Search and Selection
An estimated 2,462 abstracts and papers were considered for inclusion. Abstracts were screened with a set of initial criteria. First, nonexperimental research was excluded, such as survey, observations, policy, review, and epidemiological studies. Second, studies were required to experimentally examine cell phones and driving. A total of 320 studies remained after this initial screening. These studies were then examined in-depth through a more rigorous set of a priori criteria. First, a study had to measure driving performance, defined as controlling a vehicle, simulation, or proxy task with some representation of the traffic environment. Second, the study design had to include a driving-while-talking condition, as well as a baseline (BL) or control condition, with the appropriate statistical comparison. Figure 1 shows the abstract and paper review, the exclusions, and the additional screening steps.

Flow of search, exclusion, and inclusion of studies.
A study was selected for coding if it included a driving task (laboratory, simulation, or on road/test track) or a conversation task (naturalistic or cognitive) carried out over a device (HH or HF), required dialing, or involved a passenger present in the vehicle. Naturalistic conversation required some flow of information between the driver and conversant, which typically was initiated and sustained by open questions. A cognitive task, of which many varieties were used, often required a precise answer and typically required operations in working memory, such as math, verbal, or spatial tasks.
Third, coded dependent measures of driving performance included RT, target detection, lateral control, longitudinal control, eye movements, and collisions. (See Caird and Horrey [2011] and Society of Automotive Engineers [2015] for concise definitions and discussions of these measures.) Prior to 2007, eye movements, collisions, and dialing could not be meta-analyzed because of the rarity of studies, but these variables were coded in the current meta-analysis. An additional 48 studies were excluded because measures other than driving performance were used, such as electroencephalography, electrocardiography, and event-related brain potentials.
Fourth, the most complete source of information for a specific study was used for coding purposes—for example, a peer-reviewed publication or technical report. Identification of duplicate reports and papers (i.e., those based on the same data or analyses) required detailed comparisons among methods, measures, and statistical analyses. A total of 44 papers were excluded because they were duplicates.
Data Extraction and Coding
Figure 2 illustrates general measurement categories. Within these broad categories, variants of each category measure were coded from studies that met inclusion criteria. The first three authors double coded variables based on multiple checks for coding agreement. The following were coded: drivers’ RT to hazards or emergency events (e.g., pedestrian, vehicle), detection of a target (e.g., signs, secondary probes [RT, % detection]), lateral control (e.g., standard deviation of lane position [SDLP], lane exceedances [centerline, edge], other [root mean square error, lateral deviation, etc.], longitudinal control (e.g., speed, speed variance, headway [time or distance], headway variance), eye movements (e.g., horizontal or vertical scanning, proportion of glances to objects [mirrors, speedometer], and eyes off road), and collisions.

Categories of dependent variables and specific measures coded into the meta-analysis. RT = reaction time; SDLP = standard deviation of lane position.
Specific definitions of each variable follow. RT was defined as a response with a brake pedal to an event that required a response in the traffic environment, such as a pedestrian or a car pullout event. Target detection (RT, %), in contrast, did not necessarily require a response from a driver and could have been delayed or shed if conversation was prioritized. Both RT and percentage correct were coded for target detection based on responses to stimulus presentation. Common stimuli included peripheral detection tasks, where a button is pressed when a light appears in central or peripheral vision, such as on the hood, dashboard, or rearview mirror. Lateral control measures are essentially proxies for lane keeping (for discussion, see Green, 2012). A variety of measures were used across studies, with SDLP being the most common. Similarly, longitudinal measures were numerous but were most commonly related to either speed (velocity, variance, compliance) or headway (time, distance, variance). Eye movement measures were quite idiosyncratic and difficult to categorize and code. A priori hypotheses guided selection and coding of these variables. Specifically, horizontal or vertical scanning may be constrained if a driver is absorbed in conversation, and the proportion of glances off road is also likely to be reduced (e.g., Recarte & Nunes, 2003). Finally, collisions with other vehicles, pedestrians, and infrastructure objects were coded.
Many studies included multiple measures within a given dependent variable category, such as lateral position (e.g., SDLP, root mean square error, lateral deviation). When this occurs, Schmidt and Hunter (2015) recommend averaging multiple effect sizes so that participants in a single study are represented only once by the aggregate effect size. Multiple measures with similar events or trials were collapsed (e.g., RT to lead vehicle braking, pedestrian).
A variety of statistical values (e.g., F, t, p, M, SD, SE) were extracted from experiments and converted to effect sizes, r and d (Borenstein, Hedges, Higgins, & Rothstein, 2009; Lakens, 2013). As a coding strategy, means and standard deviations were most frequently extracted and coded, which typically results in conservative effect size estimates (Moser & Stevens, 1992). When necessary, estimates of means and standard deviations were extracted from figures using a painting software program, with size computed in Excel. Unequal group sizes were corrected per the methods of Borenstein et al. (2009). An additional 41 studies were excluded because they did not contain sufficient statistical information or because conditions were collapsed when comparisons were made. Many included studies did not contain sufficient statistical information for every reported dependent measure. For example, due to insufficient statistical information, collisions could not be coded from a number of studies. Coded dependent variables represent those measures with sufficient information. Nine studies were excluded because the description of experimental methods was insufficient or a study was quasi-experimental.
In our previous meta-analysis (Caird et al., 2008), we made numerous efforts to obtain additional statistical information from authors. Authors were also contacted from this round of newly identified studies if statistical information was collapsed across groups (e.g., HH and HF), omitted (e.g., null results), inappropriately analyzed, or insufficiently reported. In the current analysis, 13 of 38 author contacts missed our multiple requests, did not provide any data, could not locate their data, or provided inaccurate data, which resulted in excluding these studies. The availability of data from authors tended to not exceed 5 years, which is about the length of time that institutional review boards suggest holding data before deleting it.
Statistical Analysis
Effect sizes were computed in MetaExcel (Steel, 2014) if at least two studies used the same dependent and independent variables. Effect sizes (r, d), 95% confidence intervals (95% CIs), and 95% credibility intervals were computed. CIs specify the precision of the weighted mean effect size and indicate the sampling error in the effect size estimate. Credibility intervals estimate the generalizability of an effect size range and indicate whether moderators are likely present in the effect size estimate (Whitener, 1990). Large credibility intervals indicate the presence of significant moderator effects, which requires further exploration and interpretation (Schmidt & Hunter, 2015).
Results
Study Characteristics
Studies and experiments that were included in the meta-analysis are listed in Table 1. A total of 93 papers containing 106 experiments met all the inclusion criteria. The earliest included study was conducted in 1991, and the most recent was published in 2015. A total of 4,382 participants compose the overall sample of drivers, with ages ranging from 14 to 84 years (M = 25.5, SD = 5.2). More males (n = 1,877) than females (n = 1,623) participated in studies in which sex was reported by researchers.
Studies and Effects Coded Included in the Meta-Analysis
Note. BL = baseline; Exp = experiment; F = female; H = horizontal; HH = handheld; HF = hands-free; LTP = lateral position; M = male; RT = reaction time; SDLP = standard deviation of lane position; V = vertical.
Values include age in years, range and M (SD), as well as number of participants by group.
Headway includes time or distance headway. Other includes lateral deviation, root mean square error, and so on.
Quantitative Results
A total of 106 experiments contributed to this meta-analysis. Results tables that follow list the number of effect sizes or experiments (k) that were used to calculate mean effect sizes for cell phone, passenger, and dialing effects. In addition, the number of participants, mean effect sizes (r) and Cohen’s d, 95% CIs, and credibility intervals are listed. To reiterate, credibility intervals estimate the generalizability of an effect size range and indicate whether moderators are likely present in the effect size estimate (Schmidt & Hunter 2015; Whitener, 1990). The summary statistic r is primarily used to interpret the dependent variables. Interpretations and comparisons of listed effect sizes within the tables are made with respect to low (.1), moderate (.3), and large (.5) effect sizes (Cohen, 1988).
Cell phone effects
In experiments that used RT to hazard or emergency events, such as a pedestrian or a braking lead vehicle, HF and HH phones had moderate effect sizes of .25 and .27, respectively (see Table 2). The effect size differences between HH and HF phones were negligible. Detection RT to a target (e.g., a sign) or a secondary probe (e.g., a light-emitting diode) resulted in large effect sizes for HF (.49) and HH (.61), with the latter having a slightly greater cost. The percentage of targets detected had a large decline with use of an HF phone (–.52), whereas the effects between HH and BL (–.40) and between HH and HF (–.05) were medium and negligible, respectively. Thus, conversation resulted in increased RT, detection RT, and decreased detection percentage for HF and HH phones.
Meta-Analyses of the Effects of Cell Phone Use on Driving Performance Variables
Note. k = number of samples; N = total number of participants; r = weighted mean correlations; d = Cohen’s d effect size transformed from r; CI = confidence interval; CrdI = credibility interval; HF = hands-free; HH = handheld; SDLP = standard deviation of lane position.
Lateral position—which includes the variables SDLP, lane exceedance, and other (i.e., lateral position variables other than the previous two)—had small or negligible effect sizes for HF, HH, and differences between the two phones. Conversation with use of either phone did not affect speed. However, in studies that compared both phone types, those who used an HH phone decreased their speed more (–.16) than those who used an HF phone. When compared with BL driving, speed variance was greater with HF phones (.22) than HH phones (.10). Drivers who were talking on an HF phone were less compliant with speed limits (–.35) than they were during BL driving. Conversation did not affect lateral position over a number of variables, whereas speed, speed variance, and compliance had small or medium effect sizes.
Measures of headway, either time or distance, resulted in small effect size differences from BL for HH (.21) but not HF (.06) phones. Headway variance was greater while talking on an HF than during BL driving (.31). More collisions resulted while conversing on HF (.31) and HH (.20) phones than during BL driving. However, just three HH studies had sufficient collision information.
The eye movement measures of horizontal (–.36) and vertical (.26) dispersion and off-road glances (–.27) had small to medium effects. Across five experiments, drivers scanned the horizontal field of view to a lesser degree while conversing than during BL driving.
Passenger effects
In terms of conversation, far fewer studies directly examined passenger with HF or HH phone than with cell phone. For RT, a negligible effect size difference was found between passenger and HH or HF conversation (Table 3). When compared with BL driving, passenger conversation resulted in slower detection RT (–.21) and a decline in the percentage of targets detected (.55). Speed, speed variance, and headway had small effect size differences between HH, passenger, and BL and passenger differences. When measures of lateral positioning were pooled (i.e., SDLP, exceedance, and other), HH phones had a small effect size than that of passenger conversation, which had a negligible effect size. Lateral position, speed, speed variance, headway, and collisions each had only two or three experiments, which decreases the stability of the results. Collisions occurred more frequently while using an HF phone than when talking with the passenger (.44), and BL driving had fewer collisions than when talking with the passenger (–.31).
Meta-Analyses of the Effects of Passenger on Driving Performance Variables
Note. k = number of samples; N = total number of participants; r = weighted mean correlations; d = Cohen’s d effect size transformed from r; CI = confidence interval; CrdI = credibility interval; HF = hands-free; HH = handheld.
Standard deviation of lane position, exceedance, other.
Dialing effects
A limited number of studies met the inclusion criteria for dialing. Across coded variables, effect sizes were either large or medium size (Table 4). Specifically, the effect of dialing on the detection RT (.80) was large, which indicated that reactions were prolonged to targets while dialing. When compared with BL, the effect of dialing on lateral position–SDLP (.57), headway variance (.62), speed (–.66), and glances off road (.92) was also large. Lateral position–lane exceedance (.34) and speed variance (.21) had medium and small effect sizes, respectively. The effects of dialing on driving performance were larger and more adverse than those found for conversation. Dialing requires a driver to divert one’s eyes from the roadway, and the results are similar to the pattern of effect sizes found for typing texts and driving (Caird, Johnston, Willness, Asbridge, & Steel, 2014).
Meta-Analyses of the Effects of Dialing on Driving Performance Variables
Note. k = number of samples; N = total number of participants; r = weighted mean correlations; d = Cohen’s d effect size transformed from r; CI = confidence interval; CrdI = credibility interval; SDLP = standard deviation of lane position.
Publication Bias
Sample size metaregression
Metaregression, which is the regression of effect sizes, was used to check for potential publication biases (Schmidt & Hunter, 2015). To assess whether sample size influenced the magnitude of the meta-analyzed effect size, weighted least squares regressions were run with sample size, N, set as predictor and with effect size, r, set as criterion. Only comparisons for which at least 10 studies contributed to the meta-analyzed effect size were assessed. In total, 10 weighted least squares regressions were run, with adjustments for multiple comparisons. Based on the cutoff p value of .001, the relationship between sample size and effect size did not reach statistical significance in eight of 10 analyses: F(1, 11) = 0.770, p = .399 (HH vs. BL, RT); F(1, 25) = 0.355, p = .556 (HF vs. BL, SDLP); F(1, 8) = 6.024, p = .040 (HH vs. BL, speed); F(1, 17) = 0.031, p = .863 (HF vs. BL, detection RT); F(1, 11) = 0.072, p = .793 (HF vs. BL, detection percentage); F(1, 15) = 3.661, p = .075 (HF vs. BL, lateral position, other); F(1, 10) = 1.575, p = .238 (HF vs. BL, speed variance); and F(1, 13) = 3.390, p = .089 (HF vs. BL, headway). Sample size was found to predict effect size magnitude for RT, F(1, 33) = 14.521, p = .001 (HF vs. BL), and for speed, F(1, 28) = 22.862, p < .001 (HF vs. BL), in a negative direction (r = –.553 and r = –.670, respectively), indicating that smaller sample sizes tended to be associated with larger reported effect sizes. These effects are suspected to be the result of the inclusion of outliers in the analysis. After removal of effect sizes flagged as outliers (i.e., >3 SD above the mean), the results were no longer significant for either RT, F(1, 29) = 0.374, p = .546, or speed, F(1, 24) = 0.123, p = .728. Exclusion of outliers did not affect the meta-analyzed effect sizes in either case.
Date of publication metaregression
To assess whether effect sizes for RT, lateral control, and speed have changed in magnitude over time, weighted least squares regressions were run with year of publication set as predictor and effect size, r, set as criterion. Comparisons were assessed only for which at least 10 studies contributed to the meta-analyzed effect size and for which RT, SDLP, and speed were compared with BL. Five weighted least squares regressions were run in total, with adjustments for multiple comparisons. Based on the cutoff P value of .01, the relationship between year of publication and effect size failed to reach statistical significance in four of five cases: F(1, 33) = 4.673, p = .038 (HF vs. BL, RT); F(1, 11) = 4.472, p = .058 (HH vs. BL, RT); F(1, 25) = 3.733, p = .065 (HF vs. BL, SDLP); and F(1, 8) = 2.708, p = .138 (HH vs. BL, speed). Year of publication predicted effect size magnitude for speed (HF vs. BL), F(1, 28) = 16.376, p < .001, in a negative direction (r = –.607), indicating that smaller effect sizes tend to be reported in more recent publications. However, after removal of the same four outliers as in the weighted least squares regression analysis of publication bias, the results were no longer significant, F(1, 24) = .000, p = .991. The meta-analyzed effect sizes did not change appreciably.
Moderator Analyses
The widths of a number of credibility intervals (see Table 2) indicate the presence of a number of potential moderator variables. Moderator analyses for research setting and conversation type for the variables of RT (Table 5), detection RT (Table 6), and speed (Table 7) are summarized. For HF phone RT, the effects across research settings were similar. Detection RT indicated larger effects as the setting moved from laboratory to test track or on road. In this setting, speed was decreased, whereas in the setting of driving simulators, it was not. Effect size differences between naturalistic and cognitive conversation indicate that the latter had a greater performance cost than the former for RT and detection RT but no differences for speed.
Moderator Analyses for the Effects of Hands-Free Cell Phone Use on Reaction Time
Note. k = number of samples; N = total number of participants; r = weighted mean correlations; d = Cohen’s d effect size transformed from r; CI = confidence interval; CrdI = credibility interval.
Moderator Analyses for the Effects of Hands-Free Cell Phone Use on Detection Reaction Time
Note. k = number of samples; N = total number of participants; r = weighted mean correlations; d = Cohen’s d effect size transformed from r; CI = confidence interval; CrdI = credibility interval.
Moderator Analyses for the Effects of Hands-Free Cell Phone Use on Driving Speed
Note. k = number of samples; N = total number of participants; r = weighted mean correlations; d = Cohen’s d effect size transformed from r; CI = confidence interval; CrdI = credibility interval.
Discussion
The purpose of this meta-analysis was to update and extend previous meta-analyses of cell phone and passenger conversation while driving. Discussion of the pattern of effect sizes across the coded dependent measures is organized per the eight questions listed at the end of the introduction.
Question 1: Does Cell Phone Conversation Affect Driving Performance?
Conversation with an HH or HF phone was compared with driving without conversing on a phone. Medium and large effect sizes were found for RT and detection RT, respectively. When conversing, drivers responded somewhat slower to important events in the driving environment, such as a lead vehicle braking or a pedestrian suddenly entering a crosswalk. Drivers also detected and responded slower to targets, such as secondary probes and traffic signs that did not necessarily require an immediate response. Collectively, conversation on a cell phone did not result in compensatory performance adjustments, such as increasing headway or reducing speed. The performance costs to horizontal and vertical scanning and the proportion of glances to off-road locations, such as mirrors and the speedometer, while conversing were low or moderate. Collisions with vehicles, pedestrians, and infrastructure were greater while talking than when not.
In our previous meta-analysis, RT to emergency events and RT to secondary targets were coded together, which resulted in effect size estimates of .46 for HF phones and .55 for HH phones (Caird et al., 2008). In the present study, responses to emergency events were separated from secondary target responses for methodological and theoretical reasons. Secondary tasks may be ignored or shed by drivers because responses were not necessarily required. Greater decrements occurred when dealing with nonessential tasks, such as detection RT (.61 HH), whereas smaller interference effects occurred for driving-relevant RT tasks, such as lead vehicle braking (.27 HH). The pattern of effect sizes observed in the current analysis would seem to corroborate this modest prioritization and preservation of emergency event responses. Lateral and longitudinal control, as well as hazard perception, is likely prioritized and somewhat preserved, whereas nonessential secondary tasks may be attended less frequently.
The division of RT effects into emergency and secondary task categories separates the performance costs of conversation into essential and nonessential tasks, which have different performance costs and theoretical implications (Trick, Enns, Mills, & Varvik, 2004). In the context of task adaption or workload management, drivers will apportion the decrement according to the relative priority of tasks, shedding or modulating those that are least critical to successful driving (e.g., Wickens, Gutzwiller, & Santamaria, 2015). In hierarchical models of driving (e.g., Michon, 1985), the schedule of task shedding in light of increasing levels of load should, in theory, begin with nonrelevant tasks, followed by driving subtasks that are not immediately necessary for safe vehicle control and hazard avoidance (which should remain as the most critical tasks for safe driving). Models of demand regulation and calibration in driving have elaborated on the interplay between momentary driving and task demands, drivers’ capabilities, behavioral adaptation to keep the two in alignment, and situations where this alignment breaks down (e.g., Fuller, 2005; Horrey, Lesch, Mitsopoulos-Rubens, & Lee, 2015; Kuiken & Twisk, 2001).
Question 2: Do HH and HF Phones Have Similar Performance Costs?
Similar RT effect sizes resulted for HF and HH phones. HF and HH performance while conversing was similar for the variables of RT, detection RT, detection percentage, lateral positioning (SDLP), and speed. One practical difference between HH and HF phones is that the former requires the driver to take one hand off the steering wheel and hold the phone to an ear or pin the phone between the chin and shoulder. Interactions with a HF phone may require the driver to manipulate a device or wires if the phone is not automatically integrated into the vehicle. If in motion, connecting a phone by interacting with hardware and software requires the driver to take his or her eyes off the road, which may increase crash risk.
There is some indication that after the introduction of restrictions on HH cell phone use, HF use rose (McCartt, Kidd, & Teoh, 2014). Observed headset use, as a means of HF interaction, has been relatively constant at about 0.5% of observed drivers during daytime hours in the United States (Pickrell, Li, & KC, 2016). The practical enforcement issue associated with HF phone use is one of obtaining sufficient visual evidence that talking to oneself in the vehicle is in fact over an HF device. What constitutes sufficient and reliable evidence before a court requires clarification. Because HH and HF phone conversation produces similar driving performance costs, existing legislation that targets only HH phones may require reconsideration.
Question 3: Does Passenger Conversation Have Similar Costs as Cell Phone Conversation?
Conversation with passengers is generally socially accepted and nearly universally common. In the current analysis, effect size differences between cell phone and passenger conversations were minimal across a number of dependent variables. Conversation with a passenger or with someone on a cell phone produced similar RT, detection percentage, lateral position, speed, speed variance, headway, and collision costs (see Tables 2 and 3). The cognitive effects of conversation affect the availability of attention, depending on degree of processing or absorption. The cognitive demands of concurrent driving and passenger conversations may produce slowing or modulation of speech patterns, affective absorption, and social learning effects that are as yet not sufficiently described (Briggs, Hole, & Land, 2011; Caird & Horrey, 2017) but may contribute to collision risk in young drivers (Chen et al., 2000; Ouimet et al., 2015). Based on available studies to date, the cognitive costs of conversation on driving performance are similar to those exerted by cell phone conversation. The pattern of effect sizes across studies in this meta-analysis stands in contrast to a number of frequently cited studies (e.g., Drews, Pasupathi, & Strayer, 2008).
Question 4: Do Drivers Engage in Compensatory Behaviors While Conversing?
For HF and HH phones, lateral position, speed, and headway were minimally affected by conversation. These results do not support the conclusion that drivers compensate while using a mobile phone (McCartt et al., 2006; WHO, 2011). For example, models of demand regulation suggest that drivers adjust their behavior to keep driving demand or difficulty within a tolerable range (Fuller, 2005). This can be manifest through decreases in speed or increases in headway as driver workload increases (e.g., slower speeds and longer headways might offer drivers more time and space to react to braking vehicles or other events). Determination of whether drivers make intentional longitudinal adjustments to maintain performance or whether inattention to vehicle control slightly degrades with increases in cognitive demand requires further investigation.
Some studies found that lateral vehicle control seems to improve when drivers are talking on the phone or engaged in a cognitive task (e.g., Carsten & Brookhuis, 2005; Horrey & Simons, 2007; Remier, Mehler, Coughlin, Roy, & Dusek, 2011), whereas others found an increase in lane variability (He, 2012). However, the results of this meta-analysis show negligible effects of conversation on lateral position, which suggests that conversation neither degrades nor improves lane control. The absence of degradation could be due in part to the relative automaticity of lane keeping and drivers’ abilities to process this information with noncentral (peripheral) resources (e.g., Horrey & Wickens, 2006; Summala, Nieminen, & Punto, 1996). The number of poorly defined lateral control variables that were used in a number of included studies do not necessarily help to resolve what is likely a minimal effect size.
Question 5: Does Dialing Affect Driving Performance?
Dialing had large costs for detection RT, lateral position (SDLP), headway variance, speed, and glances off road. Repeatedly taking the eyes off the road to enter numbers and initiate a call increases RT to targets, decreases lateral control and headway maintenance, and reduces sampling of mirrors and the speedometer. Based on other studies that examined eye movements, dialing a long-distance number, which requires longer and more frequent glances away from the road, is likely to increase crash risk (Horrey & Wickens, 2007; Simons-Morton et al., 2014; Tivesten & Dozza, 2014). The nature of dialing has changed as mobile phones have technologically evolved, with speed-dialing options becoming the dominant way to call known contacts. Calling someone can also be achieved by voice recognition, which has a number of other distraction costs (Simmons, Caird, & Steel, 2017). Connecting or coupling of cell phones to vehicles requires drivers to interact with the vehicle to dial contacts. Minimization of the number of steps required to dial a number with a vehicle will reduce the visual-manual load on the driver. Overall, drivers who take their eyes off the road can miss important events and may require lateral and longitudinal corrections, which is similar to typing short text messages (Caird et al., 2014).
Question 6: Are the Results of Studies Similar Across Laboratory, Simulator, and Test Track Settings?
Researchers make decisions to investigate mobile phone distraction questions with different methodological resources. Previously, no effect size differences were found across the methodological settings of laboratory, simulator, and test track (Caird et al., 2008; Horrey & Wickens, 2006). In the present meta-analysis, RT, detection RT, and speed were compared across settings. RT effects were similar across settings. Detection RT indicated larger effects as the setting moved from laboratory to test track or on road. In the setting of test track or on road, speed decreased, whereas in the setting of driving simulators, it did not. In general, conversation task demands yield similar effect sizes for certain variables across laboratory, simulator, and test track settings, with some minor variations.
Question 7: Do Cognitive Tasks Such as Adding Two Digits Have Similar Driving Performance Effects as Naturalistic Conversation?
Are cognitive tasks valid proxies for actual conversation, or do cognitive tasks represent the high end of workload demand (e.g., Strayer et al., 2015)? Previous meta-analyses found no effect size differences between cognitive tasks and naturalistic conversation (Caird et al., 2008). The results of the current meta-analysis are somewhat similar to those previously found. Naturalistic conversation had a greater performance cost than did cognitive tasks for RT and detection RT, but no differences were found for speed.
Question 8: Are the Results on Mobile Phone Conversation From Naturalistic, Epidemiological, and Driving Simulation Studies Convergent or Divergent?
Across naturalistic, epidemiological, and experimental simulator studies, are the effects of talking and driving similar? Each methodological approach and setting has strengths and weaknesses. The belief that a single methodological approach will yield results with more absolute truth is scientifically naïve. A number of biases and study quality limitations of each method affect the generalizability of research, whether conducted through epidemiological (Elvik, 2011), naturalistic (Simmons, Hicks, & Caird, 2016), or driving simulator methods (Caird & Horrey, 2011).
A meta-analysis of 12 epidemiological studies of cell phone use and driving by Elvik (2011)—which contained a number of studies of low quality with several exceptions (McEvoy et al., 2005; Redelmeier & Tibshirani, 1997)—found a crash involvement odds ratio of 2.86, 95% CI [1.72–4.75], for conversation and driving. A number of naturalistic studies reported that cell phone conversation does not increase crash risk and may, in some circumstances, provide a protective effect (Dingus, 2014; Farmer et al., 2015a; Fitch et al., 2013; Klauer, Dingus, Neale, Sudweeks, & Ramsey, 2006; Klauer et al., 2014). More recently, Dingus et al. (2016) analyzed the SHRP 2 (Second Strategic Highway Research Program) naturalistic data set for distraction contributors. The SHRP 2 naturalistic data set contains data for 3,542 drivers aged 16 to 98 years collected over 3 years who were located near six centers across the United States. Specifically, conversation on an HH phone produced an odds ratio of 2.2, 95% CI [1.6–3.1], for injury and property damage crashes.
Across meta-analyses of naturalistic and epidemiological studies, moderate increases in crash risk are evident for conversing and driving. In the present meta-analysis, medium effect sizes were found for the driving performance variables of RT, detection RT, detection percentage, and collisions. Moderate increases in crash risk and costs to driving performance when conversing on a mobile phone and driving are evident. Thus, meta-analytic results across naturalistic, epidemiological, and experimental methods are convergent.
Implications for Countermeasures
A number of organizations have assembled lists of actual and potential countermeasures aimed at reducing driver distraction (e.g., GHSA, 2011; WHO, 2011). Media campaigns (print, television, social), passenger restrictions (e.g., during graduated licensing), technology constraints (e.g., apps that divert calls), driver monitoring (e.g., feedback systems), and social norms (e.g., drink driving = distracted driving) are some of the types of interventions that may be effective at reducing distracted driving (Buckley, Chapman, & Sheehan, 2014; Caird et al., 2014; GHSA, 2011; National Highway Traffic Safety Administration, 2013; Phillips, Ullberg, & Vaa, 2011; WHO, 2011). To date, the relative effectiveness of distraction countermeasures is likely modest at best (e.g., Caird & Horrey, 2017). Among various organizations, calls to action, even without evidence, may be required to reduce injuries and deaths in a timely manner (Pless & Pless, 2014). Most researchers agree that the target of countermeasures should be the collection of behaviors associated with using a cell or smart phone, which includes conversation, texting, dialing, and other forms of interaction—especially those that take the eyes from the roadway (Caird et al., 2014; Dingus et al., 2016).
Legislation and enforcement
As of 2015, 131 countries prohibit HH phone use while driving, whereas another 31 prohibit HH and HF (WHO, 2015). The GHSA (2017) maintains the status of U.S. legislation by state and type of law, and the Canadian Council on Motor Transport Administrators (2014) lists Canadian provincial legislation.
Legislation and enforcement appear to reduce observed cell phone use in state population samples (McCartt, Hellinga, Strouse, & Farmer, 2010) but not necessarily among young drivers (Goodwin, O’Brien, & Foss, 2012). The outcomes of legislation on collisions, injuries, or fatalities are exceptionally difficult to determine (McCartt et al., 2014). High-quality studies are few, with some finding reductions, some no effects, and others increases in crashes. Reasons for not finding effects may include not fully implementing HF bans, the difficulties in performing these types of epidemiological studies, and variance of statistical approaches. Many drivers purposefully or ignorantly violate these cell phone use restrictions. Making the consequence of doing so sufficiently unpalatable requires punishment (McCartt et al., 2014). Fines, demerit points, increases in insurance premiums, and deductibles may increase compliance. Despite legislation and social restrictions on cell phone use while driving, many drivers still talk, text, and interact with their phones, and changing these behaviors will be a monumental challenge.
Global impact
Research and interventions on cell phone use and driving have been a focus in industrialized countries. However, the impact of mobile phone use on traffic safety worldwide is fundamentally not known (WHO, 2015). Translation of effective and appropriate countermeasures based on high-income countries to comparable low- and middle-income countries is likely to have an unknown effect on collisions, injuries, and fatalities. Effects will be unknown unless cell or smart phone use as a crash contributor is tracked in a national or regional database. As we know, however, driver reports of whether a mobile phone was in use when a crash occurred are underreported and potentially unreliable (National Safety Council, 2013, 2017). Establishing that a smart phone was in use when a crash occurred may be possible through electronic queries of mobile phones and vehicle black boxes, but doing so raises numerous privacy concerns.
Limitations
Many dependent variables had wide credibility intervals, which is indicative of the presence of moderating variables (Schmidt & Hunter, 2015). Potential moderators include method setting (laboratory, simulator, test track), year of publication (1991–2015), age (teen, adult, older), traffic flow level (low, medium, high), road type (urban, rural, suburban, highway), intersections (e.g., yellow lights, stop signs, uncontrolled intersections), individual skill sets (e.g., pilots, gamers versus non), chronic diseases (e.g., attention-deficit/hyperactivity disorder, traumatic brain injury), phone type (e.g., make, model, type of interactivity), language spoken (English, French, etc.), and sex (male, female). Each variable represents potential contributions to heterogeneity when multiple studies are combined. Some of these variables were tested through metaregression, including research setting, conversation type, and year of publication. Other studies used variables that did not fit with inclusion criteria, such as responses to being distracted while trying to stop at yellow lights (Ohlhauser, Boyle, Marshall, & Ahmad, 2011; Xiong, Narayanaswamy, Bao, Flannagan, & Sayer, 2016). However, many of these variables could not be quantitatively examined.
For example, a number of teen and older driver studies were considered for analysis in the present study. The prevailing hypotheses are that teen and older driver groups are affected by conversation to a greater degree than adult drivers. In this meta-analysis, we coded teens (Experiments 14, 56, 94, 104) and older drivers (Experiments 1, 24, 45, 46, 47, 71, 72, 81, 85, 87, 93, 96, 102, 104; see Table 1). However, a number of statistical limitations made analysis prohibitive, such as collapsing of means, missing post hoc tests, and reporting omissions. Qualitatively, older drivers in the studies considered for meta-analysis performed similarly to adults across variables. In general, the sample of older drivers were relatively young (65–74 years). Differences between teen and adult driver performance measures were minimal (e.g., Chisholm, Caird, Teteris, Lockhart, & Smiley, 2006; Stavrinos et al., 2013). Comparisons across age groups in the SHRP 2 naturalistic data set (Guo et al., 2017) found that younger (16–29 years) and older (65–98 years) groups had increased crash risk when engaged in talking on an HH phone.
The timing of talking and listening measures of driving performance are collected but rarely reported or sufficiently controlled experimentally. In many studies, the assumption is that participants were actively engaged in conversation when performance measures were recorded. In studies that do gather data for measures of cognitive task performance, the manner of aggregation and analysis makes it difficult to determine whether the conversation was shed or modulated at the time that other (critical) driving events took place.
Notwithstanding the ability to evaluate nuances and trade-offs in performance, study quality varied across the studies listed in Table 1. Remarkably, many studies failed to report the basic characteristics of drivers who participated, including sex, age, and whether they possessed valid driver’s licenses or wore appropriate corrective lenses. Operational definitions of independent and dependent variables, such as the HF task and lateral position measurement, were frequently difficult to determine. Future studies that seek to fill in gaps in cell phone research need to adopt common definitions and measures. Specifically, definitions of driving performance (Society of Automotive Engineers, 2015) and eye movements (International Standards Organization, 2001, 2002; Society of Automotive Engineers, 2000) should be adopted, as requirements for publication and studies that do not adhere to these minimal requirements should be rejected or required to be revised.
Many studies did not provide sufficient statistical information, including null effects, interactions, and post hoc comparisons to satisfy the needs of the current analysis, which may reflect a bias in research sophistication (Rosenthal & DiMatteo, 2001). Reviewers and editors should be cognizant of these important elements in newly submitted works and help focus on improving the quality of studies that are either replications of existing research, which are sorely needed, or those that address new and interesting questions, which tend to be overemphasized. Replication of empirical studies is the cornerstone of advances in safety science, and inadequate quality and lax reporting deter from this important goal (Gilbert, King, Pettigrew, & Wilson, 2016; Ioannidis, 2005; Open Science Collaboration, 2015).
Future Directions
The evolution of technology has decreased phone size and vastly increased functionality. Although the user experience of interacting with a cell phone has changed over the past 30 years, talking with someone likely has not. As different social and communication needs and means arise, drivers will engage in numerous interrelated extracurricular behaviors that require a variety of interactions in addition to talking. Focusing on one task type, such as talking or texting, ignores the reality of momentary changes in communication preferences of the driver, such as talking to a passenger, texting a friend, sharing a video or picture, posting a comment to social media, or calling a client. Our complete understanding of these complex patterns of behavior that evolve with technology may require a broader (and perhaps philosophical) consideration of the underlying motivations (Hancock, Mouloua, & Senders, 2009). Many studies have conceived of each task as being separate with distinct costs and not one of multiple interrelated distracting behaviors that meet specific driver needs. The collection of behaviors and associated risks also requires estimation of what happens when drivers replace a relatively safe task with one that is less safe or vice versa (e.g., Farmer et al., 2015b; Foss & Goodwin, 2014; Funkhauser & Sayer, 2012; Sayer et al., 2005). The rate of adoption and use within the vehicle of new means of personal or group communication is influenced by the age and gender of the driver, pressures from work and peer networks, cultural norms, and personal safety values. Researchers must continue to uncover and explore these intricate patterns of interdependencies to expand our understanding of these complex and important road safety concerns.
Key Points
The effects of cell phones on driving continue to be a contentious societal issue.
All available studies that met inclusion criteria were meta-analyzed.
Conversation on either a handheld or hands-free phone produced moderate performance costs for a number of variables.
Passenger conversation had a similar pattern of results.
Dialing and driving had large performance costs.
Drivers did not appreciably compensate by reducing speed or increasing headway while conversing.
These meta-analytic results have implications for policy, legislation, countermeasures, and future research.
Footnotes
Acknowledgements
We graciously thank the authors who provided additional statistical and methodological information to us. The AUTO21 Network of Centres of Excellence funded this research through the Convergent Evidence From Naturalistic, Simulation and Epidemiological Data Network. An abstract of preliminary results from this meta-analysis was presented at the Third International Conference on Driver Distraction and Inattention in Sydney, Australia. W. Horrey’s contributions came, in part, while he was at the Liberty Mutual Research Institute for Safety (Hopkinton, MA).
Jeff K. Caird is a professor in the Department of Psychology and an adjunct professor in the Department of Community Health Sciences at the University of Calgary, Canada. He received his PhD from the University of Minnesota in 1994.
Sarah M. Simmons is a PhD student in psychology at the University of Calgary. She received her MSc in psychology at the University of Calgary in 2016.
Katelyn Wiley is a PhD student in human-computer interaction at the University of Saskatchewan. She received her MEDes and BSc from the University of Calgary.
Kate A. Johnston, human factors specialist, General Dynamics Mission Systems—Canada, MSc, University of Calgary. She currently researches tactical communications for military and public safety populations.
William J. Horrey is the traffic research group leader at the AAA Foundation for Traffic Safety in Washington, DC. He earned his PhD in engineering psychology from the University of Illinois at Urbana-Champaign in 2005.
