Abstract
In this article, we combine narrative disclosures in the Management Discussion and Analysis (MD&A) Section of 10-K reports with financial variables to generate explicit firm-level forecasts of 1-year-ahead return on equity (ROE). Testing our forecasts out-of-sample, we find that models enhanced with MD&A disclosures are more accurate than models using quantitative financial variables alone. We also find that text-enhanced models are as good as or better than analyst consensus forecasts for small firms and firms with low analyst following. We provide evidence on the informativeness of the MD&A section across different firm characteristics. We find that firms with large changes in future performance, negative future performance, high investor scrutiny, high distress risk, high positive accruals, and high value relevance of earnings have more informative MD&A disclosures, whereas younger firms and firms facing greater market risk and litigation exposure have less informative MD&As. Our results are robust to alternative empirical choices regarding benchmarks and specifications. Overall, our inferences are consistent with MD&A disclosures positively contributing to firms’ information environments.
Introduction
Accurately forecasting future earnings has been at the center of accounting research for over four decades. Studies have developed earnings forecasting models that employ line items or ratios from financial statements, decompose earnings into components, and incorporate cost information in periods of sales declines (e.g., Banker & Chen, 2006; Fairfield, Sweeney, & Yohn, 1996; Ou & Penman, 1989; Sloan, 1996). In this article, we further enhance earnings forecasting models by incorporating textual (i.e., qualitative) disclosures into earnings forecasts. Rather than focusing on specific topics in Management Discussion and Analysis (MD&A) or aggregate sentiment, our forecasts allow the text to “speak for itself.” 1 We assess whether the text from the MD&A section of 10-K reports improves the accuracy of out-of-sample earnings forecasts relative to traditional quantitative models or analysts’ consensus forecasts. We provide evidence on the firm characteristics and settings in which textual information has the greatest predictive power.
We focus on management qualitative disclosures contained in the MD&A section of 10-K reports for several reasons. First, all publicly traded firms in the United States must provide MD&A disclosures to allow investors to see the company and its prospects through the eyes of managers and have been required to do so since 1980. Second, survey evidence in prior studies suggests that MD&A is arguably the most read section of the annual report - out of all disclosure items of the annual report, financial analysts and investors rely on MD&A statements the most (Rogers & Grant, 1997; Tavcar, 1998). Third, unlike voluntary disclosures in conference calls, press releases, and investor meetings, MD&A disclosures are subject to regulatory requirements and are more likely to be scrutinized by auditors, analysts, investors, and regulators. Finally, MD&A disclosures are available for all publicly traded firms over a relatively long time period, allowing us to draw meaningful inferences about the benefits of MD&A content.
Many empirical studies find an association between MD&A properties (e.g., the amount, tone, and readability of disclosure) and firm characteristics or market returns (see Li, 2010b; Loughran & McDonald, 2016, for review). Yet, documenting an association between variables does not imply that the related variables have predictive out-of-sample power (Shmueli, 2010). Indeed, when we add MD&A tone to quantitative forecasting models, there is no improvement in forecasting accuracy. Therefore, whether and to what extent MD&A information helps predict future earnings in an out-of-sample setting is an open empirical question. Tone, an aggregate measure, may be too coarse to capture the predictive content in MD&A. In this article, we develop a methodology that allows us to incorporate MD&A text into forecasts of future earnings with less information loss. Our methodology does not introduce any human biases that come from ex ante word classifications, and we allow the words themselves to determine their informativeness.
Our sample of MD&A disclosures contains all 10-K reports from the Security and Exchange Commission’s (SEC) EDGAR website with a match in COMPUSTAT. We extract the MD&A section from the annual filing and use the bag-of-words approach to incorporate MD&A detailed text into firm-level forecasts of future (scaled) earnings. Specifically, using Loughran and McDonald’s (2011) dictionary of words relevant in financial settings (hereafter, LMD), we count the occurrences of each of these words in the MD&A. Hence, each observation is a vector of unique LMD word counts in that MD&A document. Because there are many words in LMD, we need a technique that accommodates a large number of predictor variables and possible multicollinearity among them. We employ ridge regression to estimate the parameters of our models over a 4-year rolling window.
Our next step is to calculate 1-year-ahead earnings forecasts using information in financial statements alone and information in both financial statements and the MD&A. If the information provided in MD&A statements is relevant for forecasting future earnings, models that incorporate MD&A disclosures will be more accurate. While we might not expect MD&A disclosures to be detrimental to forecasting, it is not obvious that MD&A disclosures will improve the accuracy of earnings forecasts because of their low signal-to-noise ratio. The SEC has repeatedly raised concerns regarding generic “boilerplate” statements and immaterial and/or redundant details in the section. Consistent with these concerns, Pava and Epstein (1993) find that MD&A sections of 25 randomly selected companies mostly describe past performance (i.e., redundant information). Relatedly, Brown and Tucker (2011) find evidence of increasing MD&A length and decreasing MD&A modifications suggesting that MD&A contains a large portion of redundant information.
Comparing forecasts of models with and without the textual content in MD&A, we show that MD&A statements improve the accuracy of earnings forecasts above and beyond traditional financial variables. The improvement is economically meaningful: Incorporating MD&A disclosures enhances the accuracy of earnings forecasts by about 5%, on average, over models that use only financial information. 2 Furthermore, this result is robust across different financial benchmarks, different estimation methods, and different scaling of variables.
Many academic papers proxy for investors’ expectations of earnings with analysts’ consensus forecasts. The presumption is that analysts use multiple sources of information (public and private) and their expertise to generate superior forecasts. Indeed, considerable empirical evidence exists showing that analysts’ forecasts are better than quantitative time-series models (see Bradshaw, Drake, Myers, & Myers, 2012, for review). This leaves us with two questions: Do MD&A disclosures help improve the accuracy of earnings forecasts for firms that are not followed by analysts? For followed firms, do analysts genuinely have superior skills or better information, which they use to forecast future earnings? Consistent with our earlier findings, we find that MD&A text improves forecast accuracy for firms with no analyst following, highlighting the value of incorporating text when alternative proxies for earnings expectations are not available. Moreover, we find that analysts are better than text-enhanced models only for large firms and for firms with high analyst following. In contrast, text-enhanced models are as good as or better than analysts for smaller firms and firms with low analyst following. This suggests that in settings where analysts and time-series models have similar information sets, analysts appear to have no accuracy advantage (i.e., no superior processing or inference skills).
Perhaps it is unsurprising that utilizing the content of MD&A in forecasting models improves their forecasting accuracy. After all, who is better equipped than top management to interpret current performance and anticipate future performance. 3 Therefore, a natural question is whether there are specific firm characteristics that lead to more (less) useful MD&A disclosures. To answer this question, we examine differences in the predictive power of MD&A content in earnings forecasting. There are both benefits and costs to managers of informative MD&A disclosure. On one hand, consistent with the SEC requirements, managers may use disclosures to reduce information asymmetry and to signal firm operating performance and prospects to investors (see Healy & Palepu, 2001, for review). To encourage more disclosure, in 1979-1980, the SEC introduced safe harbor provisions that protect managers if the forward-looking information they provide turns out to be inaccurate. On the other hand, a large body of literature argues that managers’ decision to disclose is largely influenced by firm’s political and litigation costs (Dye, 2001; Francis, Philbrick, & Schipper, 1994; Skinner, 1994, 1997; Verrecchia, 2001). For example, disclosures about firm’s operations, performance, strategies, and so on, may reduce its competitive position, or, alternatively, increase firm’s litigation exposure if disclosed information turns out to be inaccurate. As a result, managers may have incentives to withhold information relevant to their businesses. Therefore, in our cross-sectional tests, we consider measures that proxy for both the costs and benefits of disclosure.
The SEC’s MD&A guidance calls for firms to provide information about the quality and potential variability of earnings and cash flows so that investors can determine whether the past is indicative of the future, and suggests that firms with more complexity have a greater need to focus their MD&A on material information. Using a cross-sectional model with the accuracy improvement from MD&A text as the dependent variable, we test whether managers respond to a greater demand for disclosure when firm performance is less stable, operations are more complex, earnings are more value relevant, and accruals are larger. We find that firms with large changes in future performance and with poor future performance provide more informative MD&A disclosures (i.e., the addition of text improves accuracy more). Furthermore, we find no evidence that complex firms and firms with more volatile past earnings have more informative MD&As. Although these firms may have less predictable earnings and would be good candidates for helpful disclosure, they may also be harder for managers themselves to describe or forecast. When earnings are more value relevant, we find that MD&A disclosures have greater predictive value, helping to avoid large market movements from earnings surprises. When accruals are large, current and future earnings are more likely to be negatively related due to their reversal. We find that MD&A disclosures are more informative for firms with larger positive accruals, while we find the opposite for larger negative accruals. Finally, we examine various proxies of disclosure cost and investor scrutiny and find that MD&A sections are more informative for larger firms, older firms, and firms with higher financial distress risk. In addition, firms facing greater litigation risk and market risk provide less informative MD&As.
Our article contributes to the literature in several ways. First, we develop a methodology to incorporate narrative disclosures in the MD&A section to increase the accuracy of future earnings forecasts in an out-of-sample setting. Although we do not rely on a wider array of firm disclosures in this article, our methodology can be used to combine a broader range of textual and numerical information. Second, we provide information about the usefulness (for predicting future earnings) of MD&A disclosures for distressed firms, firms with high litigation costs and market risk, firms with large changes in future performance and increased uncertainty, firms with high investor scrutiny, and firms with high accruals and high value relevance of earnings. While the SEC requires disclosures of all registrants, the disclosure rules may not lead to equally meaningful disclosures across firms. We provide evidence on whether and when MD&A disclosures indeed provide context for operating performance and prospects. Finally, we show that incorporating MD&A text into earnings forecast helps level the playing field between analysts and time-series earnings forecasting models.
The rest of the article is organized as follows: “Models of Earnings Forecasts” section outlines earnings prediction models and develops our hypotheses. “Data and Model Estimation” section discusses data, variable measurements, and methodology. “Results” and “Conclusion” sections report the results and conclusions of our analysis.
Models of Earnings Forecasts
Text as a Predictor of Future Earnings
To understand whether and how textual disclosures improve investors’ expectations of future earnings, we begin with two traditional time-series models of annual earnings. These models rely only on quantitative variables from financial statements. Given the considerable evidence to suggest that simple models of annual earnings perform at least as well as more complex models (e.g., Albrecht, Lookabill, & McKeown, 1977; Gerakos & Gramacy, 2013), we use the following:
The letter Q denotes a model in which all predictor variables are traditional, quantitative measures. Model 1Q uses historical return on equity (ROE; net income before extraordinary items divided by average book value of owners’ equity) to predict 1-year-ahead ROE. Model 2Q disaggregates income into its operating and nonoperating components scaled by the average book value of equity, thus allowing the coefficients to vary based on the relative informativeness of each.
Advances in computational technologies and access to electronic filings (e.g., the SEC’s EDGAR repository) have made it possible to incorporate textual information into forecasting models for a large sample of firms, yet there is no universal rule on how to do so. We incorporate MD&A text in two ways. First, we use the positive and negative word lists from Loughran and McDonald’s (2011) dictionary (LMD). Similar to Loughran and McDonald (2011), we take the total count of all positive LMD words less the total count of all negative LMD words, scaled by the total number of words in the MD&A section. Our forecast models that incorporate tone (as a measure of MD&A content) in addition to quantitative financial variables are given below:
where the letter C denotes that along with the quantitative variables, a variable based on the predetermined categorization of words into positive and negative, Tone, is included. 4
The tone measure captures only one dimension of MD&A content and it does not allow the statistical model to differentiate among individual words. All positive (or negative) word counts are grouped together, suppressing the possibility that some words are more relevant for earnings forecasting than others, may have different strength (e.g., “weakness” vs. “failure”), or may have been frequently used in a context different from its prespecified category. 5
One of the most popular approaches in linguistics is called bag-of-words (see Henry & Leone, 2015; Li, 2010b; Loughran & McDonald, 2016, for review). This method counts the occurrences of words in a document, allowing each document to be represented as a vector of word counts that can be used in statistical models. Our second approach includes the word count matrix directly, allowing the model itself to determine the importance of individual words for earnings prediction. We use the letter T for models with quantitative and detailed text:
Analysts’ Forecasts Versus Time-Series Models
It is common practice to proxy for earnings expectations using analysts’ consensus forecasts. For a large sample of firms, consensus forecasts are readily available and analysts have been shown to be superior to time-series models (Fried & Givoly, 1982; O’Brien, 1988). Our model incorporates both quantitative and qualitative information into forecasts, so a natural question is whether analysts are still better at forecasting than statistical models.
About 37% of the firms in our sample are not followed by analysts. To ensure that our initial findings are not driven by firms which already have forecasts that incorporate other information (namely, analysts’ forecasts), we evaluate whether MD&A content helps improve earnings forecasts for the subsample of firms with no analyst following. Then turning to the sample with analyst following, we examine (a) whether analysts have an advantage over text-enhanced models and (b) for which firms the advantage is most significant. We compute the nonstale consensus forecast as the mean of analysts’ forecasts of earnings per share reported on I/B/E/S 9 months before the fiscal year end, divided by the beginning-of-year average book value of equity per share (Banker & Chen, 2006; Frankel & Lee, 1998) and compare with forecasts from Models 1T and 2T. In this manner, the timing of consensus forecasts roughly coincides with the availability of 10-K reports on the SEC’s website.
Determinants of MD&A Informativeness
Different firms are likely to have different quality MD&A sections (for predicting future earnings). The SEC’s MD&A requirements guide our predictions about determinants of disclosure informativeness, where informativeness is measured as the difference in forecast accuracy between models that include quantitative information only and models that include both quantitative information and MD&A content.
The SEC states that MD&A statements should allow investors to understand the effects of material changes, trends, and uncertainties on firm’s performance and liquidity. As the past is less indicative of the future for firms with large expected changes in profitability, we expect a positive relation between MD&A informativeness and absolute changes in future earnings (AbsFutEarnCh). Moreover, as investors dislike negative surprises more than positive ones (e.g., Matsumoto, 2002; Skinner & Sloan, 2002), we expect firms anticipating negative future performance (FutLoss) to provide more relevant MD&A disclosures.
Large firms face greater scrutiny from the SEC and analysts; firms with high institutional holdings face greater scrutiny from their actively involved shareholders. Therefore, we expect that large firms and firms with large institutional ownership issue more informative MD&A disclosures. In addition, MD&A should provide information about the quality and potential variability of a company’s earnings and cash flows. Related to earnings variability is earnings persistence. Based on Hayn (1995) and Dechow and Ge (2006), respectively, we use two proxies for persistence: whether the firm has experienced a loss (Loss) and the magnitude of their accruals (AbsAccr). Although all large accruals tend to be less persistent, managers may be more likely to explain accruals that are persistent and income increasing (i.e., in their favor) than those that are persistent and income decreasing. To allow for a difference in the slope that depends on the sign of accruals, we include a dummy variable on the sign of the accruals measure.
Firms with more complex operations need more description in their disclosures to enhance existing financial statements. The SEC also expects that all information in MD&A be presented in clear and understandable manner, regardless of firm complexity. 6 Therefore, we predict that firms with more geographic segments or more business segments have more informative disclosures. The “one size fits all” nature of Generally Accepted Accounting Principles (GAAP) standards may diminish the value relevance of some earnings reports relative to others. The more value relevant earnings (i.e., explanatory power of earnings for stock prices), the more important it is for investors to accurately forecast earnings and avoid large surprises. Therefore, we expect a positive association between earnings value relevance (EarnValRel) and MD&A informativeness.
Turning to risk, Altman’s (1968)Z-score (ZScore), initially designed to predict bankruptcy, can also rank the financial health of companies, where high (low) scores indicate good (bad) financial position. The Capital Asset Pricing Model (CAMP) beta (Beta) is a measure of a company’s market risk, where high values of Beta indicate that the company’s stock price is more volatile than the overall market and investors would benefit from a better disclosure. Therefore, we expect to find a negative (positive) relation between the firm’s Z-score (Beta) and MD&A informativeness.
A large body of literature on corporate disclosures argues that managers are reluctant to provide informative disclosures when there are high litigation concerns, political costs, and increased investor/regulatory scrutiny. This literature is reviewed in Healy and Palepu (2001), Verrecchia (2001), and Dye (2001). Despite significant regulatory efforts to improve communication between managers and investors by reducing such managerial concerns (e.g., safe harbor provisions, see Item 303(c) of Regulation S-K [17 CFR 229.303(c)]), disclosure costs are still considered a primary driver of firms’ disclosure decisions. We predict that the higher the litigation risk (LitigRisk), the less likely informative the MD&A.
We summarize the predictions on the relation between firm characteristics and text informativeness below.
Hypothesis 1
Data and Model Estimation
Data Collection
Our quantitative models use earnings, its components, and the book value of equity. We calculate ROE as earnings divided by the average book value of owners’ equity (computed from period t− 1 to period t). Model 2Q requires information on the operating (OPINC) and nonoperating (NOPINC) components of income. We calculate OPINC as the current operating income after depreciation, net of interest expense, special items, and minority interest, and NOPINC as the current nonoperating income, net of income taxes, scaled by the average book value of owners’ equity. These quantitative variables are extracted from COMPUSTAT over the period 1994-2012, for the entire population of firms. For a subsample of firms with analyst following, we use analyst nonstale forecasts of earnings per share reported on I/B/E/S 9 months before the fiscal year end, divided by the beginning-of-year average book value of equity per share. After we remove firms in financial industries (SICs 6000-6999), outliers and suspected data errors (following Banker & Chen, 2006), we have 51,935 firm-year observations with all the required accounting-based quantitative variables.
To combine these observations with textual information in the MD&A section, we download all 10-K reports submitted to the SEC on EDGAR and extract the MD&A section (Items 7 and 7a). Because 10-K reports do not have a standardized structure of text, MD&A sections without clear designations are harder to extract. We successfully extract the MD&A section for approximately 91% of the total 10-K downloads. 7 Following Li (2010a), we delete all the HTML tags, special symbols, stop words, tables, and numbers from all MD&A documents. These cleaned MD&A files are used to build the qualitative dataset, which is merged, on Central Index Key (CIK) and date, with COMPUSTAT and I/B/E/S. Table 1 summarizes our data processing steps. Our final sample consists of 26,487 firm-year observations over the period 1994-2012. 8
Sample Creation.
Note. MD&A = Management Discussion and Analysis. SIC = Standard Industrial Classification
We require at least 250 words to appear in the MD&A section to eliminate those 10-K reports in which the MD&A section is only incorporated by reference (i.e., pointing shareholders to a different filing).
In 1993, we require only book values of owners’ equity. In 2012, we require only earnings and book value of equity of companies in our sample.
To create textual variables for use in statistical models, we employ the bag-of-words approach, which identifies each word and counts the number of times it appears in a document. To restrict attention to financially relevant words, we use a stemmed version of Loughran and McDonald’s (2011) financial sentiment dictionary (LMD) as the primary source for our word counts. 9 LMD contains 3,532 distinct words that are grouped into six categories: positive, negative, uncertain, litigious, modal strong, and modal weak. The stemmed version of LMD contains 1,389 words.
Table 2 provides descriptive statistics of both quantitative and text-based variables for our sample (number of observations ranges from 19,215 to 26,487). The mean (median) of ROE is 0.052 (0.085). The mean (median) of operating income scaled by the average owners’ equity, OPINC, is 0.078 (0.108) and the mean (median) of nonoperating income scaled by the average owners’ equity, NOPINC, is −0.025 (−0.025). The magnitudes are similar to those reported in Banker and Chen (2006). Furthermore, the average (median) logged firm size is 5.686 (5.754) with the standard deviation of 2.08, indicating that both large and small firms are present in our sample. Turning to MD&A, the mean (median) of total words in the MD&A section is 5,074 (4,447) and approximately 10% of MD&A words are LMD words. 10 The Length of the MD&A section is consistent with previous research (e.g., Li, 2010a).
Descriptive Statistics: 1994-2011.
Note. This table shows the descriptive statistics for main variables used in the article. Variable definitions: ROE is net income before extraordinary items divided by the average book value of owners’ equity; OPINC is operating income after depreciation, net of interest expense, special items, and minority interest, divided by the average book value of owners’ equity; NOPINC is nonoperating income, net of income taxes, divided by the average book value of owners’ equity; Size is the natural logarithm of firm’s market value of equity; FirmAge is the natural logarithm of 1 plus the number of years since a company appears in the CRSP monthly file; MTB is the market value of equity divided by the book value of equity; EarnVol is the standard deviation of earnings, calculated using annual operating earnings scaled by equity in the last 5 fiscal years, with a minimum of three observations required; EarnQualRank is the categorical ranking of earnings based on the S&P’s Quality Ranking; ZScore is the Altman’s Z-score; Beta is the company’s CAPM Beta estimated using daily returns for the [−140, 60] period around the fiscal year end date; EarnValRel is the value relevance of earnings estimated at the SIC2 industry level; Loss equals 1 if current earnings are negative (i.e., loss) and 0 otherwise; Accr is earnings minus cash flow from operations divided by the average book value of owners’ equity; NBSeg is the natural logarithm of 1 plus the number of business segments; NGSeg is the natural logarithm of 1 plus the number of geographic segments; LitigRisk is the estimated litigation risk for a company in a given year using coefficients from Kim and Skinner (2012). MD&A Words (LMD Words) is the number of all words in the MD&A section (number of words in the MD&A section that are in the Loughran & McDonald, 2011, dictionary [LMD]); and Perc LMD is the number of LMD words in MD&A divided by the total words in MD&A.
Estimation Method: Ridge Regression
Consider the traditional ordinary least squares (OLS) model with N predictor variables,
where
In the stemmed version of LMD, there are 1,389 unique words. However, it is likely that not of LMD words are relevant for future earnings. To avoid introducing human biases and selecting LMD words we think are relevant for earnings forecasting, we allow data and statistical relations between variables to speak for themselves. It is likely that many of our text variables based on LMD word counts are highly correlated. Therefore, we need a methodology that handles a large number of predictor variables and multicollinearity at the same time; estimating Equation 1 with OLS poses a challenge.
We use ridge regression, a regularized regression method, which effectively addresses the problems of multicollinearity and many predictor variables. Specifically, ridge regression minimizes the residual sum of squares subject to a penalty on the magnitude of the estimated coefficients (Hoerl & Kennard, 1970). Ridge regression achieves its better prediction performance through a bias-variance trade-off. In contrast to OLS, ridge regression minimizes the loss function:
where λ is a positive number that penalizes large weights in ω (often referred to as the regularization parameter). By introducing the regularization parameter λ, ridge regression balances the trade-off between the bias and variance of the estimator. 11 Ridge regression shrinks the coefficient estimates, but it always keeps all the predictors in the model. Our results are similar when we use alternative methods of estimation (see Section Robustness Checks). Figure 1a summarizes our text extraction and model estimation procedures.

Text extraction and estimation procedures.
Table 3 reports the list of words with the highest predictive power for earnings forecasting over the period 1998-2012. Negative words account for more than half of that list, while positive words make up only one quarter of the top 50 (this is consistent with the much larger proportion of negative words in LMD). However, we note that while typically classified as negative, the word “decrease” may be positive if placed in front of “costs.” In untabulated analyses, we find that the coefficients of over 28% of the words have signs opposite from those predicted by their LMD category. For instance, negative words “devalue” and “disapprove” have positive signs whereas positive words “achieve” and “assure” have negative signs. The opposite sign on positive/negative words could be the result of positive (negative) words occurring in more negative (positive) contexts. Alternatively, it could indicate drawbacks of subjective word classifications into categories. This finding is consistent with findings in Jegadeesh and Wu (2013), who conclude that not all positive (negative) words are viewed positively (negatively) by investors. Therefore, we have an early indication that looking at individual words may better capture specific word usage in financial disclosures and may avoid subjectivity in word classifications. 12 Furthermore, Figure 2b and 2c illustrates that individual word usage in MD&A texts is not constant over time, whereas aggregate counts of positive, negative, litigious, and uncertain word change very little over time (Figure 2a). Greater variability across frequencies of individual words suggests potential changes in informativeness of each word over time.
List of Top 50 Words Selected in Forecasting Models.
Note. LMD = Loughran and McDonald Dictionary.

Mean normalized frequencies: Categories and individual words.
Results
Measures of Forecast Accuracy
In our analyses, we consider two types of earnings forecast models: (a) forecast models based on quantitative information and (b) forecast models based on both quantitative and textual information. We use data from year t to forecast earnings in year t+ 1 based on regression coefficients estimated in prior 4 years. We compute forecast errors as the squared difference (i.e., error) between the realized ROE in year t+ 1 and predicted
Our empirical specifications are based on pairs of nested models. For instance, our quantitative Model 1Q (2Q) incorporates
We then calculate the accuracy improvement for an individual firm as the forecast error of quantitative models minus the (adjusted) forecast error of text-enhanced models:
Positive (negative) values of AI in Equation 3 indicate that the text-enhanced model has lower (higher) prediction errors, that is, more (less) accurate forecasts. In our main tests, we calculate the average accuracy improvements across all firms. We use the one-sided t test and the Wilcoxon’s signed-rank test to draw inferences about the significance of our results. In addition to statistical significance, we calculate the economic significance of the forecast differences following Fairfield et al. (1996). In their approach, an improvement in forecast accuracy of 5% or more is considered economically significant and meaningful to investors. Therefore, we compare the fraction of forecasts with improvement in accuracy of 5% or more when the text-enhanced model is used to the fraction of forecasts with improvement in accuracy of less than 5%. We then use the binomial test to compare these two proportions.
Forecast Accuracy of Text-Enhanced Versus Quantitative Models
We estimate the coefficients of our model using a 4-year rolling window making 1998 our first year of out-of-sample forecasts. Figure 1b summarizes our estimation timeline. In Table 4, we provide descriptive statistics of the mean (median) squared prediction errors of our quantitative models and text-enhanced models over the whole sample period. Panel A compares quantitative benchmarks with corresponding models that include tone variable based on positive and negative word counts. The mean (median) squared prediction errors for both quantitative (SPEQ) models and models enhanced with positive and negative (i.e., Tone) word counts (SPEC) are the same: 0.041 (0.007) for Model 1 and 0.040 (0.006) for Model 2. This result suggests that aggregating positive and negative words counts into a single variable and including that variable in the forecasting model does not improve the model’s out-of-sample forecast accuracy. The lack of predictive power of Tone may be surprising given the evidence in prior literature that MD&A tone and future earnings are correlated (Li, 2010). However, out-of-sample predictive accuracy and in-sample explanatory power are different constructs and a correlation does not imply that the variable has predictive relevance.
Mean, Median, and Pairwise Differences in Squared Prediction Errors: Period 1998-2012, All Firms.
Note.
Percentage of observations for which text improves forecast accuracy by 5% or more (first number) and the percentage of observations for which text reduces forecast accuracy by 5% or more (second number).
Proportion of observations with improved accuracy exceeds the proportion of observations with reduced accuracy using the binomial test at the 1% significance level.
***, **, * (+++, ++, +) indicate significance at the 1%, 5%, and 10% levels, respectively, using the one-sided t test (Wilcoxon’s signed-rank test). Number of observations: 26,487.
Panel B of Table 4 reports the mean (median) squared prediction errors for quantitative models (Q) and models enhanced with individual words (T). We find that text enhancements significantly improve the accuracy of forecasts, both statistically and economically. Specifically, the mean (median) squared prediction error is 0.002 (0.001) lower for text-enhanced models than for quantitative models, or an average increase in forecast accuracy of 5% after adding MD&A text Figure 3 plots forecast accuracy improvements after adding text for each year in the sample. Further, we find that Text-enhanced forecasts are at least 5% better than quantitative forecasts 53% of the time and less than 5% better than quantitative forecasts 37% of the time (p < .01). Alternatively, we can interpret the economic significance by taking the absolute difference between actual and predicted ROE and multiplying the difference by the average book value of equity. In this manner, we can calculate the amount that each forecasting model misses on the level of earnings. Text-enhanced models predict earnings by around US$7.6 million better than the quantitative models, where the average earnings in our sample is US$107 million (or a 7% improvement). Overall, our results indicate that MD&A detailed content, rather than the crude measure of tone, is indeed relevant for earnings forecasting. This finding also contributes to the debate in prior literature about the quantity and quality of financial disclosures that matter (Beretta & Bozzolan, 2008).

Percentage differences in accuracy: Q and T models, year-by-year.
Analysts Versus Text-Enhanced Models
To provide further evidence on the usefulness of MD&A text for earnings forecasting, we examine the difference in accuracy of text-enhanced and quantitative models for the subset of firms in our sample with no analyst following (9,801 observations). For these firms, investors do not have the option of using an analyst forecast to guide expectations of future earnings and must generate forecasts on their own. Panel A of Table 5 reports the mean and median of squared prediction errors of our quantitative and text-enhanced models. The mean (median) squared prediction errors of quantitative (SPEQ) and text-enhanced (SPET) models are 0.057 (0.008) and 0.054 (0.007) for Model 1, and 0.054 (0.007) and 0.052 (0.006) for Model 2. In both cases, the mean and median accuracy improvements (AI) are positive and statistically significant (p < .01), indicating that text-enhanced models are more accurate than quantitative models for firms with no analyst following. Overall, these results indicate the benefit of using narrative disclosures to improve expectations for firms where there is no intermediary to provide forecasts.
Accuracy of Text-Enhanced Models and Analysts’ Consensus Forecasts Around 10-K Filing Dates.
Note.
Number of observations: 9,801.
Number of observations: 16,686.
***, **, * (+++, ++, +) indicate significance at the 1%, 5%, and 10% levels, respectively, using the one-sided t test (Wilcoxon’s signed-rank test).
Turning to the subsample of firms for which analyst forecasts are available (16,686 observations), we compare the accuracy of text-enhanced models to the analysts’ consensus forecasts issued around 10-K filing dates. In untabulated tests, we find that on average analysts’ consensus forecasts are better than text-enhanced forecasts. This evidence is consistent with findings in Bradshaw et al. (2012) that analysts are better than time-series models. Next, to try to understand whether this superiority is related to genuine expertise or simply the use of information beyond 10-K disclosures, we examine whether analysts’ superiority persists across large and small firms. The larger the firm, the more frequently it appears in the financial press, social media, and other outlets. Therefore, analysts may utilize contemporaneous disclosures (like news articles) into their predictions that our text-enhanced models do not. 14 Panel B of Table 5 reports the mean (median) squared prediction errors of text-enhanced (SPET) models and analysts’ consensus forecasts (SPEA) for large, medium, and small firms. We find that analysts are better than text-enhanced models for larger firms (0.005, significant at 1%). However, text-enhanced models are better than analysts for smaller firms—The error difference between text-enhanced models and analyst is −0.011, significant at 1%.
As an alternative to firm size, we also use the number of analysts following each firm to partition our sample into high, medium, and low analyst following firms. Firm size and analyst following are highly correlated and ultimately capture the relative amount of information available about a firm (in addition to MD&A). Similar to Panel B, in Panel C of Table 5, we find that analysts’ forecasting superiority persists for firms with high analyst following, whereas text-enhanced models are better than analysts for firms with low analyst following. Overall, these results indicate that analysts’ advantage increases in the amount of information available to them. By restricting our models to MD&A only, we capture only a small fraction of that information.
Cross-Sectional Variation in MD&A Informativeness
Having established that MD&A information helps improve forecasts of future earnings, we next identify different firm characteristics that explain the cross-sectional predictive power of MD&A. In Table 6, we regress our measure of forecast accuracy improvement after adding MD&A text, AI(Q-T), on various firm characteristics related to firms’ incentives to provide or withhold relevant information in MD&A. Positive values of AI(Q-T) indicate that forecast accuracy of text-enhanced models is higher than forecast accuracy of quantitative models. Equation 4 summarizes our regression specification. We include industry (using Fama–French 12 industry classification) and year fixed effects and we cluster standard errors at the firm level. 15
Determinants of MD&A Informativeness: 1998-2012.
Note. This table shows the estimated coefficients from a regression of MD&A informativeness on firm characteristics. The dependent variable is the accuracy improvement after adding MD&A text into a prediction model (AI(Q-T)), calculated as the difference in squared prediction errors of quantitative and text-enhanced models, multiplied by 100. Independent variables are as follows: AbsFutEarnCh is the absolute change in future ROE, calculated as the absolute value of the difference between return on equity in years t+ 1 and t; FutLoss equals 1 if future earnings are negative and 0 otherwise; Size is the natural logarithm of firm’s market value of equity; FirmAge is the natural logarithm of 1 plus the number of years since a company appears in the CRSP monthly file; MTB is the market value of equity divided by the book value of equity; EarnVol is the standard deviation of earnings, calculated using annual operating earnings scaled by equity in the last 5 fiscal years, with a minimum of three observations required; EarnQualRank is the categorical ranking of earnings based on the S&P’s Quality Ranking; ZScore is the Altman’s Z-score; Beta is the company’s CAPM Beta estimated using daily returns for the [−140, 60] period around the fiscal year end date; EarnValRel is the value relevance of earnings estimated at the SIC2 industry level; Loss equals 1 if current earnings are negative (i.e., loss) and 0 otherwise; AbsAccr is the absolute value of earnings minus cash flow from operations scaled by the average book value of owners’ equity; PosAccr equals to AbsAccr if accruals are positive and 0 otherwise; NBSeg is the natural logarithm of 1 plus the number of business segments; NGSeg is the natural logarithm of 1 plus the number of geographic segments; LitigRisk is the estimated litigation risk for a company in a given year using coefficients from Kim and Skinner (2012). Each of the continuous variables is winsorized at 1% and 99% to mitigate outliers. Year fixed effects, Fama–French 12 industry fixed effects, and the constant are included in the regressions but are not reported; t statistics based on clustering by firm are reported in parentheses. MD&A = Management Discussion and Analysis; AI = accuracy improvement. SIC2 = 2-digit Standard Industrial Classification Code.
***, **, * indicate significance at the 1%, 5%, and 10% levels, respectively.
We provide coefficient estimates in Table 6. First, we consider variables related to absolute changes in future earnings (AbsFutEarnCh) and future losses (FutLoss). Controlling for different firm characteristics, industry and time fixed effects, we find that the coefficients on AbsFutEarnCh and FutLoss are positive and highly significant. These results are consistent with our predictions that MD&A disclosures are more useful for enhancing expectations when the market is most likely to need it, specifically, in helping to anticipate future losses and large changes in performance.
In Hypothesis 1b, we posit that MD&A disclosures are more informative for firms with high investor scrutiny, uncertainty, bankruptcy risk, and complexity. In addition, we expect that firms with greater value relevance of earnings and lower earnings persistence provide more context for reported earnings. We proxy for investor scrutiny using firm size (Size), measured as the natural log of market capitalization and find that Size is strongly positively related to MD&A informativeness. In untabulated tests, we also use institutional ownership (i.e., the number/proportion of institutional investors) as an alternative measure of investor scrutiny. Firm size and institutional ownership are highly correlated (around 90%) and our inferences remain the same—Firms with more institutional ownership provide more informative MD&A disclosures. Overall, our evidence is consistent with firms using MD&A to respond to investors’ greater demand for information.
We find mixed results on the usefulness of MD&A for high uncertainty firms. Our measures of uncertainty are the market-to-book ratio (MTB), earnings volatility (EarnVol), firm age (FirmAge), and earnings quality rank (EarnQualRank)—an Standard & Poor’s (S&P) measure based on analysts’ assessment of earnings stability. In Table 6, EarnVol is generally not associated with MD&A informativeness (with exception of marginal significance in column (1)), whereas MTB is marginally positively associated with MD&A informativeness. Furthermore, we find that FirmAge and EarnQualRank are positively associated with MD&A informativeness, suggesting that older firms and firms with more stable earnings provide more useful narrative disclosures in MD&A. These results suggest that it may be difficult for managers to provide useful information when they themselves do not have it or have incentives to withhold it. Managers of younger firms or firms with unstable earnings operate in very volatile and risky environments and may be unable or unwilling to provide relevant information in MD&A due to the difficulty of firm’s environment or higher disclosure concerns.
Turning to our predictions about MD&A informativeness and risk and complexity, we find that firms with greater bankruptcy risk (as measured by the Altman’s Z-score) provide more useful information in their MD&As, whereas firms that are more sensitive to stock market fluctuations (i.e., high Beta firms) provide less informative MD&A disclosures. This evidence is consistent with managers providing useful information about firm-specific distress risk, yet being less able to predict and provide useful information resulting from changes in the overall market. We do not find evidence that complex firms (as measured by the number of business and geographic segments, NBSeg and NGSeg) have more informative disclosures. One potential explanation is that despite numerous SEC’s efforts to improve disclosure, complex firms’ disclosures are too intricate for simple word counts to capture their informativeness.
Measuring a company’s earnings is a complex process that builds on many estimates and judgments, due to the accrual basis of accounting. We posit that companies with high value relevance of earnings (i.e., earnings are more closely related to valuation) have more informative MD&A disclosures. Following Banker, Huang, and Natarajan (2009), we calculate an industry-specific measure of value relevance of earnings by estimating the following model at the SIC two-digit level 16 :
where Pt is firm price per share at the end of the third month after fiscal year end t and EPSt is firm earnings per share during year t. R2s of this model at the industry level are used to proxy for the value relevance of earnings. Table 6 shows a positive and significant coefficient on EarnValRel, which suggests that firms operating in industries where earnings are more value relevant have more informative MD&A sections for predicting future earnings.
Next, we look at the association between earnings persistence (measured by negative current performance and accruals) and MD&A informativeness. We find a positive association between current losses (Loss) and MD&A predictive value in column (3) of Table 6 (significant at 10%). However, Loss is not statistically significant in the other two columns of Table 6. Therefore, we interpret this result as only a weak evidence of loss firms having more informative MD&A disclosures. Earnings with a high accrual component are less persistent than earnings with a high cash flow component, due to the estimates involved in the accrual component. We use absolute accruals to capture the amount of accruals in earnings (AbsAccr). In addition, we separately control for positive accruals (PosAccr) to examine if managers provide more informative MD&As to explain large positive accruals which might otherwise be discounted. In Table 6, we find a consistent negative association between AbsAccr and MD&A informativeness, while the association between PosAccr and MD&A informativeness is consistently positive and significant. The magnitude of the coefficient estimate on PosAccr is much higher than that on AbsAccr (0.40 vs. −0.15). Overall, these results suggest that managers provide more (less) informative MD&As when accruals are positive (negative) and large. This is consistent with managers seeking to reassure investors about income increasing accruals, but providing less informative MD&As when the accruals are income decreasing.
Our final set of results on Hypothesis 1c relates to firms’ litigation costs and their role in corporate disclosures. Litigation cost has been a widely debated topic in the literature (e.g., Francis et al., 1994; Rogers, Buskirk, & Zechman, 2011; Skinner, 1997), and the consensus is that managers may avoid disclosing relevant information to avoid potential lawsuits pertaining to such disclosure. Alternatively, disclosing relevant information may mitigate future lawsuits. It may prove easier to argue that investors misunderstood or misinterpreted a narrative disclosure than to defend nondisclosure. We use Kim and Skinner’s (2012) model, which combines firm size, sales growth, stock return, return skewness, return volatility, share turnover, and indicator for a litigious industry to measure litigation risk (LitigRisk). Higher values of LitigRisk indicate greater litigation exposure. By construction, LitigRisk incorporates a significant portion of a firm’s market risk. Indeed, the correlation between LitigRisk and Beta is .38, significant at the 1% level. Therefore, in our tests, we include only Beta and only LitigRisk in columns (1) and (2), respectively, and include both variables in the same model in columns (3) and (4). Table 6 shows a significant and negative coefficient estimate on LitigRisk in column (2), where Beta is omitted, but when it is included, the coefficient estimate on LitigRisk is not significant. Overall, our evidence here is consistent with managers avoiding informative MD&A disclosure due to high litigation concerns and market risk.
Taken together, our results on the determinants of MD&A informativeness paint a picture of managers using narrative disclosures to improve their firms’ information environments when it is advantageous (and possible) to do so. Disclosure costs, such as those relating to firms’ litigation risk, seem to limit informative disclosures.
Robustness Checks
We conduct several additional tests to bolster support for our empirical methodology. First, we use two other popular time-series models for forecasting earnings. The first model splits earnings into its cash flows and accruals (following Sloan, 1996) and the second incorporates cost stickiness with sales declines into future earnings forecasts (following Banker & Chen, 2006). Furthermore, we estimate our models with all quantitative variables scaled by the average total assets rather than book equity. Our inferences remain the same.
Second, we estimate our text-enhanced models using only forward-looking sentences in the MD&A section. On one hand, restricting our analysis to forward-looking sentences may eliminate redundant information. On the other hand, a detailed description of current performance may be helpful in determining how current performance translates into future performance. For example, critical accounting policies are typically summarized in non-forward-looking sentences; however, the SEC strongly believes that their inclusion in the MD&A section helps investors to understand relevant accounting issues that management faces. We find that models enhanced with forward-looking information have greater forecasting accuracy than quantitative models, but lower forecasting accuracy than models enhanced by the entire MD&A section.
Finally, we estimate all models using Lasso and Kernel’s Ridge Regression methods. We also apply a two-stage technique, where we first select a subset of relevant words using a stepwise regression and then estimate all text-enhanced models with these preselected words. We also account for basic negation in text and use different scaling of individual word counts. All our results are similar to those reported in the article.
Conclusion
In this study, we develop techniques that allow us to combine MD&A content with financial variables to come up with explicit firm-level forecasts of future earnings. Specifically, we combine traditional measures of operating performance with MD&A text, allowing us to assess the incremental information value of MD&A. We analyze accuracy improvements out-of-sample and find that forecasting models based on financial and textual information are more accurate than models using financial variables alone.
The main motivation for using analysts’ forecasts as a proxy for earnings expectations is analysts’ superiority over time-series models and the availability of consensus forecasts at relatively low cost. However, many firms do not have analyst following, and for those firms, we find that text-enhanced models are more accurate than quantitative models. Furthermore, for firms with analyst coverage, we find that analysts are only superior to time-series models for larger and heavily followed firms. However, for smaller and less followed firms, text-enhanced models are as good as or better than analysts.
Finally, we look at the determinants of MD&A informativeness. Using different proxies for firm’s changes in future performances, uncertainty, risk, investor scrutiny, and litigation costs, we find that MD&A generally helps improve firms’ information environments. Specifically, we find that firms with large changes in future performance, negative future performance, high investor scrutiny, high financial distress risk, high positive accruals, and high value relevance of earnings have more informative MD&A sections In contrast, younger firms, firms subject to greater market risk, and firms facing greater litigation exposure have less informative MD&As.
In our analyses, we limit our attention to parsimonious earnings forecasting models that only include data from financial statements and MD&A text. Although we do not rely on a wider array of firm disclosures in the article, our methodology could be extended to incorporate a broader range of financial and textual information to answer questions pertaining to the relevance of various information to the market.
Footnotes
Acknowledgements
We thank Michael Alles, Roman Chychyla, Valentin Dimitrov, Jeffrey Hales, Alex Kogan, Andy Leone, Miguel Minutti-Meza, Sundaresh Ramnath, Glenn Shafer, Michael Smith, Miklos Vasarhelyi, and Vladimir Vovk; seminar participants at the University of Miami, Rutgers University, Claremont-McKenna College, Stevens Institute of Technology, and Baruch College; and participants of 2012 AAA Regional Meeting, 2012/2013 AAA Annual Meeting, 2013 AAA SET Section Midyear Meeting, and 2013 Text Analytics in Accounting Conference for their helpful comments and suggestions.
Authors’ Note
The data used in this study are available from the sources indicated in the text. All errors are the authors’ alone.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
