Abstract
This study describes the development and validation of pedestrian intersection crossing volume models for the seven-county Milwaukee metropolitan region. The set of three models, among the first developed at a multi-county scale, can be used to estimate the total number of pedestrian crossings per year at four-leg intersections along state highways and other major thoroughfares. Outputs are appropriate for annual volumes ranging from 1,000 to 650,000. We used negative binomial regression to relate annual pedestrian volumes at 260 intersections to roadway and surrounding neighborhood socioeconomic and land-use variables. The three models include seven variables that have significant positive associations with annual pedestrian volume: population density within 400 m of the intersection; employment density within 400 m; number of bus stops within 100 m; number of retail businesses within 100 m; number of restaurant and bar businesses within 100 m; presence of a school within 400 m; and proportion of households without a motor vehicle within 400 m. Results suggest that square root or cube root transformations of continuous explanatory variables could potentially improve model fit. The models have fair accuracy, with each of the three model formulations predicting 60% or more of validation intersection counts to within half or double the observed value. Future research could address overprediction by creating new variables to better represent the number of lanes on each intersection leg and low socioeconomic status of adjacent neighborhoods.
Transportation agencies have used statistical modeling methods to estimate motor vehicle volumes along roadways for more than half a century ( 1 ), but pedestrian volume models have only received significant attention in the last decade ( 2 ). Estimating pedestrian volumes is important for a variety of reasons, including pedestrian safety analysis (represent exposure in crash risk assessments), project prioritization, facility design, and development impact assessment ( 3 , 4 ). Better pedestrian count data and more accurate pedestrian volume predictions are important for pedestrian planning in all communities, including a growing number with goals to increase pedestrian activity and reduce pedestrian injuries as a part of broader sustainability, health, equity, and safety initiatives ( 5 , 6 ).
Pedestrian travel occurs in all types of place, including urban, suburban, small town, and rural communities. Yet, there is a wide range of pedestrian activity levels across these contexts, with urban areas generally having the highest pedestrian volumes. Even within cities, pedestrian volumes can vary greatly. For example, the number of pedestrian intersection crossings per day in San Francisco ranges from fewer than 100 to more than 100,000 ( 7 ). Like motor vehicle traffic volumes, these varying levels of pedestrian activity are important to quantify and forecast to create transportation systems that achieve community goals.
The practice of pedestrian volume modeling has improved over the last decade, and there are now more than a dozen models described in peer-reviewed publications. Yet, most models have been developed within a single city or jurisdictional boundary and have undergone little validation testing. This study seeks to address the following questions: ( 1 ) Where does it make sense to estimate pedestrian intersection crossing volumes in a multi-county region that has a wide range of development patterns from very urban to very rural? ( 2 ) What variables have significant associations with pedestrian volumes across a large region? and ( 3 ) How accurate are these pedestrian volume models? We answer these questions by developing pedestrian volume models for the seven-county Milwaukee region of Southeastern Wisconsin.
Literature Review
We used a direct demand modeling approach to estimate pedestrian volumes in the Milwaukee region. These types of model are constructed by associating characteristics of the built and social environment in the vicinity of count locations with pedestrian volumes. They use these built and social environment characteristics to estimate pedestrian volumes at specific locations throughout a study area. Collecting pedestrian counts directly and using them to create predictive models overcomes key limitations of travel-survey-based approaches, including the ability to quantify pedestrian activity at specific locations (i.e., intersections, street segments) and fully capture secondary pedestrian movements (e.g., walking to and from bus stops or parked cars) ( 8 ). Previous direct demand pedestrian volume models have been reviewed in detail by Munira and Sener ( 4 ). Other summaries have been included in broader pedestrian and bicycle demand modeling guidebooks ( 2 , 9 , 10 ).
Some of the earliest direct demand pedestrian volume models were developed by Pushkarev and Zupan ( 11 ) and Benham and Patel ( 12 ). However, the majority of pedestrian demand models have been created since 2010 as more agencies have realized pedestrian travel is central to achieving community safety, health, livability, economic, environment, and equity goals. Pedestrian demand models are most common at the local level ( 7 , 13 – 18 ), but they could also be applied across metropolitan regions and states ( 10 ). To our knowledge, California currently has the only statewide pedestrian volume model ( 19 ).
Direct demand pedestrian volume models have been developed using ordinary least squares regression and loglinear regression, but Poisson or negative binomial models may be most appropriate for these types of count data ( 4 , 16 ). Dependent variables in these models are typically pedestrian counts representing different time periods from a few hours during specific days of the week to annual volumes ( 18 ).
Common explanatory variables include measures of population or housing density, employment density, commercial retail density, and transit stop proximity and service frequency ( 4 , 7 ). Explanatory variables are often collected at several different distances from the count location (i.e., different buffer widths) to identify the most influential scale of a particular characteristic on pedestrian activity ( 3 , 4 , 15 , 17 – 20 ). Many models include these common variables, but the magnitude and direction of influence of these variables in specific models can be different depending on the study area ( 4 ).
Pedestrian volume model accuracy is often assessed using internal validation methods (e.g., analyze the distribution of residuals and overall model fit). Some direct demand models have used external validation (e.g., compare model-predicted volumes with volumes collected at locations that were not used to build the model) ( 3 , 7 , 18 ). In general, validation analyses have found predicted-levels of pedestrian activity to be intuitive at a community-wide scale but somewhat inaccurate at specific locations.
Direct demand pedestrian models have several benefits. They are relatively simple to develop and apply. Most are based on data that are already available to a transportation agency and can be estimated using common statistical packages. Many provide statistical evidence of theoretical relationships between pedestrian activity levels and land-use, transportation infrastructure, and neighborhood socioeconomic characteristics. Disadvantages of these models include a lack of accuracy when applied at some specific locations, difficulty accounting for special pedestrian generators (e.g., tourist attractions, stadiums, and other unique land uses), and challenges transferring results to other study areas where the model was not developed.
Table 1 summarizes a sample of previous direct demand models that predict pedestrian intersection crossing volumes. This table extends the literature review tables of Schneider et al. ( 7 ) and Munira and Sener ( 4 ), providing updated examples of common pedestrian demand model variables.
Examples of Direct Demand Pedestrian Volume Models.
Note: h = hour; mi = miles; na = not applicable.
Our study builds on previous pedestrian modeling research in several ways. First, by using counts from seven counties, we create one of the first pedestrian volume models to cover such a large area with a wide range of urban to rural development patterns. Second, we test several roadway variables (intersection signalization, traffic volume, multi-lane roadways) and specific land use variables (retail, restaurant, and bar businesses) that have been used in few other pedestrian demand models. Third, we apply a separate set of counts from outside of the modeling dataset to conduct external validation of the model estimates.
Method
We used pedestrian counts from 260 four-leg intersections along state highways and major thoroughfares in the seven-county Milwaukee region to develop our direct demand models and 45 separate intersections to test their accuracy. This section describes our data and statistical methods.
Study Area
We conducted our study in the seven-county Milwaukee region (population 2 million), which is located in the southeastern corner of Wisconsin. The region contains urban, suburban, exurban, and rural communities that are represented by 155 different local governments ( 24 ). The City of Milwaukee, located on Lake Michigan in the east-central part of the region, is the largest city (population 600,000), but other cities include Kenosha (100,000), Racine (80,000), and Waukesha (70,000). Population densities range from more than 10,000 people per square mile in the urban core of Milwaukee to fewer than 100 people per square mile in the rural parts of Kenosha, Racine, Walworth, and Washington Counties.
Count Data
We gathered pedestrian count data from two primary sources. First, we compiled counts collected by consultants in the field for the Wisconsin Department of Transportation (DOT) Southeast Region office between 2013 and 2018. Each of the Wisconsin DOT counts was recorded in one of 1,252 individual spreadsheets, though some spreadsheets covered the same intersection at different times during these six years. In addition to motor vehicle and bicycle counts, each spreadsheet included the number of times pedestrians crossed each leg of an intersection in 15-min increments during a given study period. The most common type of study period covered 13 h from 6 a.m. to 7 p.m. There were a variety of other study period durations, and 99% of the 1,252 counts were at least four hours long. These counts were geocoded based on the intersecting streets recorded in the spreadsheet. Since the Wisconsin DOT counts generally excluded areas of central Milwaukee County, we supplemented these counts with 38 counts collected using a similar method along major thoroughfares in the City of Milwaukee ( 25 ).
Not all of the 1,290 counts were suitable for analysis. We removed counts without geocoded locations, not on a major roadway (e.g., intersections of two local roadways), at three-leg and five-leg intersections, and at freeway ramps or minor driveways (e.g., driveways to single-family homes). We also removed counts taken on days with rain or snow (indicated by field data collectors) and counts taken between November and March (more variability from day to day during winter months). Finally, we removed counts with zero pedestrians since these were either erroneous or in locations where pedestrian volumes are too low to predict reliably in a statistical model.
After applying these criteria, 520 counts remained. To identify unique locations with counts, we grouped all counts that were located within 50 m of another count. This produced 348 intersections that had at least one count. Of the 348 intersections, 223 had one count, 97 had two counts, 18 had three counts, nine had four counts, and one had six counts.
We initially selected 50 intersections that could be used for validation testing. Seventeen of these intersections were chosen because they were located within 200 m of an intersection that was in the modeling dataset. Designating these 17 intersections for validation helped us avoid including intersections that were very close to each other (with similar surrounding environment explanatory variables) in the model estimation process. We selected the additional 33 validation intersections randomly.
Our preliminary model analysis database included 298 intersections, and our preliminary validation database included 50 intersections. After removing outliers through an initial modeling step (described later), our final model analysis database included 260 intersections and validation database included 45 intersections. Approximately 90% of these intersections were on the state highway system, and the other 10% were on other major roadways.
Annual Pedestrian Volume Estimates
Our counts were collected at different times of day, week, and year, so we applied expansion factors to estimate comparable annual volumes at each intersection. Annualization is a common traffic engineering technique. It is important because pedestrian activity peaks at different times in different types of location and is particularly seasonal in Wisconsin. We used three factors to expand the counts taken at specific times on specific days to annual pedestrian volume estimates: hour to weekday, weekday to week, and week to year. The factoring process that we followed, including tables of specific hourly, daily, and weekly factors is described in detail in the City of Milwaukee Pedestrian Plan ( 25 ).
Annual pedestrian volume ranges for the final 260 model intersections and 45 validation intersections are shown in Table 2. Figure 1 shows the locations of the model and validation intersections and their annual pedestrian volume estimates.
Annual Pedestrian Volume Estimates at Final Study Intersections

Pedestrian volume model study intersections across the seven-county Milwaukee region.
Explanatory Variables
Table 3 describes our explanatory variables, and Table 4 provides summary statistics for our model and validation databases. We tested more than 30 theoretically important variables to identify the best predictors of annual pedestrian crossing volumes at the study intersections. For ease of application and transferability, we gathered most of our explanatory variables from publicly available data sources, including the American Community Survey and local, regional, and state transportation agency data. These variables were collected between 2014 and 2018, which is also when most of the pedestrian counts were collected.
Definitions of Variables Used in Modeling Process
The jobs data used for this variable are the number of primary jobs per census block. This excludes the secondary employment locations of people who hold multiple jobs.
Some bus stops serve multiple bus routes. However, no differentiation is made by the number of routes served or bus frequency. We did not include proximity to train stations as a variable because none of our count locations were within 400 m of a train station.
We removed retail and bar and restaurant records that appeared to be business locations that were not storefronts (or locations that would attract foot traffic). Some may have been mailing addresses. Specifically, we removed all records where REMOVE > 0 or where LOC_NAME = “Postal” or “PostalExt” from the original data source.
Parks include neighborhood parks, regional parks, green easements, state forests, wildlife areas, and conservation areas. Parks do not include recreation clubs/resorts, wildlife production areas, schools, university campuses, golf courses, tourist information centers, waysides, zoos, stadium grounds, or fairgrounds.
Schools include elementary, middle, and high schools.
College and university campuses only include institutions with enrollments of 1,500 or more students.
National Center for Education Statistics, Common Core of Data (https://nces.ed.gov/ccd/) data (public schools) and Private School Survey (https://nces.ed.gov/surveys/pss/) data (private schools) were downloaded from the US Department of Homeland Security, Homeland Infrastructure Foundation-Level Data (https://hifld-geoplatform.opendata.arcgis.com/search?groupIds=f16c582f00184cb094affff556fe57ee).
National Center for Education Statistics, Integrated Post-Secondary Education System (http://nces.ed.gov/ipeds/) data were downloaded from the US Department of Homeland Security, Homeland Infrastructure Foundation-Level Data (https://hifld-geoplatform.opendata.arcgis.com/search?groupIds=f16c582f00184cb094affff556fe57ee).
To simplify data collection, AADT values for intersecting roadway legs were grouped into the following categories: 0–2499 AADT = 1250 AADT; 2500–4999 AADT = 3750 AADT; 5000–9999 AADT = 7500 AADT; 10000–14999 AADT = 12500 AADT; 15000–19999 AADT = 17500 AADT; 20000–24999 AADT = 22500 AADT; 25000–29999 AADT = 27500 AADT; 30000–39999 AADT = 35000 AADT; 40000–49999 AADT = 45000 AADT; 50000–59999 AADT = 55000 AADT.
Descriptive Statistics for Final Model and Validation Datasets
Note: SD = standard deviation; Min. = minimum; Max. = maximum.
We measured the variables at different buffer distances from our intersections. All variables were measured at a 400 m distance. We also tested a 100 m radius to represent proximity to specific destinations, such as schools, parks, bus stops, and commercial properties, since we assumed that walking activity could be high on the blocks adjacent to key pedestrian attractors but taper off rapidly beyond this distance. We explored an 800 m radius to test for a broad-area influence of population and employment density variables. Note that we used straight-line buffers for the simplicity of creating and applying the variables in the model.
We transformed some explanatory variables using the square root function (e.g., square root of the population density within 400 m of the intersection) and the cube root function. We did this to test whether or not the marginal impact of a particular surrounding land use characteristic diminishes at higher values (the cube root function reduces the influence of large variable values even more than the square root function). For example, adding 100 more people around an intersection that currently has 500 people living within 400 m may have a large impact on its pedestrian volume, but adding 100 more people around an intersection that currently has 3,000 people living within 400 m may have a smaller impact on its pedestrian volume.
The process we used to select the validation intersections produced a validation dataset that was somewhat more urban (e.g., higher mean population and employment densities; higher percentage of renters) than the model dataset. Still, the mean values of many variables were similar for both datasets, meaning that the validation intersections also provided a good representation of the region as a whole.
Regression Model Structure
Pedestrian volumes are count data, so we considered Poisson and negative binomial models to represent the relationship between total annual pedestrian crossings and our explanatory variables. The variance of annual pedestrian volume estimates across all intersections was much larger than the mean, so we used a negative binomial model. This type of structure has been used in previous pedestrian volume models (15, 16, 20).
The following equation shows the model structure:
where
PedVolumei = estimated annual pedestrian crossings at intersection i,
Xij = quantitative measure of each explanatory variable j associated with intersection i,
βj = model coefficient for explanatory variable j to be determined by negative binomial regression, and
β0 = constant to be determined by negative binomial regression.
Initial Regression Modeling Process
We developed a series of models with different combinations of explanatory variables using our initial dataset (n = 298). For each series of models, we used a stepwise process to identify a subset of statistically significant explanatory variables. In general, this stepwise process first tested a model with many explanatory variables. Then, the least significant variable was removed, and the remaining variables were tested in a second model. This process continued until the model only included variables that were statistically significant at the 90% confidence level (p < 0.10). This represented the final result of one series of models. We repeated this process more than 20 times, starting each series with a slightly different set of explanatory variables. All explanatory variables were tested in multiple models during the process. Some of the variables in the modeling dataset were highly correlated (|r| > 0.7), so we generally avoided including these pairs of variables in the same model. We also avoided including the characteristic measured at different buffer distances in the same model.
After completing this exploratory process, we created a set of three preliminary models using seven variables that consistently showed statistically significant associations (p < 0.05) with annual pedestrian crossing estimates. Our base model (Model A) included direct measurement values for each variable. We also explored how transformations of our explanatory variables related to annual crossing volumes. To represent relationships with diminishing returns, we developed models using the square root transformation (Model B) and cube root transformation (Model C) of each continuous explanatory variable.
Removal of Outliers and Final Model Estimation
For the set of three preliminary models, we examined the model-predicted versus observed annual pedestrian counts to identify potential outliers. We removed 38 intersections from the modeling dataset, including:
Two intersections with very high observed pedestrian counts in small towns (likely to be overcounts that were unrepresentative of other times of year).
Four intersections with very low observed pedestrian counts in cities (likely to be undercounts that were unrepresentative of other times of year).
Two intersections where the observed annual pedestrian count was more than 2 million. These high-volume intersections (averaging more than 5,000 daily crossings) had a strong influence on the fit of the regression model. Since there were only two intersections in this category, we decided that the model would be more representative of the region as a whole without them.
All 30 intersections where the observed annual pedestrian count was less than 1,000. As a whole, these volumes were very difficult to predict. These locations average fewer than four pedestrian crossings a day, meaning that each pedestrian crossing on the data collection day has a large impact on the annual pedestrian volume estimate.
We re-estimated the three preliminary models using the cleaned dataset (n = 260). Note that there were two pairs of explanatory variables in these models with moderate to high correlations. Population density within 400 m of the intersection was correlated with both the percentage of households with no motor vehicles within 400 m (r between 0.7 and 0.8, depending on the transformation of population density used) and the number of bus stops (r between 0.6 and 0.7). In both cases, we tested versions of the models with and without one of the two correlated variables and then repeated the process with and without the other correlated variable. We found that these correlated variables still contributed explanatory power to the models.
Results
Our final models are shown in Table 5. All three models had a good overall statistical fit and had nearly all of their explanatory variables significant at the 95% confidence level. Comparing the three models, we found that Model C had the best overall model fit since its log-likelihood, AIC, and BIC measures all had the smallest absolute values. These measures of overall fit also suggested that both the square root transformations (Model B) and cube root transformations (Model C) of the continuous explanatory variables produced better statistical models than the base model (Model A).
Final Annual Pedestrian Crossing Volume Models
Note: na = not applicable.
Lower absolute values of log-likelihood, AIC, and BIC indicate better overall model fit.
The following variables describing the area surrounding an intersection had statistically significant, positive associations with annual pedestrian volumes:
population density within 400 m;
employment density within 400 m;
number of bus stops within 100 m;
number of retail businesses within 100 m;
number of restaurant and bar businesses within 100 m;
presence of a school within 400 m; and
proportion of households without a motor vehicle within 400 m.
These variables are consistent with previous pedestrian volume modeling studies. Population and employment density represent the total number of people who spend large amounts of time in the vicinity of the intersection, likely generating more pedestrian activity. Bus stops are likely to be associated with pedestrian volumes because walking is the most common form of bus access and egress. Retail businesses, restaurant and bar businesses, and schools are likely to attract customer and student pedestrians. Finally, intersections surrounded by neighborhoods with lower rates of household vehicle ownership are likely to have more people who walk for transportation.
Two additional variables, intersection signalization and presence of a park within 400 m, tended to have consistent positive associations with annual pedestrian volumes during the modeling process. Yet, they were not significant at the 95% confidence level when included with the other seven variables in the final models.
Several theoretically important variables did not show significance in the modeling process. For example, we expected university campuses to have a positive association with pedestrian volumes. However, our sample only had six intersections within 800 m of a campus (three within 400 m) which likely prevented precise estimation of this parameter. High motor vehicle traffic volumes can be a barrier to pedestrians, so we expected that mainline roadway traffic volume would have a negative relationship with pedestrian volumes. Yet, we did not find this relationship, possibly because high-volume roadways often serve areas with many jobs, retail stores, restaurants, and bars, despite being barriers to pedestrians.
External Validation
We used several measures to evaluate how well model-predicted pedestrian counts compared with the observed pedestrian counts at our 45 validation intersections (Table 6). The first two measures, mean absolute error (MAE) and root mean squared error (RMSE), are common methods of comparing overall model prediction accuracy. Smaller values indicate that the model-predicted counts are generally closer to the observed counts across the whole set of validation intersections. According to both measures, the square root model (Model B) and the cube root model (Model C) have lower prediction error than the base model (Model A). The square root model (Model B) provides the most accurate predictions of annual pedestrian crossing volume across the full set of validation intersections.
Comparison of Model Accuracy
Lower values of mean absolute error (MAE) and root mean squared error (RMSE) indicate better overall model prediction across all validation intersections.
We also evaluated the prediction accuracy of the three models at each individual validation intersection. First, we calculated the ratio of the model-estimated count to the observed count and assessed the distribution of these ratio values (Table 6). Ratios close to one are the most accurate. The base model (Model A) has the most ratio values between 0.67 and 1.49, and the square root model (Model B) has the most ratio values between 0.50 and 1.99. All three models were able to predict at least 60% of the validation intersection counts to between 0.50 and 1.99 times the observed value, indicating fair accuracy. As a whole, all three models overpredict pedestrian volumes at more intersections than they underpredict.
Figure 2 displays how each observed value (x-axis) corresponds with the model-predicted value (y-axis). Visually, the square root model (Model B) appears to provide the most accurate representation of the validation counts.

Observed counts versus model-predicted counts at validation intersections.
Assessment of Model Overestimation and Underestimation
Figure 3 maps where the model overestimates and underestimates annual pedestrian crossing volumes at the 45 validation intersections. Roadway intersections with four or more legs across one leg of the intersection tended to be overestimated. Fourteen (93%) of the 15 intersections overestimated by more than two times in Model A had four or more lanes across one leg of the intersection. All intersections overestimated by Model B and Model C had this characteristic (compared with 90% across all validation intersections). We had considered the four-or-more lanes variable in the modeling process, but it was not significant. Therefore, different versions of this variable (e.g., the total number of lanes on each roadway approach to the intersection) could be tested in future models to help reduce overestimation.

Model prediction accuracy at validation intersections.
Intersections overestimated by more than two times also tended to have somewhat lower socioeconomic status. The 15 intersections overestimated by Model A had median incomes within 400 m of $78,100 (compared with an average of $80,300 for the entire validation dataset), 44% rental households (compared with 40%) and 11% of households under the poverty level (compared with 9%). This association between pedestrian volume overestimation and lower economic status indicators was even more prominent for the square root and cube root models. We tested each socioeconomic variable individually in the modeling process, but they were highly correlated and the percent of households with no vehicle variable provided the best representation of low socioeconomic status in the final model. However, future models should consider combining several socioeconomic status measures into a single variable to help address overestimation. This could be done through methods such as principal component analysis.
Of the 10 validation intersections where the model estimate was more than three times too high in the square root model (Model B) and cube root model (Model C), eight had observed annual volumes less than 10,000 (fewer than 28 total crossings on an average day). While we did not consider these locations to have low enough volumes to be removed from the modeling process, they are still likely to experience a large amount of day-to-day variation, making them more difficult to predict accurately.
All three models underestimated pedestrian volumes less frequently than they overestimated. There were no clear trends in the characteristics of underestimated validation intersections.
Discussion
This is the first set of pedestrian intersection crossing volume models developed for the multi-county Milwaukee region. To our knowledge, only California has a pedestrian volume model that can be applied at the intersection level across a larger area. Our models include seven statistically significant explanatory variables that are consistent with previous research and have intuitive relationships with pedestrian volumes. We reviewed our dataset carefully and removed outlier intersections that either had unreliable counts or were so high or low that they would be difficult to predict accurately. Still, we were able to use a broad range of counts from nearly all parts of the multi-county region.
Based on the range of pedestrian volumes in our model dataset, the model outputs are appropriate for annual volumes ranging from 1,000 to 650,000. Therefore, when the model is applied in practice, users should note that intersections with model-estimated values of less than 1,000 or more than 650,000 crossings per year are outside the range of the model. To map estimated annual volumes across the seven-county Milwaukee region, intersections estimated to be below the model range (i.e., in the most rural areas) could all be grouped into a “less than 1,000” category, and intersections estimated to be above the model range (i.e., adjacent to major university campuses and in the heart of the Milwaukee central business district) could be grouped into a “more than 650,000” category.
Our models predicted pedestrian volumes at validation intersections with fair accuracy. The majority of model-predicted counts are within one-half to two times the actual count value. Considering a region with the full spectrum from dense urban to sparse rural land uses and the extensive range of annual pedestrian intersection crossing volumes analyzed, the models are useful for showing broad differences in pedestrian activity between neighborhoods across many parts of the region. Yet, pedestrian volume estimates at some specific intersections are imprecise. The best option to generate a reliable volume estimate at a specific location may still be to count pedestrians in the field.
As a relatively new approach, there are opportunities to improve the models in the future:
Collect pedestrian count data at more intersections throughout the region. It is likely that more data points will make the parameter estimates more precise. More data points could also potentially produce a model that includes more variables that have a statistically significant relationship with intersection pedestrian volumes.
Test new models using variables that could improve overestimation and underestimation errors. Our validation analysis suggests that variables representing the number of roadway lanes at the intersection and low socioeconomic status of adjacent neighborhoods (e.g., median income, poverty, rental properties) could improve future model fit.
Consider other potential explanatory variables such as special attractors (e.g., sports arenas, tourist destinations) or unmeasured roadway characteristics (e.g., posted speed limits or actual traffic speeds, trash, lack of street trees, or high crime rates) that might influence pedestrian activity. Also consider more detailed land use measures, such as the square footage of retail stores, restaurants, and bars.
Add counts at three-way intersections. Our models do not apply to three-way intersections because there were not enough reliable counts available from three-way intersections to include in the modeling dataset. However, this type of intersection could be included if they were counted consistently and represented by an indicator variable in the model.
Develop separate models for each crosswalk. Individual crosswalk volumes may differ based on the land uses on each corner of the intersection. Locations where a major thoroughfare intersects a minor street may also have different numbers of pedestrians crossing parallel and perpendicular to the thoroughfare.
Improve the expansion factors used to estimate annual volumes from short counts at study intersections. Installing more automated pedestrian counters at sidewalk locations in areas with different land-use characteristics will help refine hourly, daily, and seasonal patterns of activity and create an even more accurate dependent variable.
Even with future improvements, model prediction accuracy will likely remain a key challenge for pedestrian volume models. There is a tradeoff between using input variables that are relatively easy to collect and expending many resources to gather fine-grained data that might predict better. Therefore, it is also worth exploring other methodological approaches ( 2 ). For example, some methods are similar to the four-step urban transportation modeling system approach (26, 27). Some of these tools may ultimately provide more accurate pedestrian volume predictions than direct demand models. However, few of these other approaches have undergone extensive validation or been compared directly against direct demand models.
Conclusion
We developed pedestrian volume models using 260 intersections and tested their prediction accuracy at 45 other intersections across the seven-county Milwaukee region. The models can be used to estimate volumes at intersections with as few as 1,000 pedestrian crossings and as many as 650,000 pedestrian crossings per year. Our models include seven statistically significant explanatory variables that have intuitive relationships with pedestrian volumes. The models predicted pedestrian volumes at validation intersections with fair accuracy. Detailed validation analysis showed that the model tended to overpredict volumes at intersections of roadways with four or more lanes and at locations with several indicators of low socioeconomic status. This can inform variable choices in future pedestrian volume models. We hope that our study will help other multi-jurisdiction regions develop better pedestrian volume models for use in safety analysis, project prioritization, facility design, and development impact assessment.
Footnotes
Acknowledgements
The authors would like to thank the Wisconsin DOT Southeast Region office, Southeastern Wisconsin Regional Planning Commission, and other local agencies for sharing data.
Author Contributions
The authors confirm contribution to the paper as follows: study conception and design: Schneider, Qin; data collection: Schmidt, Schneider; literature review: Schneider; analysis and interpretation of results: Schneider, Schmidt; draft manuscript preparation: Schneider, Schmidt, Qin. All authors reviewed the results and approved the final version of the manuscript.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This study was supported by a grant from the Wisconsin Department of Transportation (WisDOT) Bureau of Transportation Safety (FG-2020-UW-MILWA-05068).
