Abstract
With the rise of cycling as a mode choice for commuting and short-distance delivery, as well as policy objectives encouraging this trend, bike count models are increasingly critical to transportation planning and investment. Studies have found that network connectivity plays a role in such models, but there remains a lack of measure for the connectivity of a link in a multimodal trip context. This study proposes a connectivity measure that captures the importance of a link in connecting the origins of cyclists and nearby subway stations, and incorporates it in a negative binomial regression model to forecast bike counts at links. Representative bike trips are generated with regard to bike-friendliness using the New York City transit trip planner and used to determine the deviation from the shortest path via the designated link. The measure is shown to improve model fitness with a significance level within 10%. Insights are also drawn for income levels, bike lanes, subway station availability, and average commute time of travelers.
Cycling is garnering greater attention in the context of global warming as an environmentally friendly mode of transportation. Studies show that the continued increase in cycling will improve public health while simultaneously reducing air and noise pollution in congested areas ( 1 ).
With the rise of bike sharing, researchers have been interested in prediction of bike demand (bike share, personal bikes, cargo bikes, etc.) at both the network and facility-level ( 2 ). Accelerating this trend are new initiatives to use cargo e-bikes for freight delivery as well as the rise of short-distance ecommerce and food delivery. In late 2019, Mayor de Blasio of New York City (NYC) announced the Commercial Cargo Bike Program, in partnership with major freight delivery services such as Amazon, UPS, and DHL, to reduce congestion and carbon emissions by replacing trucks with cargo bikes ( 3 ). The rise of short-distance delivery such as food delivery from the likes of UberEats, Grubhub, Postmates, and so forth, also poses an increase in bike traffic, especially in densely populated cities. Bike deliveries through these services have increased in the U.S.A. since 2016, when Uber introduced bike service in congested parts of Washington D.C., NYC, San Francisco, and Chicago, as bikes were more efficient than cars ( 4 ).
States, regions, and cities increasingly need to account for bike trips in planning and investment decisions. There is a greater need to develop models to forecast bike count at the link level in a road network as well as to identify critical roadways for the bike mode. With a more granular understanding of demand and network structure under the influence of population and land use, effective decision making and de-densifying procedures can be implemented, such as determining on which roadways to grant access to cargo bikes to ensure personal bike riders are not deterred by increased bike density during commuting hours. There is also a greater need to quantify the connectivity of a roadway from a network perspective. Such measures can be incorporated into bike count models as independent variables and be used to rank roadway segments while choosing roadways for bike facility investments or for cargo bike access.
Many studies have tackled station-level and city-level demand forecasting for bike-share systems ( 5 , 6 ), while others have investigated the factors contributing to mode choice and individual ridership, including the effect of built environment and proximity to transit stations, weather, and perceived bike safety ( 7 , 8 ). As for connectivity measures, several network-level measures have been proposed and proven effective, incorporating factors like bike-friendliness (9–11), traffic stress ( 12 ), and travel demand ( 9 ). However, no link-level connectivity measure for bikes has been studied, to the best of the authors’ knowledge, and none that incorporate multimodal routes that include both bike and transit.
The paper proposes a method of evaluating the connectivity of a link with respect to multimodal trips with the bike mode. The effectiveness of the measure is tested with a negative binomial regression. Demographic, socio-economic, and land use attributes, as well as the connectivity measures that are proposed, are fed into the count model as independent variables. Bike count data were collected from 112 locations in Brooklyn and Queens, NYC, by NYC Department of Transportation (NYCDOT) on three weekdays and three weekends as the dependent variable. Estimation results prove the effectiveness of the connectivity measure in capturing link connectivity in the network with respect to the bike mode. Given the limited availability of count data (only 112 locations) and the focus on testing the effectiveness of the connectivity measure, this paper leaves model validation for prediction modeling to future research when more count data are available.
Literature Review
Factors Influencing Bike Usage
Regression modeling has been shown to adequately predict bike counts across given areas and time periods. Early analyses found factors like population density, percentage of land devoted to residential and employment uses, percentage of residents under the age of 16, number of active workers, and number of hours worked per week, to be closely related to bike usage in a given traffic survey zone ( 13 ). The focus has since been drawn to forecasting demand for bike-sharing systems, especially concerning the location of docking stations. A close relationship was established between station-level usage and built environment characteristics like bike lanes and proximity of public transit stations ( 5 ), while further demographic and land use attributes like income, education level, and retail density were investigated more generally in relation to bike usage ( 5 , 14 ).
At the link level, a study found success in limiting the regression problem to weather (temperature, precipitation) and calendar (day of week, time of year, holiday) inputs to predict the numbers registered by a bike counter in Sweden ( 15 ). More generally, Barnes and Krizek ( 16 ) posited that behavioral attributes like commute-to-work data are more likely to provide accurate estimates of bike demand than transportation and land use factors.
A positive relationship has been identified between bike usage and connectivity of bike facility networks and general road networks ( 17 ). The impact of the network on bike demand can be considered from link, node, and network perspectives. From the link perspective, the presence of bike lanes and high level of separation are found to contribute to bike commuting ( 18 , 19 ). Off-road bike paths contribute more to the level of bike usage ( 20 ). Some cyclists are willing to trade longer travel times for off-road bike paths ( 21 ). From a node perspective, the design and traffic control of intersections affect mode decisions ( 17 ). Certain intersection treatments including bike-specific signal phasing ( 22 ), advanced stop line for bikes ( 23 , 24 ), and crossing markings that guide cyclists through the intersection ( 25 ) promote bike usage. At a network level, studies show that extensive networks of separated bike facilities and traffic-calmed streets lead to higher bike usage (26–28). Continuous or connected bike facilities contribute more than general bike facilities ( 29 ). High density of bike facilities also adds to bike usage ( 30 ).
Network Connectivity
Some studies found that network connectivity contributes more to bike usage that cannot be captured by density measures (31–34). To measure network connectivity, a procedure adopted by many studies is as follows. First, indices capturing the bike-friendliness of links are developed. Then, a single connectivity measure of the network is generated based on the indices of links.
Schoner and Levinson ( 35 ) developed several measures of the overall connectivity of bike networks and found that the density of bike facilities has the largest elasticity among connectivity measures developed from graph theory including size, connectivity, fragmentation, and directness of the bike network. However, these measures do not incorporate travel demand, as well as cyclists’ preference for facilities. Klobucar and Fricker ( 36 ) weighted link length by a bike compatibility index, which measures perceived safety as a function of bike-friendly factors such as infrastructure and motorized traffic volume to get a “safe length” for each link. The overall network measure is a sum of the modeled number of bicyclists on each link multiplied by its safe length.
Alternatively, Lowry et al. ( 11 ) used the concept of link-level bike level of service (BLOS) from the Highway Capacity Manual 2010 (HCM), which measures bike-friendliness based on the conditions of infrastructure and motorized traffic. Their measure of connectivity is zonal, with accessibility of the zone considered. As well as bike-friendliness, the level of stress and discomfort can also be related to connectivity. Mekuria et al. ( 12 ) developed the level of traffic stress framework which incorporates the conditions of infrastructure and motorized traffic. They considered two nodes to be connected if they can be reached using only links of a given stress level while limiting the detour to be less than 25% beyond the shortest path. Percent trip connected and percent nodes connected are used to represent network connectivity.
Model Structures
Two types of models are used in this field: those based on travel surveys and those based on count data (“direct demand” models). Among the former, Zhao ( 37 ) analyzed mode choices in Beijing including measures like the jobs/housing ratio and entropy-based diversity index. For the latter, Hankey et al. ( 38 ) proposed linear and negative binomial regressions to forecast bike trips. They found that the negative binomial regression is a better fit than linear models. Fagnant and Kockelman ( 39 ) did Poisson and negative binomial regressions and concluded that negative binomial regression fits better. Chen et al. ( 40 ) proposed a generalized linear mixed model with Poisson distribution to fit to five-year bike count data in Seattle. This addressed one of the shortcomings that Hankey et al. ( 38 ) faced: panel data and temporal autocorrelations. Spatial autocorrelation is also considered, in this case for bike-share trips, by Noland et al. ( 41 ) using a negative binomial conditional autoregressive model.
There are hardly any studies that link the connectivity measures to the outcomes of ridership. Existing measures tend to be network-level measures. At link level, only bike-friendliness is measured, while no connectivity index is developed at link level.
Data
Bike Count Data
The bike count data used as the dependent variable is provided by NYCDOT, and was collected at 119 locations in Brooklyn and Queens, NYC. The data were collected on three weekdays (June 12, 19, and 26, 2018) and three weekends (June 10, 17, and 24, 2018) through camera recordings. Bikes were counted in intervals of 15 mins from 7:00 a.m. to 7:00 p.m., and then aggregated to daily counts. Each location had one weekday count and one weekend count correspondingly. The average daily bike count among the sites on weekdays was 170.14, with a standard deviation of 193.56. The average bike count on weekends was 153.12, with a standard deviation of 179.99. The locations and bike counts are shown in Figure 1. However, because of lack of data availability of some independent variables, only 112 of them were kept for estimation.

Data collection locations with bar charts of average daily bike counts.
The bike counts on both weekdays and weekends are found to be spatially correlated (weekday count: Moran’s index = 0.20, z-score = 5.97; weekend count: Moran’s index = 0.27, z-score = 7.93). However, based on the findings from Noland et al. ( 41 ), there are only minor differences in a model with and without spatial autocorrelation. In keeping with the focus on exploring the structure of the connectivity measure, spatial autocorrelation is not used in the proposed model.
Other Data Sources
The data sources that are used to generate the independent variables are listed in Table 1.
Other Data Sources Used to Generate Independent Variables
Methodology
Count Model and Analysis
Count Model
Negative binomial regression was found to be a better fit for bike count models compared with linear regression and Poisson regression. The distribution of weekday and weekend bike counts approximately follow negative binomial distributions, as shown in Figure 2, a and b , respectively. Negative binomial regression was thus selected as the method of modeling bike count.

Negative binomial distribution fitted to bike count data for (a) weekdays and (b) weekends.
To model bike count as well as testing the effectiveness of the connectivity measure on weekdays and weekends, four negative binomial regression models were estimated:
Weekday bike count regression without a connectivity measure,
Weekday bike count regression with a connectivity measure,
Weekend bike count regression without a connectivity measure,
Weekend bike count regression with a connectivity measure.
Backward elimination is used to select attributes.
Estimation Result Analysis
To evaluate whether the proposed connectivity measure improves on the model fit, a likelihood ratio test was conducted. The test statistic equation shown in Equation 1 represents the fitness improvement that the connectivity measure brought.
The above statistic approximately follows a Chi-square distribution. The null hypothesis for both weekday and weekend models is that the origin-to-subway connectivity measure does not improve the model fit significantly. Given a confidence level, we can determine whether to reject the null hypothesis or not.
A travel time coefficient −4.31 as in Equation 2,which is from the synthetic population of NYC (listed in Table 1), is assumed. The value of this coefficient may influence the significance of the connectivity measure. A sensitivity test of the coefficient was thus conducted to see how much the coefficient affects the estimate of connectivity measure coefficient. The travel time coefficient
Connectivity Measure
Origin-to-Subway Connectivity Measure
It was observed that bike counts tend to be higher near subway stations, as shown in Figure 1. Since passenger transportation in NYC is heavily reliant on the subway system, the significance of “a connectivity measure that considers the importance of a link in connecting an origin zone to a subway station as part of a bike-and-ride multimodal trip” is tested, which is named here the “origin-to-subway connectivity measure.” The authors did not choose to generate an “origin-to-bus-stop connectivity measure.” The reason is that the bus stops in the study area are dense and close to all the data collection locations, so it was supposed that the “origin-to-bus-stop connectivity measure” is unlikely to be a significant influencer of bike usage. With the above motivation, the following method is proposed to compute the origin-to-subway connectivity measure. The idea is to evaluate how close a link is to the shortest bike paths of all origin–destination (OD) pairs with subway stations as destinations, weighted by the OD demand. The demand is taken from the synthetic population of NYC ( 43 ). Note that the paths found are meant to be normative measures included as attributes in the model, not descriptive measures predicting paths that people would be taking.

Traffic analysis zones (N = 230) that are identified as origin areas of the cyclists at the data collection locations.

Area frequency distribution of the 230 identified traffic analysis zones.

Subway stations used by cyclists at the data collection locations.
A travel time parameter of

Screenshot of 511NY rideshare transit trip planner ( 49 ).
Bike-Related Path Generation
To determine the shortest path time for various trips mentioned in “Origin-to-Subway Connectivity Measure” trips between OD pairs were generated using the “511NY” TTP ( 49 ), as illustrated in Figure 6. The TTP was built using OpenTripPlanner (OTP) and covers the transportation system and services throughout New York State. It offers an interface specific to bike use in NYC. The user can search for routes involving a combination of transit + personal bike, transit + bikeshare, or bike only. Multimodal trips were generated using the transit + personal bike option, with transit limited to the subway, while bike-only trips were simply designated as such. TTP also offers the option to indicate a maximum biking distance, bike speed, and additional optimization of routes for bike-friendly travel. The average cycling speed in midtown Manhattan is around 7 mph ( 50 ), so a faster cycling speed, 10 mph, is assumed in the lower density research area in Queens. The maximum cycling distance is assumed to be 10 mi, since that is a doable distance for non-professional cyclists ( 51 ). Bike-friendly routes favor multi-use paths, low-traffic streets, and streets with bike infrastructure ( 52 ). Streets are weighted based on various factors such as bike lanes (0.60), residential areas (0.98), pedestrian sidewalks and crossings (1.1–2.5), and so forth. Weights greater than one serve to elongate the distance of that link within the routing algorithm, while weights shorter than one do the opposite. In this study, the above default weights from TTP were used.
Using these options, multimodal trips originating from the study area to various locations in NYC, and bike-only trips from the study area to relevant subway stations, were generated. Trip times and relevant subway stations were obtained from the website using Selenium WebDriver, which is a browser automation software library that allowed the authors to automate the process of inputting thousands of OD pairs into the TTP. Selenium automatically navigates a browser, as a user would, selects the relevant trip options such as bike speed and transit + bike modes, and inputs the latitude and longitude of the origin and destination points into the TTP. It then automatically retrieves the resulting details of the trip (e.g., total trip duration, bike time, transit time, names of transit stations used, etc.), thus giving the necessary inputs to calculate the connectivity measure.
The measure is sensitive to the location of subway stations so it can evaluate subway station investments. It is also sensitive to bike-friendliness, so improving bike-friendliness measures used by the TTP (e.g., multi-use paths, low-traffic streets, and streets with bike infrastructure) on the link would alter its forecast as well.
Alternative Independent Variable Generation
Census Tract Data
For demographic and commute factors including population density, gender breakdown (male percentage), median age, average commute time, percentage commuting by transit, and land use density, census data derived from the American Community Survey 2014 to 2018 five-year estimate was used (
42
). The data was broken down by census tract so that the attribute values were more closely associated with the exact location of each link. Each link was assigned to a census tract if the two intersected, and the demographic data of that census tract was in turn assigned to the link. Attributes such as population density and land use density (i.e., total housing units per square foot) were expressed with reference to the area of the tract. If a link fell within two census tracts, the data from each tract were combined and expressed as a function of the combined area of the two tracts. The average commute time of a census tract is processed as four dummy variables, with the lowest level as the reference. The five levels are [0, 30) minutes (reference level), [30, 35) minutes, [35, 40) minutes, [40, 45) minutes, and [45,
Synthetic Population
Household vehicle ownership and income data used in the analysis are from the synthetic population ( 43 ). The household vehicle ownership is averaged in each TAZ. Average household vehicle ownership of a TAZ is merged to a data collection location if the location intersects with the TAZ.
Average household income of a TAZ is processed as three dummy variables, with the highest level as reference. The four levels are [$35,000, $45,000) per year, [$45,000, $60,000) per year, [$60,000, $100,000) per year and [$100,000, $150,000) per year (reference level). Only these ranges are used because the synthetic population is derived from samples obtained from the NYMTC Household Travel Survey, which only included household incomes in those ranges in the study area. For each traveler in the synthetic population of the study area, the median of his/her income range was taken as the income of that traveler. The average household income of each TAZ was then computed and the range that the average value falls in taken as the income range of the TAZ. The average income range of a TAZ is merged to the data collection location if the location intersects with the TAZ.
Land Use Data
Data used were from the Primary Land Use Tax Lot Output (PLUTO) (
44
) of NYC Department of Planning, which is an extensive land use and geographic dataset at the tax lot level. Land use entropy was computed from PLUTO, which is an index of land use diversity. PLUTO data classifies all the tax lots into 11 land use classes: (1) One and Two Family Buildings; (2) Multi-Family Walk-Up Buildings; (3) Multi-Family Elevator Buildings; (4) Mixed Residential and Commercial Buildings; (5) Commercial and Office Buildings; (6) Industrial and Manufacturing; (7) Transportation and Utility; (8) Public Facilities and Institutions; (9) Open Space and Outdoor Recreation; (10) Parking Facilities; and (11) Vacant Land. ArcGIS was used to intersect the TAZs with the tax lots to identify the number of land use classes in each TAZ as well as to compute the area of each class. The computation is shown in Equation 3, adopted from Zhang et al. (
53
), where J indicates the set of land use classes in a TAZ and
Transit Data
The proportions of bike trips used to access transit stations are relatively low. For example, the 2018 NYC Mobility Report ( 54 ) indicates only 0.3% of transit trips are accessed by bike. With 6.1 million daily transit trips in 2016 and 460,000 daily bike trips, the 0.3% suggests only 18,000 bike trips (~4%) are access trips to transit. However, other studies in the literature review have indicated that the presence of transit stations does play a role. The subway station and bus stop data that were used are from NYC OpenData ( 45 , 46 ). ArcGIS was used to generate a half-mile buffer of the data collection locations and the number of subway stations and bus stations that fall in that buffer was counted. The counts are correspondingly merged to the locations as alternative attributes: number of subway stations within a half-mile radius of the link and number of bus stops within a half-mile radius of the link.
Crime Data
Crime data that were used are the NYPD complaint data of 2019 ( 47 ). The complaint records are aggregated to each TAZ, and then divided by the population in the TAZ. Crime data of TAZs are merged to a data collection location if the location falls in that TAZ.
Bike Facility Data for NYC
Bike facility data that were used are from NYC OpenData ( 48 ). Three attributes were derived from the bike route dataset, which are: the existence of a bike lane, the number of bike lanes, and the class of bike lanes. Existence of a bike lane is processed as a dummy variable, indicating whether or not there are bike lanes at the location. Bike lanes in NYC are classified into three classes: physically separated (class I), marked with paint and signage (class II), and shared with vehicles (class III), which is also processed as a set of three dummy variables, with class I as the reference level.
Summary of Independent Variables
Table 2 summarizes the alternative attributes that were collected from the above data sources. For the sets of dummy variables that are decomposed from categorical variables, the level omitted in the models is labeled as the “reference” level.
Alternative Attributes
Note: NA = not available.
Model Estimation and Discussion
The results of the weekday and weekend model estimation are shown in Tables 3 and 4, respectively. Two different models (model without connectivity measurements, model with origin-to-subway station connectivity measurement) were compared respectively on weekdays and weekends.
Estimation Results of Weekday Models (Dependent Variable: Daily Weekday Bike Counts)
Level of Significance (*< 0.1, **<0.01, ***<0.001).
Estimation Results of Weekend Models (Dependent Variable: Daily Weekend Bike Counts)
Level of Significance (*< 0.1, **<0.01, ***<0.001).
The values of the Chi-square distribution with p = 0.05 for the weekday models (df = 106) and the weekend models (df = 105) are 131.03 and 129.92 respectively, which are smaller than the Pearson Chi-square statistics of the model. This suggests more data is needed to produce a more effective predictive model in future research. Nonetheless, the estimated parameters provide valuable insights to the relationships between the bike counts and the various attributes including network connectivity.
Testing the Effectiveness of the Connectivity Measure
Likelihood Ratio Test
For weekday and weekend models, the likelihood ratio test statistics are 2.70 and 3.04, respectively. Given a confidence interval of 10%, by referring to the Chi-square distribution with degree of freedom = 1, the null hypothesis is rejected for both weekday and weekend models: the origin-to-subway connectivity measure does not improve the model fit significantly. Therefore, the model with connectivity measure is better fitted than the base model for weekdays and weekends respectively. Adding origin-to-subway connectivity significantly improves model fitness for both weekday and weekend bike count models. This measure captures the transit access network structure as well as the distribution of bike facilities across the network. Connectivity in weekday and weekend models both have positive signs and high significance levels, which is in line with expectations. From the above observation, it can be concluded that the effectiveness of a link connecting trip origins and subway stations is a significant contributor to the bike count on the link, and this can be effectively captured by the origin-to-subway connectivity measure.
The count of nearby subway stations and origin-to-subway connectivity is not highly correlated, with a correlation of −0.29. This indicates that the connectivity index that was defined captures information that cannot be captured by simple density measures.
Sensitivity to Travel Time Parameter
Because the travel time coefficient −4.31 that was used in Equation 2 is from the synthetic population (
45
), a sensitivity test of the coefficient is performed to see how much the coefficient affects the estimate of connectivity measure coefficient. The travel time coefficient
For the weekday models, varying the travel time coefficient by +10% and −10% leads to +0.086% and −0.136% change of the connectivity measure coefficient. The changes in p-values are −0.001 and 0.000, respectively. The changes in log-likelihood value are less than 0.01. For the weekend models, variation by +10% and −10% leads to +0.015% and −0.023% of change of the connectivity measure coefficient. The changes in p-values are 0.000 and +0.001. The changes in log-likelihood value are less than 0.05. With very small changes caused by the travel time coefficient change, the connectivity measurement is stable as a significant contributor to the bike counts.
Other Attributes
The correlations of selected attributes are shown in Figure 7. All correlations are between [−0.3, 0.3). The variables that are insignificant are not included. Variables that are not included for this reason are “gender_male,”“median_age,”“HH_inc_45k_60k,”“HH_inc_60k_100k,”“avg_commute_30_35min,”“avg_commute_35_40min,”“avg_commute_>45min,” and “per_commute_transit.” The variables that are highly correlated (correlation >0.3 or <−0.3) are not included at the same time. Variables that are not included for this reason are “landuse_dens,”“landuse_entropy” (highly correlated with “pop_dens”), “crime_per_person” (highly correlated with “SubwayCount”), “HH_veh_ownership” (highly corretated with the income variables), “BusCount” (highly correlated with “SubwayCount”), “BikeLane_count,”“BikeLane_class_none,”“BikeLane_class_type3,” and “BikeLane_class_type2” (highly correlated with “BikeLane_exist”).

Correlation heatmap of independent variables included in the count model.
Comparing the significant variables with those in the literature, the findings are mostly the same. High population density has been found to be a contributor to bike usage in the literature ( 5 , 13 ), which is in line with the positive coefficient of population density. Among all the income levels, only income below $45,000/year contributes significantly to bike usage. Compared with higher income levels, lower income levels lead to higher levels of bike usage. However, in the literature, higher median income of an area is found to contribute to the usage of bike share ( 5 ). It is inferred that bike-share users and self-owned bike users have different income effects. Subway station count within a half-mile radius contributes to bike flow, which coincides with the findings in the literature that proximity to public transit stations (in their cases, bus stops) contributes to bike usage ( 5 , 15 ). Existence of bike lane leads to higher bike flow, which is in line with the finding in Rixey ( 5 ) that bike lane width contributes to bike flow.
Considering bike facilities, the existence of bike lanes contributes to bike flow. However, if we choose to include the number of bike lanes and the class of bike lanes and leave out the existence of bike lanes, they are not significant. This might be because 83 of the 112 data collection locations do not have bike lanes of any type. The influence of more bike lanes and different types of separation may not be sufficiently reflected in the data.
As for transit availability, comparing the count of subway stations and bus stations in a half-mile, the former is significant while the latter is not that significant. This can be explained by the conclusion that cyclists mainly cycle to subway stations in this area.
Average commute time and percentage of commuting by transit represent people’s mode choice as well as the proximity between the origin zone and the Central Business District area of the city. Average commute time of between 40 and 45 min is a significant contributor to bike usage on weekends. Considering the conclusion that cyclists mainly cycle to subway stations in this area, it is inferred that in NYC longer trips (>40 min) are more likely to be bike-and-ride trips. (The reason for “average commute time greater than 45 min” not being significant might be that the locations with this variable set as 1 are too few.) However, it is not significant in the weekday model. The reason might be the energy-consuming nature of the bike mode or the limitation of the sample.
Attributes that are not significant in the model are not necessarily irrelevant. Collinearity is a major reason. Population density, land use density, and land use entropy are highly correlated in this area, which may not be the case elsewhere in the city. Crime rate has a positive coefficient when included, which might be caused by its high correlation with the count of subway stations. Household vehicle ownership is highly correlated with income level.
Conclusion
This study proposed a link-level origin-to-subway connectivity measure which captures the “power” of a link connecting origins and subway stations. The idea is to measure connectivity of a link by computing the demand-weighted probability for a cyclist to go through the link, given the choices of the shortest path and the path with the link. Bike-friendliness was incorporated through its inclusion in the path generation using TTP. Negative binomial regression was conducted to model bike counts at 112 locations in Brooklyn and Queens, NY to test the significance of the measure. Independent variables include demographic, land use, and infrastructure attributes as well as the origin-to-subway connectivity measure.
Through the study, it was found that the proposed connectivity measure reliably improves model fit. Using a likelihood ratio test, the significance level is within 10% that the measure improves model fit for both models. Other significant attributes are shown to align with the findings from the literature. The sensitivity test shows that the coefficient and significance of the connectivity measure stay stable with variations in the travel time coefficient used computing the connectivity measure.
The proposed framework of measuring link-level connectivity can be applied to other topics and modes. For example, if the planners of a bike-share system wished to identify proper bike station locations, given the OD matrix, roadway network, and bike facility network, the connectivity of the links could be computed, which would be a valuable reference for identifying bike station locations. However, this study focuses on analyzing and evaluating the effectiveness of the connectivity measure. Before applying the model to prediction, further steps are needed. More data and a more effective predictive model might be needed in future research. Because of the limitation of sample size and data collection area, the reliability of the connectivity measurement and the conclusions from estimation results should be further tested and validated with larger samples and various types of urban form. For example, in an urban area that is less reliant on the subway system, the connectivity measure may be more effective as origin-to-destination. Furthermore, with count data of different types of bikes (shared bikes, personal bikes, and cargo bikes), which is not available in this study, different count models can be estimated for different types of bikes for more detailed analysis and prediction. For the evaluation of bike-friendliness while searching for bike-related paths, the weighting method adopted by TTP should be effective but has not been proven effective compared with other methods. Alternative ways of evaluating bike-friendliness can be compared while computing connectivity. For model structure, a geographically weighted regression could be considered using inverse distances between sample locations as the weight matrix. While the model would not be directly transferable to other cities because of the use of only NYC data, the encouraging results provide a benchmark to compare with models of other cities in future studies.
Footnotes
Acknowledgements
Data shared by NYC DOT (Mark Seaman) is gratefully acknowledged.
Author Contributions
The authors confirm contribution to the paper as follows: study conception and design: J. Y. J. Chow, B. Liu; data collection: B. Liu, D. Bade; analysis and interpretation of results: B. Liu, D. Bade, J. Y. J. Chow; draft manuscript preparation: B. Liu, D. Bade, J. Y. J. Chow. All authors reviewed the results and approved the final version of the manuscript.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by C2SMART University Transportation Center (U.S. DOT #69A3551747124) and the NYU Summer Undergraduate Research Program.
The opinions expressed in this paper, and any errors, are those of the authors. This paper does not constitute a standard, specification, or regulation.
