Abstract
Much of the housing sub-market literature has focused on establishing methods that allow the partitioning of data into distinct market segments. This paper seeks to move the focus on to the question of how best to model sub-markets once they have been identified. It focuses on evaluating the effectiveness of multilevel models as a technique for modelling sub-markets. The paper uses data on housing transactions from Perth, Western Australia, to develop and compare three competing sub-market modelling strategies. Model 1 consists of a city-wide ‘benchmark’; model 2 provides a series of sub-market-specific hedonic estimates (this is the ‘industry standard’) and models 3 and 4 provide two variants on the multilevel model (differentiated by variation in the degrees of spatial granularity embedded in the model structure). The results suggest that the more granular multilevel specification enhances empirical performance and reduces the incidence of non-random spatial errors.
1. Introduction
Housing sub-markets arise as a result of the co-existence of a high degree of heterogeneity of preferences in relation to house types, sizes and locations on the demand side of the market and an extremely variegated and indivisible stock of properties on the supply side (see Grigsby, 1963; Maclennan, 1982; Watkins, 2008). The way in which segmented demand is matched on to the differentiated stock gives rise to identifiable sub-markets. Each sub-market is quasi-independent and exhibits an equilibrium price that remains distinct from that of other market segments even in the long run.
This pervasiveness and durability mean that the existence of sub-markets is of considerable analytical significance. As Galster (1996) explains, sub-markets provide a framework from which to understand market dynamics and the way in which policy interventions work through the housing system. He argues that changes in one sub-market have important but predictable repercussions for price changes and migration flows in other sub-markets. It is argued that an understanding of the sub-market structure can assist the decision-making of a variety of housing-sector stakeholders. This might include improving the effectiveness of public-sector expenditure (Bates, 2006), directing the use of tax instruments (Berry et al., 2003), enhancing private-sector investment and mortgage lending strategies by allowing more robust risk pricing (Goodman and Thibodeau, 2007), enriching estate agents’ marketing strategies (Palm, 1978) and helping to refine housing consumers’ search strategies (Maclennan et al., 1987). Significantly, it is also clear that failure adequately to accommodate housing sub-markets can undermine the performance of housing market models by limiting predictive accuracy. This has important practical implications for the methods used in constructing house price indices (Spinney et al., 2011), applying mass appraisal techniques (Adair et al., 1996), undertaking environmental impact assessment (Michaels and Smith, 1990) and measuring the implicit value of public infrastructure programmes (McGreal et al., 2000).
Watkins (2012) suggests that the sub-markets literature has emerged in three waves. The first occurred in the 1950s and 1960s, led by a group of institutional economists who identified the potential of the sub-market as an analytical construct that could be used to track housing market change (see, for instance, Fisher and Winnick, 1952; Grigsby, 1963). This work was motivated by a desire to engage in debates about efficiency and equity of housing policy interventions and, to date, continues to frame most conceptual discussions. The second wave, during the 1970s and early to mid 1980s, saw the development of a series of standard econometric tests for sub-market existence (Schnare and Struyk, 1976; Goodman, 1978). This was motivated largely by concerns that the coefficients in market-wide hedonic models were subject to aggregation bias (Straszheim, 1975). The third wave has been the most voluminous in terms of published outputs, largely as a result of improvements in the availability of impressively detailed micro datasets. This work has focused on how best to use statistical methods to reveal clusters in the data (see later for a more detailed discussion). This has seen an emerging consensus around two ways of partitioning data to reveal sub-market formulations: the first uses statistical methods, including for example techniques such as principal components and cluster analysis (exemplified by Bourassa et al., 1999) and the estimation of isotropic semi-variograms (see Tu et al., 2007) while the second uses markets experts, such as estate agents and valuers, to define segments (see for instance, Keskin, 2010; Bourassa et al., 2003).
This paper seeks to contribute to the development of a fourth phase in the evolution of the sub-market literature. It seeks to build on the emerging consensus about how best to partition data by shifting the focus on to how best to accommodate the sub-markets revealed within house price models. As Costello et al. (2010) note, to date, there have been few to attempts systematically appraise alternative ways of modelling sub-markets. The dominant approaches have been based on simply including sub-market dummies within hedonic models (for example, Fletcher et al., 2000; Butler, 1982) or estimating a set of sub-market-specific hedonic equations (for example, Bourassa et al., 2003; Goodman and Thibodeau, 2003). The former has been criticised for failing to allow the implicit price of individual attributes (such as a parking space) to vary between sub-markets (Maclennan et al., 1987). The latter addresses this but suffers from an inability to differentiate between the effects of hard boundaries such as school catchment areas and softer and more fluid spatial influences such as neighbourhood quality (see Clapp and Wang, 2006). As we argue later in this paper, in operational terms the utility of both of these methods is highly constrained by the need to impose hard sub-market boundaries that draw on pre-determined partitions.
This has spawned an interest in the application of multilevel modelling strategies as an alternative basis for modelling housing sub-markets and capturing the fluidity of sub-market boundaries over time (Leishman, 2009; Orford, 1999; Goodman and Thibodeau, 1998). Multilevel models are advised when the observations being analysed are clustered and correlated, the causal processes underlying the relationships operate simultaneously at multiple spatial scales and there is value in seeking to disentangle the spatial effects (Subramanian, 2010). Their use has begun to expand within the quantitative human geography literature and the technique has been used to explore a range of complex spatial impacts and interactions including the composition of public health outcomes and measurement of social well-being (see Moon et al., 2005; and Ballas and Tranmer, 2008, respectively). This clearly resonates with the challenges associated with modelling housing sub-markets.
Thus the main aim of this paper is to undertake an appraisal of the performance of multilevel models of housing sub-markets. Specifically, the paper analyses the outputs of two different variants of a multilevel house price model and compares the results with those generated by employing the more standard approach of estimating a series of individual house price functions for each separate sub-market. Both modelling strategies employ agent-based definitions of sub-markets that have been shown to be superior to other partitioning schemes (see Costello et al., 2011, for evidence). The empirical analysis is designed as a comparative experiment. It applies different methods to data from Perth, Western Australia, covering the period between early 2007 and the end of 2008. The research design is adapted from a series of previous studies that explore the empirical performance of competing sub-market formulations (Costello et al., 2010; Bourassa et al., 2003; Goodman and Thibodeau, 2003; Watkins, 2001). There are three stages to our performance evaluation. The first stage seeks to develop a robust hedonic house price model to act as an ‘industry-standard’ benchmark against which the performance of the alternative sub-market modelling strategies can be compared. The second stage parameterises the competing models: specifically it estimates the set of sub-market-specific hedonics and the two variants on the multilevel model. The third stage explores the predictive accuracy of the models.
The paper has four main sections. The next section explores the existing literature to outline the nature of sub-markets. It establishes the need for modelling strategies to accommodate sub-markets and, if possible, be able to deal with dynamic change within sub-market structures. Section 3 describes the data and methods of estimation used in the paper. Section 4 presents the main modelling results and discusses the comparative performance. The final section sets out some conclusions.
2. The Nature of Sub-markets and the Case for Multilevel Modelling Strategies
Research on housing sub-markets has, hitherto, focused on the development of consistent methods of identifying their boundaries. This reflects concerns by some commentators that the lack of a common approach to the definition and identification of sub-markets contributed to lack of a consensus about their importance in the analysis of metropolitan housing markets (see Rothenberg et al., 1991). The explanation for sub-market existence set out in the opening paragraph of this article, based implicitly on the contribution of Grigsby (1963), emphasises that potential sub-markets are clusters of dwellings that are relatively close substitutes in the view of those who demand housing, although not necessarily in close spatial proximity (see Galster, 1996, for a detailed discussion).
Maclennan and Tu (1996) emphasise the fact that neighbourhood and environmental attributes are traded with housing, alongside physical attributes. They point to the indivisibility of some housing attributes, and impossibility of replication of others, as root causes of sub-market creation. Examples of non-divisible attributes include those typically measured by researchers using dummy variables, such as property type. Non-replicable attributes are more likely to relate to a property vintage. For example, stone-built properties were constructed at lower cost in the past than they can be today, hence, the existing stock of such properties is difficult to replicate.
From these facts, Maclennan and Tu (1996) develop an argument first articulated clearly by Schnare and Struyk (1976), that consumers’ demand for non-divisible, non-replicable attributes may be price inelastic. This may give rise, in essence, to a two-stage housing choice in which consumers restrict their potential choices to those possessing a particular attribute, or bundle of attributes. This might reflect an overriding desire to locate in the catchment area of a highly ranked school, as in the Schnare and Struyk example, or an overriding desire to consume a desirable bundle of environmental and neighbourhood attributes. The result, in either case, is that consumers seek to maximise utility from the available bundles of physical attributes only after restricting the potential options to exclude those that do not reflect their overriding desires (those for which their demand is relatively price inelastic).
Despite increasing clarity about the conceptual basis for sub-market existence, there is no evident consensus about the appropriate approach for identifying or testing for sub-markets. This has spawned considerable investment in studies that explore different mechanisms for partitioning house price datasets, driven in part by a desire to move beyond the imposition of sub-market structures based on prior notions or pre-existing administrative boundaries. Bourassa et al. (1999), for instance, demonstrate one widely accepted approach to testing for spatial sub-markets: the use of a combination of principal component analysis and hedonic regressions. Chow tests and weighted standard error tests related to the latter are used to ensure that the existence of spatial sub-markets is accepted only when parameter estimates vary across the metropolitan area and stratification leads to greater predictive accuracy. The authors conclude that further research directions might include an exploration of methods to determine the optimal number of sub-markets in a metropolitan area.
Interestingly, several studies found that spatial sub-markets based on real estate agents’ definitions led to models with greater predictive accuracy compared with those based on statistically derived sub-markets. Michaels and Smith (1990) asked five agents to cluster 85 locations within suburban Boston into between five and ten mutually exclusive sub-markets. This generated three useable classifications: one with ten segments and the two others with four. The returns were then amalgamated to give a composite classification of four sub-markets. The expert-defined boundaries produced house price estimates that substantially reduced standard errors when compared with a market-wide hedonic formulation. The method used to develop the sub-market classification was, however, subject to criticism for its lack of rigour.
Bourassa et al. (2003) also explored the capacity of real estate professionals to define sub-markets. The experts consulted identified 34 sub-markets which were collapsed to 18 using statistical methods. The study compared price models estimated using the expert-based boundaries with those developed for sub-markets identified by using a combination of principal component analysis and cluster analysis to identify non-contiguous groupings of properties. Again, the models that used the boundaries identified by market experts produced the greatest predictive accuracy and led to a significant increase in the proportion of predicted prices within 20 per cent of actual values. Elsewhere, in a study of the Glasgow housing market, Watkins (2001) used the boundaries used by agents in listing service publications to delineate the market. This sub-market formulation proved superior, in terms of reducing standard error to the alternative produced using the standard PCA and cluster analysis methods. More recently, Keskin (2010) invited eight estate agents working in the Istanbul housing market to draw sub-market boundaries on a 1/200 000 scale map of the city. The interviewees sketched between five and seven sub-markets, even though there was no guidance on the number of segments thought to exist. GIS technology was used to layer the maps on top of each other and this process helped to produce a composite sub-market geography based on five distinct segments. House price models were estimated for each segment and the results were compared with competing models developed using the standard PCA and cluster analysis methods described earlier. The agent-based models produced considerable benefits in terms of predictive accuracy: 21 per cent of the estimates were within 10 per cent of the actual value, versus 15 per cent for the alternative approach.
Despite the evidence that agent-based approaches have considerable utility and can be derived simply, several commentators identify the instability of the boundaries generated by these approaches as an on-going problem (Watkins, 2011). It is clear that, irrespective of the quality of data and analytical rigour underlying some previous cross-sectional analyses of the metropolitan housing market structure, the findings of such studies have limited value if sub-market structures are subject to significant or rapid change.
A series of studies has explored the stability of sub-market boundaries with respect to migration and, specifically, the concept of filtering (see Jones et al., 2003, 2004; Rothenberg, 1991). An important argument implicit in these studies is that, while intrametropolitan differences in housing attribute prices may be interpreted as evidence of sub-markets, there is no reason to suppose that these price differences are stable over time. Differential rates of new housing supply and migration between sub-markets may act to break down sub-market boundaries, effectively smoothing attribute price differences spatially through arbitrage processes (see Jones et al., 2004). Interestingly, Jones et al. (2003) tested the temporal stability of previously identified spatial sub-market boundaries and found evidence that, although several sub-markets remained unchanged, others had been subject to some modification.
The modification of sub-market boundaries presents a challenge to the standard methods used to accommodate sub-markets within house price models. One way of dealing with this challenge is to develop a modelling strategy that builds on standard definitions but allows some fluidity in the precise definition of boundaries to be revealed emprically. Bourassa et al., (2007) seek, in part, to do this by using lattice models and geostatistical approaches, but their main finding is that a traditional hedonic model with sub-market dummies has superior predictive performance compared with their comparative models. However, they do note the potential for further comparison with the approach demonstrated by Pavlov (2000) and Fik et al. (2003) which included x/y co-ordinates in the hedonic models. The latter also used interactions between x/y co-ordinates and location dummies. These studies were motivated by the desire to improve hedonic estimation in the absence of prior knowledge of sub-market boundaries and thus, allow spatial effects to emerge in complex patterns with some decaying over varying differences from the individual homes.
This has also provided the context for the emerging interest in the potential of multilevel models that has appeared (apparently) independently in the UK and the US. The initial contributions developed from the notion that hedonic specification could be better contextualised by applying the expansion method (Can, 1992). In other words, a more complex model can be developed by expanding the parameters of the simple hedonic equation (see section 3 for more formal mathematical notation that illustrates this point). In the UK, Jones and Bullen (1993) developed an expanded multilevel hedonic with two tiers: the property level and the sub-market level. This formulation captures the market-wide influences on property values, but also allows parameters to vary between sub-markets. Thus, the price of a property is a function of the market-wide price and a sub-market-specific differential. The approach was applied to data on individual properties, in 33 London local authorities, drawn from the 5 per cent Survey of Building Society Mortgages collected by the Department of the Environment (DoE) (see Jones and Bullen, 1994). The structure of the dataset limited the scope of fine-grained spatial analysis. With as few as 20 observations within each district, there was little scope to analyse sub-markets at the micro level employed by other analysts (see Orford, 2000, 2002).
In the US, Goodman and Thibodeau (1998) introduced a similar two-level (property and sub-market) specification. The approach involved identifying spatial sub-market areas using data on housing transactions that took place in a single school district in Dallas, Texas, between early 1995 and 1997. The housing data were augmented with information on the performance of public elementary schools and the results showed that significant price differentials existed between school catchment areas. This approach was developed further in future papers and, with access to a larger dataset covering the entire metropolitan area, the researchers were able to establish a hierarchical model with multiple levels (Goodman and Thibodeau, 2003, 2007). The rationale for the model is that all dwellings share the amenities available within their locality and thus the determinants of house prices are nested within multiple geographies: properties are located within neighbourhoods, neighbourhoods within school districts and school catchments within municipal boundaries. The analysis showed evidence of differentials at a variety of spatial scales.
The potential of this approach has been explored further elsewhere. Orford (2000) uses around 1500 housing observations collected from estate agents in Cardiff, Wales, to examine how a multilevel approach might explicitly incorporate spatial market segments. The paper showed evidence of price differentials for sub-markets that reflected segmentation associated with particular communities, reinforced by institutional factors including the influence of agents and the significance of structural heterogeneity within the housing stock. More recently, Leishman (2009) develops a multilevel approach that demonstrates the possibility of modelling a unitary metropolitan housing market, but allowing coefficients to vary between small, census-derived geographies within the city. The paper shows that, using a multilevel hedonic estimation approach, sub-market boundaries in Glasgow changed significantly within a relatively short time-period (of three to four years). It is this basic model structure that provides the general framework for the empirical analysis that follows in this paper.
3. Research Questions and Approach
The empirical analysis is motivated by a number of research questions that emerge from our review of the literature
—Is it appropriate to model the metropolitan housing market without accounting for the possibility of spatial divisions?
—Does a spatially segmented model based on real estate agents’ definitions of ‘sub-markets’ out-perform a unitary model?
—Does a multilevel hedonic model out-perform the spatially segmented model?
—How do different variants of the multilevel approach perform?
To determine the empirical performance of a number of conceptual approaches to defining housing sub-markets, we estimate a number of empirical models for the city of Perth, Western Australia. We then proceed to measure prediction error, both in an aggregate sense and in terms of underlying spatial patterns, for each of these models. We use GIS and tests for spatial autocorrelation to demonstrate spatial clustering of errors (see Fik et al., 2003, for another example of this approach).
Model 1 is a simple hedonic model estimated using OLS. Its specification includes a set of continuous predictors including distance from the CBD. This model provides a benchmark with which to compare the statistical performance and predictive accuracy of the later models as well as embodying the ‘unitary housing market’ hypothesis
where,
Model 2 is really a set of hedonic models estimated separately according to pre-defined spatial divisions in the data, taking them as a representation of a priori sub-markets. The spatial sub-divisions are derived from market analysis published by the Real Estate Institute of Western Australia (REIWA). In these regular statistical publications and market commentaries, the Perth metropolitan housing market is divided, in spatial terms, into 22 sub-regional areas. In summary, our sample contains 272 individual suburb specifications, the general basis for location description. From these suburbs, REIWA aggregates suburbs into 22 sub-regional sub-market areas. In addition, we test an alternate specification of 145 individual postcode specifications in model 4 later in the empirical study.
where, Pi is the natural log of the transaction price of the ith dwelling in sub-region j; and
Models 3 and 4 represent full random coefficients multilevel estimations in which all continuous hedonic variable parameters are permitted to vary spatially, between pre-defined spatial units. In model 3, we adopt the sub-regions or potential spatial sub-markets defined by REIWA. In model 4, we adopt postcodes, a considerably smaller unit of geography.
where,
The multilevel models are estimated as linear mixed models containing both fixed and random effects, requiring estimation using restricted maximum likelihood. 1 In practical terms, the estimation approach allows the decomposition of residuals to reveal random intercepts and hedonic slope parameters that are specific to each defined spatial area. A city-wide intercept and set of hedonic parameters are estimated as fixed effects and take the place of the ‘normal’ hedonic coefficients obtained from a linear regression. For a given observation, the predicted price can be obtained by multiplying out the physical attributes with the city-wide coefficients and summing with the product of attributes and the coefficients or random effects specific to the spatial area in which the dwelling is located.
The estimations are carried out using a two-year sample of housing transactions in the Perth metropolitan area, Western Australia. A period extending from the beginning of 2007 to the end of 2008 was chosen as a study period after a preliminary analysis (not reported in this paper) to determine a period of relative stability in the Perth metropolitan housing market. Perth is the capital and largest city in Western Australia and the fourth most populous city in Australia. In March 2011, the population was estimated at approximately 1.7 million persons (ABS, 2011). Perth is predominantly a monocentric city with a dominant central business district as the centre of a well-developed rail-based public transport system. The city extends along the coast of the Indian Ocean around the central Swan/Canning River systems. The greater metropolitan area spans approximately 80 km from its northern to southern extremities and extends approximately 70 km east of the CBD.
The hedonic data used for the estimations in this study were supplied on licence by Landgate, the Western Australian Land Information Authority. These data benefit from considerable detail in terms of hedonic attributes. There are dummy variables describing the presence of ensuite and other bathrooms, dining, family, living and games rooms as well as swimming pool and study or home office variables. Additional variables describe wall and roof construction, location and property age. However, previous empirical work involving this particular dataset has highlighted significant collinearity between many of these attribute variables and location. This is, of course, a common problem in hedonic studies, but it is particularly problematic in the context of this study since the main empirical objective is to construct possible spatial sub-markets from smaller geographical building-blocks. The hedonic analyses therefore focus on a reduced set of explanatory variables.
4. Estimation Results
4.1 The Benchmark Model
Descriptive statistics for the explanatory variables are provided in the rightmost three columns of Table 1. The empirical performance of the city-wide hedonic model is, unsurprisingly, relatively poor. The adjusted R2 value is 0.40 (see Table 2), although almost all of the physical attribute variables are statistically significant at the 1 per cent level. The ‘house’ property type is implicit in the constant. Distance from the CBD is significant at 1 per cent and negative, in line with prior theoretical expectations. Despite the careful choice of a comparatively stable study period, the time dummy variables indicate significant variation in transaction prices from the base period.
City-wide hedonic model with distance variable
Notes: ** indicates significant at the 5 per cent level; *** significant at the 1 per cent level.
Descriptive statistics: sub-region model coefficients
4.2 Models Segmented by Real Estate Agents’ Sub-markets
Intuitively, we might expect stronger empirical performance of hedonic models estimated with a large dataset when we move from a single, city-wide model to a set of separate estimations formed by disaggregating to smaller spatial units. This is precisely what becomes evident when the data are segmented by the sub-regional units defined by REIWA, noting that the disaggregated models still include distance from the CBD as an explanatory variable. The sub-regional spatial units defined by REIWA are illustrated in Figure 1 and the estimation results are summarised in Table 2.

Perth, Western Australia: sub-markets as defined by REIWA.
For two sub-regions, adjusted R2 values are below 0.50 (Wanneroo North East and Wanneroo North West). In the other 20 cases, adjusted R2 values range from 0.54 to 0.87. The sample sizes range from 1144 (Fremantle) to 5313 (Rockingham). Table 2 sets out descriptive statistics for the sub-region model coefficients. The magnitudes of standard deviations relative to their respective means suggests remarkable stability for the intercepts and for coefficients on bedrooms, total number of rooms, car parking, land area and the ratio of bathrooms to bedrooms. There is much more variation in the coefficients for property-type variables, as might be expected given that their incidences are likely to exhibit strong spatial patterns. Perhaps most interesting is the evident variation in the coefficient of distance from the CBD (the standard deviation is more than twice the size of the mean). The descriptive statistics therefore give some mixed messages. There is evidence of some variation in slope parameters for ubiquitous attributes and much stronger evidence for those that are likely to cluster spatially.
4.3 The Multilevel Models
As described in the previous section, we estimated two full random effects multilevel models. In the first model, hedonic attribute parameters are estimated on a city-wide basis (as fixed effects) and a full specification of random effects allows estimation of differences in slopes between different spatial units in the metropolitan area. In the first model, we use the REIWA defined sub-regions as the second of the two levels. In the second model, we use postcodes, a more fine-grained spatial unit of geography.
Given the volume of estimation results, providing a summary is challenging. We proceed by summarising the fixed effects and model fit statistics for each of the two models in Table 3. In Table 4, we provide a set of descriptive statistics for the estimated random effects. In other words, the coefficients shown in Table 3 are analogous to hedonic coefficients from an OLS estimation and the descriptive statistics give an indication of the variation in these (the adjustment resulting from application of the random effects) between the defined spatial units.
Multilevel model: fixed effects and model fit statistics (N = 60 699)
Notes: * indicates significant at the 10 per cent level; ** significant at the 5 per cent level; *** significant at the 1 per cent level.
Multilevel model: estimated random effect statistics
The ‘townhouse’ property type is not significant in the first multilevel model, but is significant at 1 per cent in the second. The ‘terrace’ property type variable is significant only at 10 per cent in the first model, but is significant at 1 per cent in the second. The results are supportive of the idea that spatial aggregation in the presence of spatially varying attribute parameters gives rise to misleading results. In the second model, the specification permits estimation of attribute parameters for much smaller spatial units. One of the benefits is that the city-wide parameter estimates appear to be more stable.
For most of the other variables, the results are generally stable between the two multilevel estimations. Differences in parameter estimates seem to affect primarily the property type variables. Interestingly, the coefficient on distance from the CBD is very stable between the two estimations. This result may appear surprising given its apparent instability in the earlier hedonic estimations (models 1 and 2). In both cases, the LR and Wald chi-squared tests suggest strong explanatory power, but at this stage little more can be said about the relative performance of models 3 and 4 given that the likelihood ratios cannot be compared directly (since model 4 is specified with a greater number of random effects parameters).
Given the impracticality of presenting coefficients for all defined spatial units, Table 4 summarises the mean and standard deviation of the estimated random effects. As discussed in the previous section, these can be interpreted as location-specific differences in attribute parameters (compared with the corresponding city-wide coefficients). The descriptive statistics in Table 4 reveal instability in parameter estimates between the two estimation approaches. In particular, in moving from a multilevel model with relatively large spatial units (REIWA sub-regions), to that with smaller spatial units (postcodes) reveals that the mean and standard deviation of the random effects differ noticeably for the ‘group house’, ‘villa’, ‘home unit’ and ‘flat’ property-type variables. The random effects for land area and total number of rooms also appear to have much more variation in the second model than the first. The more granular of the two models allows greater variation in these property-type and lot-size effects to emerge. This is indicative of the sort of aggregation bias observed in house price models with few spatial variables or that omit sub-markets from the specification.
4.4 Predictive Performance of the Models
We now turn to the predictive performance of the five models examined in the empirical analysis. Table 5 summarises the mean, standard deviation, lower and upper quartile prediction errors. The figures are percentages.
Predictive accuracy of the models
Note: figures are percentage prediction errors; for example, model 1 mean is −3.28 per cent.
The figures show an improvement in predictive accuracy between models 1 and 2 (the city-wide and sub-region models respectively). The first multilevel model (model 3), with larger spatial units, has poorer predictive power than the sub-region models. This is interesting because, of course, models 2 and 3 are conceptually similar, despite the different estimation approaches. Both models are designed to allow hedonic parameters to vary between sub-regional spatial units as defined by real estate agents. While model 2 achieves this through separate estimation of the hedonic model for each spatial unit, model 3 does so through a combination of city-wide effects and sub-regional effects. On the basis of predictive power, the multilevel approach used for model 3 appears to be less efficient, achieving slightly lower predictive accuracy than a simpler segmented OLS model. However, the second multilevel model (model 4) has superior predictive accuracy in comparison with the other three models. Mean prediction error is -1.61 per cent with a standard deviation of 19.18 per cent. Figures 2 and 3 depict the spatial patterns of predictive error. Specifically, postcode units shown in light grey are those in which predictions are within 20 per cent of the observed value for at least 75 per cent of observations. Postcodes shaded in the medium tone are those in which predictions are within 20 per cent of the observed value for at least 50 per cent (but less than 75 per cent) of the observed value. Therefore, units shown in the darkest tone are those for which the model was not able to predict within 20 per cent of the observed value for more 50 per cent or more of the observations in that spatial unit. Although the choice of these predictive performance criteria is arbitrary, the same criteria are applied to each of the four model estimations to aid comparison.

Predictive accuracy of models 1 and 2.

Predictive accuracy of models 3 and 4.
The progressive improvement between models 1 and 2 and between 2 and 3 are evident visually. Similarly, the lower incidence and more random spatial distribution of large prediction errors (i.e. suburbs in which at least 50 per cent of properties’ transactions prices are not predicted within 20 per cent of their observed values) is evident in a comparison of models 1 and 4 (the worst and best in terms of predictive power). However, it is also notable that even the best empirically performing model leads to a spatial pattern of prediction errors that is not random. Transacted properties in waterfront locations, either facing the Indian Ocean or the substantial frontage of Swan River, are associated with a much higher incidence of high prediction error. We formalise these suggestive, but impressionistic, results by presenting a series of tests for spatial autocorrelation in Table 6. The first two rows set out the results of so-called global tests for spatial autocorrelation—specifically, Moran’s I and Geary’s c coefficients. The null hypothesis is of no spatial autocorrelation in the residuals of each of the four model estimations. 2 In the case of Moran’s I, the test results clearly show that the null hypothesis is rejected for models 1 through 3, but not for model 4. For Geary’s c, we fail to reject the null for all four models. Turning to the local tests for spatial autocorrelation, reporting detailed results for all 272 spatial units would be impractical and so we summarise the results in the bottom section of Table 6 by summing the number of spatial units (postcodes) for which the null hypothesis of no spatial autocorrelation is rejected. The results clearly show that spatial autocorrelation affects fewer spatial units as we move from models 1 through 4, and that the multilevel models offer a considerable advantage from this perspective.
Tests for spatial autocorrelation in residuals
Notes: *** denotes significant at 1 per cent; ** at 5 per cent; there are 272 local spatial units (suburbs) in total. Figures in the last two rows indicate the number of suburbs for which the null hypothesis of no spatial autocorrelation is rejected. The results are summarised at the 5 per cent and the 1 per cent significance level for the Moran’s I and Geary’s c statistics.
5. Conclusions
This paper set out to examine the utility of applying multilevel strategies to modelling spatial housing sub-markets. The empirical analysis was designed to compare the predictive performance of several models. Setting up a simple, city-wide OLS hedonic model, with a simple distance variable, allowed us to establish a basic benchmark with which to compare several alternative approaches. We found that separate estimation of the hedonic models for potential sub-markets (or sub-regions) defined by real estate agents led to a model that was superior to the benchmark in terms of predictive power.
Estimation of multilevel hedonic models also led to improvement beyond the benchmark OLS model. Here, however, the predictive performance for one of the models (model 3) was slightly below that of the sub-region models on average. This is an important and interesting finding in that it implies that a spatially segmented OLS estimation approach is acceptable when there is certainty about spatial sub-market boundaries, and would be appropriate when sub-market boundaries are expected to remain stable over time. However, as we argue earlier, the best modelling strategy is not just that which produces the greatest predictive accuracy. The best approach to modelling sub-markets should use empirically based sub-markets as a starting-point but should also recognise that the (at least some) sub-market boundaries will be subject to modification over time. Thus, an effective modelling strategy should also be able to capture the fluidity in sub-market dimensions. Where the spatial extent of sub-market boundaries is less certain, or when there is an expectation of change in these boundaries over time, our analysis suggests that a multilevel approach, with more finely grained spatial units of geography, may be preferable to a segmented approach that imposes harder boundaries.
Our second multilevel model (model 4), defined with smaller spatial units, exhibits predictive performance that exceeded all other estimation approaches examined in this paper. The spatial pattern of prediction errors is less concentrated. The results imply that multilevel models have the capacity to improve predictive power and reduce spatial dependence when compared with standard hedonic methods. In addition, when applied over time, multilevel models appear better able to deal with dynamic change in the composition of sub-markets and have the potential to capture the multiple (often nested) geographies that exist within local housing systems. They can be used effectively to differentiate between neighbourhood, local (sub-regional) and regional influences. More applied research is required, however, to exploit this potential.
Footnotes
Acknowledgements
The authors gratefully acknowledge Landgate and the Valuer General of Western Australia for providing the data used in this research.
Funding
This work was supported by the Carnegie Trust for Scotland and the Curtin University Visiting Research Fellowships programme.
