Abstract
Singer was the first to draw attention to the fact that the Pareto Law (originally employed to describe income inequality) could also be applied to the size distribution of urban centres. His argument was illustrated with evidence from seven nations at various points in time. As part of his study, Singer attempted to demonstrate two propositions relating to the Pareto distribution of centre size. One was that the populations of larger centres could be described by the Pareto distribution, while the other was that the populations of the smaller (non-urban) centres would also conform to this distribution. It is argued that neither proposition can be regarded as valid.
Introduction
Application of the Pareto (1896) Law of income distribution to the size distribution of urban centres was first undertaken in a widely cited paper by Singer (1936), although this development is frequently attributed to Zipf (1949). Singer was critical of the conventional “index of urbanisation”, which indicated the percentage of a nation’s population residing in urban centres. He argued that from various economic standpoints “A town of 500,000 inhabitants is …. something different from ten towns of 50,000 inhabitants” (Singer, 1936: 254). When used for the size distribution of centres the Pareto Law measures, the relative importance “of the smaller and larger types of human agglomerations” (Singer, 1936: 254). In this way, Singer saw his courbe des populations as the counterpart of Pareto’s courbe des revenus. An interesting, though neglected, feature of Singer’s paper was his contention that the larger centres of the system, as well as rural or non-urban centres, would both conform to the Pareto distribution. Such assertions, which involve both tails of the distribution, form the basis of this note. Initial attention is given to Singer’s approach.
The pareto distribution
Singer’s test of Pareto’s Law involved the urban systems of seven nations at various dates. The data employed were in the form of sizes classes of centre, taken mostly from official published sources. One such size class would be 5000–9999, with 5000 as the lower limit. For the 23 cases considered, the number of size classes varied between 4 and 14. This variation reflected the conventions employed by the different national statistical offices at various times. The Pareto distribution has the following form
For each of the 23 cases examined, Singer (1936: 255–258, 261) employed a least-squares regression to derive the best fit for equation (1). The number of observation points in each case varied between 4 and 14. Each observation point was based on x and y, as defined above. It was thus possible to determine the expected cumulative frequency (or number of centres greater than x) for each size class. The goodness of fit was determined by the following error term E
The graph of equation (1) is referred to as the “population curve”, in keeping with Singer’s terminology. An example of this downward-sloping regression line is shown in Figure 1, where K = 0.31 m and α = 0.84. Such a line is generally valid for centres with populations above 2000. This is the minimum population of urban centres, an approximate figure that is used officially in the majority of nations. Equation (1) does not apply to the largest centre, which cannot be plotted in Figure 1. This has been an inconvenient feature of the Pareto distribution when applied to urban centres, and one that is addressed in “Larger centres of the distribution (the upper tail)” section.

Graph of the Pareto distribution.
Smaller centres of the distribution (the lower tail)
The second of Singer’s two concerns (to be considered first in this note) was with the population curve below 2000 in Figure 1. Most of Singer’s examples involved a cut-off size of approximately 2000, the threshold population of urban centres (the right-hand dotted line in Figure 1). Nevertheless, he suggested that equation (1) would be relevant for centres with populations as low as 200 (the left-hand dotted line in Figure 1) which was regarded as “the lower limit of human agglomerations” (Singer, 1936: 263). He then sought to determine the number of centres with populations greater than 200. Equation (1) was solved for x = 200 with K and α known in each of four cases (England and Wales, Germany, USA and France)
The difficulty here was that although the expected value for y (cumulative frequency) could be determined, there was no reliable observed value with which it could be compared. In an attempt to fill this void, recourse was made to data on frequencies that were drawn from certain published sources: an administrative atlas in the case England and Wales, and data from a yearbook in the case of Germany. In both cases, the fit was questionable, and no similar sources were presented either for the USA or France. It becomes apparent that for the want of adequate empirical data, this attempt at showing the conformity of non-urban centres to the Pareto distribution was impossible to sustain.
In addition, studies have shown that below a population of around 2000 (the cut-off size for urban centres in many nations), the linearity of the population curve does not continue, and size decreases at a decreasing rate, as shown by the broken line in Figure 1. This points to the existence of a cumulative lognormal distribution (Aitchison and Brown, 1957; Eeckhout, 2004; Fazio and Modica, 2015; Gibrat, 1931; Parr and Suzuki, 1973; Reggiani and Nijkamp, 2015). An indication of this is apparent from Singer’s data. If we consider the smallest size class in the various cases, we find that the observed value for frequency was significantly below the expected value in around 60% of the cases, suggesting that the Pareto distribution is not relevant over this range of centres. This is consistent with the view that size distribution of rural centres differs from that of urban centres (Sonis and Grossman, 1984). In either case, Singer’s argument that the lower range of centres would conform to the Pareto distribution is not valid.
Larger centres of the distribution (the upper tail)
The other concern of Singer (1936: 262) involved the claim that the larger centres of the system were consistent with the Pareto distribution. In the data presented, these larger centres were grouped together within a single size class, so that the individual centres could not be identified. For example, in the data for Germany from 1880 to 1933, the highest size class contained fewer than 10 centres in only one of six cases.
Singer’s approach to the populations of larger centres
Singer (1936: 262) argued that the point where the population curve intersected the x-axis would be “that number of inhabitants which is exceeded by only one town, i.e. a population intermediate between that of the largest and that of the second town.” This intermediate population xint was defined as follows
It was then shown that in four cases (one for each nation) the intermediate population lay between the official populations of the largest and second centres. This approach was unsatisfactory on two counts, however. The intersection of the population curve with the x-axis must necessarily be at the population of the second centre and not at some intermediate population. In addition, no attempt was made to formulate expressions for the populations of the largest centre and the second centre, by which adherence to the Pareto distribution could have been tested. Given the approach adopted by Singer, the claim that the larger centres conformed to the Pareto distribution cannot be sustained.
Up to now the largest centre has been excluded from explicit consideration. This is because cumulative frequency involves the number of centres greater than x. In the case of the largest centre, its population cannot be exceeded, so that this centre is excluded from Figure 1. In applications of the Pareto distribution to such phenomena, as income levels or sizes of industrial plants the issue is not especially serious. The magnitudes of highest income or largest plant are of limited significance, so that undertaking the analysis in terms of size classes suffices. By contrast, the population of the largest centre (along with other major urban centres) is of major importance in studies dealing with the relationships between urban structure and economic development (Berry and Kasarda, 1977; Mills and Song, 1979).
A modification of the pareto distribution
Starting with data in the form of size-classes, it is possible to estimate the populations of the larger centres. Figure 1 is no longer relevant. The approach begins with the construction of the regression line or population curve. Each observation point refers to x (the population of the lower limit of a size class), but the definition of cumulative frequency of x is now r, the rank or number of centres with populations greater than or equal to x. This change involves alterations in both axes, as indicated in Figure 2. The r-axis replaces the y-axis of Figure 1 (note that r always exceeds y by one), while the r = 2 abscissa of Figure 2, indicated by the dotted line, replaces the x-axis of Figure 1. Below the r = 2 abscissa is the r = 1 abscissa or the x-axis of Figure 2. Since in this example r is used to indicate cumulative frequency, the intercept C is the expected total number of centres, and the expected slope is β. In Figure 2, C = 1 m and β = 0.9.

Graph of the modified Pareto distribution.
This modified Pareto distribution is written as
In the case of the largest centre r = 1, so that its population is
In contrast to the Singer’s approach, the modification of the Pareto distribution outlined here permits the populations of the largest and second centres to be estimated, along with other centres belonging to the highest size class. Over recent decades, this modification of the Pareto distribution has come to replace the original formulation; see, for example, Eeckhout (2004); Rosen and Resnick (1980); Zipf (1949). With the axes transposed, the modified version of the Pareto distribution is equivalent to the rank-size distribution (Lotka, 1925: 306–307). However, a case can be made for regarding each distribution as distinct (Parr, 2021).
Closing comment
Singer’s application of the Pareto Law to the size distribution of centres was an important innovation which contributed to our understanding of centres systems and their development over time (Allen, 1954). Having demonstrated the similarity of the size distribution of centres to the Pareto Law, Singer went on to consider conditions at the two tails of the distribution. At the lower tail, the possibility was investigated that the smaller (non-urban) centres would conform to the Pareto distribution. For the upper tail, the possibility of centre populations conforming to the Pareto distribution was also considered. In the former case, Singer was unable to show convincingly that the smaller centres adhered to the Pareto distribution, and the possibility that an alternative distribution might have been more appropriate was not entertained. In the latter case, Singer failed to formulate the populations of the larger centres, and was thus unable to demonstrate their conformity to the Pareto distribution. It should be emphasised that Singer’s unsatisfactory findings reported here do not diminish the worth of his general contribution.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
