Twelve variance-analytical and non-parametrical coefficients of reliability for rating scales designed for rating persons, are being compared to each other theoretically and empirically. The empirical comparison is produced with the help of over 20.000 ratings, which 54 subjects delivered in two, antagonistical as far as their reliability was concerned, test instructions. The preference for two coefficients ( rkrand and rkfix') has been established. The ICC appears to be useful for the estimation of reliability as well.
Get full access to this article
View all access options for this article.
References
1.
Acock, A. C. and Martin, J. D.The undermentioned controversy: Should ordinal data be treated as interval?Sociology and Social Research, 1974, 58, 427 -433.
2.
Anderson, N. H.Scales and statistics: Parametric and nonparametric. Psychological Bulletin, 1961, 58, 305 -316.
3.
Bartko, J. J.The intraclass correlation coefficient as a measure of reliability. Psychological Reports, 1966, 19, 3 -11.
4.
Bartko, J. J.Corrective note to: 'The intraclass correlation coefficient as a Measure of reliability.'Psychological Reports, 1974, 34, 418.
5.
Bintig, A.Über die Zuverlässigkeit sozialwissenschaftlicher Beurteilungsskalen bei der Wahrnehmung und Beurteilung von Personen. Frankfurt , 1978. (On the reliability of rating scales of the perception and judgement of persons.)
6.
Boneau, C. A.The effects of violations of assumptions underlying the t-test. Psychological Bulletin , 1960, 57, 49-64.
7.
Brown, B. W., Jr., Lucero, R. J. and Foss, A. B.A situation where the Pearson correlation coefficient leads to erroneous assessments of reliability. Journal of Clinical Psychology , 1962, 18, 95 -97.
8.
Brunswick, E.Perception and the representative design of psychological experiments. Berkeley: University of California Press, 1956.
9.
Cicchetti, D. C., Aviano, S. L., and Vitale, J.Computer programs for assessing rater agreement and rater bias qualitative data. EDUCATIONAL AND PSYCHOLOGICAL MEASUREMENT, 1977, 37, 195 -202.
10.
Cohen, J.A Coefficient of agreement for nominal scales. EDUCATIONAL AND PSYCHOLOGICAL MEASUREMENT, 1960, 20, 37 -46.
11.
Cohen, J.Weighted kappa: Nominal scale agreement with provision for scaled disagreement or partial credit. Psychological Bulletin, 1968, 70 , 213-220.
12.
Cohen, J.Weighted chi-square: An extension of the kappa method. EDUCATIONAL AND PSYCHOLOGICAL MEASUREMENT , 1972, 32, 61-74.
13.
Cohen, R.Systematische Tendenzen bei Persönlich-keits-beur-teilungen . Bern: Huber, 1969.
14.
Ebel, R. L.Estimating the reliability of ratings. Psychometrica, 1951, 16, 407- 424 .
15.
Everitt, B. S.Moments of the statistic kappa and weighted kappa. British Journal of Mathematical and Statistical Psychology, 1968, 21, 97-103.
16.
Finn, R. H.A note on estimating the reliability of categorial data. EDUCATIONAL AND PSYCHOLOGICAL MEASUREMENT , 1970, 30, 71-76.
17.
Fleiss, J. L., Everitt, B. S., and Cohen, J.Large sample standard errors of kappa and weighted kappa. Psychological Bulletin, 1969, 72, 323-327.
18.
Guilford, J. P.Psychometric methods. New York: McGraw-Hill, ( 2nd ed.), 1954.
19.
Hays, W. L.Statistics. London: Holt-Rinehart & Winston, 1969.
20.
Koslowsky, M. and Bailit, H.A measure of reliability using qualitative data. EDUCATIONAL AND PSYCHOLOGICAL MEASUREMENT, 1975, 35, 843-846 .
21.
Krech, D., Crutchfield, R. S. and Ballachey, E. L.Individual in society. New York: McGraw-Hill, 1962.
22.
Lord, F. M.On the statistical treatment of football numbers. American Psychologist, 1953 , 8, 750 -751.
23.
Lu, K. H.A measure of agreement among subjective judgments. EDUCATIONAL AND PSYCHOLOGICAL MEASUREMENT , 1971, 31, 75-84. (a)
24.
Lu, K. H.Statistical control of "impurity" in the "estimation of test reliability."EDUCATIONAL AND PSYCHOLOGICAL MEASUREMENT, 1971, 31, 641-655. (b)
25.
Mann, L.Social psychology. Sidney: Wiley, 1969.
26.
Maxwell, A. E.The effect of correlated errors on estimates of reliability coefficients. EDUCATIONAL AND PSYCHOLOGICAL MEASUREMENT, 1968, 28 , 803-811.
27.
Maxwell, A. E. and Pilliner, A. E. G.Deriving coefficients of reliability and agreement for ratings. British Journal of Mathematical and Statistical Psychology, 1968, 21, 105-116.
28.
Overall, J. E.Estimating individual rater reliability from analysis of treatment effects. EDUCATIONAL AND PSYCHOLOGICAL MEASUREMENT, 1968, 28 , 255-264.
29.
Siegel, S.Nonparametric statistics for the behavioral sciences. New York: McGraw-Hill, 1956.
30.
Silverman, F. H.Correspondence between mean and median scale values for sets of stimuli scaled by the method of equal appearing intervals. Perception and Motor Skills, 1967, 25, 727 -728.
31.
Silverman, F. H.Intraclass correlation coefficient as an index of reliability of median scale values for sets of stimuli rated by equal-appearing intervals. Perception and Motor Skills , 1968, 26, 878.
32.
Stevens, S. S.Measurement, statistics and the schemapiric view. Science, 1968, 161, 849-856 .
33.
Tagiuri, R.Person perception. In Lindzey, G. and Aronson, E. (eds.) The handbook of social psychology. Vol. 3 . Reading, Massachusetts : Addison-Wesley , 1969.
34.
Taylor, J. B.Rating scales as measures of clinical judgment: A method for increasing scale reliability and sensitivity . EDUCATIONAL AND PSYCHOLOGICAL MEASUREMENT, 1968, 28, 747 -766.
35.
Torgeson, W. S.Theory and methods of scaling . New York: Wiley, 1958.
36.
Winer, B. J.Statistical principles in experimental design. New York: McGraw-Hill, 1962.