View Categories

Cohen’s kappa

Jacob Cohen invented kappa as a way of describing how good agreement was between two raters who were each using some categorical rating (Cohen, 1961). A kappa of zero indicates purely random, no better than chance, agreement between the raters and a kappa of 1.0 indicates perfect agreement. Kappa can be used for any categorisation and it is a statistic describing the sample of things rated, e.g. therapist styles, therapy “ruptures”. Where agreement is really disagreement, kappa can take values below zero.

Details #

The reason for agreement indices like kappa is that simple agreement rates can look very good even with purely chance agreement if the ratings have one category that is very common, and seen to be very common by both raters even though they have zero agreement on when they see it. See my PSYCTC.org blog post: Why kappa? or How simple agreement rates are deceptive which has pointers to other information, including to my own Rblog post which goes into more of the technicalities using R.

Most illustrations, including mine (above), look at binary, yes/no, ratings as that’s easiest to illustrate but kappa can handle any number of categories and there are extensions to weight the seriousness of disagreement (“weighted kappa”) where there are multiple categories (i.e. more than two) and when they may have some ordinal relationships between the categories. An example is AAI (Adult Attachment Interview) codings a higher weight might be attached to one rater rating the interview “autonomous” and the other “dismissing” than might be given to the disagreement in which the first rater classifies an interview as “Unresolved/Disorganized” and the second rater classifies it as “dismissing” (as there is an overarching distinction in AAI categories between “autonomous” and the other, non-autonomous categories).

There are also extensions of kappa for the situations in which there are more than two raters and there are also alternative agreement indices with claims to some advantages over kappa. There are also quite complicated explorations of how we should move from a kappa value that is a simple summary of the numbers in your dataset of ratings, to treat that kappa as an estimate of some population kappa treating your dataset as a random sample from that population. In the Null Hypothesis Significance Testing paradigm (NHST) that theory can give you a p value for the improbability that you would have seen a kappa as strong or stronger than you did were your dataset a sample from an infinite population. In the estimation paradigm that theory can give you a confidence interval (CI) around your observed kappa. As ever with the NHST / estimation models of sampling entirely at random from infinitely large populations, these p values or CIs should be treated very cautiously as our data are so rarely actually random samples from infinite populations of interest. The NHST paradigm here is particularly unhelpful as the null hypothesis of zero agreement between raters is not what interests us, what matters is how good the agreement is. There is a particular issue with the sample / population model here that we need to think whether to treat our raters as a sample from an infinite population of raters as well as to treat the set of things rated as a sample.

Despite all the issues in that last paragraph, kappa, weighted or not, for two raters or more, for a binary or a multiple option classification, remains the agreement index you are most likely to meet in our field and, if not overvalued and if one thinks carefully about the rating method used, the particular raters and the particular set of objects that were rated kappa remains a genuinely useful descriptive statistic.

Try also #

Chapters #

Kappa is mentioned in Chapter 3.

Online applications #

My Rblog post: “Chance corrected agreement“

Reference #

Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20, 37–46.

History #

Created 26.i.22, text extended and improved and more links added 27.ix.26.

Powered by BetterDocs