The Kappa paradoxes are usually framed as a problem with interpreting Kappa's magnitude. This post argues they cause a second, quieter problem: they can make it mathematically impossible to plan a study's sample size around Kappa.
The Kappa paradoxes — where Kappa can come out surprisingly low even when raters agree on the large majority of subjects — are well documented, going back at least to Feinstein and Cicchetti's 1990 discussion of the issue. Most of that literature focuses on how to interpret a low Kappa value in a specific study.
This post looks at a different angle: sample size planning. A common question researchers ask before running a reliability study is, "how many subjects do I need so that my coefficient's standard error stays below some target?" For most statistics, adding more subjects reliably shrinks the standard error toward zero. Kappa-family coefficients — Cohen's Kappa, Conger's Kappa, Fleiss' generalized Kappa, and Krippendorff's alpha all behave similarly here — don't follow that rule.
Why more subjects doesn't guarantee more precision
The post derives the maximum possible variance of Fleiss' generalized Kappa as a function of sample size, and shows that this maximum doesn't shrink to zero as the number of subjects grows — it approaches a fixed floor instead. For the two-rater case specifically, that floor works out to a standard error just above 0.31, no matter how many subjects are added. In other words, there will always exist some pattern of ratings that produces a standard error exceeding 0.3, regardless of sample size.
The practical implication: if a study is planned around achieving a target standard error for a Kappa-type coefficient, that target may simply be unreachable for certain rating patterns — a limitation that has nothing to do with how many subjects were recruited. This is one more reason, alongside the more familiar prevalence and marginal-homogeneity paradoxes, to consider coefficients like Gwet's AC1 that don't share Kappa's instability.
Want the full technical detail? The derivation of the variance bound and its constants for different numbers of raters and categories, along with the full bibliography, are on the original Blogger post.
Read the full post on Blogger ↗