Executive Summary
Kappa 0.621 (substantial) on 300 items
Your two reviewers picked the same category on 72.0% of 300 items. Cohen's kappa is 0.621 (95% CI 0.551 to 0.690), which the Landis & Koch convention calls substantial agreement and is clearly better than chance (p < 0.001). Chance alone would have produced 26.2% agreement, leaving 73.8% available above that floor, of which 62.1% was achieved. Because the categories are ordered (Poor, Fair, Good, Excellent), near misses earn partial credit: quadratic-weighted kappa is 0.830. The single most common disagreement was Reviewer A choosing "Good" while Reviewer B chose "Fair" (15 items, 5.0%). Agreement was weakest on "Fair" at 64.90%. Remember: kappa measures whether the raters make the same call, not whether the call is right.
Analysis Overview
Cohen's kappa on 300 items judged by Reviewer A Rating and Reviewer B Rating across 4 categories.
Raw agreement of 72.0% alone cannot answer your question because raters who both lean on one category collide constantly without reading carefully. Cohen's kappa fixes this by subtracting the chance agreement that each rater's own habits would produce (26.2% here) and reporting how much of the remaining headroom (73.8%) was actually won: 62.1%. That subtraction is also kappa's weakness — a heavily skewed category mix can inflate chance agreement and make kappa look weak even when raw agreement is high. Here the category distribution is reasonably even, so the two numbers track closely: 72.0% observed, 26.2% by chance, 0.621 kappa.
Data Quality
300 items used of 300 rows loaded.
All 300 rows loaded carried judgements from both raters and were analysed; no incomplete pairs were dropped. Labels were trimmed of whitespace and blanks treated as missing. The 4-by-4 confusion matrix has 6 cells with fewer than 5 items, which means the large-sample confidence interval is only indicative rather than exact. No case variants needed merging and no rare categories were pooled. The data are clean and complete for agreement analysis.
Where the Two Raters Land
Every combination of Reviewer A Rating and Reviewer B Rating's choices, counted.
The diagonal holds 72.0% of the 300 items — that is agreement. The brightest off-diagonal cell is Reviewer A Rating = "Good" against Reviewer B Rating = "Fair" with 15 items (5.0%), so that pair of categories is where definitions need work. Because the categories are ordered along a quality scale, cells near the diagonal are near misses and far cells are serious ones. The largest asymmetric gap is on "Poor", used 1.0% more often by one rater than the other — a small systematic bias. No single category is driving all the disagreement; the pattern is spread across the scale.
Agreement Statistics
Percent agreement, chance agreement, kappa with its interval, and the weighting variants.
| Statistic | Estimate | CI Low | CI High | Interpretation |
|---|---|---|---|---|
| Raw percent agreement | 0.72 | — | — | The two raters chose the same category for 72.0% of the 300 items. |
| Agreement expected by chance | 0.2621 | — | — | What two raters with these same category habits would agree on by luck alone (26.2%). |
| Cohen's kappa | 0.6206 | 0.5511 | 0.69 | 62.1% of the agreement left available above chance was achieved — substantial on the Landis & Koch convention. |
| Weighted kappa (linear) | 0.7316 | 0.6776 | 0.7856 | Near-miss disagreements between adjacent categories count as partial credit, in proportion to how far apart they are. |
| Weighted kappa (quadratic) | 0.83 | 0.7867 | 0.8734 | The same idea with distance squared: near misses are forgiven much more, far misses punished much harder. Always the most flattering of the three. |
| Prevalence-and-bias-adjusted kappa | 0.6267 | — | — | A diagnostic, not a replacement: what kappa would be if every category were equally common and both raters used them equally often. A large gap from Cohen's kappa means skewed marginals are doing the work. |
Cohen's kappa is 0.6206 (95% CI 0.5511 to 0.6900), computed as (0.720 − 0.262) / (1 − 0.262) = 0.621. The confidence interval rests on the Fleiss-Cohen-Everitt standard error of 0.035, a large-sample approximation applied here to 6 thin cells (fewer than 5 items each), so treat the width as indicative. The prevalence-and-bias-adjusted kappa is 0.6267, within 0.01 of Cohen's kappa, meaning the category mix is not distorting the headline figure. Weighted kappa (quadratic) is 0.830, more forgiving of near misses because the categories are ordered. The significance test z = 18.33, p < 0.001, shows the agreement is far better than chance.
Which Categories They Fight Over
Per-category agreement between Reviewer A Rating and Reviewer B Rating.
Per-category agreement ranges from "Excellent" at 80.00% down to "Fair" at 64.90% — a spread of 15.10 percentage points. "Poor" stands at 74.8% and "Good" at 70.4%. No single category is carrying all the disagreement; the weakness is distributed across the scale. This detail matters because a single coefficient can hide that one rubric entry is the problem. Here the raters fight over "Fair" more than others, but the issue is not confined to that category.
How Each Rater Uses the Scale
Category shares for Reviewer A Rating and Reviewer B Rating side by side.
The two raters' category distributions are nearly identical. Reviewer A Rating uses "Good" 33.3% of the time, Reviewer B Rating 33.0%. "Fair" appears in 25.7% of both raters' judgements. "Excellent" is 23.7% for A and 23.0% for B. "Poor" is the largest gap at 1.0 percentage points (17.3% vs 18.3%). Because the marginals are balanced, the chance-agreement floor stays low at 26.2%, which keeps kappa close to what the raw agreement of 72.0% suggests. Neither rater is systematically favouring a category the other avoids.
Methods & Disclosure
Every formula behind the numbers, and what they cannot decide.
| Item | Detail |
|---|---|
| Design | 300 items each judged once by Reviewer A Rating and once by Reviewer B Rating, across 4 categories. |
| Percent agreement | Diagonal of the confusion matrix over the total: 72.0%. |
| Chance agreement | Sum over categories of (Reviewer A Rating's share) x (Reviewer B Rating's share): 26.2%. |
| Cohen's kappa | (observed - expected) / (1 - expected) = (0.720 - 0.262) / (1 - 0.262) = 0.621. |
| Standard error and CI | Fleiss-Cohen-Everitt asymptotic standard error, 0.035; the 95% interval is the estimate plus or minus 1.96 standard errors. It is a large-sample approximation, and 6 of the 16 matrix cells hold fewer than 5 items, so treat the interval as indicative. |
| Significance test | z = kappa / SE under the null of chance-level agreement = 18.33, p < 0.001. The null variance is a different quantity from the one behind the confidence interval. |
| Category order | Orderedness was detected, not assumed: the labels are all points on a quality scale, so they were ordered along it. |
| Weighting | Weighted kappa is reported because the categories are ordered — the labels are all points on a quality scale, so they were ordered along it. |
| Benchmark labels | Landis & Koch (1977) labels (slight / fair / moderate / substantial / almost perfect) are a naming convention with no theoretical basis; the acceptable level of agreement depends on what the judgement is used for. |
| What kappa is not | Kappa measures agreement, not accuracy: two raters can agree completely and both be wrong. It also describes these two raters on these items, and does not generalise to raters or items outside this set. |
The analysis is closed-form arithmetic on the 4-by-4 confusion matrix. Percent agreement is the diagonal divided by the total (216 / 300 = 72.0%). Chance agreement is the sum of (Reviewer A's share) × (Reviewer B's share) across all categories, yielding 26.2%. Cohen's kappa is (observed − expected) / (1 − expected) = (0.720 − 0.262) / (1 − 0.262) = 0.621. The significance test is z = kappa / SE = 0.621 / 0.035 = 18.33, p < 0.001, under the null of chance-level agreement. Categories were detected as ordered (all points on a quality scale), licensing weighted kappa. Two hard limits: the Landis & Koch bands are a naming convention with no theoretical basis, and kappa says nothing about whether either rater is right.
Methodology
Statistical methodology and diagnostics for Rater Agreement — Cohen's Kappa
Statistical Method
Standard-library analysis: two raters, one categorical judgement per item — how much do they really agree? Map the two judgement columns and get the full confusion matrix as a heatmap, raw percent agreement, the agreement chance alone would produce, Cohen's kappa with its standard error and 95% confidence interval, a test of kappa against chance, linear- and quadratic-weighted kappa when the categories turn out to be ordered (detected from the labels, never assumed), a per-category breakdown of exactly where the raters diverge, and — when skewed marginals are depressing kappa — the kappa paradox explained with your own numbers rather than reported as a bad score.
- Each row is one item, judged once by each of two raters
- Both raters chose from the same list of categories
- The two judgements are independent — neither rater saw the other's answer
- Items are independent of each other; the same item does not appear on several rows
- Kappa measures agreement, not accuracy — two raters can agree completely and both be wrong
- Kappa is depressed when one category dominates or when the raters use the categories at different rates; this is the kappa paradox, and the analysis reports it rather than hiding it
- The confidence interval is a large-sample approximation and is only indicative when the confusion matrix has thin cells
- Weighted kappa is offered only when the labels themselves reveal an order; ordered categories with opaque labels will be treated as nominal
Analysis Code
Complete R source code for this analysis
Rater Agreement — Cohen's Kappa
Two raters, one categorical judgement per item: how much do they agree, and how much of that agreement is more than chance would have produced? The analysis builds the full confusion matrix, reports raw percent agreement alongside Cohen's kappa with its standard error and 95% confidence interval, adds linear- and quadratic-weighted kappa when the categories turn out to be ORDERED (detected from the level names, never assumed), breaks agreement down category by category, and explains the kappa paradox with this dataset's own numbers whenever skewed marginals are depressing kappa below what the raw agreement suggests.
Why This Method?
Percent agreement on its own is not interpretable: two raters who both answer "Yes" to almost everything will agree 90% of the time without looking at a single item. Kappa subtracts the agreement chance alone would deliver given each rater's own habits, and expresses what is left as a fraction of what was achievable.
What This Analysis Covers
- The k x k confusion matrix as a heatmap, plus the largest disagreement
- Raw percent agreement, chance agreement, and Cohen's kappa with SE + CI
- A z-test of kappa against zero (no better than chance)
- Linear- and quadratic-weighted kappa when the categories are ordered
- Per-category agreement — which categories the raters actually fight over
- Each rater's marginal distribution, and the kappa-paradox explanation
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {rater_1, rater_2}. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Core Analysis Pipeline
Rule 1 — the labels are numbers (1, 2, 3 / 0-10 scales).
num <- suppressWarnings(as.numeric(clean))
if (!anyNA(num) && length(unique(num)) == length(num)) {
return(list(ordered = TRUE, order = levs[order(num)],
rule = "the labels are numbers, so they were ordered by value"))
}Rule 2 — the labels start with a number ("1 - Poor", "2 = Fair").
lead <- suppressWarnings(as.numeric(sub("^\\s*([-+]?[0-9]+(\\.[0-9]+)?)\\s*[-=.):|].*$",
"\\1", clean)))
has_lead <- grepl("^\\s*[-+]?[0-9]+(\\.[0-9]+)?\\s*[-=.):|]", clean)
if (all(has_lead) && !anyNA(lead) && length(unique(lead)) == length(lead)) {
return(list(ordered = TRUE, order = levs[order(lead)],
rule = "every label starts with a number, so they were ordered by that number"))
}Rule 3 — the labels are all drawn from one recognised ordinal vocabulary.
scales <- list(
"an agreement scale" = c("strongly disagree", "disagree", "somewhat disagree",
"slightly disagree", "neutral", "neither agree nor disagree",
"slightly agree", "somewhat agree", "agree", "strongly agree"),
"a quality scale" = c("very poor", "poor", "below average", "fair", "average",
"satisfactory", "good", "very good", "excellent", "outstanding"),
"a frequency scale" = c("never", "rarely", "seldom", "sometimes", "occasionally",
"often", "frequently", "usually", "always"),
"a severity scale" = c("none", "minimal", "mild", "moderate", "severe",
"very severe", "extreme", "critical"),
"a magnitude scale" = c("very low", "low", "medium", "moderate", "high", "very high"),
"a satisfaction scale" = c("very dissatisfied", "dissatisfied", "neutral",
"satisfied", "very satisfied"),
"a likelihood scale" = c("very unlikely", "unlikely", "possible", "likely",
"very likely", "certain"),
"a priority scale" = c("trivial", "minor", "moderate", "major", "critical", "blocker")
)
low <- tolower(clean)
if (length(levs) >= 3 && !any(duplicated(low))) {
for (nm in names(scales)) {
sc <- scales[[nm]]
if (all(low %in% sc)) {
return(list(ordered = TRUE, order = levs[order(match(low, sc))],
rule = paste0("the labels are all points on ", nm,
", so they were ordered along it")))
}
}
}
list(ordered = FALSE, order = levs,
rule = "no numeric values, numeric prefixes, or recognised ordinal wording were found in the labels")
}
compute_shared <- function(df, params, col_map = list()) {
# === SHARED EXPORTS ===
# initial_rows/final_rows/rows_removed $ row accounting (items)
# r1_h / r2_h $ humanized names of the two mapped columns
# n_items / k_cats $ usable items, categories analysed
# levs $ character — categories in analysis order
# n_dropped / n_case_merged / n_lumped $ cleaning counts
# case_example $ character(2) — the merged spelling pair, or NULL
# ordered_flag / order_rule $ orderedness verdict + why
# p_o / p_e / kappa / kappa_se / ci_low / ci_high / z_stat / p_value
# kappa_linear / kappa_quadratic (+ _se) $ NA when categories are nominal
# band $ Landis & Koch label for the unweighted kappa
# pabak $ prevalence/bias-adjusted kappa (diagnostic)
# paradox $ logical — skewed marginals are depressing kappa
# max_prev / prev_cat / max_bias / bias_cat
# confusion_df / kappa_df / category_df / marginal_df / methods_df
# top_off_* $ the single largest disagreement cell
# n_sparse $ confusion cells holding fewer than 5 items
# metrics / json_output
# === /SHARED EXPORTS ===
r1_h <- humanize_semantic("rater_1", col_map)
r2_h <- humanize_semantic("rater_2", col_map)Step 1: Validate the mapped columns
initial_rows <- nrow(df)
for (req in c("rater_1", "rater_2")) {
if (is.null(df[[req]])) {
stop(sprintf("column_mapping must map both judgement columns(%s and %s) — one column per rater, one row per item.",
r1_h, r2_h))
}
}Step 2: Normalise the two judgement columns to clean category labels.
Categories are text by definition; numbers are accepted and kept as labels (a 1-5 scale is a set of five categories, not a measurement).
as_label <- function(v) {
s <- trimws(as.character(v))
s[is.na(s)] <- ""
s[tolower(s) %in% c("", "na", "n/a", "null", "none given", "-", "--", "?")] <- ""
s
}
a <- as_label(df$rater_1)
b <- as_label(df$rater_2)
keep <- a != "" & b != ""
n_dropped <- sum(!keep)
a <- a[keep]; b <- b[keep]
n_items <- length(a)
if (n_items < 20) {
stop(sprintf("Only %d item(s) have a judgement from both %s and %s — kappa needs at least 20 to be worth reporting.",
n_items, r1_h, r2_h))
}Step 3: Merge labels that differ only in capitalisation, and say so.
all_lab <- c(a, b)
by_case <- split(all_lab, tolower(all_lab))
canon <- list()
n_case_merged <- 0L
case_example <- NULL
for (key in names(by_case)) {
spellings <- by_case[[key]]
uq <- unique(spellings)
tab <- sort(table(spellings), decreasing = TRUE)
winner <- names(tab)[1]
canon[[key]] <- winner
if (length(uq) > 1) {
n_case_merged <- n_case_merged + sum(spellings != winner)
if (is.null(case_example)) {
case_example <- c(setdiff(uq, winner)[1], winner)
}
}
}
a <- unlist(canon[tolower(a)], use.names = FALSE)
b <- unlist(canon[tolower(b)], use.names = FALSE)Step 4: Reject identifier-like columns; lump a long tail into "Other".
all_lab <- c(a, b)
freq <- sort(table(all_lab), decreasing = TRUE)
n_lumped <- 0L
if (length(freq) > 25) {
stop(sprintf("%s and %s together hold %d distinct values — that looks like free text or an identifier, not a set of categories. Map the two columns holding each rater's category choice.",
r1_h, r2_h, length(freq)))
}
if (length(freq) > 12) {
keep_lab <- names(freq)[1:11]
n_lumped <- length(freq) - 11L
a[!(a %in% keep_lab)] <- "Other"
b[!(b %in% keep_lab)] <- "Other"
}Step 5: Refuse the degenerate cases with a message naming the column.
ua <- unique(a); ub <- unique(b)
if (length(ua) < 2) {
stop(sprintf("%s gave the same answer(\"%s\") to every item, so there is nothing for kappa to measure — agreement beyond chance is undefined when one rater never varies.",
r1_h, ua[1]))
}
if (length(ub) < 2) {
stop(sprintf("%s gave the same answer(\"%s\") to every item, so there is nothing for kappa to measure — agreement beyond chance is undefined when one rater never varies.",
r2_h, ub[1]))
}Step 6: Decide the category order (detected, not assumed)
levs_all <- sort(unique(c(a, b)))
det <- detect_category_order(levs_all)
if (det$ordered) {
levs <- det$order
} else {Nominal: order by how much of the data each category carries, so the heatmap reads busiest-first rather than alphabetically.
tot <- sapply(levs_all, function(l) sum(a == l) + sum(b == l))
levs <- levs_all[order(-tot, levs_all)]
}
k <- length(levs)
fa <- factor(a, levels = levs)
fb <- factor(b, levels = levs)
tab <- table(fa, fb)
N <- matrix(as.numeric(tab), nrow = k, ncol = k,
dimnames = list(levs, levs))
n <- sum(N)
p <- N / n
rowm <- rowSums(p)
colm <- colSums(p)
final_rows <- n_items
rows_removed <- initial_rows - final_rowsStep 7: Cohen's kappa, its standard error, CI, and a z-test vs chance
I <- diag(1, k)
base <- kappa_general(p, I, n)
p_o <- base$po; p_e <- base$pe
kappa <- base$kappa; kappa_se <- base$se
ci_low <- if (is.na(kappa) || is.na(kappa_se)) NA_real_ else kappa - 1.96 * kappa_se
ci_high <- if (is.na(kappa) || is.na(kappa_se)) NA_real_ else kappa + 1.96 * kappa_seVariance under the null (kappa = 0) is a different quantity from the variance used for the CI — this is the one the significance test needs.
var0 <- (p_e + p_e^2 - sum(rowm * colm * (rowm + colm))) / (n * (1 - p_e)^2)
se0 <- if (is.finite(var0) && var0 > 0) sqrt(var0) else NA_real_
z_stat <- if (is.na(kappa) || is.na(se0)) NA_real_ else kappa / se0
p_value <- if (is.na(z_stat)) NA_real_ else 2 * pnorm(-abs(z_stat))Step 8: Weighted kappa — only when the categories are actually ordered
kappa_linear <- NA_real_; kappa_linear_se <- NA_real_
kappa_quadratic <- NA_real_; kappa_quadratic_se <- NA_real_
weight_note <- ""
if (k < 3) {
weight_note <- paste0(
"Weighted kappa is not reported: with only ", k,
" categories every disagreement is already the maximum possible one, so any weighting scheme collapses back to the unweighted value.")
} else if (!det$ordered) {
weight_note <- paste0(
"Weighted kappa is not reported, because the categories do not appear to be ordered — ",
det$rule,
". Weighting assumes some disagreements are worse than others, which is only meaningful on a scale that has a direction. If your categories ARE ordered, rename them so the order is visible(for example \"1 - \", \"2 - \", \"3 - \" prefixes) and re-run.")
} else {
idx <- seq_len(k)
D <- abs(outer(idx, idx, "-")) / (k - 1)
Wl <- 1 - D
Wq <- 1 - D^2
kl <- kappa_general(p, Wl, n)
kq <- kappa_general(p, Wq, n)
kappa_linear <- kl$kappa; kappa_linear_se <- kl$se
kappa_quadratic <- kq$kappa; kappa_quadratic_se <- kq$se
weight_note <- paste0(
"Weighted kappa is reported because the categories are ordered — ",
det$rule, ".")
}
band <- landis_koch(kappa)Step 9: Prevalence and bias — the two things that move kappa without
either rater changing how well they agree.
avg_marg <- (rowm + colm) / 2
prev_i <- which.max(avg_marg)
max_prev <- avg_marg[prev_i]
prev_cat <- levs[prev_i]
bias_vec <- abs(rowm - colm)
bias_i <- which.max(bias_vec)
max_bias <- bias_vec[bias_i]
bias_cat <- levs[bias_i]
pabak <- (k * p_o - 1) / (k - 1)
paradox <- isTRUE(p_o >= 0.70 && !is.na(kappa) && kappa < 0.60 && max_prev >= 0.60)Step 10: Confusion matrix (long form for the heatmap)
confusion_df <- data.frame(
rater_1_category = rep(levs, times = k),
rater_2_category = rep(levs, each = k),
n_items = as.numeric(N[cbind(rep(seq_len(k), times = k),
rep(seq_len(k), each = k))]),
stringsAsFactors = FALSE
)
confusion_df$share_pct <- round(100 * confusion_df$n_items / n, 2)
n_sparse <- sum(N < 5)Largest off-diagonal cell, found NA-safely and only among real cells.
off <- N; diag(off) <- NA_real_
off_ok <- which(!is.na(off) & off > 0)
if (length(off_ok) > 0) {
best <- off_ok[which.max(off[off_ok])]
bi <- ((best - 1) %% k) + 1
bj <- ((best - 1) %/% k) + 1
top_off_n <- off[best]
top_off_r1 <- levs[bi]
top_off_r2 <- levs[bj]
} else {
top_off_n <- 0; top_off_r1 <- NA_character_; top_off_r2 <- NA_character_
}Step 11: Per-category agreement (proportion of specific agreement:
twice the agreed items over the times either rater used the category)
denom <- rowSums(N) + colSums(N)
spec <- ifelse(denom > 0, 200 * diag(N) / denom, NA_real_)
category_df <- data.frame(
category = levs,
agreement_pct = round(spec, 1),
stringsAsFactors = FALSE
)
ok_cat <- which(!is.na(category_df$agreement_pct))
worst_cat <- if (length(ok_cat) > 0)
category_df$category[ok_cat[which.min(category_df$agreement_pct[ok_cat])]] else NA_character_
worst_val <- if (length(ok_cat) > 0)
min(category_df$agreement_pct[ok_cat], na.rm = TRUE) else NA_real_
best_cat <- if (length(ok_cat) > 0)
category_df$category[ok_cat[which.max(category_df$agreement_pct[ok_cat])]] else NA_character_
best_val <- if (length(ok_cat) > 0)
max(category_df$agreement_pct[ok_cat], na.rm = TRUE) else NA_real_Step 12: Each rater's marginal distribution (the paradox's evidence)
marginal_df <- data.frame(
category = rep(levs, times = 2),
rater = c(rep(r1_h, k), rep(r2_h, k)),
share_pct = round(100 * c(rowm, colm), 1),
stringsAsFactors = FALSE
)Step 13: The kappa results table
kappa_rows <- list(
list("Raw percent agreement", p_o, NA_real_, NA_real_,
sprintf("The two raters chose the same category for %s of the %s items.",
pct1(p_o), format(n, big.mark = ","))),
list("Agreement expected by chance", p_e, NA_real_, NA_real_,
sprintf("What two raters with these same category habits would agree on by luck alone(%s).",
pct1(p_e))),
list("Cohen's kappa", kappa, ci_low, ci_high,
sprintf("%s of the agreement left available above chance was achieved — %s on the Landis & Koch convention.",
pct1(kappa), band))
)
if (!is.na(kappa_linear)) {
kappa_rows[[length(kappa_rows) + 1]] <- list(
"Weighted kappa(linear)", kappa_linear,
kappa_linear - 1.96 * kappa_linear_se, kappa_linear + 1.96 * kappa_linear_se,
"Near-miss disagreements between adjacent categories count as partial credit, in proportion to how far apart they are.")
kappa_rows[[length(kappa_rows) + 1]] <- list(
"Weighted kappa(quadratic)", kappa_quadratic,
kappa_quadratic - 1.96 * kappa_quadratic_se, kappa_quadratic + 1.96 * kappa_quadratic_se,
"The same idea with distance squared: near misses are forgiven much more, far misses punished much harder. Always the most flattering of the three.")
}
kappa_rows[[length(kappa_rows) + 1]] <- list(
"Prevalence-and-bias-adjusted kappa", pabak, NA_real_, NA_real_,
"A diagnostic, not a replacement: what kappa would be if every category were equally common and both raters used them equally often. A large gap from Cohen's kappa means skewed marginals are doing the work.")
kappa_df <- data.frame(
statistic = sapply(kappa_rows, function(r) r[[1]]),
estimate = round(sapply(kappa_rows, function(r) as.numeric(r[[2]])), 4),
ci_low = round(sapply(kappa_rows, function(r) as.numeric(r[[3]])), 4),
ci_high = round(sapply(kappa_rows, function(r) as.numeric(r[[4]])), 4),
interpretation = sapply(kappa_rows, function(r) r[[5]]),
stringsAsFactors = FALSE
)Step 14: Methods disclosure
methods_df <- data.frame(
item = c(
"Design",
"Percent agreement",
"Chance agreement",
"Cohen's kappa",
"Standard error and CI",
"Significance test",
"Category order",
"Weighting",
"Benchmark labels",
"What kappa is not"
),
detail = c(
sprintf("%s items each judged once by %s and once by %s, across %d categories.",
format(n, big.mark = ","), r1_h, r2_h, k),
sprintf("Diagonal of the confusion matrix over the total: %s.", pct1(p_o)),
sprintf("Sum over categories of(%s's share) x (%s's share): %s.",
r1_h, r2_h, pct1(p_e)),
sprintf("(observed - expected) / (1 - expected) = (%s - %s) / (1 - %s) = %s.",
r3(p_o), r3(p_e), r3(p_e), r3(kappa)),
sprintf("Fleiss-Cohen-Everitt asymptotic standard error, %s; the 95%% interval is the estimate plus or minus 1.96 standard errors. It is a large-sample approximation%s.",
r3(kappa_se),
if (n_sparse > 0) sprintf(", and %d of the %d matrix cells hold fewer than 5 items, so treat the interval as indicative", n_sparse, k * k) else ""),
sprintf("z = kappa / SE under the null of chance-level agreement = %s, %s. The null variance is a different quantity from the one behind the confidence interval.",
r2(z_stat), fmt_pp(p_value)),
sprintf("Orderedness was detected, not assumed: %s.", det$rule),
weight_note,
"Landis & Koch(1977) labels(slight / fair / moderate / substantial / almost perfect) are a naming convention with no theoretical basis; the acceptable level of agreement depends on what the judgement is used for.",
"Kappa measures agreement, not accuracy: two raters can agree completely and both be wrong. It also describes these two raters on these items, and does not generalise to raters or items outside this set."
),
stringsAsFactors = FALSE
)Step 15: Headline metrics + the computed one-paragraph answer
metrics <- list(
`Items Rated` = n,
`Categories` = k,
`Percent Agreement` = pct1(p_o),
`Cohen's Kappa` = round(kappa, 3),
`Kappa 95% CI` = paste0(r3(ci_low), " to ", r3(ci_high)),
`Agreement Strength` = band,
`Weighted Kappa` = if (is.na(kappa_quadratic)) "not applicable" else round(kappa_quadratic, 3),
`Kappa vs Chance p` = fmt_p(p_value)
)
paradox_sentence <- if (paradox) {
paste0(
" Read those two numbers together before quoting either: raw agreement is high(", pct1(p_o),
") while kappa is only ", r3(kappa), ", which is the well-known kappa paradox rather than a contradiction. ",
pct1(max_prev), " of all judgements fell into the single category \"", prev_cat,
"\", so two raters with these habits would already have agreed on ", pct1(p_e),
" of items by chance; only ", pct1(1 - p_e), " of the scale was left for skill to win, and ",
pct1(kappa), " of that remainder was won.")
} else {
paste0(
" Chance alone would have produced ", pct1(p_e), " agreement given each rater's own habits, leaving ",
pct1(1 - p_e), " of the scale available above chance, of which ", pct1(kappa), " was achieved.")
}
weighted_sentence <- if (!is.na(kappa_linear)) {
paste0(" The categories are ordered(", det$rule,
"), so near misses can be given partial credit: linear-weighted kappa is ", r3(kappa_linear),
" and quadratic-weighted kappa ", r3(kappa_quadratic),
" — both higher than the unweighted value because most disagreements are between neighbouring categories.")
} else {
paste0(" ", weight_note)
}
json_output <- list(
answer = paste0(
r1_h, " and ", r2_h, " agreed on ", pct1(p_o), " of ", format(n, big.mark = ","),
" items across ", k, " categories. Cohen's kappa is ", r3(kappa),
" (95% CI ", r3(ci_low), " to ", r3(ci_high), "), ", band,
" agreement on the Landis & Koch convention, and ",
if (!is.na(p_value) && p_value < 0.05)
paste0("clearly better than chance(", fmt_pp(p_value), ")")
else
paste0("not distinguishable from chance(", fmt_pp(p_value), ")"),
".", paradox_sentence, weighted_sentence,
" Agreement was weakest on \"", worst_cat, "\" (", r2(worst_val),
"% of the times either rater used it) and strongest on \"", best_cat, "\" (",
r2(best_val), "%). Kappa measures agreement, not correctness."
),
cards = lapply(
c("tldr", "overview", "preprocessing", "confusion_heatmap", "kappa_results",
"category_agreement", "rater_marginals", "methods"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
initial_rows = initial_rows, final_rows = final_rows,
rows_removed = rows_removed, n_dropped = n_dropped,
n_case_merged = n_case_merged, case_example = case_example,
n_lumped = n_lumped,
r1_h = r1_h, r2_h = r2_h,
n_items = n, k_cats = k, levs = levs,
ordered_flag = det$ordered, order_rule = det$rule,
p_o = p_o, p_e = p_e, kappa = kappa, kappa_se = kappa_se,
ci_low = ci_low, ci_high = ci_high,
z_stat = z_stat, p_value = p_value,
kappa_linear = kappa_linear, kappa_linear_se = kappa_linear_se,
kappa_quadratic = kappa_quadratic, kappa_quadratic_se = kappa_quadratic_se,
weight_note = weight_note, band = band, pabak = pabak,
paradox = paradox, paradox_sentence = paradox_sentence,
max_prev = max_prev, prev_cat = prev_cat,
max_bias = max_bias, bias_cat = bias_cat,
n_sparse = n_sparse,
top_off_n = top_off_n, top_off_r1 = top_off_r1, top_off_r2 = top_off_r2,
worst_cat = worst_cat, worst_val = worst_val,
best_cat = best_cat, best_val = best_val,
confusion_df = confusion_df, kappa_df = kappa_df,
category_df = category_df, marginal_df = marginal_df,
methods_df = methods_df,
metrics = metrics, json_output = json_output
)
}