Executive Summary
Fleiss' kappa 0.484 (moderate) across 4 raters on 200 items
The short answer
The four reviewers agreed on 63.2% of rater pairs across 200 items, with Fleiss' kappa of 0.484 (95% CI 0.426 to 0.535)—moderate agreement and clearly better than chance (p < 0.001). However, Reviewer D is notably out of step; without that rater, the panel's kappa would rise to 0.690.
The detail
Fleiss' kappa is 0.484, which the Landis & Koch convention classifies as moderate. Krippendorff's alpha gives 0.485, confirming the result. Chance alone would have produced 28.7% agreement given the panel's category habits, leaving 71.3% of the scale available; the panel achieved 48.4% of that headroom. The z-statistic is 24.45 with p < 0.001, far better than random. Reviewer A is most aligned (rater kappa 0.584); Reviewer D is least aligned (rater kappa 0.273), a gap of 0.311. Of the 200 items, 75 (37.5%) drew unanimous agreement, while 53 (26.5%) produced no majority at all. Critical caveat: these coefficients measure whether the panel agrees, not whether the panel is right.
What this can't tell you
Agreement does not imply correctness. A unanimous panel can be unanimously wrong; only a gold standard or external validation settles whether the judgements are accurate.
Analysis Overview
Fleiss' kappa and Krippendorff's alpha on 200 items judged by 4 raters across 4 categories.
The short answer
The four reviewers show moderate agreement at 0.484 (Fleiss' kappa), well above what random chance would produce. However, raw agreement of 63.2% looks higher than it is because the panel leans heavily on certain categories—a chance baseline of 28.7% means they'd collide on those categories even without reading carefully. Krippendorff's alpha reaches the same answer (0.485) from a different angle, asking about disagreement instead.
The detail
Fleiss' kappa of 0.484 (95% CI 0.426 to 0.535) captures the 48.4% of available headroom above chance that the panel actually won. Raw pairwise agreement was 63.2%, but the panel's own category habits would have produced 28.7% agreement by luck alone, leaving 71.3% of the scale available. Kappa divides the achieved agreement by that available headroom. Krippendorff's alpha (0.485) computes the same idea through observed versus expected disagreement and tolerates missing ratings. With exactly 4 raters and 4 categories, both coefficients apply; with 2 raters you would use Cohen's kappa; with continuous scores, an intraclass correlation.
What this can't tell you
These coefficients measure consistency of calls, not correctness. A panel can agree unanimously and be wrong; only comparison to a gold standard settles that.
Data Quality
200 items and 800 judgements from 200 rows loaded.
The short answer
All 200 items were complete: every one of the 4 raters judged every item, producing 800 judgements with no gaps. The data came in wide form (one row per item, one column per rater) and was correctly recognized because each rater column held only 4 distinct labels rather than unique values. Both Fleiss' kappa and Krippendorff's alpha were computed on the same full dataset.
The detail
200 rows resolved to 200 items, 4 raters, and 800 judgements across 4 categories. No ratings were missing: every rater judged every item, so both coefficients used exactly the same 200 items and 800 judgements. The 95% confidence interval around kappa runs from 0.426 to 0.535, a width of 0.109, based on these 200 items. Labels were trimmed of surrounding spaces before comparison, and blanks or placeholders were treated as missing judgements rather than as a category.
What this can't tell you
The interval width (0.109) reflects the sample size of 200 items; a smaller item pool would widen it. The completeness of the data means both kappa and alpha use identical item sets, so any visible gap between them would point to different methods rather than different data.
Agreement Statistics
Pairwise agreement, chance agreement, Fleiss' kappa and Krippendorff's alpha with bootstrap intervals.
| Statistic | Estimate | CI Low | CI High | Interpretation |
|---|---|---|---|---|
| Raw pairwise agreement | 0.6325 | — | — | Across every pair of raters who judged the same item, 63.2% of those pairs chose the same category. |
| Agreement expected by chance | 0.2871 | — | — | What a panel using the categories this often would have agreed on by luck alone (28.7%). |
| Fleiss' kappa | 0.4845 | 0.4258 | 0.5346 | 48.4% of the agreement left available above chance was achieved, on the 200 item(s) every one of the 4 raters judged — moderate on the Landis & Koch convention. |
| Krippendorff's alpha (nominal) | 0.4851 | 0.4245 | 0.5448 | The disagreement-based coefficient, computed on all 200 items with two or more judgements — it does not require every rater to have judged every item. |
| Krippendorff's alpha (ordinal) | 0.7719 | 0.7229 | 0.815 | The same coefficient with an ordinal distance between categories, so a near-miss between neighbouring levels counts as a smaller disagreement than a jump across the scale. |
The short answer
The panel agreed on 63.2% of rater pairs, but chance would have delivered 28.7% on its own given how often each category is used. Fleiss' kappa reports that 48.4% of the remaining headroom was actually won. The 95% confidence interval of 0.426 to 0.535 comes from resampling items, not from a formula, because the classical Fleiss standard error describes a null hypothesis and would misstate the interval around a non-zero estimate.
The detail
Raw pairwise agreement is 0.6325 (63.2%). Expected agreement by chance is 0.2871 (28.7%). Fleiss' kappa is 0.4845 (48.4% of available headroom), with 95% CI from 0.4258 to 0.5346 on all 200 items. Krippendorff's alpha (nominal) is 0.4851 with CI 0.4245 to 0.5448. Ordinal alpha is 0.7719 (CI 0.7229 to 0.815) because most disagreements here fall between neighbouring categories rather than across the full scale. The Fleiss and nominal alpha estimates differ by 0.001, confirming they are different arithmetic routes to the same complete dataset. Ordinal alpha will always be more flattering when categories carry an order.
What this can't tell you
The Landis & Koch bands are convention, not a statistical standard. Whether 0.484 is fit for purpose depends on what the judgements decide—that is your call, not the data's.
Which Rater Is Out of Step
Each rater's agreement with the rest of the panel, and the panel's kappa without them.
Reviewer A Rating leads at 70.3% raw agreement with the rest of the panel and a rater kappa of 0.584. Reviewer D Rating trails at 48.17% raw agreement and a rater kappa of 0.273—a gap of 0.311 of kappa, large enough that particular people, not the rubric as a whole, are holding the headline number down. Removing Reviewer D Rating would lift the panel's kappa from 0.484 to 0.690, a gain of 0.205. Treat that as diagnostic: dropping the rater who disagrees most always raises measured agreement, even when that rater is the one who is right. A low rater-level score is associated with, not proof of, misreading the rubric; the same pattern appears when one rater applies it correctly and the others do not.
Which Categories the Panel Fights Over
Category-specific kappa across 4 categories.
Disagreement is concentrated rather than spread evenly: "Excellent" is the most reliable at 0.644, while "Fair" is the least at 0.368—a spread of 0.276. Fixing "Fair" (which absorbs 25.75% of all judgements) would move the headline number more than a general calibration session. "Good" (39.25% of judgements) scores 0.4059, and "Poor" (12.25% of judgements) scores 0.6046. A category that one rater reaches for freely and others barely touch scores low here even when overall kappa looks healthy, which is exactly the detail a single summary coefficient hides. Rubric work aimed at "Fair" and "Good" would be the highest-leverage intervention.
How Often the Panel Reached a Verdict
Consensus strength across 200 items on the analysis basis.
Disagreement is not spread evenly: 75 items (37.5%) drew unanimous verdicts, 72 (36.0%) produced a majority without unanimity, and 53 (26.5%) produced no majority at all. Those 53 no-majority items are the concrete work list—they are where the rubric is silent or the item is genuinely ambiguous. Re-reading a sample of them usually explains the headline number faster than the coefficient itself does. Consensus bands are set from each item's largest block of agreeing raters, so an evenly split item needs no arbitrary winner. With only 4 raters the possible values are coarse; empty bands mean the arithmetic cannot land there rather than that no item did.
How Each Rater Uses the Scale
Category shares per rater, side by side.
The busiest category, "Good", takes 39.2% of all judgements—a reasonably even spread across raters (41.5%, 41%, 40.5%, 34%) that keeps the chance-agreement floor at 28.7% and lets kappa stay closer to the raw agreement of 63.2%. Reviewer D Rating stands visibly apart: it reaches for "Poor" (22%) and "Fair" (34%) much more often than the others (9–11% and 18.5–28.5% respectively), and "Excellent" much less often (10% vs 24–29%). This is a different threshold, not random error, and will depress agreement even when each individual judgement looks defensible. Reviewers A, B, and C are much more aligned in their use of the scale.
Methods & Disclosure
Every formula behind the numbers, and what they cannot decide.
| Item | Detail |
|---|---|
| Design | 200 item(s) judged by 4 rater(s) across 4 categories, 800 judgement(s) in total. |
| Input shape | Detected from the data, not from the mapping: the 4 mapped columns were read as one column per rater, because each holds only a handful of repeated labels (4, 4, 4, 4 distinct) rather than one value per row. |
| Raw pairwise agreement | Mean over items of the share of rater PAIRS on that item choosing the same category: 63.2%. |
| Chance agreement | Sum of the squared overall category shares (Poor 12.2%, Fair 25.8%, Good 39.2%, Excellent 22.8%): 28.7%. |
| Fleiss' kappa | (observed - expected) / (1 - expected) = (0.632 - 0.287) / (1 - 0.287) = 0.484, computed on the 200 item(s) every one of the 4 raters judged. |
| Significance test | Fleiss' null-hypothesis standard error is 0.020, giving z = 24.45 and p < 0.001 against the hypothesis that agreement is no better than chance. That variance describes kappa when kappa is truly zero, so it belongs to this test and NOT to the confidence interval. |
| Confidence intervals | A seeded nonparametric bootstrap resampling ITEMS with replacement (500 resamples, fixed seed), taking the 2.5th and 97.5th percentiles. The bootstrap standard error of kappa is 0.030. This is used rather than a closed form because the null variance above is the wrong quantity for an interval around a non-zero estimate, and alpha has no simple closed-form variance. |
| Missing ratings | No ratings are missing: every rater judged every item, so Fleiss' kappa and Krippendorff's alpha are computed on exactly the same data. |
| Krippendorff's alpha | Built from the coincidence matrix: alpha = 1 - observed disagreement / expected disagreement, with (n-1) weighting so it is defined for small samples. Nominal alpha treats every disagreement as equal. Ordinal alpha weights a disagreement by the distance between the categories along the detected order. |
| Category order | Orderedness was detected, not assumed: the labels are all points on a quality scale, so they were ordered along it. |
| Per-category kappa | Fleiss' category-specific coefficient, generalized to a variable number of raters per item: one minus the observed splits on that category over the splits expected from its overall share. |
| Per-rater figures | A rater's agreement figure is the share of the OTHER raters' judgements on the same items that matched this rater's own, chance-corrected against the same 28.7% floor. The leave-one-out column recomputes Fleiss' kappa for the panel with that rater removed. |
| Benchmark labels | Landis & Koch (1977) labels (slight / fair / moderate / substantial / almost perfect) are a naming convention with no theoretical basis; the acceptable level of agreement depends on what the judgement is used for. |
| What agreement is not | Agreement is not correctness: a panel can agree unanimously and be unanimously wrong. These figures describe this panel on these items and do not generalise to other raters or other items. |
The short answer
Kappa is computed from a simple formula on the 200-by-4 matrix of category counts: (observed agreement – expected agreement) / (1 – expected agreement). The confidence intervals come from resampling items 500 times, not from a closed-form formula, because the null-hypothesis standard error (0.020, giving z = 24.45) describes what kappa looks like when it is truly zero and does not apply to an interval around a non-zero estimate.
The detail
Raw pairwise agreement: mean share of rater pairs per item choosing the same category = 63.2%. Chance agreement: sum of squared category shares (Poor 12.2%, Fair 25.8%, Good 39.2%, Excellent 22.8%) = 28.7%. Fleiss' kappa = (0.632 − 0.287) / (1 − 0.287) = 0.484 on all 200 items. Significance: null standard error 0.020, z = 24.45, p < 0.001. Bootstrap intervals: 500 resamples of items with replacement, fixed seed, 2.5th and 97.5th percentiles; bootstrap SE = 0.030. Krippendorff's alpha uses the coincidence matrix: alpha = 1 − (observed disagreement / expected disagreement), with (n−1) weighting. Ordinal alpha applies category distance weights. Category order was detected from labels, not assumed.
What this can't tell you
Agreement is not correctness; no coefficient here validates whether the panel is right. These figures describe this panel on these items only—different raters or a different item pool with different category prevalence will produce different kappa values. The confidence intervals are approximate and widen sharply with small item counts.
Methodology
Statistical methodology and diagnostics for Multi-Rater Agreement — Fleiss Kappa
Statistical Method
Standard-library analysis: three or more raters, one categorical judgement per item — how much do they really agree, and who is out of step? Map one column per rater (or, if your file has one row per rating, map the item, rater and judgement columns — the shape is detected from the data) and get Fleiss' kappa with a bootstrap confidence interval and a test against chance, Krippendorff's alpha in nominal and — when the labels turn out to be ordered — ordinal form, a per-category kappa showing exactly which categories the panel cannot pin down, a per-rater agreement with the consensus plus a leave-one-rater-out kappa that names the outlier, the item-level distribution of how strong the consensus actually was, and an explicit account of what each coefficient did with any missing ratings instead of quietly dropping them.
- Each item is judged once by each rater, from the same list of categories
- The judgements are independent — no rater saw another's answer before deciding
- Items are independent of each other; the same item is not counted twice
- Fleiss' kappa additionally assumes every item carries the same number of raters; where that fails the analysis says so and reports Krippendorff's alpha alongside
- Agreement is not correctness — a panel can agree unanimously and be unanimously wrong; only a gold standard can settle that
- Kappa is depressed when one category dominates or when raters use the categories at different rates; this is the kappa paradox, and the analysis reports it rather than hiding it
- The confidence intervals come from resampling items and are an approximation; with few items they are wide, and the analysis reports the item count next to the verdict
- Fleiss' kappa is computed on items every rater judged; when ratings are missing that is a different item set from the one Krippendorff's alpha uses, and both are shown rather than blended
Analysis Code
Complete R source code for this analysis
Multi-Rater Agreement — Fleiss Kappa
Three or more raters, one categorical judgement per item: how much do they agree, how much of that is more than chance would have produced, and which rater is pulling the group apart? The analysis computes Fleiss' kappa with a standard error and confidence interval, Krippendorff's alpha (nominal, and ordinal when the categories turn out to be ordered), a per-category kappa showing which categories the raters fight over, a per-rater agreement with the consensus plus a leave-one-rater-out kappa that names the outlier, and the item-level distribution of how strong the consensus actually was.
Why This Method?
Raw agreement across a panel is not interpretable on its own: a panel that answers "Yes" to almost everything will agree constantly without reading a single item. Fleiss' kappa subtracts the agreement chance alone would deliver given how often each category is used overall. Krippendorff's alpha answers the same question through a different route and, unlike Fleiss', keeps working when not every rater judged every item.
What This Analysis Covers
- Raw pairwise agreement, chance agreement, and Fleiss' kappa with SE + CI
- Krippendorff's alpha, nominal and (when the labels are ordered) ordinal
- Explicit accounting for missing ratings — what each statistic used
- Per-category kappa: which categories the panel cannot pin down
- Per-rater agreement with the consensus and leave-one-rater-out kappa
- The item-level consensus distribution, and the kappa paradox when it fires
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {rater_1 .. rater_N}. Both real-world layouts are accepted and the shape is detected from the data, never assumed. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Core Analysis Pipeline
Rule 1 — the labels are numbers (1, 2, 3 / 0-10 scales).
num <- suppressWarnings(as.numeric(clean))
if (!anyNA(num) && length(unique(num)) == length(num)) {
return(list(ordered = TRUE, order = levs[order(num)],
rule = "the labels are numbers, so they were ordered by value"))
}Rule 2 — the labels start with a number ("1 - Poor", "2 = Fair").
lead <- suppressWarnings(as.numeric(sub("^\\s*([-+]?[0-9]+(\\.[0-9]+)?)\\s*[-=.):|].*$",
"\\1", clean)))
has_lead <- grepl("^\\s*[-+]?[0-9]+(\\.[0-9]+)?\\s*[-=.):|]", clean)
if (all(has_lead) && !anyNA(lead) && length(unique(lead)) == length(lead)) {
return(list(ordered = TRUE, order = levs[order(lead)],
rule = "every label starts with a number, so they were ordered by that number"))
}Rule 3 — the labels are all drawn from one recognised ordinal vocabulary.
scales <- list(
"an agreement scale" = c("strongly disagree", "disagree", "somewhat disagree",
"slightly disagree", "neutral", "neither agree nor disagree",
"slightly agree", "somewhat agree", "agree", "strongly agree"),
"a quality scale" = c("very poor", "poor", "below average", "fair", "average",
"satisfactory", "good", "very good", "excellent", "outstanding"),
"a frequency scale" = c("never", "rarely", "seldom", "sometimes", "occasionally",
"often", "frequently", "usually", "always"),
"a severity scale" = c("none", "minimal", "mild", "moderate", "severe",
"very severe", "extreme", "critical"),
"a magnitude scale" = c("very low", "low", "medium", "moderate", "high", "very high"),
"a satisfaction scale" = c("very dissatisfied", "dissatisfied", "neutral",
"satisfied", "very satisfied"),
"a likelihood scale" = c("very unlikely", "unlikely", "possible", "likely",
"very likely", "certain"),
"a decision scale" = c("reject", "major revision", "revise", "minor revision",
"accept with changes", "accept"),
"a priority scale" = c("trivial", "minor", "moderate", "major", "critical", "blocker")
)
low <- tolower(clean)
if (length(levs) >= 3 && !any(duplicated(low))) {
for (nm in names(scales)) {
sc <- scales[[nm]]
if (all(low %in% sc)) {
return(list(ordered = TRUE, order = levs[order(match(low, sc))],
rule = paste0("the labels are all points on ", nm,
", so they were ordered along it")))
}
}
}
list(ordered = FALSE, order = levs,
rule = "no numeric values, numeric prefixes, or recognised ordinal wording were found in the labels")
}Step 1: Collect the mapped judgement columns, in mapped order
rc <- grep("^rater_[0-9]+$", names(df), value = TRUE)
rc <- rc[order(as.numeric(sub("^rater_", "", rc)))]
ch <- humanize_semantic(rc, col_map)
if (length(rc) < 3) {
stop(sprintf("Only %d judgement column(s) were mapped(%s). This analysis needs three or more raters. For exactly two raters the right tool is Cohen's kappa, not Fleiss' — or, if your file has one row per rating, map the three columns holding the item, the rater and the judgement.",
length(rc),
if (length(ch) > 0) paste(ch, collapse = ", ") else "none"))
}
cols <- lapply(rc, function(cn) as_label(df[[cn]]))
names(cols) <- rc
d_distinct <- vapply(cols, function(v) length(unique(v[v != ""])), integer(1))Step 2: Detect the input SHAPE from the data, not from the mapping.
Wide means one column per rater and one row per item. Long means one row per rating, with an item column, a rater column and a judgement column. A long file is unmistakable in the counts: the item column carries many distinct values, each repeating once per rater, while the other two carry only a handful.
shape <- "wide"
shape_rule <- ""
li <- ri <- ji <- NA_integer_
if (length(rc) == 3) {
ii <- which.max(d_distinct)
others <- setdiff(seq_len(3), ii)
avg_rep <- if (d_distinct[ii] > 0) initial_rows / d_distinct[ii] else 0
if (d_distinct[ii] >= 10 && avg_rep >= 2.5 && min(d_distinct[others]) >= 2 &&
d_distinct[ii] > 3 * max(d_distinct[others])) {
item_v <- cols[[ii]]Which of the remaining two is the RATER? The one for which each item appears at most once — a rater judges an item once, whereas several raters give an item the same judgement all the time.
dupr <- vapply(others, function(o)
sum(duplicated(paste0(item_v, "\r", cols[[o]]))), numeric(1))
ri <- others[which.min(dupr)]
ji <- setdiff(others, ri)
li <- ii
shape <- "long"
shape_rule <- sprintf(
"%s holds %d distinct values each repeating about %s times, while %s and %s hold only %d and %d distinct values — that is one row per rating, not one column per rater",
ch[li], d_distinct[li], r2(avg_rep), ch[ri], ch[ji],
d_distinct[ri], d_distinct[ji])
}
}Step 3: Flatten either shape into one table of (item, rater, label)
n_dup_ratings <- 0L
if (shape == "long") {
item_v <- cols[[li]]; rater_v <- cols[[ri]]; lab_v <- cols[[ji]]
ok <- item_v != "" & rater_v != ""
n_blank_rows <- sum(!ok)
item_v <- item_v[ok]; rater_v <- rater_v[ok]; lab_v <- lab_v[ok]The same rater judging the same item twice is a data error, not a second opinion; the first judgement is kept and the count is reported.
key <- paste0(item_v, "\r", rater_v)
dupd <- duplicated(key)
n_dup_ratings <- sum(dupd)
item_v <- item_v[!dupd]; rater_v <- rater_v[!dupd]; lab_v <- lab_v[!dupd]
if (length(unique(rater_v)) > 20) {
stop(sprintf("%s holds %d distinct raters. Above 20 raters this looks like an identifier column rather than a panel; map the column naming who gave each judgement.",
ch[ri], length(unique(rater_v))))
}
row_unit <- "rating"
shape_rule <- paste0("the three mapped columns were read as a long file because ", shape_rule)
} else {Wide: a column holding a distinct value on nearly every row is an identifier or free text, not a category, and is refused by name. Both an absolute cap and a proportional test are needed — a key column in a short file has few distinct values in absolute terms but repeats nothing.
n_filled <- vapply(cols, function(v) sum(v != ""), integer(1))
bad <- which(d_distinct > 25 | (d_distinct >= 10 & d_distinct > 0.5 * n_filled))
if (length(bad) > 0) {
b <- bad[1]
why <- if (d_distinct[b] > 25)
"which is too many to be a set of categories"
else
sprintf("across %d judgement(s) — nearly one per row, so almost nothing repeats",
n_filled[b])
stop(sprintf("%s holds %d distinct values, %s. That makes it an identifier or free text rather than a rater's category choice. Map one column per rater, each holding that rater's category. If your file instead has one row per rating, map exactly three columns: the item, the rater, and the judgement.",
ch[b], d_distinct[b], why))
}
n_blank_rows <- 0L
item_v <- rep(as.character(seq_len(initial_rows)), times = length(rc))
rater_v <- rep(ch, each = initial_rows)
lab_v <- unlist(cols, use.names = FALSE)
keep <- lab_v != ""
item_v <- item_v[keep]; rater_v <- rater_v[keep]; lab_v <- lab_v[keep]
row_unit <- "item"
shape_rule <- sprintf(
"the %d mapped columns were read as one column per rater, because each holds only a handful of repeated labels(%s distinct) rather than one value per row",
length(rc), paste(d_distinct, collapse = ", "))
}Step 4: Merge labels that differ only in capitalisation, and say so.
by_case <- split(lab_v, tolower(lab_v))
canon <- list(); n_case_merged <- 0L; case_example <- NULL
for (keyc in names(by_case)) {
spellings <- by_case[[keyc]]
uq <- unique(spellings)
tab <- sort(table(spellings), decreasing = TRUE)
winner <- names(tab)[1]
canon[[keyc]] <- winner
if (length(uq) > 1) {
n_case_merged <- n_case_merged + sum(spellings != winner)
if (is.null(case_example)) case_example <- c(setdiff(uq, winner)[1], winner)
}
}
lab_v <- unlist(canon[tolower(lab_v)], use.names = FALSE)Step 5: Refuse free text; lump a long tail of rare labels into "Other".
freq <- sort(table(lab_v), decreasing = TRUE)
n_lumped <- 0L
if (length(freq) > 25) {
stop(sprintf("The judgement columns(%s) together hold %d distinct values — that looks like free text rather than a set of categories. Map the columns holding each rater's category choice.",
paste(ch, collapse = ", "), length(freq)))
}
if (length(freq) > 12) {
keep_lab <- names(freq)[1:11]
n_lumped <- length(freq) - 11L
lab_v[!(lab_v %in% keep_lab)] <- "Other"
}
rater_names <- sort(unique(rater_v))
n_raters <- length(rater_names)
if (n_raters < 3) {
stop(sprintf("Only %d rater(s) were found in %s. Fleiss' kappa needs three or more; with exactly two raters use Cohen's kappa instead.",
n_raters, if (shape == "long") ch[ri] else paste(ch, collapse = ", ")))
}
levs_all <- sort(unique(lab_v))
if (length(levs_all) < 2) {
stop(sprintf("Every judgement in %s is \"%s\", so there is only one category and agreement beyond chance is undefined — nothing varies for kappa to explain.",
paste(ch, collapse = ", "), levs_all[1]))
}Step 6: Decide the category order (detected, not assumed)
det <- detect_category_order(levs_all)
if (det$ordered) {
levs <- det$order
} else {
tot <- vapply(levs_all, function(l) sum(lab_v == l), numeric(1))
levs <- levs_all[order(-tot, levs_all)]
}
k <- length(levs)Step 7: Build the item-by-category count matrix
items <- unique(item_v)
Nmat <- matrix(0, nrow = length(items), ncol = k,
dimnames = list(NULL, levs))
im <- match(item_v, items)
cm <- match(lab_v, levs)
tabm <- table(factor(im, levels = seq_along(items)),
factor(cm, levels = seq_len(k)))
Nmat[] <- as.numeric(tabm)
m_i <- rowSums(Nmat)
n_ratings <- sum(m_i)
all_idx <- which(m_i >= 2)
n_items_dropped <- length(items) - length(all_idx)
complete_idx <- which(m_i == n_raters)
n_items_partial <- sum(m_i >= 2 & m_i < n_raters)
n_missing <- n_raters * length(items) - n_ratings
if (length(all_idx) < 20) {
stop(sprintf("Only %d item(s) carry judgements from at least two of %s. Multi-rater agreement needs at least 20 such items before the numbers mean anything.",
length(all_idx), paste(ch, collapse = ", ")))
}Step 8: Fleiss' kappa. It is a complete-cases statistic: the classical
coefficient assumes every item was judged by the same number of raters. When ratings are missing the module reports BOTH the complete-case value (the one other software produces) and a generalized value over every item with two or more ratings — and says which items each one used.
fleiss_calc <- function(ii) {
m <- m_i[ii]
sq <- rowSums(Nmat[ii, , drop = FALSE]^2)
P_i <- (sq - m) / (m * (m - 1))
P_bar <- mean(P_i)
tot <- sum(m)
p <- colSums(Nmat[ii, , drop = FALSE]) / tot
P_e <- sum(p^2)
kap <- if ((1 - P_e) > 1e-12) (P_bar - P_e) / (1 - P_e) else NA_real_
list(P_bar = P_bar, P_e = P_e, kappa = kap, p = p, tot = tot,
n = length(ii), m_const = if (length(unique(m)) == 1) m[1] else NA_real_)
}
basis <- if (length(complete_idx) >= 20) "complete" else "all"
basis_idx <- if (basis == "complete") complete_idx else all_idx
fb <- fleiss_calc(basis_idx)
fa <- fleiss_calc(all_idx)
p_bar <- fb$P_bar; p_e <- fb$P_e; kappa <- fb$kappa
p_j <- fb$p
kappa_all <- fa$kappa
n_basis <- fb$nFleiss' null-hypothesis standard error (Fleiss 1971), which exists only when every item on the basis carries the same number of raters. It is the variance of kappa when kappa is truly zero, so it belongs to the significance test and NOT to the confidence interval.
se_null <- NA_real_; z_stat <- NA_real_; p_value <- NA_real_
if (!is.na(fb$m_const) && fb$m_const >= 2) {
R <- fb$m_const
S2 <- sum(p_j^2); S3 <- sum(p_j^3)
inner <- S2 - (2 * R - 3) * S2^2 + 2 * (R - 2) * S3
if (is.finite(inner) && inner > 0 && (1 - S2) > 1e-12) {
se_null <- sqrt(2 / (n_basis * R * (R - 1)) * inner) / (1 - S2)
z_stat <- kappa / se_null
p_value <- 2 * pnorm(-abs(z_stat))
}
}Step 9: Krippendorff's alpha, via the coincidence matrix. Alpha is the
statistic that tolerates missing ratings, so it uses EVERY item carrying at least two judgements — including the ones Fleiss had to set aside.
Ocon <- t(apply(Nmat, 1, function(row) {
m <- sum(row)
if (m < 2) return(rep(0, k * k))
as.vector((outer(row, row) - diag(row, nrow = k)) / (m - 1))
}))
if (k == 1) Ocon <- matrix(Ocon, ncol = 1)
LO <- outer(seq_len(k), seq_len(k), pmin)
HI <- outer(seq_len(k), seq_len(k), pmax)
ordinal_ok <- det$ordered && k >= 3
alpha_calc <- function(ii) {
o <- matrix(colSums(Ocon[ii, , drop = FALSE]), nrow = k, ncol = k)
nc <- rowSums(o)
ntot <- sum(nc)
a_nom <- NA_real_; a_ord <- NA_real_
De_n <- ntot^2 - sum(nc^2)
if (is.finite(De_n) && De_n > 1e-12 && ntot > 1) {
Do_n <- ntot - sum(diag(o))
a_nom <- 1 - (ntot - 1) * Do_n / De_n
}
if (ordinal_ok && ntot > 1) {
cum0 <- c(0, cumsum(nc))
S <- matrix(cum0[HI + 1] - cum0[LO], nrow = k, ncol = k)
Dm <- (S - outer(nc, nc, function(x, y) (x + y) / 2))^2
De_o <- sum(outer(nc, nc) * Dm)
if (is.finite(De_o) && De_o > 1e-12) {
a_ord <- 1 - (ntot - 1) * sum(o * Dm) / De_o
}
}
list(nominal = a_nom, ordinal = a_ord)
}
aa <- alpha_calc(all_idx)
alpha_nom <- aa$nominal; alpha_ord <- aa$ordinal
alpha_nom_basis <- alpha_calc(basis_idx)$nominalStep 10: Confidence intervals by resampling ITEMS. The closed-form
variance above is a null-hypothesis quantity and is wrong for an interval around a non-zero estimate; alpha has no simple closed form at all. A seeded nonparametric bootstrap over items answers both, and treats items as the sampling unit, which is what they are.
B <- 500L
set.seed(42)
nb <- length(basis_idx); na_ <- length(all_idx)
bk <- numeric(B); ban <- numeric(B); bao <- numeric(B)
for (b in seq_len(B)) {
ib <- basis_idx[sample.int(nb, nb, replace = TRUE)]
bk[b] <- fleiss_calc(ib)$kappa
ia <- all_idx[sample.int(na_, na_, replace = TRUE)]
ab <- alpha_calc(ia)
ban[b] <- ab$nominal; bao[b] <- ab$ordinal
}
qci <- function(v) {
v <- v[is.finite(v)]
if (length(v) < 50) return(c(NA_real_, NA_real_, NA_real_))
c(as.numeric(quantile(v, 0.025, names = FALSE)),
as.numeric(quantile(v, 0.975, names = FALSE)), sd(v))
}
qk <- qci(bk); qan <- qci(ban); qao <- qci(bao)
ci_low <- qk[1]; ci_high <- qk[2]; se_boot <- qk[3]
alpha_nom_lo <- qan[1]; alpha_nom_hi <- qan[2]
alpha_ord_lo <- qao[1]; alpha_ord_hi <- qao[2]
band <- landis_koch(kappa)Step 11: Per-category kappa — which categories the panel fights over.
Fleiss' category-specific coefficient, generalized to a variable number of raters per item.
mb <- m_i[basis_idx]
Nb <- Nmat[basis_idx, , drop = FALSE]
tot_b <- sum(mb)
cat_kappa <- vapply(seq_len(k), function(j) {
pj <- sum(Nb[, j]) / tot_b
den <- tot_b * pj * (1 - pj)
if (!is.finite(den) || den <= 1e-12) return(NA_real_)
obs <- sum(Nb[, j] * (mb - Nb[, j]) / (mb - 1))
1 - obs / den
}, numeric(1))
category_df <- data.frame(
category = levs,
category_kappa = round(cat_kappa, 4),
share_pct = round(100 * colSums(Nb) / tot_b, 2),
n_uses = as.numeric(colSums(Nb)),
stringsAsFactors = FALSE
)
ok_cat <- which(!is.na(cat_kappa))
worst_cat <- if (length(ok_cat) > 0) levs[ok_cat[which.min(cat_kappa[ok_cat])]] else NA_character_
worst_val <- if (length(ok_cat) > 0) min(cat_kappa[ok_cat]) else NA_real_
best_cat <- if (length(ok_cat) > 0) levs[ok_cat[which.max(cat_kappa[ok_cat])]] else NA_character_
best_val <- if (length(ok_cat) > 0) max(cat_kappa[ok_cat]) else NA_real_With exactly two categories every disagreement is the same disagreement, so both category-specific kappas equal each other and the overall coefficient. Naming a "worst" category there would be an artefact of tie-breaking.
cat_degenerate <- k == 2 ||
(length(ok_cat) > 1 && (best_val - worst_val) < 1e-9)Step 12: Per-rater agreement with the rest of the panel, and the
leave-one-rater-out kappa. The second is the actionable one: it says what the panel's agreement would be if this rater were not in it.
basis_items <- items[basis_idx]
in_basis <- item_v %in% basis_items
bi_item <- item_v[in_basis]; bi_rater <- rater_v[in_basis]; bi_lab <- lab_v[in_basis]
bi_row <- match(bi_item, basis_items)
bi_col <- match(bi_lab, levs)
m_of_row <- mb[bi_row]
n_same <- Nb[cbind(bi_row, bi_col)]
rater_obs <- numeric(n_raters); rater_kap <- numeric(n_raters)
rater_wo <- numeric(n_raters); rater_n <- numeric(n_raters)
for (r in seq_len(n_raters)) {
sel <- bi_rater == rater_names[r]
den <- sum(m_of_row[sel] - 1)
rater_n[r] <- sum(sel)
rater_obs[r] <- if (den > 0) sum(n_same[sel] - 1) / den else NA_real_
rater_kap[r] <- if (is.na(rater_obs[r]) || (1 - p_e) <= 1e-12) NA_real_
else (rater_obs[r] - p_e) / (1 - p_e)Leave-one-rater-out: strip this rater's ratings and recompute Fleiss on the items that still carry two or more judgements.
Nw <- Nb
idxw <- which(sel)
if (length(idxw) > 0) {
dec <- table(factor(bi_row[idxw], levels = seq_len(nrow(Nb))),
factor(bi_col[idxw], levels = seq_len(k)))
Nw <- Nw - matrix(as.numeric(dec), nrow = nrow(Nb), ncol = k)
}
mw <- rowSums(Nw)
keepw <- which(mw >= 2)
if (length(keepw) >= 5) {
sqw <- rowSums(Nw[keepw, , drop = FALSE]^2)
Pw <- (sqw - mw[keepw]) / (mw[keepw] * (mw[keepw] - 1))
totw <- sum(mw[keepw])
pw <- colSums(Nw[keepw, , drop = FALSE]) / totw
Pew <- sum(pw^2)
rater_wo[r] <- if ((1 - Pew) > 1e-12) (mean(Pw) - Pew) / (1 - Pew) else NA_real_
} else {
rater_wo[r] <- NA_real_
}
}
rater_df <- data.frame(
rater = rater_names,
agreement_pct = round(100 * rater_obs, 2),
rater_kappa = round(rater_kap, 4),
kappa_without_rater = round(rater_wo, 4),
items_rated = rater_n,
stringsAsFactors = FALSE
)
ok_r <- which(!is.na(rater_kap))
outlier_rater <- if (length(ok_r) > 0) rater_names[ok_r[which.min(rater_kap[ok_r])]] else NA_character_
outlier_kappa <- if (length(ok_r) > 0) min(rater_kap[ok_r]) else NA_real_
best_rater <- if (length(ok_r) > 0) rater_names[ok_r[which.max(rater_kap[ok_r])]] else NA_character_
best_rater_kappa <- if (length(ok_r) > 0) max(rater_kap[ok_r]) else NA_real_
outlier_gain <- if (!is.na(outlier_rater) && !is.na(rater_wo[match(outlier_rater, rater_names)]))
rater_wo[match(outlier_rater, rater_names)] - kappa else NA_real_
rater_spread <- if (length(ok_r) > 1) max(rater_kap[ok_r]) - min(rater_kap[ok_r]) else NA_real_Step 13: Item-level consensus. Banded by each item's MODAL share —
the largest number of raters landing on one category, over the raters who judged it. Only the maximum COUNT is used, never its position, so an evenly split item needs no arbitrary winner.
modal_share <- apply(Nb, 1, max) / mb
band_labels <- c("No majority — the raters split",
"Simple majority — more than half",
"Strong majority — more than two thirds",
"Unanimous — every rater agreed")
band_idx <- ifelse(modal_share >= 1, 4L,
ifelse(modal_share > 2 / 3, 3L,
ifelse(modal_share > 0.5, 2L, 1L)))
item_df <- data.frame(
consensus_band = band_labels,
n_items = as.numeric(tabulate(band_idx, nbins = 4)),
stringsAsFactors = FALSE
)
item_df$share_pct <- round(100 * item_df$n_items / n_basis, 2)
n_unanimous <- item_df$n_items[4]
n_split <- item_df$n_items[1]Step 14: Per-rater marginals — how each rater uses the scale. These
are the habits behind both the chance floor and any outlier.
marg <- matrix(0, nrow = n_raters, ncol = k, dimnames = list(rater_names, levs))
tabr <- table(factor(bi_rater, levels = rater_names),
factor(bi_lab, levels = levs))
marg[] <- as.numeric(tabr)
marg_share <- marg / pmax(rowSums(marg), 1)
overall_share <- colSums(marg) / sum(marg)
divergence <- rowSums(abs(sweep(marg_share, 2, overall_share)))
show_r <- rater_names
n_raters_hidden <- 0L
if (n_raters > 8) {
show_r <- rater_names[order(-divergence)][1:8]
n_raters_hidden <- n_raters - 8L
}
marginal_df <- data.frame(
category = rep(levs, times = length(show_r)),
rater = rep(show_r, each = k),
share_pct = round(100 * as.vector(t(marg_share[show_r, , drop = FALSE])), 2),
stringsAsFactors = FALSE
)
prev_i <- which.max(p_j)
max_prev <- p_j[prev_i]
prev_cat <- levs[prev_i]
paradox <- isTRUE(p_bar >= 0.70 && !is.na(kappa) && kappa < 0.60 && max_prev >= 0.60)Step 15: The results table
basis_phrase <- if (basis == "complete")
sprintf("the %s item(s) every one of the %d raters judged", fmt_n(n_basis), n_raters)
else
sprintf("all %s item(s) carrying at least two judgements", fmt_n(n_basis))
rows <- list(
list("Raw pairwise agreement", p_bar, NA_real_, NA_real_,
sprintf("Across every pair of raters who judged the same item, %s of those pairs chose the same category.",
pct1(p_bar))),
list("Agreement expected by chance", p_e, NA_real_, NA_real_,
sprintf("What a panel using the categories this often would have agreed on by luck alone(%s).",
pct1(p_e))),
list("Fleiss' kappa", kappa, ci_low, ci_high,
sprintf("%s of the agreement left available above chance was achieved, on %s — %s on the Landis & Koch convention.",
pct1(kappa), basis_phrase, band))
)
if (n_missing > 0 && basis == "complete") {
rows[[length(rows) + 1]] <- list(
"Fleiss' kappa (all items, generalized)", kappa_all, NA_real_, NA_real_,
sprintf("The same coefficient generalized to a variable number of raters per item, so it uses all %s items with two or more judgements instead of only the %s complete ones.",
fmt_n(length(all_idx)), fmt_n(n_basis)))
}
rows[[length(rows) + 1]] <- list(
"Krippendorff's alpha (nominal)", alpha_nom, alpha_nom_lo, alpha_nom_hi,
sprintf("The disagreement-based coefficient, computed on all %s items with two or more judgements — it does not require every rater to have judged every item.",
fmt_n(length(all_idx))))
if (!is.na(alpha_ord)) {
rows[[length(rows) + 1]] <- list(
"Krippendorff's alpha (ordinal)", alpha_ord, alpha_ord_lo, alpha_ord_hi,
"The same coefficient with an ordinal distance between categories, so a near-miss between neighbouring levels counts as a smaller disagreement than a jump across the scale.")
}
if (n_missing > 0 && basis == "complete" && !is.na(alpha_nom_basis)) {
rows[[length(rows) + 1]] <- list(
"Krippendorff's alpha (complete items only)", alpha_nom_basis, NA_real_, NA_real_,
"Alpha restricted to the same complete items Fleiss' kappa used, so the two headline coefficients can be compared on identical data.")
}
agreement_df <- data.frame(
statistic = vapply(rows, function(r) r[[1]], character(1)),
estimate = round(vapply(rows, function(r) as.numeric(r[[2]]), numeric(1)), 4),
ci_low = round(vapply(rows, function(r) as.numeric(r[[3]]), numeric(1)), 4),
ci_high = round(vapply(rows, function(r) as.numeric(r[[4]]), numeric(1)), 4),
interpretation = vapply(rows, function(r) r[[5]], character(1)),
stringsAsFactors = FALSE
)Step 16: Methods disclosure
se_line <- if (is.na(se_null)) {
sprintf("The classical null-hypothesis standard error is not available here, because the items on the analysis basis do not all carry the same number of raters. The significance test is therefore not reported; read the bootstrap interval instead, which does not need that assumption.")
} else {
sprintf("Fleiss' null-hypothesis standard error is %s, giving z = %s and %s against the hypothesis that agreement is no better than chance. That variance describes kappa when kappa is truly zero, so it belongs to this test and NOT to the confidence interval.",
r3(se_null), r2(z_stat), fmt_pp(p_value))
}
methods_df <- data.frame(
item = c(
"Design",
"Input shape",
"Raw pairwise agreement",
"Chance agreement",
"Fleiss' kappa",
"Significance test",
"Confidence intervals",
"Missing ratings",
"Krippendorff's alpha",
"Category order",
"Per-category kappa",
"Per-rater figures",
"Benchmark labels",
"What agreement is not"
),
detail = c(
sprintf("%s item(s) judged by %d rater(s) across %d categories, %s judgement(s) in total.",
fmt_n(length(items)), n_raters, k, fmt_n(n_ratings)),
sprintf("Detected from the data, not from the mapping: %s.", shape_rule),
sprintf("Mean over items of the share of rater PAIRS on that item choosing the same category: %s.",
pct1(p_bar)),
sprintf("Sum of the squared overall category shares(%s): %s.",
paste(sprintf("%s %s", levs, vapply(p_j, pct1, character(1))), collapse = ", "),
pct1(p_e)),
sprintf("(observed - expected) / (1 - expected) = (%s - %s) / (1 - %s) = %s, computed on %s.",
r3(p_bar), r3(p_e), r3(p_e), r3(kappa), basis_phrase),
se_line,
sprintf("A seeded nonparametric bootstrap resampling ITEMS with replacement(%d resamples, fixed seed), taking the 2.5th and 97.5th percentiles. The bootstrap standard error of kappa is %s. This is used rather than a closed form because the null variance above is the wrong quantity for an interval around a non-zero estimate, and alpha has no simple closed-form variance.",
B, r3(se_boot)),
if (n_missing > 0)
sprintf("%s of the %s possible rater-by-item judgements are absent. Fleiss' kappa is a complete-cases statistic and used %s; Krippendorff's alpha used all %s items with at least two judgements; %s item(s) carrying a single judgement contribute to neither, because agreement needs at least two opinions.",
fmt_n(n_missing), fmt_n(n_raters * length(items)), basis_phrase,
fmt_n(length(all_idx)), fmt_n(n_items_dropped))
else
"No ratings are missing: every rater judged every item, so Fleiss' kappa and Krippendorff's alpha are computed on exactly the same data.",
sprintf("Built from the coincidence matrix: alpha = 1 - observed disagreement / expected disagreement, with(n-1) weighting so it is defined for small samples. Nominal alpha treats every disagreement as equal. %s",
if (!is.na(alpha_ord))
"Ordinal alpha weights a disagreement by the distance between the categories along the detected order."
else
"Ordinal alpha is not reported for these categories."),
sprintf("Orderedness was detected, not assumed: %s.", det$rule),
"Fleiss' category-specific coefficient, generalized to a variable number of raters per item: one minus the observed splits on that category over the splits expected from its overall share.",
sprintf("A rater's agreement figure is the share of the OTHER raters' judgements on the same items that matched this rater's own, chance-corrected against the same %s floor. The leave-one-out column recomputes Fleiss' kappa for the panel with that rater removed.",
pct1(p_e)),
"Landis & Koch(1977) labels(slight / fair / moderate / substantial / almost perfect) are a naming convention with no theoretical basis; the acceptable level of agreement depends on what the judgement is used for.",
"Agreement is not correctness: a panel can agree unanimously and be unanimously wrong. These figures describe this panel on these items and do not generalise to other raters or other items."
),
stringsAsFactors = FALSE
)Step 17: Headline metrics + the computed one-paragraph answer
metrics <- list(
`Items Rated` = as.numeric(length(items)),
`Raters` = as.numeric(n_raters),
`Categories` = as.numeric(k),
`Raw Agreement` = pct1(p_bar),
`Fleiss Kappa` = round(kappa, 3),
`Fleiss 95% CI` = paste0(r3(ci_low), " to ", r3(ci_high)),
`Krippendorff Alpha` = round(alpha_nom, 3),
`Agreement Strength` = band,
`Least Aligned Rater` = if (is.na(outlier_rater)) "not identifiable" else outlier_rater,
`Kappa vs Chance p` = fmt_p(p_value)
)
paradox_sentence <- if (paradox) {
paste0(
" Read the two agreement numbers together before quoting either: raw pairwise agreement is high(",
pct1(p_bar), ") while kappa is only ", r3(kappa),
", which is the well-known kappa paradox rather than a contradiction. ",
pct1(max_prev), " of all judgements fell into the single category \"", prev_cat,
"\", so a panel with these habits would already have agreed on ", pct1(p_e),
" of rater pairs by chance; only ", pct1(1 - p_e),
" of the scale was left for skill to win, and ", pct1(kappa), " of that remainder was won.")
} else {
paste0(
" Chance alone would have produced ", pct1(p_e),
" pairwise agreement given how often each category is used, leaving ",
pct1(1 - p_e), " of the scale available above chance, of which ",
pct1(kappa), " was achieved.")
}
missing_sentence <- if (n_missing > 0) {
paste0(" ", fmt_n(n_missing), " of the ", fmt_n(n_raters * length(items)),
" possible judgements are missing, which the two coefficients handle differently: Fleiss' kappa is a complete-cases statistic and used ",
basis_phrase, ", while Krippendorff's alpha used all ", fmt_n(length(all_idx)),
" items carrying at least two judgements",
if (n_items_dropped > 0)
paste0(" and ", fmt_n(n_items_dropped),
" item(s) with a single judgement were used by neither")
else "",
". Where ratings are missing, alpha is the coefficient to quote.")
} else ""
ordinal_sentence <- if (!is.na(alpha_ord)) {
paste0(" The categories are ordered(", det$rule,
"), so ordinal alpha(", r3(alpha_ord),
") is also reported; it counts a disagreement between neighbouring levels as smaller than a jump across the scale, which is why it sits above the nominal value of ",
r3(alpha_nom), ".")
} else {
paste0(" Ordinal alpha is not reported because ", det$rule,
"; treating unordered labels as if some disagreements were milder than others would invent structure the data does not carry.")
}
outlier_sentence <- if (!is.na(outlier_rater) && !is.na(outlier_gain)) {
paste0(" ", outlier_rater, " is the least aligned member of the panel(rater kappa ",
r3(outlier_kappa), " against ", r3(best_rater_kappa), " for ", best_rater,
"); with that rater excluded the panel's kappa would be ",
r3(kappa + outlier_gain),
if (outlier_gain > 0) paste0(", ", r3(outlier_gain), " higher than it is now")
else paste0(", ", r3(abs(outlier_gain)), " lower than it is now"),
".")
} else ""
json_output <- list(
answer = paste0(
fmt_n(n_raters), " raters judged ", fmt_n(length(items)), " items across ", k,
" categories, agreeing on ", pct1(p_bar),
" of rater pairs. Fleiss' kappa is ", r3(kappa),
" (95% CI ", r3(ci_low), " to ", r3(ci_high), "), ", band,
" agreement on the Landis & Koch convention, and ",
if (!is.na(p_value) && p_value < 0.05)
paste0("clearly better than chance(", fmt_pp(p_value), ")")
else if (!is.na(p_value))
paste0("not distinguishable from chance(", fmt_pp(p_value), ")")
else
"could not be tested against chance because the raters did not all judge the same items",
"; Krippendorff's alpha is ", r3(alpha_nom), ".",
paradox_sentence, missing_sentence, ordinal_sentence, outlier_sentence,
if (cat_degenerate)
paste0(" With only ", k,
" categories every disagreement is the same disagreement, so the category-specific kappas are necessarily equal to each other and to the overall coefficient — there is no per-category story to tell here.")
else
paste0(" The panel is least consistent on the category \"", worst_cat,
"\" (category kappa ", r3(worst_val), ") and most consistent on \"",
best_cat, "\" (", r3(best_val), ")."),
" Agreement is not correctness: a panel can agree unanimously and be unanimously wrong."
),
cards = lapply(
c("tldr", "overview", "preprocessing", "agreement_results", "rater_consensus",
"category_kappa", "item_consensus", "rater_marginals", "methods"),
function(cid) list(id = cid, metrics = metrics)
)
)
final_rows <- if (row_unit == "item") length(all_idx) else n_ratings
list(
initial_rows = initial_rows, final_rows = final_rows,
rows_removed = max(0, initial_rows - final_rows),
row_unit = row_unit, shape = shape, shape_rule = shape_rule,
n_blank_rows = n_blank_rows, n_dup_ratings = n_dup_ratings,
n_case_merged = n_case_merged, case_example = case_example, n_lumped = n_lumped,
col_names = ch, rater_names = rater_names, n_raters = n_raters,
levs = levs, k_cats = k, n_items = length(items), n_ratings = n_ratings,
n_missing = n_missing, n_items_partial = n_items_partial,
n_items_dropped = n_items_dropped, n_all = length(all_idx),
basis = basis, n_basis = n_basis, n_basis_ratings = tot_b,
basis_phrase = basis_phrase,
p_bar = p_bar, p_e = p_e, p_j = p_j, kappa = kappa, kappa_all = kappa_all,
se_null = se_null, se_boot = se_boot, z_stat = z_stat, p_value = p_value,
ci_low = ci_low, ci_high = ci_high,
alpha_nom = alpha_nom, alpha_nom_lo = alpha_nom_lo, alpha_nom_hi = alpha_nom_hi,
alpha_ord = alpha_ord, alpha_nom_basis = alpha_nom_basis,
ordered_flag = det$ordered, order_rule = det$rule, band = band,
paradox = paradox, paradox_sentence = paradox_sentence,
missing_sentence = missing_sentence, ordinal_sentence = ordinal_sentence,
outlier_sentence = outlier_sentence,
max_prev = max_prev, prev_cat = prev_cat,
outlier_rater = outlier_rater, outlier_kappa = outlier_kappa,
outlier_gain = outlier_gain, best_rater = best_rater,
best_rater_kappa = best_rater_kappa, rater_spread = rater_spread,
n_raters_hidden = n_raters_hidden,
worst_cat = worst_cat, worst_val = worst_val,
best_cat = best_cat, best_val = best_val, cat_degenerate = cat_degenerate,
n_unanimous = n_unanimous, n_split = n_split,
agreement_df = agreement_df, rater_df = rater_df, category_df = category_df,
item_df = item_df, marginal_df = marginal_df, methods_df = methods_df,
metrics = metrics, json_output = json_output, boot_B = B
)
}