Executive Summary
Rater reliability across 30 Sample ID values and 4 raters
The short answer
Your raters show good absolute agreement at ICC(2,1) = 0.782, and that reliability is statistically confirmed (p = 8.36e-29). Use this headline number if raters are a random sample and absolute Quality Score values matter. If decisions rest on the average of all 4 raters, reliability jumps to excellent at ICC(2,k) = 0.935.
The detail
Across 30 Sample IDs and 4 raters (120 ratings total), the headline intraclass correlation is ICC(2,1) = 0.782 (good, Koo & Li). The F-test (F = 21.91, p = 8.36e-29) confirms raters distinguish Sample ID values well above chance. Which ICC to report: ICC(2,1) if raters are a random sample and absolute scores matter (safest default); ICC(3,1) = 0.839 if these exact 4 raters are your permanent panel; ICC(2,k) = 0.935 if decisions use the panel average instead of any single rater.
What this can't tell you
The choice between ICC(2,1) and ICC(3,1) depends on your operational setup—whether you expect to hire new raters or keep these four. The data supports both interpretations; the decision is yours.
Analysis Overview
Intraclass correlation from 120 ratings: 30 Sample ID values x 4 raters.
The short answer
A single rater's reliability is good at 0.782 when absolute scores matter (ICC(2,1)), but rises to excellent at 0.935 when you average all 4 raters. The gap of 0.058 between consistency and absolute-agreement forms shows that systematic rater bias—some raters run high, others low—is eating into agreement. Choose ICC(2,1) = 0.782 as your headline unless these exact 4 raters are your permanent panel.
The detail
The analysis computed five ICC forms from 120 ratings of 30 Sample IDs by 4 raters. ICC(2,1) = 0.782 (good, per Koo & Li) is the most conservative choice and assumes raters are a random sample. ICC(3,1) = 0.839 forgives systematic offsets, showing the ordering is consistent even when absolute values drift. The 0.058 gap between them quantifies the cost of rater bias. ICC(2,k) = 0.935 (excellent) and ICC(3,k) = 0.954 show that pooling all four raters dramatically improves reliability. The F-test (F = 21.91, p = 8.36e-29) confirms raters are detecting real subject differences, not noise.
What this can't tell you
The ICC forms assume a two-way random or mixed model; if your use case fixes both raters and subjects, ICC(1,1) = 0.778 may be more appropriate, though the data structure here supports the two-way forms.
Data Quality
Ratings used, duplicates averaged, incomplete subjects excluded.
The short answer
All 120 rows were retained: no duplicates, no incomplete subjects. Every Sample ID was rated by all 4 raters in a complete crossing, so the 30 × 4 analysis grid has no gaps and no data loss.
The detail
The dataset loaded with 120 rows and 120 ratings were used in the final analysis. Zero duplicate ratings were found, so no averaging was needed. Every Sample ID value was rated by all 4 raters—a complete two-way design with no missing cells. The analysis requires one row per rating, with each Sample ID identified consistently across raters; that requirement was met. No subjects were excluded from the two-way ICCs.
What this can't tell you
The export does not reveal whether any individual rating was a typo, outlier, or data-entry error that passed consistency checks; a domain review of extreme or unusual scores would sharpen confidence in the raw values.
ICC Results — Which One to Report
All ICC forms with interpretation bands and when to use each.
| Icc Type | Value | Interpretation | Use When |
|---|---|---|---|
| ICC(1,1) — one-way random, single rater | 0.778 | good | Each Sample ID is rated by a DIFFERENT random set of raters (no shared panel). |
| ICC(2,1) — absolute agreement, single rater | 0.782 | good | Raters are a random sample and absolute Quality Score values must match — the default to report. |
| ICC(3,1) — consistency, single rater | 0.839 | good | These exact 4 raters are the only ones of interest; systematic differences between them are forgiven. |
| ICC(2,k) — absolute agreement, average of all raters | 0.935 | excellent | Decisions rest on the AVERAGE Quality Score of all 4 raters, and absolute values matter. |
| ICC(3,k) — consistency, average of all raters | 0.954 | excellent | Decisions rest on the average of these exact 4 raters; only relative ordering matters. |
The short answer
Report ICC(2,1) = 0.782 as your headline—good absolute agreement for a single rater. The 0.058 gap between ICC(3,1) = 0.839 (consistency) and ICC(2,1) = 0.782 (absolute) reveals that systematic rater bias is the problem. Averaging all 4 raters lifts reliability to excellent at ICC(2,k) = 0.935.
The detail
Five ICC forms emerge from the same ANOVA mean squares. ICC(1,1) = 0.778 applies only if each Sample ID is rated by a different random set of raters. ICC(2,1) = 0.782 is the default: raters as a random sample, absolute Quality Score values must match (good, Koo & Li). ICC(3,1) = 0.839 forgives systematic rater offsets; the 0.058 gap shows rater bias is real and worth addressing. ICC(2,k) = 0.935 (excellent) and ICC(3,k) = 0.954 (excellent) show that averaging all 4 raters recovers nearly perfect reliability. Interpretation bands: below 0.5 poor, 0.5–0.75 moderate, 0.75–0.9 good, above 0.9 excellent.
What this can't tell you
Which ICC form matches your operational use case is a business decision: whether raters are interchangeable, whether you use single or averaged scores, and whether these exact 4 raters are permanent or temporary.
Systematic Rater Bias
Mean Quality Score per rater on the same Sample ID values.
The short answer
Dr. Adams runs high at 51.79 and Dr. Chen runs low at 46.92—a spread of 4.87 points of pure leniency/severity bias. This systematic gap accounts for 6.9% of total variance, so rater calibration would meaningfully improve absolute agreement.
The detail
Each rater's mean Quality Score over the same 30 Sample IDs reveals systematic bias (leniency or severity), not differences in what was rated. Dr. Adams: 51.79, Dr. Diaz: 49.42, Dr. Baker: 48.15, Dr. Chen: 46.92. The 4.87-point spread between highest and lowest is pure bias. This systematic difference accounts for 6.9% of total variance in Quality Score. Calibrating raters against a shared standard (anchored examples, rubric review) would raise absolute agreement without changing the underlying ordering.
What this can't tell you
Whether the observed bias reflects true differences in rater stringency, differences in which samples each rater happened to rate first (order effects), or both; a randomized presentation order on a new round would clarify this.
Subject-Level Agreement
One point per Sample ID: the first two raters' scores, diagonal = perfect agreement.
The short answer
Dr. Adams and Dr. Baker correlate at r = 0.876, showing a tight upward pattern. Most points cluster near the diagonal, but several subjects show large vertical scatter—places where the two raters genuinely disagree on absolute value, even though their ordering is consistent.
The detail
Each of 30 points is one Sample ID: Dr. Adams's score (horizontal) vs Dr. Baker's score (vertical), with the diagonal = perfect agreement. The points form a tight band running upward and roughly parallel to the diagonal, consistent with the r = 0.876 correlation. Points hugging the diagonal mean absolute agreement; a band parallel but offset means consistent ordering with systematic bias—the ICC(3,1) vs ICC(2,1) story. Several subjects scatter far vertically from the diagonal (e.g., Sample IDs at Adams = 44.1, Baker = 36.5; Adams = 59.2, Baker = 49.1), indicating genuine disagreement on absolute value despite overall consistency.
What this can't tell you
Whether the scattered subjects reflect ambiguous rubric language, rater inattention, or genuine difficulty in the samples themselves; a qualitative review of high-scatter subjects would clarify whether rubric refinement or rater training is the better lever.
Where the Disagreement Comes From
Variance in Quality Score split between subjects, raters, and noise.
The short answer
Your raters are detecting real differences in quality reliably. Between-subject signal accounts for 78.2% of total variance—the largest share—which means most disagreement reflects genuine differences in the samples being rated, not rater inconsistency. Systematic rater bias is minimal at 6.9%, and residual noise is 15%, leaving room for improvement in rubric clarity but no urgent calibration crisis.
The detail
Variance decomposition shows: Between subjects 78.2%, Between raters 6.9%, Residual 15%. The dominance of between-subject variance (78.2%) indicates raters are successfully discriminating signal. The between-rater systematic component (6.9%) is small, suggesting raters are not systematically lenient or severe relative to one another. Residual noise at 15% captures one-off disagreements—the floor for any rating system without perfect rubric specification.
What this can't tell you
The residual 15% could reflect ambiguous rubric language, genuine sample difficulty that splits rater judgment, or rating carelessness; a breakdown by sample would show which specific items drive this noise. Consider a rubric audit or anchored scale examples to test whether clearer guidance would shrink that 15% further.
Intraclass Correlation — Rater Reliability
How consistent are your raters, instruments, or repeated measurements? From long-format ratings (one row per rating) this computes the full ICC family — ICC(1,1), ICC(2,1), ICC(3,1) plus the average-measure versions — from two-way ANOVA mean squares, with Koo & Li interpretation bands, a rater-bias check, a subject-level agreement plot, and a variance breakdown showing where the disagreement comes from.
Why This Method?
The ICC is the standard reliability coefficient for continuous ratings: it asks how much of the total variation comes from real differences between the things being rated rather than from rater disagreement. Computing every common form at once answers the perennial question of WHICH ICC to report for a given study design.
What This Analysis Covers
- Single-measure and average-measure ICCs, consistency and agreement
- Koo & Li (2016) interpretation bands
- Systematic rater bias (mean score per rater)
- Subject-level agreement between the first two raters
- Variance decomposition: subjects vs raters vs residual noise
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {subject, rater, score}. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Core Analysis Pipeline
compute_shared <- function(df, params, col_map = list()) {
# === SHARED EXPORTS ===
# initial_rows/final_rows/rows_removed $ row accounting (ratings)
# subj_h / rater_h / score_h $ humanized names of the three mapped columns
# n_subjects / k_raters $ complete-crossing design dimensions
# rater_levels $ character — the actual rater level values
# n_duplicates $ duplicate ratings averaged
# n_incomplete $ subjects excluded (not rated by all raters)
# incomplete_subjects $ character — the excluded subject ids
# icc / icc_p $ named list of ICC values + F-test p for ICC(3,1)
# icc_results_df $ data.frame(icc_type, value, interpretation, use_when)
# rater_means_df $ data.frame(rater_name, mean_score)
# agreement_df $ data.frame(rater_1_score, rater_2_score) — complete subjects
# agreement_r $ numeric — Pearson r between the first two raters
# variance_df $ data.frame(source, variance_pct)
# headline_band $ Koo & Li band for ICC(2,1)
# metrics / json_output
# === /SHARED EXPORTS ===
subj_h <- humanize_semantic("subject", col_map)
rater_h <- humanize_semantic("rater", col_map)
score_h <- humanize_semantic("score", col_map)Step 1: Check the mapped columns
initial_rows <- nrow(df)
for (req in c("subject", "rater", "score")) {
if (is.null(df[[req]])) {
stop(sprintf("column_mapping must map the '%s' column (%s / %s / %s).",
req, subj_h, rater_h, score_h))
}
}Step 2: Coerce score numeric (95% rule); clean subject/rater ids
v <- df$score
if (!is.numeric(v)) {
conv <- suppressWarnings(as.numeric(as.character(v)))
n_orig <- sum(!is.na(v) & as.character(v) != "")
if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) {
v <- conv
} else {
stop(sprintf("The rating column(%s) must be numeric — fewer than 95%% of its values convert cleanly.",
score_h))
}
}
df$score <- v
df$subject <- trimws(as.character(df$subject))
df$rater <- trimws(as.character(df$rater))
keep <- !is.na(df$score) & !is.na(df$subject) & df$subject != "" &
!is.na(df$rater) & df$rater != ""
n_invalid <- sum(!keep)
df <- df[keep, , drop = FALSE]
if (nrow(df) < 10) {
stop(sprintf("Only %d usable ratings after removing rows with missing %s, %s, or %s — need at least 10.",
nrow(df), subj_h, rater_h, score_h))
}
rater_levels <- sort(unique(df$rater))
k <- length(rater_levels)
if (k < 2) {
stop(sprintf("ICC needs at least 2 distinct values in the %s column; found %d. Each row should be one rating by one rater.",
rater_h, k))
}
if (k > 30) {
stop(sprintf("The %s column has %d distinct values — that looks like an identifier, not a set of raters. Map the column naming who/what gave each rating.",
rater_h, k))
}Step 3: Average duplicate ratings per subject x rater
agg <- aggregate(score ~ subject + rater, data = df, FUN = mean)
n_duplicates <- nrow(df) - nrow(agg)Step 4: Keep subjects rated by ALL raters (two-way ICCs need a
complete crossing); report the incomplete ones
cnt <- table(agg$subject)
complete_subjects <- names(cnt)[cnt == k]
incomplete_subjects <- names(cnt)[cnt < k]
n_incomplete <- length(incomplete_subjects)
n <- length(complete_subjects)
if (n < 5) {
stop(sprintf("ICC needs at least 5 %s values rated by every %s; only %d of %d have a complete set of %d ratings.",
subj_h, rater_h, n, length(cnt), k))
}
dat <- agg[agg$subject %in% complete_subjects, , drop = FALSE]
if (is.na(var(dat$score)) || isTRUE(var(dat$score) == 0)) {
stop(sprintf("The %s values are constant — reliability cannot be estimated.", score_h))
}
final_rows <- nrow(dat)
rows_removed <- initial_rows - final_rows
dat$subject_f <- factor(dat$subject)
dat$rater_f <- factor(dat$rater, levels = rater_levels)Step 5: ICCs from ANOVA mean squares
Two-way: aov(score ~ subject + rater) -> MSR (subjects), MSC (raters), MSE
a2 <- aov(score ~ subject_f + rater_f, data = dat)
ms2 <- summary(a2)[[1]][["Mean Sq"]]
MSR <- ms2[1]; MSC <- ms2[2]; MSE <- ms2[3]
icc21 <- (MSR - MSE) / (MSR + (k - 1) * MSE + k * (MSC - MSE) / n)
icc31 <- (MSR - MSE) / (MSR + (k - 1) * MSE)
icc2k <- (MSR - MSE) / (MSR + (MSC - MSE) / n)
icc3k <- (MSR - MSE) / MSROne-way: aov(score ~ subject) -> MSB, MSW
a1 <- aov(score ~ subject_f, data = dat)
ms1 <- summary(a1)[[1]][["Mean Sq"]]
MSB <- ms1[1]; MSW <- ms1[2]
icc11 <- (MSB - MSW) / (MSB + (k - 1) * MSW)F-test for ICC(3,1): MSR/MSE against F(n-1, (n-1)(k-1))
f_stat <- MSR / MSE
f_df1 <- n - 1
f_df2 <- (n - 1) * (k - 1)
f_p <- pf(f_stat, f_df1, f_df2, lower.tail = FALSE)Koo & Li (2016) bands: <0.5 poor, 0.5-0.75 moderate, 0.75-0.9 good, >0.9 excellent
koo_li <- function(x) {
if (is.na(x)) return("not estimable")
if (x < 0.5) "poor" else if (x < 0.75) "moderate" else if (x < 0.9) "good" else "excellent"
}
headline_band <- koo_li(icc21)
icc_results_df <- data.frame(
icc_type = c("ICC(1,1) — one-way random, single rater",
"ICC(2,1) — absolute agreement, single rater",
"ICC(3,1) — consistency, single rater",
"ICC(2,k) — absolute agreement, average of all raters",
"ICC(3,k) — consistency, average of all raters"),
value = round(c(icc11, icc21, icc31, icc2k, icc3k), 3),
interpretation = sapply(c(icc11, icc21, icc31, icc2k, icc3k), koo_li),
use_when = c(
sprintf("Each %s is rated by a DIFFERENT random set of raters(no shared panel).", subj_h),
sprintf("Raters are a random sample and absolute %s values must match — the default to report.", score_h),
sprintf("These exact %d raters are the only ones of interest; systematic differences between them are forgiven.", k),
sprintf("Decisions rest on the AVERAGE %s of all %d raters, and absolute values matter.", score_h, k),
sprintf("Decisions rest on the average of these exact %d raters; only relative ordering matters.", k)
),
stringsAsFactors = FALSE
)Step 6: Variance components (from the two-way mean squares)
var_subject <- max(0, (MSR - MSE) / k)
var_rater <- max(0, (MSC - MSE) / n)
var_error <- max(0, MSE)
var_total <- var_subject + var_rater + var_error
if (var_total <= 0) var_total <- 1
variance_df <- data.frame(
source = c("Between subjects", "Between raters", "Residual"),
variance_pct = round(100 * c(var_subject, var_rater, var_error) / var_total, 1),
stringsAsFactors = FALSE
)Step 7: Rater means (systematic bias) on the complete-crossing data
rm_agg <- aggregate(score ~ rater_f, data = dat, FUN = mean)
rater_means_df <- data.frame(
rater_name = as.character(rm_agg$rater_f),
mean_score = round(rm_agg$score, 2),
stringsAsFactors = FALSE
)NA-safe extremes (never which.max over possibly-NA vectors)
ok_idx <- which(!is.na(rater_means_df$mean_score))
hi_idx <- ok_idx[which.max(rater_means_df$mean_score[ok_idx])]
lo_idx <- ok_idx[which.min(rater_means_df$mean_score[ok_idx])]
rater_spread <- rater_means_df$mean_score[hi_idx] - rater_means_df$mean_score[lo_idx]Step 8: Subject-level agreement — first two raters, one point per subject
d1 <- dat[dat$rater == rater_levels[1], c("subject", "score")]
d2 <- dat[dat$rater == rater_levels[2], c("subject", "score")]
names(d1)[2] <- "rater_1_score"
names(d2)[2] <- "rater_2_score"
merged <- merge(d1, d2, by = "subject")
merged <- merged[order(merged$rater_1_score), , drop = FALSE]
set.seed(42)
if (nrow(merged) > 1000) merged <- merged[sample(nrow(merged), 1000), , drop = FALSE]
agreement_df <- data.frame(
rater_1_score = round(merged$rater_1_score, 3),
rater_2_score = round(merged$rater_2_score, 3),
stringsAsFactors = FALSE
)
agreement_r <- if (nrow(agreement_df) >= 3)
suppressWarnings(cor(agreement_df$rater_1_score, agreement_df$rater_2_score))
else NA_real_
icc <- list(icc11 = icc11, icc21 = icc21, icc31 = icc31,
icc2k = icc2k, icc3k = icc3k)
metrics <- list(
`Subjects` = n,
`Raters` = k,
`Ratings Used` = final_rows,
`ICC(2,1)` = round(icc21, 3),
`Reliability` = headline_band,
`F-test p` = signif(f_p, 3)
)
json_output <- list(
answer = paste0(
"Intraclass correlation from ", format(final_rows, big.mark = ","),
" ratings of ", n, " ", subj_h, " values by ", k, " raters(", rater_h,
"): ICC(2,1) = ", round(icc21, 3), " — ", headline_band,
" absolute agreement for a single rater(Koo & Li). Consistency ICC(3,1) = ",
round(icc31, 3), "; average-measure ICC(2,k) = ", round(icc2k, 3),
". The F-test for subject discrimination is ",
ifelse(is.na(f_p), "not estimable",
ifelse(f_p < 0.05, paste0("significant(p = ", signif(f_p, 3), ")"),
paste0("not significant(p = ", signif(f_p, 3), ")"))),
". Between-", subj_h, " differences explain ", variance_df$variance_pct[1],
"% of the variance in ", score_h, "."
),
cards = lapply(
c("tldr", "overview", "preprocessing", "icc_results",
"rater_means", "subject_agreement", "variance_breakdown"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
initial_rows = initial_rows, final_rows = final_rows,
rows_removed = rows_removed, n_invalid = n_invalid,
subj_h = subj_h, rater_h = rater_h, score_h = score_h,
n_subjects = n, k_raters = k, rater_levels = rater_levels,
n_duplicates = n_duplicates, n_incomplete = n_incomplete,
incomplete_subjects = incomplete_subjects,
icc = icc, f_stat = f_stat, f_df1 = f_df1, f_df2 = f_df2, f_p = f_p,
icc_results_df = icc_results_df,
rater_means_df = rater_means_df,
rater_hi = rater_means_df$rater_name[hi_idx],
rater_lo = rater_means_df$rater_name[lo_idx],
rater_spread = rater_spread,
agreement_df = agreement_df, agreement_r = agreement_r,
variance_df = variance_df,
headline_band = headline_band,
metrics = metrics, json_output = json_output
)
}Your turn
Bring your own data and the question you actually need answered.
CympleData Scientist Send me your data and question, I’ll send you the analytics. ds@mcpanalytics.ai