Standard Icc
Executive Summary

Executive Summary

Rater reliability across 30 Sample ID values and 4 raters

Subjects
30
Raters
4
Ratings Used
120
ICC(2,1)
0.782
Reliability
good
F-test p
8.36e-29
The headline reliability is ICC(2,1) = 0.782 — GOOD absolute agreement (Koo & Li) for a single rater's Quality Score across 30 Sample ID values and 4 raters. That the raters distinguish Sample ID values at all is statistically confirmed (F = 21.91, p = 8.36e-29). Which ICC to report: use ICC(2,1) if your raters are a random sample and absolute Quality Score values matter (the safest default); ICC(3,1) = 0.839 if these exact 4 raters are the only ones of interest; ICC(2,k) = 0.935 if decisions rest on the average of all 4 raters rather than any single one.
What this means

The short answer

Your raters show good absolute agreement at ICC(2,1) = 0.782, and that reliability is statistically confirmed (p = 8.36e-29). Use this headline number if raters are a random sample and absolute Quality Score values matter. If decisions rest on the average of all 4 raters, reliability jumps to excellent at ICC(2,k) = 0.935.

The detail

Across 30 Sample IDs and 4 raters (120 ratings total), the headline intraclass correlation is ICC(2,1) = 0.782 (good, Koo & Li). The F-test (F = 21.91, p = 8.36e-29) confirms raters distinguish Sample ID values well above chance. Which ICC to report: ICC(2,1) if raters are a random sample and absolute scores matter (safest default); ICC(3,1) = 0.839 if these exact 4 raters are your permanent panel; ICC(2,k) = 0.935 if decisions use the panel average instead of any single rater.

What this can't tell you

The choice between ICC(2,1) and ICC(3,1) depends on your operational setup—whether you expect to hire new raters or keep these four. The data supports both interpretations; the decision is yours.

Overview

Analysis Overview

Intraclass correlation from 120 ratings: 30 Sample ID values x 4 raters.

N Subjects30
N Raters4
N Ratings120
Icc 2 10.782
What this means

The short answer

A single rater's reliability is good at 0.782 when absolute scores matter (ICC(2,1)), but rises to excellent at 0.935 when you average all 4 raters. The gap of 0.058 between consistency and absolute-agreement forms shows that systematic rater bias—some raters run high, others low—is eating into agreement. Choose ICC(2,1) = 0.782 as your headline unless these exact 4 raters are your permanent panel.

The detail

The analysis computed five ICC forms from 120 ratings of 30 Sample IDs by 4 raters. ICC(2,1) = 0.782 (good, per Koo & Li) is the most conservative choice and assumes raters are a random sample. ICC(3,1) = 0.839 forgives systematic offsets, showing the ordering is consistent even when absolute values drift. The 0.058 gap between them quantifies the cost of rater bias. ICC(2,k) = 0.935 (excellent) and ICC(3,k) = 0.954 show that pooling all four raters dramatically improves reliability. The F-test (F = 21.91, p = 8.36e-29) confirms raters are detecting real subject differences, not noise.

What this can't tell you

The ICC forms assume a two-way random or mixed model; if your use case fixes both raters and subjects, ICC(1,1) = 0.778 may be more appropriate, though the data structure here supports the two-way forms.

Data Preparation

Data Quality

Ratings used, duplicates averaged, incomplete subjects excluded.

Initial Rows120
Final Rows120
Rows Removed0
N Duplicates0
N Incomplete Subjects0
What this means

The short answer

All 120 rows were retained: no duplicates, no incomplete subjects. Every Sample ID was rated by all 4 raters in a complete crossing, so the 30 × 4 analysis grid has no gaps and no data loss.

The detail

The dataset loaded with 120 rows and 120 ratings were used in the final analysis. Zero duplicate ratings were found, so no averaging was needed. Every Sample ID value was rated by all 4 raters—a complete two-way design with no missing cells. The analysis requires one row per rating, with each Sample ID identified consistently across raters; that requirement was met. No subjects were excluded from the two-way ICCs.

What this can't tell you

The export does not reveal whether any individual rating was a typo, outlier, or data-entry error that passed consistency checks; a domain review of extreme or unusual scores would sharpen confidence in the raw values.

Data Table

ICC Results — Which One to Report

All ICC forms with interpretation bands and when to use each.

Icc TypeValueInterpretationUse When
ICC(1,1) — one-way random, single rater0.778goodEach Sample ID is rated by a DIFFERENT random set of raters (no shared panel).
ICC(2,1) — absolute agreement, single rater0.782goodRaters are a random sample and absolute Quality Score values must match — the default to report.
ICC(3,1) — consistency, single rater0.839goodThese exact 4 raters are the only ones of interest; systematic differences between them are forgiven.
ICC(2,k) — absolute agreement, average of all raters0.935excellentDecisions rest on the AVERAGE Quality Score of all 4 raters, and absolute values matter.
ICC(3,k) — consistency, average of all raters0.954excellentDecisions rest on the average of these exact 4 raters; only relative ordering matters.
What this means

The short answer

Report ICC(2,1) = 0.782 as your headline—good absolute agreement for a single rater. The 0.058 gap between ICC(3,1) = 0.839 (consistency) and ICC(2,1) = 0.782 (absolute) reveals that systematic rater bias is the problem. Averaging all 4 raters lifts reliability to excellent at ICC(2,k) = 0.935.

The detail

Five ICC forms emerge from the same ANOVA mean squares. ICC(1,1) = 0.778 applies only if each Sample ID is rated by a different random set of raters. ICC(2,1) = 0.782 is the default: raters as a random sample, absolute Quality Score values must match (good, Koo & Li). ICC(3,1) = 0.839 forgives systematic rater offsets; the 0.058 gap shows rater bias is real and worth addressing. ICC(2,k) = 0.935 (excellent) and ICC(3,k) = 0.954 (excellent) show that averaging all 4 raters recovers nearly perfect reliability. Interpretation bands: below 0.5 poor, 0.5–0.75 moderate, 0.75–0.9 good, above 0.9 excellent.

What this can't tell you

Which ICC form matches your operational use case is a business decision: whether raters are interchangeable, whether you use single or averaged scores, and whether these exact 4 raters are permanent or temporary.

Visualization

Systematic Rater Bias

Mean Quality Score per rater on the same Sample ID values.

What this means

The short answer

Dr. Adams runs high at 51.79 and Dr. Chen runs low at 46.92—a spread of 4.87 points of pure leniency/severity bias. This systematic gap accounts for 6.9% of total variance, so rater calibration would meaningfully improve absolute agreement.

The detail

Each rater's mean Quality Score over the same 30 Sample IDs reveals systematic bias (leniency or severity), not differences in what was rated. Dr. Adams: 51.79, Dr. Diaz: 49.42, Dr. Baker: 48.15, Dr. Chen: 46.92. The 4.87-point spread between highest and lowest is pure bias. This systematic difference accounts for 6.9% of total variance in Quality Score. Calibrating raters against a shared standard (anchored examples, rubric review) would raise absolute agreement without changing the underlying ordering.

What this can't tell you

Whether the observed bias reflects true differences in rater stringency, differences in which samples each rater happened to rate first (order effects), or both; a randomized presentation order on a new round would clarify this.

Visualization

Subject-Level Agreement

One point per Sample ID: the first two raters' scores, diagonal = perfect agreement.

What this means

The short answer

Dr. Adams and Dr. Baker correlate at r = 0.876, showing a tight upward pattern. Most points cluster near the diagonal, but several subjects show large vertical scatter—places where the two raters genuinely disagree on absolute value, even though their ordering is consistent.

The detail

Each of 30 points is one Sample ID: Dr. Adams's score (horizontal) vs Dr. Baker's score (vertical), with the diagonal = perfect agreement. The points form a tight band running upward and roughly parallel to the diagonal, consistent with the r = 0.876 correlation. Points hugging the diagonal mean absolute agreement; a band parallel but offset means consistent ordering with systematic bias—the ICC(3,1) vs ICC(2,1) story. Several subjects scatter far vertically from the diagonal (e.g., Sample IDs at Adams = 44.1, Baker = 36.5; Adams = 59.2, Baker = 49.1), indicating genuine disagreement on absolute value despite overall consistency.

What this can't tell you

Whether the scattered subjects reflect ambiguous rubric language, rater inattention, or genuine difficulty in the samples themselves; a qualitative review of high-scatter subjects would clarify whether rubric refinement or rater training is the better lever.

Visualization

Where the Disagreement Comes From

Variance in Quality Score split between subjects, raters, and noise.

What this means

The short answer

Your raters are detecting real differences in quality reliably. Between-subject signal accounts for 78.2% of total variance—the largest share—which means most disagreement reflects genuine differences in the samples being rated, not rater inconsistency. Systematic rater bias is minimal at 6.9%, and residual noise is 15%, leaving room for improvement in rubric clarity but no urgent calibration crisis.

The detail

Variance decomposition shows: Between subjects 78.2%, Between raters 6.9%, Residual 15%. The dominance of between-subject variance (78.2%) indicates raters are successfully discriminating signal. The between-rater systematic component (6.9%) is small, suggesting raters are not systematically lenient or severe relative to one another. Residual noise at 15% captures one-off disagreements—the floor for any rating system without perfect rubric specification.

What this can't tell you

The residual 15% could reflect ambiguous rubric language, genuine sample difficulty that splits rater judgment, or rating carelessness; a breakdown by sample would show which specific items drive this noise. Consider a rubric audit or anchored scale examples to test whether clearer guidance would shrink that 15% further.

Rate this report Was this the answer you needed?
The exact source that produced this report — yours to keep, read, and re-run.
Download PDF
How this was computed method · R source · citation
The code that did it

Intraclass Correlation — Rater Reliability

How consistent are your raters, instruments, or repeated measurements? From long-format ratings (one row per rating) this computes the full ICC family — ICC(1,1), ICC(2,1), ICC(3,1) plus the average-measure versions — from two-way ANOVA mean squares, with Koo & Li interpretation bands, a rater-bias check, a subject-level agreement plot, and a variance breakdown showing where the disagreement comes from.

Why This Method?

The ICC is the standard reliability coefficient for continuous ratings: it asks how much of the total variation comes from real differences between the things being rated rather than from rater disagreement. Computing every common form at once answers the perennial question of WHICH ICC to report for a given study design.

What This Analysis Covers

  • Single-measure and average-measure ICCs, consistency and agreement
  • Koo & Li (2016) interpretation bands
  • Systematic rater bias (mean score per rater)
  • Subject-level agreement between the first two raters
  • Variance decomposition: subjects vs raters vs residual noise

Standard Library

Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {subject, rater, score}. All narrative is derived from the user's own column names and computed values.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Core Analysis Pipeline

compute_shared <- function(df, params, col_map = list()) {
  # === SHARED EXPORTS ===
  #   initial_rows/final_rows/rows_removed  $ row accounting (ratings)
  #   subj_h / rater_h / score_h  $ humanized names of the three mapped columns
  #   n_subjects / k_raters       $ complete-crossing design dimensions
  #   rater_levels                $ character — the actual rater level values
  #   n_duplicates                $ duplicate ratings averaged
  #   n_incomplete                $ subjects excluded (not rated by all raters)
  #   incomplete_subjects         $ character — the excluded subject ids
  #   icc / icc_p                 $ named list of ICC values + F-test p for ICC(3,1)
  #   icc_results_df   $ data.frame(icc_type, value, interpretation, use_when)
  #   rater_means_df   $ data.frame(rater_name, mean_score)
  #   agreement_df     $ data.frame(rater_1_score, rater_2_score) — complete subjects
  #   agreement_r      $ numeric — Pearson r between the first two raters
  #   variance_df      $ data.frame(source, variance_pct)
  #   headline_band    $ Koo & Li band for ICC(2,1)
  #   metrics / json_output
  # === /SHARED EXPORTS ===

  subj_h  <- humanize_semantic("subject", col_map)
  rater_h <- humanize_semantic("rater", col_map)
  score_h <- humanize_semantic("score", col_map)

Step 1: Check the mapped columns

initial_rows <- nrow(df)
  for (req in c("subject", "rater", "score")) {
    if (is.null(df[[req]])) {
      stop(sprintf("column_mapping must map the &#x27;%s' column (%s / %s / %s).",
                   req, subj_h, rater_h, score_h))
    }
  }

Step 2: Coerce score numeric (95% rule); clean subject/rater ids

v <- df$score
  if (!is.numeric(v)) {
    conv <- suppressWarnings(as.numeric(as.character(v)))
    n_orig <- sum(!is.na(v) & as.character(v) != "")
    if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) {
      v <- conv
    } else {
      stop(sprintf("The rating column(%s) must be numeric — fewer than 95%% of its values convert cleanly.",
                   score_h))
    }
  }
  df$score   <- v
  df$subject <- trimws(as.character(df$subject))
  df$rater   <- trimws(as.character(df$rater))
  keep <- !is.na(df$score) & !is.na(df$subject) & df$subject != "" &
          !is.na(df$rater) & df$rater != ""
  n_invalid <- sum(!keep)
  df <- df[keep, , drop = FALSE]
  if (nrow(df) < 10) {
    stop(sprintf("Only %d usable ratings after removing rows with missing %s, %s, or %s — need at least 10.",
                 nrow(df), subj_h, rater_h, score_h))
  }

  rater_levels <- sort(unique(df$rater))
  k <- length(rater_levels)
  if (k < 2) {
    stop(sprintf("ICC needs at least 2 distinct values in the %s column; found %d. Each row should be one rating by one rater.",
                 rater_h, k))
  }
  if (k > 30) {
    stop(sprintf("The %s column has %d distinct values — that looks like an identifier, not a set of raters. Map the column naming who/what gave each rating.",
                 rater_h, k))
  }

Step 3: Average duplicate ratings per subject x rater

agg <- aggregate(score ~ subject + rater, data = df, FUN = mean)
  n_duplicates <- nrow(df) - nrow(agg)

Step 4: Keep subjects rated by ALL raters (two-way ICCs need a

complete crossing); report the incomplete ones

cnt <- table(agg$subject)
  complete_subjects   <- names(cnt)[cnt == k]
  incomplete_subjects <- names(cnt)[cnt < k]
  n_incomplete <- length(incomplete_subjects)
  n <- length(complete_subjects)
  if (n < 5) {
    stop(sprintf("ICC needs at least 5 %s values rated by every %s; only %d of %d have a complete set of %d ratings.",
                 subj_h, rater_h, n, length(cnt), k))
  }
  dat <- agg[agg$subject %in% complete_subjects, , drop = FALSE]
  if (is.na(var(dat$score)) || isTRUE(var(dat$score) == 0)) {
    stop(sprintf("The %s values are constant — reliability cannot be estimated.", score_h))
  }
  final_rows <- nrow(dat)
  rows_removed <- initial_rows - final_rows

  dat$subject_f <- factor(dat$subject)
  dat$rater_f   <- factor(dat$rater, levels = rater_levels)

Step 5: ICCs from ANOVA mean squares

Two-way: aov(score ~ subject + rater) -> MSR (subjects), MSC (raters), MSE

a2  <- aov(score ~ subject_f + rater_f, data = dat)
  ms2 <- summary(a2)[[1]][["Mean Sq"]]
  MSR <- ms2[1]; MSC <- ms2[2]; MSE <- ms2[3]

  icc21 <- (MSR - MSE) / (MSR + (k - 1) * MSE + k * (MSC - MSE) / n)
  icc31 <- (MSR - MSE) / (MSR + (k - 1) * MSE)
  icc2k <- (MSR - MSE) / (MSR + (MSC - MSE) / n)
  icc3k <- (MSR - MSE) / MSR

One-way: aov(score ~ subject) -> MSB, MSW

a1  <- aov(score ~ subject_f, data = dat)
  ms1 <- summary(a1)[[1]][["Mean Sq"]]
  MSB <- ms1[1]; MSW <- ms1[2]
  icc11 <- (MSB - MSW) / (MSB + (k - 1) * MSW)

F-test for ICC(3,1): MSR/MSE against F(n-1, (n-1)(k-1))

f_stat <- MSR / MSE
  f_df1  <- n - 1
  f_df2  <- (n - 1) * (k - 1)
  f_p    <- pf(f_stat, f_df1, f_df2, lower.tail = FALSE)

Koo & Li (2016) bands: <0.5 poor, 0.5-0.75 moderate, 0.75-0.9 good, >0.9 excellent

koo_li <- function(x) {
    if (is.na(x)) return("not estimable")
    if (x < 0.5) "poor" else if (x < 0.75) "moderate" else if (x < 0.9) "good" else "excellent"
  }
  headline_band <- koo_li(icc21)

  icc_results_df <- data.frame(
    icc_type = c("ICC(1,1) — one-way random, single rater",
                 "ICC(2,1) — absolute agreement, single rater",
                 "ICC(3,1) — consistency, single rater",
                 "ICC(2,k) — absolute agreement, average of all raters",
                 "ICC(3,k) — consistency, average of all raters"),
    value = round(c(icc11, icc21, icc31, icc2k, icc3k), 3),
    interpretation = sapply(c(icc11, icc21, icc31, icc2k, icc3k), koo_li),
    use_when = c(
      sprintf("Each %s is rated by a DIFFERENT random set of raters(no shared panel).", subj_h),
      sprintf("Raters are a random sample and absolute %s values must match — the default to report.", score_h),
      sprintf("These exact %d raters are the only ones of interest; systematic differences between them are forgiven.", k),
      sprintf("Decisions rest on the AVERAGE %s of all %d raters, and absolute values matter.", score_h, k),
      sprintf("Decisions rest on the average of these exact %d raters; only relative ordering matters.", k)
    ),
    stringsAsFactors = FALSE
  )

Step 6: Variance components (from the two-way mean squares)

var_subject <- max(0, (MSR - MSE) / k)
  var_rater   <- max(0, (MSC - MSE) / n)
  var_error   <- max(0, MSE)
  var_total   <- var_subject + var_rater + var_error
  if (var_total <= 0) var_total <- 1
  variance_df <- data.frame(
    source = c("Between subjects", "Between raters", "Residual"),
    variance_pct = round(100 * c(var_subject, var_rater, var_error) / var_total, 1),
    stringsAsFactors = FALSE
  )

Step 7: Rater means (systematic bias) on the complete-crossing data

rm_agg <- aggregate(score ~ rater_f, data = dat, FUN = mean)
  rater_means_df <- data.frame(
    rater_name = as.character(rm_agg$rater_f),
    mean_score = round(rm_agg$score, 2),
    stringsAsFactors = FALSE
  )

NA-safe extremes (never which.max over possibly-NA vectors)

ok_idx <- which(!is.na(rater_means_df$mean_score))
  hi_idx <- ok_idx[which.max(rater_means_df$mean_score[ok_idx])]
  lo_idx <- ok_idx[which.min(rater_means_df$mean_score[ok_idx])]
  rater_spread <- rater_means_df$mean_score[hi_idx] - rater_means_df$mean_score[lo_idx]

Step 8: Subject-level agreement — first two raters, one point per subject

d1 <- dat[dat$rater == rater_levels[1], c("subject", "score")]
  d2 <- dat[dat$rater == rater_levels[2], c("subject", "score")]
  names(d1)[2] <- "rater_1_score"
  names(d2)[2] <- "rater_2_score"
  merged <- merge(d1, d2, by = "subject")
  merged <- merged[order(merged$rater_1_score), , drop = FALSE]
  set.seed(42)
  if (nrow(merged) > 1000) merged <- merged[sample(nrow(merged), 1000), , drop = FALSE]
  agreement_df <- data.frame(
    rater_1_score = round(merged$rater_1_score, 3),
    rater_2_score = round(merged$rater_2_score, 3),
    stringsAsFactors = FALSE
  )
  agreement_r <- if (nrow(agreement_df) >= 3)
    suppressWarnings(cor(agreement_df$rater_1_score, agreement_df$rater_2_score))
  else NA_real_

  icc <- list(icc11 = icc11, icc21 = icc21, icc31 = icc31,
              icc2k = icc2k, icc3k = icc3k)

  metrics <- list(
    `Subjects`     = n,
    `Raters`       = k,
    `Ratings Used` = final_rows,
    `ICC(2,1)`     = round(icc21, 3),
    `Reliability`  = headline_band,
    `F-test p`     = signif(f_p, 3)
  )

  json_output <- list(
    answer = paste0(
      "Intraclass correlation from ", format(final_rows, big.mark = ","),
      " ratings of ", n, " ", subj_h, " values by ", k, " raters(", rater_h,
      "): ICC(2,1) = ", round(icc21, 3), " — ", headline_band,
      " absolute agreement for a single rater(Koo & Li). Consistency ICC(3,1) = ",
      round(icc31, 3), "; average-measure ICC(2,k) = ", round(icc2k, 3),
      ". The F-test for subject discrimination is ",
      ifelse(is.na(f_p), "not estimable",
             ifelse(f_p < 0.05, paste0("significant(p = ", signif(f_p, 3), ")"),
                    paste0("not significant(p = ", signif(f_p, 3), ")"))),
      ". Between-", subj_h, " differences explain ", variance_df$variance_pct[1],
      "% of the variance in ", score_h, "."
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "icc_results",
        "rater_means", "subject_agreement", "variance_breakdown"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  list(
    initial_rows = initial_rows, final_rows = final_rows,
    rows_removed = rows_removed, n_invalid = n_invalid,
    subj_h = subj_h, rater_h = rater_h, score_h = score_h,
    n_subjects = n, k_raters = k, rater_levels = rater_levels,
    n_duplicates = n_duplicates, n_incomplete = n_incomplete,
    incomplete_subjects = incomplete_subjects,
    icc = icc, f_stat = f_stat, f_df1 = f_df1, f_df2 = f_df2, f_p = f_p,
    icc_results_df = icc_results_df,
    rater_means_df = rater_means_df,
    rater_hi = rater_means_df$rater_name[hi_idx],
    rater_lo = rater_means_df$rater_name[lo_idx],
    rater_spread = rater_spread,
    agreement_df = agreement_df, agreement_r = agreement_r,
    variance_df = variance_df,
    headline_band = headline_band,
    metrics = metrics, json_output = json_output
  )
}
Your data has more stories to tell.Run any analysis on your own data: R modules you own, interactive reports, AI insights, and PDF export. 500 free credits when you finish onboarding.
Try Free — No SignupSign Up Free

Your turn

Bring your own data and the question you actually need answered.

CympleData Scientist Send me your data and question, I’ll send you the analytics. ds@mcpanalytics.ai

Cite this analysis

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing