Upload your ratings, map who rated what, and get the full intraclass correlation report — ICC(2,1), ICC(3,1), average-measure versions, Koo & Li bands, and exactly which ICC to report. Free.
Free analyses run on up to 10,000 rows. Larger files are randomly sampled to that size — sign up to analyze your full dataset.
Computing intraclass correlations...
Sent to — the full ICC table with Koo & Li bands, rater bias chart, subject-level agreement plot, variance breakdown, R code, and AI insights.
Analyze another fileThe analysis builds the subject-by-rater grid (averaging duplicate ratings, excluding subjects not rated by everyone), then computes the whole ICC family from ANOVA mean squares: one-way ICC(1,1); two-way ICC(2,1) for absolute agreement and ICC(3,1) for consistency; and the average-measure ICC(2,k) and ICC(3,k). An F-test checks that raters distinguish subjects at all, Koo & Li (2016) bands translate each value into poor/moderate/good/excellent, and a variance decomposition splits disagreement into systematic rater bias versus random noise.
Use it for any reliability study — multiple raters scoring the same subjects, one instrument measured repeatedly, or several devices measuring the same samples — whenever the rating is numeric.
Not for categorical ratings (use Cohen's or Fleiss' kappa), for a single rater with no repeats, or when different subjects are rated on different scales.
Built for: Researchers, clinicians, QA leads, and ML teams measuring rater or instrument reliability
Typical data source: A long-format spreadsheet of ratings: what was rated, who rated it, and the score
Long format — one row per rating. For example, four clinicians scoring the same 30 patient scans:
Minimum 10 rows · Best with 5-500 subjects and 2-10 raters (10-5,000 rows)
Standard-library analysis: how consistent are your raters, instruments, or repeated measurements? Upload long-format ratings (one row per rating) and get the full intraclass correlation family — ICC(1,1), ICC(2,1), ICC(3,1) and the average-measure versions — computed from ANOVA variance components, with Koo & Li interpretation bands, a systematic rater-bias check, a subject-level agreement plot, and a variance breakdown showing exactly where the disagreement comes from. Built for reliability studies: inter-rater agreement, test-retest, instrument comparison.
Every ICC form side by side — single vs average measure, agreement vs consistency — each with its Koo & Li band and a rule for when to report it.
Each rater's average score on the same subjects; any gap is pure systematic leniency or severity.
The first two raters plotted subject by subject against the perfect-agreement diagonal — offsets and outliers are visible at a glance.
How much of the variation is real subject differences versus rater bias versus noise — the anatomy of your reliability number.
Plain-English interpretation — what the numbers mean, what's significant, and what to do next.
Are my raters interchangeable?
Map what was rated, who rated it, and the score. You get the full ICC family with Koo & Li bands, a rater-bias chart that shows who scores systematically high or low, and a plain-language rule for which ICC to put in your paper.
See our FAQ for details on pricing, data privacy, and how the analysis works. Every report includes a Methodology section showing the statistical test, assumptions checked, and diagnostics run.
Run any analysis on your own data — validated R analyses, interactive reports, AI insights, and PDF export.
Try Free — No Credit CardTell us what went wrong, in your own words. We capture the page you're on automatically, so no need to describe where you are.