Free — no account required

Do Your Two Raters Really Agree? Beyond Chance

Upload a CSV, map the two judgement columns, and get the confusion matrix, percent agreement, Cohen's kappa with a confidence interval, weighted kappa for ordered categories, and a per-category breakdown of where they diverge. Free.

24,000+ analyses run
Encrypted & deleted in 7 days
PDF & citation included

Free analyses run on up to 10,000 rows. Larger files are randomly sampled to that size — sign up to analyze your full dataset.

📊
-
Rows
-
Columns
-
Numeric

Running rater agreement — cohen's kappa analysis...

Building the confusion matrix and computing kappa...

Your report is ready

Sent to — the confusion-matrix heatmap, percent and chance agreement, Cohen's kappa with its interval and benchmark band, weighted kappa where the categories are ordered, the per-category breakdown, R code, and AI insights.

Analyze another file
Sample Output

Every report includes interactive charts, tables, and AI insights

Upload your data to get your own report

View all case studies See all free tools

How it works

The analysis builds the k-by-k confusion matrix of the two raters' choices, then works entirely in closed form on it: raw percent agreement is the diagonal share; expected agreement is the sum over categories of the product of the two raters' marginal shares; Cohen's kappa is the observed excess over that expectation as a fraction of the amount available. The standard error is the Fleiss-Cohen-Everitt asymptotic form (which also covers the weighted case), and a separate null variance gives a z-test of kappa against chance. Whether the categories are ordered is detected from the labels themselves — numeric values, numeric prefixes, or recognised ordinal wording — and only then are linear- and quadratic-weighted kappa computed. A prevalence-and-bias-adjusted kappa is reported as a diagnostic of how much the category mix is depressing the headline value.

Use it whenever two people, systems, or passes assign a category to the same items and you need to know how much of their agreement is real: annotation QA, diagnostic coding, content moderation double-review, survey coding, or grading against a rubric.

Not for continuous or near-continuous ratings — scores, measurements, times — where an intraclass correlation is the right tool and kappa would throw away the scale. Not for three or more raters, which needs Fleiss' kappa. Not for comparing a rater against a known-correct answer, which is an accuracy question, not an agreement one.

Built for: Annotation and ML teams, clinical and research coders, QA and moderation leads, and anyone reporting inter-rater reliability for a categorical rubric

Typical data source: A spreadsheet with one row per item and two columns holding each rater's category choice

Machine LearningHealthcareResearchEducationTrust and SafetyMarket Research

What data do you need?

One row per item, one column per rater. For example, two reviewers grading the same submissions:

item_id (text) reviewer_a (categorical) reviewer_b (categorical)
IT0001 Good Good
IT0002 Fair Good
IT0003 Excellent Excellent

Minimum 20 rows · Best with 50-10,000 items and 2-8 categories

What's in the report?

Standard-library analysis: two raters, one categorical judgement per item — how much do they really agree? Map the two judgement columns and get the full confusion matrix as a heatmap, raw percent agreement, the agreement chance alone would produce, Cohen's kappa with its standard error and 95% confidence interval, a test of kappa against chance, linear- and quadratic-weighted kappa when the categories turn out to be ordered (detected from the labels, never assumed), a per-category breakdown of exactly where the raters diverge, and — when skewed marginals are depressing kappa — the kappa paradox explained with your own numbers rather than reported as a bad score.

🟧

Where the Two Raters Land

Every combination of the two raters' choices, counted — the diagonal is agreement and the brightest cell off it is the disagreement worth fixing first.

📋

Agreement Statistics

Percent agreement, chance agreement, kappa with its 95% interval, the weighted variants where the categories are ordered, and a prevalence-and-bias diagnostic.

📊

Which Categories They Fight Over

Per-category agreement — which parts of the rubric the two raters read the same way and which they do not.

📊

How Each Rater Uses the Scale

How often each rater reaches for each category; the two distributions that kappa's chance correction is built from, and the evidence behind any kappa paradox.

📋

Methods & Disclosure

The formulas in full, the orderedness rule that fired, and the two things kappa cannot tell you.

🤖

AI Insights

Plain-English interpretation — what the numbers mean, what's significant, and what to do next.

The Question This Answers

Do our two reviewers actually agree?

Map each reviewer's category column. You get the confusion matrix, percent agreement, Cohen's kappa with a confidence interval and a benchmark reading, and a per-category breakdown showing which parts of the rubric the two of them read differently.

Questions?

See our FAQ for details on pricing, data privacy, and how the analysis works. Every report includes a Methodology section showing the statistical test, assumptions checked, and diagnostics run.

Your data has more stories to tell

Run any analysis on your own data — validated R analyses, interactive reports, AI insights, and PDF export.

Try Free — No Credit Card
Powered by MCP Analytics