Free — no account required

Did Crossing The Threshold Cause The Change? Find Out Without An Experiment

Upload a CSV, map the outcome and the score your cutoff rule is applied to, and get the sharp regression discontinuity estimate with a confidence interval, the plot it comes from, the effective sample near the cutoff, and the three checks that decide whether it means anything. Free.

24,000+ analyses run
Encrypted & deleted in 7 days
PDF & citation included

Free analyses run on up to 10,000 rows. Larger files are randomly sampled to that size — sign up to analyze your full dataset.

📊
-
Rows
-
Columns
-
Numeric

Running regression discontinuity analysis...

Fitting local linear regressions either side of the cutoff...

Your report is ready

Sent to — The discontinuity estimate with a robust confidence interval, the RD plot with binned means and both fitted lines, manipulation and covariate-balance tests, bandwidth sensitivity, R code, and AI insights.

Analyze another file
Sample Output

Every report includes interactive charts, tables, and AI insights

Upload your data to get your own report

View all case studies See all free tools

How it works

The cutoff is taken from the 'cutoff' module parameter, else read off the treatment column, else assumed and flagged loudly. The discontinuity is estimated by local linear regression fitted separately either side of the cutoff, weighting each row by a triangular kernel that falls from 1 at the cutoff to 0 at the edge of the bandwidth; the gap between the two fitted lines at the cutoff is the estimate, with a heteroskedasticity-robust (HC1) standard error and a t-based 95% interval. The bandwidth comes from an Imbens-Kalyanaraman plug-in rule implemented directly — pilot density and residual variance at the cutoff, a global cubic for the third derivative, side-specific quadratics for the curvature difference, and the IK regularisation terms, with the triangular-kernel constant 3.4375 — then widened or narrowed so at least 15 rows sit each side. Three diagnostics follow: a McCrary-style density test (fine histogram with the cutoff on a bin edge, triangular-kernel local linear fits of the bin frequencies either side extrapolated to the cutoff, log-difference with its asymptotic standard error); covariate balance, running each mapped pre-determined column through the same estimator at the same bandwidth; and bandwidth sensitivity, recomputing the estimate from half to double the chosen bandwidth. If a treatment column is supplied and treatment is not a deterministic function of the cutoff, the design is reported as fuzzy and the sharp estimate is withheld.

Use it whenever a rule on a numeric score, measure, or index decides who gets something — an eligibility threshold, a discount tier, an audit trigger, a class-size cap — and you want the causal effect of crossing that line without running an experiment.

Not when treatment was assigned by anything other than a threshold on a measurable running variable (use the group-comparison, propensity-matching, or difference-in-differences tools), not when the change happened at a point in TIME across everyone at once (use interrupted time series), and not when compliance with the cutoff is partial — that is a fuzzy design needing a two-stage estimator, which this analysis detects and declines rather than approximating.

Built for: Economists, policy analysts, growth and pricing teams, and researchers evaluating threshold rules they did not randomise

Typical data source: Any spreadsheet with one row per person, account, or order: the score or measure a rule is applied to, and what happened afterwards

EducationPublic PolicyFinanceHealthcareSaaSRetailResearch

What data do you need?

One row per unit: the score the rule is applied to, whether the unit got the treatment, anything fixed beforehand, and the outcome:

student_id (text) exam_score (numeric) scholarship_awarded (text) prior_gpa (numeric) enrollment_days (numeric)
S0001 58.4 No 2.58 98.4
S0002 61.2 Yes 2.61 113.6
S0003 73.9 Yes 2.77 126.1

Minimum 30 rows · Best with 500-20,000 rows, with at least a few hundred within reach of the cutoff

What's in the report?

Standard-library analysis: did crossing the threshold cause the change? Map the outcome and the running variable a cutoff rule is applied to — an exam score, a revenue band, an eligibility index, a queue position — and get the sharp regression discontinuity estimate: local linear regression with a triangular kernel on each side of the cutoff at a data-driven bandwidth, the discontinuity with a robust 95% confidence interval, the scatter of outcome against the running variable with binned means and both fitted lines, the effective sample size actually near the cutoff, and the three diagnostics that decide whether the design is credible — a McCrary-style manipulation test, covariate balance at the cutoff, and bandwidth sensitivity.

🔵

Outcome Against the Running Variable

Binned means of the outcome across the running variable with both local linear fits and the cutoff marked — the gap where the lines meet the cutoff is the estimate.

📋

The Discontinuity Estimate

The cutoff, the bandwidth, the fitted level either side, the jump with its robust interval, and the effective sample the whole result rests on.

📋

Manipulation of the Running Variable

A density test for units sorting across the threshold — the failure that kills the design outright.

📋

Covariate Balance at the Cutoff

Pre-determined columns run through the same estimator; any of them jumping at the cutoff means the design is compromised.

🔵

Bandwidth Sensitivity

The estimate and interval recomputed from half to double the chosen bandwidth — a finding should survive the whole range.

📋

Method & Assumptions

The design, the estimator, the bandwidth rule, what each diagnostic found, and the limits — including that the estimate applies at the cutoff only.

🤖

AI Insights

Plain-English interpretation — what the numbers mean, what's significant, and what to do next.

The Question This Answers

Did the eligibility rule actually change anything?

Map the outcome and the score the rule is applied to. You get the jump at the cutoff with a confidence interval, the picture that jump comes from, and the three checks that decide whether it means anything — whether people gamed the score, whether other characteristics jump at the line too, and whether the number survives a range of bandwidths.

Questions?

See our FAQ for details on pricing, data privacy, and how the analysis works. Every report includes a Methodology section showing the statistical test, assumptions checked, and diagnostics run.

Your data has more stories to tell

Run any analysis on your own data — validated R analyses, interactive reports, AI insights, and PDF export.

Try Free — No Credit Card
Powered by MCP Analytics