Executive Summary
Did crossing Enrollment Days = 94.59 change Exam Score?
The short answer
Crossing the Enrollment Days = 94.59 threshold shows no statistically reliable change in exam scores: the estimate is -1.508 (95% CI -3.217 to 0.201, p = 0.084). The result rests on 712 of 2,000 rows (35.60%) within the 9.01-day bandwidth. Diagnostics found no evidence of manipulation (p = 0.458), but covariate balance was not tested and bandwidth sensitivity is a concern.
The detail
Jump at cutoff: -1.508. Confidence interval includes zero. Effective sample: 712 rows (35.60%). Bandwidth: 9.0102. Manipulation test (density ratio 1.1448, p = 0.458): no detectable sorting. Covariate balance: not tested (no covariates mapped). Bandwidth sensitivity: estimates range from -0.754 to -2.798 across the tested range; 3 of 6 confidence intervals exclude zero.
What this can't tell you
The estimate is not statistically significant and is sensitive to bandwidth choice. No pre-determined covariates were checked for balance, leaving open the possibility that something other than the threshold rule changes at the cutoff. The design estimates an effect only at Enrollment Days = 94.59 and makes no claim about exam scores for units far above or below the threshold.
Analysis Overview
Regression discontinuity in Exam Score at Enrollment Days = 94.59.
The short answer
Crossing the 94.59 enrollment threshold is associated with a decline of about 1.5 points in exam scores, but this pattern is not statistically distinguishable from zero at the threshold itself. The regression discontinuity design compares units just barely above and below the cutoff—where they are otherwise alike—to isolate the threshold rule's effect. This comparison rests on only 712 of 2,000 rows (35.60%), the ones within 9.01 enrollment days of the line.
The detail
The estimate is -1.508 (95% CI -3.217 to 0.201, p = 0.084). The fitted exam score just below the cutoff is 50.950; just above it is 49.442. The bandwidth of 9.0102 was set by the Imbens-Kalyanaraman rule, which balances bias and variance at the cutoff. This design estimates an effect only at Enrollment Days = 94.59 and makes no claim about units far above or below it.
What this can't tell you
The confidence interval includes zero, so no reliable step appears at the threshold. Bandwidth sensitivity is a concern: across the tested range (4.51 to 18.02), estimates vary from -0.754 to -2.798, and only 3 of 6 confidence intervals exclude zero. No pre-determined covariates were mapped, leaving a gap in the evidence that something else might jump at the threshold.
Data Preparation
How the cutoff was set, what was dropped, and how the rows split.
The short answer
All 2,000 rows were usable; no rows were dropped for missing data. The cutoff of 94.59 was assumed from the median of Enrollment Days because no threshold and no treatment column were supplied. This splits the sample into 1,000 rows below and 1,000 at or above the line. Missing outcome values were never imputed near the cutoff, preserving the integrity of the jump being measured.
The detail
Initial rows: 2,000; final rows: 2,000; rows removed: 0. The analysis assumes the median enrollment day (94.59) is the true threshold. This is a placeholder: the real cutoff must be supplied, because every number below changes with it. The split is 1,000 below the cutoff and 1,000 at or above it.
What this can't tell you
The assumed cutoff is not verified against any administrative rule or policy. If the true threshold differs from 94.59, all estimates and diagnostics are invalid. No imputation near the cutoff is correct practice for RD, but it means any missing exam scores near the line are excluded from the analysis entirely.
Outcome Against the Running Variable
Binned means with the local linear fit either side of the cutoff.
The short answer
The scatter of binned means shows a strong upward trend in exam scores across enrollment days, with a fitted line on each side of the 94.59 cutoff. The two lines meet at a gap of -1.508, but the pattern is consistent with gentle curvature that straight lines cannot fully capture, which is why bandwidth sensitivity matters here.
The detail
The 32 binned-mean points trace exam scores from about 3.6 at enrollment day 76.18 to 98.4 at day 107.5, with a clear positive association. The left-hand local linear fit (below 94.59) extrapolates to 50.950; the right-hand fit (above 94.59) extrapolates to 49.442. Both lines are drawn only within the 9.01-day bandwidth on each side of the cutoff. The vertical reference line marks the cutoff at 94.59. The gap between the lines at that point is the discontinuity estimate.
What this can't tell you
The visual pattern does not rule out curvature near the cutoff that the straight-line fit cannot follow. The plot suggests a smooth upward trend rather than a sharp step, which aligns with the non-significant result and the bandwidth sensitivity finding that the estimate changes substantially with the window width.
The Discontinuity Estimate
The cutoff, the bandwidth, the effective sample, and the jump with its interval.
| Measure | Estimate | Detail |
|---|---|---|
| Cutoff on Enrollment Days | 94.59 | No cutoff and no treatment column were supplied, so the median of Enrollment Days (94.59) was ASSUMED to be the threshold. This is a placeholder, not a discovered rule: supply the real cutoff, because every number below changes with it. |
| Bandwidth | 9.01 | The bandwidth 9.01 came from the Imbens-Kalyanaraman plug-in rule, which trades the bias of a wide window against the noise of a narrow one using the estimated density, residual variance, and curvature at the cutoff. |
| Rows inside the bandwidth | 712 | 712 of the 2,000 usable rows (35.60%) — 514 below the cutoff and 198 at or above it. This is the sample the estimate actually rests on. |
| Fitted Exam Score just below the cutoff | 50.95 | The left-hand local linear line extrapolated to Enrollment Days = 94.59 |
| Fitted Exam Score just above the cutoff | 49.44 | The right-hand local linear line extrapolated to Enrollment Days = 94.59 |
| Discontinuity estimate at the cutoff | -1.508 | 95% CI -3.217 to 0.201; p = 0.084; not statistically significant (p = 0.084) |
| Standard error | 0.8705 | Heteroskedasticity-robust (HC1) around the weighted local linear fit |
| Effect relative to the level just below | -2.96 | The jump as a percentage of the fitted Exam Score just below the cutoff (50.950) |
The short answer
The discontinuity at Enrollment Days = 94.59 is -1.508 exam score points (95% CI -3.217 to 0.201), not statistically significant. This estimate comes from 712 rows (35.60% of the full sample) within 9.01 enrollment days of the cutoff. The fitted exam score just below the line is 50.950; just above it is 49.442.
The detail
Cutoff: 94.59 (assumed from the median; no cutoff or treatment column supplied). Bandwidth: 9.0102 (Imbens-Kalyanaraman plug-in rule). Effective sample: 712 rows (514 below, 198 at or above). Left-side fitted level: 50.950. Right-side fitted level: 49.442. Discontinuity: -1.5077 (p = 0.084). Standard error: 0.8705 (heteroskedasticity-robust, HC1). Effect relative to the level just below: -2.96%.
What this can't tell you
The confidence interval includes zero, so no reliable step is detected. The estimate rests on a minority of the full sample (35.60%), meaning most rows contribute no weight to this result. Whatever this estimates, it applies only at the threshold itself and says nothing about exam scores for units far above or below Enrollment Days = 94.59.
Manipulation of the Running Variable
Density test for sorting across the cutoff in Enrollment Days.
| Measure | Estimate | Detail |
|---|---|---|
| Density of Enrollment Days just below the cutoff | 0.0214 | Local linear fit of the binned frequencies of Enrollment Days below the cutoff, extrapolated to it |
| Density of Enrollment Days just above the cutoff | 0.0244 | Local linear fit of the binned frequencies of Enrollment Days at or above the cutoff, extrapolated to it |
| Density ratio (above divided by below) | 1.145 | 1.00 means no pile-up on either side |
| Log density difference | 0.1352 | The quantity actually tested — zero under no manipulation |
| z statistic | 0.742 | Log difference divided by its standard error (0.1822) |
| p-value | 0.458 | 0.458 — no detectable pile-up on either side |
| Histogram bin width | 0.6488 | 2 times the spread of Enrollment Days divided by the square root of the row count (McCrary's rule) |
| Bins used each side | 20 | 10 below and 10 above, inside a density bandwidth of 6.35 |
The short answer
No evidence of manipulation: students did not systematically sort themselves across the enrollment threshold to gain or avoid treatment. The density of enrollment days is nearly identical on both sides of the cutoff.
The detail
The density test fits the distribution of enrollment days on each side of the cutoff and extrapolates both to 94.59. Density just below the cutoff was 0.0214; just above, 0.0244 (ratio 1.1448). The log density difference is 0.1352 with a z-statistic of 0.742 and p-value of 0.458, with a standard error of 0.1822. The test used a histogram bin width of 0.6488 and 20 bins (10 below, 10 above) within a density bandwidth of 6.35.
What this can't tell you
The manipulation test has limited power in small samples and absence of evidence is not proof of absence of sorting. The test does not rule out that units manipulated enrollment days in ways the density test cannot detect (e.g., clustering at round numbers rather than piling up exactly at the cutoff).
Covariate Balance at the Cutoff
Do pre-determined columns jump at the threshold as well?
| Covariate | Jump | Std Error | P Value | Rows Used | Verdict |
|---|---|---|---|---|---|
| (no covariates mapped) | — | — | n/a | 0 | not tested — map pre-determined columns to run this check |
The short answer
No pre-determined columns were mapped, so covariate balance could not be tested. This is a gap in the evidence, not a pass. Without checking whether age, prior scores, tenure, or other fixed characteristics jump at the threshold, we cannot rule out that something other than the enrollment rule changes at the cutoff.
The detail
Covariate balance test: not conducted. No pre-determined columns were supplied. The analysis cannot assess whether any baseline characteristics differ across the threshold, which is the main way a spurious discontinuity reveals itself.
What this can't tell you
Map columns that were fixed before the threshold was applied (age, prior score, tenure, region, prior performance) and they should show no jump at the cutoff. Their absence here leaves the design vulnerable to the claim that some unmeasured confounder, not the enrollment rule itself, drives the exam score difference. This is not a minor gap—covariate balance is a primary check on the validity of the causal interpretation.
Bandwidth Sensitivity
The estimate and its interval recomputed across a range of bandwidths.
The short answer
The estimated effect is sensitive to bandwidth choice. As the bandwidth widens from half to double the chosen value, the estimate strengthens from -0.7537 to -2.798 points, and the pattern of statistical significance shifts across that range.
The detail
The headline estimate of -1.508 at bandwidth 9.01 is one point in a wider picture. At 0.50× bandwidth (4.51), the estimate is -0.7537 (95% CI: -3.1461 to 1.6386). At 1.25× (11.26), it is -1.916 (CI: -3.4793 to -0.3528). At 2.00× (18.02), it is -2.798 (CI: -4.1728 to -1.4233). Three of the six confidence intervals exclude zero; the others include it. The spread of estimates is 135.59% of the headline figure.
What this can't tell you
Bandwidth sensitivity this wide signals that the result is borderline and dependent on the window chosen. A narrower bandwidth reduces bias but increases noise; a wider one does the reverse. The plug-in rule balances these, but the practical implication is that the true effect — if one exists — is not pinned down by this data at this sample size.
Method & Assumptions
The design, the identification choices, the diagnostics, and the limits.
| Item | Detail |
|---|---|
| Design | Sharp is ASSUMED: no treatment column was mapped, so the analysis takes it that everything at or above 94.59 on Enrollment Days was treated and nothing below it was. If compliance was imperfect, this assumption is wrong and the number below is not a sharp RD estimate — map the treatment column to have it checked. |
| Cutoff | No cutoff and no treatment column were supplied, so the median of Enrollment Days (94.59) was ASSUMED to be the threshold. This is a placeholder, not a discovered rule: supply the real cutoff, because every number below changes with it. |
| Estimator | Local linear regression fitted separately either side of the cutoff, weighting each row by a triangular kernel (weight 1 at the cutoff falling to 0 at the edge of the bandwidth). The estimate is the gap between the two lines where they meet Enrollment Days = 94.59. |
| Bandwidth | The bandwidth 9.01 came from the Imbens-Kalyanaraman plug-in rule, which trades the bias of a wide window against the noise of a narrow one using the estimated density, residual variance, and curvature at the cutoff. |
| Inference | Heteroskedasticity-robust (HC1) standard errors on the weighted fit; the 95% interval uses the t distribution on 708 degrees of freedom. The interval is NOT bias-corrected: a bandwidth chosen to minimise mean squared error deliberately accepts some bias in exchange for precision, so the true coverage of this interval is below 95% whenever Exam Score is curved near the cutoff. The bandwidth sensitivity row is the practical check — a narrower bandwidth carries less bias and a wider interval. |
| Effective sample | 712 of 2,000 rows sit inside the bandwidth (35.60%). A regression discontinuity on a large table can still rest on a small handful of rows near the threshold, which is why this number is reported rather than the row count. |
| Manipulation test | The density of Enrollment Days does not jump detectably at the cutoff (0.0244 just above versus 0.0214 just below, ratio 1.14, p = 0.458), so there is no sign that units sorted themselves across the threshold. Absence of evidence is not proof: the test has limited power in small samples. |
| Covariate balance | No pre-determined columns were mapped, so covariate balance could not be tested. That is a gap in the evidence, not a pass: map columns that were fixed before the threshold was applied (age, prior score, tenure, region) and they should show no jump at the cutoff. |
| Bandwidth sensitivity | The estimate moves with the bandwidth: 3 of the 6 confidence intervals exclude zero and the estimates range from -2.798 to -0.754, a spread of 135.59% of the headline figure. Treat the headline number as one point in that range rather than a settled quantity. |
| What this cannot tell you | Whatever this estimates, it estimates it only at Enrollment Days = 94.59. It describes units sitting right at that threshold and says nothing about Exam Score for units far above or below it — a regression discontinuity buys credibility at the cutoff by giving up everything else. |
| Key assumptions | Units just below and just above Enrollment Days = 94.59 are otherwise comparable; nothing else changes at the same threshold; Exam Score and any other pre-determined characteristic evolve smoothly through it; and the relationship between Enrollment Days and Exam Score is well approximated by a straight line on each side within the bandwidth. |
The short answer
The analysis uses local linear regression on each side of the cutoff, weighted by distance, with a bandwidth of 9.01 days chosen by the Imbens-Kalyanaraman plug-in rule. The design assumes a sharp threshold at the enrollment median (94.59) and no sorting across it.
The detail
No cutoff or treatment column were supplied, so the analysis assumes the median enrollment day (94.59) is the threshold and that all students at or above it were treated. The estimator fits separate lines on each side, weighting rows by a triangular kernel, and reads the gap at the cutoff. The bandwidth 9.01 minimises mean squared error; inference uses heteroskedasticity-robust (HC1) standard errors on 708 degrees of freedom, with a t-distribution 95% interval that is not bias-corrected. The effective sample is 712 rows (35.60% of 2,000). Manipulation was tested via density fit (p = 0.458); covariate balance could not be tested because no pre-threshold covariates were mapped. The interval's true coverage is below 95% if exam score is curved near the cutoff — the bandwidth sensitivity check is the practical guard against this.
What this can't tell you
The design estimates an effect only at enrollment days = 94.59. If the supplied treatment column shows compliance was imperfect, this is a fuzzy design, not sharp, and the number is not a sharp RD estimate. If pre-treatment covariates exist, they should be mapped to check balance; their absence is a gap in evidence, not a pass.
Methodology
Statistical methodology and diagnostics for Regression Discontinuity
Statistical Method
Standard-library analysis: did crossing the threshold cause the change? Map the outcome and the running variable a cutoff rule is applied to — an exam score, a revenue band, an eligibility index, a queue position — and get the sharp regression discontinuity estimate: local linear regression with a triangular kernel on each side of the cutoff at a data-driven bandwidth, the discontinuity with a robust 95% confidence interval, the scatter of outcome against the running variable with binned means and both fitted lines, the effective sample size actually near the cutoff, and the three diagnostics that decide whether the design is credible — a McCrary-style manipulation test, covariate balance at the cutoff, and bandwidth sensitivity.
- A threshold rule on a numeric running variable decides who is treated, and units just either side of it are otherwise comparable
- Nothing else changes at the same threshold — no other rule, programme, or reporting boundary shares the cutoff
- The running variable is not manipulable by the units themselves (the manipulation test looks for violations)
- Pre-determined characteristics evolve smoothly through the cutoff (the balance check looks for violations)
- Within the bandwidth, the relationship between the outcome and the running variable is well approximated by a straight line on each side
- The estimate is LOCAL: it applies at the cutoff only and says nothing about units far above or below it
- The confidence interval is not bias-corrected — a bandwidth chosen to minimise mean squared error deliberately accepts bias, so coverage falls below 95% wherever the outcome is curved near the cutoff
- Only the rows inside the bandwidth contribute; a large table can still yield an estimate resting on a few dozen observations
- A fuzzy design (imperfect compliance with the cutoff) needs a different, two-stage estimator — this analysis detects that case and declines to report a sharp estimate rather than reporting one anyway
Analysis Code
Complete R source code for this analysis
Regression Discontinuity — Did Crossing the Threshold Cause the Change?
Sharp regression discontinuity. Units on either side of a cutoff in a running variable receive different treatment, and the jump in the outcome AT the cutoff estimates the causal effect for units sitting right at the threshold. The estimate is a local linear regression with a triangular kernel fitted separately on each side, at a data-driven bandwidth, with a heteroskedasticity-robust confidence interval.
Why This Method?
When a rule ("scholarship if score >= 60", "audit if revenue >= 5m") decides who is treated, units just below and just above the line are otherwise alike — the assignment near the threshold is as good as random. Comparing them recovers a causal effect from observational data, with no experiment.
What This Analysis Covers
- The discontinuity estimate with a 95% confidence interval
- Outcome against the running variable: binned means plus both fitted lines
- Manipulation of the running variable (McCrary-style density test)
- Covariate balance at the cutoff
- Bandwidth sensitivity across a range of bandwidths
- The effective sample size that actually carries the estimate
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {outcome, running, treatment, covariate_1..N}. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Estimation Machinery
Everything below is implemented directly on base R / stats: the local linear RD estimator with a triangular kernel and HC1-robust standard errors, the Imbens-Kalyanaraman plug-in bandwidth, and the McCrary density test. No RD-specific package is used.
A non-positive extrapolated density means the side has been emptied out right at the cutoff — the strongest possible sorting signal, not a reason to decline the test. Floor it at half of one observation's worth of density and say so, so the log ratio and its standard error stay defined.
f_floor <- 0.5 / (n * b)
floored <- lo$f <= f_floor || hi$f <= f_floor
lo$f <- max(lo$f, f_floor)
hi$f <- max(hi$f, f_floor)
theta <- log(hi$f) - log(lo$f)
se <- sqrt((1 / (n * hd)) * (24 / 5) * (1 / hi$f + 1 / lo$f))
if (!is.finite(se) || se <= 0) {
return(list(testable = FALSE, bin_width = b, bandwidth = hd,
n_bins_below = lo$nbin, n_bins_above = hi$nbin))
}
z <- theta / se
list(testable = TRUE, f_below = lo$f, f_above = hi$f, floored = floored,
ratio = hi$f / lo$f, theta = theta, se = se, z = z,
p = 2 * stats::pnorm(-abs(z)), bin_width = b, bandwidth = hd,
n_bins_below = lo$nbin, n_bins_above = hi$nbin)
}Core Analysis Pipeline
compute_shared <- function(df, params, col_map = list()) {
# === SHARED EXPORTS ===
# initial_rows/final_rows/rows_removed $ row accounting
# outcome_h/running_h/treatment_h $ humanized user column names
# cutoff / cutoff_rule $ the threshold and how it was set
# n_below/n_above/n_at_cutoff $ counts either side of the cutoff
# h / bw_rule $ bandwidth and the rule that chose it
# fit $ local_linear_rd() at h
# tau/tau_se/tau_p/ci_low/ci_high $ the discontinuity estimate
# n_eff/n_eff_left/n_eff_right/eff_pct $ effective sample near the cutoff
# design_kind/design_label/first_stage $ sharp vs fuzzy detection
# dens/dens_verdict/dens_short $ manipulation (McCrary) test
# bal_df/bal_short/bal_verdict $ covariate balance
# bw_df/bw_short/bw_verdict $ bandwidth sensitivity
# locality $ the local-effect honesty sentence
# rdd_fit_df/estimate_df/manip_df/methods_df $ card datasets
# metrics / json_output
# === /SHARED EXPORTS ===Step 1: Resolve the mapped columns (humanized for every sentence)
initial_rows <- nrow(df)
outcome_h <- humanize_semantic("outcome", col_map)
running_h <- humanize_semantic("running", col_map)
treatment_h <- humanize_semantic("treatment", col_map)
for (req in c("outcome", "running")) {
if (!(req %in% names(df))) {
stop(sprintf("Regression discontinuity needs '%s' (%s) mapped.",
humanize_semantic(req, col_map),
c(outcome = "the numeric outcome",
running = "the running variable that the cutoff applies to")[[req]]))
}
}Step 2: Coerce outcome and running variable to numeric (95% rule)
y_all <- coerce_num_95(df$outcome)
if (is.null(y_all)) {
stop(sprintf("The outcome column '%s' does not look numeric — fewer than 95%% of its values parse as numbers. Pick a numeric column.",
outcome_h))
}
x_all <- coerce_num_95(df$running)
if (is.null(x_all)) {
stop(sprintf("The running variable '%s' does not look numeric — fewer than 95%% of its values parse as numbers. A regression discontinuity needs a numeric score, date-number, or measure that the cutoff is applied to.",
running_h))
}Step 3: Keep rows complete on the outcome and the running variable
keep <- is.finite(y_all) & is.finite(x_all)
n_dropped <- sum(!keep)
y <- y_all[keep]; x <- x_all[keep]
final_rows <- length(y)
rows_removed <- initial_rows - final_rows
if (final_rows < 30) {
stop(sprintf("Only %d rows have both %s and %s — regression discontinuity needs at least 30, because the estimate is built from the rows near the cutoff alone.",
final_rows, outcome_h, running_h))
}
if (length(unique(x)) < 10) {
stop(sprintf("The running variable '%s' takes only %d distinct values — regression discontinuity needs a running variable that varies continuously around the cutoff.",
running_h, length(unique(x))))
}
if (!is.finite(stats::sd(x)) || stats::sd(x) == 0) {
stop(sprintf("The running variable '%s' is constant — there is no threshold to compare across.", running_h))
}Step 4: The treated flag, when supplied — and the cutoff
treat <- NULL
treat_note <- ""
if ("treatment" %in% names(df)) {
tv <- df$treatment[keep]
tn <- coerce_num_95(tv)
if (!is.null(tn) && length(unique(stats::na.omit(tn))) == 2) {
lv <- sort(unique(stats::na.omit(tn)))
treat <- as.numeric(tn == lv[2])
} else {
ch <- trimws(as.character(tv))
lv <- sort(unique(ch[!is.na(ch) & ch != ""]))
if (length(lv) == 2) {
pos <- c("yes", "y", "true", "t", "1", "treated", "treatment", "enrolled",
"eligible", "awarded", "in", "on", "approved")
hit <- tolower(lv) %in% pos
pos_lvl <- if (sum(hit) == 1) lv[hit] else lv[2]
treat <- as.numeric(ch == pos_lvl)
treat_note <- sprintf("'%s' was read as the treated value of %s. ", pos_lvl, treatment_h)
} else {
treat_note <- sprintf("The treatment column %s has %d distinct values rather than two, so it could not be read as a treated/untreated flag and was ignored. ",
treatment_h, length(lv))
}
}
if (!is.null(treat) && any(!is.finite(treat))) treat[!is.finite(treat)] <- 0
}
cutoff_param <- suppressWarnings(as.numeric(params$cutoff %||% params$threshold %||% NA))
if (length(cutoff_param) != 1) cutoff_param <- NA_real_
if (is.finite(cutoff_param)) {
cutoff <- cutoff_param
cutoff_rule <- sprintf("The cutoff was supplied as %s on %s.", rs(cutoff), running_h)
} else if (!is.null(treat)) {Sharp separation gives the cutoff exactly: the smallest treated value.
if (max(c(-Inf, x[treat == 0])) < min(c(Inf, x[treat == 1]))) {
cutoff <- min(x[treat == 1])
cutoff_rule <- sprintf("The cutoff was read off %s: every treated unit has %s of at least %s and every untreated unit is below it.",
treatment_h, running_h, rs(cutoff))
} else {
v <- sort(unique(x))
cnt_all <- as.numeric(table(factor(x, levels = v)))
cnt_trt <- as.numeric(tapply(treat, factor(x, levels = v), sum))
cnt_trt[!is.finite(cnt_trt)] <- 0
cnt_unt <- cnt_all - cnt_trt
before_trt <- c(0, cumsum(cnt_trt)[-length(v)])
before_unt <- c(0, cumsum(cnt_unt)[-length(v)])
acc <- (sum(cnt_trt) - before_trt + before_unt) / length(x)
acc[!is.finite(acc)] <- -Inf
cutoff <- v[which.max(acc)]
cutoff_rule <- sprintf("No cutoff was supplied and %s does not separate cleanly, so the threshold on %s that best reproduces it(%s, matching %s of rows) was used.",
treatment_h, running_h, rs(cutoff),
paste0(r2(100 * max(acc)), "%"))
}
} else {
zero_ok <- min(x) < 0 && max(x) > 0 &&
sum(x < 0) >= 10 && sum(x >= 0) >= 10
if (zero_ok) {
cutoff <- 0
cutoff_rule <- sprintf("No cutoff and no treatment column were supplied, so zero was ASSUMED to be the threshold on %s because the values straddle it. If the real rule sits elsewhere, supply the cutoff — every number below changes with it.",
running_h)
} else {
cutoff <- stats::median(x)
cutoff_rule <- sprintf("No cutoff and no treatment column were supplied, so the median of %s(%s) was ASSUMED to be the threshold. This is a placeholder, not a discovered rule: supply the real cutoff, because every number below changes with it.",
running_h, rs(cutoff))
}
}
if (!is.finite(cutoff)) {
stop(sprintf("The cutoff on %s could not be determined. Supply it as the 'cutoff' module parameter.", running_h))
}
xc <- x - cutoff
n_below <- sum(xc < 0); n_above <- sum(xc >= 0)
n_at_cutoff <- sum(xc == 0)
if (n_below < 10 || n_above < 10) {
stop(sprintf("Only %d rows fall below the cutoff(%s on %s) and %d at or above it — at least 10 are needed on each side for a regression discontinuity.",
n_below, rs(cutoff), running_h, n_above))
}Step 5: Bandwidth — Imbens-Kalyanaraman plug-in, then guard rails
side_h <- function(v, k) {
if (length(v) == 0) return(Inf)
v <- sort(v)
if (length(v) < k) return(v[length(v)])
v[k]
}
h_min <- max(side_h(abs(xc[xc < 0]), 15), side_h(xc[xc >= 0], 15)) * 1.0001
h_max <- max(abs(xc))
h_param <- suppressWarnings(as.numeric(params$bandwidth %||% NA))
if (length(h_param) != 1) h_param <- NA_real_
h_ik <- ik_bandwidth(xc, y)
if (is.finite(h_param) && h_param > 0) {
h_raw <- h_param
bw_source <- sprintf("The bandwidth %s was supplied as a module parameter.", rs(h_param))
} else if (is.finite(h_ik)) {
h_raw <- h_ik
bw_source <- sprintf("The bandwidth %s came from the Imbens-Kalyanaraman plug-in rule, which trades the bias of a wide window against the noise of a narrow one using the estimated density, residual variance, and curvature at the cutoff.",
rs(h_ik))
} else {
h_raw <- 1.84 * stats::sd(xc) * final_rows^(-1 / 5)
bw_source <- sprintf("The plug-in bandwidth rule could not be evaluated on this data, so the pilot rule(1.84 times the spread of %s divided by the fifth root of the row count) was used instead, giving %s.",
running_h, rs(h_raw))
}
h <- min(max(h_raw, h_min), h_max)
bw_clamped <- abs(h - h_raw) > 1e-9
bw_rule <- paste0(
bw_source,
if (bw_clamped)
sprintf(" It was then widened/narrowed to %s so that at least 15 rows sit on each side of the cutoff inside the window.", rs(h))
else "")
fit <- local_linear_rd(xc, y, h)
if (is.null(fit)) {
stop(sprintf("A local linear fit around the cutoff(%s on %s) could not be estimated — too few rows, or no variation in %s, on one side of the threshold.",
rs(cutoff), running_h, running_h))
}
tau <- fit$est; tau_se <- fit$se; tau_p <- fit$p
ci_low <- fit$ci_low; ci_high <- fit$ci_high
n_eff <- fit$n_eff; n_eff_left <- fit$n_left; n_eff_right <- fit$n_right
eff_pct <- 100 * n_eff / final_rows
rel_pct <- if (abs(fit$level_below) > 1e-8) 100 * tau / abs(fit$level_below) else NA_real_Step 6: Sharp or fuzzy? The sharp estimator needs deterministic assignment
The first stage is the jump in the probability of treatment at the cutoff, estimated with the same local linear machinery at the same bandwidth.
implied <- as.numeric(xc >= 0)
design_kind <- "sharp_assumed"
n_crossovers <- NA_integer_
first_stage <- NA_real_
if (!is.null(treat)) {
n_crossovers <- sum(treat != implied)
if (n_crossovers == 0) {
design_kind <- "sharp"
first_stage <- 1
} else {
fs <- local_linear_rd(xc, treat, h)
first_stage <- if (is.null(fs)) NA_real_ else fs$est
design_kind <- if (is.finite(first_stage) && abs(first_stage) >= 0.2)
"fuzzy" else "no_first_stage"
}
}
design_label <- switch(
design_kind,
sharp = "sharp",
sharp_assumed = "sharp(assumed — no treatment column supplied)",
fuzzy = "fuzzy — sharp estimate does not apply",
no_first_stage = "no first stage — the cutoff does not change treatment"
)
is_sharp <- design_kind %in% c("sharp", "sharp_assumed")Step 7: Diagnostic 1 — manipulation of the running variable
dens <- mccrary_density(xc)
if (is.null(dens) || !isTRUE(dens$testable)) {
dens_short <- "not testable"
dens_verdict <- sprintf("The density of %s around the cutoff could not be estimated(too few distinct values or too few bins on one side), so manipulation could not be tested — treat that as a gap in the evidence, not as a clean bill of health.",
running_h)
} else if (dens$p < 0.05) {
dens_short <- "detected"
dens_verdict <- sprintf("The density of %s jumps at the cutoff: %s just above versus %s just below, a ratio of %s(%s).%s Units appear able to place themselves on the favourable side of the threshold, and when that is true the units either side are no longer comparable — the design is not credible and the estimate below should not be read as causal.",
running_h, rs(dens$f_above), rs(dens$f_below),
r2(dens$ratio), fmt_pp(dens$p),
if (isTRUE(dens$floored))
" One side extrapolates to a density at or below zero at the cutoff — effectively a hole in the data right where the rule bites — so it was floored at half of one observation's worth of density to keep the ratio finite; the true gap is wider than the number shown."
else "")
} else {
dens_short <- "no evidence"
dens_verdict <- sprintf("The density of %s does not jump detectably at the cutoff(%s just above versus %s just below, ratio %s, %s), so there is no sign that units sorted themselves across the threshold. Absence of evidence is not proof: the test has limited power in small samples.",
running_h, rs(dens$f_above), rs(dens$f_below),
r2(dens$ratio), fmt_pp(dens$p))
}Step 8: Diagnostic 2 — covariate balance at the cutoff
cov_cols <- grep("^covariate_[0-9]+$", names(df), value = TRUE)
cov_cols <- cov_cols[order(as.integer(sub("^covariate_", "", cov_cols)))]
bal_rows <- list(); dropped_cov <- character(0)
for (cc in cov_cols) {
lab <- humanize_semantic(cc, col_map)
v <- df[[cc]][keep]
cn <- coerce_num_95(v)
if (is.null(cn)) {
ch <- trimws(as.character(v))
lv <- sort(unique(ch[!is.na(ch) & ch != ""]))
if (length(lv) == 2) {
cn <- as.numeric(ch == lv[2])
lab <- sprintf("%s is %s", lab, lv[2])
} else {
dropped_cov <- c(dropped_cov, sprintf("%s(%d distinct text values)", lab, length(lv)))
next
}
}
if (sum(is.finite(cn)) < 20) {
dropped_cov <- c(dropped_cov, sprintf("%s(fewer than 20 usable values)", lab)); next
}
vv <- suppressWarnings(stats::var(cn, na.rm = TRUE))
if (!is.finite(vv) || vv == 0) {
dropped_cov <- c(dropped_cov, sprintf("%s(constant)", lab)); next
}
cfit <- local_linear_rd(xc, cn, h)
if (is.null(cfit)) {
dropped_cov <- c(dropped_cov, sprintf("%s(no usable values near the cutoff)", lab)); next
}
bal_rows[[length(bal_rows) + 1]] <- data.frame(
covariate = lab,
jump = round(cfit$est, 4),
std_error = round(cfit$se, 4),
p_value = fmt_p(cfit$p),
rows_used = cfit$n_eff,
verdict = if (cfit$p < 0.05) "jumps at the cutoff" else "no detectable jump",
stringsAsFactors = FALSE
)
}
if (length(bal_rows) > 0) {
bal_df <- do.call(rbind, bal_rows)
rownames(bal_df) <- NULL
n_bad <- sum(bal_df$verdict == "jumps at the cutoff")
if (n_bad > 0) {
bal_short <- "imbalanced"
bal_verdict <- sprintf("%d of the %d pre-determined column(s) tested(%s) jump at the cutoff themselves. Things that were fixed before the threshold was applied should not change at it — when they do, the units either side differ in more than the treatment and the design is compromised. Read the estimate as a description of the two groups, not as the effect of crossing the threshold.",
n_bad, nrow(bal_df),
paste(bal_df$covariate[bal_df$verdict == "jumps at the cutoff"], collapse = ", "))
} else {
bal_short <- "balanced"
bal_verdict <- sprintf("None of the %d pre-determined column(s) tested(%s) jump detectably at the cutoff, which is what a credible discontinuity looks like: only the treatment changes at the threshold. This supports the design without proving it — an unmeasured characteristic could still jump.",
nrow(bal_df), paste(bal_df$covariate, collapse = ", "))
}
} else {
bal_df <- data.frame(
covariate = "(no covariates mapped)", jump = NA_real_, std_error = NA_real_,
p_value = "n/a", rows_used = 0L,
verdict = "not tested — map pre-determined columns to run this check",
stringsAsFactors = FALSE)
bal_short <- "not tested"
bal_verdict <- "No pre-determined columns were mapped, so covariate balance could not be tested. That is a gap in the evidence, not a pass: map columns that were fixed before the threshold was applied(age, prior score, tenure, region) and they should show no jump at the cutoff."
}Step 9: Diagnostic 3 — bandwidth sensitivity
mults <- c(0.50, 0.75, 1.00, 1.25, 1.50, 2.00)
bw_list <- list(); used_h <- numeric(0)
for (mm in mults) {
hh <- min(h * mm, h_max)Multipliers that clamp to the same window would repeat a row verbatim.
if (any(abs(used_h - hh) < 1e-9)) next
used_h <- c(used_h, hh)
f2 <- local_linear_rd(xc, y, hh)
if (is.null(f2)) next
bw_list[[length(bw_list) + 1]] <- data.frame(
bandwidth_label = sprintf("%s× h(%s)", r2(mm), rs(hh)),
estimate = round(f2$est, 4),
ci_low = round(f2$ci_low, 4),
ci_high = round(f2$ci_high, 4),
stringsAsFactors = FALSE
)
}
bw_df <- if (length(bw_list) > 0) do.call(rbind, bw_list) else
data.frame(bandwidth_label = character(0), estimate = numeric(0),
ci_low = numeric(0), ci_high = numeric(0), stringsAsFactors = FALSE)
rownames(bw_df) <- NULL
n_bw <- nrow(bw_df)
n_bw_sig <- if (n_bw > 0) sum(bw_df$ci_low > 0 | bw_df$ci_high < 0) else 0L
sign_consistent <- n_bw > 0 && all(sign(bw_df$estimate) == sign(tau))Spread across the bandwidth grid relative to the headline figure — a sharper stability statistic than the distance of any single bandwidth.
spread_pct <- if (n_bw > 0 && abs(tau) > 1e-12)
100 * (max(bw_df$estimate) - min(bw_df$estimate)) / abs(tau) else NA_real_
main_sig <- is.finite(tau_p) && tau_p < 0.05
if (n_bw < 2) {
bw_short <- "not testable"
bw_verdict <- "Fewer than two bandwidths could be estimated on this data, so the estimate's stability across bandwidths could not be checked."
} else if (n_bw >= 4 && n_bw_sig == 1) {
bw_short <- "fragile"
bw_verdict <- sprintf("The jump reaches significance at exactly one of the %d bandwidths tried and at none of the others; across the grid the estimate runs from %s to %s. An effect that appears at a single bandwidth and disappears at every neighbouring one is not an effect, it is a bandwidth artefact, and it should not be reported as a finding.",
n_bw, r3(min(bw_df$estimate)), r3(max(bw_df$estimate)))
} else if (!sign_consistent) {
bw_short <- "unstable"
bw_verdict <- sprintf("The estimate changes sign across the %d bandwidths tried, from %s to %s. A discontinuity whose direction depends on how wide a window you look through is not identified by this data.",
n_bw, r3(min(bw_df$estimate)), r3(max(bw_df$estimate)))
} else if (main_sig && n_bw_sig >= ceiling(0.8 * n_bw) &&
is.finite(spread_pct) && spread_pct <= 25) {
bw_short <- "stable"
bw_verdict <- sprintf("Across %d bandwidths from %s to %s the estimate stays between %s and %s — a spread of %s of the headline figure — and %d of the %d confidence intervals exclude zero. The finding does not depend on the bandwidth the rule chose.",
n_bw, rs(h * min(mults)), rs(min(h * max(mults), h_max)),
r3(min(bw_df$estimate)), r3(max(bw_df$estimate)),
paste0(r2(spread_pct), "%"), n_bw_sig, n_bw)
} else if (!main_sig && n_bw_sig == 0) {
bw_short <- "consistently null"
bw_verdict <- sprintf("No bandwidth among the %d tried produces a confidence interval that excludes zero, so the absence of a jump is not an artefact of the bandwidth either. The estimates run from %s to %s.",
n_bw, r3(min(bw_df$estimate)), r3(max(bw_df$estimate)))
} else {
bw_short <- "sensitive"
bw_verdict <- sprintf("The estimate moves with the bandwidth: %d of the %d confidence intervals exclude zero and the estimates range from %s to %s%s. Treat the headline number as one point in that range rather than a settled quantity.",
n_bw_sig, n_bw, r3(min(bw_df$estimate)), r3(max(bw_df$estimate)),
if (is.finite(spread_pct))
sprintf(", a spread of %s of the headline figure", paste0(r2(spread_pct), "%"))
else "")
}Step 10: The locality statement — repeated every run, in the user's names
locality <- sprintf("Whatever this estimates, it estimates it only at %s = %s. It describes units sitting right at that threshold and says nothing about %s for units far above or below it — a regression discontinuity buys credibility at the cutoff by giving up everything else.",
running_h, rs(cutoff), outcome_h)