Executive Summary
Multiple-testing correction across 40 tests
The short answer
Of 40 tests, 11 were significant before correction; 10 survive both Holm-Bonferroni (strict family-wise control) and Benjamini-Hochberg (false-discovery-rate control). One uncorrected finding did not survive either correction, flagged as a likely multiple-testing artifact.
The detail
Of 40 tests evaluated, 11 showed p < 0.05 uncorrected. Holm-Bonferroni retained 10 (keeping family-wise error rate at α = 0.05), and Benjamini-Hochberg also retained 10 (keeping false-discovery rate at α = 0.05). The gap of 1 test between uncorrected and corrected counts represents findings the corrections removed: Metric 22 (raw p = 0.0466, Holm p = 1, BH p = 0.1694) crossed the raw threshold but failed both corrections. The strongest survivor is Metric 9 (raw p = 0.0001, Holm p = 0.003, BH p = 0.0019).
What this can't tell you
All 40 p-values were valid; no rows were dropped due to missing or out-of-range data. The analysis captures the full test family as submitted.
Analysis Overview
Holm-Bonferroni and Benjamini-Hochberg corrections applied to 40 p-values from Test Name.
The short answer
Running 40 independent tests inflates false positives: at α = 0.05, random noise alone would produce about 2 false positives by chance. Two corrections recover the signal. Holm-Bonferroni is stricter—it controls the risk of even one false positive anywhere in the family. Benjamini-Hochberg is more lenient—it tolerates a small fraction of false positives among discoveries you report, trading precision for more discoveries.
The detail
The data contains 40 hypothesis tests, each with a raw p-value. Under the all-null scenario, 11 tests would be expected to cross p < 0.05 by chance alone at this family size. Holm-Bonferroni adjusts each p-value to control the family-wise error rate (probability of any false positive ≤ α = 0.05), yielding 10 significant tests. Benjamini-Hochberg adjusts to control the false-discovery rate (expected fraction of false positives among reported discoveries ≤ α = 0.05), also yielding 10 significant tests. Both methods are applied to the same 40 tests.
What this can't tell you
This analysis does not evaluate the practical or clinical importance of the surviving results—only their statistical detectability after correcting for multiple comparisons. Importance depends on effect size and domain context, which lie outside this correction framework.
Data Quality
P-value validation and exclusions.
The short answer
All 40 p-values loaded were valid (between 0 and 1); no rows were dropped. The full test family of 40 entered the correction.
The detail
40 rows were loaded from the dataset with test names and p-values. Validation checks found 0 missing values and 0 out-of-range entries (p-values outside [0, 1]). All 40 tests were valid and included in both Holm and Benjamini-Hochberg adjustments. No exclusions were necessary.
What this can't tell you
Data quality checks confirm structural validity (no missing or impossible values) but do not assess whether the underlying hypothesis tests themselves were correctly specified or whether the p-values reflect appropriate statistical models for the research questions.
Adjusted Results
Every test with raw, Holm, and BH p-values and significance verdicts.
| Test | Raw P | Holm P | Bh P | Sig Raw | Sig Holm | Sig Bh |
|---|---|---|---|---|---|---|
| Metric 9 | 1.00e-04 | 0.003 | 0.0019 | Yes | Yes | Yes |
| Metric 7 | 1.00e-04 | 0.0045 | 0.0019 | Yes | Yes | Yes |
| Metric 4 | 1.00e-04 | 0.0055 | 0.0019 | Yes | Yes | Yes |
| Metric 2 | 3.00e-04 | 0.0112 | 0.003 | Yes | Yes | Yes |
| Metric 1 | 6.00e-04 | 0.0233 | 0.0048 | Yes | Yes | Yes |
| Metric 6 | 7.00e-04 | 0.0256 | 0.0048 | Yes | Yes | Yes |
| Metric 10 | 9.00e-04 | 0.0295 | 0.0048 | Yes | Yes | Yes |
| Metric 8 | 0.001 | 0.0335 | 0.0048 | Yes | Yes | Yes |
| Metric 5 | 0.0011 | 0.0343 | 0.0048 | Yes | Yes | Yes |
| Metric 3 | 0.0013 | 0.0404 | 0.0052 | Yes | Yes | Yes |
| Metric 22 | 0.0466 | 1 | 0.1694 | Yes | No | No |
| Metric 35 | 0.0596 | 1 | 0.1932 | No | No | No |
| Metric 34 | 0.0628 | 1 | 0.1932 | No | No | No |
| Metric 11 | 0.0699 | 1 | 0.1996 | No | No | No |
| Metric 12 | 0.0907 | 1 | 0.2419 | No | No | No |
| Metric 26 | 0.1178 | 1 | 0.2913 | No | No | No |
| Metric 15 | 0.1238 | 1 | 0.2913 | No | No | No |
| Metric 25 | 0.1443 | 1 | 0.3206 | No | No | No |
| Metric 29 | 0.1807 | 1 | 0.3805 | No | No | No |
| Metric 36 | 0.206 | 1 | 0.4119 | No | No | No |
| Metric 16 | 0.2232 | 1 | 0.4252 | No | No | No |
| Metric 24 | 0.2896 | 1 | 0.5236 | No | No | No |
| Metric 27 | 0.3085 | 1 | 0.5236 | No | No | No |
| Metric 39 | 0.3141 | 1 | 0.5236 | No | No | No |
| Metric 32 | 0.3724 | 1 | 0.5958 | No | No | No |
| Metric 20 | 0.3967 | 1 | 0.6103 | No | No | No |
| Metric 13 | 0.4245 | 1 | 0.6108 | No | No | No |
| Metric 38 | 0.4276 | 1 | 0.6108 | No | No | No |
| Metric 33 | 0.5477 | 1 | 0.732 | No | No | No |
| Metric 19 | 0.5771 | 1 | 0.732 | No | No | No |
| Metric 30 | 0.5816 | 1 | 0.732 | No | No | No |
| Metric 40 | 0.5856 | 1 | 0.732 | No | No | No |
| Metric 17 | 0.6274 | 1 | 0.7517 | No | No | No |
| Metric 31 | 0.6389 | 1 | 0.7517 | No | No | No |
| Metric 37 | 0.6804 | 1 | 0.7776 | No | No | No |
| Metric 28 | 0.8161 | 1 | 0.8939 | No | No | No |
| Metric 14 | 0.8269 | 1 | 0.8939 | No | No | No |
| Metric 23 | 0.8585 | 1 | 0.9037 | No | No | No |
| Metric 18 | 0.9477 | 1 | 0.972 | No | No | No |
| Metric 21 | 0.9763 | 1 | 0.9763 | No | No | No |
The short answer
The 10 strongest results all survive both corrections. Metric 9 leads with raw p = 0.0001 (Holm p = 0.003, BH p = 0.0019). The 11th uncorrected finding, Metric 22 (raw p = 0.0466), fails both corrections, marked as a likely artifact.
The detail
Results are sorted smallest p-value first. The top 10 tests—from Metric 9 (raw p = 0.0001) through Metric 3 (raw p = 0.0013)—all carry "Yes" under both Holm and Benjamini-Hochberg columns. Metric 22 (raw p = 0.0466) is marked "Yes" uncorrected but "No" under both Holm (p = 1) and BH (p = 0.1694), indicating it crossed the raw threshold but did not survive correction. The remaining 29 tests showed no raw significance (p ≥ 0.05).
What this can't tell you
This table shows which tests are detectable at α = 0.05 after correction but does not rank them by effect size or practical magnitude. A test marked "Yes" under BH but "No" under Holm is a plausible discovery that meets false-discovery-rate control but not the stricter family-wise standard; deciding whether to act on it requires domain judgment.
What Survives Correction
Significant test counts: uncorrected vs Holm vs BH.
The short answer
Eleven tests crossed p < 0.05 uncorrected; both corrections retained 10. The single removal represents a plausible multiple-testing artifact, leaving a robust core of 10 findings.
The detail
The three counts are: 11 significant uncorrected, 10 after Holm-Bonferroni, and 10 after Benjamini-Hochberg. The gap of 1 test between uncorrected and both corrected methods reflects the filtering of likely false positives. Holm and BH retained identical counts (10) in this case, though they use different standards: Holm controls family-wise error rate (risk of any false positive), while BH controls false-discovery rate (expected fraction of false positives among discoveries).
What this can't tell you
Identical counts under Holm and BH do not mean the methods are equivalent; they reflect this particular dataset's p-value structure. With different data, the counts may diverge. The 10 survivors are detectable; whether they are important depends on effect size and domain context.
P-Value Distribution
10-bin histogram of raw p-values — the real-effects-vs-noise diagnostic.
The short answer
The lowest p-value bin (0.0–0.1) holds 15 of 40 tests, roughly 3.8 times the 4 that a uniform (all-null) distribution would predict. This spike near zero is the hallmark of genuinely non-null effects among your tests.
The detail
P-values were binned into ten 0.1-width intervals. Under the all-null hypothesis, p-values follow a uniform distribution and each bin would hold about 4 tests. The 0.0–0.1 bin observed 15 tests, the 0.1–0.2 bin held 4, and the 0.2–0.3 bin held 3. The remaining seven bins ranged from 0 to 4. The concentration of p-values near zero is inconsistent with a purely null dataset and consistent with a mixture: some tests detecting true effects (small p-values) and others capturing noise (uniform tail).
What this can't tell you
The histogram confirms a signal is present but does not identify which individual tests are true effects versus false positives—that is the role of the multiple-testing corrections. The shape suggests genuine effects exist, but the corrections are still necessary to flag which ones are reliable.
Multiple Comparisons Correction — What Survives?
Takes a set of raw p-values from many tests, applies Holm-Bonferroni and Benjamini-Hochberg corrections, and reports which results are still significant under each method at alpha = 0.05 — plus the classic p-value distribution diagnostic.
Why This Method?
Running many tests inflates false positives: at alpha = 0.05, 20 null tests yield one "significant" result by chance on average. Holm-Bonferroni controls the family-wise error rate (the chance of ANY false positive); Benjamini-Hochberg controls the false discovery rate (the expected fraction of false positives among discoveries). Reporting both, per test, is the standard defensible answer to "which results are real?"
What This Analysis Covers
- Holm and BH adjusted p-values for every test, with significance verdicts
- Significant counts under no correction / Holm / BH
- A 10-bin p-value histogram diagnosing real effects vs mostly-null noise
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {test_label, p_value}. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Core Analysis Pipeline
compute_shared <- function(df, params, col_map = list()) {
# === SHARED EXPORTS ===
# initial_rows/final_rows/rows_removed $ row accounting (rows_removed = invalid p dropped)
# p_name / label_name $ humanized user column names
# n_tests $ integer — valid tests entering the correction
# n_na / n_out_of_range $ integer — dropped: missing vs outside [0,1]
# alpha $ 0.05
# results_all $ data.frame(test, raw_p, holm_p, bh_p, sig_raw, sig_holm, sig_bh) sorted by raw_p
# adjusted_results_df $ results_all capped at 50 rows
# results_truncated $ logical — TRUE when results_all > 50 rows
# n_sig_raw/n_sig_holm/n_sig_bh $ integer — significant counts per method
# significance_summary_df $ data.frame(method, significant_count)
# p_value_distribution_df $ data.frame(p_range, count) — 10 bins of width 0.1
# dist_diagnosis $ character — computed narrative for the histogram shape
# metrics / json_output
# === /SHARED EXPORTS ===Step 1: Locate the mapped columns
initial_rows <- nrow(df)
p_name <- humanize_semantic("p_value", col_map)
label_name <- humanize_semantic("test_label", col_map)
if (!("p_value" %in% names(df))) {
stop(sprintf("column_mapping must map p_value to your p-value column('%s' not found).", p_name))
}Step 2: Labels — OPTIONAL (LAT-2188). The method needs p-values;
labels are presentation. A bare column of p-values is the canonical input for this tool, and demanding a label column is what forced the 2026-08-11 user into mapping one column to both keys. Absent -> positional names.
if ("test_label" %in% names(df)) {
lab <- as.character(df$test_label)
blank <- is.na(lab) | trimws(lab) == ""
if (any(blank)) lab[blank] <- paste0("Test ", which(blank))
} else {
lab <- paste0("Test ", seq_len(nrow(df)))
}Step 3: Coerce p-values (95% rule) and check they fall in [0, 1]
v <- df$p_value
n_censored <- 0L
if (!is.numeric(v)) {Censored notation (LAT-2188): '<0.0001' is how p-values arrive from stats software and published tables. Strip a leading comparison operator and analyse at the reported bound — for '<x' the bound is an UPPER bound on p, so significance conclusions stay conservative. The count is disclosed in the method note below; five of the twenty values in the file that motivated this were '<0.0001'.
raw_chr <- trimws(as.character(v))
stripped <- sub("^(<=|>=|\u2264|\u2265|<|>)\\s*", "", raw_chr)
was_censored <- raw_chr != stripped & !is.na(suppressWarnings(as.numeric(stripped)))
n_censored <- sum(was_censored, na.rm = TRUE)
conv <- suppressWarnings(as.numeric(stripped))
n_orig <- sum(!is.na(v) & raw_chr != "")
if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) {
v <- conv
} else {
stop(sprintf("The mapped p-value column('%s') is not numeric — map a column of raw p-values between 0 and 1 (censored forms like '<0.0001' are accepted).", p_name))
}
}
n_na <- sum(is.na(v))
n_out_of_range <- sum(!is.na(v) & (v < 0 | v > 1))
keep <- !is.na(v) & v >= 0 & v <= 1
p <- as.numeric(v[keep])
labels <- lab[keep]
final_rows <- length(p)
rows_removed <- initial_rows - final_rows
if (final_rows < 5) {
stop(sprintf(
"Only %d valid p-values remained in '%s' after dropping %d invalid (missing or outside 0-1) — need at least 5 tests to correct.",
final_rows, p_name, rows_removed))
}
n_tests <- final_rows
alpha <- 0.05Step 4: Adjust — Holm (family-wise) + Benjamini-Hochberg (FDR)
holm <- p.adjust(p, method = "holm")
bh <- p.adjust(p, method = "BH")
sig_raw <- p < alpha
sig_holm <- holm < alpha
sig_bh <- bh < alpha
n_sig_raw <- sum(sig_raw)
n_sig_holm <- sum(sig_holm)
n_sig_bh <- sum(sig_bh)
results_all <- data.frame(
test = labels,
raw_p = signif(p, 4),
holm_p = signif(holm, 4),
bh_p = signif(bh, 4),
sig_raw = ifelse(sig_raw, "Yes", "No"),
sig_holm = ifelse(sig_holm, "Yes", "No"),
sig_bh = ifelse(sig_bh, "Yes", "No"),
stringsAsFactors = FALSE
)
results_all <- results_all[order(results_all$raw_p), , drop = FALSE]
rownames(results_all) <- NULL
results_truncated <- nrow(results_all) > 50
adjusted_results_df <- head(results_all, 50)
significance_summary_df <- data.frame(
method = c("Uncorrected", "Holm-Bonferroni", "Benjamini-Hochberg"),
significant_count = c(n_sig_raw, n_sig_holm, n_sig_bh),
stringsAsFactors = FALSE
)Step 5: P-value distribution — 10 bins of width 0.1
breaks <- seq(0, 1, by = 0.1)
bin_idx <- pmin(pmax(findInterval(p, breaks, rightmost.closed = TRUE), 1), 10)
bin_labels <- paste0(format(breaks[-11], nsmall = 1), "-", format(breaks[-1], nsmall = 1))
counts <- as.integer(table(factor(bin_idx, levels = 1:10)))
p_value_distribution_df <- data.frame(
p_range = bin_labels,
count = counts,
stringsAsFactors = FALSE
)Classic diagnostic, narrated from the computed bin counts: a spike in the lowest bin suggests true effects; a near-uniform histogram suggests mostly null effects.
uniform_expected <- n_tests / 10
first_bin <- counts[1]
dist_diagnosis <- if (first_bin >= 2 * uniform_expected && first_bin >= 3) {
sprintf(paste0(
"The lowest bin(0.0-0.1) holds %d of %d p-values — about %.1fx the ~%.0f a ",
"uniform(all-null) distribution would put there. That spike near zero is the ",
"classic signature of genuinely non-null effects among your tests."),
first_bin, n_tests, first_bin / max(uniform_expected, 1e-9), uniform_expected)
} else if (first_bin > uniform_expected) {
sprintf(paste0(
"The lowest bin(0.0-0.1) holds %d of %d p-values, modestly above the ~%.0f a ",
"uniform(all-null) distribution would produce — weak evidence of a few true ",
"effects mixed into mostly null tests."),
first_bin, n_tests, uniform_expected)
} else {
sprintf(paste0(
"The histogram is close to uniform(lowest bin: %d of %d p-values vs ~%.0f ",
"expected under all-null) — consistent with mostly null effects, so treat any ",
"uncorrected significant results with suspicion."),
first_bin, n_tests, uniform_expected)
}
metrics <- list(
`Tests Evaluated` = n_tests,
`Invalid P-Values Dropped` = as.integer(rows_removed),
`Significant(Uncorrected)` = as.integer(n_sig_raw),
`Significant(Holm)` = as.integer(n_sig_holm),
`Significant(BH)` = as.integer(n_sig_bh),
`Smallest P-Value` = signif(min(p), 3)
)
json_output <- list(
answer = paste0(
"Of ", n_tests, " tests, ", n_sig_raw, " were significant uncorrected(p < 0.05); ",
"Holm-Bonferroni keeps ", n_sig_holm, " (strict family-wise control) and ",
"Benjamini-Hochberg keeps ", n_sig_bh, " (false-discovery-rate control). ",
dist_diagnosis,
if (n_censored > 0) paste0(
" Note: ", n_censored, " censored p-value",
if (n_censored > 1) "s were" else " was",
" reported as a bound(e.g. '<0.0001') and analysed at that bound — ",
"conservative for significance.") else ""
),
cards = lapply(
c("tldr", "overview", "preprocessing", "adjusted_results",
"significance_summary", "p_value_distribution"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
initial_rows = initial_rows, final_rows = final_rows,
rows_removed = rows_removed,
p_name = p_name, label_name = label_name,
n_tests = n_tests, n_na = n_na, n_out_of_range = n_out_of_range,
n_censored = n_censored,
alpha = alpha,
results_all = results_all,
adjusted_results_df = adjusted_results_df,
results_truncated = results_truncated,
n_sig_raw = n_sig_raw, n_sig_holm = n_sig_holm, n_sig_bh = n_sig_bh,
significance_summary_df = significance_summary_df,
p_value_distribution_df = p_value_distribution_df,
dist_diagnosis = dist_diagnosis,
metrics = metrics, json_output = json_output
)
}
# Card: tldr (tldr)
card_tldr <- function(shared, df, params) {
survival_word <- if (shared$n_sig_raw == 0) {
"No result was significant even before correction."
} else if (shared$n_sig_holm == shared$n_sig_raw) {
"Every uncorrected finding survived even the strictest correction — these results are robust to multiple testing."
} else if (shared$n_sig_holm == 0 && shared$n_sig_bh == 0) {
"No result survived either correction — the uncorrected 'significant' findings are consistent with multiple-testing luck."
} else {
sprintf("%d uncorrected finding(s) did not survive Holm — likely multiple-testing artifacts.",
shared$n_sig_raw - shared$n_sig_holm)
}
text <- paste0(
"Of ", shared$n_tests, " tests, ", shared$n_sig_raw,
" were significant uncorrected(p < ", shared$alpha, "); Holm-Bonferroni keeps ",
shared$n_sig_holm, " (strict family-wise control) and Benjamini-Hochberg keeps ",
shared$n_sig_bh, " (false-discovery-rate control). ", survival_word
)
list(
title = "Executive Summary",
description = paste0("Multiple-testing correction across ", shared$n_tests, " tests"),
metrics = shared$metrics,
text = text
)
}Your turn
Bring your own data and the question you actually need answered.
CympleData Scientist Send me your data and question, I’ll send you the analytics. ds@mcpanalytics.ai