Standard Multiple Comparisons
Executive Summary

Executive Summary

Multiple-testing correction across 40 tests

Tests Evaluated
40
Invalid P-Values Dropped
0
Significant (Uncorrected)
11
Significant (Holm)
10
Significant (BH)
10
Smallest P-Value
1.00e-04
Of 40 tests, 11 were significant uncorrected (p < 0.05); Holm-Bonferroni keeps 10 (strict family-wise control) and Benjamini-Hochberg keeps 10 (false-discovery-rate control). 1 uncorrected finding(s) did not survive Holm — likely multiple-testing artifacts.
What this means

The short answer

Of 40 tests, 11 were significant before correction; 10 survive both Holm-Bonferroni (strict family-wise control) and Benjamini-Hochberg (false-discovery-rate control). One uncorrected finding did not survive either correction, flagged as a likely multiple-testing artifact.

The detail

Of 40 tests evaluated, 11 showed p < 0.05 uncorrected. Holm-Bonferroni retained 10 (keeping family-wise error rate at α = 0.05), and Benjamini-Hochberg also retained 10 (keeping false-discovery rate at α = 0.05). The gap of 1 test between uncorrected and corrected counts represents findings the corrections removed: Metric 22 (raw p = 0.0466, Holm p = 1, BH p = 0.1694) crossed the raw threshold but failed both corrections. The strongest survivor is Metric 9 (raw p = 0.0001, Holm p = 0.003, BH p = 0.0019).

What this can't tell you

All 40 p-values were valid; no rows were dropped due to missing or out-of-range data. The analysis captures the full test family as submitted.

Overview

Analysis Overview

Holm-Bonferroni and Benjamini-Hochberg corrections applied to 40 p-values from Test Name.

N Tests40
Alpha0.05
N Significant Raw11
N Significant Holm10
N Significant Bh10
What this means

The short answer

Running 40 independent tests inflates false positives: at α = 0.05, random noise alone would produce about 2 false positives by chance. Two corrections recover the signal. Holm-Bonferroni is stricter—it controls the risk of even one false positive anywhere in the family. Benjamini-Hochberg is more lenient—it tolerates a small fraction of false positives among discoveries you report, trading precision for more discoveries.

The detail

The data contains 40 hypothesis tests, each with a raw p-value. Under the all-null scenario, 11 tests would be expected to cross p < 0.05 by chance alone at this family size. Holm-Bonferroni adjusts each p-value to control the family-wise error rate (probability of any false positive ≤ α = 0.05), yielding 10 significant tests. Benjamini-Hochberg adjusts to control the false-discovery rate (expected fraction of false positives among reported discoveries ≤ α = 0.05), also yielding 10 significant tests. Both methods are applied to the same 40 tests.

What this can't tell you

This analysis does not evaluate the practical or clinical importance of the surviving results—only their statistical detectability after correcting for multiple comparisons. Importance depends on effect size and domain context, which lie outside this correction framework.

Data Preparation

Data Quality

P-value validation and exclusions.

Initial Rows40
Final Rows40
Rows Removed0
N Missing0
N Out Of Range0
What this means

The short answer

All 40 p-values loaded were valid (between 0 and 1); no rows were dropped. The full test family of 40 entered the correction.

The detail

40 rows were loaded from the dataset with test names and p-values. Validation checks found 0 missing values and 0 out-of-range entries (p-values outside [0, 1]). All 40 tests were valid and included in both Holm and Benjamini-Hochberg adjustments. No exclusions were necessary.

What this can't tell you

Data quality checks confirm structural validity (no missing or impossible values) but do not assess whether the underlying hypothesis tests themselves were correctly specified or whether the p-values reflect appropriate statistical models for the research questions.

Data Table

Adjusted Results

Every test with raw, Holm, and BH p-values and significance verdicts.

TestRaw PHolm PBh PSig RawSig HolmSig Bh
Metric 91.00e-040.0030.0019YesYesYes
Metric 71.00e-040.00450.0019YesYesYes
Metric 41.00e-040.00550.0019YesYesYes
Metric 23.00e-040.01120.003YesYesYes
Metric 16.00e-040.02330.0048YesYesYes
Metric 67.00e-040.02560.0048YesYesYes
Metric 109.00e-040.02950.0048YesYesYes
Metric 80.0010.03350.0048YesYesYes
Metric 50.00110.03430.0048YesYesYes
Metric 30.00130.04040.0052YesYesYes
Metric 220.046610.1694YesNoNo
Metric 350.059610.1932NoNoNo
Metric 340.062810.1932NoNoNo
Metric 110.069910.1996NoNoNo
Metric 120.090710.2419NoNoNo
Metric 260.117810.2913NoNoNo
Metric 150.123810.2913NoNoNo
Metric 250.144310.3206NoNoNo
Metric 290.180710.3805NoNoNo
Metric 360.20610.4119NoNoNo
Metric 160.223210.4252NoNoNo
Metric 240.289610.5236NoNoNo
Metric 270.308510.5236NoNoNo
Metric 390.314110.5236NoNoNo
Metric 320.372410.5958NoNoNo
Metric 200.396710.6103NoNoNo
Metric 130.424510.6108NoNoNo
Metric 380.427610.6108NoNoNo
Metric 330.547710.732NoNoNo
Metric 190.577110.732NoNoNo
Metric 300.581610.732NoNoNo
Metric 400.585610.732NoNoNo
Metric 170.627410.7517NoNoNo
Metric 310.638910.7517NoNoNo
Metric 370.680410.7776NoNoNo
Metric 280.816110.8939NoNoNo
Metric 140.826910.8939NoNoNo
Metric 230.858510.9037NoNoNo
Metric 180.947710.972NoNoNo
Metric 210.976310.9763NoNoNo
What this means

The short answer

The 10 strongest results all survive both corrections. Metric 9 leads with raw p = 0.0001 (Holm p = 0.003, BH p = 0.0019). The 11th uncorrected finding, Metric 22 (raw p = 0.0466), fails both corrections, marked as a likely artifact.

The detail

Results are sorted smallest p-value first. The top 10 tests—from Metric 9 (raw p = 0.0001) through Metric 3 (raw p = 0.0013)—all carry "Yes" under both Holm and Benjamini-Hochberg columns. Metric 22 (raw p = 0.0466) is marked "Yes" uncorrected but "No" under both Holm (p = 1) and BH (p = 0.1694), indicating it crossed the raw threshold but did not survive correction. The remaining 29 tests showed no raw significance (p ≥ 0.05).

What this can't tell you

This table shows which tests are detectable at α = 0.05 after correction but does not rank them by effect size or practical magnitude. A test marked "Yes" under BH but "No" under Holm is a plausible discovery that meets false-discovery-rate control but not the stricter family-wise standard; deciding whether to act on it requires domain judgment.

Visualization

What Survives Correction

Significant test counts: uncorrected vs Holm vs BH.

What this means

The short answer

Eleven tests crossed p < 0.05 uncorrected; both corrections retained 10. The single removal represents a plausible multiple-testing artifact, leaving a robust core of 10 findings.

The detail

The three counts are: 11 significant uncorrected, 10 after Holm-Bonferroni, and 10 after Benjamini-Hochberg. The gap of 1 test between uncorrected and both corrected methods reflects the filtering of likely false positives. Holm and BH retained identical counts (10) in this case, though they use different standards: Holm controls family-wise error rate (risk of any false positive), while BH controls false-discovery rate (expected fraction of false positives among discoveries).

What this can't tell you

Identical counts under Holm and BH do not mean the methods are equivalent; they reflect this particular dataset's p-value structure. With different data, the counts may diverge. The 10 survivors are detectable; whether they are important depends on effect size and domain context.

Visualization

P-Value Distribution

10-bin histogram of raw p-values — the real-effects-vs-noise diagnostic.

What this means

The short answer

The lowest p-value bin (0.0–0.1) holds 15 of 40 tests, roughly 3.8 times the 4 that a uniform (all-null) distribution would predict. This spike near zero is the hallmark of genuinely non-null effects among your tests.

The detail

P-values were binned into ten 0.1-width intervals. Under the all-null hypothesis, p-values follow a uniform distribution and each bin would hold about 4 tests. The 0.0–0.1 bin observed 15 tests, the 0.1–0.2 bin held 4, and the 0.2–0.3 bin held 3. The remaining seven bins ranged from 0 to 4. The concentration of p-values near zero is inconsistent with a purely null dataset and consistent with a mixture: some tests detecting true effects (small p-values) and others capturing noise (uniform tail).

What this can't tell you

The histogram confirms a signal is present but does not identify which individual tests are true effects versus false positives—that is the role of the multiple-testing corrections. The shape suggests genuine effects exist, but the corrections are still necessary to flag which ones are reliable.

Rate this report Was this the answer you needed?
The exact source that produced this report — yours to keep, read, and re-run.
Download PDF
How this was computed method · R source · citation
The code that did it

Multiple Comparisons Correction — What Survives?

Takes a set of raw p-values from many tests, applies Holm-Bonferroni and Benjamini-Hochberg corrections, and reports which results are still significant under each method at alpha = 0.05 — plus the classic p-value distribution diagnostic.

Why This Method?

Running many tests inflates false positives: at alpha = 0.05, 20 null tests yield one "significant" result by chance on average. Holm-Bonferroni controls the family-wise error rate (the chance of ANY false positive); Benjamini-Hochberg controls the false discovery rate (the expected fraction of false positives among discoveries). Reporting both, per test, is the standard defensible answer to "which results are real?"

What This Analysis Covers

  • Holm and BH adjusted p-values for every test, with significance verdicts
  • Significant counts under no correction / Holm / BH
  • A 10-bin p-value histogram diagnosing real effects vs mostly-null noise

Standard Library

Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {test_label, p_value}. All narrative is derived from the user's own column names and computed values.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Core Analysis Pipeline

compute_shared <- function(df, params, col_map = list()) {
  # === SHARED EXPORTS ===
  #   initial_rows/final_rows/rows_removed  $ row accounting (rows_removed = invalid p dropped)
  #   p_name / label_name    $ humanized user column names
  #   n_tests                $ integer — valid tests entering the correction
  #   n_na / n_out_of_range  $ integer — dropped: missing vs outside [0,1]
  #   alpha                  $ 0.05
  #   results_all            $ data.frame(test, raw_p, holm_p, bh_p, sig_raw, sig_holm, sig_bh) sorted by raw_p
  #   adjusted_results_df    $ results_all capped at 50 rows
  #   results_truncated      $ logical — TRUE when results_all > 50 rows
  #   n_sig_raw/n_sig_holm/n_sig_bh $ integer — significant counts per method
  #   significance_summary_df $ data.frame(method, significant_count)
  #   p_value_distribution_df $ data.frame(p_range, count) — 10 bins of width 0.1
  #   dist_diagnosis          $ character — computed narrative for the histogram shape
  #   metrics / json_output
  # === /SHARED EXPORTS ===

Step 1: Locate the mapped columns

initial_rows <- nrow(df)
  p_name <- humanize_semantic("p_value", col_map)
  label_name <- humanize_semantic("test_label", col_map)
  if (!("p_value" %in% names(df))) {
    stop(sprintf("column_mapping must map p_value to your p-value column(&#x27;%s' not found).", p_name))
  }

Step 2: Labels — OPTIONAL (LAT-2188). The method needs p-values;

labels are presentation. A bare column of p-values is the canonical input for this tool, and demanding a label column is what forced the 2026-08-11 user into mapping one column to both keys. Absent -> positional names.

if ("test_label" %in% names(df)) {
    lab <- as.character(df$test_label)
    blank <- is.na(lab) | trimws(lab) == ""
    if (any(blank)) lab[blank] <- paste0("Test ", which(blank))
  } else {
    lab <- paste0("Test ", seq_len(nrow(df)))
  }

Step 3: Coerce p-values (95% rule) and check they fall in [0, 1]

v <- df$p_value
  n_censored <- 0L
  if (!is.numeric(v)) {

Censored notation (LAT-2188): '<0.0001' is how p-values arrive from stats software and published tables. Strip a leading comparison operator and analyse at the reported bound — for '<x' the bound is an UPPER bound on p, so significance conclusions stay conservative. The count is disclosed in the method note below; five of the twenty values in the file that motivated this were '<0.0001'.

raw_chr <- trimws(as.character(v))
    stripped <- sub("^(<=|>=|\u2264|\u2265|<|>)\\s*", "", raw_chr)
    was_censored <- raw_chr != stripped & !is.na(suppressWarnings(as.numeric(stripped)))
    n_censored <- sum(was_censored, na.rm = TRUE)
    conv <- suppressWarnings(as.numeric(stripped))
    n_orig <- sum(!is.na(v) & raw_chr != "")
    if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) {
      v <- conv
    } else {
      stop(sprintf("The mapped p-value column(&#x27;%s') is not numeric — map a column of raw p-values between 0 and 1 (censored forms like '<0.0001' are accepted).", p_name))
    }
  }
  n_na <- sum(is.na(v))
  n_out_of_range <- sum(!is.na(v) & (v < 0 | v > 1))
  keep <- !is.na(v) & v >= 0 & v <= 1
  p <- as.numeric(v[keep])
  labels <- lab[keep]
  final_rows <- length(p)
  rows_removed <- initial_rows - final_rows
  if (final_rows < 5) {
    stop(sprintf(
      "Only %d valid p-values remained in &#x27;%s' after dropping %d invalid (missing or outside 0-1) — need at least 5 tests to correct.",
      final_rows, p_name, rows_removed))
  }
  n_tests <- final_rows
  alpha <- 0.05

Step 4: Adjust — Holm (family-wise) + Benjamini-Hochberg (FDR)

holm <- p.adjust(p, method = "holm")
  bh   <- p.adjust(p, method = "BH")
  sig_raw  <- p    < alpha
  sig_holm <- holm < alpha
  sig_bh   <- bh   < alpha
  n_sig_raw  <- sum(sig_raw)
  n_sig_holm <- sum(sig_holm)
  n_sig_bh   <- sum(sig_bh)

  results_all <- data.frame(
    test     = labels,
    raw_p    = signif(p, 4),
    holm_p   = signif(holm, 4),
    bh_p     = signif(bh, 4),
    sig_raw  = ifelse(sig_raw,  "Yes", "No"),
    sig_holm = ifelse(sig_holm, "Yes", "No"),
    sig_bh   = ifelse(sig_bh,   "Yes", "No"),
    stringsAsFactors = FALSE
  )
  results_all <- results_all[order(results_all$raw_p), , drop = FALSE]
  rownames(results_all) <- NULL
  results_truncated <- nrow(results_all) > 50
  adjusted_results_df <- head(results_all, 50)

  significance_summary_df <- data.frame(
    method = c("Uncorrected", "Holm-Bonferroni", "Benjamini-Hochberg"),
    significant_count = c(n_sig_raw, n_sig_holm, n_sig_bh),
    stringsAsFactors = FALSE
  )

Step 5: P-value distribution — 10 bins of width 0.1

breaks <- seq(0, 1, by = 0.1)
  bin_idx <- pmin(pmax(findInterval(p, breaks, rightmost.closed = TRUE), 1), 10)
  bin_labels <- paste0(format(breaks[-11], nsmall = 1), "-", format(breaks[-1], nsmall = 1))
  counts <- as.integer(table(factor(bin_idx, levels = 1:10)))
  p_value_distribution_df <- data.frame(
    p_range = bin_labels,
    count = counts,
    stringsAsFactors = FALSE
  )

Classic diagnostic, narrated from the computed bin counts: a spike in the lowest bin suggests true effects; a near-uniform histogram suggests mostly null effects.

uniform_expected <- n_tests / 10
  first_bin <- counts[1]
  dist_diagnosis <- if (first_bin >= 2 * uniform_expected && first_bin >= 3) {
    sprintf(paste0(
      "The lowest bin(0.0-0.1) holds %d of %d p-values — about %.1fx the ~%.0f a ",
      "uniform(all-null) distribution would put there. That spike near zero is the ",
      "classic signature of genuinely non-null effects among your tests."),
      first_bin, n_tests, first_bin / max(uniform_expected, 1e-9), uniform_expected)
  } else if (first_bin > uniform_expected) {
    sprintf(paste0(
      "The lowest bin(0.0-0.1) holds %d of %d p-values, modestly above the ~%.0f a ",
      "uniform(all-null) distribution would produce — weak evidence of a few true ",
      "effects mixed into mostly null tests."),
      first_bin, n_tests, uniform_expected)
  } else {
    sprintf(paste0(
      "The histogram is close to uniform(lowest bin: %d of %d p-values vs ~%.0f ",
      "expected under all-null) — consistent with mostly null effects, so treat any ",
      "uncorrected significant results with suspicion."),
      first_bin, n_tests, uniform_expected)
  }

  metrics <- list(
    `Tests Evaluated`         = n_tests,
    `Invalid P-Values Dropped` = as.integer(rows_removed),
    `Significant(Uncorrected)` = as.integer(n_sig_raw),
    `Significant(Holm)`      = as.integer(n_sig_holm),
    `Significant(BH)`        = as.integer(n_sig_bh),
    `Smallest P-Value`        = signif(min(p), 3)
  )

  json_output <- list(
    answer = paste0(
      "Of ", n_tests, " tests, ", n_sig_raw, " were significant uncorrected(p < 0.05); ",
      "Holm-Bonferroni keeps ", n_sig_holm, " (strict family-wise control) and ",
      "Benjamini-Hochberg keeps ", n_sig_bh, " (false-discovery-rate control). ",
      dist_diagnosis,
      if (n_censored > 0) paste0(
        " Note: ", n_censored, " censored p-value",
        if (n_censored > 1) "s were" else " was",
        " reported as a bound(e.g. &#x27;<0.0001') and analysed at that bound — ",
        "conservative for significance.") else ""
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "adjusted_results",
        "significance_summary", "p_value_distribution"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  list(
    initial_rows = initial_rows, final_rows = final_rows,
    rows_removed = rows_removed,
    p_name = p_name, label_name = label_name,
    n_tests = n_tests, n_na = n_na, n_out_of_range = n_out_of_range,
    n_censored = n_censored,
    alpha = alpha,
    results_all = results_all,
    adjusted_results_df = adjusted_results_df,
    results_truncated = results_truncated,
    n_sig_raw = n_sig_raw, n_sig_holm = n_sig_holm, n_sig_bh = n_sig_bh,
    significance_summary_df = significance_summary_df,
    p_value_distribution_df = p_value_distribution_df,
    dist_diagnosis = dist_diagnosis,
    metrics = metrics, json_output = json_output
  )
}

# Card: tldr (tldr)
card_tldr <- function(shared, df, params) {
  survival_word <- if (shared$n_sig_raw == 0) {
    "No result was significant even before correction."
  } else if (shared$n_sig_holm == shared$n_sig_raw) {
    "Every uncorrected finding survived even the strictest correction — these results are robust to multiple testing."
  } else if (shared$n_sig_holm == 0 && shared$n_sig_bh == 0) {
    "No result survived either correction — the uncorrected &#x27;significant' findings are consistent with multiple-testing luck."
  } else {
    sprintf("%d uncorrected finding(s) did not survive Holm — likely multiple-testing artifacts.",
            shared$n_sig_raw - shared$n_sig_holm)
  }
  text <- paste0(
    "Of ", shared$n_tests, " tests, ", shared$n_sig_raw,
    " were significant uncorrected(p < ", shared$alpha, "); Holm-Bonferroni keeps ",
    shared$n_sig_holm, " (strict family-wise control) and Benjamini-Hochberg keeps ",
    shared$n_sig_bh, " (false-discovery-rate control). ", survival_word
  )
  list(
    title = "Executive Summary",
    description = paste0("Multiple-testing correction across ", shared$n_tests, " tests"),
    metrics = shared$metrics,
    text = text
  )
}
Your data has more stories to tell.Run any analysis on your own data: R modules you own, interactive reports, AI insights, and PDF export. 500 free credits when you finish onboarding.
Try Free — No SignupSign Up Free

Your turn

Bring your own data and the question you actually need answered.

CympleData Scientist Send me your data and question, I’ll send you the analytics. ds@mcpanalytics.ai

Cite this analysis

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing