Standard Ab Sequential
Executive Summary

Can I Stop?

Stop — a winner has separated

Decision
Stop — a winner has separated
Observations Analysed
6000
Effect Estimate
3.84
Always-Valid Low
1.34
Always-Valid High
6.34
Always-Valid p
< 0.001
Fixed-Horizon p
< 0.001
Boundary Crossed At
observation 1,671
Analysis Points
5973
Fixed-Horizon Crossings
5219
Fixed-Horizon Excursions
12
Always-Valid Crossings
4290
False Alarms Avoided
929
Width Ratio
1.56
You can stop for a winner: Variant B has separated from Control by 3.84 pp, and the always-valid confidence sequence (1.34 pp to 6.34 pp) excludes zero. That conclusion holds even though the experiment was re-tested after every one of its 6,000 observations. The boundary was first crossed at observation 1,671. An ordinary fixed-horizon test on the same data would have shown p below 0.05 at 929 point(s) where the always-valid boundary was not crossed, across 12 separate excursion(s). The guarantee is not free: the always-valid interval is 1.56 times wider than the fixed-horizon interval at this sample size, and it would take about 15,512 observations (2.59 times the current sample) for it to become as tight as the fixed-horizon interval already is. The guarantee covers this one comparison as specified; it does not protect a metric, a segment or an arm picked after seeing the results.
Suggested Interpretation

The short answer

You can stop for a winner: Variant B has separated from Control by 3.84 pp, and the always-valid confidence sequence (1.34 pp to 6.34 pp) excludes zero—even though the experiment was re-tested after every one of its 6,000 observations.

The detail

The boundary was first crossed at observation 1,671. A fixed-horizon test on the same data would have shown p below 0.05 at 929 points where the always-valid boundary was not crossed, across 12 separate excursions. The always-valid p-value is < 0.001 and the fixed-horizon p-value is also < 0.001 at this sample size. The guarantee costs precision: the always-valid interval is 1.56 times wider than the fixed-horizon interval (2.50 pp vs 1.60 pp half-width), and reaching today's fixed-horizon tightness would require about 15,512 observations (2.59 times the current sample).

What this can't tell you

The guarantee covers this one comparison as specified. It does not protect a metric, segment, or arm picked after seeing the results.

Overview

Analysis Overview

Always-valid sequential monitoring of Variant B against Control across 6,000 observations.

N Observations6000
N Control2975
N Treatment3025
N Analysis Points5973
Suggested Interpretation

The short answer

A mixture sequential probability ratio test replaces the ordinary 5% test with one valid at every look. The fixed-horizon test earns its 5% false-positive rate only if you decide your sample size before seeing data and look once. Here, 6,000 observations were ordered by Signup Date and re-tested after every observation; the control arm was identified by name, not outcome. The cumulative effect and both boundaries—always-valid and fixed-horizon—are computed on your data so you can compare them yourself.

The detail

The analysis monitored 6,000 observations (n_control = 2,975; n_treatment = 3,025) at 5,973 analysis points. The effect is the difference in mean Converted between Variant B and Control. A mixture sequential probability ratio test with a normal mixture over the true effect produces a boundary valid at all simultaneous looks. An ordinary Welch t-test is also run at every point so both corridors can be displayed on the same data.

What this can't tell you

This boundary protects this one comparison as specified. It does not protect a metric, segment, or arm chosen after seeing results.

Data Preparation

Data Quality

Row accounting, arm sizes, and how the ordering was resolved.

Initial Rows6000
Final Rows6000
Rows Removed0
Suggested Interpretation

The short answer

All 6,000 rows were analysed; none were incomplete. The outcome Converted is already binary (0 or 1), so the effect is a difference in conversion rates. Control has 2,975 rows and Variant B has 3,025. When 5,850 rows tied on arrival time, they kept their original file order and stepped through the monitoring path one at a time.

The detail

6,000 rows loaded; 6,000 analysed; 0 rows removed. The Converted column holds only 0 and 1 values, making the effect a difference in proportions. Control contributed 2,975 observations and Variant B contributed 3,025. Tied arrival values (5,850 rows) were resolved by preserving the original file order, so each observation enters the monitoring sequence individually.

What this can't tell you

This analysis cannot verify from the file that assignment to Assigned Group was actually random, or that each row represents a distinct person. Both are inherited assumptions about how the experiment was run.

Visualization

Effect and Boundaries Over the Experiment

The cumulative effect with the always-valid and fixed-horizon boundaries drawn over it.

Suggested Interpretation

The short answer

The cumulative effect traces an upward pattern from observation 1,671 onward, leaving the always-valid corridor at that point and remaining outside it through the final observation. Both corridors narrow as data accumulates, but the always-valid boundary (outer pair) narrows more slowly because it pays for every look you might take.

The detail

The cumulative effect estimate (Variant B minus Control) starts at 4.4444 pp at observation 28 and drifts downward through observation 154 (−3.8961 pp), then climbs steadily from observation 218 onward. At observation 1,671 the effect exits the always-valid corridor on the upper side and stays outside. The inner corridor (fixed-horizon 0.05 boundary) is narrower everywhere; this is why watching it makes teams stop too early. Both corridors narrow as n increases, but the always-valid corridor widens more slowly because it is protecting against all possible looks, not just one.

What this can't tell you

The pattern shown is descriptive of the monitoring path as data arrived in calendar order. It does not explain why the effect moved as it did.

Visualization

The Two p-Values Side by Side

The fixed-horizon p-value and the always-valid p-value at every analysis point.

Suggested Interpretation

The short answer

The fixed-horizon p-value wanders below 0.05 and back above it across 929 analysis points where the always-valid p-value does not cross that threshold. The always-valid p-value only ever moves down and is a legitimate stopping criterion at any point.

The detail

Both p-values were computed at all 5,973 analysis points on the same cumulative data. The fixed-horizon p-value fell below 0.05 at 5,219 points in 12 separate excursions; the always-valid p-value (the running minimum of the inverse mixture likelihood ratio) crossed below 0.05 at 4,290 points. The 929 points where only the fixed-horizon test fired are moments this experiment could have been stopped on evidence that had not actually arrived. The fixed-horizon series is free to wander because it is recomputed from scratch at each point; the always-valid series is monotonic because it records the best evidence seen so far.

What this can't tell you

These counts assume a look after every observation—the most aggressive peeking schedule. A weekly review cadence would see fewer opportunities to be misled, but not zero.

Data Table

The Decision, Line by Line

Everything the stopping decision rests on.

ItemDetail
DecisionStop — a winner has separated
Always-valid boundary crossedyes, first at observation 1,671
Observations analysed6,000
Effect (Variant B minus Control)3.84 pp
Always-valid 95% confidence sequence1.34 pp to 6.34 pp
Always-valid p-value< 0.001
Fixed-horizon p-value at this sample size< 0.001
Smallest effect this test was set up to find3.18 pp
Control armControl (n = 2,975)
Treatment armVariant B (n = 3,025)
Suggested Interpretation

The short answer

Stop—a winner has separated. The always-valid confidence sequence (1.34 pp to 6.34 pp) excludes zero, and the boundary was crossed at observation 1,671.

The detail

The effect estimate is 3.84 pp (Variant B minus Control). The always-valid 95% confidence sequence is 1.34 pp to 6.34 pp; the always-valid p-value is < 0.001 and the fixed-horizon p-value at this sample size is also < 0.001. The smallest effect this test was set up to find is 3.18 pp (tau). Control has n = 2,975 and Variant B has n = 3,025. The decision rule has three exits and one refusal: stop for a winner when the confidence sequence excludes zero (true here), stop for futility when it lies entirely within 3.18 pp (false), keep running when neither is true but both are reachable (not applicable), or refuse when the interval is wider than 3.18 pp and no boundary is reachable (not applicable).

What this can't tell you

Nothing in this table settles whether assignment was random or whether each row is a distinct person.

Data Table

What Peeking Would Have Cost You

The fixed-horizon test and the always-valid boundary, counted over the same analysis points.

CheckValue
Analysis points examined (one after every observation)5,973
Points where the fixed-horizon p-value was below 0.055,219
Separate excursions below 0.05 by the fixed-horizon p-value12
Points where the always-valid boundary was crossed4,290
Points where the fixed-horizon test said stop but the always-valid boundary did not929
First point the fixed-horizon test would have stoppedobservation 701
First point the always-valid boundary was crossedobservation 1,671
Suggested Interpretation

The short answer

Over 5,973 analysis points, the fixed-horizon test would have stopped at 929 moments where the always-valid boundary had not yet been crossed. The fixed-horizon p-value fell below 0.05 at 5,219 points in 12 separate excursions; the always-valid boundary crossed at 4,290.

The detail

The fixed-horizon test would have first stopped at observation 701; the always-valid boundary was first crossed at observation 1,671—a gap of 970 observations. This assumes the most aggressive schedule possible: a look after every observation. Both tests were evaluated at all 5,973 analysis points on the same cumulative data. The 929 points where only the fixed-horizon test fired represent the concrete cost of watching a dashboard that is not always-valid.

What this can't tell you

A weekly or monthly review cadence would see fewer opportunities to be misled by the fixed-horizon test, but the gap between the two stopping points (observation 701 vs 1,671) would still exist.

Data Table

The Price of Being Allowed to Peek

How much extra data the always-valid guarantee costs on this experiment.

MeasureValue
Always-valid half-width at this sample size2.50 pp
Fixed-horizon half-width at this sample size1.60 pp
How much wider the always-valid interval is1.56 times
Observations needed for the always-valid interval to be as tight as the fixed-horizon one is now15,512
That as a multiple of the current sample size2.59 times
Suggested Interpretation

The short answer

The always-valid interval is 1.56 times wider than the fixed-horizon interval at this sample size. Reaching today's fixed-horizon precision under the always-valid boundary would require about 15,512 observations—2.59 times what you have.

The detail

The always-valid half-width is 2.50 pp; the fixed-horizon half-width is 1.60 pp. The ratio is 1.56. If you genuinely can commit to a single analysis at a sample size fixed in advance, the fixed-horizon test is the more efficient choice. This tool is for the common case where that commitment is not realistic and you need the right to look whenever you like.

What this can't tell you

These calculations assume the same power and type-I error rate for both methods. The trade-off is inherent to always-valid inference and cannot be eliminated.

Data Table

Does the Verdict Survive a Different Alternative?

The same monitoring path re-run against three pre-specified alternatives.

AlternativeAlternative SizeHalf WidthDecisionEver Crossed
0.05 x the outcome's standard deviation1.592.53stop for a winneryes
0.10 x the outcome's standard deviation3.182.50stop for a winneryes
0.20 x the outcome's standard deviation6.352.62stop for a winneryes
Suggested Interpretation

The short answer

The boundary's tuning constant (tau) was set to 0.10 times the outcome's standard deviation. Re-running the entire path at half and double that value (0.05 and 0.20 times the standard deviation) yields the same verdict—stop for a winner—under all three. The conclusion does not hinge on the choice of tau.

The detail

Tau defines the effect size the test is aimed at, set here to 0.10 × outcome standard deviation (3.18 pp). At tau = 0.05 × SD (1.59 pp), the half-width is 2.53 pp and the verdict is stop for a winner. At tau = 0.10 × SD (3.18 pp), the half-width is 2.50 pp and the verdict is stop for a winner. At tau = 0.20 × SD (6.35 pp), the half-width is 2.62 pp and the verdict is stop for a winner. All three alternatives crossed the boundary. Tau should strictly be fixed before the experiment; it is read from the outcome's spread here rather than from the observed gap between arms, so it does not borrow the answer.

What this can't tell you

Tau sensitivity does not test robustness to violations of the normal-mixture assumption or to departures from the Welch t-test model. The choice of tau before the experiment is a design decision, not a data-driven one.

Data Table

Methods and Disclosure

Every formula, constant and assumption behind the decision.

ItemDetail
TestMixture sequential probability ratio test (mSPRT) on the difference in means between the two arms, implemented directly from the normal mixture — no group-sequential package is used.
BoundaryRobbins' normal-mixture boundary. The mixture likelihood ratio is Lambda = sqrt(V/(V+tau^2)) exp(delta^2 tau^2 / (2 V (V+tau^2))) with V the sampling variance of the effect; Ville's inequality bounds the chance that Lambda ever reaches 1/alpha at alpha, which is what makes continuous monitoring safe.
Confidence sequenceInverting the boundary gives effect +/- h with h = sqrt(2 V (V+tau^2)/tau^2 log(sqrt((V+tau^2)/V)/alpha)); this interval covers the true effect at every sample size at once, not just at one pre-chosen sample size.
Pre-specified alternative (tau)3.18 pp — set at 0.10 times the outcome's standard deviation (31.76). It is the effect size the boundary is tuned to find, and it is read from how much the outcome varies, not from the difference between the arms.
Significance levelalpha = 0.05, applied to the whole monitoring path rather than to a single look.
Analysis scheduleThe test is evaluated after every single observation from the point where both arms reach 10 rows — the most aggressive peeking possible — giving 5,973 analysis points.
Fixed-horizon comparisonAt each of those points a Welch two-sample t-test is computed on the same cumulative data, purely so the two boundaries can be compared on this experiment's own numbers.
Effect definitionMean of Variant B minus mean of Control, both being conversion rates, reported in percentage points.
Control arm chosen byControl — its name identifies it as the control.
Arrival order read ascalendar dates in the column 'Signup Date'; 5,850 row(s) share an arrival value with an earlier row and keep their original file order.
What this cannot doThe guarantee is about the stopping rule, not about everything else. It holds for this one pre-specified comparison of these two arms on this one outcome. It does not cover choosing the metric, the segment or the arm after seeing the results, it cannot detect repeat visits by the same person, and it does not know whether the assignment was actually random.
Suggested Interpretation

The short answer

The test is a mixture sequential probability ratio test (mSPRT) using Robbins' normal-mixture boundary. Ville's inequality guarantees that the chance the mixture likelihood ratio ever reaches 1/0.05 is at most 0.05, no matter how many times you look. The confidence sequence inverts this boundary and covers the true effect at every sample size simultaneously.

The detail

The mixture likelihood ratio is Lambda = sqrt(V/(V+tau²)) exp(δ²τ²/(2V(V+τ²))) with V the sampling variance of the effect and tau = 3.18 pp. The confidence sequence is effect ± h where h = sqrt(2V(V+τ²)/τ² log(sqrt((V+τ²)/V)/α)) with α = 0.05. The analysis evaluates the test after every observation from the point both arms reach 10 rows, giving 5,973 analysis points. A Welch two-sample t-test is computed at each point for comparison. The control arm was identified by name; arrival order was read as calendar dates from Signup Date, with 5,850 ties resolved by preserving file order.

What this can't tell you

The normal mixture is asymptotic and improves with sample size. Variance is estimated from the same data being tested, so coverage at very small samples is optimistic. The guarantee covers the stopping rule and this one pre-specified comparison only—it does not protect against choosing the metric, segment, or arm after seeing results, detecting repeat visits, or verifying random assignment.

Methodology

Methodology

Statistical methodology and diagnostics for Sequential A/B Test — Peek Safely

Statistical Method

Sequential A/B Test — Peek Safely

Standard-library analysis: can I stop this experiment early? Upload an ordered experiment log — the variant each row saw, what happened, and the date or sequence the observation arrived — and get inference that stays valid under continuous monitoring. The analysis re-tests after every single observation using a mixture sequential probability ratio test, and returns the always-valid confidence sequence over time, the current decision (stop for a winner, stop for futility, keep running, or not enough data for any boundary to be reachable), the boundary that was crossed and when, and the cumulative effect charted with both boundaries drawn over it. The naive fixed-horizon p-value is computed alongside at the same points, so the report can show — from your own data, not from an assertion — how many times it would have crossed 0.05 while the always-valid boundary did not. The report also states what the guarantee costs in power, in this experiment's own numbers.

Data
N = 6000 observations
Assumptions
  • Each row is one independent observation assigned to one of exactly two arms
  • Assignment to the arms was random, so the arms are comparable on everything except the variant
  • The arrival column reflects the order the observations actually arrived — sequential inference is entirely about that order
  • The outcome is numeric or a 0/1 flag, and the two arms' means are the quantity of interest
  • The comparison — these two arms, this outcome — was decided before the data was inspected
Limitations
  • Always-valid inference buys the right to peek by widening the interval, so it needs more data than a fixed-horizon test to reach the same conclusion; the report states that cost in your own numbers rather than presenting peeking as free
  • The guarantee covers the stopping rule only. It does not protect against choosing the metric, the segment or the arm after seeing the data, and no single file can reveal that this happened
  • The boundary has one tuning constant, the effect size it is aimed at. It should strictly be fixed before the experiment; here it is read from the outcome's spread, and the report re-runs the verdict at half and double that value
  • The mixture boundary uses a normal approximation with the variance estimated from the same data, so at small samples the stated coverage is optimistic
Software & Citation
MCP Analytics · mcpanalytics.ai
Code Appendix

Analysis Code

Complete R source code for this analysis

Sequential A/B Test — Peek Safely

Reads an ordered experiment log (who saw which variant, what happened, and when the observation arrived) and answers the question a running experiment actually raises: given everything so far, can I stop?

Why This Method?

A fixed-horizon p-value is only valid if you look once, at a sample size fixed before the data existed. Look after every visitor and the chance of seeing p < 0.05 somewhere along the way is far above 5%. A mixture sequential probability ratio test (mSPRT) replaces the fixed critical value with a boundary that is valid at every sample size simultaneously, so monitoring the experiment continuously does not inflate the false-positive rate. The price is power, and this module states that price in the user's own numbers rather than selling peeking as free.

What This Analysis Covers

  • The always-valid confidence sequence for the treatment effect over time
  • The current decision: stop for a winner, stop for futility, keep running,

or not enough data for any boundary to be reachable

  • The naive fixed-horizon p-value computed alongside, and how many times it

would have crossed 0.05 while the always-valid boundary did not

  • The cost of the guarantee: how much wider the always-valid interval is and

how many observations it needs to match today's fixed-horizon precision

Standard Library

Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {variant, outcome, arrival}. The mSPRT, the confidence sequence and the boundary crossing are implemented directly from the normal mixture — no group-sequential package is used.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Core Analysis Pipeline

Step 1: Confirm the mapping and humanize the user's column names

variant_h <- humanize_semantic("variant", col_map)
  outcome_h <- humanize_semantic("outcome", col_map)
  arrival_h <- humanize_semantic("arrival", col_map)
  needed <- c("variant", "outcome", "arrival")
  missing_keys <- needed[!(needed %in% names(df))]
  if (length(missing_keys) > 0) {
    labels <- c(variant = variant_h, outcome = outcome_h, arrival = arrival_h)
    stop(sprintf(
      paste0("Sequential testing needs three columns mapped: the variant each ",
             "row saw, the outcome, and the order or date the observation ",
             "arrived. Missing: %s."),
      paste(labels[missing_keys], collapse = ", ")))
  }

Step 2: Read the arrival column as dates, or as a numeric order

arr_ch <- trimws(as.character(df$arrival))
  arr_ch[arr_ch == "NA"] <- NA_character_
  n_nonblank_arr <- sum(!is.na(arr_ch) & arr_ch != "")
  parse_dates_vec <- function(ch) {
    ok <- function(dd) sum(!is.na(dd)) >= 0.95 * sum(!is.na(ch) & trimws(ch) != "")
    d <- suppressWarnings(as.Date(ch, format = "%Y-%m-%d"))
    if (!ok(d)) d <- suppressWarnings(as.Date(lubridate::ymd(ch, quiet = TRUE)))
    if (!ok(d)) d <- suppressWarnings(as.Date(lubridate::mdy(ch, quiet = TRUE)))
    if (!ok(d)) d <- suppressWarnings(as.Date(lubridate::dmy(ch, quiet = TRUE)))
    if (ok(d)) d else NULL
  }
  arrival_kind <- ""
  ord_key <- NULL
  if (n_nonblank_arr > 0) {
    dts <- parse_dates_vec(arr_ch)
    if (!is.null(dts)) {
      ord_key <- as.numeric(dts)
      arrival_kind <- "calendar dates"
    } else {
      nv <- suppressWarnings(as.numeric(arr_ch))
      if (sum(!is.na(nv)) >= 0.95 * n_nonblank_arr) {
        ord_key <- nv
        arrival_kind <- "a numeric arrival order"
      }
    }
  }
  if (is.null(ord_key)) {
    stop(sprintf(
      paste0("The arrival column &#x27;%s' could not be read either as dates or as ",
             "a numeric order, so the observations cannot be put in the order ",
             "they arrived. Sequential testing is entirely about that order — ",
             "map a date, a timestamp, or a running sequence number."),
      arrival_h))
  }

Step 3: Read the outcome as a number (95% rule, then yes/no words)

y_raw <- df$outcome
  outcome_read <- ""
  if (is.numeric(y_raw)) {
    y <- as.numeric(y_raw)
    outcome_read <- "already numeric"
  } else {
    ch <- trimws(as.character(y_raw))
    ch[ch == "NA"] <- NA_character_
    nonblank <- !is.na(ch) & ch != ""
    conv <- suppressWarnings(as.numeric(ch))
    if (sum(nonblank) > 0 && sum(!is.na(conv)) >= 0.95 * sum(nonblank)) {
      y <- conv
      outcome_read <- "read as numbers"
    } else {
      lut <- c("yes" = 1, "no" = 0, "true" = 1, "false" = 0, "y" = 1, "n" = 0,
               "t" = 1, "f" = 0, "converted" = 1, "not converted" = 0,
               "1" = 1, "0" = 0, "success" = 1, "failure" = 0)
      conv2 <- unname(lut[tolower(ch)])
      if (sum(nonblank) > 0 && sum(!is.na(conv2)) >= 0.95 * sum(nonblank)) {
        y <- conv2
        outcome_read <- "read as yes/no labels turned into 1 and 0"
      } else {
        stop(sprintf(
          paste0("The outcome column &#x27;%s' is not numeric and is not a yes/no ",
                 "flag — fewer than 95%% of its values could be turned into ",
                 "numbers. Map a numeric outcome(revenue, time, a 0/1 ",
                 "conversion flag) so the two arms can be compared."),
          outcome_h))
      }
    }
  }

Step 4: Keep complete rows only

g_ch <- trimws(as.character(df$variant))
  g_ch[g_ch == "" | g_ch == "NA"] <- NA_character_
  keep <- !is.na(g_ch) & !is.na(y) & !is.na(ord_key)
  n_dropped <- sum(!keep)
  g_k <- g_ch[keep]; y_k <- y[keep]; o_k <- ord_key[keep]
  final_rows <- length(y_k)
  rows_removed <- initial_rows - final_rows
  if (final_rows < 30) {
    stop(sprintf(
      paste0("Only %d complete row(s) remained after dropping rows missing ",
             "%s, %s or %s. Sequential monitoring needs at least 30 ",
             "observations before any boundary means anything."),
      final_rows, variant_h, outcome_h, arrival_h))
  }

Step 5: Resolve the two arms — never chosen after seeing the outcome

tab <- sort(table(g_k), decreasing = TRUE)
  tiny_arms <- names(tab)[tab < MIN_ARM_KEEP]
  if (length(tiny_arms) > 0) {
    keep2 <- !(g_k %in% tiny_arms)
    g_k <- g_k[keep2]; y_k <- y_k[keep2]; o_k <- o_k[keep2]
    final_rows <- length(y_k)
    rows_removed <- initial_rows - final_rows
    tab <- sort(table(g_k), decreasing = TRUE)
  }
  arms <- names(tab)
  if (length(arms) < 2) {
    stop(sprintf(
      paste0("The variant column &#x27;%s' holds only %d group%s with at least %d ",
             "observation%s(%s). A sequential A/B test compares exactly two ",
             "arms."),
      variant_h, length(arms), if (length(arms) == 1) "" else "s",
      MIN_ARM_KEEP, "s", paste(arms, collapse = ", ")))
  }
  if (length(arms) > 2) {
    stop(sprintf(
      paste0("The variant column &#x27;%s' holds %d arms (%s). This analysis will ",
             "not pick two of them for you: the always-valid guarantee covers ",
             "the comparison you specified before looking at the data, and ",
             "choosing which arms to compare after seeing the results is ",
             "exactly what it does not protect against. Filter the file to the ",
             "two arms you want to compare and re-run."),
      variant_h, length(arms), paste(arms, collapse = ", ")))
  }

Control is chosen by NAME, not by outcome — a fixed, stated rule.

ctrl_pat <- "^(control|baseline|a|original|holdout|current|old)$"
  ctrl_hit <- arms[grepl(ctrl_pat, tolower(trimws(arms)))]
  control_arm <- if (length(ctrl_hit) >= 1) ctrl_hit[1] else sort(arms)[1]
  control_rule <- if (length(ctrl_hit) >= 1)
    "its name identifies it as the control" else
      "no arm was named as a control, so the alphabetically first label was used"
  treat_arm <- setdiff(arms, control_arm)[1]

Step 6: Order by arrival (ties keep their original file order)

ordv <- order(o_k, seq_along(o_k))
  g_s <- g_k[ordv]; y_s <- y_k[ordv]; o_s <- o_k[ordv]
  n_tied <- sum(duplicated(o_s))
  n_total <- length(y_s)

  sd_all <- stats::sd(y_s)
  if (!is.finite(sd_all) || sd_all <= 0) {
    stop(sprintf(
      paste0("Every retained row has the same value in the outcome column ",
             "&#x27;%s', so there is no difference between the arms to test and no ",
             "boundary to cross."),
      outcome_h))
  }
  tau <- TAU_MULT * sd_all
  t2 <- tau^2

  is_binary <- all(y_s %in% c(0, 1))
  unit <- if (is_binary) "percentage points" else "units of the outcome"

fmt_eff already carries "pp" for a binary outcome, so the spelled-out unit is appended only when it would not be a duplicate.

unit_suffix <- if (is_binary) "" else paste0(" ", unit)
  to_disp <- function(x) if (is_binary) 100 * x else x
  disp_digits <- if (is_binary) 2 else 4
  fmt_eff <- function(x) trimws(paste0(.fmt_n(to_disp(x), disp_digits),
                                       if (is_binary) " pp" else ""))

Step 7: Cumulative statistics after every single observation

is_c <- as.integer(g_s == control_arm)
  is_t <- 1L - is_c
  nA <- cumsum(is_c);        nB <- cumsum(is_t)
  sA <- cumsum(y_s * is_c);  sB <- cumsum(y_s * is_t)
  qA <- cumsum(y_s^2 * is_c); qB <- cumsum(y_s^2 * is_t)
  mA <- ifelse(nA > 0, sA / nA, NA_real_)
  mB <- ifelse(nB > 0, sB / nB, NA_real_)
  vA <- ifelse(nA > 1, (qA - nA * mA^2) / (nA - 1), NA_real_)
  vB <- ifelse(nB > 1, (qB - nB * mB^2) / (nB - 1), NA_real_)
  vA[!is.na(vA) & vA < 0] <- 0
  vB[!is.na(vB) & vB < 0] <- 0
  V <- vA / nA + vB / nB
  delta <- mB - mA

  valid <- nA >= MIN_PER_ARM & nB >= MIN_PER_ARM &
    is.finite(V) & V > 0 & is.finite(delta)
  idx <- which(valid)
  if (length(idx) == 0) {
    stop(sprintf(
      paste0("Neither arm of &#x27;%s' reached %d observations with any variation ",
             "in &#x27;%s', so the experiment cannot be monitored yet. %s has %d ",
             "row(s) and %s has %d."),
      variant_h, MIN_PER_ARM, outcome_h,
      control_arm, sum(is_c), treat_arm, sum(is_t)))
  }

  h_path <- .av_halfwidth(V[idx], t2, ALPHA)
  d_path <- delta[idx]
  lam_path <- .av_lambda(d_path, V[idx], t2)
  se_path <- sqrt(V[idx])
  dfw <- (V[idx])^2 / ((vA[idx] / nA[idx])^2 / pmax(nA[idx] - 1, 1) +
                         (vB[idx] / nB[idx])^2 / pmax(nB[idx] - 1, 1))
  dfw[!is.finite(dfw) | dfw < 1] <- 1
  t_path <- d_path / se_path
  p_fh_path <- 2 * stats::pt(-abs(t_path), df = dfw)
  fh_hw_path <- stats::qt(1 - ALPHA / 2, df = dfw) * se_path
  p_av_path <- cummin(pmin(1, 1 / lam_path))

  crossed_av <- is.finite(h_path) & abs(d_path) > h_path
  naive_sig <- is.finite(p_fh_path) & p_fh_path < ALPHA

  n_points <- length(idx)
  av_points <- sum(crossed_av)
  naive_points <- sum(naive_sig)
  naive_excursions <- sum(naive_sig & !c(FALSE, head(naive_sig, -1)))
  naive_only <- sum(naive_sig & !crossed_av)
  cross_idx <- if (av_points > 0) which(crossed_av)[1] else NA_integer_
  cross_n <- if (!is.na(cross_idx)) idx[cross_idx] else NA_integer_
  naive_first_idx <- if (naive_points > 0) which(naive_sig)[1] else NA_integer_
  naive_first_n <- if (!is.na(naive_first_idx)) idx[naive_first_idx] else NA_integer_

Step 8: The current decision

last <- length(idx)
  effect <- d_path[last]
  h_final <- h_path[last]
  se_final <- se_path[last]
  fh_hw_final <- fh_hw_path[last]
  p_av <- p_av_path[last]
  p_fh <- p_fh_path[last]
  av_low <- effect - h_final
  av_high <- effect + h_final
  fh_low <- effect - fh_hw_final
  fh_high <- effect + fh_hw_final

  decide <- function(eff, hw) {
    if (!is.finite(hw)) return("insufficient")
    if (abs(eff) > hw) return("winner")
    if (hw > tau) return("insufficient")
    if (abs(eff) + hw <= tau) return("futility")
    "continue"
  }
  decision <- decide(effect, h_final)
  decision_label <- switch(
    decision,
    winner      = "Stop — a winner has separated",
    futility    = "Stop for futility — no effect worth acting on remains possible",
    continue    = "Keep running — no boundary has been crossed yet",
    insufficient = "Not enough data yet — no boundary is reachable at this sample size")

Step 9: The price of the guarantee, in this experiment's own numbers

frac_c <- sum(is_c) / n_total
  frac_c <- min(max(frac_c, 1e-6), 1 - 1e-6)
  vA_f <- vA[idx][last]; vB_f <- vB[idx][last]
  h_at_n <- function(n) {
    Vn <- vA_f / (n * frac_c) + vB_f / (n * (1 - frac_c))
    .av_halfwidth(Vn, t2, ALPHA)
  }
  grid <- unique(round(exp(seq(log(n_total), log(n_total * 5000), length.out = 6000))))
  grid <- grid[is.finite(grid) & grid >= n_total]
  hg <- h_at_n(grid)
  first_at_or_below <- function(target) {
    okv <- which(is.finite(hg) & hg <= target)
    if (length(okv) == 0) NA_real_ else grid[okv[1]]
  }
  n_equiv <- first_at_or_below(fh_hw_final)
  n_needed_futility <- if (decision == "insufficient")
    first_at_or_below(tau) else NA_real_
  width_ratio <- if (is.finite(fh_hw_final) && fh_hw_final > 0)
    h_final / fh_hw_final else NA_real_

Step 10: Sensitivity to the pre-specified alternative

tau_mults <- c(0.05, 0.10, 0.20)
  tau_rows <- lapply(tau_mults, function(mm) {
    tt <- mm * sd_all
    hh <- .av_halfwidth(V[idx], tt^2, ALPHA)
    dd <- decide(effect, hh[last])
    data.frame(
      alternative = paste0(.fmt_n(mm, 2), " x the outcome&#x27;s standard deviation"),
      alternative_size = .fmt_n(to_disp(tt), disp_digits),
      half_width = .fmt_n(to_disp(hh[last]), disp_digits),
      decision = switch(dd, winner = "stop for a winner",
                        futility = "stop for futility",
                        continue = "keep running",
                        insufficient = "not enough data"),
      ever_crossed = if (sum(is.finite(hh) & abs(d_path) > hh) > 0) "yes" else "no",
      stringsAsFactors = FALSE)
  })
  tau_df <- do.call(rbind, tau_rows)
  tau_agreement <- length(unique(tau_df$decision)) == 1

Step 12: One computed paragraph for the JSON answer

contrast_txt <- if (naive_only > 0) {
    paste0("Along the way the fixed-horizon p-value fell below 0.05 at ",
           format(naive_only, big.mark = ","), " analysis point(s) where the ",
           "always-valid boundary was not crossed, in ",
           format(naive_excursions, big.mark = ","),
           " separate excursion(s) — that gap is the false-positive risk ",
           "peeking at a fixed-horizon test would have exposed you to here.")
  } else if (naive_points > 0) {
    paste0("The fixed-horizon p-value was below 0.05 at ",
           format(naive_points, big.mark = ","),
           " analysis point(s), all of them points where the always-valid ",
           "boundary had also been crossed, so on this data the two tests ",
           "never disagreed.")
  } else {
    paste0("The fixed-horizon p-value never fell below 0.05 at any of the ",
           format(n_points, big.mark = ","),
           " analysis points, so on this data peeking would not have misled ",
           "you — which is luck, not protection.")
  }
  json_output <- list(
    answer = paste0(
      decision_label, ". After ", format(n_total, big.mark = ","),
      " observations the effect of ", treat_arm, " over ", control_arm,
      " is ", fmt_eff(effect), " with an always-valid 95% confidence sequence of ",
      fmt_eff(av_low), " to ", fmt_eff(av_high), unit_suffix,
      " (always-valid p-value ", .fmt_p(p_av), "). ", contrast_txt,
      " The guarantee costs power: the always-valid interval is ",
      if (is.finite(width_ratio)) paste0(.fmt_n(width_ratio, 2), " times") else "wider than",
      " the fixed-horizon interval at this sample size."
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "boundary_path", "pvalue_paths",
        "decision", "peek_comparison", "power_cost", "tau_sensitivity", "methods"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  list(
    initial_rows = initial_rows, final_rows = final_rows,
    rows_removed = rows_removed, n_dropped = n_dropped, n_tied = n_tied,
    variant_h = variant_h, outcome_h = outcome_h, arrival_h = arrival_h,
    arrival_kind = arrival_kind, outcome_read = outcome_read,
    tiny_arms = tiny_arms,
    control_arm = control_arm, treat_arm = treat_arm, control_rule = control_rule,
    n_control = sum(is_c), n_treat = sum(is_t), n_total = n_total,
    is_binary = is_binary, unit = unit, unit_suffix = unit_suffix,
    to_disp = to_disp, fmt_eff = fmt_eff,
    sd_all = sd_all, tau = tau, alpha = ALPHA, tau_mult = TAU_MULT,
    min_per_arm = MIN_PER_ARM,
    decision = decision, decision_label = decision_label,
    effect = effect, av_low = av_low, av_high = av_high,
    fh_low = fh_low, fh_high = fh_high,
    h_final = h_final, fh_hw_final = fh_hw_final, se_final = se_final,
    p_av = p_av, p_fh = p_fh,
    cross_idx = cross_idx, cross_n = cross_n,
    naive_first_idx = naive_first_idx, naive_first_n = naive_first_n,
    n_points = n_points, naive_points = naive_points,
    naive_excursions = naive_excursions, av_points = av_points,
    naive_only = naive_only,
    width_ratio = width_ratio, n_equiv = n_equiv,
    n_needed_futility = n_needed_futility, grid_max = max(grid),
    tau_df = tau_df, tau_agreement = tau_agreement,
    boundary_path_df = boundary_path_df, pvalue_paths_df = pvalue_paths_df,
    decision_df = decision_df, peek_df = peek_df, power_df = power_df,
    methods_df = methods_df,
    metrics = metrics, json_output = json_output
  )
}
Your data has more stories to tell. Run any analysis on your own data — validated R modules, interactive reports, AI insights, and PDF export. 500 free credits on signup.
Try Free — No Signup Sign Up Free

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing