# Survival analysis: agent brief

> **Audience: an agent deciding whether to run this analysis on a dataset.** Not marketing
> copy. Terse, complete, and honest about failure. If you are asked "should I run survival
> analysis here", you should be able to answer from this file alone, **including saying no**.
>
> Tool: `standard_survival` · free page: `/free/standard_survival` ·
> worked example: `/worked-examples/survival` · lesson: https://youtu.be/9jOl9KRj3lA

---

## 1. What it answers

**How long until an event happens, when some subjects have not had it yet.**

Churn, failure, relapse, conversion, attrition, tenure. The distinguishing feature is not
the subject matter, it is the *incompleteness*: at the moment you analyse, some subjects
are still running and their final duration is unknown.

**Questions it is mistaken for:**

| Actually asked | Right tool |
|---|---|
| *What share churned this month?* | a rate, not this. A rate is a snapshot and says nothing about duration. |
| *Which customers will churn next?* | classification / a risk score. Survival gives timing and group differences, not per-account predictions. |
| *Did group A convert more than group B?* (everyone observed) | `standard_proportions` |
| *Is the mean of A different from B?* (all durations complete) | `standard_group_comparison` |
| *How many are left at each funnel step?* | `standard_funnel` |
| *Did the launch change the metric?* | `standard_event_impact` |

## 2. When it applies, and when it does not

**Apply it when: some subjects have not had the event yet.** That is the whole
discriminating condition. Those subjects are *censored*: their duration is a **lower
bound**, not a missing value.

**Do not apply it when:**

- **Every duration is complete.** If every subject has had the event, there is no censoring
  and an ordinary comparison of means or medians is correct and simpler.
- **You only have aggregates.** A monthly churn rate or a count per period cannot be turned
  back into per-subject durations. This analysis needs rows.
- **The "event" can happen more than once per subject** and you care about all occurrences.
  Standard Kaplan-Meier assumes one terminal event per subject.
- **Subjects left for reasons unrelated to the event and that relates to their risk.**
  Censoring must be *non-informative*. If accounts vanish from your export precisely because
  they were about to churn, the estimate is biased and no method here fixes it.

## 3. What the data must look like

`column_mapping` requires **all three**:

| key | type | meaning |
|---|---|---|
| `time` | **numeric** | duration until the event, or until the observation ended |
| `event` | 1/0 or yes/no | did the event happen |
| `group` | categorical | what to compare; a constant column degrades to one overall curve |

Constraints: min 30 rows, max 100,000, max null rate 0.5.

**The transformation users almost always have to do first, and the most likely reason an
upload fails to attach:** an export carries *dates* (`started_at`, `churned_at`), and the
tool needs a *number*. There is no date-arithmetic step. Compute it:

- event happened → `time = end_date - start_date`
- still running → `time = today - start_date`, and `event = 0`

**The mistake that silently produces a wrong answer:** filtering out the rows that have not
had the event. It reads as cleaning up blanks. It is the single thing this analysis exists
to prevent, and it biases the result in one direction only (see §5).

## 4. What it returns, and how to read each piece

| Output | Read it as | The trap |
|---|---|---|
| **Kaplan-Meier curve** | share still event-free at time *t*, counting each subject for as long as it was observed | the tail is thin; late portions rest on few subjects and are wide |
| **Median survival + CI** | the time by which half the subjects have had the event | "not reached" is a legitimate result: over half were still running at the end of observation. It is not an error |
| **Log-rank test** | do the whole curves differ, using every event and every censored subject | tests the *curves*, not one summary point. Curves that cross can differ while the test is unimpressive |
| **Hazard ratio + CI** | relative risk of the event at any given moment, one group against another | assumes proportional hazards, a roughly constant ratio over time. If curves cross, the single ratio is a fiction |
| **Restricted mean (RMST)** | average survival **within a stated window** | meaningless without its horizon. Always report the window with the number |

## 5. How it fails

**The naive alternative and its direction.** Averaging only subjects that had the event
discards every censored subject. Censored durations are lower bounds, so excluding them can
**only shorten** the estimate. The error never points the other way. On the worked example
it is a 44.26% understatement of the median.

**Failure modes that yield a plausible wrong answer rather than an error:**

- **Informative censoring** (§2). Nothing in the output reveals it. It must be reasoned about.
- **Crossing curves.** The log-rank test loses power and the hazard ratio stops meaning
  anything, while both still return numbers. Look at the curves.
- **Tiny risk sets in the tail.** Late survival estimates can swing on one or two subjects.
- **A group column with too many levels.** Each curve is estimated on a fraction of the data;
  precision collapses quietly.
- **Ties.** Many events on the same day change the Cox estimate depending on the tie-handling
  convention. The worked example has 21 tied event days and uses Efron (R's default).

## 6. Verified numbers you may cite

From `/worked-examples/survival`, agreed three ways (the notebook's R run, R's canonical
`survival` functions called directly, and an independent from-scratch Python
reimplementation of Kaplan-Meier, log-rank and Efron-ties Cox). All agree to every digit.

| Quantity | Value |
|---|---|
| accounts / events / censored | 320 / 210 / 110 (34.38% censored) |
| naive median (event-only subjects) | 301 days |
| Kaplan-Meier median | **540 days** (95% CI 451–637) |
| naive understatement | **44.26%** |
| survival at 1 / 2 / 3 years | 0.6137 / 0.3671 / 0.2305 |
| median, onboarding completed | 703 days (599–800) |
| median, not completed | 257 days (190–374) |
| log-rank p | 1.834e-12 |
| hazard ratio, not-completed vs completed | 2.6578 (2.0044–3.5242) |
| restricted mean, 730-day horizon | 470.03 days |

**Do not say:** that the 44.26% is a claim about *means*. It compares **medians** (301 against
540). A mean lifetime for the whole book is not identifiable without a stated horizon, which
is what the restricted mean is for. Do not quote a mean lifetime without its window.

## 7. Where everything is

| | |
|---|---|
| Tool | `standard_survival` |
| Free page | https://mcpanalytics.ai/free/standard_survival |
| Worked example | https://mcpanalytics.ai/worked-examples/survival |
| Dataset | `/worked-examples/files/survival_accounts.csv` (320 accounts) |
| Notebook / source | `/worked-examples/files/survival.html` · `survival.Rmd` |
| Independent checks | `survival_check.R` · `survival_check.py` |
| Lesson (7 min) | https://youtu.be/9jOl9KRj3lA |
| Elementary (2.5 min) | https://youtu.be/Q_Lw8KoNtGc |
| Validation record | `lattice/v2/refs/LAT-2369-rmd-lesson-ladder/worked-example-survival/VALIDATION.md` |
| Our tool's own run on this data | https://api.mcpanalytics.ai/rpt/rpt_b0bVwESJlm1ACQB9Lw26RA (job `mcp_standard_survival_55670a6caba8`, 2026-08-23, success) |

## 8. Routing shortcut

```
Does every subject have a completed duration?
├── yes → NOT this. Compare means/medians directly.
└── no, some are still running
    ├── do you have one row per subject with a duration and an event flag?
    │   ├── no, only aggregates → cannot run. Ask for row-level data.
    │   └── yes → run standard_survival
    └── could subjects have dropped out BECAUSE they were about to have the event?
        └── yes → run it, but flag the estimate as biased (informative censoring)
```
