Executive Summary
Top driver and the priority opportunity for Overall Satisfaction across 400 responses.
Price Value is the strongest driver of Overall Satisfaction, holding 40.4% of the combined driver importance, ahead of Support Quality at 26.2%. The clearest opportunity for action is Wait Time Rating: it is strongly associated with Overall Satisfaction yet rated only 20.3 out of 100 today, placing it in the Improve / Priority quadrant. Lifting Wait Time Rating is where effort is most likely to move Overall Satisfaction. Across 400 responses, importance reflects association with the outcome, not proof of causation—treat these findings as starting points for testing, not established levers.
Analysis Overview
What Key Driver Analysis measures for Overall Satisfaction across 5 drivers.
Key Driver Analysis identifies which attributes most closely track Overall Satisfaction by measuring importance—how tightly a driver correlates with the outcome—against performance, its current rating. Crossing these two dimensions produces an action quadrant: drivers that are important yet under-performing are the priorities, while important, well-rated drivers are strengths to protect. Across 400 responses, the five drivers analyzed explain 56.5% of the variation in Overall Satisfaction. This analysis is correlational; it shows which attributes are associated with satisfaction, not which ones would move it if changed. The quadrant placement guides action more reliably than importance alone, since it reflects both the driver's association with the outcome and how it is rated today.
Data Quality
Outcome, drivers used and dropped, imputation, and model fit.
Data quality is strong: all 400 responses yielded a usable Overall Satisfaction value, with no rows removed. Five drivers were carried into the model after mapping; numeric drivers with missing values were imputed using the column median. All mapped driver columns were usable. Driver ratings were detected as a fixed rating scale, so performance was rescaled against that scale's range to produce a 0-to-100 metric. The resulting model explains 56.5% of the variation in Overall Satisfaction, setting the context for the importance shares. This R-squared indicates the five drivers capture more than half the factors that move satisfaction, though substantial variation remains unexplained.
What Drives the Outcome
Each driver's share of the combined importance, ranked.
Relative importance shows each driver's share of the combined standardized effect on Overall Satisfaction. Price Value leads at 40.4%, followed by Support Quality at 26.2%, making these two drivers account for roughly two-thirds of the tracked influence. Wait Time Rating contributes 16.2%, Ease of Use 13.6%, and Brand Perception the weakest link at 3.6%. A high relative importance score means the driver moves closely with Overall Satisfaction; it does not prove that changing the driver will move Overall Satisfaction. The distribution reveals that satisfaction is most tightly bound to perceived value and support quality, with wait time, usability, and brand perception playing smaller but measurable roles.
Driver Detail
Importance, correlation, performance, and quadrant for each driver of Overall Satisfaction.
| Driver | Relative Importance | Correlation | Std Beta | Mean Performance | Quadrant |
|---|---|---|---|---|---|
| Price Value | 40.4 | 0.565 | 0.565 | 74.6 | Maintain |
| Support Quality | 26.2 | 0.387 | 0.366 | 79.7 | Maintain |
| Wait Time Rating | 16.2 | 0.266 | 0.226 | 20.3 | Improve / Priority |
| Ease of Use | 13.6 | 0.217 | 0.191 | 49 | Monitor |
| Brand Perception | 3.6 | 0.05 | 0.051 | 86.8 | Reduce effort / Possible over-invest |
Each driver shows its relative importance, plain correlation with Overall Satisfaction, standardized effect, current performance (0-100), and action quadrant. Price Value has the highest correlation at 0.565 and performance of 74.6, placing it in Maintain. Support Quality follows with a correlation of 0.387 and performance of 79.7, also Maintain. Wait Time Rating has a correlation of 0.266 but the lowest performance at 20.3, making it the sole Improve / Priority driver. Ease of Use (0.217 correlation, 49 performance) sits in Monitor, while Brand Perception (0.05 correlation, 86.8 performance) is flagged as a possible over-investment. The quadrant, not importance alone, is the guide to action.
Importance vs Performance
Action quadrant: relative importance against current performance for each driver of Overall Satisfaction.
The importance-performance matrix plots all five drivers by relative importance (horizontal) against current performance (vertical), split by median lines into four quadrants. Wait Time Rating stands alone in the upper-left Improve / Priority region: it has meaningful importance (16.2%) but is severely under-performing at 20.3. Price Value and Support Quality occupy the upper-right Maintain quadrant, both important and well-rated. Ease of Use sits in the lower-left Monitor quadrant, weak on both dimensions. Brand Perception is in the lower-right Reduce effort region—highly rated (86.8) but only weakly associated with satisfaction (3.6% importance). Read the quadrant placement, not individual points, and remember it reflects association, not proven causation.
Action Summary
Recommended next step for each importance-performance quadrant.
| Quadrant | Drivers | Action |
|---|---|---|
| Maintain | Price Value, Support Quality | Protect these strengths. They are strongly associated with the outcome and already rated well. |
| Improve / Priority | Wait Time Rating | Act here first. These are strongly associated with the outcome but rated low today, so gains here are most likely to move the outcome. |
| Monitor | Ease of Use | Low priority. Weakly associated with the outcome and rated low; keep an eye on them but do not over-invest. |
| Reduce effort / Possible over-invest | Brand Perception | Possible over-investment. Rated highly but only weakly associated with the outcome, so extra effort here may not pay off. |
Actions are grouped by quadrant. Maintain: Price Value and Support Quality are strongly associated with satisfaction and already rated well (74.6 and 79.7 respectively); protect these strengths. Improve / Priority: Wait Time Rating is strongly associated with satisfaction but rated only 20.3; act here first, as gains offer the most headroom to move the outcome. Monitor: Ease of Use is weakly associated and rated low (49); keep a light watch but do not over-invest. Reduce effort: Brand Perception is rated highly (86.8) but only weakly associated (3.6% importance); consider easing effort here since extra investment may not pay off. Every recommendation rests on association with Overall Satisfaction, so treat these as places to test, not proven levers.
Methodology
Statistical methodology and diagnostics for Key Driver Analysis — What Moves Satisfaction
Statistical Method
Standard-library analysis: which of the attributes you measure most move an outcome you care about — satisfaction, NPS, retention, or spend? Key Driver Analysis combines statistical IMPORTANCE (how tightly each attribute tracks the outcome, via standardized regression effects and correlations) with current PERFORMANCE (how each attribute is rated today) and places every driver on an action quadrant: Maintain, Improve / Priority, Monitor, or Reduce effort. It tells you not just what matters, but where to invest first. Works on any dataset: map a numeric outcome and up to 10 numeric driver ratings.
- The outcome is numeric (satisfaction, NPS, spend, or a score)
- Drivers are numeric ratings or measures
- Driver effects on the outcome are approximately linear
- Association, not causation — importance describes co-movement, not the effect of an intervention
- Highly correlated drivers share importance approximately (flagged in the narrative, not hidden)
- Importance and performance are relative to the drivers you map and the responses you provide
Analysis Code
Complete R source code for this analysis
Key Driver Analysis — What Moves Satisfaction
The marquee CX / survey analysis: of the attributes you measure, which ones most move an outcome you care about (satisfaction, NPS, spend, retention)? Key Driver Analysis combines two lenses — statistical IMPORTANCE (how tightly each attribute tracks the outcome) and current PERFORMANCE (how well each attribute is rated today) — into an action quadrant that tells you where to invest first.
Why This Method?
Ranking drivers by raw correlation alone tells you what matters but not where you are weak. Ranking by performance alone tells you where you are weak but not whether it matters. Plotting importance against performance resolves both at once: the drivers that are important AND under-performing are the priorities; important-and-strong drivers are strengths to protect.
What This Analysis Covers
- Relative importance of each driver (standardized effect + correlation)
- Current performance of each driver on a 0-100 scale
- The importance-vs-performance action quadrant
- A per-quadrant action summary
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {outcome, driver_1..driver_N}. All narrative is derived from the user's own column names and computed values. This is a CORRELATIONAL analysis — it reports association, never proven causation.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Step 1: Row accounting + semantic column discovery
initial_rows <- nrow(df)
if (!"outcome" %in% names(df)) {
stop("column_mapping must map an 'outcome' column (the numeric outcome to explain, e.g. satisfaction, NPS, or spend)")
}
driver_cols <- grep("^driver_[0-9]+$", names(df), value = TRUE)
driver_cols <- driver_cols[order(as.integer(sub("^driver_", "", driver_cols)))]
if (length(driver_cols) == 0) {
stop("column_mapping must map at least one driver column(driver_1)")
}
outcome_name <- humanize_semantic("outcome", col_map)
driver_names <- setNames(humanize_semantic(driver_cols, col_map), driver_cols)Step 2: Coerce outcome to numeric; drop rows with a missing outcome
df$outcome <- suppressWarnings(as.numeric(df$outcome))
df <- df[!is.na(df$outcome), , drop = FALSE]
if (nrow(df) < 10) {
stop(sprintf(
"Only %d rows have a usable numeric value in the outcome column '%s'. At least 10 are required for Key Driver Analysis.",
nrow(df), outcome_name))
}Step 3: Coerce each driver to numeric (95%% rule); impute NA with median
dropped_drivers <- character(0)
for (dc in driver_cols) {
v <- df[[dc]]
if (!is.numeric(v)) {
conv <- suppressWarnings(as.numeric(as.character(v)))
n_orig <- sum(!is.na(v) & as.character(v) != "")
if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) {
df[[dc]] <- conv
} else {
dropped_drivers <- c(dropped_drivers, dc); next
}
}
v <- df[[dc]]
med <- median(v, na.rm = TRUE)
if (is.na(med)) { dropped_drivers <- c(dropped_drivers, dc); next }
v[is.na(v)] <- med
df[[dc]] <- v
}Step 4: Drop zero-variance (constant) drivers — report them
for (dc in setdiff(driver_cols, dropped_drivers)) {
v <- df[[dc]]
if (isTRUE(var(v, na.rm = TRUE) == 0) || is.na(var(v, na.rm = TRUE))) {
dropped_drivers <- c(dropped_drivers, dc)
}
}
model_drivers <- setdiff(driver_cols, dropped_drivers)
if (length(model_drivers) == 0) {
stop("No usable driver columns remained after cleaning(all were constant, empty, or non-numeric).")
}
df_clean <- df[, c("outcome", model_drivers), drop = FALSE]
final_rows <- nrow(df_clean)
rows_removed <- initial_rows - final_rowsStep 5: Guard — need clearly more rows than drivers (p < n)
while (length(model_drivers) >= final_rows - 2 && length(model_drivers) > 1) {
drop_dc <- model_drivers[length(model_drivers)]
dropped_drivers <- c(dropped_drivers, drop_dc)
model_drivers <- model_drivers[-length(model_drivers)]
df_clean <- df_clean[, c("outcome", model_drivers), drop = FALSE]
}Step 6: Importance — standardized regression coefficients + correlations
Standardize outcome + drivers, fit OLS on the z-scores. |standardized beta| is each driver's independent effect on a common scale. Relative importance = each driver's share of the total |beta|, times 100 (a practical relative-weights proxy). Perfectly collinear drivers are aliased by lm (NA beta) — they carry no independent share, so their |beta| is treated as 0 and a collinearity note is raised.
zdf <- as.data.frame(scale(df_clean[, c("outcome", model_drivers), drop = FALSE]))
z_model <- lm(outcome ~ ., data = zdf)
z_coef <- coef(z_model)
std_betas_raw <- z_coef[model_drivers] # named by semantic; NA if aliased
aliased_any <- any(is.na(std_betas_raw))
abs_beta <- abs(std_betas_raw)
abs_beta[is.na(abs_beta)] <- 0
total_beta <- sum(abs_beta)
rel_importance <- if (total_beta > 0) 100 * abs_beta / total_beta else rep(0, length(abs_beta))
rel_importance <- round(as.numeric(rel_importance), 1)
r_squared <- summary(z_model)$r.squared
if (is.na(r_squared)) r_squared <- 0
correlations <- sapply(model_drivers, function(dc) {
suppressWarnings(cor(df_clean[[dc]], df_clean$outcome, use = "complete.obs"))
})
correlations[is.na(correlations)] <- 0Collinearity check — exact aliasing or any driver pair above 0.9 |r|.
max_pair_cor <- 0
if (length(model_drivers) >= 2) {
dm <- suppressWarnings(cor(df_clean[, model_drivers, drop = FALSE],
use = "pairwise.complete.obs"))
dm[!is.finite(dm)] <- 0
diag(dm) <- 0
max_pair_cor <- max(abs(dm))
}
collinear <- aliased_any || (max_pair_cor > 0.9)Step 7: Performance — mean rating normalized to 0-100
mean_ratings <- sapply(model_drivers, function(dc) mean(df_clean[[dc]], na.rm = TRUE))
all_vals <- unlist(df_clean[, model_drivers], use.names = FALSE)
vmax <- max(all_vals, na.rm = TRUE)
vmin <- min(all_vals, na.rm = TRUE)
scale_detected <- is.finite(vmax) && vmax <= 10 && vmin >= 0
if (scale_detected) {
scale_min <- if (vmin < 1) 0 else 1
scale_max <- if (vmax <= 5) 5 else if (vmax <= 7) 7 else 10
performance <- 100 * (mean_ratings - scale_min) / (scale_max - scale_min)
} else {
performance <- sapply(model_drivers, function(dc) {
lo <- min(df_clean[[dc]], na.rm = TRUE)
hi <- max(df_clean[[dc]], na.rm = TRUE)
if (hi > lo) 100 * (mean(df_clean[[dc]], na.rm = TRUE) - lo) / (hi - lo) else 50
})
}
performance <- round(pmin(100, pmax(0, as.numeric(performance))), 1)Step 8: Quadrants — median split of importance x performance
imp_median <- median(rel_importance)
perf_median <- median(performance)
high_imp <- rel_importance >= imp_median
high_perf <- performance >= perf_median
quadrant <- ifelse(high_imp & high_perf, "Maintain",
ifelse(high_imp & !high_perf, "Improve / Priority",
ifelse(!high_imp & !high_perf, "Monitor",
"Reduce effort / Possible over-invest")))
drivers_df <- data.frame(
semantic = model_drivers,
driver = unname(driver_names[model_drivers]),
relative_importance = rel_importance,
correlation = round(as.numeric(correlations), 3),
std_beta = round(as.numeric(ifelse(is.na(std_betas_raw), 0, std_betas_raw)), 3),
mean_rating = round(as.numeric(mean_ratings), 2),
mean_performance = performance,
quadrant = quadrant,
stringsAsFactors = FALSE
)
drivers_df <- drivers_df[order(-drivers_df$relative_importance,
-abs(drivers_df$correlation)), , drop = FALSE]
rownames(drivers_df) <- NULLStep 10: Headline drivers
top_driver_name <- drivers_df$driver[1]
top_rel_importance <- drivers_df$relative_importance[1]
priority_rows <- drivers_df[drivers_df$quadrant == "Improve / Priority", , drop = FALSE]
priority_driver_name <- if (nrow(priority_rows) > 0) priority_rows$driver[1] else NA_character_
priority_driver_perf <- if (nrow(priority_rows) > 0) priority_rows$mean_performance[1] else NA_real_
n_priority <- nrow(priority_rows)Step 11: KPI metrics
metrics <- list(
`Responses` = final_rows,
`Drivers Analysed` = length(model_drivers),
`Top Driver` = top_driver_name,
`Top Relative Importance` = round(top_rel_importance, 1),
`Model R Squared` = round(r_squared, 3),
`Priority Drivers` = n_priority
)Step 12: json_output machine channel
priority_clause <- if (!is.na(priority_driver_name)) {
paste0(priority_driver_name, " is the clearest priority — important yet rated ",
round(priority_driver_perf, 0), " out of 100.")
} else {
"No driver falls in the Improve / Priority quadrant."
}
json_output <- list(
answer = paste0(
"Key Driver Analysis of ", outcome_name, " across ",
format(final_rows, big.mark = ","), " responses on ",
n_things(length(model_drivers), "driver"), ": ", top_driver_name,
" is the most influential, holding about ", round(top_rel_importance, 0),
"% of the combined driver importance. ", priority_clause,
" Model R-squared is ", round(r_squared, 3),
". These are associations, not proof that a driver changes the outcome."
),
cards = lapply(
c("tldr", "overview", "preprocessing", "importance_chart",
"driver_table", "priority_matrix", "action_summary"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
initial_rows = initial_rows, final_rows = final_rows, rows_removed = rows_removed,
outcome_name = outcome_name, driver_names = driver_names,
model_drivers = model_drivers, dropped_drivers = dropped_drivers,
df_clean = df_clean, r_squared = r_squared,
scale_detected = scale_detected, collinear = collinear,
drivers_df = drivers_df, importance_df = importance_df,
driver_details_df = driver_details_df,
importance_performance_df = importance_performance_df,
quadrant_summary_df = quadrant_summary_df,
top_driver_name = top_driver_name, top_rel_importance = top_rel_importance,
priority_driver_name = priority_driver_name,
priority_driver_perf = priority_driver_perf, n_priority = n_priority,
metrics = metrics, json_output = json_output
)
}