Standard Gap Analysis
Executive Summary

Executive Summary

The largest importance-performance gap across 10 attributes.

Responses
43
Attributes Analysed
10
Importance Source
derived
Top Priority
sound oral rulings
Top Priority Score
3.13
Concentrate Here
5
Importance Boundary
0.941
Performance Boundary
7.584
Reclassified Under Alternative Boundary
10
Across 43 responses and 10 attributes, sound oral rulings has the widest gap between how much it matters and how well it is rated (priority score 3.13 standard deviations, importance 0.98 against performance 7.29). At the grand-mean boundaries (importance 0.94, performance 7.58), 5 attributes sit in Concentrate Here: sound oral rulings, sound written rulings, preparation, familiarity with law, demeanor. The weakest case for attention is physical ability. The picture is boundary-sensitive: under a fixed correlation cut of 0.30 on importance and the scale midpoint (5.5) on performance, judicial integrity, demeanor, diligence, case flow management, decisiveness, preparation, familiarity with law, sound oral rulings, sound written rulings and physical ability would change quadrant. Importance here is association, not a proven lever: the map says where the unmet need is, not that closing it will raise the outcome.
What this means

The short answer

Sound oral rulings has the largest gap between importance and performance: it matters most (importance 0.98) but is rated lowest (7.29 out of 10), a gap of 3.13 standard deviations. Five attributes sit in Concentrate Here at the grand-mean boundaries; all 10 would move quadrants under a fixed alternative boundary, so any recommendation depends on which boundary you choose.

The detail

At the data-driven boundaries (importance 0.941, performance 7.584), five attributes occupy Concentrate Here: sound oral rulings, sound written rulings, preparation, familiarity with law, and demeanor. Sound oral rulings leads the priority ranking at 3.13 standard deviations. Under a fixed alternative boundary (importance ≥0.30, performance ≥5.5), all 10 attributes would shift to Keep Up The Good Work, erasing the Concentrate Here quadrant entirely. Importance is derived from correlation with overall worthy of retention; no stated-importance data were supplied, so the analysis cannot check whether respondents themselves would agree on what matters.

What this can't tell you

The picture is boundary-sensitive: the choice of where to cross the axes is a methodological judgment, not a fact in the data. Use the grand-mean boundary to rank within your own list; use the fixed alternative when comparing across surveys or time. The derived importance figures are vulnerable to halo effects and shared variance—attributes rated highly by satisfied respondents inflate all correlations at once.

Overview

Analysis Overview

How 10 attributes are placed on the importance-performance map.

N Observations43
N Attributes10
N Concentrate5
N Reclassified10
What this means

The short answer

This analysis maps 10 judicial performance attributes against how much they matter for retention decisions. Importance is measured by correlation with overall worthy-of-retention ratings; performance is the current average rating each attribute receives. The map reveals which gaps are largest and which quadrant each attribute falls into, but the boundaries dividing the quadrants are judgment calls rather than fixed facts.

The detail

The analysis places each of 10 attributes on two axes: derived importance (correlation with overall retention worthiness) and current performance (mean rating on a 1–10 scale). The reference lines sit at the grand mean of each axis—importance 0.94 and performance 7.58—creating four quadrants. This data-driven boundary splits the 43 responses into relative winners and losers. The four quadrants are: Concentrate Here (important and weak), Keep Up The Good Work (important and strong), Low Priority (unimportant and weak), and Possible Overkill (unimportant and strong). All 10 attributes were usable; no responses were dropped and no ratings were imputed.

What this can't tell you

The boundaries are relative to this attribute set. Adding or removing attributes shifts the means and can move other attributes across quadrant lines. Neither the grand-mean boundary nor the fixed alternative (0.30 importance, 5.5 performance) is objectively correct; the choice determines the work order. Derived importance is vulnerable to halo effects—respondents satisfied overall rate everything higher, inflating all correlations simultaneously. No stated-importance column was mapped, so the analysis cannot check whether respondents themselves would call these attributes important. Comparing stated and derived importance would be the strongest design; consider mapping a direct importance rating per attribute on the next survey."

Data Preparation

Data Quality

Rows and attributes used, exclusions, imputation, and the rating scale.

Initial Rows43
Final Rows43
Rows Removed0
N Attributes10
N Imputed0
What this means

The short answer

All 43 responses and all 10 attribute columns were usable. No ratings were missing, so no imputation was needed. The rating scale runs 1–10, making the midpoint 5.5 for alternative boundary calculations.

The detail

43 rows loaded; all 43 were usable. All mapped attribute columns were usable with no missing values and zero median-imputed ratings. The 10 attributes carried onto the map are: judicial integrity, demeanor, diligence, case flow management, decisiveness, preparation, familiarity with law, sound oral rulings, sound written rulings, and physical ability. Ratings span 1 to 10.

What this can't tell you

This export is a snapshot at one point in time and does not reveal whether ratings or their relationships change over time. No demographic or judge-level variables are present, so the analysis cannot examine whether the importance-performance gaps differ by judge type, caseload, or tenure."

Visualization

Importance vs Performance Map

Each attribute placed by derived importance against current performance, split into four quadrants.

What this means

The short answer

Five attributes cluster in the upper-left (important but underperforming): sound oral rulings, sound written rulings, preparation, familiarity with law, and demeanor. Three sit in the lower-right (well-rated but less important): diligence, judicial integrity, and physical ability. No attributes occupy the upper-right (important and strong). Attributes near the boundary lines are not meaningfully different from those just across them.

The detail

The 10 attributes are plotted with derived importance (0.907–0.982) on the horizontal axis and current performance (7.293–8.021) on the vertical axis. The reference lines cross at importance 0.9411 and performance 7.5842. Three quadrants are populated. Concentrate Here contains sound oral rulings (0.982, 7.293), sound written rulings (0.968, 7.384), preparation (0.95, 7.467), familiarity with law (0.942, 7.488), and demeanor (0.944, 7.516). Low Priority holds case flow management (0.927, 7.479) and decisiveness (0.925, 7.565). Possible Overkill holds diligence (0.93, 7.693), judicial integrity (0.937, 8.021), and physical ability (0.907, 7.935). Keep Up The Good Work is empty.

What this can't tell you

The axes measure association, not causation. A high correlation between an attribute and overall retention worthiness does not prove that improving that attribute will raise retention. The map is relative to this attribute set; adding or removing attributes shifts both the reference lines and the positions of other attributes. Attributes near the boundary lines (e.g., demeanor at importance 0.944, just above the 0.9411 line) are not meaningfully separated from those just across it and should not be treated as discretely different."

Visualization

Priority Ranking

Attributes ranked by the size of the importance-performance gap.

What this means

The short answer

Sound oral rulings leads by a wide margin with a priority score of 3.13 standard deviations—meaning it matters much more than its performance rating warrants. Sound written rulings follows at 2.07. Only five of the ten attributes have positive scores; the other five are rated better than their importance would predict, with physical ability trailing at −3.092.

The detail

Priority score measures how far above average an attribute sits on importance minus how far above average it sits on performance, both in standard deviations. Sound oral rulings (3.13) and sound written rulings (2.07) lead; preparation (0.922), familiarity with law (0.435), and demeanor (0.41) round out the positive five. Case flow management (−0.194), decisiveness (−0.659), diligence (−0.979), judicial integrity (−2.05), and physical ability (−3.092) are rated above what their importance predicts. The score is relative to this attribute set: it re-ranks if you add or drop attributes.

What this can't tell you

The score subtracts standard deviations rather than raw scale points because derived importance is a correlation and performance is a rating—they are on different scales. This means the ranking captures relative priority within your attribute set but does not measure absolute performance shortfall. A different set of attributes would produce different scores and rankings.

Data Table

Attribute Detail

Importance, performance, gap, and quadrant for each of 10 attributes.

AttributeImportancePerformancePriority ScoreQuadrant
sound oral rulings0.9827.2933.133Concentrate Here
sound written rulings0.9687.3842.074Concentrate Here
preparation0.957.4670.922Concentrate Here
familiarity with law0.9427.4880.435Concentrate Here
demeanor0.9447.5160.41Concentrate Here
case flow management0.9277.479-0.194Low Priority
decisiveness0.9257.565-0.659Low Priority
diligence0.937.693-0.979Possible Overkill
judicial integrity0.9378.021-2.05Possible Overkill
physical ability0.9077.935-3.092Possible Overkill
What this means

The short answer

Sound oral rulings stands apart with the highest importance (0.982) and lowest performance (7.293), yielding the largest priority gap (3.133). Five attributes sit in Concentrate Here; none sit in Keep Up The Good Work. Physical ability has the lowest importance (0.907) and lowest priority score (−3.092).

The detail

Importance ranges from 0.907 (physical ability) to 0.982 (sound oral rulings). Performance ranges from 7.293 (sound oral rulings) to 8.021 (judicial integrity). Priority scores range from 3.133 (sound oral rulings) to −3.092 (physical ability). The five Concentrate Here attributes are sound oral rulings, sound written rulings, preparation, familiarity with law, and demeanor. Two sit in Low Priority (case flow management, decisiveness), three in Possible Overkill (diligence, judicial integrity, physical ability), and zero in Keep Up The Good Work. The largest gap within Concentrate Here is between sound oral rulings (3.133) and demeanor (0.41).

What this can't tell you

Two attributes can share a quadrant while sitting at opposite ends of it. For example, demeanor (priority score 0.41) and sound oral rulings (priority score 3.133) are both in Concentrate Here, but sound oral rulings needs much more urgent attention. The quadrant assignment alone does not convey this difference; the priority score must be read alongside it."

Data Table

Stated vs Derived Importance

Where the two ways of measuring importance agree, and where they contradict each other.

AttributeStated ImportanceDerived ImportanceDerived EvidenceStated RankDerived RankRank Shift
sound oral rulings0.982p < 0.0011
sound written rulings0.968p < 0.0012
preparation0.95p < 0.0013
familiarity with law0.942p < 0.0015
demeanor0.944p < 0.0014
case flow management0.927p < 0.0018
decisiveness0.925p < 0.0019
diligence0.93p < 0.0017
judicial integrity0.937p < 0.0016
physical ability0.907p < 0.00110
What this means

The short answer

Only derived importance is available; no respondents stated which attributes matter. This means the analysis cannot verify whether respondents themselves would agree these attributes are important, and derived importance is vulnerable to halo bias—a respondent satisfied overall rates every attribute higher, inflating all correlations at once.

The detail

No stated-importance columns were mapped, so importance is inferred entirely from each attribute's correlation with overall worthy of retention. All 10 attributes show significant correlations (p < 0.001). Sound oral rulings ranks first with derived importance 0.982; physical ability ranks last at 0.907. The attributes correlate with each other up to 0.99, so each attribute's bivariate importance shares credit with the others. In a joint standardized regression, sound oral rulings' unique contribution is 0.53, against its bivariate correlation of 0.98—the 0.45-point gap is the credit it shares with other attributes. This distance illustrates the halo effect: high overall satisfaction inflates all attribute ratings together, making it hard to isolate which attribute independently drives retention decisions.

What this can't tell you

The analysis cannot determine which attributes respondents themselves would call important, because no stated-importance rating was collected. Derived importance and stated importance often diverge—respondents may state one attribute matters while their retention decisions track another. Mapping a direct importance rating per attribute on the next survey would allow the two sources to be compared, revealing whether the current rankings reflect what respondents truly prioritize or whether halo bias is distorting the picture."

Data Table

Boundary Sensitivity

How the quadrant picture changes under the other boundary convention.

AttributeQuadrant Grand MeanQuadrant AlternativeChanged
sound oral rulingsConcentrate HereKeep Up The Good Workyes
sound written rulingsConcentrate HereKeep Up The Good Workyes
preparationConcentrate HereKeep Up The Good Workyes
familiarity with lawConcentrate HereKeep Up The Good Workyes
demeanorConcentrate HereKeep Up The Good Workyes
case flow managementLow PriorityKeep Up The Good Workyes
decisivenessLow PriorityKeep Up The Good Workyes
diligencePossible OverkillKeep Up The Good Workyes
judicial integrityPossible OverkillKeep Up The Good Workyes
physical abilityPossible OverkillKeep Up The Good Workyes
What this means

The short answer

All ten attributes would move quadrants if you shift the boundary from the data-driven grand means (importance 0.94, performance 7.58) to a fixed alternative (importance ≥0.30, performance ≥5.5). Sound oral rulings, sound written rulings, preparation, familiarity with law, and demeanor would move from Concentrate Here to Keep Up The Good Work—erasing the need to fix anything.

The detail

Under the grand-mean boundary, five attributes sit in Concentrate Here, two in Low Priority, and three in Possible Overkill. Under the fixed alternative, all ten move to Keep Up The Good Work. Sound oral rulings, sound written rulings, preparation, familiarity with law, and demeanor shift from Concentrate Here. Case flow management and decisiveness shift from Low Priority. Diligence, judicial integrity, and physical ability shift from Possible Overkill. Every recommendation that rests on one of these ten attributes depends on which boundary you choose, not on the data itself.

What this can't tell you

Neither boundary is correct—they answer different questions. Use the grand-mean boundary (0.94, 7.58) when you want to rank within your own attribute set and identify relative winners and losers. Use the fixed alternative (0.30, 5.5) when you need the map to mean the same thing across different surveys or over time. The choice is methodological, not empirical.

Data Table

Prioritized Action List

The recommended next step for each quadrant, ordered by urgency.

QuadrantN AttributesAttributesAction
Concentrate Here5sound oral rulings, sound written rulings, preparation, familiarity with law, demeanorFix these first. They matter more than average and are rated below average, so this is where the largest unmet need sits.
Keep Up The Good Work0NoneProtect these. They matter and are already rated well — they are the strengths worth defending rather than improving further.
Low Priority2case flow management, decisivenessLeave these alone for now. They are rated below average but also matter less than average, so weak scores here cost comparatively little.
Possible Overkill3diligence, judicial integrity, physical abilityConsider easing off. These are rated well but matter less than average, so effort spent here may be buying little.
What this means

The short answer

Start with the five Concentrate Here attributes—sound oral rulings, sound written rulings, preparation, familiarity with law, and demeanor—beginning with sound oral rulings, which has the widest gap (3.13 standard deviations). Before committing budget, check the boundary-sensitivity card: all ten attributes sit close enough to a boundary that a different convention would move them, so the action plan depends on which boundary you choose.

The detail

Concentrate Here (5 attributes): Fix these first. They matter more than average and are rated below average. Sound oral rulings carries the widest gap. Keep Up The Good Work (0 attributes): Protect these—they matter and are already rated well. Low Priority (2 attributes: case flow management, decisiveness): Leave alone. They are rated below average but matter less than average, so weak scores here cost comparatively little. Possible Overkill (3 attributes: diligence, judicial integrity, physical ability): Consider easing off. These are rated well but matter less than average, so effort here may buy little.

What this can't tell you

Every recommendation rests on association between attribute ratings and overall retention worthiness, not proven causation. Treat each as a hypothesis worth testing rather than a proven lever. All ten attributes would shift quadrants under the alternative boundary, so verify your boundary choice before acting on any single recommendation.

Rate this report Was this the answer you needed?
The exact source that produced this report — yours to keep, read, and re-run.
Download PDF
How this was computed method · R source · citation
The code that did it

Importance-Performance Gap Analysis

What matters most that we do worst? Classic Importance-Performance Analysis (IPA): every attribute is placed on a four-quadrant map by how important it is against how well it is currently rated, ranked by the gap between the two, and turned into a prioritized action list.

Why This Method?

Ranking attributes by performance alone tells you where you are weak but not whether anyone cares. Ranking by importance alone tells you what matters but not where you are failing. Crossing the two is the whole point: the attributes that are important AND under-performing are the ones worth money, and the map makes that visible in one picture.

Two ways to measure importance — and they disagree

STATED importance is what respondents said mattered. DERIVED importance is inferred from how strongly each attribute tracks an overall satisfaction score. They routinely disagree, and the disagreement is informative rather than an error: people under-report what actually drives them and over-report what they think they should care about. This module detects which inputs it was given, computes BOTH when both are available, and shows where they contradict each other instead of quietly picking one.

What This Analysis Covers

  • Attribute-level importance and performance on a four-quadrant map
  • The priority ranking by importance-performance gap
  • Stated versus derived importance, and where they disagree
  • How the quadrant picture changes under the other boundary convention
  • A per-quadrant action list

Sibling tool

standard_key_drivers answers "what drives satisfaction?" — it ranks drivers by derived importance and reports model fit. This tool answers "what should we fix first?" — it takes importance as given (stated or derived), crosses it with current performance, and produces a prioritized gap list. Same vocabulary, different question.

Standard Library

Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {performance_1..N, importance_1..N, overall}. All narrative is derived from the user's own column names and computed values. Derived importance is CORRELATIONAL — it reports association, never proven causation.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Longest shared leading run of whole words.

n_lead <- 0
  repeat {
    k <- n_lead + 1
    if (any(sapply(parts, length) <= k)) break
    w <- sapply(parts, function(p) p[k])
    if (length(unique(w)) != 1) break
    n_lead <- k
  }

Longest shared trailing run of whole words.

n_tail <- 0
  repeat {
    k <- n_tail + 1
    if (any(sapply(parts, length) <= n_lead + k)) break
    w <- sapply(parts, function(p) p[length(p) - k + 1])
    if (length(unique(w)) != 1) break
    n_tail <- k
  }
  if (n_lead == 0 && n_tail == 0) return(labels)

  out <- sapply(parts, function(p) {
    keep <- p[(n_lead + 1):(length(p) - n_tail)]
    trimws(paste(keep, collapse = " "))
  }, USE.NAMES = FALSE)

Refuse the strip if it empties or collides any label.

if (any(nchar(out) == 0) || anyDuplicated(out) > 0) return(labels)
  out
}

Step 1: Row accounting + semantic column discovery

initial_rows <- nrow(df)
  perf_cols <- grep("^performance_[0-9]+$", names(df), value = TRUE)
  perf_cols <- perf_cols[order(as.integer(sub("^performance_", "", perf_cols)))]
  imp_cols <- grep("^importance_[0-9]+$", names(df), value = TRUE)
  imp_cols <- imp_cols[order(as.integer(sub("^importance_", "", imp_cols)))]
  has_overall_col <- "overall" %in% names(df)

  if (length(perf_cols) == 0) {
    stop(paste0("column_mapping must map at least ", MIN_ATTRS,
                " performance columns(performance_1, performance_2, ...) — ",
                "one per attribute you rate."))
  }

  perf_names_raw <- humanize_semantic(perf_cols, col_map)
  imp_names_raw  <- if (length(imp_cols) > 0) humanize_semantic(imp_cols, col_map) else character(0)
  overall_name   <- if (has_overall_col) humanize_semantic("overall", col_map) else NA_character_

Attribute labels come from the PERFORMANCE columns, with any wording shared by all of them removed ("Satisfaction: Price" -> "Price").

attr_labels_all <- strip_common_affix(perf_names_raw)
  affix_stripped <- !identical(attr_labels_all, perf_names_raw)
  names(attr_labels_all) <- perf_cols

Stated importance is paired to performance BY POSITION: importance_1 describes the same attribute as performance_1. A partial or mismatched set cannot be paired safely, so it is refused rather than guessed.

paired_stated <- length(imp_cols) > 0 && length(imp_cols) == length(perf_cols) &&
    identical(sub("^importance_", "", imp_cols), sub("^performance_", "", perf_cols))
  if (length(imp_cols) > 0 && !paired_stated) {
    stop(paste0(
      "Stated importance must be mapped one-for-one with performance: ",
      n_things(length(imp_cols), "importance column"), " were mapped(",
      paste(imp_names_raw, collapse = ", "), ") against ",
      n_things(length(perf_cols), "performance column"), " (",
      paste(perf_names_raw, collapse = ", "),
      "). Map the same attributes in the same order, or map none and supply ",
      "an overall satisfaction column instead."))
  }

Step 2: Coerce every mapped rating to numeric (95% rule)

A column that will not convert cleanly, or that has no usable values, is dropped and reported rather than silently coerced to NA.

dropped_cols <- character(0)
  coerce <- function(dd, cc) {
    v <- dd[[cc]]
    if (is.numeric(v)) return(v)
    conv <- suppressWarnings(as.numeric(as.character(v)))
    n_orig <- sum(!is.na(v) & as.character(v) != "")
    if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) conv else NULL
  }
  bad_perf <- character(0)
  for (pc in perf_cols) {
    conv <- coerce(df, pc)
    if (is.null(conv)) { bad_perf <- c(bad_perf, pc); next }
    df[[pc]] <- conv
  }
  bad_imp <- character(0)
  if (paired_stated) {
    for (ic in imp_cols) {
      conv <- coerce(df, ic)
      if (is.null(conv)) { bad_imp <- c(bad_imp, ic); next }
      df[[ic]] <- conv
    }
  }
  overall_usable <- FALSE
  if (has_overall_col) {
    conv <- coerce(df, "overall")
    if (!is.null(conv)) { df$overall <- conv; overall_usable <- TRUE }
  }

Step 3: Drop constant / all-missing attributes — and their partner

An attribute is only usable if BOTH sides that will be plotted survive.

for (i in seq_along(perf_cols)) {
    pc <- perf_cols[i]
    if (pc %in% bad_perf) next
    v <- df[[pc]]
    if (all(is.na(v))) { bad_perf <- c(bad_perf, pc); next }
    if (isTRUE(stats::var(v, na.rm = TRUE) == 0) ||
        is.na(stats::var(v, na.rm = TRUE))) {
      bad_perf <- c(bad_perf, pc)
    }
  }
  drop_idx <- which(perf_cols %in% bad_perf)
  if (paired_stated) drop_idx <- union(drop_idx, which(imp_cols %in% bad_imp))
  keep_idx <- setdiff(seq_along(perf_cols), drop_idx)
  if (length(drop_idx) > 0) {
    dropped_cols <- unname(attr_labels_all[perf_cols[drop_idx]])
  }
  perf_use <- perf_cols[keep_idx]
  imp_use  <- if (paired_stated) imp_cols[keep_idx] else character(0)
  attr_labels <- unname(attr_labels_all[perf_use])

  if (length(perf_use) < MIN_ATTRS) {
    kept_note <- if (length(perf_use) > 0)
      paste0(" The attributes read were: ", paste(attr_labels, collapse = ", "), ".")
    else
      paste0(" The columns mapped were: ", paste(perf_names_raw, collapse = ", "), ".")
    stop(sprintf(paste0(
      "Importance-Performance Analysis needs at least %d usable attributes; ",
      "only %d of the %d mapped performance columns survived cleaning%s.%s ",
      "A four-quadrant map of fewer than %d attributes is not meaningful."),
      MIN_ATTRS, length(perf_use), length(perf_cols),
      if (length(dropped_cols) > 0)
        paste0(" (excluded as constant, empty, or non-numeric: ",
               paste(dropped_cols, collapse = ", "), ")") else "",
      kept_note, MIN_ATTRS))
  }

Step 4: Which importance sources do we actually have?

have_stated  <- paired_stated && length(imp_use) == length(perf_use)
  have_derived <- overall_usable && sum(!is.na(df$overall)) >= MIN_ROWS
  if (!have_stated && !have_derived) {
    stop(paste0(
      "Importance-Performance Analysis needs importance from somewhere. ",
      "Either map an importance column for each attribute(",
      paste(attr_labels, collapse = ", "),
      "), or map an overall satisfaction column so importance can be ",
      "derived from how each attribute tracks it."))
  }

Step 5: Rows — derived importance needs a usable overall value

if (have_derived) df <- df[!is.na(df$overall), , drop = FALSE]
  keep_row <- rowSums(!is.na(df[, perf_use, drop = FALSE])) > 0
  df <- df[keep_row, , drop = FALSE]
  final_rows <- nrow(df)
  rows_removed <- initial_rows - final_rows
  if (final_rows < MIN_ROWS) {
    stop(sprintf(paste0(
      "Only %d usable responses remain after cleaning%s — at least %d are ",
      "required before an importance-performance map means anything. ",
      "Attributes read: %s."),
      final_rows,
      if (have_derived) sprintf(" (rows with no &#x27;%s' value are dropped)", overall_name) else "",
      MIN_ROWS, paste(attr_labels, collapse = ", ")))
  }

Step 6: Impute remaining missing ratings with each column's median

n_imputed <- 0L
  for (cc in c(perf_use, imp_use)) {
    v <- df[[cc]]
    miss <- is.na(v)
    if (any(miss)) {
      med <- stats::median(v, na.rm = TRUE)
      if (!is.na(med)) { v[miss] <- med; n_imputed <- n_imputed + sum(miss) }
      df[[cc]] <- v
    }
  }

Step 7: Rating scale — needed for the midpoint boundary convention

rating_vals <- unlist(df[, c(perf_use, imp_use), drop = FALSE], use.names = FALSE)
  rating_vals <- rating_vals[is.finite(rating_vals)]
  vmin <- min(rating_vals); vmax <- max(rating_vals)
  scale_detected <- vmin >= 0 && vmax <= 10
  if (scale_detected) {
    scale_min <- if (vmin < 1) 0 else 1
    scale_max <- if (vmax <= 5) 5 else if (vmax <= 7) 7 else 10
  } else {
    scale_min <- vmin; scale_max <- vmax
  }
  scale_mid <- (scale_min + scale_max) / 2

Step 8: Performance — the mean rating per attribute

perf_mean <- sapply(perf_use, function(pc) mean(df[[pc]], na.rm = TRUE))
  perf_mean <- as.numeric(perf_mean)

Step 9: Stated importance — the mean importance rating per attribute

stated_imp <- if (have_stated) {
    as.numeric(sapply(imp_use, function(ic) mean(df[[ic]], na.rm = TRUE)))
  } else rep(NA_real_, length(perf_use))

Step 10: Derived importance — how each attribute tracks the overall

score. Bivariate Pearson correlation is the standard derived-importance statistic; a joint standardized regression is also fitted purely to disclose how much of that association is shared rather than unique.

derived_imp <- rep(NA_real_, length(perf_use))
  derived_p   <- rep(NA_real_, length(perf_use))
  max_attr_cor <- NA_real_
  top_beta <- NA_real_
  if (have_derived) {
    for (i in seq_along(perf_use)) {
      ct <- tryCatch(stats::cor.test(df[[perf_use[i]]], df$overall),
                     error = function(e) NULL)
      if (!is.null(ct) && is.finite(ct$estimate)) {
        derived_imp[i] <- as.numeric(ct$estimate)
        derived_p[i]   <- as.numeric(ct$p.value)
      } else {
        derived_imp[i] <- 0
      }
    }
    if (length(perf_use) >= 2) {
      cm <- suppressWarnings(stats::cor(df[, perf_use, drop = FALSE],
                                        use = "pairwise.complete.obs"))
      cm[!is.finite(cm)] <- 0
      diag(cm) <- 0
      max_attr_cor <- max(abs(cm))
    }
    zfit <- tryCatch({
      zd <- as.data.frame(scale(df[, c("overall", perf_use), drop = FALSE]))
      stats::lm(overall ~ ., data = zd)
    }, error = function(e) NULL)
    if (!is.null(zfit)) {
      bb <- stats::coef(zfit)[perf_use]
      ok <- which(!is.na(derived_imp))
      if (length(ok) > 0) {
        lead <- ok[order(-abs(derived_imp[ok]))][1]
        top_beta <- unname(bb[lead])
      }
    }
  }

Step 11: Which importance source drives the primary map?

Stated is preferred when present because it shares the performance scale, which makes the raw gap directly readable. The choice is stated in the prose and the other source is reported alongside it — never silently dropped.

req_source <- tolower(as.character(params$importance_source %||% "auto"))
  imp_source <- if (req_source == "derived" && have_derived) "derived"
                else if (req_source == "stated" && have_stated) "stated"
                else if (have_stated) "stated" else "derived"
  importance <- if (imp_source == "stated") stated_imp else derived_imp
  imp_source_h <- if (imp_source == "stated")
    "stated importance(the average importance rating respondents gave)"
  else
    paste0("derived importance(each attribute&#x27;s correlation with ", overall_name, ")")

Short form, for sentences that already carry their own parentheses.

imp_source_short <- if (imp_source == "stated") "stated importance" else "derived importance"

Step 12: Boundaries — the methodological choice that moves the map

Primary: the data-driven grand mean of each axis (the cross-hair sits at the average attribute). Alternative: the scale midpoint, which is fixed and independent of this attribute set. For derived importance the scale midpoint does not exist, so the alternative is the conventional moderate-association cut of 0.30.

imp_boundary  <- mean(importance, na.rm = TRUE)
  perf_boundary <- mean(perf_mean, na.rm = TRUE)
  perf_boundary_alt <- scale_mid
  imp_boundary_alt  <- if (imp_source == "stated") scale_mid else DERIVED_ALT_CUT
  alt_label <- if (imp_source == "stated") {
    paste0("the scale midpoint(", fmt_n(scale_mid, 1), " on a ",
           fmt_n(scale_min, 0), "-to-", fmt_n(scale_max, 0), " scale)")
  } else {
    paste0("a fixed correlation cut of ", fmt_n(DERIVED_ALT_CUT, 2),
           " on importance and the scale midpoint(", fmt_n(scale_mid, 1),
           ") on performance")
  }

  quad_of <- function(imp, perf, bi, bp) {
    ifelse(imp >= bi & perf <  bp, "Concentrate Here",
    ifelse(imp >= bi & perf >= bp, "Keep Up The Good Work",
    ifelse(imp <  bi & perf <  bp, "Low Priority",
                                   "Possible Overkill")))
  }
  quadrant     <- quad_of(importance, perf_mean, imp_boundary, perf_boundary)
  quadrant_alt <- quad_of(importance, perf_mean, imp_boundary_alt, perf_boundary_alt)
  reclassified <- quadrant != quadrant_alt
  n_reclassified <- sum(reclassified)
  reclassified_names <- attr_labels[reclassified]

Step 13: The gap. Two measures, both computed.

priority_score standardizes each axis ACROSS THE ATTRIBUTE SET, so it works whether importance is a rating or a correlation. raw_gap is the plain importance-minus-performance difference in scale points, and is only defined when both axes are on the same rating scale.

z_of <- function(v) {
    s <- stats::sd(v, na.rm = TRUE)
    if (!is.finite(s) || s == 0) rep(0, length(v))
    else (v - mean(v, na.rm = TRUE)) / s
  }
  priority_score <- z_of(importance) - z_of(perf_mean)
  imp_range <- diff(range(importance, na.rm = TRUE))
  scale_span <- if (is.finite(scale_max - scale_min) && (scale_max - scale_min) > 0)
    scale_max - scale_min else NA_real_
  imp_tied <- if (imp_source == "stated" && is.finite(scale_span))
    imp_range < 0.05 * scale_span else FALSE
  raw_gap_available <- imp_source == "stated"
  raw_gap <- if (raw_gap_available) stated_imp - perf_mean else rep(NA_real_, length(perf_use))

Step 14: Master attribute frame, ranked by the gap

attrs_df <- data.frame(
    attribute      = attr_labels,
    importance     = round(importance, 3),
    performance    = round(perf_mean, 3),
    priority_score = round(priority_score, 3),
    raw_gap        = round(raw_gap, 3),
    quadrant       = quadrant,
    quadrant_alt   = quadrant_alt,
    stated_importance  = round(stated_imp, 3),
    derived_importance = round(derived_imp, 3),
    derived_p          = derived_p,
    stringsAsFactors = FALSE
  )
  attrs_df <- attrs_df[order(-attrs_df$priority_score, attrs_df$attribute), , drop = FALSE]
  rownames(attrs_df) <- NULL

  top_priority_name  <- attrs_df$attribute[1]
  top_priority_score <- attrs_df$priority_score[1]
  bottom_name        <- attrs_df$attribute[nrow(attrs_df)]
  concentrate_names  <- attrs_df$attribute[attrs_df$quadrant == "Concentrate Here"]
  n_concentrate      <- length(concentrate_names)

Step 15: Stated versus derived — ranks, shifts, and quadrant flips

have_both <- have_stated && have_derived
  rank_rho <- NA_real_
  n_source_flips <- 0L
  source_flip_names <- character(0)
  max_shift_name <- NA_character_
  max_shift <- NA_real_
  stated_rank  <- rep(NA_real_, nrow(attrs_df))
  derived_rank <- rep(NA_real_, nrow(attrs_df))
  if (have_stated) stated_rank  <- rank(-attrs_df$stated_importance, ties.method = "min")
  if (have_derived) derived_rank <- rank(-attrs_df$derived_importance, ties.method = "min")
  rank_shift <- if (have_both) abs(stated_rank - derived_rank) else rep(NA_real_, nrow(attrs_df))
  if (have_both) {
    rank_rho <- suppressWarnings(stats::cor(attrs_df$stated_importance,
                                            attrs_df$derived_importance,
                                            method = "spearman"))
    if (!is.finite(rank_rho)) rank_rho <- NA_real_
    ord <- order(-rank_shift, attrs_df$attribute)
    max_shift_name <- attrs_df$attribute[ord[1]]
    max_shift <- rank_shift[ord[1]]
    q_stated  <- quad_of(attrs_df$stated_importance, attrs_df$performance,
                         mean(attrs_df$stated_importance), perf_boundary)
    q_derived <- quad_of(attrs_df$derived_importance, attrs_df$performance,
                         mean(attrs_df$derived_importance), perf_boundary)
    flips <- q_stated != q_derived
    n_source_flips <- sum(flips)
    source_flip_names <- attrs_df$attribute[flips]
  }

Derived importance is an estimate, so it carries its own evidence: an attribute whose correlation is not distinguishable from zero must not be read as "unimportant with confidence".

derived_evidence <- if (have_derived) {
    sapply(seq_len(nrow(attrs_df)), function(i) {
      pv <- attrs_df$derived_p[i]
      if (is.na(pv)) return("not available")
      paste0(fmt_p(pv), if (pv < 0.05) "" else " (not distinguishable from zero)")
    })
  } else rep(NA_character_, nrow(attrs_df))

  importance_sources_df <- data.frame(
    attribute          = attrs_df$attribute,
    stated_importance  = attrs_df$stated_importance,
    derived_importance = attrs_df$derived_importance,
    derived_evidence   = derived_evidence,
    stated_rank        = stated_rank,
    derived_rank       = derived_rank,
    rank_shift         = rank_shift,
    stringsAsFactors   = FALSE
  )
  boundary_sensitivity_df <- data.frame(
    attribute             = attrs_df$attribute,
    quadrant_grand_mean   = attrs_df$quadrant,
    quadrant_alternative  = attrs_df$quadrant_alt,
    changed               = ifelse(attrs_df$quadrant != attrs_df$quadrant_alt,
                                   "yes", "no"),
    stringsAsFactors = FALSE
  )
  action_plan_df <- do.call(rbind, lapply(QUADS, function(q) {
    members <- attrs_df$attribute[attrs_df$quadrant == q]
    data.frame(
      quadrant     = q,
      n_attributes = length(members),
      attributes   = if (length(members) > 0) paste(members, collapse = ", ") else "None",
      action       = unname(QUAD_ACTIONS[q]),
      stringsAsFactors = FALSE
    )
  }))
  rownames(action_plan_df) <- NULL

Step 17: KPI metrics

metrics <- list(
    `Responses`            = final_rows,
    `Attributes Analysed`  = length(perf_use),
    `Importance Source`    = if (imp_source == "stated") "stated" else "derived",
    `Top Priority`         = top_priority_name,
    `Top Priority Score`   = round(top_priority_score, 2),
    `Concentrate Here`     = as.integer(n_concentrate),
    `Importance Boundary`  = round(imp_boundary, 3),
    `Performance Boundary` = round(perf_boundary, 3),
    `Reclassified Under Alternative Boundary` = as.integer(n_reclassified)
  )

Step 18: json_output machine channel

concentrate_clause <- if (n_concentrate > 0) {
    paste0(n_things(n_concentrate, "attribute"), " ",
           vform(n_concentrate, "falls", "fall"), " in Concentrate Here(",
           paste(concentrate_names, collapse = ", "), ")")
  } else {
    "No attribute falls in Concentrate Here"
  }
  disagree_clause <- if (have_both) {
    paste0(" Stated and derived importance rank the attributes differently ",
           "(Spearman rank correlation ", fmt_n(rank_rho, 2), "); ",
           max_shift_name, " moves the most, by ",
           n_things(as.integer(max_shift), "rank position"),
           ", and ", n_things(as.integer(n_source_flips), "attribute"), " ",
           vform(n_source_flips, "changes", "change"),
           " quadrant depending on which importance source is used.")
  } else if (imp_source == "derived") {
    paste0(" Importance was derived from ", overall_name,
           " because no stated-importance columns were mapped.")
  } else {
    " Importance is as stated by respondents; no overall satisfaction column was mapped to cross-check it."
  }
  json_output <- list(
    answer = paste0(
      "Importance-Performance Analysis of ",
      n_things(length(perf_use), "attribute"), " across ",
      n_things(final_rows, "response"), ", using ", imp_source_h, ". ",
      top_priority_name, " has the largest importance-performance gap ",
      "(priority score ", fmt_n(top_priority_score, 2),
      " standard deviations), and ", bottom_name, " the smallest. ",
      concentrate_clause, " at the grand-mean boundaries(importance ",
      fmt_n(imp_boundary, 2), ", performance ", fmt_n(perf_boundary, 2), "); ",
      "under ", alt_label, ", ",
      n_things(as.integer(n_reclassified), "attribute"), " ",
      vform(n_reclassified, "changes", "change"), " quadrant.",
      disagree_clause,
      " These are associations between attribute ratings and priority, not ",
      "proof that improving an attribute will move the outcome."
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "quadrant_map", "gap_ranking",
        "attribute_table", "importance_comparison", "boundary_sensitivity",
        "action_plan"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  list(
    initial_rows = initial_rows, final_rows = final_rows, rows_removed = rows_removed,
    n_imputed = n_imputed,
    attr_labels = attr_labels, perf_cols = perf_use, dropped_cols = dropped_cols,
    affix_stripped = affix_stripped, perf_names_raw = perf_names_raw,
    have_stated = have_stated, have_derived = have_derived, have_both = have_both,
    imp_source = imp_source, imp_source_h = imp_source_h,
    imp_source_short = imp_source_short,
    overall_name = overall_name,
    scale_detected = scale_detected, scale_min = scale_min,
    scale_max = scale_max, scale_mid = scale_mid,
    imp_boundary = imp_boundary, perf_boundary = perf_boundary,
    imp_boundary_alt = imp_boundary_alt, perf_boundary_alt = perf_boundary_alt,
    alt_label = alt_label,
    attrs_df = attrs_df,
    quadrant_points_df = quadrant_points_df, gap_ranking_df = gap_ranking_df,
    attribute_detail_df = attribute_detail_df,
    importance_sources_df = importance_sources_df,
    boundary_sensitivity_df = boundary_sensitivity_df,
    action_plan_df = action_plan_df,
    n_reclassified = n_reclassified, reclassified_names = reclassified_names,
    n_source_flips = n_source_flips, source_flip_names = source_flip_names,
    rank_rho = rank_rho, max_shift_name = max_shift_name, max_shift = max_shift,
    top_priority_name = top_priority_name, top_priority_score = top_priority_score,
    bottom_name = bottom_name, concentrate_names = concentrate_names,
    n_concentrate = n_concentrate, imp_tied = imp_tied,
    raw_gap_available = raw_gap_available,
    max_attr_cor = max_attr_cor, top_beta = top_beta,
    metrics = metrics, json_output = json_output
  )
}
Your data has more stories to tell.Run any analysis on your own data — validated R modules, interactive reports, AI insights, and PDF export. 500 free credits on signup.
Try Free — No SignupSign Up Free

Cite this analysis

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing