Executive Summary
Top-box picture across 20 item(s) on a 6-point scale.
The short answer
A2 leads at 68.3% top-2-box and C4 trails at 10.6%—a 57.7 percentage-point gap so wide that the items measure fundamentally different constructs and cannot be collapsed into a single score.
The detail
2,436 respondents answered 20 items on a 6-point scale (codes 1, 2, 3, 4, 5, 6). A2 scores highest with 31.5% at "6" and only 1.7% at "1"; C4 scores lowest at 10.6% top-2-box with a median answer of "2". The gap of 57.7 percentage points between the highest and lowest top-2-box shares confirms the items are not interchangeable. No group column was mapped, so these are whole-sample figures. Reading the mean of the response codes as a score assumes the steps between response options are equally spaced, which an ordinal scale does not guarantee.
What this can't tell you
The analysis cannot estimate a single overall personality score without assuming equal spacing between scale points, which the ordinal structure does not justify. A finer-grained item-by-item interpretation is more defensible.
Analysis Overview
Ordinal analysis of 20 survey item(s) across 2,436 respondents on a 6-point scale.
This analysis summarizes 2,436 respondents' answers to 20 personality statements on a 6-point scale. Because the scale is ordinal—the ranking of options is real, but the spacing between them is not—each item is reported using its full distribution, median category, and top-box shares, which rely only on ordering. The mean code is shown for reference but does not assume equal spacing. This approach preserves what the data actually supplies: the direction and concentration of responses, not assumed distances between points.
Data Quality
Which mapped columns became scale items, and what was excluded.
The short answer
All 2,436 respondents and all 20 items loaded cleanly with zero blanks, so every item's percentages are directly comparable and no rows were removed.
The detail
2,436 rows loaded; 2,436 respondents answered at least one analysed item. All mapped columns were usable as scale items. 0 individual answers were blank and were skipped; each item's percentages use that item's own answered count as the denominator. Every analysed item spans the same lowest and highest option (1 and 6), so top-box and bottom-box shares are directly comparable between items. Columns were classified as numeric codes by the 95 percent coercion rule, and the scale was ordered from the data itself rather than assumed.
What this can't tell you
The data export contains only aggregate response counts, not order-level or respondent-level detail, so it is not possible to cross-check individual response patterns or flag unusual answering sequences. A respondent-level export would allow deeper diagnostics.
The Response Scale
The scale detected from the data, and how often each option was chosen.
| Rank | Level | Responses | Share |
|---|---|---|---|
| 1 | 1 | 6069 | 12.5 |
| 2 | 2 | 7767 | 15.9 |
| 3 | 3 | 5729 | 11.8 |
| 4 | 4 | 10038 | 20.6 |
| 5 | 5 | 11278 | 23.1 |
| 6 | 6 | 7839 | 16.1 |
The short answer
Respondents used all six options, with "5" the most popular at 23.1% of all answers across all items, and the numeric codes read in ascending order from lowest to highest.
The detail
A 6-point scale was detected from the data: codes 1, 2, 3, 4, 5, 6, read in ascending numeric order. The distribution across all analysed items is: "1" (6,069 responses, 12.5%), "2" (7,767 responses, 15.9%), "3" (5,729 responses, 11.8%), "4" (10,038 responses, 20.6%), "5" (11,278 responses, 23.1%), and "6" (7,839 responses, 16.1%). Every analysed item uses the same lowest and highest option, so the items can be compared directly on top-box and bottom-box shares.
What this can't tell you
The numeric codes are ordered but their spacing is unknown—the distance from "1" to "2" may not equal the distance from "5" to "6". This is why median and top-box shares (which use only the order) are preferred over means.
Item Summary
Top-box, top-2-box, bottom-box, median and modal category for every item.
| Item | Responses | Top Box PCT | Top2 Box PCT | Bottom Box PCT | Median Level | Mode Level | Mean Code |
|---|---|---|---|---|---|---|---|
| A2 | 2436 | 31.5 | 68.3 | 1.7 | 5 | 5 | 4.8 |
| A4 | 2436 | 40.8 | 64.6 | 4.7 | 5 | 6 | 4.69 |
| A3 | 2436 | 27.1 | 62.8 | 3.5 | 5 | 5 | 4.6 |
| A5 | 2436 | 24.7 | 59.5 | 2.3 | 5 | 5 | 4.54 |
| E4 | 2436 | 25.8 | 59.4 | 5.3 | 5 | 5 | 4.41 |
| C1 | 2436 | 22 | 59.1 | 2.5 | 5 | 5 | 4.53 |
| E5 | 2436 | 21.2 | 55.6 | 3.7 | 5 | 5 | 4.39 |
| C2 | 2436 | 20 | 54.5 | 3.1 | 5 | 5 | 4.37 |
| C3 | 2436 | 17.2 | 50.3 | 3 | 5 | 5 | 4.3 |
| E3 | 2436 | 12.3 | 38.8 | 5.5 | 4 | 4 | 3.98 |
| N2 | 2436 | 10.7 | 29.2 | 11.8 | 4 | 4 | 3.52 |
| C5 | 2436 | 10.3 | 27.5 | 17.9 | 3 | 4 | 3.31 |
| N3 | 2436 | 8.9 | 25 | 17.4 | 3 | 2 | 3.22 |
| E2 | 2436 | 9.6 | 23.4 | 18.9 | 3 | 2 | 3.15 |
| N4 | 2436 | 9.3 | 23 | 16.6 | 3 | 2 | 3.2 |
| E1 | 2436 | 8.7 | 22 | 23.6 | 3 | 2 | 2.98 |
| N5 | 2436 | 8.8 | 20.7 | 23.6 | 3 | 2 | 2.97 |
| N1 | 2436 | 7.2 | 19.5 | 23.1 | 3 | 2 | 2.94 |
| A1 | 2436 | 3 | 10.9 | 33.3 | 2 | 1 | 2.41 |
| C4 | 2436 | 2.4 | 10.6 | 27.7 | 2 | 2 | 2.55 |
Nine of the 20 items have a median of "5", clustering most responses at the upper end of the scale. A2 leads with 68.3% top-2-box (choosing "5" or "6"), while C4 trails at 10.6%. The top 9 items all clear 50% top-2-box; the remaining 11 fall below it, with the largest gaps among the N and C items. A4 shows the highest single top-box at 40.8%, while E3 and below show weaker concentration in the top category. Bottom-box shares are small for most high-ranking items (A2 at 1.7%) but rise sharply for lower-ranking ones: C5 at 17.9%, N3 at 17.4%, and A1 at 33.3%—the largest bottom-box share across all items.
Response Distribution by Item
Diverging stacked bars: each item's answers, negative to the left of zero and positive to the right.
The diverging bars reveal which items push responses decisively toward the top of the scale and which split across the range. A2, A4, A3, and A5 (the A items) show strong rightward skew, with level-5 and level-6 responses dominating. The C and E items in the middle ranks are more evenly spread: C1 has 37.07% at level-5 but only 22.04% at level-6, and E4 shows 33.66% at level-5 and 25.78% at level-6. The N items and the lowest-ranking C and E items (N1, N5, E1, E2, N4, N3, C5, E3) carry substantial left-side shares: N3 shows 17.4% at level-1 and only 8.9% at level-6, creating a split distribution rather than clustering. A1 is the most polarized: 33.3% bottom-box against modest top-box, suggesting divergent respondent views on that statement.
Items Ranked by Top-2-Box
The share choosing one of the two highest options, item by item.
A2 leads decisively at 68.3% top-2-box, followed by A4 (64.6%), A3 (62.8%), and A5 (59.5%)—the A group dominates the top tier. The top 9 items all exceed 50% top-2-box, a natural break point; C3 sits just above at 50.3%. Below that threshold, E3 (38.8%) stands alone, then a steeper drop to N2 (29.2%) and C5 (27.5%). The bottom tier—N3, E2, N4, E1, N5, N1, A1, and C4—clusters between 10.6% and 25%, with C4 and A1 the weakest at 10.6% and 10.9%. The 57.7 percentage-point spread from A2 to C4 shows strong differentiation across items.
Differences Between Groups
Rank-based comparison of each item's answers across the grouping column.
| Item | Test | Statistic | P Value | P Adjusted | Effect Size | Highest Group | Lowest Group | Interpretation |
|---|---|---|---|---|---|---|---|---|
| (no comparison run) | not run | — | n/a | n/a | — | n/a | n/a | No group column was mapped, so the responses were described for the whole sample only. Map a categorical column to compare response distributions between groups with a rank-based test. |
The short answer
No grouping column was mapped, so the entire sample is treated as one group and no between-group comparisons are available.
The detail
The group_comparison card reports: "No group column was mapped, so the responses were described for the whole sample only." To compare response distributions between groups, map a categorical column (e.g. demographic, segment, or treatment indicator) and the analysis will apply a rank-based test. A rank-based test is appropriate for ordinal data because it compares the order of answers and does not assume equal spacing between options.
What this can't tell you
The analysis cannot detect whether any personality statement differs across subgroups. Consider mapping a categorical grouping variable—such as respondent segment, department, or cohort—to run between-group comparisons and identify which items separate the groups most strongly.
Methods & Disclosure
How every figure is computed, and what an ordinal scale can and cannot support.
| Aspect | Detail |
|---|---|
| Detected scale | A 6-point scale was detected from the data: the response codes 1, 2, 3, 4, 5, 6, read in ascending order. |
| Top box | The share of answers at the highest option, "6". |
| Top-2 box | The share of answers at the two highest options, "5" and "6". |
| Bottom box | The share of answers at the lowest option, "1". |
| Median category | The lowest option whose cumulative share of answers reaches 50 percent. It is defined for ties and for an even number of respondents alike, and it uses only the ordering of the options. |
| Modal category | The single most-chosen option. Where two options tie, the lower of the two is reported and the tie is flagged. |
| Mean code | Reading the mean of the response codes as a score assumes the steps between response options are equally spaced, which an ordinal scale does not guarantee: nothing in the data says the distance from "5" to "6" matches the distance from "1" to "2". The median category and the top-box shares use only the ordering, which is all the scale supplies. |
| Diverging bar | Shares at or below the scale midpoint are drawn left of zero and shares above it to the right, with items ordered by their top-2-box share. |
| Group comparison | No group column was mapped, so the responses were described for the whole sample only. Map a categorical column to compare response distributions between groups with a rank-based test. |
| Missing answers | 0 blank answers were skipped; each item's percentages use that item's own answered count as the denominator, so items with different response rates stay comparable. |
The short answer
Top-2-box, top-box, bottom-box, and the median category are computed from the ordering of the scale options alone and need no assumption about spacing. The mean of the codes is shown for reference but assumes equal spacing, which the ordinal scale does not guarantee.
The detail
Detected scale: 6-point numeric codes (1, 2, 3, 4, 5, 6) in ascending order. Top-box is the share choosing "6"; top-2-box is the share choosing "5" or "6"; bottom-box is the share choosing "1". The median category is the lowest option whose cumulative share reaches 50 percent and is well-defined for ties and even sample sizes. The modal category is the single most-chosen option. The mean code assumes equal spacing between consecutive options—a claim the ordinal scale does not make. Diverging bars show shares at or below the scale midpoint left of zero and shares above it to the right, ordered by top-2-box. No group column was mapped. 0 blank answers were skipped; each item's percentages use that item's own answered count (2,436) as the denominator.
What this can't tell you
The analysis does not adjust for survey sampling or weighting, which lie outside the scope of the data export. A respondent-level export with sampling weights would allow precision adjustments if the original survey design calls for them.
Likert / Top-Box Survey Analysis — Ordinal Items Done Right
Analyses a set of Likert-scale survey items as the ordinal data they are, not as if they were continuous measurements. For every item: the full response distribution, top-box / top-2-box / bottom-box percentages, the median response category, the modal category, and a diverging stacked bar across items so the reader can rank them at a glance. An optional grouping column triggers a rank-based comparison between groups (Mann-Whitney for two groups, Kruskal-Wallis for three or more), Holm-corrected across items.
Why This Method?
A Likert response is an ordered label, not a number. "Strongly agree" is above "Agree", but nothing guarantees that the step from "Agree" to "Strongly agree" is the same size as the step from "Neutral" to "Agree". Averaging the codes quietly assumes it is. Top-box shares, the median category, and rank-based tests use only the ordering, which is all the scale actually gives you. The mean of the codes is still reported — it is familiar and often what the reader expects — but always with the equal-spacing assumption stated next to it.
What This Analysis Covers
- The response scale detected from the data (numeric codes or text labels)
- Per item: full distribution, top-box, top-2-box, bottom-box
- Per item: median response category and modal category
- A diverging stacked bar across all items, ranked
- Optional between-group comparison with Mann-Whitney / Kruskal-Wallis
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {item_1..item_N, group}. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Ordinal scale detection
Core Analysis Pipeline
compute_shared <- function(df, params, col_map = list()) {
# === SHARED EXPORTS ===
# initial_rows/final_rows/rows_removed $ row accounting
# item_names $ named character — semantic -> humanized user names
# used_items $ character — semantic item columns actually analysed
# dropped_df $ data.frame(item, reason) — excluded mapped columns
# scale_type $ "numeric codes" or "text labels"
# scale_family $ the matched wording family, or "" for numeric
# common_levels $ character — the response options, lowest first
# k_levels $ integer — number of scale points
# top_level/bottom_level $ the top-box and bottom-box option labels
# scale_consistent $ TRUE when every item spans the same lowest/highest option
# scale_levels_df $ rank, level, responses, share (pooled across items)
# item_summary_df $ per item: n, top/top2/bottom box, median, mode, mean code
# item_dist_df $ item, level, share (SIGNED — the diverging bar)
# item_rank_df $ item, top2_box_pct (ranked, the at-a-glance bar)
# group_tests_df $ per item rank-based test vs the grouping column
# methods_df $ aspect, detail — full disclosure
# has_group / group_h / group_labels / group_test_name
# best_item/worst_item + their top-2-box shares
# mean_caveat $ the equal-spacing sentence used wherever a mean appears
# n_missing_total $ blank answers skipped across items
# metrics / json_output
# === /SHARED EXPORTS ===Step 1: Discover the mapped items
initial_rows <- nrow(df)
item_cols <- grep("^item_[0-9]+$", names(df), value = TRUE)
item_cols <- item_cols[order(as.integer(sub("^item_", "", item_cols)))]
if (length(item_cols) < 1) {
stop("Likert analysis needs at least one survey item mapped(item_1). Map the columns holding the scale answers.")
}
item_names <- setNames(humanize_semantic(item_cols, col_map), item_cols)
all_h <- unname(item_names[item_cols])
if (initial_rows < 10) {
stop(sprintf("Only %d rows of survey responses were found for %s — at least 10 responses are needed to describe a response distribution.",
initial_rows, oxford(all_h)))
}Step 2: Classify every mapped column — numeric codes or text labels
The 95% coercion rule decides; blanks never count against a column.
dropped_items <- character(0)
dropped_reasons <- character(0)
drop_item <- function(ic, reason) {
dropped_items <<- c(dropped_items, ic)
dropped_reasons <<- c(dropped_reasons, reason)
}
raw_vals <- list(); kind <- character(0)
for (ic in item_cols) {
v <- df[[ic]]
ch <- as.character(v)
ch[!is.na(ch) & trimws(ch) == ""] <- NA_character_
non_blank <- !is.na(ch)
if (sum(non_blank) == 0) {
drop_item(ic, "every response is blank"); next
}
conv <- suppressWarnings(as.numeric(ch))
if (sum(!is.na(conv)) >= 0.95 * sum(non_blank)) {
conv[!non_blank] <- NA_real_
raw_vals[[ic]] <- conv; kind[ic] <- "numeric"
} else {
raw_vals[[ic]] <- ch; kind[ic] <- "text"
}
}Step 3: Per-column scale checks — constant, too many options, unorderable
A Likert item is a short ordered list. A column with more than 12 distinct answers is a measurement or an identifier, not a scale, and text answers that no known wording family can order are left out rather than guessed at.
ok_items <- character(0); ok_rank <- list(); ok_family <- character(0)
for (ic in names(raw_vals)) {
v <- raw_vals[[ic]]
obs <- v[!is.na(v)]
lv <- unique(obs)
if (length(lv) < 2) {
drop_item(ic, "every response is the same, so there is no distribution to describe"); next
}
if (length(lv) > 12) {
drop_item(ic, sprintf("%d distinct answers — more than 12 distinct response options, so this reads as a measurement or an identifier rather than a rating scale",
length(lv)))
next
}
if (kind[[ic]] == "numeric") {
ok_items <- c(ok_items, ic)
ok_rank[[ic]] <- sort(lv)
ok_family[ic] <- ""
} else {
ord <- order_text_labels(lv)
if (is.null(ord)) {
drop_item(ic, "the response labels could not be placed in a meaningful order, so the answers cannot be treated as an ordered scale")
next
}
ok_items <- c(ok_items, ic)
ok_rank[[ic]] <- lv[order(ord$rank)]
ok_family[ic] <- ord$family
}
}Step 4: One scale for all items — the majority answer format wins
if (length(ok_items) == 0) {
stop(sprintf("None of the mapped columns(%s) look like a rating scale: each was blank, constant, had more than 12 distinct answers, or used labels that could not be ordered. Map the columns holding the survey scale answers.",
oxford(all_h)))
}
kinds_ok <- kind[ok_items]
n_num <- sum(kinds_ok == "numeric"); n_txt <- sum(kinds_ok == "text")
scale_kind <- if (n_num >= n_txt) "numeric" else "text"
minority <- ok_items[kinds_ok != scale_kind]
for (ic in minority) {
drop_item(ic, sprintf("answers are %s while the other mapped items use %s, and two different answer formats cannot share one scale",
if (scale_kind == "numeric") "text labels" else "numeric codes",
if (scale_kind == "numeric") "numeric codes" else "text labels"))
}
used_items <- setdiff(ok_items, minority)
scale_type <- if (scale_kind == "numeric") "numeric codes" else "text labels"Step 5: Build the common ordered scale from the union of the answers
if (scale_kind == "numeric") {
all_codes <- sort(unique(unlist(lapply(used_items, function(ic) ok_rank[[ic]]))))
common_levels <- fmt_code(all_codes)
level_code <- all_codes
scale_family <- ""
} else {
lab_union <- unique(unlist(lapply(used_items, function(ic) ok_rank[[ic]])))
ord <- order_text_labels(lab_union)
if (is.null(ord)) {
stop(sprintf("The mapped items(%s) use response labels from different wording families, so they cannot be placed on one shared scale. Analyse items that share a response scale together.",
oxford(unname(item_names[used_items]))))
}
common_levels <- lab_union[order(ord$rank)]
level_code <- rep(NA_real_, length(common_levels))
scale_family <- ord$family
}
k_levels <- length(common_levels)
if (k_levels < 2) {
stop(sprintf("Only one distinct response option was found across %s — a rating scale needs at least two.",
oxford(unname(item_names[used_items]))))
}
top_level <- common_levels[k_levels]
bottom_level <- common_levels[1]Step 6: Index every response onto the common scale
as_index <- function(ic) {
v <- raw_vals[[ic]]
if (scale_kind == "numeric") match(fmt_code(v), common_levels)
else match(as.character(v), common_levels)
}
idx_list <- lapply(used_items, as_index)
names(idx_list) <- used_items
answered_any <- Reduce(`|`, lapply(idx_list, function(z) !is.na(z)))
final_rows <- sum(answered_any)
rows_removed <- initial_rows - final_rows
if (final_rows < 10) {
stop(sprintf("Only %d respondents answered at least one of %s — at least 10 are needed to describe a response distribution.",
final_rows, oxford(unname(item_names[used_items]))))
}
n_missing_total <- sum(vapply(idx_list, function(z) sum(is.na(z)), numeric(1)))Step 7: Is every item on the same stretch of the scale?
item_min <- vapply(idx_list, function(z) {
zz <- z[!is.na(z)]; if (length(zz) == 0) NA_real_ else min(zz)
}, numeric(1))
item_max <- vapply(idx_list, function(z) {
zz <- z[!is.na(z)]; if (length(zz) == 0) NA_real_ else max(zz)
}, numeric(1))
scale_consistent <- all(!is.na(item_min)) &&
length(unique(item_min)) == 1 && length(unique(item_max)) == 1Step 8: Per-item ordinal statistics
median category = the lowest option whose cumulative share reaches 50%, which is well defined for ties and even sample sizes alike.
hn <- unname(item_names[used_items])
counts_mat <- matrix(0, nrow = length(used_items), ncol = k_levels,
dimnames = list(hn, common_levels))
n_vec <- integer(length(used_items))
median_lv <- character(length(used_items))
mode_lv <- character(length(used_items))
mode_tied <- logical(length(used_items))
mean_code <- rep(NA_real_, length(used_items))
for (i in seq_along(used_items)) {
z <- idx_list[[used_items[i]]]
zz <- z[!is.na(z)]
n_vec[i] <- length(zz)
cnt <- tabulate(zz, nbins = k_levels)
counts_mat[i, ] <- cnt
cum <- cumsum(cnt) / max(1, sum(cnt))
hit <- which(cum >= 0.5)
median_lv[i] <- if (length(hit) > 0) common_levels[hit[1]] else NA_character_
best <- which(cnt == max(cnt))
mode_lv[i] <- common_levels[best[1]]
mode_tied[i] <- length(best) > 1
mean_code[i] <- if (scale_kind == "numeric") mean(level_code[zz]) else mean(zz)
}
share_of <- function(i, cols) 100 * sum(counts_mat[i, cols]) / max(1, n_vec[i])
top_box <- vapply(seq_along(used_items), function(i) share_of(i, k_levels), numeric(1))
top2_box <- vapply(seq_along(used_items), function(i)
share_of(i, unique(c(max(1, k_levels - 1), k_levels))), numeric(1))
bot_box <- vapply(seq_along(used_items), function(i) share_of(i, 1), numeric(1))
mean_label <- if (scale_kind == "numeric") "Mean code" else "Mean position on the scale"
mean_caveat <- paste0(
"Reading the mean of the ", if (scale_kind == "numeric") "response codes" else "response positions",
" as a score assumes the steps between response options are equally spaced, ",
"which an ordinal scale does not guarantee: nothing in the data says the ",
"distance from \"", common_levels[max(1, k_levels - 1)], "\" to \"", top_level,
"\" matches the distance from \"", bottom_level, "\" to \"", common_levels[min(2, k_levels)],
"\". The median category and the top-box shares use only the ordering, which is all the scale supplies."
)
item_summary_df <- data.frame(
item = hn,
responses = n_vec,
top_box_pct = round(top_box, 1),
top2_box_pct = round(top2_box, 1),
bottom_box_pct = round(bot_box, 1),
median_level = median_lv,
mode_level = mode_lv,
mean_code = round(mean_code, 2),
stringsAsFactors = FALSE
)
ord_items <- order(-item_summary_df$top2_box_pct, item_summary_df$item)
item_summary_df <- item_summary_df[ord_items, , drop = FALSE]
rownames(item_summary_df) <- NULL
best_item <- item_summary_df$item[1]
best_top2 <- item_summary_df$top2_box_pct[1]
worst_item <- item_summary_df$item[nrow(item_summary_df)]
worst_top2 <- item_summary_df$top2_box_pct[nrow(item_summary_df)]
item_rank_df <- data.frame(
item = item_summary_df$item,
top2_box_pct = item_summary_df$top2_box_pct,
stringsAsFactors = FALSE
)Step 9: Pooled scale usage — which options respondents actually chose
pooled <- colSums(counts_mat)
scale_levels_df <- data.frame(
rank = seq_len(k_levels),
level = common_levels,
responses = as.integer(pooled),
share = round(100 * pooled / max(1, sum(pooled)), 1),
stringsAsFactors = FALSE
)Step 10: The diverging stacked bar
Shares below the scale midpoint are drawn to the left of zero and shares above it to the right. On an odd-length scale the middle option sits on the left of zero so that the right-hand side reads as agreement only.
midpoint <- (k_levels + 1) / 2
sign_of <- ifelse(seq_len(k_levels) <= midpoint, -1, 1)
mid_on_left <- (k_levels %% 2) == 1
dist_rows <- list()
for (i in seq_along(used_items)) {
ii <- match(hn[i], item_summary_df$item)
for (j in seq_len(k_levels)) {
share <- 100 * counts_mat[i, j] / max(1, n_vec[i])
dist_rows[[length(dist_rows) + 1]] <- data.frame(
item = hn[i], level = common_levels[j],
share = round(sign_of[j] * share, 2),
order_key = ii, level_rank = j,
stringsAsFactors = FALSE
)
}
}
item_dist_df <- do.call(rbind, dist_rows)
item_dist_df <- item_dist_df[order(item_dist_df$order_key, item_dist_df$level_rank), , drop = FALSE]
item_dist_df <- item_dist_df[, c("item", "level", "share")]
rownames(item_dist_df) <- NULLStep 11: Optional between-group comparison — rank based, never a t-test
group_h <- humanize_semantic("group", col_map)
has_group <- "group" %in% names(df)
group_labels <- character(0); group_test_name <- ""
group_note <- ""
gvec <- NULL
if (has_group) {
g <- as.character(df$group)
g[!is.na(g) & trimws(g) == ""] <- NA_character_
tab <- table(g)
small <- names(tab)[tab < 5]
if (length(small) > 0) g[g %in% small] <- NA_character_
tab <- sort(table(g), decreasing = TRUE)
if (length(tab) > 8) {
keep <- names(tab)[1:8]
g[!is.na(g) & !(g %in% keep)] <- NA_character_
group_note <- sprintf("Only the 8 largest %s values were compared; smaller ones were set aside. ", group_h)
}
if (length(small) > 0) {
group_note <- paste0(group_note, sprintf(
"%d %s value(s) with fewer than 5 responses were set aside as too small to compare. ",
length(small), group_h))
}
gvec <- g
group_labels <- sort(unique(g[!is.na(g)]))
if (length(group_labels) < 2) {
has_group <- FALSE
group_note <- paste0(group_note, sprintf(
"Fewer than two %s values had enough responses, so no comparison was run. ", group_h))
} else {
group_test_name <- if (length(group_labels) == 2)
"Mann-Whitney U(Wilcoxon rank-sum)" else "Kruskal-Wallis"
}
}
if (has_group) {
rows <- list(); praw <- numeric(0)
for (i in seq_along(used_items)) {
z <- idx_list[[used_items[i]]]
ok <- !is.na(z) & !is.na(gvec)
zz <- z[ok]; gg <- factor(gvec[ok])
stat <- NA_real_; pv <- NA_real_; eff <- NA_real_
hi <- "n/a"; lo <- "n/a"; tname <- group_test_name
if (length(zz) >= 10 && nlevels(gg) >= 2 &&
all(table(gg) >= 3) && length(unique(zz)) > 1) {
mr <- tapply(zz, gg, mean)
mr_ok <- which(!is.na(mr))
if (length(mr_ok) > 0) {
hi <- names(mr)[mr_ok[which.max(mr[mr_ok])]]
lo <- names(mr)[mr_ok[which.min(mr[mr_ok])]]
}
if (nlevels(gg) == 2) {
tt <- suppressWarnings(tryCatch(stats::wilcox.test(zz ~ gg), error = function(e) NULL))
if (!is.null(tt)) {
n1 <- sum(gg == levels(gg)[1]); n2 <- sum(gg == levels(gg)[2])
stat <- unname(tt$statistic); pv <- tt$p.value
eff <- 2 * stat / (n1 * n2) - 1
}
} else {
tt <- tryCatch(stats::kruskal.test(zz ~ gg), error = function(e) NULL)
if (!is.null(tt)) {
stat <- unname(tt$statistic); pv <- tt$p.value
eff <- stat / max(1, length(zz) - 1)
}
}
} else {
tname <- "not run"
}
praw <- c(praw, pv)
rows[[i]] <- list(item = hn[i], test = tname, statistic = stat,
p_value = pv, effect = eff, hi = hi, lo = lo,
n = length(zz))
}
padj <- rep(NA_real_, length(praw))
okp <- which(!is.na(praw))
if (length(okp) > 0) padj[okp] <- stats::p.adjust(praw[okp], method = "holm")
eff_label <- if (length(group_labels) == 2) "rank-biserial correlation" else "epsilon squared"
group_tests_df <- do.call(rbind, lapply(seq_along(rows), function(i) {
r <- rows[[i]]
sig <- !is.na(padj[i]) && padj[i] < 0.05
interp <- if (r$test == "not run") {
sprintf("Too few usable responses to compare %s across %s.", r$item, group_h)
} else if (sig) {
sprintf("%s answer %s differently across %s(%s after Holm correction across %d items): %s rates it highest and %s lowest. The test compares the ORDER of the answers, so it needs no assumption about the spacing between response options.",
r$item, "is answered", group_h, fmt_pp(padj[i]), length(used_items), r$hi, r$lo)
} else {
sprintf("No difference in how %s is answered across %s survives correction(%s after Holm correction across %d items).",
r$item, group_h, fmt_pp(padj[i]), length(used_items))
}
data.frame(item = r$item, test = r$test,
statistic = if (is.na(r$statistic)) NA_real_ else round(r$statistic, 3),
p_value = fmt_p(r$p_value), p_adjusted = fmt_p(padj[i]),
effect_size = if (is.na(r$effect)) NA_real_ else round(r$effect, 3),
highest_group = r$hi, lowest_group = r$lo,
interpretation = interp, stringsAsFactors = FALSE)
}))
rownames(group_tests_df) <- NULL
group_sig_items <- group_tests_df$item[!is.na(padj) & padj < 0.05]
n_group_sig <- length(group_sig_items)
} else {
eff_label <- "n/a"
group_tests_df <- data.frame(
item = "(no comparison run)", test = "not run", statistic = NA_real_,
p_value = "n/a", p_adjusted = "n/a", effect_size = NA_real_,
highest_group = "n/a", lowest_group = "n/a",
interpretation = paste0(
if (nzchar(group_note)) group_note else
sprintf("No %s column was mapped, so the responses were described for the whole sample only. ", group_h),
"Map a categorical column to compare response distributions between groups with a rank-based test."),
stringsAsFactors = FALSE)
group_sig_items <- character(0); n_group_sig <- 0
}Step 12: Excluded columns and the methods disclosure
dropped_df <- if (length(dropped_items) > 0) {
data.frame(item = unname(item_names[dropped_items]),
reason = dropped_reasons, stringsAsFactors = FALSE)
} else {
data.frame(item = character(0), reason = character(0), stringsAsFactors = FALSE)
}
scale_sentence <- if (scale_kind == "numeric") {
sprintf("A %d-point scale was detected from the data: the response codes %s, read in ascending order.",
k_levels, paste(common_levels, collapse = ", "))
} else {
sprintf("A %d-point scale was detected from the data: the labels %s, ordered as a %s scale from the wording itself rather than assumed.",
k_levels, paste0("\"", paste(common_levels, collapse = "\", \""), "\""), scale_family)
}
methods_df <- data.frame(
aspect = c("Detected scale", "Top box", "Top-2 box", "Bottom box",
"Median category", "Modal category", mean_label,
"Diverging bar", "Group comparison", "Missing answers"),
detail = c(
scale_sentence,
sprintf("The share of answers at the highest option, \"%s\".", top_level),
sprintf("The share of answers at the two highest options, \"%s\" and \"%s\".",
common_levels[max(1, k_levels - 1)], top_level),
sprintf("The share of answers at the lowest option, \"%s\".", bottom_level),
"The lowest option whose cumulative share of answers reaches 50 percent. It is defined for ties and for an even number of respondents alike, and it uses only the ordering of the options.",
"The single most-chosen option. Where two options tie, the lower of the two is reported and the tie is flagged.",
mean_caveat,
sprintf("Shares at or below the scale midpoint are drawn left of zero and shares above it to the right, with items ordered by their top-2-box share%s.",
if (mid_on_left) sprintf("; the middle option \"%s\" is drawn on the left so the right-hand side reads as the top half of the scale only",
common_levels[ceiling(midpoint)]) else ""),
if (has_group)
sprintf("%s on the ranked answers, one test per item, %s correction across the %d items, with %s as the effect size. A rank-based test is used rather than a t-test on the raw codes because the codes are labels in order, not measured quantities.",
group_test_name, "Holm", length(used_items), eff_label)
else
group_tests_df$interpretation[1],
sprintf("%s blank answers were skipped; each item's percentages use that item's own answered count as the denominator, so items with different response rates stay comparable.",
comma(n_missing_total))
),
stringsAsFactors = FALSE
)
metrics <- list(
`Items Analysed` = length(used_items),
`Respondents` = final_rows,
`Scale Points` = k_levels,
`Scale Type` = scale_type,
`Top Box` = top_level,
`Highest Top-2-Box` = paste0(best_item, " (", pct1(best_top2), ")"),
`Lowest Top-2-Box` = paste0(worst_item, " (", pct1(worst_top2), ")"),
`Group Differences` = if (!has_group) "not tested" else
paste0(n_group_sig, " of ", length(used_items), " items")
)
group_clause <- if (has_group) {
if (n_group_sig > 0)
paste0(" Across ", group_h, ", ", oxford(group_sig_items),
" differ(s) by a ", group_test_name,
" test after Holm correction; the remaining ",
length(used_items) - n_group_sig, " do not.")
else
paste0(" No item differs across ", group_h,
" once the ", group_test_name,
" results are Holm-corrected across the ", length(used_items), " items.")
} else ""
json_output <- list(
answer = paste0(
"Ordinal analysis of ", length(used_items), " survey item(s) answered by ",
comma(final_rows), " respondents on a ", k_levels, "-point scale of ",
scale_type, " (", paste(common_levels, collapse = " < "), "). ",
best_item, " scores highest with ", pct1(best_top2),
" in the top two options and ", worst_item, " lowest with ",
pct1(worst_top2), "; the top box is \"", top_level, "\".",
group_clause,
" Percentages and median categories use only the ordering of the response options; ",
"where a mean of the codes is reported it assumes equal spacing between options, which the scale does not guarantee."
),
cards = lapply(
c("tldr", "overview", "preprocessing", "scale_detection", "item_summary",
"diverging_distribution", "top_box_ranking", "group_comparison", "methods"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
initial_rows = initial_rows, final_rows = final_rows, rows_removed = rows_removed,
item_names = item_names, used_items = used_items, item_h = hn,
dropped_df = dropped_df, n_missing_total = n_missing_total,
scale_type = scale_type, scale_kind = scale_kind, scale_family = scale_family,
common_levels = common_levels, k_levels = k_levels,
top_level = top_level, bottom_level = bottom_level,
scale_consistent = scale_consistent, scale_sentence = scale_sentence,
mid_on_left = mid_on_left, mean_caveat = mean_caveat, mean_label = mean_label,
scale_levels_df = scale_levels_df, item_summary_df = item_summary_df,
item_dist_df = item_dist_df, item_rank_df = item_rank_df,
group_tests_df = group_tests_df, methods_df = methods_df,
has_group = has_group, group_h = group_h, group_labels = group_labels,
group_test_name = group_test_name, group_note = group_note,
group_sig_items = group_sig_items, n_group_sig = n_group_sig,
mode_tied_any = any(mode_tied),
best_item = best_item, best_top2 = best_top2,
worst_item = worst_item, worst_top2 = worst_top2,
metrics = metrics, json_output = json_output
)
}