Executive Summary
Does assigned to workshop relate to depression score through job search selfefficacy?
The workshop assignment is not associated with a reduction in depression through job-search confidence. The indirect pathway (a × b) is −0.0157 with a 95% bootstrap confidence interval of −0.044 to 0.008, which includes zero. Workshop assignment shows a weak total association with depression (−0.056, p = 0.215), and job-search confidence is strongly associated with lower depression among those with the same workshop assignment (b = −0.240, p < 0.001), but the workshop does not reliably raise confidence in the first place (a = 0.0656, p = 0.203). The data do not support a mediated effect through confidence, and the proportion mediated is not reported because the total association is too weak. Interpreting any of these paths as causal requires ruling out unmeasured confounding, correct causal ordering, and no reverse causation—none of which the data can verify.
Analysis Overview
Mediation analysis of assigned to workshop and depression score through job search selfefficacy, across 899 rows.
This analysis uses mediation analysis to decompose the association between workshop assignment and depression score into two pathways: one running through job-search confidence (the indirect effect) and one that does not (the direct effect). The data contain 899 rows. Three separate least-squares regressions estimate the path coefficients: workshop assignment predicts job-search confidence (a), job-search confidence predicts depression holding workshop assignment fixed (b), and workshop assignment predicts depression directly (c-prime). The product of a and b yields the indirect effect; its confidence interval comes from percentile bootstrap resampling (2,000 resamples) rather than normal-theory approximation because the product of two coefficients is skewed. The analysis makes no claim about causation without three untestable assumptions: no unmeasured confounding of any pair among the three variables, correct causal ordering (workshop → confidence → depression), and no reverse causation from depression back to confidence.
Data Quality
Row completeness, column typing, and what was excluded.
The short answer
All 899 rows contained values for every mapped variable and required no imputation. Three covariates—economic hardship, sex, and age—entered every model to adjust for observed differences. No columns were dropped, and complete-case analysis preserved the full dataset.
The detail
Initial and final counts: 899 rows loaded, 899 rows analysed, 0 rows removed. The three covariate terms (economic hardship; sex; age) were included identically in all regressions. No constant or collinear columns were excluded, and no missing values were imputed in the mediator or outcome, because filling in a mediator would invent the quantity the analysis measures.
What this can't tell you
This export contains only the columns entered into the model. Any unmeasured variable—such as prior employment history, mental health baseline, or study site—cannot be controlled for and remains a source of potential confounding. Consider whether the covariates measured are sufficient to block the main pathways of confounding between workshop assignment and depression.
Effect Paths
Each arrow of the mediation model, with its estimate and uncertainty.
| Path | Estimate | Std Error | P Value | CI Low | CI High | Interpretation |
|---|---|---|---|---|---|---|
| a: assigned to workshop → job search selfefficacy | 0.0656 | 0.0515 | 0.203 | — | — | A one-unit-higher assigned to workshop is associated with a 0.066 higher job search selfefficacy. |
| b: job search selfefficacy → depression score (holding assigned to workshop fixed) | -0.24 | 0.0282 | < 0.001 | — | — | Among rows with the same assigned to workshop, a one-unit-higher job search selfefficacy is associated with a 0.240 lower depression score. |
| c-prime (direct): assigned to workshop → depression score (holding job search selfefficacy fixed) | -0.0403 | 0.0435 | 0.355 | — | — | The part of the assigned to workshop-to-depression score association that does not run through job search selfefficacy. |
| c (total): assigned to workshop → depression score | -0.056 | 0.0452 | 0.215 | — | — | The whole assigned to workshop-to-depression score association, before splitting it up. |
| a×b (indirect, through job search selfefficacy) | -0.0157 | — | 0.207 (Sobel) | -0.0444 | 0.0078 | The part that runs through job search selfefficacy. Bootstrap 95% CI -0.044 to 0.008, which includes zero. |
The short answer
The strongest link is between job-search confidence and depression: a one-unit higher confidence is associated with a 0.240-unit lower depression score (p < 0.001). The workshop's link to confidence is much weaker (a = 0.066, p = 0.203), so the indirect route through confidence does not reach statistical significance.
The detail
Path a (workshop → confidence): 0.0656 (SE 0.0515, p = 0.203). Path b (confidence → depression, holding workshop fixed): −0.240 (SE 0.0282, p < 0.001). Path c-prime (workshop → depression, holding confidence fixed): −0.0403 (SE 0.0435, p = 0.355). Path c (total, workshop → depression): −0.056 (SE 0.0452, p = 0.215). Indirect effect a × b: −0.0157 with bootstrap 95% CI −0.0444 to 0.0078, which includes zero. The Sobel test (p = 0.207) is reported for completeness; its normal approximation to a product of coefficients is poor, which is why the percentile bootstrap interval is the one to read.
What this can't tell you
These are associations measured on columns recorded together, not demonstrated mechanisms. The workshop's weak and uncertain link to confidence (p = 0.203) means the indirect effect is underpowered to detect even if a true pathway exists. Consider whether the workshop actually moved confidence in the direction expected, or whether confidence is the operative mechanism at all.
Effect Decomposition
Total, direct, and indirect association side by side.
The total association between workshop assignment and depression (−0.056) decomposes into a direct component (−0.0403) and an indirect component through job-search confidence (−0.0157). These sum exactly by construction under least squares. The indirect effect's bootstrap confidence interval (−0.0444 to 0.0078) overlaps zero, indicating the mediated pathway is not statistically distinguishable from no effect. The total association itself is weak (p = 0.215), so no proportion-mediated figure is reported; a ratio computed from such a small or uncertain total would be unstable and uninterpretable. The error bars on the direct and total components come from normal-theory intervals; the indirect bar uses percentile bootstrap, so their widths are not directly comparable.
Bootstrap Distribution
Where the indirect effect of 'assigned to workshop' on 'depression score' through 'job search selfefficacy' lands across 2,000 resamples of your rows.
The indirect effect was resampled 2,000 times by drawing rows with replacement, refitting both regressions on each resample, and recording the product a × b. The resulting distribution is centered near −0.016 but is skewed, with the 2.5th percentile at −0.0444 and the 97.5th percentile at 0.0078—a range that straddles zero. This skewness is precisely why the percentile bootstrap interval is used instead of the Sobel test: the product of two estimated coefficients is not normally distributed, especially in smaller samples, so a symmetric normal-theory interval would be too narrow and mis-centered. The bootstrap distribution shows that the indirect effect could plausibly be negative, zero, or slightly positive depending on the resample, confirming that the data do not reliably detect a mediated pathway.
What This Rests On
The assumptions a causal reading requires, and which of them this dataset can check.
| Assumption | What It Means | Checkable Here |
|---|---|---|
| No unmeasured confounding | Nothing outside the model causes both assigned to workshop and job search selfefficacy, both assigned to workshop and depression score, or both job search selfefficacy and depression score. | No — requires knowledge outside the dataset |
| Causal ordering | assigned to workshop comes before job search selfefficacy, which comes before depression score. | No — requires knowledge outside the dataset |
| No reverse causation | depression score does not feed back into job search selfefficacy. | No — requires knowledge outside the dataset |
| Design | Each row carries one observation of 3 columns measured together; the analysis has no record of anyone assigning assigned to workshop. | No — the dataset records no assignment mechanism |
| Model form | Each path is a straight line and the errors are well behaved; assigned to workshop and depression score are treated as numeric. | Partly — the fitted models assume it |
| What this analysis can conclude | That assigned to workshop, job search selfefficacy and depression score are associated in the pattern a mediation model would produce — not that assigned to workshop causes depression score by way of job search selfefficacy. | Yes — this is what was computed |
The short answer
Three assumptions are required to read these numbers as causal effects: no unmeasured variable drives both the workshop and confidence, or both the workshop and depression, or both confidence and depression; the causal order is workshop → confidence → depression; and depression does not feed back into confidence. None can be checked from this dataset alone.
The detail
The assumption table shows what each causal claim requires and what this dataset can verify. No unmeasured confounding: cannot check without external knowledge. Causal ordering: cannot check; the columns were measured together, so ordering comes from your knowledge of the process. No reverse causation: cannot check; depression could theoretically reduce confidence. Design: the dataset records no assignment mechanism, only values at one time. Model form: the analysis assumes linear paths and well-behaved errors; this is partly checkable through fitted models. Practical consequence: if you can name a plausible variable that drives both workshop assignment and confidence—such as prior motivation or employment history—and you have not mapped it as a covariate, the estimates above are the wrong size and the report cannot tell you by how much.
What this can't tell you
The data cannot rule out unmeasured confounding, reverse causation, or misspecified causal order. The covariates entered (economic hardship, sex, age) may not block all confounding pathways. Consider whether motivation, baseline depression, or prior job-search experience could drive both assignment and confidence, biasing the estimates.
Methods & Disclosure
Every model, formula, and setting behind the numbers.
| Item | Detail |
|---|---|
| Path a | Ordinary least squares of job search selfefficacy on assigned to workshop and the covariates: a = 0.066 (SE 0.051, p = 0.203). |
| Path b and c-prime | Ordinary least squares of depression score on assigned to workshop and job search selfefficacy and the covariates: b = -0.240 (SE 0.028, p < 0.001) and c-prime = -0.040 (SE 0.044, p = 0.355). |
| Path c (total) | Ordinary least squares of depression score on assigned to workshop and the covariates: c = -0.056 (SE 0.045, p = 0.215). |
| Indirect effect | a times b = -0.016, with percentile bootstrap 95% CI -0.044 to 0.008 from 2,000 resamples (2,000 usable). |
| Bootstrap | Rows resampled with replacement, both regressions refitted on each resample, and the 2.5th and 97.5th percentiles of the resulting a times b values reported. Seed 20260728, so the interval is reproducible. |
| Sobel test | z = -1.261 (p = 0.207), computed as a times b divided by the square root of (b squared times the variance of a, plus a squared times the variance of b). Reported for completeness; its normality assumption is poor for a product of coefficients, so the bootstrap interval is the one to read. |
| Proportion mediated | Not reported as a headline: the ratio of indirect to total is unstable here (total effect -0.056, p = 0.215), and a proportion computed from a small or oppositely-signed total effect can exceed one or turn negative without meaning anything. |
| Covariates | 3 covariate term(s) entered every regression identically: economic hardship; sex; age. |
| Decomposition check | With least squares on one sample and one covariate set, c must equal c-prime plus a times b exactly; the observed gap is 0.00000000, confirming the decomposition is arithmetically consistent. |
| What it does not establish | A mediation model fitted to columns measured together cannot establish causation. Reading this as "assigned to workshop changes job search selfefficacy, which in turn changes depression score" requires three assumptions that no amount of arithmetic on this table can check: (1) no unmeasured variable causes both assigned to workshop and job search selfefficacy, or both assigned to workshop and depression score, or both job search selfefficacy and depression score; (2) the causal order really is assigned to workshop then job search selfefficacy then depression score, and not some other ordering of the same three columns; (3) no reverse causation — depression score does not feed back into job search selfefficacy. These columns were measured together on the same rows, so the ordering comes from your knowledge of the process, not from the data. If any of the three fails, the numbers here are still an accurate description of the associations, but they are not causal effects. |
The short answer
Three ordinary least-squares regressions and a percentile bootstrap with 2,000 resamples were used to estimate the pathways. The indirect effect (−0.0157) was computed as a × b, with its 95% confidence interval (−0.044 to 0.008) drawn from the 2.5th and 97.5th percentiles of the resampled distribution rather than from the Sobel test.
The detail
Path a: OLS of confidence on workshop and covariates; a = 0.066 (SE 0.051, p = 0.203). Path b and c-prime: OLS of depression on workshop, confidence, and covariates; b = −0.240 (SE 0.028, p < 0.001), c-prime = −0.0403 (SE 0.044, p = 0.355). Path c: OLS of depression on workshop and covariates; c = −0.056 (SE 0.045, p = 0.215). Indirect effect: a × b = −0.0157, percentile bootstrap 95% CI −0.044 to 0.008 from 2,000 resamples (all usable), seed 20260728. Sobel test: z = −1.261 (p = 0.207), reported for completeness; its normality assumption is poor for a product of coefficients. Covariates (economic hardship, sex, age) entered every regression identically. Decomposition check: c = c-prime + a × b exactly (observed gap 0.00000000).
What this can't tell you
The bootstrap reproduces the same interval on identical data but does not address unmeasured confounding or violations of causal assumptions. The Sobel test is included for readers who expect it; its weakness is stated rather than hidden. Consider whether the model form (linear, no interactions) is appropriate for your substantive question.
Mediation & Moderation Analysis
Does a predictor relate to an outcome through a third variable (mediation), or does its relationship with the outcome depend on a third variable (moderation)? Map a predictor, an outcome, and either a mediator or a moderator (plus optional covariates).
Mediation decomposes the total association into a direct path and an indirect (mediated) path: the a path (predictor to mediator), the b path (mediator to outcome, holding the predictor fixed), the direct path c-prime, and the total path c. The indirect effect ab gets a percentile bootstrap confidence interval computed by resampling rows and refitting both models — the product ab is markedly non-normal, so a bootstrap interval is preferred over the Sobel normal-theory test, which is reported alongside for completeness.
Moderation fits the interaction model and reports the conditional (simple) slope of the predictor at low, mean, and high values of the moderator (or within each moderator level when the moderator is categorical).
Statistical honesty
On observational cross-sectional data this analysis does NOT establish that the predictor causes the outcome through the mediator. Every narrative it emits states the assumption set the causal reading rests on — no unmeasured confounding of any of the three paths, correct causal ordering, and no reverse causation — in plain language, using the user's own column names.
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {predictor, outcome, mediator | moderator, covariate_1..N}. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Core Analysis Pipeline
compute_shared <- function(df, params, col_map = list()) {
# === SHARED EXPORTS ===
# mode $ "mediation" | "moderation"
# initial_rows/final_rows/rows_removed/n $ row accounting
# x_h / y_h / m_h / w_h $ humanized user names of the mapped columns
# cov_labels $ humanized covariate labels actually in the model
# dropped_notes $ character — every column dropped and why
# a/b/c/c_prime + *_se/*_p $ mediation path estimates
# indirect / ind_lo / ind_hi $ indirect effect + percentile bootstrap CI
# prop_med / prop_med_ok $ proportion mediated + whether interpretable
# sobel_z / sobel_p $ Sobel normal-theory test (reported, not preferred)
# inter_coef/inter_se/inter_p $ moderation interaction estimate
# slopes_df $ conditional (simple) slopes with CIs
# boot_B / boot_valid / boot_seed $ bootstrap accounting
# path_df / effect_df / draws_df / assume_df / methods_df $ card datasets
# assumption_text / headline $ computed prose blocks reused across cards
# metrics / json_output
# === /SHARED EXPORTS ===Step 1: Resolve the mapped columns and decide the analysis mode
initial_rows <- nrow(df)
x_h <- humanize_semantic("predictor", col_map)
y_h <- humanize_semantic("outcome", col_map)
m_h <- humanize_semantic("mediator", col_map)
w_h <- humanize_semantic("moderator", col_map)
if (!("predictor" %in% names(df))) {
stop(sprintf("No predictor column was mapped. Map the column you think does the driving(requested as '%s').", x_h))
}
if (!("outcome" %in% names(df))) {
stop(sprintf("No outcome column was mapped. Map the numeric result you want explained(requested as '%s').", y_h))
}
has_m <- "mediator" %in% names(df)
has_w <- "moderator" %in% names(df)
if (!has_m && !has_w) {
stop(sprintf("Neither a mediator nor a moderator was mapped, so there is no third variable to analyse. Map a mediator(the step '%s' is supposed to work through on its way to '%s') for a mediation analysis, or a moderator (the condition under which the link between '%s' and '%s' changes) for a moderation analysis.",
x_h, y_h, x_h, y_h))
}
mode <- if (has_m) "mediation" else "moderation"
both_mapped <- has_m && has_wStep 2: Coerce the modelled columns (95% numeric-coercion rule)
dropped_notes <- character(0)
as_num_95 <- function(v) {
if (is.numeric(v)) return(v)
ch <- as.character(v)
non_blank <- !is.na(ch) & trimws(ch) != ""
conv <- suppressWarnings(as.numeric(ch))
if (sum(non_blank) == 0) return(NULL)
if (sum(!is.na(conv[non_blank])) < 0.95 * sum(non_blank)) return(NULL)
conv
}The predictor may legitimately be a two-level treatment flag; anything else non-numeric is refused by name rather than silently coerced.
xv <- as_num_95(df$predictor)
predictor_encoding <- ""
if (is.null(xv)) {
ch <- as.character(df$predictor)
lv <- sort(unique(trimws(ch[!is.na(ch) & trimws(ch) != ""])))
if (length(lv) == 2) {
xv <- ifelse(trimws(ch) == lv[2], 1, ifelse(trimws(ch) == lv[1], 0, NA_real_))
predictor_encoding <- sprintf("'%s' has exactly two values and was encoded as 0 for '%s' and 1 for '%s', so its coefficient is the difference between those two groups. ",
x_h, lv[1], lv[2])
} else {
stop(sprintf("The predictor column '%s' is neither numeric nor a two-value indicator (it has %d distinct values), so it cannot be used as a predictor here. Map a numeric column, or a column with exactly two values.",
x_h, length(lv)))
}
}
yv <- as_num_95(df$outcome)
if (is.null(yv)) {
stop(sprintf("The outcome column '%s' does not look numeric — fewer than 95%% of its values parse as numbers. Mediation and moderation here need a numeric outcome.", y_h))
}
mv <- NULL
if (mode == "mediation") {
mv <- as_num_95(df$mediator)
if (is.null(mv)) {
stop(sprintf("The mediator column '%s' does not look numeric — fewer than 95%% of its values parse as numbers. The mediator must be a numeric quantity that '%s' can move and that can in turn move '%s'.",
m_h, x_h, y_h))
}
}Step 3: Prepare the moderator (numeric, or a lumped categorical)
wv <- NULL; w_levels <- character(0); w_is_numeric <- TRUE
if (mode == "moderation") {
wv <- as_num_95(df$moderator)
if (is.null(wv)) {
w_is_numeric <- FALSE
ch <- trimws(as.character(df$moderator))
ch[is.na(ch) | ch == ""] <- "Missing"
tab <- sort(table(ch), decreasing = TRUE)
if (length(tab) > 8) {
keep_lv <- names(tab)[seq_len(8)]
ch[!(ch %in% keep_lv)] <- "Other"
dropped_notes <- c(dropped_notes, sprintf(
"'%s' had %d distinct values; the 8 most common were kept and the rest pooled into 'Other'.",
w_h, length(tab)))
}
if (length(unique(ch)) < 2) {
stop(sprintf("The moderator column '%s' has only one distinct value, so there is nothing for the link between '%s' and '%s' to vary across.",
w_h, x_h, y_h))
}
wv <- ch
}
}Step 4: Covariates — coerce, dummy-code, drop the unusable by name
cov_keys <- grep("^covariate_[0-9]+$", names(df), value = TRUE)
cov_keys <- cov_keys[order(as.integer(sub("^covariate_", "", cov_keys)))]
cov_mat_list <- list(); cov_labels <- character(0)
n_rows_all <- nrow(df)
for (ck in cov_keys) {
ch_lab <- humanize_semantic(ck, col_map)
v <- df[[ck]]
num <- as_num_95(v)
if (!is.null(num)) {
if (all(is.na(num))) {
dropped_notes <- c(dropped_notes, sprintf("'%s' was dropped: every value is missing.", ch_lab))
next
}
if (isTRUE(stats::var(num, na.rm = TRUE) == 0)) {
dropped_notes <- c(dropped_notes, sprintf("'%s' was dropped: it is constant, so it cannot explain any variation.", ch_lab))
next
}
cm <- matrix(num, ncol = 1); colnames(cm) <- ch_lab
cov_mat_list[[length(cov_mat_list) + 1]] <- cm
cov_labels <- c(cov_labels, ch_lab)
next
}
ch <- trimws(as.character(v))
ch[is.na(ch) | ch == ""] <- "Missing"
u <- unique(ch)
if (length(u) < 2) {
dropped_notes <- c(dropped_notes, sprintf("'%s' was dropped: it is constant, so it cannot explain any variation.", ch_lab))
next
}
if (length(u) > 0.9 * n_rows_all) {
dropped_notes <- c(dropped_notes, sprintf("'%s' was dropped: it has a near-unique value on almost every row, which makes it an identifier rather than a covariate.", ch_lab))
next
}
tab <- sort(table(ch), decreasing = TRUE)
if (length(tab) > 8) {
keep_lv <- names(tab)[seq_len(8)]
ch[!(ch %in% keep_lv)] <- "Other"
dropped_notes <- c(dropped_notes, sprintf("'%s' had %d categories; the 8 most common were kept and the rest pooled into 'Other'.", ch_lab, length(tab)))
}
lv <- sort(unique(ch))
ref <- lv[1]
for (l in lv[-1]) {
cm <- matrix(as.numeric(ch == l), ncol = 1)
colnames(cm) <- sprintf("%s: %s vs %s", ch_lab, l, ref)
cov_mat_list[[length(cov_mat_list) + 1]] <- cm
cov_labels <- c(cov_labels, colnames(cm))
}
}
C <- if (length(cov_mat_list) > 0) do.call(cbind, cov_mat_list) else NULLStep 5: Row-wise complete cases across everything the models use
ok <- !is.na(xv) & !is.na(yv)
if (mode == "mediation") ok <- ok & !is.na(mv)
if (mode == "moderation") ok <- ok & (if (w_is_numeric) !is.na(wv) else !is.na(wv) & wv != "")
if (!is.null(C)) ok <- ok & stats::complete.cases(C)
n_dropped <- sum(!ok)
xv <- xv[ok]; yv <- yv[ok]
if (mode == "mediation") mv <- mv[ok]
if (mode == "moderation") wv <- if (w_is_numeric) wv[ok] else wv[ok]
if (!is.null(C)) C <- C[ok, , drop = FALSE]
n <- length(xv)
final_rows <- n
rows_removed <- initial_rows - final_rows
MIN_ROWS <- 30
if (n < MIN_ROWS) {
stop(sprintf("Only %d rows have a value for every mapped column('%s', '%s'%s) — at least %d complete rows are needed before a mediation or moderation model can be fitted, and a percentile bootstrap on fewer rows would be meaningless.",
n, x_h, y_h,
if (mode == "mediation") sprintf(", '%s'", m_h) else sprintf(", '%s'", w_h),
MIN_ROWS))
}Step 6: Degenerate-input guards, each naming the user's own column
if (isTRUE(stats::var(xv) == 0)) {
stop(sprintf("The predictor column '%s' is constant — every row has the same value — so it cannot be related to anything.", x_h))
}
if (isTRUE(stats::var(yv) == 0)) {
stop(sprintf("The outcome column '%s' is constant — every row has the same value — so there is no variation to explain.", y_h))
}
if (mode == "mediation" && isTRUE(stats::var(mv) == 0)) {
stop(sprintf("The mediator column '%s' is constant — every row has the same value — so nothing can travel through it from '%s' to '%s'.",
m_h, x_h, y_h))
}
if (mode == "moderation" && w_is_numeric && isTRUE(stats::var(wv) == 0)) {
stop(sprintf("The moderator column '%s' is constant — every row has the same value — so the link between '%s' and '%s' has nothing to vary across.",
w_h, x_h, y_h))
}Drop covariate columns that are collinear with the rest of the design.
if (!is.null(C)) {
keep_c <- rep(TRUE, ncol(C))
for (j in seq_len(ncol(C))) {
if (isTRUE(stats::var(C[, j]) == 0)) {
dropped_notes <- c(dropped_notes, sprintf("'%s' was dropped: it is constant on the rows that survived cleaning.", colnames(C)[j]))
keep_c[j] <- FALSE
}
}
C <- C[, keep_c, drop = FALSE]
if (ncol(C) > 0) {
base_cols <- cbind(1, xv, if (mode == "mediation") mv else NULL)
probe <- cbind(base_cols, C)
qp <- qr(probe)
if (qp$rank < ncol(probe)) {
keep_idx <- sort(qp$pivot[seq_len(qp$rank)])
drop_idx <- setdiff(seq_len(ncol(probe)), keep_idx)
drop_cov <- drop_idx[drop_idx > ncol(base_cols)] - ncol(base_cols)
for (j in drop_cov) {
dropped_notes <- c(dropped_notes, sprintf("'%s' was dropped: it is an exact linear combination of the other columns in the model.", colnames(C)[j]))
}
C <- C[, setdiff(seq_len(ncol(C)), drop_cov), drop = FALSE]
}
}
if (ncol(C) == 0) C <- NULL
}
cov_labels <- if (is.null(C)) character(0) else colnames(C)Step 7: Bootstrap settings — deterministic and disclosed
boot_seed <- 20260728L
boot_B <- suppressWarnings(as.integer(params$bootstrap_samples %||% 2000L))
if (is.na(boot_B)) boot_B <- 2000L
boot_B <- max(500L, min(5000L, boot_B))
if (n > 5000) boot_B <- min(boot_B, 1000L)
a <- b <- c_tot <- c_prime <- NA_real_
a_se <- b_se <- c_se <- cp_se <- NA_real_
a_p <- b_p <- c_p <- cp_p <- NA_real_
indirect <- ind_lo <- ind_hi <- NA_real_
prop_med <- NA_real_; prop_med_ok <- FALSE
sobel_z <- sobel_p <- NA_real_
inter_coef <- inter_se <- inter_p <- NA_real_
slopes_df <- NULL
boot_valid <- 0L
boot_draws <- numeric(0)
boot_label <- ""
identity_gap <- NA_real_
if (mode == "mediation") {Step 8 (mediation): the three regressions behind a, b, c and c-prime
Am <- cbind(1, xv); if (!is.null(C)) Am <- cbind(Am, C) # M ~ X + covariates
Ay <- cbind(1, xv, mv); if (!is.null(C)) Ay <- cbind(Ay, C) # Y ~ X + M + covariates
At <- cbind(1, xv); if (!is.null(C)) At <- cbind(At, C) # Y ~ X + covariates
fit_m <- ols_fit(Am, mv)
fit_y <- ols_fit(Ay, yv)
fit_t <- ols_fit(At, yv)
if (is.null(fit_m) || is.null(fit_y) || is.null(fit_t)) {
stop(sprintf("The mediation models for '%s', '%s' and '%s' could not be fitted — the mapped columns (or the covariates) are collinear, so the paths are not separately identified.",
x_h, m_h, y_h))
}
a <- fit_m$coef[2]; a_se <- fit_m$se[2]; a_p <- fit_m$p[2]
c_prime <- fit_y$coef[2]; cp_se <- fit_y$se[2]; cp_p <- fit_y$p[2]
b <- fit_y$coef[3]; b_se <- fit_y$se[3]; b_p <- fit_y$p[3]
c_tot <- fit_t$coef[2]; c_se <- fit_t$se[2]; c_p <- fit_t$p[2]
indirect <- a * bWith OLS on one sample and one covariate set, c = c-prime + a*b exactly.
identity_gap <- abs(c_tot - (c_prime + indirect))Step 9 (mediation): percentile bootstrap for the indirect effect
set.seed(boot_seed)
boot_draws <- replicate(boot_B, {
idx <- sample.int(n, n, replace = TRUE)
fm <- .lm.fit(Am[idx, , drop = FALSE], mv[idx])
fy <- .lm.fit(Ay[idx, , drop = FALSE], yv[idx])
if (fm$rank < ncol(Am) || fy$rank < ncol(Ay)) return(NA_real_)
ca <- fm$coefficients[2]; cb <- fy$coefficients[3]
if (is.na(ca) || is.na(cb)) return(NA_real_)
ca * cb
})
boot_label <- sprintf("indirect effect of '%s' on '%s' through '%s'", x_h, y_h, m_h)
good <- boot_draws[is.finite(boot_draws)]
boot_valid <- length(good)
if (boot_valid >= max(200L, as.integer(0.8 * boot_B))) {
qs <- stats::quantile(good, c(0.025, 0.975), names = FALSE, type = 7)
ind_lo <- qs[1]; ind_hi <- qs[2]
}Step 10 (mediation): Sobel test — reported, but not the preferred CI
if (!is.na(a_se) && !is.na(b_se)) {
denom <- sqrt(b^2 * a_se^2 + a^2 * b_se^2)
if (is.finite(denom) && denom > 0) {
sobel_z <- indirect / denom
sobel_p <- 2 * stats::pnorm(-abs(sobel_z))
}
}Proportion mediated is only interpretable when the total path is non-trivial and the direct and indirect paths point the same way.
if (!is.na(c_tot) && abs(c_tot) > 1e-9) {
prop_med <- indirect / c_tot
prop_med_ok <- is.finite(prop_med) && prop_med > 0 && prop_med <= 1 &&
!is.na(c_p) && c_p < 0.05
}
} else {Step 8 (moderation): the interaction model
if (w_is_numeric) {
Aw <- cbind(1, xv, wv, xv * wv)
lab_w <- c("(Intercept)", x_h, w_h, sprintf("%s × %s", x_h, w_h))
} else {
lv <- sort(unique(wv)); ref <- lv[1]
D <- sapply(lv[-1], function(l) as.numeric(wv == l))
D <- matrix(as.numeric(D), nrow = n)
XD <- D * xv
Aw <- cbind(1, xv, D, XD)
lab_w <- c("(Intercept)", x_h,
sprintf("%s: %s vs %s", w_h, lv[-1], ref),
sprintf("%s × %s: %s vs %s", x_h, w_h, lv[-1], ref))
}
if (!is.null(C)) { Aw <- cbind(Aw, C); lab_w <- c(lab_w, cov_labels) }
fit_w <- ols_fit(Aw, yv)
if (is.null(fit_w)) {
stop(sprintf("The moderation model for '%s', '%s' and '%s' could not be fitted — the mapped columns (or the covariates) are collinear, so the interaction is not identified.",
x_h, w_h, y_h))
}Step 9 (moderation): conditional (simple) slopes with their own CIs
tq <- stats::qt(0.975, df = fit_w$df)
simple_slope <- function(k) {
# k is a contrast vector over the coefficient vector
est <- sum(k * fit_w$coef)
va <- as.numeric(t(k) %*% fit_w$vcov %*% k)
se <- sqrt(max(va, 0))
tv <- if (se > 0) est / se else NA_real_
pv <- if (is.na(tv)) NA_real_ else 2 * stats::pt(-abs(tv), df = fit_w$df)
c(est = est, se = se, lo = est - tq * se, hi = est + tq * se, p = pv)
}
p_all <- length(fit_w$coef)
if (w_is_numeric) {
mw <- mean(wv); sw <- stats::sd(wv)
w_points <- c(mw - sw, mw, mw + sw)
w_names <- c(sprintf("Low %s(%s)", w_h, r2(w_points[1])),
sprintf("Mean %s(%s)", w_h, r2(w_points[2])),
sprintf("High %s(%s)", w_h, r2(w_points[3])))
rows <- lapply(seq_along(w_points), function(i) {
k <- numeric(p_all); k[2] <- 1; k[4] <- w_points[i]
simple_slope(k)
})
inter_coef <- fit_w$coef[4]; inter_se <- fit_w$se[4]; inter_p <- fit_w$p[4]
} else {
lv <- sort(unique(wv))
w_names <- sprintf("%s = %s", w_h, lv)
rows <- lapply(seq_along(lv), function(i) {
k <- numeric(p_all); k[2] <- 1
if (i > 1) k[2 + length(lv) - 1 + (i - 1)] <- 1
simple_slope(k)
})With a categorical moderator the "interaction" is a set of terms; the headline number is the spread between the strongest and weakest slope.
ests <- sapply(rows, function(r) r[["est"]])
hi_i <- which(ests == max(ests))[1]; lo_i <- which(ests == min(ests))[1]
k <- numeric(p_all)
if (hi_i > 1) k[2 + length(lv) - 1 + (hi_i - 1)] <- 1
if (lo_i > 1) k[2 + length(lv) - 1 + (lo_i - 1)] <- k[2 + length(lv) - 1 + (lo_i - 1)] - 1
sp <- simple_slope(k)
inter_coef <- sp[["est"]]; inter_se <- sp[["se"]]; inter_p <- sp[["p"]]
}
slopes_df <- data.frame(
condition = w_names,
slope = round(sapply(rows, function(r) r[["est"]]), 4),
std_error = round(sapply(rows, function(r) r[["se"]]), 4),
ci_low = round(sapply(rows, function(r) r[["lo"]]), 4),
ci_high = round(sapply(rows, function(r) r[["hi"]]), 4),
p_value = sapply(rows, function(r) fmt_p(r[["p"]])),
stringsAsFactors = FALSE
)Step 10 (moderation): bootstrap the moderation quantity itself
inter_idx <- if (w_is_numeric) 4L else NA_integer_
kv <- if (w_is_numeric) NULL else {
lv <- sort(unique(wv)); ests <- slopes_df$slope
hi_i <- which(ests == max(ests))[1]; lo_i <- which(ests == min(ests))[1]
k <- numeric(p_all)
if (hi_i > 1) k[2 + length(lv) - 1 + (hi_i - 1)] <- 1
if (lo_i > 1) k[2 + length(lv) - 1 + (lo_i - 1)] <- k[2 + length(lv) - 1 + (lo_i - 1)] - 1
k
}
set.seed(boot_seed)
boot_draws <- replicate(boot_B, {
idx <- sample.int(n, n, replace = TRUE)
fw <- .lm.fit(Aw[idx, , drop = FALSE], yv[idx])
if (fw$rank < ncol(Aw)) return(NA_real_)
cf <- fw$coefficients
if (any(is.na(cf))) return(NA_real_)
if (!is.na(inter_idx)) cf[inter_idx] else sum(kv * cf)
})
boot_label <- if (w_is_numeric)
sprintf("interaction between '%s' and '%s'", x_h, w_h)
else
sprintf("gap between the strongest and weakest '%s' slope across levels of '%s'", x_h, w_h)
good <- boot_draws[is.finite(boot_draws)]
boot_valid <- length(good)
if (boot_valid >= max(200L, as.integer(0.8 * boot_B))) {
qs <- stats::quantile(good, c(0.025, 0.975), names = FALSE, type = 7)
ind_lo <- qs[1]; ind_hi <- qs[2]
}
indirect <- inter_coef
}
boot_sig <- !is.na(ind_lo) && !is.na(ind_hi) && (ind_lo > 0 || ind_hi < 0)Step 11: The assumption block — computed, and printed every time
assumption_text <- if (mode == "mediation") {
paste0(
"Reading this as \"", x_h, " changes ", m_h, ", which in turn changes ", y_h,
"\" requires three assumptions that no amount of arithmetic on this table can check: ",
"(1) no unmeasured variable causes both ", x_h, " and ", m_h, ", or both ", x_h,
" and ", y_h, ", or both ", m_h, " and ", y_h, "; ",
"(2) the causal order really is ", x_h, " then ", m_h, " then ", y_h,
", and not some other ordering of the same three columns; ",
"(3) no reverse causation — ", y_h, " does not feed back into ", m_h, ". ",
"These columns were measured together on the same rows, so the ordering comes from ",
"your knowledge of the process, not from the data. If any of the three fails, the ",
"numbers here are still an accurate description of the associations, but they are ",
"not causal effects."
)
} else {
paste0(
"Reading this as \"", w_h, " changes how ", x_h, " affects ", y_h,
"\" requires assumptions this data cannot check: ",
"(1) no unmeasured variable causes both ", x_h, " and ", y_h,
" in a way that happens to vary with ", w_h, "; ",
"(2) ", w_h, " is a condition that precedes the ", x_h, "-to-", y_h,
" relationship rather than a consequence of it; ",
"(3) no reverse causation — ", y_h, " does not feed back into ", x_h, " or ", w_h, ". ",
"These columns were measured together on the same rows, so the ordering comes from ",
"your knowledge of the process, not from the data. An interaction is a statement ",
"about how an association varies, not proof that either variable causes the other."
)
}
boot_pref_text <- paste0(
"The confidence interval above is a percentile bootstrap: the rows were resampled with ",
"replacement ", format(boot_B, big.mark = ","), " times, the model refitted on each ",
"resample, and the middle 95% of the resulting estimates reported. This is used in ",
"preference to the Sobel test because ",
if (mode == "mediation")
"the product of two coefficients is skewed and not normally distributed, so a normal-theory interval on it is too narrow and mis-centred, especially in smaller samples. The Sobel test is reported alongside for completeness, not as the verdict."
else
"the sampling distribution of a conditional effect can be skewed, and the bootstrap does not require it to be normal."
)Step 13: Metrics and the computed one-paragraph answer
metrics <- if (mode == "mediation") {
list(
`Rows Analysed` = n,
`Path a` = round(a, 4),
`Path b` = round(b, 4),
`Direct Effect` = round(c_prime, 4),
`Total Effect` = round(c_tot, 4),
`Indirect Effect` = round(indirect, 4),
`Indirect 95% CI` = paste0(r3(ind_lo), " to ", r3(ind_hi)),
`Indirect CI Excludes Zero` = if (boot_sig) "yes" else "no",
`Proportion Mediated` = if (prop_med_ok) pct1(100 * prop_med) else "not interpretable",
`Bootstrap Resamples` = as.integer(boot_B)
)
} else {
list(
`Rows Analysed` = n,
`Moderation Estimate` = round(inter_coef, 4),
`Moderation 95% CI` = paste0(r3(ind_lo), " to ", r3(ind_hi)),
`Moderation CI Excludes Zero` = if (boot_sig) "yes" else "no",
`Lowest Conditional Slope` = round(min(slopes_df$slope), 4),
`Highest Conditional Slope` = round(max(slopes_df$slope), 4),
`Bootstrap Resamples` = as.integer(boot_B)
)
}
headline <- if (mode == "mediation") {
paste0(
"Across ", format(n, big.mark = ","), " rows, ", x_h, " is associated with ", m_h,
" (a = ", r3(a), ", ", fmt_pp(a_p), "), and ", m_h, " is associated with ", y_h,
" among rows with the same ", x_h, " (b = ", r3(b), ", ", fmt_pp(b_p), "). ",
"The indirect path a×b is ", r3(indirect), " with a percentile bootstrap 95% CI of ",
r3(ind_lo), " to ", r3(ind_hi), ", which ",
if (boot_sig) "excludes zero" else "includes zero",
"; the direct path c-prime is ", r3(c_prime), " and the total is ", r3(c_tot), ". ",
if (prop_med_ok)
paste0(pct1(100 * prop_med), " of the total association runs through ", m_h, ". ")
else
paste0("A proportion-mediated figure is not reported: the total association(",
r3(c_tot), ", ", fmt_pp(c_p),
") is not solid enough for that ratio to mean anything. ")
)
} else {
paste0(
"Across ", format(n, big.mark = ","), " rows, the association between ", x_h,
" and ", y_h, " varies with ", w_h, ": the conditional slope runs from ",
r3(min(slopes_df$slope)), " to ", r3(max(slopes_df$slope)),
" across the reported conditions. The moderation estimate is ", r3(inter_coef),
" (bootstrap 95% CI ", r3(ind_lo), " to ", r3(ind_hi), "), which ",
if (boot_sig) "excludes zero" else "includes zero", ". "
)
}
json_output <- list(
answer = paste0(
if (mode == "mediation") "Mediation analysis. " else "Moderation analysis. ",
headline,
"These are regression decompositions of associations, not established causal effects: ",
assumption_text
),
cards = lapply(
c("tldr", "overview", "preprocessing", "effect_paths", "effect_decomposition",
"bootstrap_distribution", "assumptions", "methods"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
mode = mode, both_mapped = both_mapped,
initial_rows = initial_rows, final_rows = final_rows,
rows_removed = rows_removed, n = n, n_dropped = n_dropped,
x_h = x_h, y_h = y_h, m_h = m_h, w_h = w_h,
w_is_numeric = w_is_numeric,
cov_labels = cov_labels, dropped_notes = dropped_notes,
predictor_encoding = predictor_encoding,
a = a, b = b, c_tot = c_tot, c_prime = c_prime,
a_se = a_se, b_se = b_se, c_se = c_se, cp_se = cp_se,
a_p = a_p, b_p = b_p, c_p = c_p, cp_p = cp_p,
indirect = indirect, ind_lo = ind_lo, ind_hi = ind_hi, boot_sig = boot_sig,
prop_med = prop_med, prop_med_ok = prop_med_ok,
sobel_z = sobel_z, sobel_p = sobel_p, identity_gap = identity_gap,
inter_coef = inter_coef, inter_se = inter_se, inter_p = inter_p,
slopes_df = slopes_df,
boot_B = boot_B, boot_valid = boot_valid, boot_seed = boot_seed,
boot_label = boot_label, boot_pref_text = boot_pref_text,
assumption_text = assumption_text, headline = headline,
path_df = path_df, effect_df = effect_df, draws_df = draws_df,
assume_df = assume_df, methods_df = methods_df,
metrics = metrics, json_output = json_output
)
}