Executive Summary
Relationship map across 12 columns
The short answer
The 12 chemical properties form a densely interconnected web: 55 of 66 pairs show statistically significant associations at p < 0.05. The strongest relationship is fixed acidity and pH, which move in opposite directions (r = −0.683, p = 4.06e-220).
The detail
Across 1,599 observations and 66 column pairs tested, fixed acidity ~ pH leads with r = −0.683 and p = 4.06e-220, marking a highly reliable negative association. No pair reaches r ≥ 0.7 in absolute value. All top relationships carry p < 0.001, confirming they exceed chance.
What this can't tell you
Significance at p < 0.05 means the pattern is real; it does not indicate practical importance or how tightly the relationship holds in individual cases. Moderate correlations (r ≈ 0.67) still leave substantial unexplained variance.
Analysis Overview
Pairwise Pearson correlations across 12 columns and 1,599 observations.
The short answer
Pearson correlation measures how two variables move together on a scale from −1 (perfect opposition) to +1 (perfect co-movement). This analysis tests all 66 possible pairs among 12 chemical properties and quality, separating real associations from chance variation.
The detail
The analysis computed pairwise Pearson r for 1,599 observations across fixed acidity, volatile acidity, citric acid, residual sugar, chlorides, free sulfur dioxide, total sulfur dioxide, density, pH, sulphates, alcohol, and quality. A significance test per pair flags relationships where p < 0.05. Correlation measures association only; it does not imply one variable drives another.
What this can't tell you
Pearson r assumes a linear relationship. If the true pattern is curved or if outliers dominate, r may misrepresent the strength of association; the scatter plot for the strongest pair tests this assumption.
Data Quality
Column typing, imputation, and exclusions.
The short answer
All 1,599 observations were retained and all 12 mapped columns were usable. Missing values were imputed with each column's median, and pairwise complete observations were used for each correlation.
The detail
Initial rows: 1,599; final rows: 1,599; rows removed: 0. No columns were excluded as non-numeric or constant. Median imputation filled gaps, preserving the full dataset for analysis.
What this can't tell you
The preprocessing report does not disclose how many missing values existed before imputation. If missingness was substantial and non-random, median imputation may have introduced bias. Consider requesting a missing-data summary to assess whether re-analysis with listwise deletion or multiple imputation would change the findings.
Correlation Matrix
Every pairwise correlation across the mapped columns.
The short answer
The heatmap reveals blocks of moderate correlation rather than strong redundancy. Fixed acidity clusters moderately with citric acid (r = 0.672) and density (r = 0.668), and moves opposite to pH (r = −0.683). No pair achieves r ≥ 0.7, so no two properties are near-redundant.
The detail
The color scale ranges from deep blue (r = −1, perfect opposition) to deep red (r = +1, perfect co-movement). The hottest cells are fixed acidity ~ pH at r = −0.683, fixed acidity ~ citric acid at r = 0.672, and fixed acidity ~ density at r = 0.668. Row-wise, fixed acidity shows the strongest gradients. The coldest (most negative) cells include volatile acidity ~ citric acid (r = −0.552) and citric acid ~ pH (r = −0.542). 0 pairs show strong positive correlation (r ≥ 0.7).
What this can't tell you
A heatmap displays pairwise correlations only; it does not reveal multivariate patterns or whether one variable mediates the relationship between two others. Partial correlation (controlling for a third variable) would clarify whether associations are direct or indirect.
Strongest Relationships
Top column pairs ranked by absolute correlation, with significance.
| Pair | Correlation | P Value | N | Strength | Significance |
|---|---|---|---|---|---|
| fixed acidity ~ pH | -0.683 | 4.06e-220 | 1599 | moderate | *** |
| fixed acidity ~ citric acid | 0.672 | 2.54e-210 | 1599 | moderate | *** |
| fixed acidity ~ density | 0.668 | 3.07e-207 | 1599 | moderate | *** |
| free sulfur dioxide ~ total sulfur dioxide | 0.668 | 6.4e-207 | 1599 | moderate | *** |
| volatile acidity ~ citric acid | -0.552 | 1.81e-128 | 1599 | moderate | *** |
| citric acid ~ pH | -0.542 | 1.01e-122 | 1599 | moderate | *** |
| density ~ alcohol | -0.496 | 3.94e-100 | 1599 | moderate | *** |
| alcohol ~ quality | 0.476 | 2.83e-91 | 1599 | moderate | *** |
| volatile acidity ~ quality | -0.391 | 2.05e-59 | 1599 | weak | *** |
| chlorides ~ sulphates | 0.371 | 1.99e-53 | 1599 | weak | *** |
| citric acid ~ density | 0.365 | 1.48e-51 | 1599 | weak | *** |
| residual sugar ~ density | 0.355 | 9.01e-49 | 1599 | weak | *** |
| density ~ pH | -0.342 | 5.12e-45 | 1599 | weak | *** |
| citric acid ~ sulphates | 0.313 | 1.27e-37 | 1599 | weak | *** |
| chlorides ~ pH | -0.265 | 4.15e-27 | 1599 | weak | *** |
The short answer
Two pairs tie for second-strongest: fixed acidity ~ citric acid (r = 0.672) and fixed acidity ~ density (r = 0.668), both highly significant. Among pairs involving quality, alcohol shows the strongest link (r = 0.476, p = 2.83e-91), followed by volatile acidity (r = −0.391, p = 2.05e-59).
The detail
The top five by absolute correlation are: fixed acidity ~ pH (r = −0.683, p = 4.06e-220), fixed acidity ~ citric acid (r = 0.672, p = 2.54e-210), fixed acidity ~ density (r = 0.668, p = 3.07e-207), free sulfur dioxide ~ total sulfur dioxide (r = 0.668, p = 6.4e-207), and volatile acidity ~ citric acid (r = −0.552, p = 1.81e-128). All carry p < 0.001 and n = 1,599. None reach |r| ≥ 0.7. For quality, alcohol (r = 0.476) and volatile acidity (r = −0.391) dominate; all other quality pairs fall below |r| = 0.3.
What this can't tell you
Moderate correlations with quality do not establish which chemical properties are most important for predicting or improving quality. Multivariate methods (regression, feature importance) would rank predictive power. The strongest pairs among chemical properties may reflect measurement overlap rather than independent drivers of quality.
Strongest Pair
The strongest relationship plotted point by point.
The short answer
Fixed acidity and pH form a tight downward-sloping band, confirming the linear negative relationship (r = −0.683). Points cluster around the trend with modest scatter; no extreme outliers visibly dominate the correlation.
The detail
The scatter shows fixed acidity (horizontal) against pH (vertical) for 1,599 observations. As fixed acidity increases from ~4.6 to ~14, pH decreases from ~4 to ~3, tracing a clear negative linear pattern. The cloud is narrow enough to confirm that r = −0.683 accurately describes the relationship; curvature or clusters are not apparent. Scatter around the line is consistent across the range, suggesting homogeneous variance.
What this can't tell you
The scatter plot shows association, not causation. Tightness of the band confirms the linear model fits; it does not explain why acidity and pH move together (they may both reflect fermentation chemistry or be mechanically linked by the buffer system). A residual plot would reveal whether the linear model systematically mispredicts at any acidity level.
Correlation Analysis — What Moves Together
Computes pairwise Pearson correlations across the numeric columns the user selects: full matrix heatmap, strongest relationships ranked with significance tests, and a scatter of the dominant pair.
Why This Method?
Correlation is the fastest map of a dataset's relationships: one number per pair, directly comparable, with a significance test separating real co-movement from noise. It is the standard first step before regression, forecasting, or KPI pruning.
What This Analysis Covers
- Full correlation matrix (heatmap)
- Strongest relationships ranked with p-values
- The dominant pair plotted point by point
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {feature_1..feature_N}. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Core Analysis Pipeline
compute_shared <- function(df, params, col_map = list()) {
# === SHARED EXPORTS ===
# initial_rows/final_rows/rows_removed $ row accounting
# feature_names $ named character — semantic -> humanized names
# used_features $ character — semantic names used
# dropped_features $ character — excluded columns
# cor_long_df $ data.frame(feature_a, feature_b, correlation) — full matrix
# pairs_df $ data.frame(pair, correlation, p_value, n, strength, significance)
# top_pair_df $ data.frame(value_a, value_b) — <=1000 sample of strongest pair
# top_pair_names $ character(2) — humanized names of the strongest pair
# top_r / top_p $ numeric — strongest pair stats
# n_sig $ integer — pairs significant at p<0.05
# metrics / json_output
# === /SHARED EXPORTS ===Step 1: Discover mapped features
initial_rows <- nrow(df)
feat_cols <- grep("^feature_[0-9]+$", names(df), value = TRUE)
feat_cols <- feat_cols[order(as.integer(sub("^feature_", "", feat_cols)))]
if (length(feat_cols) < 2) {
stop("column_mapping must map at least two feature columns(feature_1, feature_2)")
}
feature_names <- setNames(humanize_semantic(feat_cols, col_map), feat_cols)Step 2: Coerce numeric (95% rule); impute median; drop unusable
dropped_features <- character(0)
for (fc in feat_cols) {
v <- df[[fc]]
if (!is.numeric(v)) {
conv <- suppressWarnings(as.numeric(as.character(v)))
n_orig <- sum(!is.na(v) & as.character(v) != "")
if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) {
df[[fc]] <- conv
} else {
dropped_features <- c(dropped_features, fc); next
}
}
v <- df[[fc]]
med <- median(v, na.rm = TRUE)
if (is.na(med)) { dropped_features <- c(dropped_features, fc); next }
v[is.na(v)] <- med
df[[fc]] <- v
if (isTRUE(var(v) == 0) || is.na(var(v))) {
dropped_features <- c(dropped_features, fc)
}
}
used_features <- setdiff(feat_cols, dropped_features)
if (length(used_features) < 2) {
stop(paste0("Correlation needs at least two usable numeric columns; only ",
length(used_features), " remained after cleaning."))
}
X <- df[, used_features, drop = FALSE]
final_rows <- nrow(X)
rows_removed <- initial_rows - final_rows
if (final_rows < 10) stop(sprintf("Only %d usable rows — need at least 10.", final_rows))Step 3: Correlation matrix + per-pair tests
C <- suppressWarnings(cor(X, use = "pairwise.complete.obs", method = "pearson"))
hn <- unname(feature_names[used_features])Long format for the heatmap (diagonal included — reads as the standard matrix)
k <- length(used_features)
cor_long_df <- do.call(rbind, lapply(seq_len(k), function(i) {
data.frame(feature_a = hn[i], feature_b = hn,
correlation = round(as.numeric(C[i, ]), 3),
stringsAsFactors = FALSE)
}))
rownames(cor_long_df) <- NULLPairwise tests (upper triangle only)
pair_rows <- list()
for (i in seq_len(k - 1)) {
for (j in (i + 1):k) {
r_ij <- C[i, j]
if (is.na(r_ij)) next
ct <- tryCatch(cor.test(X[[i]], X[[j]]), error = function(e) NULL)
p_ij <- if (!is.null(ct)) ct$p.value else NA_real_
pair_rows[[length(pair_rows) + 1]] <- data.frame(
pair = paste0(hn[i], " ~ ", hn[j]),
a = used_features[i], b = used_features[j],
correlation = round(r_ij, 3),
p_value = signif(p_ij, 3),
n = sum(complete.cases(X[[i]], X[[j]])),
stringsAsFactors = FALSE
)
}
}
pairs_all <- do.call(rbind, pair_rows)
if (is.null(pairs_all) || nrow(pairs_all) == 0) {
stop("No valid correlation pairs could be computed from the mapped columns.")
}
pairs_all <- pairs_all[order(-abs(pairs_all$correlation)), , drop = FALSE]
rownames(pairs_all) <- NULL
pairs_all$strength <- ifelse(abs(pairs_all$correlation) >= 0.7, "strong",
ifelse(abs(pairs_all$correlation) >= 0.4, "moderate", "weak"))
pairs_all$significance <- ifelse(is.na(pairs_all$p_value), "",
ifelse(pairs_all$p_value < 0.001, "***",
ifelse(pairs_all$p_value < 0.01, "**",
ifelse(pairs_all$p_value < 0.05, "*", ""))))
n_sig <- sum(pairs_all$p_value < 0.05, na.rm = TRUE)
pairs_df <- head(pairs_all[, c("pair", "correlation", "p_value", "n",
"strength", "significance")], 15)Step 4: Strongest pair scatter — <=1000 sample
top <- pairs_all[1, ]
top_pair_names <- c(feature_names[[top$a]], feature_names[[top$b]])
set.seed(42)
sidx <- if (final_rows > 1000) sample(final_rows, 1000) else seq_len(final_rows)
top_pair_df <- data.frame(
value_a = X[[top$a]][sidx],
value_b = X[[top$b]][sidx],
stringsAsFactors = FALSE
)
top_pair_df <- top_pair_df[order(top_pair_df$value_a), ]
rownames(top_pair_df) <- NULL
top_r <- top$correlation
top_p <- top$p_value
metrics <- list(
`Observations` = final_rows,
`Columns Analysed` = length(used_features),
`Pairs Tested` = nrow(pairs_all),
`Significant Pairs` = as.integer(n_sig),
`Strongest |r|` = round(abs(top_r), 3),
`Strongest Pair` = top$pair
)
json_output <- list(
answer = paste0(
"Pairwise Pearson correlations on ", length(used_features),
" columns across ", format(final_rows, big.mark = ","), " rows: ",
"strongest pair is ", top$pair, " (r = ", top_r,
if (!is.na(top_p)) paste0(", p = ", top_p) else "", "); ",
n_sig, " of ", nrow(pairs_all), " pairs significant at p<0.05."
),
cards = lapply(
c("tldr", "overview", "preprocessing", "correlation_heatmap",
"strongest_pairs", "top_pair_scatter"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
initial_rows = initial_rows, final_rows = final_rows,
rows_removed = rows_removed,
feature_names = feature_names, used_features = used_features,
dropped_features = dropped_features,
cor_long_df = cor_long_df, pairs_df = pairs_df, pairs_all = pairs_all,
top_pair_df = top_pair_df, top_pair_names = top_pair_names,
top_r = top_r, top_p = top_p, n_sig = n_sig,
metrics = metrics, json_output = json_output
)
}