Standard Correlation
Executive Summary

Executive Summary

Relationship map across 12 columns

Observations
1599
Columns Analysed
12
Pairs Tested
66
Significant Pairs
55
Strongest |r|
0.683
Strongest Pair
fixed acidity ~ pH
Across 1,599 observations and 66 column pairs, the data is densely interconnected — most metrics move together. The strongest relationship is fixed acidity ~ pH, which move negatively together (r = -0.683, p = 4.06e-220). 55 pair(s) are statistically significant at p < 0.05.
What this means

The short answer

The 12 chemical properties form a densely interconnected web: 55 of 66 pairs show statistically significant associations at p < 0.05. The strongest relationship is fixed acidity and pH, which move in opposite directions (r = −0.683, p = 4.06e-220).

The detail

Across 1,599 observations and 66 column pairs tested, fixed acidity ~ pH leads with r = −0.683 and p = 4.06e-220, marking a highly reliable negative association. No pair reaches r ≥ 0.7 in absolute value. All top relationships carry p < 0.001, confirming they exceed chance.

What this can't tell you

Significance at p < 0.05 means the pattern is real; it does not indicate practical importance or how tightly the relationship holds in individual cases. Moderate correlations (r ≈ 0.67) still leave substantial unexplained variance.

Overview

Analysis Overview

Pairwise Pearson correlations across 12 columns and 1,599 observations.

N Observations1599
N Columns12
N Pairs66
N Significant55
What this means

The short answer

Pearson correlation measures how two variables move together on a scale from −1 (perfect opposition) to +1 (perfect co-movement). This analysis tests all 66 possible pairs among 12 chemical properties and quality, separating real associations from chance variation.

The detail

The analysis computed pairwise Pearson r for 1,599 observations across fixed acidity, volatile acidity, citric acid, residual sugar, chlorides, free sulfur dioxide, total sulfur dioxide, density, pH, sulphates, alcohol, and quality. A significance test per pair flags relationships where p < 0.05. Correlation measures association only; it does not imply one variable drives another.

What this can't tell you

Pearson r assumes a linear relationship. If the true pattern is curved or if outliers dominate, r may misrepresent the strength of association; the scatter plot for the strongest pair tests this assumption.

Data Preparation

Data Quality

Column typing, imputation, and exclusions.

Initial Rows1599
Final Rows1599
Rows Removed0
What this means

The short answer

All 1,599 observations were retained and all 12 mapped columns were usable. Missing values were imputed with each column's median, and pairwise complete observations were used for each correlation.

The detail

Initial rows: 1,599; final rows: 1,599; rows removed: 0. No columns were excluded as non-numeric or constant. Median imputation filled gaps, preserving the full dataset for analysis.

What this can't tell you

The preprocessing report does not disclose how many missing values existed before imputation. If missingness was substantial and non-random, median imputation may have introduced bias. Consider requesting a missing-data summary to assess whether re-analysis with listwise deletion or multiple imputation would change the findings.

Visualization

Correlation Matrix

Every pairwise correlation across the mapped columns.

What this means

The short answer

The heatmap reveals blocks of moderate correlation rather than strong redundancy. Fixed acidity clusters moderately with citric acid (r = 0.672) and density (r = 0.668), and moves opposite to pH (r = −0.683). No pair achieves r ≥ 0.7, so no two properties are near-redundant.

The detail

The color scale ranges from deep blue (r = −1, perfect opposition) to deep red (r = +1, perfect co-movement). The hottest cells are fixed acidity ~ pH at r = −0.683, fixed acidity ~ citric acid at r = 0.672, and fixed acidity ~ density at r = 0.668. Row-wise, fixed acidity shows the strongest gradients. The coldest (most negative) cells include volatile acidity ~ citric acid (r = −0.552) and citric acid ~ pH (r = −0.542). 0 pairs show strong positive correlation (r ≥ 0.7).

What this can't tell you

A heatmap displays pairwise correlations only; it does not reveal multivariate patterns or whether one variable mediates the relationship between two others. Partial correlation (controlling for a third variable) would clarify whether associations are direct or indirect.

Data Table

Strongest Relationships

Top column pairs ranked by absolute correlation, with significance.

PairCorrelationP ValueNStrengthSignificance
fixed acidity ~ pH-0.6834.06e-2201599moderate***
fixed acidity ~ citric acid0.6722.54e-2101599moderate***
fixed acidity ~ density0.6683.07e-2071599moderate***
free sulfur dioxide ~ total sulfur dioxide0.6686.4e-2071599moderate***
volatile acidity ~ citric acid-0.5521.81e-1281599moderate***
citric acid ~ pH-0.5421.01e-1221599moderate***
density ~ alcohol-0.4963.94e-1001599moderate***
alcohol ~ quality0.4762.83e-911599moderate***
volatile acidity ~ quality-0.3912.05e-591599weak***
chlorides ~ sulphates0.3711.99e-531599weak***
citric acid ~ density0.3651.48e-511599weak***
residual sugar ~ density0.3559.01e-491599weak***
density ~ pH-0.3425.12e-451599weak***
citric acid ~ sulphates0.3131.27e-371599weak***
chlorides ~ pH-0.2654.15e-271599weak***
What this means

The short answer

Two pairs tie for second-strongest: fixed acidity ~ citric acid (r = 0.672) and fixed acidity ~ density (r = 0.668), both highly significant. Among pairs involving quality, alcohol shows the strongest link (r = 0.476, p = 2.83e-91), followed by volatile acidity (r = −0.391, p = 2.05e-59).

The detail

The top five by absolute correlation are: fixed acidity ~ pH (r = −0.683, p = 4.06e-220), fixed acidity ~ citric acid (r = 0.672, p = 2.54e-210), fixed acidity ~ density (r = 0.668, p = 3.07e-207), free sulfur dioxide ~ total sulfur dioxide (r = 0.668, p = 6.4e-207), and volatile acidity ~ citric acid (r = −0.552, p = 1.81e-128). All carry p < 0.001 and n = 1,599. None reach |r| ≥ 0.7. For quality, alcohol (r = 0.476) and volatile acidity (r = −0.391) dominate; all other quality pairs fall below |r| = 0.3.

What this can't tell you

Moderate correlations with quality do not establish which chemical properties are most important for predicting or improving quality. Multivariate methods (regression, feature importance) would rank predictive power. The strongest pairs among chemical properties may reflect measurement overlap rather than independent drivers of quality.

Visualization

Strongest Pair

The strongest relationship plotted point by point.

What this means

The short answer

Fixed acidity and pH form a tight downward-sloping band, confirming the linear negative relationship (r = −0.683). Points cluster around the trend with modest scatter; no extreme outliers visibly dominate the correlation.

The detail

The scatter shows fixed acidity (horizontal) against pH (vertical) for 1,599 observations. As fixed acidity increases from ~4.6 to ~14, pH decreases from ~4 to ~3, tracing a clear negative linear pattern. The cloud is narrow enough to confirm that r = −0.683 accurately describes the relationship; curvature or clusters are not apparent. Scatter around the line is consistent across the range, suggesting homogeneous variance.

What this can't tell you

The scatter plot shows association, not causation. Tightness of the band confirms the linear model fits; it does not explain why acidity and pH move together (they may both reflect fermentation chemistry or be mechanically linked by the buffer system). A residual plot would reveal whether the linear model systematically mispredicts at any acidity level.

Rate this report Was this the answer you needed?
The exact source that produced this report — yours to keep, read, and re-run.
Download PDF
How this was computed method · R source · citation
The code that did it

Correlation Analysis — What Moves Together

Computes pairwise Pearson correlations across the numeric columns the user selects: full matrix heatmap, strongest relationships ranked with significance tests, and a scatter of the dominant pair.

Why This Method?

Correlation is the fastest map of a dataset's relationships: one number per pair, directly comparable, with a significance test separating real co-movement from noise. It is the standard first step before regression, forecasting, or KPI pruning.

What This Analysis Covers

  • Full correlation matrix (heatmap)
  • Strongest relationships ranked with p-values
  • The dominant pair plotted point by point

Standard Library

Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {feature_1..feature_N}. All narrative is derived from the user's own column names and computed values.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Core Analysis Pipeline

compute_shared <- function(df, params, col_map = list()) {
  # === SHARED EXPORTS ===
  #   initial_rows/final_rows/rows_removed  $ row accounting
  #   feature_names     $ named character — semantic -> humanized names
  #   used_features     $ character — semantic names used
  #   dropped_features  $ character — excluded columns
  #   cor_long_df       $ data.frame(feature_a, feature_b, correlation) — full matrix
  #   pairs_df          $ data.frame(pair, correlation, p_value, n, strength, significance)
  #   top_pair_df       $ data.frame(value_a, value_b) — <=1000 sample of strongest pair
  #   top_pair_names    $ character(2) — humanized names of the strongest pair
  #   top_r / top_p     $ numeric — strongest pair stats
  #   n_sig             $ integer — pairs significant at p<0.05
  #   metrics / json_output
  # === /SHARED EXPORTS ===

Step 1: Discover mapped features

initial_rows <- nrow(df)
  feat_cols <- grep("^feature_[0-9]+$", names(df), value = TRUE)
  feat_cols <- feat_cols[order(as.integer(sub("^feature_", "", feat_cols)))]
  if (length(feat_cols) < 2) {
    stop("column_mapping must map at least two feature columns(feature_1, feature_2)")
  }
  feature_names <- setNames(humanize_semantic(feat_cols, col_map), feat_cols)

Step 2: Coerce numeric (95% rule); impute median; drop unusable

dropped_features <- character(0)
  for (fc in feat_cols) {
    v <- df[[fc]]
    if (!is.numeric(v)) {
      conv <- suppressWarnings(as.numeric(as.character(v)))
      n_orig <- sum(!is.na(v) & as.character(v) != "")
      if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) {
        df[[fc]] <- conv
      } else {
        dropped_features <- c(dropped_features, fc); next
      }
    }
    v <- df[[fc]]
    med <- median(v, na.rm = TRUE)
    if (is.na(med)) { dropped_features <- c(dropped_features, fc); next }
    v[is.na(v)] <- med
    df[[fc]] <- v
    if (isTRUE(var(v) == 0) || is.na(var(v))) {
      dropped_features <- c(dropped_features, fc)
    }
  }
  used_features <- setdiff(feat_cols, dropped_features)
  if (length(used_features) < 2) {
    stop(paste0("Correlation needs at least two usable numeric columns; only ",
                length(used_features), " remained after cleaning."))
  }
  X <- df[, used_features, drop = FALSE]
  final_rows <- nrow(X)
  rows_removed <- initial_rows - final_rows
  if (final_rows < 10) stop(sprintf("Only %d usable rows — need at least 10.", final_rows))

Step 3: Correlation matrix + per-pair tests

C <- suppressWarnings(cor(X, use = "pairwise.complete.obs", method = "pearson"))
  hn <- unname(feature_names[used_features])

Long format for the heatmap (diagonal included — reads as the standard matrix)

k <- length(used_features)
  cor_long_df <- do.call(rbind, lapply(seq_len(k), function(i) {
    data.frame(feature_a = hn[i], feature_b = hn,
               correlation = round(as.numeric(C[i, ]), 3),
               stringsAsFactors = FALSE)
  }))
  rownames(cor_long_df) <- NULL

Pairwise tests (upper triangle only)

pair_rows <- list()
  for (i in seq_len(k - 1)) {
    for (j in (i + 1):k) {
      r_ij <- C[i, j]
      if (is.na(r_ij)) next
      ct <- tryCatch(cor.test(X[[i]], X[[j]]), error = function(e) NULL)
      p_ij <- if (!is.null(ct)) ct$p.value else NA_real_
      pair_rows[[length(pair_rows) + 1]] <- data.frame(
        pair = paste0(hn[i], " ~ ", hn[j]),
        a = used_features[i], b = used_features[j],
        correlation = round(r_ij, 3),
        p_value = signif(p_ij, 3),
        n = sum(complete.cases(X[[i]], X[[j]])),
        stringsAsFactors = FALSE
      )
    }
  }
  pairs_all <- do.call(rbind, pair_rows)
  if (is.null(pairs_all) || nrow(pairs_all) == 0) {
    stop("No valid correlation pairs could be computed from the mapped columns.")
  }
  pairs_all <- pairs_all[order(-abs(pairs_all$correlation)), , drop = FALSE]
  rownames(pairs_all) <- NULL
  pairs_all$strength <- ifelse(abs(pairs_all$correlation) >= 0.7, "strong",
                        ifelse(abs(pairs_all$correlation) >= 0.4, "moderate", "weak"))
  pairs_all$significance <- ifelse(is.na(pairs_all$p_value), "",
                            ifelse(pairs_all$p_value < 0.001, "***",
                            ifelse(pairs_all$p_value < 0.01, "**",
                            ifelse(pairs_all$p_value < 0.05, "*", ""))))
  n_sig <- sum(pairs_all$p_value < 0.05, na.rm = TRUE)
  pairs_df <- head(pairs_all[, c("pair", "correlation", "p_value", "n",
                                 "strength", "significance")], 15)

Step 4: Strongest pair scatter — <=1000 sample

top <- pairs_all[1, ]
  top_pair_names <- c(feature_names[[top$a]], feature_names[[top$b]])
  set.seed(42)
  sidx <- if (final_rows > 1000) sample(final_rows, 1000) else seq_len(final_rows)
  top_pair_df <- data.frame(
    value_a = X[[top$a]][sidx],
    value_b = X[[top$b]][sidx],
    stringsAsFactors = FALSE
  )
  top_pair_df <- top_pair_df[order(top_pair_df$value_a), ]
  rownames(top_pair_df) <- NULL
  top_r <- top$correlation
  top_p <- top$p_value

  metrics <- list(
    `Observations`        = final_rows,
    `Columns Analysed`    = length(used_features),
    `Pairs Tested`        = nrow(pairs_all),
    `Significant Pairs`   = as.integer(n_sig),
    `Strongest |r|`       = round(abs(top_r), 3),
    `Strongest Pair`      = top$pair
  )

  json_output <- list(
    answer = paste0(
      "Pairwise Pearson correlations on ", length(used_features),
      " columns across ", format(final_rows, big.mark = ","), " rows: ",
      "strongest pair is ", top$pair, " (r = ", top_r,
      if (!is.na(top_p)) paste0(", p = ", top_p) else "", "); ",
      n_sig, " of ", nrow(pairs_all), " pairs significant at p<0.05."
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "correlation_heatmap",
        "strongest_pairs", "top_pair_scatter"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  list(
    initial_rows = initial_rows, final_rows = final_rows,
    rows_removed = rows_removed,
    feature_names = feature_names, used_features = used_features,
    dropped_features = dropped_features,
    cor_long_df = cor_long_df, pairs_df = pairs_df, pairs_all = pairs_all,
    top_pair_df = top_pair_df, top_pair_names = top_pair_names,
    top_r = top_r, top_p = top_p, n_sig = n_sig,
    metrics = metrics, json_output = json_output
  )
}
Your data has more stories to tell.Run any analysis on your own data — validated R modules, interactive reports, AI insights, and PDF export. 500 free credits on signup.
Try Free — No SignupSign Up Free

Cite this analysis

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing