Standard Rfm
Executive Summary

Executive Summary

RFM segments across 4,338 customers

Customers
4338
Orders
18532
Segments
8
Biggest Segment
Champions
Top Segment Revenue %
64.8
At-Risk Revenue
921923.19
Across 4,338 customers and 18,532 orders, the biggest segment is Champions (23.7% of customers). Revenue is concentrated in Champions, which accounts for 64.8% of total OrderValue. Champions are 1,028 customers (23.7%) holding 64.8% of revenue. 921,923 of revenue (10.3%) sits with At Risk / Can't Lose Them customers — historically valuable buyers who have gone quiet; that is the revenue at stake if they are not won back.
What this means

The short answer

Champions dominate both customer count and revenue: 1,028 customers (23.7%) account for 64.8% of total OrderValue. A secondary priority is the 765 At Risk and Can't Lose Them customers holding 921,923 in revenue (10.3%)—historically valuable buyers now silent, representing the revenue at stake if not re-engaged.

The detail

Champions total 5,772,227 in OrderValue across 1,028 customers. At Risk customers (730) contribute 870,108, and Can't Lose Them (35) contribute 51,815, for a combined 921,923. At Risk has an average recency of 149.3 days; Can't Lose Them average 243.4 days—both are lapsed. Their high historical value (1,192 and 1,480 per customer, respectively) makes recovery high-leverage compared to acquiring new customers.

What this can't tell you

This snapshot captures the customer state as of 2011-12-09. Whether these At Risk and Can't Lose Them customers can realistically be re-engaged depends on factors outside this transaction data—competitive pressure, product changes, or customer lifecycle stage—that would require external context or a longitudinal follow-up to assess.

Overview

Analysis Overview

RFM segmentation of 4,338 customers from 18,532 orders.

N Customers4338
N Orders18532
N Segments8
Reference Date2011-12-09
What this means

The short answer

RFM segmentation divides your 4,338 customers into eight behavioural groups based on how recently they ordered, how often, and how much they spent. Each customer receives a quintile score (1–5) on each dimension, then maps into a segment: Champions (recent, frequent) down to Hibernating (lapsed, infrequent). This framework identifies where your revenue lives and which segments are at risk of churn.

The detail

The analysis scored 4,338 customers against 18,532 orders, with the reference date of 2011-12-09. Recency measures days since last order; Frequency counts the number of orders per customer; Monetary sums total OrderValue. Each customer is ranked 1–5 against the rest of the base on each metric, then placed into one of eight segments by their Recency and Frequency scores alone. The eight segments are Champions, Loyal, Potential Loyalist, New Customers, At Risk, Need Attention, Can't Lose Them, and Hibernating. This is the standard RFM segment map.

What this can't tell you

Segment membership is determined by R and F scores only; Monetary value is reported separately per segment but does not drive assignment. This approach reveals who is at risk but not why — consider order-level data (product category, channel, discount) to understand what drives lapse or loyalty within each segment.

Data Preparation

Data Quality

Rows, customers, date range, and cleaning decisions.

Initial Rows18532
Final Rows18532
Rows Removed0
N Customers4338
Unparseable Dates0
Negative Orders0
What this means

The short answer

All 18,532 orders from 4,338 customers loaded without loss or parsing errors. The order date range is 2010-12-01 to 2011-12-09, with no rows removed during quality checks. Negative order values (refunds) are included in the Monetary calculation, so revenue figures reflect net spend.

The detail

Initial and final row counts are both 18,532; zero rows were excluded. Zero dates were unparseable. The preprocessing preserved all negative OrderValues, meaning a customer's total monetary value is their cumulative spend minus any refunds. This approach is standard for RFM and ensures that net customer value is captured accurately.

What this can't tell you

The data does not distinguish between refunds due to returns, cancellations, or other causes. If refund patterns differ meaningfully across segments—for example, if At Risk customers have higher refund rates—that signal is absorbed into the Monetary score and not visible separately. A transaction-level export showing order status (completed, refunded, cancelled) would clarify whether segment behavior includes systematic differences in return propensity.

Visualization

Segment Sizes

Customer count per RFM segment.

What this means

The short answer

Champions lead the customer base at 1,028 customers (23.7%), followed by Hibernating at 962 (22.2%). Together, the top two segments account for about 46% of all customers. The remaining six segments range from 35 to 730 customers, with At Risk (730) and Need Attention (599) forming a substantial second tier.

The detail

Ranked by customer count: Champions 1,028; Hibernating 962; At Risk 730; Need Attention 599; Potential Loyalist 484; Loyal 264; New Customers 236; Can't Lose Them 35. The distribution is driven by Recency and Frequency quintile scores: a healthy base clusters toward Champions (recent, frequent) and away from Hibernating (lapsed, infrequent). Here, Hibernating is nearly as large as Champions, and At Risk represents 16.8% of the customer base, signalling a material retention challenge.

What this can't tell you

Segment membership reflects the current snapshot only. To assess whether the base is decaying or stabilizing, compare these sizes to a prior period — consider a historical export to track segment drift over time.

Data Table

Segment Value Breakdown

Size, share, revenue, and recency per segment.

SegmentCustomersShare PCTTotal ValueAvg ValueAvg Recency Days
Champions102823.75,772,2275,61512.2
At Risk73016.8870,1081,192149.3
Loyal2646.1742,2972,81249.1
Potential Loyalist48411.2586,2371,21115.9
Hibernating96222.2420,701437221.9
Need Attention59913.8391,12465351.6
New Customers2365.476,89832618.4
Can't Lose Them350.851,8151,480243.4
What this means

The short answer

Champions generate far more revenue than any other segment (5,772,227, 64.8% of total), with an average customer value of 5,615. At Risk is the second-largest revenue holder at 870,108 (16.8% of customers) but with a much lower average value per customer (1,192). The win-back priority is At Risk and Can't Lose Them: both are dormant (avg recency 149.3 and 243.4 days respectively) but retain high per-customer value.

The detail

Total revenue by segment: Champions 5,772,227 (23.7% of customers); At Risk 870,108 (16.8%); Loyal 742,297 (6.1%); Potential Loyalist 586,237 (11.2%); Hibernating 420,701 (22.2%); Need Attention 391,124 (13.8%); New Customers 76,898 (5.4%); Can't Lose Them 51,815 (0.8%). Average customer value: Champions 5,615; Loyal 2,812; Can't Lose Them 1,480; Potential Loyalist 1,211; At Risk 1,192; Need Attention 653; Hibernating 437; New Customers 326. Average recency (days): Champions 12.2; Potential Loyalist 15.9; New Customers 18.4; Loyal 49.1; Need Attention 51.6; At Risk 149.3; Hibernating 221.9; Can't Lose Them 243.4.

What this can't tell you

High average value does not guarantee recovery success — it only signals potential upside per customer. Segment-level response rates to reactivation campaigns are not provided; consider testing a small win-back cohort to measure conversion before scaling investment.

Visualization

Recency x Frequency Grid

Customer counts across all 25 R x F score combinations.

What this means

The short answer

The densest cell is R1×F1 (619 customers) — the most lapsed, lowest-frequency group. The top-right corner (R5×F5 = 439 customers) is your Champions core. High-frequency lapsed customers (left column, F4–F5) represent a concentrated win-back opportunity: 11 at R1×F5, 53 at R2×F5, and 121 at R3×F5 are frequent buyers who have gone quiet.

The detail

Of 25 possible R×F combinations, 20 are occupied. The R1×F1 cell (longest lapsed, fewest orders) is the largest at 619 customers, reflecting a heavy tail of one-time or very-low-frequency purchasers. The top-right quadrant (R4–R5 combined with F4–F5) clusters 852 customers — your active, engaged base. The left column (R1–R3 with F5) totals 185 customers: these are historically frequent buyers now dormant, the highest-ROI reactivation targets. The F2 column is entirely empty (0 customers), indicating no customers fall into the F2 quintile — the frequency distribution may be bimodal or skewed.

What this can't tell you

The grid shows distribution but not why lapse occurred. Overlaying order-date patterns (e.g., seasonal gaps, category shifts) would clarify whether dormancy is temporary seasonality or permanent churn.

Visualization

Average Customer Value by Segment

Mean per-customer monetary value in each segment.

What this means

The short answer

Champions customers are worth 5,615 on average—17.2x more than New Customers at 326. Loyal customers average 2,812. The two lapsed high-value segments, Can't Lose Them (1,480) and At Risk (1,192), are worth 4.5 and 3.7 times a typical New Customer, making each recovered customer far more valuable than acquiring a replacement.

The detail

Champions: 5,615.01 per customer. Loyal: 2,811.73. Can't Lose Them: 1,480.42. Potential Loyalist: 1,211.23. At Risk: 1,191.93. Need Attention: 652.96. Hibernating: 437.32. New Customers: 325.84. The ratio of Champions to New Customers is 5,615.01 ÷ 325.84 = 17.2. At Risk (1,191.93) ÷ New Customers (325.84) = 3.7. Can't Lose Them (1,480.42) ÷ New Customers (325.84) = 4.5.

What this can't tell you

These are mean values per segment; individual customer values within each segment will vary. A segment average does not predict whether a specific At Risk or Can't Lose Them customer can be re-engaged or at what cost. Lifetime value projections and win-back success rates would require historical re-engagement data or cohort analysis beyond this cross-sectional snapshot.

Rate this report Was this the answer you needed?
The exact source that produced this report — yours to keep, read, and re-run.
Download PDF
How this was computed method · R source · citation
The code that did it

RFM Customer Segmentation — Who Are Your Best Customers?

Segments customers by Recency, Frequency, and Monetary value computed from a raw transactions table (one row per order): quintile R/F/M scores, the standard industry segment map (Champions, At Risk, ...), and per-segment size, value, and recency statistics.

Why This Method?

RFM is the workhorse of customer-base analysis: three behavioral signals every business already records (when, how often, how much), scored against the customer base itself, so it needs no model fitting and works on any transaction log. The segment names carry the action: reward Champions, win back At Risk, welcome New Customers.

What This Analysis Covers

  • Per-customer Recency / Frequency / Monetary + quintile R/F/M scores
  • Standard segment map from R and F scores
  • Segment sizes, revenue concentration, and at-risk revenue
  • The full Recency x Frequency grid

Standard Library

Platform standard-library module (LAT-1441): runs on ANY transaction dataset via the semantic mapping {customer_id, order_date, order_value}. All narrative is derived from the user's own column names and computed values.

suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Core Analysis Pipeline

Multi-format date parser

Tries each candidate format across the whole vector, keeps the format that parses the most values, then back-fills stragglers per value from the remaining formats. Returns a Date vector (NA where unparseable).

parse_dates_multi <- function(x) {
  if (inherits(x, "Date")) return(x)
  if (inherits(x, "POSIXct") || inherits(x, "POSIXlt")) return(as.Date(x))
  s <- trimws(as.character(x))
  s[is.na(s) | s == ""] <- NA_character_
  fmts <- c("%Y-%m-%d", "%Y/%m/%d", "%m/%d/%Y", "%d/%m/%Y", "%m-%d-%Y",
            "%d-%m-%Y", "%d.%m.%Y", "%b %d, %Y", "%B %d, %Y",
            "%d %b %Y", "%d %B %Y")
  parsed <- lapply(fmts, function(f) suppressWarnings(as.Date(s, format = f)))
  counts <- vapply(parsed, function(p) sum(!is.na(p)), integer(1))
  best <- parsed[[which.max(counts)]]   # counts is never all-NA (integer counts)
  for (p in parsed[order(-counts)]) {
    fill <- is.na(best) & !is.na(p)
    if (any(fill)) best[fill] <- p[fill]
  }
  best
}

Quintile scorer (5 = best)

Primary path: quantile breaks with type = 1. When ties or too few distinct values make the breaks collapse, fall back to ntile-style ranking on average ranks so tied customers always share a score; a single distinct value collapses gracefully to the neutral score 3.

score_quintile <- function(x, higher_is_better = TRUE) {
  n <- length(x)
  ux <- unique(x)
  if (length(ux) <= 1) return(rep(3L, n))
  sc <- NULL
  if (length(ux) >= 5) {
    br <- quantile(x, probs = seq(0, 1, 0.2), type = 1, na.rm = TRUE)
    if (length(unique(br)) == 6) {
      sc <- as.integer(cut(x, breaks = br, include.lowest = TRUE, labels = FALSE))
    }
  }
  if (is.null(sc)) {
    r <- rank(x, ties.method = "average")
    sc <- as.integer(ceiling(r * 5 / n))
  }
  sc <- pmin(pmax(sc, 1L), 5L)
  if (!higher_is_better) sc <- 6L - sc
  sc
}

Segment map (standard industry names), from R and F scores.

Assigned lowest-priority first so later, more specific rules overwrite: "Can't Lose Them" (R=1, F>=4) takes precedence over the broader "At Risk" (R<=2, F>=3); Champions wins over Loyal.

assign_segments <- function(R, F_) {
  seg <- rep("Need Attention", length(R))
  seg[R <= 2 & F_ <= 2] <- "Hibernating"
  seg[R <= 2 & F_ >= 3] <- "At Risk"
  seg[R == 1 & F_ >= 4] <- "Can&#x27;t Lose Them"
  seg[R >= 4 & F_ == 1] <- "New Customers"
  seg[R >= 4 & F_ %in% 2:3] <- "Potential Loyalist"
  seg[R >= 3 & F_ >= 4] <- "Loyal"
  seg[R >= 4 & F_ >= 4] <- "Champions"
  seg
}

SEGMENT_ORDER <- c("Champions", "Loyal", "Potential Loyalist", "New Customers",
                   "At Risk", "Can&#x27;t Lose Them", "Need Attention", "Hibernating")

compute_shared <- function(df, params, col_map = list()) {
  # === SHARED EXPORTS ===
  #   initial_rows/final_rows/rows_removed   $ row accounting
  #   n_missing_id / n_bad_dates / n_missing_value / n_negative  $ preprocessing counts
  #   negative_total    $ numeric — summed value of negative (refund) orders
  #   id_name/date_name/value_name  $ humanized user column names
  #   ref_date / date_min / date_max $ Date — reference + observed range
  #   customer_df       $ data.frame(customer, recency_days, frequency,
  #                       monetary, recency_score, frequency_score,
  #                       monetary_score, segment) — one row per customer
  #   seg_stats         $ data.frame(segment, customers, share_pct,
  #                       total_value, avg_value, avg_recency_days)
  #   segment_sizes_df / segment_value_df / rf_grid_df / monetary_by_segment_df
  #   n_customers       $ integer
  #   total_revenue     $ numeric — sum of all order values
  #   biggest_segment / biggest_share          $ by customer count
  #   top_rev_segment / top_rev_share          $ revenue concentration
  #   at_risk_revenue / at_risk_rev_share      $ At Risk + Can't Lose Them
  #   metrics / json_output
  # === /SHARED EXPORTS ===

Step 1: Required semantic columns + humanized names

initial_rows <- nrow(df)
  id_name    <- humanize_semantic("customer_id", col_map)[1]
  date_name  <- humanize_semantic("order_date", col_map)[1]
  value_name <- humanize_semantic("order_value", col_map)[1]
  for (need in c("customer_id", "order_date", "order_value")) {
    if (!need %in% names(df)) {
      stop(sprintf("Required column &#x27;%s' is not mapped.",
                   humanize_semantic(need, col_map)[1]))
    }
  }

Step 2: Customer id — character, drop blank/missing

cid <- trimws(as.character(df$customer_id))
  keep_id <- !is.na(cid) & cid != ""
  n_missing_id <- sum(!keep_id)
  df <- df[keep_id, , drop = FALSE]
  cid <- cid[keep_id]

Step 3: Dates — multi-format parse; stop (humanized) if >5% unparseable

dts <- parse_dates_multi(df$order_date)
  n_bad_dates <- sum(is.na(dts))
  if (nrow(df) > 0 && n_bad_dates / nrow(df) > 0.05) {
    stop(sprintf(
      paste0("Could not read %s of %s values in &#x27;%s' as dates (%.1f%%). ",
             "More than 5%% of the order dates are unparseable — please ",
             "check the date format of that column."),
      format(n_bad_dates, big.mark = ","), format(nrow(df), big.mark = ","),
      date_name, 100 * n_bad_dates / nrow(df)))
  }
  keep_dt <- !is.na(dts)
  df <- df[keep_dt, , drop = FALSE]
  cid <- cid[keep_dt]
  dts <- dts[keep_dt]

Step 4: Order value — 95% numeric coercion rule; drop NA values;

KEEP negatives (refunds) but count them for the narrative.

val <- df$order_value
  if (!is.numeric(val)) {
    conv <- suppressWarnings(as.numeric(as.character(val)))
    n_orig <- sum(!is.na(val) & trimws(as.character(val)) != "")
    if (n_orig == 0 || sum(!is.na(conv)) < 0.95 * n_orig) {
      stop(sprintf(
        "Column &#x27;%s' is not numeric enough to use as an order amount — fewer than 95%% of its values could be read as numbers.",
        value_name))
    }
    val <- conv
  }
  keep_val <- !is.na(val)
  n_missing_value <- sum(!keep_val)
  cid <- cid[keep_val]
  dts <- dts[keep_val]
  val <- val[keep_val]
  n_negative <- sum(val < 0)
  negative_total <- if (n_negative > 0) sum(val[val < 0]) else 0

  final_rows <- length(val)
  rows_removed <- initial_rows - final_rows
  if (final_rows < 30) {
    stop(sprintf(
      "Only %d usable orders remain after cleaning &#x27;%s', '%s', and '%s' — need at least 30.",
      final_rows, id_name, date_name, value_name))
  }

Step 5: Reference date + per-customer R/F/M

ref_date <- max(dts)
  date_min <- min(dts)
  date_max <- ref_date

  tx <- data.frame(customer = cid, order_date = dts, order_value = val,
                   stringsAsFactors = FALSE)
  customer_df <- tx %>%
    group_by(customer) %>%
    summarise(
      recency_days = as.numeric(ref_date - max(order_date)),
      frequency    = n(),
      monetary     = sum(order_value),
      .groups = "drop"
    ) %>%
    as.data.frame(stringsAsFactors = FALSE)

  n_customers <- nrow(customer_df)
  if (n_customers < 5) {
    stop(sprintf(
      "Only %d distinct customers found in &#x27;%s' — RFM segmentation needs at least 5.",
      n_customers, id_name))
  }

Step 6: Quintile scores (5 = best) + segment map

customer_df$recency_score   <- score_quintile(customer_df$recency_days,
                                                higher_is_better = FALSE)
  customer_df$frequency_score <- score_quintile(customer_df$frequency)
  customer_df$monetary_score  <- score_quintile(customer_df$monetary)
  customer_df$segment <- assign_segments(customer_df$recency_score,
                                         customer_df$frequency_score)

Step 7: Segment statistics

total_revenue <- sum(customer_df$monetary)
  seg_stats <- customer_df %>%
    group_by(segment) %>%
    summarise(
      customers        = n(),
      total_value      = round(sum(monetary), 2),
      avg_value        = round(mean(monetary), 2),
      avg_recency_days = round(mean(recency_days), 1),
      .groups = "drop"
    ) %>%
    mutate(share_pct = round(100 * customers / n_customers, 1)) %>%
    as.data.frame(stringsAsFactors = FALSE)
  seg_stats$segment <- factor(seg_stats$segment, levels = SEGMENT_ORDER)
  seg_stats <- seg_stats[order(seg_stats$segment), , drop = FALSE]
  seg_stats$segment <- as.character(seg_stats$segment)
  rownames(seg_stats) <- NULL

Headline facts (all computed; guard against all-NA/max on empty)

by_size <- seg_stats[order(-seg_stats$customers), , drop = FALSE]
  biggest_segment <- by_size$segment[1]
  biggest_share   <- by_size$share_pct[1]
  by_rev <- seg_stats[order(-seg_stats$total_value), , drop = FALSE]
  top_rev_segment <- by_rev$segment[1]
  top_rev_share   <- if (total_revenue != 0)
    round(100 * by_rev$total_value[1] / total_revenue, 1) else NA_real_
  at_risk_segments <- c("At Risk", "Can&#x27;t Lose Them")
  at_risk_revenue  <- round(sum(seg_stats$total_value[
    seg_stats$segment %in% at_risk_segments]), 2)
  at_risk_rev_share <- if (total_revenue != 0)
    round(100 * at_risk_revenue / total_revenue, 1) else NA_real_

Format value columns as plain rounded strings — large numerics would otherwise serialize in scientific notation (e.g. 1.922e+05) in the table.

segment_value_df$total_value <- format(round(segment_value_df$total_value, 0),
                                         big.mark = ",", scientific = FALSE,
                                         trim = TRUE)
  segment_value_df$avg_value <- format(round(segment_value_df$avg_value, 0),
                                       big.mark = ",", scientific = FALSE,
                                       trim = TRUE)
  rownames(segment_value_df) <- NULL

  grid <- expand.grid(r = 1:5, f = 1:5)
  rf_grid_df <- data.frame(
    recency_score   = paste0("R", grid$r),
    frequency_score = paste0("F", grid$f),
    customers = mapply(function(r, f) {
      sum(customer_df$recency_score == r & customer_df$frequency_score == f)
    }, grid$r, grid$f),
    stringsAsFactors = FALSE
  )

  monetary_by_segment_df <- seg_stats[, c("segment", "avg_value")]
  names(monetary_by_segment_df) <- c("segment", "avg_customer_value")
  monetary_by_segment_df <- monetary_by_segment_df[
    order(-monetary_by_segment_df$avg_customer_value), , drop = FALSE]
  rownames(monetary_by_segment_df) <- NULL

  metrics <- list(
    `Customers`             = n_customers,
    `Orders`                = final_rows,
    `Segments`              = nrow(seg_stats),
    `Biggest Segment`       = biggest_segment,
    `Top Segment Revenue %` = top_rev_share,
    `At-Risk Revenue`       = at_risk_revenue
  )

  json_output <- list(
    answer = paste0(
      "RFM segmentation of ", format(n_customers, big.mark = ","),
      " customers across ", format(final_rows, big.mark = ","),
      " orders(", format(date_min), " to ", format(date_max), "): ",
      "the biggest segment is ", biggest_segment, " (", biggest_share,
      "% of customers); ", top_rev_segment, " concentrates ", top_rev_share,
      "% of total revenue; ",
      format(round(at_risk_revenue), big.mark = ","), " (",
      at_risk_rev_share, "% of revenue) sits with At Risk / Can&#x27;t Lose Them ",
      "customers and is the immediate win-back opportunity."
    ),
    cards = lapply(
      c("tldr", "overview", "preprocessing", "segment_sizes",
        "segment_value", "rf_grid", "monetary_by_segment"),
      function(cid) list(id = cid, metrics = metrics)
    )
  )

  list(
    initial_rows = initial_rows, final_rows = final_rows,
    rows_removed = rows_removed,
    n_missing_id = n_missing_id, n_bad_dates = n_bad_dates,
    n_missing_value = n_missing_value,
    n_negative = n_negative, negative_total = negative_total,
    id_name = id_name, date_name = date_name, value_name = value_name,
    ref_date = ref_date, date_min = date_min, date_max = date_max,
    customer_df = customer_df, seg_stats = seg_stats,
    segment_sizes_df = segment_sizes_df,
    segment_value_df = segment_value_df,
    rf_grid_df = rf_grid_df,
    monetary_by_segment_df = monetary_by_segment_df,
    n_customers = n_customers, total_revenue = total_revenue,
    biggest_segment = biggest_segment, biggest_share = biggest_share,
    top_rev_segment = top_rev_segment, top_rev_share = top_rev_share,
    at_risk_revenue = at_risk_revenue, at_risk_rev_share = at_risk_rev_share,
    metrics = metrics, json_output = json_output
  )
}
Your data has more stories to tell.Run any analysis on your own data — validated R modules, interactive reports, AI insights, and PDF export. 500 free credits on signup.
Try Free — No SignupSign Up Free

Cite this analysis

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing