Executive Summary
Rater reliability across 182 subject values and 2 raters
The short answer
Self-reported weight and scale measurement agree at ICC(2,1) = 0.986—excellent reliability (Koo & Li band) for a single rater across 182 subjects. The F-test (F = 136.59, p = 1.37e-141) confirms raters reliably distinguish subjects from one another.
The detail
ICC(2,1) = 0.986 quantifies absolute agreement between the two measurement methods. The F-statistic of 136.59 with p = 1.37e-141 is far below any conventional threshold, establishing that subject-level variance dominates measurement noise. For reporting: use ICC(2,1) = 0.986 if raters are treated as a random sample and absolute weight kg values matter (standard default); ICC(3,1) = 0.985 if these exact two raters are the only ones of interest; ICC(2,k) = 0.993 if conclusions rest on the average of both raters rather than either alone.
What this can't tell you
The analysis does not distinguish between random measurement error and systematic bias in either method. A Bland-Altman plot or repeated measurement under identical conditions would clarify whether disagreement is noise or a consistent offset.
Analysis Overview
Intraclass correlation from 364 ratings: 182 subject values x 2 raters.
The short answer
Intraclass correlation quantifies how much of the variation in weight measurements comes from real differences between people rather than disagreement between the two measurement methods. The key choice is whether to report reliability for a single rater or for the average of both; here, a single rater reaches 0.986, while the average reaches 0.993.
The detail
The analysis uses two-way ANOVA to compute ICC from 364 ratings across 182 subjects and 2 raters (scale-measured and self-reported). ICC(2,1) = 0.986 is the standard report for a random rater panel where absolute weight values matter. ICC(3,1) = 0.985 applies when only these exact two raters are of interest. The gap between them (−0.000) reflects systematic rater bias; here it is negligible. ICC(2,k) = 0.993 is the reliability of the average of both raters, always higher than the single-rater forms because averaging reduces noise.
What this can't tell you
This analysis does not address whether 0.986 is adequate for your specific use case—that is a practical tolerance decision. Consider whether measurement error of this magnitude would affect decisions downstream.
Data Quality
Ratings used, duplicates averaged, incomplete subjects excluded.
The short answer
All 364 ratings were retained and properly structured: 182 subjects, each rated by both raters, with no duplicate ratings and no missing cells. The data is complete and ready for two-way ICC analysis.
The detail
Initial load: 364 rows. Final analysis: 364 rows used. Rows removed: 0. Duplicate ratings found: 0. Incomplete subjects (not rated by all raters): 0. Every subject value was rated by all 2 raters, forming a complete crossing required for two-way ANOVA. One row per rating with consistent subject identifiers across raters.
What this can't tell you
This export does not show whether the same person's weight was measured on the same day or across different occasions, or whether the self-reported value was collected before or after the scale measurement. Temporal separation could affect apparent agreement if weight changed between measurements.
ICC Results — Which One to Report
All ICC forms with interpretation bands and when to use each.
| Icc Type | Value | Interpretation | Use When |
|---|---|---|---|
| ICC(1,1) — one-way random, single rater | 0.986 | excellent | Each subject is rated by a DIFFERENT random set of raters (no shared panel). |
| ICC(2,1) — absolute agreement, single rater | 0.986 | excellent | Raters are a random sample and absolute weight kg values must match — the default to report. |
| ICC(3,1) — consistency, single rater | 0.985 | excellent | These exact 2 raters are the only ones of interest; systematic differences between them are forgiven. |
| ICC(2,k) — absolute agreement, average of all raters | 0.993 | excellent | Decisions rest on the AVERAGE weight kg of all 2 raters, and absolute values matter. |
| ICC(3,k) — consistency, average of all raters | 0.993 | excellent | Decisions rest on the average of these exact 2 raters; only relative ordering matters. |
The short answer
All ICC forms exceed 0.985, placing agreement in the excellent range. The single-rater forms (ICC(2,1) and ICC(3,1)) are nearly identical at 0.986 and 0.985, showing negligible systematic bias between scale-measured and self-reported. The average-measure forms (ICC(2,k) and ICC(3,k)) both reach 0.993.
The detail
ICC(2,1) = 0.986 (excellent) is the headline: absolute agreement for a single rater, assuming the two methods are a random sample. ICC(3,1) = 0.985 (excellent) applies when only these exact two methods are of interest; the gap of 0 between ICC(2,1) and ICC(3,1) means systematic rater differences are negligible. ICC(2,k) = 0.993 (excellent) and ICC(3,k) = 0.993 (excellent) reflect the average of both methods. ICC(1,1) = 0.986 is included for reference but applies only to designs where each subject is rated by a different set of raters.
What this can't tell you
High ICC does not guarantee that individual predictions are precise enough for clinical or commercial thresholds. Consider whether the residual scatter around the regression line meets your tolerance.
Systematic Rater Bias
Mean weight kg per rater on the same subject values.
The short answer
Both raters (scale-measured and self-reported) report the same average weight: 65.68 kg. There is no systematic bias—neither method is consistently higher or lower than the other.
The detail
Scale-measured mean: 65.68 kg. Self-reported mean: 65.68 kg. Spread: 0 kg. The two methods have identical means across the 182 subjects, meaning there is no leniency or severity bias to correct. This perfect alignment is consistent with the negligible gap between ICC(2,1) and ICC(3,1) observed in the ICC table.
What this can't tell you
Identical means do not rule out subject-level disagreement. Some individuals may report their weight higher or lower than the scale measures, even though the panel averages match. The scatter plot in the subject-agreement card reveals those individual discrepancies.
Subject-Level Agreement
One point per subject: the first two raters' scores, diagonal = perfect agreement.
The short answer
The two methods form a tight upward band hugging the diagonal: subjects where scale-measured and self-reported agree exactly sit on the line, while scattered points slightly off the diagonal show small individual disagreements. The correlation is r = 0.986, indicating nearly perfect ranking agreement.
The detail
Points cluster tightly around the perfect-agreement diagonal across the observed range (39–59 kg). Most subjects show agreement within 1–2 kg; scattered outliers include a subject at (39, 41), another at (50, 55), and (54, 59), where self-reported exceeds scale-measured by 5 kg. The r = 0.986 correlation confirms that when scale-measured is high, self-reported is also high, preserving rank order. The minimal offset parallel to the diagonal explains why ICC(3,1) and ICC(2,1) are nearly identical.
What this can't tell you
This scatter does not distinguish between measurement error and true weight change. If self-reported was collected days or weeks after scale measurement, disagreement could reflect actual weight gain or loss rather than measurement unreliability.
Where the Disagreement Comes From
Variance in weight kg split between subjects, raters, and noise.
The short answer
Of the total variance in weight kg, 98.5% comes from real differences between subjects, 0% from systematic rater bias, and 1.5% from residual disagreement. This dominance of between-subject signal is why reliability is so high.
The detail
Variance decomposition from two-way ANOVA shows between-subject variance at 98.5%—the signal raters are supposed to detect. Between-rater systematic variance is 0%, meaning neither self-report nor scale measurement is consistently offset relative to the other. Residual variance (one-off disagreements) is 1.5%. Reliability rises when the first bar dominates, which it does here. To improve further: calibration sessions would target the rater bar (currently zero, so no gain available there); clearer measurement protocol or repeated assessments would compress the residual bar from 1.5%.
What this can't tell you
Variance decomposition describes consistency but not accuracy. A high between-subject percentage confirms raters agree on who is heavier or lighter, not whether reported or scale values are correct in absolute terms.
Intraclass Correlation — Rater Reliability
How consistent are your raters, instruments, or repeated measurements? From long-format ratings (one row per rating) this computes the full ICC family — ICC(1,1), ICC(2,1), ICC(3,1) plus the average-measure versions — from two-way ANOVA mean squares, with Koo & Li interpretation bands, a rater-bias check, a subject-level agreement plot, and a variance breakdown showing where the disagreement comes from.
Why This Method?
The ICC is the standard reliability coefficient for continuous ratings: it asks how much of the total variation comes from real differences between the things being rated rather than from rater disagreement. Computing every common form at once answers the perennial question of WHICH ICC to report for a given study design.
What This Analysis Covers
- Single-measure and average-measure ICCs, consistency and agreement
- Koo & Li (2016) interpretation bands
- Systematic rater bias (mean score per rater)
- Subject-level agreement between the first two raters
- Variance decomposition: subjects vs raters vs residual noise
Standard Library
Platform standard-library module (LAT-1441): runs on ANY dataset via the semantic mapping {subject, rater, score}. All narrative is derived from the user's own column names and computed values.
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))Core Analysis Pipeline
compute_shared <- function(df, params, col_map = list()) {
# === SHARED EXPORTS ===
# initial_rows/final_rows/rows_removed $ row accounting (ratings)
# subj_h / rater_h / score_h $ humanized names of the three mapped columns
# n_subjects / k_raters $ complete-crossing design dimensions
# rater_levels $ character — the actual rater level values
# n_duplicates $ duplicate ratings averaged
# n_incomplete $ subjects excluded (not rated by all raters)
# incomplete_subjects $ character — the excluded subject ids
# icc / icc_p $ named list of ICC values + F-test p for ICC(3,1)
# icc_results_df $ data.frame(icc_type, value, interpretation, use_when)
# rater_means_df $ data.frame(rater_name, mean_score)
# agreement_df $ data.frame(rater_1_score, rater_2_score) — complete subjects
# agreement_r $ numeric — Pearson r between the first two raters
# variance_df $ data.frame(source, variance_pct)
# headline_band $ Koo & Li band for ICC(2,1)
# metrics / json_output
# === /SHARED EXPORTS ===
subj_h <- humanize_semantic("subject", col_map)
rater_h <- humanize_semantic("rater", col_map)
score_h <- humanize_semantic("score", col_map)Step 1: Validate the mapped columns
initial_rows <- nrow(df)
for (req in c("subject", "rater", "score")) {
if (is.null(df[[req]])) {
stop(sprintf("column_mapping must map the '%s' column (%s / %s / %s).",
req, subj_h, rater_h, score_h))
}
}Step 2: Coerce score numeric (95% rule); clean subject/rater ids
v <- df$score
if (!is.numeric(v)) {
conv <- suppressWarnings(as.numeric(as.character(v)))
n_orig <- sum(!is.na(v) & as.character(v) != "")
if (n_orig > 0 && sum(!is.na(conv)) >= 0.95 * n_orig) {
v <- conv
} else {
stop(sprintf("The rating column(%s) must be numeric — fewer than 95%% of its values convert cleanly.",
score_h))
}
}
df$score <- v
df$subject <- trimws(as.character(df$subject))
df$rater <- trimws(as.character(df$rater))
keep <- !is.na(df$score) & !is.na(df$subject) & df$subject != "" &
!is.na(df$rater) & df$rater != ""
n_invalid <- sum(!keep)
df <- df[keep, , drop = FALSE]
if (nrow(df) < 10) {
stop(sprintf("Only %d usable ratings after removing rows with missing %s, %s, or %s — need at least 10.",
nrow(df), subj_h, rater_h, score_h))
}
rater_levels <- sort(unique(df$rater))
k <- length(rater_levels)
if (k < 2) {
stop(sprintf("ICC needs at least 2 distinct values in the %s column; found %d. Each row should be one rating by one rater.",
rater_h, k))
}
if (k > 30) {
stop(sprintf("The %s column has %d distinct values — that looks like an identifier, not a set of raters. Map the column naming who/what gave each rating.",
rater_h, k))
}Step 3: Average duplicate ratings per subject x rater
agg <- aggregate(score ~ subject + rater, data = df, FUN = mean)
n_duplicates <- nrow(df) - nrow(agg)Step 4: Keep subjects rated by ALL raters (two-way ICCs need a
complete crossing); report the incomplete ones
cnt <- table(agg$subject)
complete_subjects <- names(cnt)[cnt == k]
incomplete_subjects <- names(cnt)[cnt < k]
n_incomplete <- length(incomplete_subjects)
n <- length(complete_subjects)
if (n < 5) {
stop(sprintf("ICC needs at least 5 %s values rated by every %s; only %d of %d have a complete set of %d ratings.",
subj_h, rater_h, n, length(cnt), k))
}
dat <- agg[agg$subject %in% complete_subjects, , drop = FALSE]
if (is.na(var(dat$score)) || isTRUE(var(dat$score) == 0)) {
stop(sprintf("The %s values are constant — reliability cannot be estimated.", score_h))
}
final_rows <- nrow(dat)
rows_removed <- initial_rows - final_rows
dat$subject_f <- factor(dat$subject)
dat$rater_f <- factor(dat$rater, levels = rater_levels)Step 5: ICCs from ANOVA mean squares
Two-way: aov(score ~ subject + rater) -> MSR (subjects), MSC (raters), MSE
a2 <- aov(score ~ subject_f + rater_f, data = dat)
ms2 <- summary(a2)[[1]][["Mean Sq"]]
MSR <- ms2[1]; MSC <- ms2[2]; MSE <- ms2[3]
icc21 <- (MSR - MSE) / (MSR + (k - 1) * MSE + k * (MSC - MSE) / n)
icc31 <- (MSR - MSE) / (MSR + (k - 1) * MSE)
icc2k <- (MSR - MSE) / (MSR + (MSC - MSE) / n)
icc3k <- (MSR - MSE) / MSROne-way: aov(score ~ subject) -> MSB, MSW
a1 <- aov(score ~ subject_f, data = dat)
ms1 <- summary(a1)[[1]][["Mean Sq"]]
MSB <- ms1[1]; MSW <- ms1[2]
icc11 <- (MSB - MSW) / (MSB + (k - 1) * MSW)F-test for ICC(3,1): MSR/MSE against F(n-1, (n-1)(k-1))
f_stat <- MSR / MSE
f_df1 <- n - 1
f_df2 <- (n - 1) * (k - 1)
f_p <- pf(f_stat, f_df1, f_df2, lower.tail = FALSE)Koo & Li (2016) bands: <0.5 poor, 0.5-0.75 moderate, 0.75-0.9 good, >0.9 excellent
koo_li <- function(x) {
if (is.na(x)) return("not estimable")
if (x < 0.5) "poor" else if (x < 0.75) "moderate" else if (x < 0.9) "good" else "excellent"
}
headline_band <- koo_li(icc21)
icc_results_df <- data.frame(
icc_type = c("ICC(1,1) — one-way random, single rater",
"ICC(2,1) — absolute agreement, single rater",
"ICC(3,1) — consistency, single rater",
"ICC(2,k) — absolute agreement, average of all raters",
"ICC(3,k) — consistency, average of all raters"),
value = round(c(icc11, icc21, icc31, icc2k, icc3k), 3),
interpretation = sapply(c(icc11, icc21, icc31, icc2k, icc3k), koo_li),
use_when = c(
sprintf("Each %s is rated by a DIFFERENT random set of raters(no shared panel).", subj_h),
sprintf("Raters are a random sample and absolute %s values must match — the default to report.", score_h),
sprintf("These exact %d raters are the only ones of interest; systematic differences between them are forgiven.", k),
sprintf("Decisions rest on the AVERAGE %s of all %d raters, and absolute values matter.", score_h, k),
sprintf("Decisions rest on the average of these exact %d raters; only relative ordering matters.", k)
),
stringsAsFactors = FALSE
)Step 6: Variance components (from the two-way mean squares)
var_subject <- max(0, (MSR - MSE) / k)
var_rater <- max(0, (MSC - MSE) / n)
var_error <- max(0, MSE)
var_total <- var_subject + var_rater + var_error
if (var_total <= 0) var_total <- 1
variance_df <- data.frame(
source = c("Between subjects", "Between raters", "Residual"),
variance_pct = round(100 * c(var_subject, var_rater, var_error) / var_total, 1),
stringsAsFactors = FALSE
)Step 7: Rater means (systematic bias) on the complete-crossing data
rm_agg <- aggregate(score ~ rater_f, data = dat, FUN = mean)
rater_means_df <- data.frame(
rater_name = as.character(rm_agg$rater_f),
mean_score = round(rm_agg$score, 2),
stringsAsFactors = FALSE
)NA-safe extremes (never which.max over possibly-NA vectors)
ok_idx <- which(!is.na(rater_means_df$mean_score))
hi_idx <- ok_idx[which.max(rater_means_df$mean_score[ok_idx])]
lo_idx <- ok_idx[which.min(rater_means_df$mean_score[ok_idx])]
rater_spread <- rater_means_df$mean_score[hi_idx] - rater_means_df$mean_score[lo_idx]Step 8: Subject-level agreement — first two raters, one point per subject
d1 <- dat[dat$rater == rater_levels[1], c("subject", "score")]
d2 <- dat[dat$rater == rater_levels[2], c("subject", "score")]
names(d1)[2] <- "rater_1_score"
names(d2)[2] <- "rater_2_score"
merged <- merge(d1, d2, by = "subject")
merged <- merged[order(merged$rater_1_score), , drop = FALSE]
set.seed(42)
if (nrow(merged) > 1000) merged <- merged[sample(nrow(merged), 1000), , drop = FALSE]
agreement_df <- data.frame(
rater_1_score = round(merged$rater_1_score, 3),
rater_2_score = round(merged$rater_2_score, 3),
stringsAsFactors = FALSE
)
agreement_r <- if (nrow(agreement_df) >= 3)
suppressWarnings(cor(agreement_df$rater_1_score, agreement_df$rater_2_score))
else NA_real_
icc <- list(icc11 = icc11, icc21 = icc21, icc31 = icc31,
icc2k = icc2k, icc3k = icc3k)
metrics <- list(
`Subjects` = n,
`Raters` = k,
`Ratings Used` = final_rows,
`ICC(2,1)` = round(icc21, 3),
`Reliability` = headline_band,
`F-test p` = signif(f_p, 3)
)
json_output <- list(
answer = paste0(
"Intraclass correlation from ", format(final_rows, big.mark = ","),
" ratings of ", n, " ", subj_h, " values by ", k, " raters(", rater_h,
"): ICC(2,1) = ", round(icc21, 3), " — ", headline_band,
" absolute agreement for a single rater(Koo & Li). Consistency ICC(3,1) = ",
round(icc31, 3), "; average-measure ICC(2,k) = ", round(icc2k, 3),
". The F-test for subject discrimination is ",
ifelse(is.na(f_p), "not estimable",
ifelse(f_p < 0.05, paste0("significant(p = ", signif(f_p, 3), ")"),
paste0("not significant(p = ", signif(f_p, 3), ")"))),
". Between-", subj_h, " differences explain ", variance_df$variance_pct[1],
"% of the variance in ", score_h, "."
),
cards = lapply(
c("tldr", "overview", "preprocessing", "icc_results",
"rater_means", "subject_agreement", "variance_breakdown"),
function(cid) list(id = cid, metrics = metrics)
)
)
list(
initial_rows = initial_rows, final_rows = final_rows,
rows_removed = rows_removed, n_invalid = n_invalid,
subj_h = subj_h, rater_h = rater_h, score_h = score_h,
n_subjects = n, k_raters = k, rater_levels = rater_levels,
n_duplicates = n_duplicates, n_incomplete = n_incomplete,
incomplete_subjects = incomplete_subjects,
icc = icc, f_stat = f_stat, f_df1 = f_df1, f_df2 = f_df2, f_p = f_p,
icc_results_df = icc_results_df,
rater_means_df = rater_means_df,
rater_hi = rater_means_df$rater_name[hi_idx],
rater_lo = rater_means_df$rater_name[lo_idx],
rater_spread = rater_spread,
agreement_df = agreement_df, agreement_r = agreement_r,
variance_df = variance_df,
headline_band = headline_band,
metrics = metrics, json_output = json_output
)
}