Standard Interrupted Time Series
Executive Summary

Executive Summary

Did the intervention shift the level or the slope of drivers killed?

Time Points
192
Level Change
-32.919
Slope Change
1.551
p (level)
0.001
p (slope)
0.037
Lag-1 Autocorrelation
0.56
Gap at Series End
1.208
At 1983-02-01, drivers killed shows both an immediate drop of 32.92 (95% CI -52.93 to -12.91, p = 0.001) and a change in trend of 1.551 per period (95% CI 0.095 to 3.008, p = 0.037) — a pattern consistent with an intervention effect. By the end of the series the fitted values sit 1.21 (1.05%) away from where the pre-intervention trend was headed. CAUTION: the residuals are strongly autocorrelated (lag-1 r = 0.56), so these classical confidence intervals and p-values are anti-conservative — too narrow. This is a single-series observational design: any other event at the same time would be indistinguishable from the intervention.
What this means

The short answer

Monthly driver deaths dropped 32.92 immediately when the seatbelt law took effect (p = 0.001), but the trend then reversed, rising 1.551 per month after the initial dip—a pattern consistent with an intervention effect. By series end, the fitted values sat only 1.21 deaths away from where the pre-law trend would have gone, but strong autocorrelation means these confidence intervals are too narrow.

The detail

At 1983-02-01, drivers killed shows a level change of −32.919 (95% CI −52.93 to −12.91, p = 0.001) and a slope change of 1.551 per period (95% CI 0.095 to 3.008, p = 0.037). The lag-1 residual autocorrelation is 0.56, exceeding the 0.3 threshold. Classical OLS standard errors are anti-conservative under this autocorrelation, so the reported intervals are too narrow: the true uncertainty is larger. The gap at series end is 1.208 (1.05% of the counterfactual level). This is a single-series observational design; any concurrent event is indistinguishable from the intervention.

What this can't tell you

The strong autocorrelation (r = 0.56) means the effective sample size is smaller than 192 points suggest. Confidence intervals and p-values here overstate precision; a robust standard-error correction would widen them. The causal claim rests entirely on the assumption that no other event shifted driver deaths at the same time.

Overview

Analysis Overview

Interrupted time series on drivers killed: segmented regression split at 1983-02-01.

N Points192
Level Change-32.919
Slope Change1.551
Gap At End1.208
What this means

The short answer

Segmented regression isolates the intervention's effect by estimating the pre-law trend and measuring what changed exactly at the break: an immediate drop and a trend reversal. This beats naive before-after averages because it doesn't confuse the existing downward trend with the law's impact.

The detail

The pre-intervention trend ran at −0.102 drivers killed per period. At 1983-02-01, the series shows a level change of −32.919 and a slope change of 1.551 per period. By extending the pre-law trend forward as a counterfactual, the model can say how far the actual series diverged from where it would have gone untouched. The fitted series ends 1.208 deaths away from that counterfactual projection—a gap that grows post-intervention because the slope bent upward even as the level dropped.

What this can't tell you

This design separates trend from intervention only if nothing else moved the series at the same time. A concurrent campaign, seasonal shift, or data collection change would look identical to the law's effect.

Data Preparation

Data Preparation

How the series was assembled and where the break was placed.

Initial Rows192
Final Rows192
Rows Removed0
What this means

The short answer

All 192 monthly observations were retained; no rows dropped for missing data. The intervention date was fixed at 1983-02-01, placing 169 points before the law and 23 after, ensuring the break sits at the exact date the law took effect.

The detail

The dataset loaded 192 rows and used all 192, forming 192 time points with zero removals. The intervention date was provided (not assumed at the series midpoint) as 1983-02-01: observations on or after that date count as post-intervention. This split yields 169 pre-intervention points and 23 post-intervention points, with the first post point at 1983-02-01.

What this can't tell you

The data spans January 1969 to December 1984. Any shifts in how driver deaths were recorded or defined during this window would alter the series but would not appear as a separate variable; consider whether measurement consistency is documented elsewhere.

Visualization

Actual, Fitted, and Counterfactual

The series split at 1983-02-01, with the pre-trend projected forward.

What this means

The short answer

From 1969 through early 1983, monthly driver deaths trended downward at −0.102 per period. At the law's introduction in February 1983, the series dropped 32.92 deaths immediately, then reversed course and began rising at 1.450 per period afterward—the two changes together define a pattern consistent with the law's effect.

The detail

The actual series began at 107 deaths in January 1969 and fluctuated around a declining trend through January 1983. The fitted model captures this pre-trend at −0.102 per period (95% CI −0.175 to −0.029). At 1983-02-01 the fitted line jumps down by 32.92, then the post-intervention slope becomes 1.450 per period (95% CI −0.005 to 2.904). The vertical separation at the break is the level change; the widening gap afterward reflects the slope change. By the final point (1984-12-01), the fitted and counterfactual lines have diverged to 1.21 deaths. Residuals are strongly autocorrelated, so the fitted line's apparent precision should be read cautiously.

What this can't tell you

The autocorrelation caution applies here: the fitted trajectory may appear smoother than the true variability warrants. Any concurrent event—enforcement campaign, reporting change, seasonal shift—would produce the same visual separation at the break.

Data Table

Level and Slope Changes

Pre and post slopes, the jump at the break, and the trend break — each with its confidence interval.

MeasureEstimateDetail
Pre-intervention slope-0.102Change in drivers killed per period before 1983-02-01 (95% CI -0.175 to -0.029)
Post-intervention slope1.45Change in drivers killed per period from 1983-02-01 on (95% CI -0.005 to 2.904)
Level change at intervention-32.92Immediate jump at 1983-02-01; 95% CI -52.93 to -12.91; p = 0.001
Slope change at intervention1.551Trend break at 1983-02-01; 95% CI 0.095 to 3.008; p = 0.037
What this means

The short answer

The level change of −32.919 deaths (p = 0.001) is the clear winner: the law is associated with an immediate drop. The trend also bent upward by 1.551 per period (p = 0.037), reversing the pre-law decline. Both intervals exclude zero, but the strong autocorrelation means these intervals are too narrow.

The detail

Pre-intervention slope: −0.102 (95% CI −0.175 to −0.029). Post-intervention slope: 1.450 (95% CI −0.005 to 2.904). Level change at 1983-02-01: −32.919 (95% CI −52.93 to −12.91, p = 0.001). Slope change: 1.551 (95% CI 0.095 to 3.008, p = 0.037). The level change is statistically significant; the slope change is also significant, though its lower confidence bound (0.095) is close to zero. The lag-1 autocorrelation of 0.56 means these classical intervals are anti-conservative—the true confidence intervals are wider than shown, and the effective sample is smaller than 192 points.

What this can't tell you

The autocorrelation violation means significance claims should be treated cautiously. A robust standard-error method would widen these intervals and likely reduce the slope change's statistical strength. The post-intervention slope's lower bound (−0.005) nearly touches zero, suggesting trend reversal is less firmly established than the level drop.

Data Table

The Counterfactual Gap

Actual versus the pre-trend projection at the end of the series, in units and percent.

MeasureEstimateDetail
Counterfactual drivers killed at series end115Where the pre-intervention trend would have put drivers killed by 1984-12-01 had nothing changed
Model-fitted drivers killed at series end116.2Where the segmented model places drivers killed at 1984-12-01 given the level and slope changes
Gap at series end (units)1.208Fitted minus counterfactual at 1984-12-01 (1.21)
Gap at series end (% of counterfactual)1.05Relative to the counterfactual level of 115.00 (1.05%)
Average post-period gap (units)-15.86Mean fitted-minus-counterfactual gap across the 23 post-intervention points
What this means

The short answer

If the pre-law downward trend had continued unchanged, driver deaths would have reached 114.999 by December 1984. Instead, the model places them at 116.207—a gap of only 1.208 deaths (1.05% of the counterfactual). The average post-intervention gap is −15.855, showing the level drop dominated the first months after the law.

The detail

Counterfactual drivers killed at series end (1984-12-01): 114.999. Model-fitted drivers killed at series end: 116.207. Gap at series end (units): 1.208. Gap at series end (percent of counterfactual): 1.05%. Average post-period gap across all 23 post-intervention points: −15.855. The negative average gap reflects the large initial level drop (−32.919); the small positive gap at the end shows the upward trend bend (1.551 per period) has partially closed that deficit over 23 months. The projection assumes the pre-law trend would have held indefinitely—the further past the break it extends, the more speculative it becomes.

What this can't tell you

The counterfactual is built on the assumption that the pre-law trend would have persisted. Any structural change in traffic patterns, enforcement, or economic conditions independent of the law would alter the true counterfactual and make this projection unreliable.

Data Table

Method & Assumptions

The design, the model, the autocorrelation check, and the assumptions the causal read rests on.

ItemDetail
DesignObservational interrupted time series (segmented regression) on a single series — no control group. Not a randomized experiment.
Intervention pointThe intervention date was provided as 1983-02-01: points on or after it count as post-intervention.
ModelOLS regression of drivers killed on time, a post-intervention indicator, and time-since-intervention: the indicator's coefficient is the level change and the time-since coefficient is the slope change.
InferenceClassical OLS standard errors and 95% confidence intervals. No autocorrelation-robust (e.g. Newey-West) correction is applied — when residual autocorrelation is high, the intervals here are too narrow.
Autocorrelation checkCAUTION: the lag-1 residual autocorrelation is 0.56 — beyond the 0.3 threshold. This analysis uses classical OLS standard errors, so the confidence intervals and p-values reported here are anti-conservative (too narrow): treat any significance claim with caution, as the effective sample is smaller than the point count suggests.
CounterfactualThe pre-intervention trend extended forward unchanged; the gap between it and the fitted series at the end of the data is the cumulative divergence.
Key assumptionsThe pre-intervention trend would have continued unchanged absent the intervention; no other event shifted the series at the same time — any concurrent event is indistinguishable from the intervention itself; the series' measurement is consistent throughout.
What this means

The short answer

This is an observational interrupted time series — a single monthly series fitted before and after the law took effect in February 1983, with no control group. The design can detect whether the series shifted at the break, but cannot rule out coinciding events. Strong autocorrelation (0.56) means the reported confidence intervals are too narrow.

The detail

The model is segmented OLS regression of drivers killed on time, a post-intervention indicator (1983-02-01 onward), and time-since-intervention. The level and slope changes are the coefficients on the indicator and time-since terms respectively. Classical OLS standard errors are reported; no autocorrelation-robust correction (e.g. Newey-West) is applied. Lag-1 residual autocorrelation is 0.56, exceeding the 0.3 threshold, making the 95% confidence intervals anti-conservative (too narrow).

What this can't tell you

No control series guards against concurrent events — any shift in enforcement, media coverage, seasonal patterns, or data collection at February 1983 is indistinguishable from the seatbelt law itself. The pre-intervention trend is assumed to continue unchanged; no other structural break is tested. The strong autocorrelation means effective sample size is smaller than the 192 monthly observations suggest, so significance claims require caution. Consider a longer post-intervention window or a control series (e.g. non-driver road deaths) to strengthen causal inference.

Rate this report Was this the answer you needed?
The exact source that produced this report — yours to keep, read, and re-run.
Download PDF
How this was computed method · R source · citation
The code that did it

Core Analysis Pipeline

compute_shared <- function(df, params, col_map = list()) {
  # === SHARED EXPORTS ===
  #   initial_rows/final_rows/rows_removed  $ row accounting
  #   n_dropped            $ rows dropped for missing date/value
  #   date_h/value_h       $ humanized user names for prose
  #   n_points             $ distinct time points used
  #   agg_note             $ note when several rows per date were averaged
  #   split_rule           $ how the intervention point was chosen
  #   midpoint_used        $ TRUE when the midpoint fallback was assumed
  #   int_date_str         $ ISO date used as the break (or midpoint date)
  #   t0/n_pre/n_post      $ break index + points each side
  #   b0/b1 (pre trend)    $ intercept + pre slope
  #   level/level_se/level_p/level_lo/level_hi  $ LEVEL change (jump)
  #   slope/slope_se/slope_p/slope_lo/slope_hi  $ SLOPE change (trend break)
  #   post_slope/post_slope_lo/post_slope_hi    $ pre slope + slope change
  #   r1/ac_flag/ac_verdict $ lag-1 residual autocorrelation check
  #   cf_end/fit_end/gap/gap_pct/gap_pct_txt    $ counterfactual gap at end
  #   avg_gap              $ mean post-period gap (actual fit - counterfactual)
  #   its_trends_df        $ period_date (ISO), series_value, series_label
  #   segments_df          $ measure, estimate, detail
  #   gap_df               $ measure, estimate, detail
  #   methods_df           $ item, detail
  #   metrics / json_output
  # === /SHARED EXPORTS ===

Step 1: Resolve mapped columns (humanized for all prose)

initial_rows <- nrow(df)
  date_h  <- humanize_semantic("date", col_map)
  value_h <- humanize_semantic("value", col_map)
  for (req in c("date", "value")) {
    if (!(req %in% names(df))) {
      stop(sprintf("Interrupted time series needs &#x27;%s' (%s) mapped.",
                   humanize_semantic(req, col_map),
                   c(date = "the date/time column",
                     value = "the numeric series to analyse")[[req]]))
    }
  }

Step 2: Coerce the value to numeric (95% rule)

y_raw <- df$value
  if (!is.numeric(y_raw)) {
    ch <- as.character(y_raw)
    non_blank <- !is.na(ch) & trimws(ch) != ""
    conv <- suppressWarnings(as.numeric(ch))
    if (sum(non_blank) == 0 ||
        sum(!is.na(conv[non_blank])) < 0.95 * sum(non_blank)) {
      stop(sprintf("The value column &#x27;%s' does not look numeric — fewer than 95%% of its values parse as numbers. Pick a numeric column.",
                   value_h))
    }
    y_raw <- conv
  }

Step 3: Parse the date column (ISO, mdy, dmy, or integer years)

parse_dates_vec <- function(ch) {
    ok <- function(dd) sum(!is.na(dd)) >= 0.95 * sum(!is.na(ch) & trimws(ch) != "")
    d <- suppressWarnings(as.Date(ch, format = "%Y-%m-%d"))
    if (!ok(d)) d <- suppressWarnings(as.Date(lubridate::ymd(ch, quiet = TRUE)))
    if (!ok(d)) d <- suppressWarnings(as.Date(lubridate::mdy(ch, quiet = TRUE)))
    if (!ok(d)) d <- suppressWarnings(as.Date(lubridate::dmy(ch, quiet = TRUE)))
    if (!ok(d)) {
      nv <- suppressWarnings(as.numeric(ch))
      nv_ok <- stats::na.omit(nv)
      if (length(nv_ok) >= 0.95 * sum(!is.na(ch) & trimws(ch) != "") &&
          length(nv_ok) > 0 && all(nv_ok == round(nv_ok)) &&
          all(nv_ok >= 1900 & nv_ok <= 2100)) {
        d <- as.Date(ifelse(is.na(nv), NA, sprintf("%04d-01-01", nv)),
                     format = "%Y-%m-%d")
      }
    }
    if (ok(d)) d else NULL
  }
  parse_date_one <- function(s) {
    s <- trimws(as.character(s))
    d <- suppressWarnings(as.Date(s, format = "%Y-%m-%d"))
    if (is.na(d)) d <- suppressWarnings(as.Date(lubridate::ymd(s, quiet = TRUE)))
    if (is.na(d)) d <- suppressWarnings(as.Date(lubridate::mdy(s, quiet = TRUE)))
    if (is.na(d)) d <- suppressWarnings(as.Date(lubridate::dmy(s, quiet = TRUE)))
    if (is.na(d)) {
      ny <- suppressWarnings(as.numeric(s))
      if (!is.na(ny) && ny == round(ny) && ny >= 1900 && ny <= 2100)
        d <- as.Date(sprintf("%04d-01-01", ny))
    }
    d
  }

  d_ch <- trimws(as.character(df$date))
  dts <- parse_dates_vec(d_ch)
  if (is.null(dts)) {
    stop(sprintf("The date column &#x27;%s' could not be read as dates — fewer than 95%% of its values parse as calendar dates (or integer years). Map a date column so the series can be ordered in time.",
                 date_h))
  }

Step 4: Keep rows complete on date and value; average duplicates

keep <- !is.na(y_raw) & !is.na(dts)
  n_dropped <- sum(!keep)
  y_k <- y_raw[keep]; d_k <- dts[keep]
  if (length(y_k) < 4) {
    stop(sprintf("Only %d usable rows remained after dropping rows missing %s or %s — far too few for an interrupted time series.",
                 length(y_k), date_h, value_h))
  }
  agg <- stats::aggregate(list(y = y_k), by = list(date_iso = format(d_k, "%Y-%m-%d")), FUN = mean)
  agg <- agg[order(agg$date_iso), , drop = FALSE]
  n_points <- nrow(agg)
  final_rows <- length(y_k)
  rows_removed <- initial_rows - final_rows
  agg_note <- if (n_points < final_rows) {
    sprintf("Several rows share the same %s value, so the %s values were first averaged to one point per date(%s rows became %d points).",
            date_h, value_h, format(final_rows, big.mark = ","), n_points)
  } else ""
  if (n_points < 16) {
    stop(sprintf("The series in &#x27;%s' has only %d distinct time points — interrupted time series needs at least 16 (8 on each side of the intervention).",
                 date_h, n_points))
  }

  dd <- as.Date(agg$date_iso)
  dnum <- as.numeric(dd)
  yv <- agg$y
  idx <- seq_len(n_points)

Step 5: Place the intervention point (parameter or midpoint ASSUMPTION)

int_param <- params$intervention_date %||% params$treatment_date %||% NULL
  midpoint_used <- FALSE
  if (!is.null(int_param)) {
    tsd <- parse_date_one(int_param)
    if (is.na(tsd)) {
      stop(sprintf("The intervention_date parameter(&#x27;%s') could not be read as a date.",
                   as.character(int_param)))
    }
    is_post <- dnum >= as.numeric(tsd)
    int_date_str <- format(tsd, "%Y-%m-%d")
    split_rule <- sprintf("The intervention date was provided as %s: points on or after it count as post-intervention.",
                          int_date_str)
  } else {
    midpoint_used <- TRUE
    cut_v <- min(dnum) + (max(dnum) - min(dnum)) / 2
    is_post <- dnum > cut_v
    int_date_str <- format(as.Date(cut_v, origin = "1970-01-01"), "%Y-%m-%d")
    split_rule <- sprintf("No intervention_date parameter was provided, so the intervention point was ASSUMED at the midpoint of the observed date range(%s) — points after it count as post-intervention. This midpoint is an assumption, not data: provide an intervention_date parameter to place the break where the intervention actually happened.",
                          int_date_str)
  }

  n_pre <- sum(!is_post)
  n_post <- sum(is_post)
  if (n_pre < 8) {
    stop(sprintf("Only %d point(s) in &#x27;%s' fall before the intervention date (%s) — interrupted time series needs at least 8 points on each side of the break to separate a level shift from the pre-existing trend.",
                 n_pre, date_h, int_date_str))
  }
  if (n_post < 8) {
    stop(sprintf("Only %d point(s) in &#x27;%s' fall on or after the intervention date (%s) — interrupted time series needs at least 8 points on each side of the break to estimate the post-intervention trend.",
                 n_post, date_h, int_date_str))
  }
  t0 <- min(idx[is_post])
  break_date_str <- agg$date_iso[t0]

Step 6: Segmented OLS — value ~ time + post + time_since_intervention

post01 <- as.integer(is_post)
  tsi <- pmax(0L, idx - t0)
  af <- data.frame(y = yv, t = idx, post = post01, tsi = tsi)
  fit <- stats::lm(y ~ t + post + tsi, data = af)
  co <- summary(fit)$coefficients
  if (!all(c("t", "post", "tsi") %in% rownames(co)) ||
      any(is.na(coef(fit)[c("t", "post", "tsi")]))) {
    stop("The segmented regression could not be estimated — the series may be too short or degenerate around the intervention point.")
  }
  b0 <- unname(coef(fit)["(Intercept)"])
  b1 <- unname(coef(fit)["t"])
  level    <- unname(coef(fit)["post"])
  slope    <- unname(coef(fit)["tsi"])
  level_se <- co["post", "Std. Error"]; level_p <- co["post", 4]
  slope_se <- co["tsi", "Std. Error"];  slope_p <- co["tsi", 4]
  ci <- stats::confint(fit)
  level_lo <- ci["post", 1]; level_hi <- ci["post", 2]
  slope_lo <- ci["tsi", 1];  slope_hi <- ci["tsi", 2]
  b1_lo <- ci["t", 1]; b1_hi <- ci["t", 2]

Post slope = pre slope + slope change, with a combined-variance CI

V <- stats::vcov(fit)
  post_slope <- b1 + slope
  ps_se <- sqrt(V["t", "t"] + V["tsi", "tsi"] + 2 * V["t", "tsi"])
  tcrit <- stats::qt(0.975, df = fit$df.residual)
  post_slope_lo <- post_slope - tcrit * ps_se
  post_slope_hi <- post_slope + tcrit * ps_se

Step 7: Lag-1 residual autocorrelation check

res <- stats::residuals(fit)[order(idx)]
  r1 <- if (length(res) >= 10) {
    suppressWarnings(stats::cor(res[-1], res[-length(res)]))
  } else NA_real_
  ac_flag <- !is.na(r1) && abs(r1) > 0.3
  ac_verdict <- if (is.na(r1)) {
    "The lag-1 residual autocorrelation could not be computed."
  } else if (ac_flag) {
    sprintf("CAUTION: the lag-1 residual autocorrelation is %s — beyond the 0.3 threshold. This analysis uses classical OLS standard errors, so the confidence intervals and p-values reported here are anti-conservative(too narrow): treat any significance claim with caution, as the effective sample is smaller than the point count suggests.",
            r2(r1))
  } else {
    sprintf("The lag-1 residual autocorrelation is %s, low enough that classical OLS standard errors are a reasonable approximation for this series.",
            r2(r1))
  }

Step 8: Counterfactual projection and the gap at series end

t_end <- n_points
  cf_end  <- b0 + b1 * t_end
  fit_end <- b0 + b1 * t_end + level + slope * (t_end - t0)
  gap <- fit_end - cf_end
  gap_pct <- if (abs(cf_end) > 1e-8) 100 * gap / abs(cf_end) else NA_real_
  gap_pct_txt <- if (is.na(gap_pct)) "not defined(counterfactual level near zero)" else paste0(r2(gap_pct), "%")
  cf_v  <- b0 + b1 * idx
  fit_v <- as.numeric(stats::fitted(fit))
  avg_gap <- mean(fit_v[is_post] - cf_v[is_post])

Step 9: Chart dataset — actual, fitted, counterfactual (ISO dates)

keep_i <- idx
  if (n_points > 500) {
    keep_i <- sort(unique(c(1, n_points, t0 - 1, t0,
                            round(seq(1, n_points, length.out = 500)))))
  }
  cf_i <- keep_i[keep_i >= (t0 - 1)]   # counterfactual from the last pre point on
  its_trends_df <- rbind(
    data.frame(period_date = agg$date_iso[keep_i],
               series_value = round(yv[keep_i], 3),
               series_label = "Actual", stringsAsFactors = FALSE),
    data.frame(period_date = agg$date_iso[keep_i],
               series_value = round(fit_v[keep_i], 3),
               series_label = "Fitted", stringsAsFactors = FALSE),
    data.frame(period_date = agg$date_iso[cf_i],
               series_value = round(cf_v[cf_i], 3),
               series_label = "Counterfactual", stringsAsFactors = FALSE)
  )
  rownames(its_trends_df) <- NULL
Your data has more stories to tell.Run any analysis on your own data — validated R modules, interactive reports, AI insights, and PDF export. 500 free credits on signup.
Try Free — No SignupSign Up Free

Cite this analysis

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing