Decomposing credit risk into age, cohort and period

A step-by-step decomposition of rolling three-month credit risk using public Bondora loan data.
R
credit risk
Author

Joshua Kunst

Published

September 12, 2026

Credit vintages are usually read over horizons longer than one month. A twelve-month default rate is natural for regulatory and annual portfolio views, while a shorter horizon is often more useful for early delinquency monitoring. In this post we use three months: long enough for early deterioration to emerge, but short enough to retain useful timing information.

Prefer a visual explanation?

If you would rather begin with the intuition, the interactive story shows the same journey visually: first the loans and their three-month outcomes, then the AGE, COHORT and PERIOD decomposition. This post is the complete explanation; the visualization is its shorter companion.

Explore the interactive RD3M decomposition →

The quantity of interest is the rolling three-month default risk, or RD3M. It asks a direct question: among the loans alive today, how many default during the next three months?

Every observation has three coordinates:

  • Age \(a\): months since origination at the beginning of the window.
  • Cohort \(c\): origination month.
  • Period \(p\): first calendar month of the three-month window.

With the convention used here, they satisfy:

\[ p=c+a. \]

We will construct RD3M from individual loans, arrange it as a vintage table, decompose it sequentially into AGE, COHORT and PERIOD effects, and finally return to the observed portfolio risk.

What exactly does RD3M measure?

RD3M is simply a ratio:

\[ RD3M= \frac{\text{loans that default during the next three months}} {\text{loans alive at the beginning of the window}}. \]

The denominator includes only loans for which the complete three-month outcome can be observed. Otherwise a recent loan could be counted as a non-default even though its full outcome window has not happened yet.

The windows are rolling. A default at age 6 belongs to the windows starting at ages 4, 5 and 6:

Starting age Window Default at age 6?
3 3–5 No
4 4–6 Yes
5 5–7 Yes
6 6–8 Yes

This overlap is intentional. We are asking the same forward-looking question at each possible starting month. It also means consecutive RD3M observations share two calendar months and must not be interpreted as independent measurements.

Public loan data

Go & Grow’s public statistics page links a downloadable Bondora loan dataset. The workbook contains one row per loan and a second sheet with its data dictionary.

For this example we use Estonian loans. The source workbook is large, so the preparation chunk stores a small local RDS containing only the required columns. Later renders reuse that file.

Download and prepare Bondora data
library(dplyr)
library(ggplot2)
library(lubridate)
library(tidyr)

data_url <- "https://sabanners001.blob.core.windows.net/statistics/public/loan_dataset_investor.xlsx"
country_name <- "Estonia"
data_dir <- "blog/posts/2026-09-12-decomposing-credit-vintages/data"
data_file <- file.path(data_dir, "bondora-apc-data.rds")

dir.create(data_dir, recursive = TRUE, showWarnings = FALSE)

needed_columns <- c(
  "loan_id",
  "country",
  "loan_issued_at",
  "loan_status",
  "is_default",
  "months_on_book"
)

if (!file.exists(data_file)) {
  raw_file <- tempfile(fileext = ".xlsx")
  download.file(data_url, raw_file, mode = "wb")

  header <- openxlsx::read.xlsx(
    raw_file,
    sheet = "Loan Dataset",
    rows = 1,
    colNames = FALSE,
    skipEmptyCols = FALSE
  ) |>
    unlist(use.names = FALSE) |>
    as.character()

  column_index <- match(needed_columns, header)

  if (anyNA(column_index)) {
    stop(
      "Missing Bondora fields: ",
      paste(needed_columns[is.na(column_index)], collapse = ", ")
    )
  }

  loans_raw <- openxlsx::read.xlsx(
    raw_file,
    sheet = "Loan Dataset",
    cols = column_index,
    detectDates = TRUE,
    check.names = FALSE
  ) |>
    tibble::as_tibble() |>
    filter(country == country_name)

  loan_dictionary <- openxlsx::read.xlsx(
    raw_file,
    sheet = "Dataset Dictionary"
  ) |>
    tibble::as_tibble() |>
    filter(Column %in% needed_columns) |>
    select(Column, Description)

  saveRDS(
    list(
      loans = loans_raw,
      dictionary = loan_dictionary
    ),
    data_file,
    compress = "gzip"
  )

  unlink(raw_file)
}

bondora_data <- readRDS(data_file)
loans_raw <- bondora_data$loans |>
  filter(country == country_name)
loan_dictionary <- bondora_data$dictionary

glimpse(loans_raw)
Rows: 240,987
Columns: 6
$ loan_id        <chr> "0000B9FA-5598-4130-9143-AB63012DC8D8", "0001463B-E6E2-…
$ country        <chr> "Estonia", "Estonia", "Estonia", "Estonia", "Estonia", …
$ loan_issued_at <dbl> 43877.85, 44377.77, 44299.71, 43699.39, 43890.12, 42863…
$ loan_status    <chr> "Repaid", "Repaid", "Defaulted", "Repaid", "Repaid", "R…
$ is_default     <lgl> FALSE, FALSE, TRUE, FALSE, FALSE, FALSE, FALSE, TRUE, F…
$ months_on_book <dbl> 12, 6, 39, 26, 0, 15, 0, 35, 8, 10, 17, 65, 9, 22, 16, …

The data dictionary clarifies the six fields used below. In particular, months_on_book records how many months the loan has remained in active status.

knitr::kable(
  loan_dictionary,
  col.names = c("Field", "Description"),
  align = c("l", "l")
)
Field Description
country Country of the loan.
is_default Boolean. Flag indicating if the loan is in default.
loan_id Unique identifier of the loan.
loan_issued_at Time when loan was issued.
loan_status Shows the status of the loan on the given date.
months_on_book Months a loan has been in active status.

Preparing complete three-month windows

We keep cohorts originated from 2018 through 2022 and follow them through at most 36 months. A three-month window can begin no later than age 34 if it must end by age 36.

cohort_from <- as.Date("2018-01-01")
cohort_to <- as.Date("2022-12-01")
max_age <- 36L
horizon_months <- 3L

# A window beginning at age 34 covers ages 34, 35 and 36.
max_start_age <- max_age - horizon_months + 1L

loans <- loans_raw |>
  mutate(
    # Excel stores this field as a serial date in the current workbook.
    loan_date = if (inherits(loan_issued_at, "Date")) {
      loan_issued_at
    } else {
      as.Date(loan_issued_at, origin = "1899-12-30")
    },
    cohort = floor_date(loan_date, unit = "month"),
    months_on_book = as.integer(months_on_book),
    is_default = as.logical(is_default),
    default_age = if_else(is_default, months_on_book, NA_integer_),
    observed_age = pmin(months_on_book, max_age),
    # Once default occurs, its three-month outcome is already known.
    last_start_age = if_else(
      !is.na(default_age) & default_age <= max_age,
      pmin(default_age, max_start_age),
      pmin(observed_age - horizon_months + 1L, max_start_age)
    )
  ) |>
  filter(
    cohort >= cohort_from,
    cohort <= cohort_to,
    observed_age >= 1L,
    last_start_age >= 1L
  ) |>
  select(
    loan_id,
    country,
    loan_status,
    cohort,
    months_on_book,
    is_default,
    default_age,
    observed_age,
    last_start_age
  )

glimpse(loans)
Rows: 112,838
Columns: 9
$ loan_id        <chr> "0000B9FA-5598-4130-9143-AB63012DC8D8", "0001463B-E6E2-…
$ country        <chr> "Estonia", "Estonia", "Estonia", "Estonia", "Estonia", …
$ loan_status    <chr> "Repaid", "Repaid", "Defaulted", "Repaid", "Defaulted",…
$ cohort         <date> 2020-02-01, 2021-06-01, 2021-04-01, 2019-08-01, 2022-0…
$ months_on_book <int> 12, 6, 39, 26, 35, 8, 65, 9, 30, 10, 20, 102, 70, 14, 3…
$ is_default     <lgl> FALSE, FALSE, TRUE, FALSE, TRUE, FALSE, FALSE, TRUE, TR…
$ default_age    <int> NA, NA, 39, NA, 35, NA, NA, 9, 30, NA, 20, NA, NA, NA, …
$ observed_age   <int> 12, 6, 36, 26, 35, 8, 36, 9, 30, 10, 20, 36, 36, 14, 36…
$ last_start_age <int> 10, 4, 34, 24, 34, 6, 34, 9, 30, 8, 20, 34, 34, 12, 34,…

For a surviving loan, last_start_age requires all three future months to be observed. For a defaulted loan, the outcome is known as soon as default occurs, so a window ending after that date does not create an unknown outcome.

One loan, window by window

Start with one loan that defaulted around age 6.

one_loan <- loans |>
  filter(is_default, default_age <= max_age) |>
  arrange(abs(default_age - 6L)) |>
  slice(1) |>
  select(loan_id, cohort, observed_age, default_age, last_start_age)

one_loan
# A tibble: 1 × 5
  loan_id                     cohort     observed_age default_age last_start_age
  <chr>                       <date>            <int>       <int>          <int>
1 0239137C-0ABA-4EA7-9EFF-AA… 2019-10-01            6           6              6

uncount() creates one row for every eligible starting age. The period is the first month of the forward window; window_end_period is its third and final month.

one_loan_windows <- one_loan |>
  uncount(last_start_age, .id = "age") |>
  mutate(
    period = cohort %m+% months(age),
    window_end_period = period %m+% months(horizon_months - 1L),
    default_3m = as.integer(
      !is.na(default_age) &
        default_age >= age &
        default_age <= age + horizon_months - 1L
    )
  ) |>
  select(
    loan_id,
    cohort,
    age,
    period,
    window_end_period,
    default_age,
    default_3m
  )

one_loan_windows
# A tibble: 6 × 7
  loan_id   cohort       age period     window_end_period default_age default_3m
  <chr>     <date>     <int> <date>     <date>                  <int>      <int>
1 0239137C… 2019-10-01     1 2019-11-01 2020-01-01                  6          0
2 0239137C… 2019-10-01     2 2019-12-01 2020-02-01                  6          0
3 0239137C… 2019-10-01     3 2020-01-01 2020-03-01                  6          0
4 0239137C… 2019-10-01     4 2020-02-01 2020-04-01                  6          1
5 0239137C… 2019-10-01     5 2020-03-01 2020-05-01                  6          1
6 0239137C… 2019-10-01     6 2020-04-01 2020-06-01                  6          1

The repeated 1 values around the default are not duplicate events in an event count. They are three valid answers to three different questions: whether the same loan will default within three months when observed from three consecutive starting ages.

From loans to the RD3M vintage table

Apply the same expansion to all loans.

loan_windows <- loans |>
  select(loan_id, cohort, default_age, last_start_age) |>
  uncount(last_start_age, .id = "age") |>
  mutate(
    period = cohort %m+% months(age),
    window_end_period = period %m+% months(horizon_months - 1L),
    default_3m = as.integer(
      !is.na(default_age) &
        default_age >= age &
        default_age <= age + horizon_months - 1L
    )
  ) |>
  select(
    loan_id,
    cohort,
    age,
    period,
    window_end_period,
    default_3m
  )

loan_windows
# A tibble: 2,399,367 × 6
   loan_id              cohort       age period     window_end_period default_3m
   <chr>                <date>     <int> <date>     <date>                 <int>
 1 0000B9FA-5598-4130-… 2020-02-01     1 2020-03-01 2020-05-01                 0
 2 0000B9FA-5598-4130-… 2020-02-01     2 2020-04-01 2020-06-01                 0
 3 0000B9FA-5598-4130-… 2020-02-01     3 2020-05-01 2020-07-01                 0
 4 0000B9FA-5598-4130-… 2020-02-01     4 2020-06-01 2020-08-01                 0
 5 0000B9FA-5598-4130-… 2020-02-01     5 2020-07-01 2020-09-01                 0
 6 0000B9FA-5598-4130-… 2020-02-01     6 2020-08-01 2020-10-01                 0
 7 0000B9FA-5598-4130-… 2020-02-01     7 2020-09-01 2020-11-01                 0
 8 0000B9FA-5598-4130-… 2020-02-01     8 2020-10-01 2020-12-01                 0
 9 0000B9FA-5598-4130-… 2020-02-01     9 2020-11-01 2021-01-01                 0
10 0000B9FA-5598-4130-… 2020-02-01    10 2020-12-01 2021-02-01                 0
# ℹ 2,399,357 more rows

Now group the rows by cohort, starting age and starting period. Each cell contains the loans eligible at the beginning of that window and the number that default during its three months.

rd3m_cells <- loan_windows |>
  summarise(
    loans_at_risk = n(),
    defaults_3m = sum(default_3m),
    .by = c(cohort, age, period, window_end_period)
  ) |>
  mutate(rd3m = defaults_3m / loans_at_risk) |>
  arrange(cohort, age)

rd3m_cells
# A tibble: 2,040 × 7
   cohort       age period     window_end_period loans_at_risk defaults_3m
   <date>     <int> <date>     <date>                    <int>       <int>
 1 2018-01-01     1 2018-02-01 2018-04-01                 1202           3
 2 2018-01-01     2 2018-03-01 2018-05-01                 1168           6
 3 2018-01-01     3 2018-04-01 2018-06-01                 1110           9
 4 2018-01-01     4 2018-05-01 2018-07-01                 1066           7
 5 2018-01-01     5 2018-06-01 2018-08-01                 1013           4
 6 2018-01-01     6 2018-07-01 2018-09-01                  961           2
 7 2018-01-01     7 2018-08-01 2018-10-01                  927           2
 8 2018-01-01     8 2018-09-01 2018-11-01                  904           3
 9 2018-01-01     9 2018-10-01 2018-12-01                  866           4
10 2018-01-01    10 2018-11-01 2019-01-01                  842           8
# ℹ 2,030 more rows
# ℹ 1 more variable: rd3m <dbl>

The denominator is a loan count, not monetary exposure or EAD. Every surviving loan therefore has the same weight.

Portfolio RD3M

Before separating AGE, COHORT and PERIOD, pool all eligible loans observed from the same starting month:

\[ RD3M_p= \frac{\text{defaults during }p,p+1,p+2} {\text{loans alive at the start of }p \text{ with an observable outcome}}. \]

portfolio_rd3m <- loan_windows |>
  summarise(
    loans_at_risk = n(),
    defaults_3m = sum(default_3m),
    .by = period
  ) |>
  mutate(rd3m = defaults_3m / loans_at_risk) |>
  arrange(period)

portfolio_rd3m
# A tibble: 93 × 4
   period     loans_at_risk defaults_3m    rd3m
   <date>             <int>       <int>   <dbl>
 1 2018-02-01          1202           3 0.00250
 2 2018-03-01          2154           6 0.00279
 3 2018-04-01          3136          16 0.00510
 4 2018-05-01          4191          18 0.00429
 5 2018-06-01          5222          25 0.00479
 6 2018-07-01          6252          28 0.00448
 7 2018-08-01          7345          40 0.00545
 8 2018-09-01          8524          44 0.00516
 9 2018-10-01          9671          66 0.00682
10 2018-11-01         11138         109 0.00979
# ℹ 83 more rows
ggplot(portfolio_rd3m, aes(period, rd3m)) +
  geom_line(linewidth = 0.9, alpha = 0.8) +
  scale_x_date(date_breaks = "1 year", date_labels = "%Y") +
  scale_y_continuous(labels = scales::label_percent(accuracy = 0.1)) +
  labs(
    title = "Rolling three-month portfolio default risk",
    subtitle = "Each point covers the starting month and the following two months",
    x = "Window start",
    y = "RD3M"
  )

This curve tells us when forward-looking portfolio risk changes. It does not yet tell us whether the movement comes from loan age, the origination cohorts present in the portfolio or conditions shared by the calendar window.

Reading RD3M as a vintage

A vintage table places cohorts in rows and starting ages in columns. Here each cell is a forward-looking three-month rate rather than a cumulative rate since origination.

triangle_cohorts <- seq.Date(
  as.Date("2019-06-01"),
  as.Date("2020-05-01"),
  by = "month"
)

rd3m_triangle <- rd3m_cells |>
  filter(cohort %in% triangle_cohorts, age <= 12L) |>
  select(cohort, age, rd3m) |>
  mutate(
    cohort = format(cohort, "%Y-%m"),
    rd3m = scales::percent(rd3m, accuracy = 0.01)
  ) |>
  pivot_wider(
    names_from = age,
    names_prefix = "M",
    values_from = rd3m
  ) |>
  arrange(cohort)

rd3m_triangle
# A tibble: 12 × 13
   cohort  M1    M2    M3    M4    M5    M6    M7    M8    M9    M10   M11   M12  
   <chr>   <chr> <chr> <chr> <chr> <chr> <chr> <chr> <chr> <chr> <chr> <chr> <chr>
 1 2019-06 1.08% 1.73% 2.63% 2.52% 2.61% 2.57% 2.07% 1.77% 1.89% 2.31% 2.65% 2.97%
 2 2019-07 1.00% 1.44% 2.44% 2.10% 3.00% 2.53% 2.82% 2.22% 2.85% 3.18% 3.48% 2.83%
 3 2019-08 1.21% 1.94% 2.66% 2.44% 2.53% 2.31% 2.27% 2.59% 2.61% 2.72% 2.58% 3.07%
 4 2019-09 1.13% 1.81% 3.29% 2.70% 2.52% 1.83% 2.28% 3.32% 3.82% 3.78% 2.94% 2.25%
 5 2019-10 0.87% 1.66% 2.54% 2.28% 2.29% 2.38% 2.71% 2.95% 3.00% 2.86% 3.18% 3.07%
 6 2019-11 0.86% 1.12% 1.27% 1.23% 1.72% 2.33% 2.44% 2.50% 2.42% 2.11% 2.23% 2.67%
 7 2019-12 0.77% 0.95% 1.25% 1.32% 2.02% 2.47% 2.44% 1.92% 1.93% 2.06% 2.47% 2.06%
 8 2020-01 0.20% 0.49% 0.83% 1.32% 2.05% 2.19% 1.99% 1.58% 1.92% 2.69% 2.74% 1.93%
 9 2020-02 0.19% 0.40% 0.81% 1.55% 2.02% 2.13% 2.11% 2.18% 2.24% 1.76% 1.65% 1.36%
10 2020-03 0.21% 0.85% 1.51% 1.88% 1.94% 1.82% 1.82% 2.01% 1.95% 1.82% 1.20% 0.87%
11 2020-04 0.44% 0.83% 1.32% 1.44% 1.92% 2.07% 2.13% 1.74% 2.18% 1.98% 2.26% 1.49%
12 2020-05 0.49% 1.16% 1.62% 1.66% 1.35% 1.40% 1.25% 1.29% 1.12% 1.16% 1.42% 1.78%

Moving horizontally follows one cohort as it ages. Moving vertically compares different origination cohorts at the same age. A calendar period cuts diagonally through the table because different cohorts reach that period at different ages.

period_example <- as.Date("2020-04-01")

rd3m_cells |>
  filter(
    period == period_example,
    cohort >= as.Date("2018-01-01"),
    cohort <= as.Date("2020-03-01")
  ) |>
  select(
    cohort,
    age,
    period,
    window_end_period,
    loans_at_risk,
    defaults_3m,
    rd3m
  ) |>
  arrange(cohort)
# A tibble: 27 × 7
   cohort       age period     window_end_period loans_at_risk defaults_3m
   <date>     <int> <date>     <date>                    <int>       <int>
 1 2018-01-01    27 2020-04-01 2020-06-01                  496           6
 2 2018-02-01    26 2020-04-01 2020-06-01                  437           1
 3 2018-03-01    25 2020-04-01 2020-06-01                  500           8
 4 2018-04-01    24 2020-04-01 2020-06-01                  571           4
 5 2018-05-01    23 2020-04-01 2020-06-01                  590           8
 6 2018-06-01    22 2020-04-01 2020-06-01                  660           8
 7 2018-07-01    21 2020-04-01 2020-06-01                  645           4
 8 2018-08-01    20 2020-04-01 2020-06-01                  772           9
 9 2018-09-01    19 2020-04-01 2020-06-01                  730          14
10 2018-10-01    18 2020-04-01 2020-06-01                 1035          18
# ℹ 17 more rows
# ℹ 1 more variable: rd3m <dbl>

All these cells begin in April 2020 and end in June 2020, but they contain loans at different ages and from different origination cohorts. That is the variation the APC decomposition tries to organise.

Raw age and cohort profiles

Pooling cohorts and periods gives a first descriptive view of how three-month risk changes with age.

age_rd3m <- loan_windows |>
  summarise(
    loans_at_risk = n(),
    defaults_3m = sum(default_3m),
    .by = age
  ) |>
  mutate(rd3m = defaults_3m / loans_at_risk)

age_rd3m
# A tibble: 34 × 4
     age loans_at_risk defaults_3m    rd3m
   <int>         <int>       <int>   <dbl>
 1     1        112838         517 0.00458
 2     2        110438        1201 0.0109 
 3     3        107602        1873 0.0174 
 4     4        104421        2127 0.0204 
 5     5        101171        2294 0.0227 
 6     6         98074        2472 0.0252 
 7     7         94770        2608 0.0275 
 8     8         91664        2646 0.0289 
 9     9         88723        2685 0.0303 
10    10         85361        2715 0.0318 
# ℹ 24 more rows
ggplot(age_rd3m, aes(age, rd3m)) +
  geom_line(linewidth = 0.9, alpha = 0.8) +
  scale_y_continuous(labels = scales::label_percent(accuracy = 0.1)) +
  labs(
    title = "Three-month default risk by starting age",
    subtitle = "All origination cohorts and calendar periods pooled",
    x = "Starting age (months)",
    y = "RD3M"
  )

The same calculation by cohort shows differences among origination months before adjusting for their age and calendar-period composition.

cohort_rd3m <- loan_windows |>
  summarise(
    loans_at_risk = n(),
    defaults_3m = sum(default_3m),
    .by = cohort
  ) |>
  mutate(rd3m = defaults_3m / loans_at_risk)

cohort_rd3m
# A tibble: 60 × 4
   cohort     loans_at_risk defaults_3m   rd3m
   <date>             <int>       <int>  <dbl>
 1 2020-02-01         44967         751 0.0167
 2 2021-06-01         54431        1761 0.0324
 3 2021-04-01         47063        1443 0.0307
 4 2019-08-01         50404        1072 0.0213
 5 2022-04-01         50541        2303 0.0456
 6 2021-09-01         42559        1699 0.0399
 7 2021-03-01         44131        1330 0.0301
 8 2021-12-01         50526        1963 0.0389
 9 2018-06-01         26712         328 0.0123
10 2021-10-01         54868        2143 0.0391
# ℹ 50 more rows
ggplot(cohort_rd3m, aes(cohort, rd3m)) +
  geom_line(linewidth = 0.9, alpha = 0.8) +
  scale_x_date(date_breaks = "1 year", date_labels = "%Y") +
  scale_y_continuous(labels = scales::label_percent(accuracy = 0.1)) +
  labs(
    title = "Three-month default risk by origination cohort",
    subtitle = "All starting ages and calendar periods pooled",
    x = "Origination cohort",
    y = "RD3M"
  )

Neither profile is an isolated causal effect. Age, cohort and period are observed together and linked exactly by the calendar identity.

Why APC cannot have a unique unconstrained answer

Suppose we write an additive model on the log-odds scale:

\[ y_{a,c,p}=\mu+A_a+C_c+P_p+\varepsilon_{a,c,p}. \]

Because \(p=c+a\), a linear trend can be transferred among AGE, COHORT and PERIOD without changing the fitted values. The data alone cannot decide which component owns that trend.

This post uses a transparent sequential allocation:

\[ \text{mean}\rightarrow AGE\rightarrow COHORT\rightarrow PERIOD \rightarrow\text{residual}. \]

AGE receives the weighted pattern left after the mean; COHORT receives what is left after AGE; PERIOD receives what remains after both. Changing the order can change the individual components, even though the same observations are being described.

Sequential APC decomposition

Some cells contain no defaults, so their raw RD3M is zero and its logit is not finite. We use the small empirical-logit correction:

\[ q=\frac{\text{three-month defaults}+0.5} {\text{eligible loans}+1}. \]

Then transform the adjusted rate to log-odds:

\[ y_{a,c,p}=\operatorname{logit}(q_{a,c,p}) =\log\left(\frac{q_{a,c,p}}{1-q_{a,c,p}}\right). \]

The weights in every step are loans_at_risk, so cells with more eligible loans have more influence.

Start from the weighted mean

apc_decomposition <- rd3m_cells |>
  mutate(
    q = (defaults_3m + 0.5) / (loans_at_risk + 1),
    y_logit = qlogis(q),
    mu = weighted.mean(y_logit, loans_at_risk),
    residual_after_mean = y_logit - mu
  )

The intercept mu is the same for every row. It is a weighted mean on the log-odds scale, not the arithmetic average of portfolio RD3M.

AGE explains the first residual

apc_decomposition <- apc_decomposition |>
  mutate(
    age_effect = weighted.mean(residual_after_mean, loans_at_risk),
    .by = age
  ) |>
  mutate(
    residual_after_age = y_logit - mu - age_effect
  )

Each age effect is the weighted mean of what the global intercept did not explain for that starting age.

COHORT explains what AGE leaves behind

apc_decomposition <- apc_decomposition |>
  mutate(
    cohort_effect = weighted.mean(residual_after_age, loans_at_risk),
    .by = cohort
  ) |>
  mutate(
    residual_after_cohort = y_logit - mu - age_effect - cohort_effect
  )

COHORT now measures systematic differences among origination months after the age profile has already been allocated.

PERIOD explains what AGE and COHORT leave behind

apc_decomposition <- apc_decomposition |>
  mutate(
    period_effect = weighted.mean(residual_after_cohort, loans_at_risk),
    .by = period
  ) |>
  mutate(
    residual_after_period = y_logit - mu - age_effect - cohort_effect -
      period_effect
  )

PERIOD captures common movements among all windows beginning in the same calendar month, after AGE and COHORT have received their sequential contributions.

Reconstruct the adjusted rate

apc_decomposition <- apc_decomposition |>
  mutate(
    fitted_logit = mu + age_effect + cohort_effect + period_effect,
    fitted_rd3m = plogis(fitted_logit),
    reconstructed_q = plogis(fitted_logit + residual_after_period)
  )

glimpse(apc_decomposition)
Rows: 2,040
Columns: 20
$ cohort                <date> 2018-01-01, 2018-01-01, 2018-01-01, 2018-01-01,…
$ age                   <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 1…
$ period                <date> 2018-02-01, 2018-03-01, 2018-04-01, 2018-05-01,…
$ window_end_period     <date> 2018-04-01, 2018-05-01, 2018-06-01, 2018-07-01,…
$ loans_at_risk         <int> 1202, 1168, 1110, 1066, 1013, 961, 927, 904, 866…
$ defaults_3m           <int> 3, 6, 9, 7, 4, 2, 2, 3, 4, 8, 10, 11, 8, 6, 5, 5…
$ rd3m                  <dbl> 0.002495840, 0.005136986, 0.008108108, 0.0065666…
$ q                     <dbl> 0.002909393, 0.005560308, 0.008550855, 0.0070290…
$ y_logit               <dbl> -5.836897, -5.186526, -4.753136, -4.950649, -5.4…
$ mu                    <dbl> -3.806735, -3.806735, -3.806735, -3.806735, -3.8…
$ residual_after_mean   <dbl> -2.03016238, -1.37979123, -0.94640163, -1.143914…
$ age_effect            <dbl> -1.707365388, -0.802500336, -0.337459829, -0.182…
$ residual_after_age    <dbl> -0.3227970, -0.5772909, -0.6089418, -0.9614343, …
$ cohort_effect         <dbl> -0.9036941, -0.9036941, -0.9036941, -0.9036941, …
$ residual_after_cohort <dbl> 0.58089707, 0.32640317, 0.29475226, -0.05774028,…
$ period_effect         <dbl> 0.58089707, -0.39650715, 0.27505240, -0.06870897…
$ residual_after_period <dbl> 0.000000000, 0.722910323, 0.019699859, 0.0109686…
$ fitted_logit          <dbl> -5.836897, -5.909436, -4.772836, -4.961618, -5.2…
$ fitted_rd3m           <dbl> 0.002909393, 0.002706372, 0.008385452, 0.0069529…
$ reconstructed_q       <dbl> 0.002909393, 0.005560308, 0.008550855, 0.0070290…

The row-level identity on the transformed scale is exact:

\[ y=\mu+AGE+COHORT+PERIOD+residual. \]

apc_decomposition |>
  summarise(max_rate_difference = max(abs(q - reconstructed_q)))
# A tibble: 1 × 1
  max_rate_difference
                <dbl>
1            5.55e-17

The difference should be zero apart from numerical precision.

Reading the three effects

The components are additive in log-odds. Exponentiating one effect gives an odds multiplier for three-month default: 1x means no change relative to the relevant sequential baseline, 1.5x means 50% higher odds and 0.5x means half the odds.

apc_decomposition |>
  distinct(age, age_effect) |>
  ggplot(aes(age, exp(age_effect))) +
  geom_hline(yintercept = 1, linetype = "dashed") +
  geom_line(linewidth = 0.9, alpha = 0.8) +
  scale_y_continuous(
    labels = scales::label_number(accuracy = 0.01, suffix = "x")
  ) +
  labs(
    title = "AGE effect",
    subtitle = "Three-month odds after removing the common mean",
    x = "Starting age (months)",
    y = "Odds multiplier"
  )

apc_decomposition |>
  distinct(cohort, cohort_effect) |>
  ggplot(aes(cohort, exp(cohort_effect))) +
  geom_hline(yintercept = 1, linetype = "dashed") +
  geom_line(linewidth = 0.9, alpha = 0.8) +
  scale_x_date(date_breaks = "1 year", date_labels = "%Y") +
  scale_y_continuous(
    labels = scales::label_number(accuracy = 0.01, suffix = "x")
  ) +
  labs(
    title = "COHORT effect",
    subtitle = "Three-month odds after AGE is removed",
    x = "Origination cohort",
    y = "Odds multiplier"
  )

apc_decomposition |>
  distinct(period, period_effect) |>
  ggplot(aes(period, exp(period_effect))) +
  geom_hline(yintercept = 1, linetype = "dashed") +
  geom_line(linewidth = 0.9, alpha = 0.8) +
  scale_x_date(date_breaks = "1 year", date_labels = "%Y") +
  scale_y_continuous(
    labels = scales::label_number(accuracy = 0.01, suffix = "x")
  ) +
  labs(
    title = "PERIOD effect",
    subtitle = "Three-month odds after AGE and COHORT are removed",
    x = "Window start",
    y = "Odds multiplier"
  )

These are relative effects, not three separate default rates. A peak in AGE identifies starting ages with higher three-month odds. A peak in COHORT identifies origination months with higher odds after AGE is removed. A peak in PERIOD marks starting windows with higher odds after both previous allocations.

Returning to the RD3M scale

To express the decomposition in probability points, apply plogis() after each sequential step and define each contribution as the change from the previous step.

apc_decomposition <- apc_decomposition |>
  mutate(
    risk_base = plogis(mu),
    risk_after_age = plogis(mu + age_effect),
    risk_after_cohort = plogis(mu + age_effect + cohort_effect),
    risk_after_period = fitted_rd3m,
    contribution_age = risk_after_age - risk_base,
    contribution_cohort = risk_after_cohort - risk_after_age,
    contribution_period = risk_after_period - risk_after_cohort,
    # The final display residual closes the gap to the unadjusted cell RD3M.
    contribution_residual = rd3m - risk_after_period,
    reconstructed_rd3m = risk_base + contribution_age +
      contribution_cohort + contribution_period + contribution_residual
  )

The first four terms reconstruct the fitted value. The final residual returns to the raw cell RD3M rather than to the continuity-adjusted q.

apc_decomposition |>
  summarise(
    max_rd3m_difference = max(abs(rd3m - reconstructed_rd3m))
  )
# A tibble: 1 × 1
  max_rd3m_difference
                <dbl>
1            3.47e-18

Reconstructing portfolio RD3M

Pool the cell-level contributions within each starting period using the same eligible-loan weights.

portfolio_apc <- apc_decomposition |>
  summarise(
    base = weighted.mean(risk_base, loans_at_risk),
    age_component = weighted.mean(contribution_age, loans_at_risk),
    cohort_component = weighted.mean(contribution_cohort, loans_at_risk),
    period_component = weighted.mean(contribution_period, loans_at_risk),
    residual_component = weighted.mean(contribution_residual, loans_at_risk),
    .by = period
  ) |>
  mutate(
    fitted_rd3m = base + age_component + cohort_component + period_component,
    reconstructed_rd3m = fitted_rd3m + residual_component
  ) |>
  left_join(
    portfolio_rd3m |>
      select(period, observed_rd3m = rd3m),
    by = "period"
  ) |>
  arrange(period)

portfolio_apc
# A tibble: 93 × 9
   period       base age_component cohort_component period_component
   <date>      <dbl>         <dbl>            <dbl>            <dbl>
 1 2018-02-01 0.0217      -0.0177          -0.00238         0.00128 
 2 2018-03-01 0.0217      -0.0146          -0.00419        -0.000976
 3 2018-04-01 0.0217      -0.0118          -0.00570         0.00133 
 4 2018-05-01 0.0217      -0.0102          -0.00649        -0.000335
 5 2018-06-01 0.0217      -0.00887         -0.00700        -0.00207 
 6 2018-07-01 0.0217      -0.00766         -0.00747        -0.00221 
 7 2018-08-01 0.0217      -0.00659         -0.00780        -0.00189 
 8 2018-09-01 0.0217      -0.00574         -0.00792        -0.00286 
 9 2018-10-01 0.0217      -0.00489         -0.00787        -0.00145 
10 2018-11-01 0.0217      -0.00446         -0.00754        -0.000414
# ℹ 83 more rows
# ℹ 4 more variables: residual_component <dbl>, fitted_rd3m <dbl>,
#   reconstructed_rd3m <dbl>, observed_rd3m <dbl>

The direct portfolio calculation and the reconstructed calculation use the same loans and outcomes, so they meet exactly after adding the residual.

portfolio_apc |>
  summarise(
    max_rd3m_difference = max(
      abs(observed_rd3m - reconstructed_rd3m)
    )
  )
# A tibble: 1 × 1
  max_rd3m_difference
                <dbl>
1            6.94e-18

First compare observed RD3M with the APC fit before the residual is added.

ggplot(portfolio_apc, aes(period)) +
  geom_line(
    aes(y = observed_rd3m, colour = "Observed RD3M"),
    linewidth = 0.9,
    alpha = 0.8
  ) +
  geom_line(
    aes(y = fitted_rd3m, colour = "APC without residual"),
    linewidth = 0.9,
    alpha = 0.8
  ) +
  scale_colour_manual(
    values = c(
      "Observed RD3M" = "#2F6690",
      "APC without residual" = "#8A8A8A"
    )
  ) +
  scale_x_date(date_breaks = "1 year", date_labels = "%Y") +
  scale_y_continuous(labels = scales::label_percent(accuracy = 0.1)) +
  labs(
    title = "Observed and APC-fitted portfolio risk",
    subtitle = "The remaining gap is the portfolio-level residual",
    x = "Window start",
    y = "RD3M",
    colour = NULL
  )

Now prepare the five additive contributions.

portfolio_components <- portfolio_apc |>
  select(
    period,
    Base = base,
    AGE = age_component,
    COHORT = cohort_component,
    PERIOD = period_component,
    Residual = residual_component
  ) |>
  pivot_longer(
    cols = -period,
    names_to = "component",
    values_to = "contribution"
  ) |>
  mutate(
    component = factor(
      component,
      levels = c("Base", "AGE", "COHORT", "PERIOD", "Residual")
    )
  )

component_colours <- c(
  "Base" = "#c9c9c9",
  "AGE" = "#4C78A8",
  "COHORT" = "#59A14F",
  "PERIOD" = "#E3A72F",
  "Residual" = "#E45756"
)
ggplot(
  portfolio_components,
  aes(period, contribution, fill = component)
) +
  geom_hline(yintercept = 0, linewidth = 0.3, colour = "#808080") +
  # Reverse the default stack so Base is always added first from zero.
  geom_col(
    width = 25,
    alpha = 0.85,
    position = position_stack(reverse = TRUE)
  ) +
  geom_line(
    data = portfolio_apc,
    aes(period, observed_rd3m),
    colour = "#2F6690",
    linewidth = 0.9,
    alpha = 0.75,
    inherit.aes = FALSE
  ) +
  scale_fill_manual(values = component_colours) +
  scale_x_date(date_breaks = "1 year", date_labels = "%Y") +
  scale_y_continuous(labels = scales::label_percent(accuracy = 0.1)) +
  labs(
    title = "What drives rolling three-month portfolio risk?",
    subtitle = "Base is added first; the other components move RD3M above or below it",
    x = "Window start",
    y = "Contribution to RD3M",
    fill = NULL
  )

The blue line is observed RD3M. Base is the first component added and therefore always occupies the same range from zero to its constant value. AGE, COHORT, PERIOD and the residual are stacked afterwards: positive contributions extend above the baseline and negative contributions extend below zero. Together, the five components add exactly to observed RD3M.

The components are easier to inspect separately.

ggplot(
  portfolio_components,
  aes(period, contribution, fill = component)
) +
  geom_hline(yintercept = 0, linewidth = 0.3, colour = "#808080") +
  geom_col(width = 25, alpha = 0.9, show.legend = FALSE) +
  facet_wrap(vars(component), ncol = 3) +
  scale_fill_manual(values = component_colours) +
  scale_x_date(date_breaks = "2 years", date_labels = "%Y") +
  scale_y_continuous(labels = scales::label_percent(accuracy = 0.1)) +
  labs(
    title = "How each APC component changes through time",
    subtitle = "All panels use the same RD3M contribution scale",
    x = "Window start",
    y = "Contribution to RD3M"
  )

How overlapping windows change the reading

Rolling three-month windows are appropriate for this target, but they change the temporal meaning of the plots.

Suppose defaults increase only in April. That event can raise the RD3M windows beginning in February, March and April because all three include April. A PERIOD movement labelled February therefore describes the February–April window; it does not prove that the underlying shock occurred in February.

Three consequences follow:

  1. Adjacent points are mechanically related. They share two of their three outcome months and many of the same loans.
  2. Short shocks appear wider. A movement concentrated in one outcome month can affect as many as three consecutive starting periods.
  3. The components describe forward risk. AGE, COHORT and PERIOD explain the probability of default over a three-month window, not an isolated monthly event rate.

None of these properties makes RD3M incorrect. They are consequences of asking a rolling forward-looking question. A rolling RD12M series has the same structure with even more overlap.

Methodological details and caveats

Window start versus outcome month

Throughout the post, period means the start of the three-month window. Its outcome is observed from period through window_end_period. Using the end month instead would shift the labels by two months without changing the underlying windows, so the convention must remain explicit when results are compared with external events.

Complete outcomes and right censoring

Non-defaulted loans enter a cell only when all three outcome months are observed. This prevents recent loans from being counted automatically as non-defaults. Nevertheless, early repayment or another non-default exit can still change which loans remain observable. If exit is related to credit quality, the resulting composition deserves separate analysis.

Descriptive decomposition versus statistical inference

The construction here is descriptive. The algebraic reconstruction does not require consecutive windows to be independent. Independence matters if standard errors, confidence intervals or hypothesis tests are later attached to the effects. In that case, treating every rolling window as an independent record would understate uncertainty. Resampling or covariance estimation should preserve the loan-level and temporal dependence.

Sequential allocation

The APC identity prevents a unique unconstrained separation. This post resolves the ambiguity by choosing AGE → COHORT → PERIOD. The result should therefore be read as one transparent allocation of observed variation, not the only possible one. Repeating the calculation under alternative orders is a useful sensitivity check.

Why the weighted effects average to zero

There is no additional centring step. mu makes the first residual have weighted mean zero. AGE is the weighted group mean of that residual, so its overall weighted mean is also zero. Subtracting AGE leaves another zero-mean residual, and the same argument repeats for COHORT and PERIOD.

apc_decomposition |>
  summarise(
    age = weighted.mean(age_effect, loans_at_risk),
    cohort = weighted.mean(cohort_effect, loans_at_risk),
    period = weighted.mean(period_effect, loans_at_risk)
  )
# A tibble: 1 × 3
        age   cohort   period
      <dbl>    <dbl>    <dbl>
1 -4.42e-16 3.22e-18 5.33e-18

Monetary weights

loans_at_risk gives every loan the same influence. Production credit-risk work often weights by balance, exposure or EAD. Bondora provides original loan amount, but a genuine balance-weighted history would require historical balances at each window start rather than a current snapshot.

A compact practical version

Once the definitions are understood, the essential calculation is short:

rd3m_practical <- loans |>
  select(loan_id, cohort, default_age, last_start_age) |>
  uncount(last_start_age, .id = "age") |>
  mutate(
    period = cohort %m+% months(age),
    default_3m = as.integer(
      !is.na(default_age) &
        default_age >= age &
        default_age <= age + 2L
    )
  ) |>
  summarise(
    loans_at_risk = n(),
    defaults_3m = sum(default_3m),
    .by = c(cohort, age, period)
  ) |>
  mutate(rd3m = defaults_3m / loans_at_risk)

rd3m_practical
# A tibble: 2,040 × 6
   cohort       age period     loans_at_risk defaults_3m    rd3m
   <date>     <int> <date>             <int>       <int>   <dbl>
 1 2020-02-01     1 2020-03-01          2054           4 0.00195
 2 2020-02-01     2 2020-04-01          2017           8 0.00397
 3 2020-02-01     3 2020-05-01          1979          16 0.00808
 4 2020-02-01     4 2020-06-01          1934          30 0.0155 
 5 2020-02-01     5 2020-07-01          1883          38 0.0202 
 6 2020-02-01     6 2020-08-01          1827          39 0.0213 
 7 2020-02-01     7 2020-09-01          1757          37 0.0211 
 8 2020-02-01     8 2020-10-01          1700          37 0.0218 
 9 2020-02-01     9 2020-11-01          1655          37 0.0224 
10 2020-02-01    10 2020-12-01          1588          28 0.0176 
# ℹ 2,030 more rows

Everything after this table concerns how its variation is allocated and interpreted. The target itself remains simple: defaults during the next three months divided by loans alive at the beginning of a fully observable window.

References