A year-over-year percent change is the most-quoted number in
open-data reporting and the least qualified. yoy() computes
it, and refuses to compute it in the three cases where the figure would
describe something other than the data.
Ontario’s inmate data is published one row per placement, by fiscal year and by group. A realistic shape:
seg <- data.frame(
EndFiscalYear = rep(2019:2023, each = 2),
Gender = rep(c("Female", "Male"), 5),
Number_Of_Placements = c(31, 402, 28, 377, 12, 190, 19, 268, 24, 331)
)Segregation placements rising is bad news, so the direction is declared. That affects colour and the verdict column, never the arithmetic.
y <- yoy(seg,
value = Number_Of_Placements,
period = EndFiscalYear,
by = "Gender",
direction = "lower_is_better"
)
y
#> Gender EndFiscalYear Number_Of_Placements previous change change % interval
#> ────── ───────────── ──────────────────── ──────── ────── ──────── ────────────
#> Female 2019 31 — — — —
#> ↳ percent withheld: no comparison period
#> Female 2020 28 31 -3 ▼ -9.7% [-48%, +56%]
#> Female 2021 12 28 -16 ▼ -57.1% [-80%, -13%]
#> Female 2022 19 12 7 — —
#> ↳ percent withheld: base below 20
#> Female 2023 24 19 5 — —
#> ↳ percent withheld: base below 20
#> Male 2019 402 — — — —
#> ↳ percent withheld: no comparison period
#> Male 2020 377 402 -25 ▼ -6.2% [-19%, +8%]
#> Male 2021 190 377 -187 ▼ -49.6% [-58%, -40%]
#> Male 2022 268 190 78 ▲ +41.1% [+17%, +71%]
#> Male 2023 331 268 63 ▲ +23.5% [+5%, +46%]
#>
#> lag 1 period · units: count · 95% exact rate-ratio interval · percent withheld below a base of 20 · lower is betterThe interval is exact. Conditional on the two periods’ total, the
current count is binomial, so the ratio of the two has a Clopper-Pearson
interval, which is the same construction
stats::poisson.test() uses:
With a year missing, comparing row i with row i - 1 compares 2023 with 2021 and labels it a one-year change. Matching on the period’s own value cannot do that:
gap <- data.frame(year = c(2019, 2020, 2022, 2023), n = c(100, 120, 140, 150))
yoy(gap, value = n, period = year)
#> year n previous change change % interval
#> ──── ─── ──────── ────── ──────── ────────────
#> 2019 100 — — — —
#> ↳ percent withheld: no comparison period
#> 2020 120 100 20 ▲ +20.0% [-9%, +58%]
#> 2021 — 120 — — —
#> 2022 140 — — — —
#> ↳ percent withheld: no comparison period
#> 2023 150 140 10 ▲ +7.1% [-15%, +36%]
#>
#> lag 1 period · units: count · 95% exact rate-ratio interval · percent withheld below a base of 202021 appears as a row with no value, and 2022 has no comparison period rather than a quietly three-year one.
Two placements becoming twenty is a 900% rise and also nothing at all. The number describes how small the denominator was:
tiny <- data.frame(year = 2019:2021, n = c(2, 20, 25))
yoy(tiny, value = n, period = year)
#> year n previous change change % interval
#> ──── ── ──────── ────── ──────── ─────────────
#> 2019 2 — — — —
#> ↳ percent withheld: no comparison period
#> 2020 20 2 18 — —
#> ↳ percent withheld: base below 20
#> 2021 25 20 5 ▲ +25.0% [-33%, +137%]
#>
#> lag 1 period · units: count · 95% exact rate-ratio interval · percent withheld below a base of 20The change of +18 is still reported, and so is the
direction: those are facts. Only the percent is withheld, with the
reason attached. Lower the gate if the small base is the point:
If a rate moves from 4% to 5% that is one percentage point. Calling it 25% answers a different question:
rate <- data.frame(year = 2019:2023, share = c(4.1, 4.6, 5.2, 5.0, 5.4))
yoy(rate, value = share, period = year, units = "percent")
#> year share previous change points
#> ──── ───── ──────── ────── ────────
#> 2019 4.1 — — —
#> ↳ percent withheld: no comparison period
#> 2020 4.6 4.1 0.5 ▲ +0.5pp
#> 2021 5.2 4.6 0.6 ▲ +0.6pp
#> 2022 5.0 5.2 -0.2 ▼ -0.2pp
#> 2023 5.4 5.0 0.4 ▲ +0.4pp
#>
#> lag 1 period · units: percent · percentage-point change, not percent of a percentThe column is named points, and no count interval is
offered, because these are not counts.
A monthly series compares with the same month a year earlier. The lag follows the series’ own frequency, so January is never put against December:
yoy_summary(y)
#> Gender from to first last total_pct cagr_pct periods up down flat
#> 1 Female 2019 2023 31 24 -22.58065 -6.197938 5 0 2 0
#> 2 Male 2019 2023 402 331 -17.66169 -4.742214 5 2 2 0cagr_pct is the compound rate per period: applying it
across the span returns total_pct exactly.
yoy() compares counts. A count is not comparable across
places of different size, or across years whose populations differ, and
the fix is a rate – but a rate is not made trustworthy by being a
rate.
stops <- data.frame(
division = c("North", "South", "East"),
stops = c(412, 77, 3),
residents = c(120000, 41000, 9500))
rate(stops, stops, residents, by = "division", per = "100k")
#> ── Rate per 100,000, 95% exact Poisson interval ──────────────────
#> division count population rate lower upper flag
#> East 3 9500 31.57895 6.512338 92.28708 <NA>
#> North 412 120000 343.33333 310.977110 378.14142 <NA>
#> South 77 41000 187.80488 148.212673 234.72387 <NA>
#> ──────────────────────────────────────────────────────────────────East’s interval is wider than its own estimate. That is the honest reading of three events, and it is the reason the interval is printed beside the rate rather than left to the reader to imagine. The interval is the exact Poisson one, so a count of zero has a lower limit of exactly zero instead of a negative rate.
A share is a different quantity. Its denominator is the total of the same events, not a population, so shares over a complete grouping sum to 100:
share(stops, stops, by = "division")
#> ── Share of 492, 95% Wilson interval ─────────────────────────────
#> division count total share lower upper
#> East 3 492 0.6097561 0.2075843 1.777215
#> North 412 492 83.7398374 80.2200225 86.736863
#> South 77 492 15.6504065 12.7074531 19.125597
#> ──────────────────────────────────────────────────────────────────“40% of stops” and “40 stops per 1,000 residents” are different claims, and a table that labels one as the other is wrong however carefully the arithmetic was done. That is why these are separate functions.
For the change in a rate, both the counts and the denominators move:
d <- data.frame(
year = rep(2021:2023, each = 2),
division = rep(c("North", "South"), 3),
stops = c(400, 70, 430, 66, 455, 61),
residents = c(120000, 41000, 122000, 41500, 125000, 42000))
rate_change(d, stops, residents, year, by = "division", per = "100k")
#> ── Change in rate per 100,000, lag 1, 95% exact conditional interval
#> division year count population rate previous_rate rate_ratio pct_change
#> North 2021 400 120000 333.3333 NA NA NA
#> North 2022 430 122000 352.4590 333.3333 1.0573770 5.737705
#> North 2023 455 125000 364.0000 352.4590 1.0327442 3.274419
#> South 2021 70 41000 170.7317 NA NA NA
#> South 2022 66 41500 159.0361 170.7317 0.9314974 -6.850258
#> South 2023 61 42000 145.2381 159.0361 0.9132395 -8.676046
#> pct_lower pct_upper flag
#> NA NA no comparison period
#> -7.937165 21.46533 <NA>
#> -9.679876 18.10199 <NA>
#> NA NA no comparison period
#> -34.473784 32.29053 <NA>
#> -36.596300 31.35600 <NA>
#> ──────────────────────────────────────────────────────────────────North’s count rose, and so did its population, so the rate change is
smaller than the count change – which is the whole reason to look at the
rate. The interval conditions on the two counts and corrects for the
ratio of the two exposures; with equal populations it reduces to exactly
the conditional-binomial interval yoy() uses above.
The same object writes to six formats. The format follows the file name.
dir <- tempdir()
for (ext in c("csv", "tsv", "json", "md", "html", "pdf")) {
f <- file.path(dir, paste0("placements.", ext))
yoy_write(y, f, title = "Placements by gender")
cat(sprintf("%-5s %6d bytes\n", ext, file.size(f)))
}
#> csv 931 bytes
#> tsv 931 bytes
#> json 3374 bytes
#> md 1630 bytes
#> html 6852 bytes
#> pdf 5098 bytesThe delimited writers lead with the settings as comments, and carry
the flag column, so a withheld percentage stays marked as
withheld wherever it lands:
cat(yoy_csv(y, NULL, digits = 1L))
#> # value: Number_Of_Placements; period: EndFiscalYear; by: Gender; lag: 1; units: count; interval: 95% exact rate ratio; percent withheld below a base of 20; direction: lower is better
#> Gender,EndFiscalYear,value,previous,change,pct_change,flag,pct_lower,pct_upper,verdict
#> Female,2019,31,,,,no comparison period,,,
#> Female,2020,28,31,-3,-9.7,,-47.8,55.6,better
#> Female,2021,12,28,-16,-57.1,,-80.1,-13,better
#> Female,2022,19,12,7,,base below 20,,,worse
#> Female,2023,24,19,5,,base below 20,,,worse
#> Male,2019,402,,,,no comparison period,,,
#> Male,2020,377,402,-25,-6.2,,-18.7,8.2,better
#> Male,2021,190,377,-187,-49.6,,-57.9,-39.8,better
#> Male,2022,268,190,78,41.1,,16.7,70.8,worse
#> Male,2023,331,268,63,23.5,,4.8,45.6,worseMarkdown, for a report:
cat(yoy_markdown(yoy(gap, value = n, period = year), NULL))
#> | year | n | previous | change | change % | interval | note |
#> | ---- | --- | -------- | -----: | -------: | -----------: | -------------------- |
#> | 2019 | 100 | — | — | — | | no comparison period |
#> | 2020 | 120 | 100 | 20 | +20.0% | [-9%, +58%] | |
#> | 2021 | — | 120 | — | — | | |
#> | 2022 | 140 | — | — | — | | no comparison period |
#> | 2023 | 150 | 140 | 10 | +7.1% | [-15%, +36%] | |
#>
#> _value: n; period: year; lag: 1; units: count; interval: 95% exact rate ratio; percent withheld below a base of 20_The HTML is a single self-contained file: it makes no network request, so the page renders later exactly as it rendered when the capsule was sealed. It defines its palette as custom properties and carries a dark variant. The PDF is drawn on R’s own device and paginates rather than truncating, so neither needs a package beyond base R.
"safe" is a blue/orange pair that survives both common
forms of colour blindness; "mono" emits no colour at all,
for print.