How DMAR’s Functions Fit Together: A Deep Dive

Ken Kelley

September 2026

Why This Vignette Exists

DMAR has grown to encompass several large families of functions: effect size point estimates, exact-noncentral distribution confidence intervals, asymptotic-variance utilities, design-stage expected-value calculators, accuracy in parameter estimation (AIPE) sample size planners, reliability coefficients, and agreement / multivariate measures. Each function is documented in isolation, but the package’s value is in how the functions compose. This vignette walks through the package’s architecture from the user’s perspective and shows the five high-traffic composition patterns.

The Four Building Blocks

For any single population quantity \(\theta\) the package supports four operations:

       point estimate      sample size                noncentral CI            asymptotic
       /  variance         planning                                            building blocks
       ┌──────────────┐    ┌──────────────────┐      ┌───────────────────┐    ┌─────────────────┐
\(\theta\) → │ smd, omega², r│ →  │ ss_aipe_smd,     │ →    │ ci_smd, ci_r,     │ ←  │ var_smd, var_r, │
       │ R², icc, ...  │    │ ss_aipe_R2, ...  │      │ ci_omega_squared, │    │ var_omega_squared│
       └──────────────┘    └──────────────────┘      │ ci_R2, ...        │    │ ...             │
              ↑                       ↑              └───────────────────┘            ↑
              │                       │                       │                       │
              │                       │                       │                       │
              │                ┌──────┴─────────────┐         │                       │
              │                │ expected_smd,      │         │                       │
              │                │ expected_r,        │ ────────┘                       │
              │                │ expected_R2 ...    │                                 │
              │                └────────────────────┘                                 │
              │                                                                       │
              └──────────────── transforms of each other ─────────────────────────────┘
              (smd ↔ cles ↔ proportion_of_superiority ↔ nnt_from_smd; r ↔ partial_r)

The four columns are tightly coupled: every AIPE planner is built from a variance utility; every CI is built from a noncentral distribution; every expected-value function is the design-stage analog of an estimator. Once you internalize the layout, asking “what is the recommended workflow for quantity X” reduces to asking “which row of X’s variance / CI / expected value / AIPE planner do I need.”

1. Variance Utility → AIPE Planner

The most pervasive composition is var_* feeding ss_aipe_*. For any quantity \(\theta\) with a known asymptotic variance \(\mathrm{Var}(\hat\theta \mid \theta, n)\), the AIPE planner solves for the smallest \(n\) such that

\[ z_{1-\alpha/2} \cdot \sqrt{\mathrm{Var}(\hat\theta \mid \theta, n)} \le w/2, \]

where \(w\) is the target full width of the CI. The AIPE framework, planning for a sufficiently narrow confidence interval rather than for power, is developed in Kelley and Maxwell (2003) and Maxwell, Kelley, and Rausch (2008); see also Maxwell, Delaney, and Kelley (2027) for the model comparison treatment. Concretely:

# SMD (Cohen's d):     var_smd        → ss_aipe_smd
# Pearson r (J = 0):   var_r          → ss_aipe_partial_r (J = 1, simple case)
# Partial r:           var_partial_r  → ss_aipe_partial_r
# Semipartial r:       var_semipartial_r → ss_aipe_semipartial_r
# Squared multi R:     var_R2         → ss_aipe_R2
# Cliff's delta:       (max-variance bound) → ss_aipe_cliff_delta
# omega²:              var_omega_squared → ss_aipe_omega_squared (noncentral F)
# Cronbach alpha:      var_alpha      → ss_aipe_reliability
# ICC:                 var_icc        → ss_aipe_icc
# CV:                  var_cv         → ss_aipe_cv
# Indirect effect ab:  var_indirect_effect → ss_aipe_indirect_effect

This pairing is the package’s most common usage pattern. The variance utility is the design-stage primitive; the AIPE planner is the inverse, returning the \(n\) that achieves a target precision.

Worked Example: Pearson r

We plan a study where we expect \(\rho \approx 0.30\) and want a 95% CI of full width 0.20.

# Step 1: ask the variance utility for the sampling variance at a candidate n.
var_r(rho = 0.30, n = 100)
term value
var_r_normal 0.00836
var_fisher_z 0.0103

# Step 2: the AIPE planner inverts the same family of formulas to give
# the n needed for the target width. DMAR has no zero-order Pearson-r
# planner; the partial-r planner at J = 1 is the simple-correlation case.
ss_aipe_partial_r(rho = 0.30, J = 1, width = 0.20)
term value
necessary_N 321
expected_width 0.2
rho 0.3
J 1
width_target 0.2
conf_level 0.95

Confidence level: 95%

Both sides use the same normal-theory \((1 - \rho^2)^2\) variance family (Fisher, 1915; Olkin & Finn, 1995), differing only by an off-by-one in the denominator degrees of freedom (\(n - 1\) for var_r(), \(n - J - 1\) for the planner), so the planner’s output is consistent with what var_r() returns at that recommended \(n\).

2. Expected-Value Function → Design-Stage Substitution

For estimators with non-trivial bias, the package provides expected_*() functions that report \(\mathrm{E}[\hat\theta \mid \theta, n]\) and the size of the bias. These are useful in two situations:

  1. Reporting: at the analysis stage, the difference between \(\hat\theta\) and \(\mathrm{E}[\hat\theta]\) tells the analyst how much of the observed value is bias.
  2. Design-stage planning: when a sample size planner takes \(\theta\) as input but the realized data will yield a biased sample \(\hat\theta\), substituting \(\mathrm{E}[\hat\theta]\) in place of \(\theta\) corrects the planner’s prediction.

The package’s expected-value functions are:

Quantity Bias Helper
Pearson \(r\) downward (small \(n\)) expected_r()
Partial \(r\) downward (grows with \(J\)) expected_partial_r()
Squared multi \(R^2\) upward (always) expected_R2()
Cohen’s \(d\) upward (small \(n\)) expected_smd()

Worked Example: The Standardized Mean Difference

# Population delta = 0.4, planned n = 40 per group.
# Naive planning assumes the observed d will average 0.40; in fact
# the sample d will average slightly larger.
expected_smd(delta = 0.4, n_1 = 40)
delta n_1 n_2 expected_smd bias j_correction
0.4 40 40 0.404 0.0039 0.99

The reported j_correction is the Hedges (1981) factor \(J(\mathit{df})\) that maps Cohen’s \(d\) to Hedges’ \(g\). The same factor is applied internally by smd(..., unbiased = TRUE), which returns the bias-corrected estimate \(g = J(\mathit{df})\,d\), so the same constant runs throughout the \(d\)-family of functions for consistency.

3. Effect Size Transforms: smd ↔︎ cles ↔︎ proportion_of_superiority ↔︎ nnt

A single standardized mean difference (Cohen’s \(d\)) can be reported on five common scales without losing information:

Scale Formula Function in DMAR
Standardized mean difference \(d\) smd(), ci_smd()
Proportion of superiority (Cohen’s \(U_3\)) \(\Phi(d)\) proportion_of_superiority()
Common language (independent groups) \(\Phi(d / \sqrt 2)\) cles()
Success-rate difference \(2 \Phi(d / \sqrt 2) - 1\) (intermediate)
Number needed to treat \(1 / [2 \Phi(d / \sqrt 2) - 1]\) nnt_from_smd()

Each downstream function accepts either a point estimate of \(d\) or the same plus group sizes for an exact noncentral \(t\) CI, and propagates the CI bounds through the monotone transform. Because the transforms are monotone, the coverage probability is preserved exactly (no extra approximation is introduced).

Walked Example at Three Reference d Values

for (d in c(0.2, 0.5, 0.8)) {
  cat(sprintf(
    "d = %.1f  ->  PS = %.2f   CLES = %.2f   NNT = %.1f\n",
    d,
    proportion_of_superiority(d)$value[2],
    cles(d)$value[2],
    nnt_from_smd(d)$value[3]
  ))
}
#> d = 0.2  ->  PS = 0.58   CLES = 0.56   NNT = 8.9
#> d = 0.5  ->  PS = 0.69   CLES = 0.64   NNT = 3.6
#> d = 0.8  ->  PS = 0.79   CLES = 0.71   NNT = 2.3

For an applied report, the same effect can be communicated on whichever scale is most natural for the audience without recomputing anything.

4. Ordinal Alternatives: cliff_delta, vargha_delaney_A, probability_of_superiority_paired

When the dependent variable is ordinal, skewed, or you want a nonparametric effect size, the same role is played by U-statistic estimators:

Independent groups Function
Probability of superiority cles() (under bivariate normality)
Stochastic dominance cliff_delta() (ordinal, U-statistic)
Vargha-Delaney \(A = (\delta + 1)/2\), vargha_delaney_A()
Paired data probability_of_superiority_paired()

Each returns an analytic CI; cliff_delta() uses the Cliff (1996) U-statistic variance, and probability_of_superiority_paired() uses the Brunner-Munzel (2000) within-pair variance, all with a Fisher-style \(\arctanh\) transformation so the bounds stay in \([-1, 1]\) (for \(\delta\)) or \([0, 1]\) (for \(P_S\)).

5. Reliability Composition

The reliability family follows the same four-block pattern but with one extra wrinkle: alpha is the workhorse, omega is the modern default, and \(H\) is the upper bound that sets the ceiling of what is achievable with the same indicators:

# Point estimate:      reliability_alpha(), reliability_omega(), reliability_H()
# CI:                  reliability() (general wrapper)
# Asymptotic variance: var_alpha()
# Design-stage AIPE:   ss_aipe_reliability()

\(H\) is uniformly \(\ge \omega \ge \alpha\) for unidimensional indicators; if reliability_H() is far larger than reliability_alpha(), the equally weighted composite is wasting reliability that an optimally weighted composite would capture.

6. The Mediation Pipeline

For a simple three-variable mediation model \(X \to M \to Y\) with indirect effect \(ab\):

# Asymptotic variance: var_indirect_effect(a, b, var_a, var_b, cov_ab)
#   - returns Sobel, Aroian, Goodman, and second-order delta variances.
# AIPE planning:       ss_aipe_indirect_effect(a, b, width, method)
#   - method = "closed_form": the delta method (Sobel, 1982) width
#   - method = "monte_carlo": the Monte Carlo interval width, simulated
#     at each candidate n (Tofighi & MacKinnon, 2011)

The two methods correspond to the two main approaches in the mediation literature; var_indirect_effect() makes the underlying SE formulas explicit so the user can see what the planner is solving against.

7. Multivariate and Agreement Utilities

The package’s measurement family is small but covers the common needs:

These are stand-alone utilities; they do not feed into AIPE planners yet, although the variance components implied by variance_components_mls() do compose with var_icc() and ss_aipe_icc() for generalizability- theory planning.

Argument-Name Conventions

When using multiple DMAR functions in a single pipeline, the argument-name conventions are worth knowing:

Quantity Argument name(s)
Population correlation rho (Greek; new code)
Sample correlation r (Roman; existing code)
Population SMD delta
Sample SMD smd
Group sample sizes n_1, n_2
Total sample size n (single-sample), N (some legacy)
Number of controls (partial r) J
Number of items p_items
Number of predictors p
Number of raters (ICC) k
Effect df / error df df_effect, df_error
Target CI width width
Which-width specifier which_width ∈ {“Full”, “Lower”, “Upper”}
Confidence level conf_level
Per-tail alphas alpha_lower, alpha_upper
Coverage assurance probability assurance (in (0.5, 1))

Pattern: lower-case Greek letters denote population values (rho, delta, alpha); Roman letters denote sample estimators (r, smd). When in doubt, the function’s @param line in the help file is explicit about which is wanted.

Two Consistent Return Shapes

DMAR functions return one of two tidy data-frame shapes:

  1. Long form: data.frame(term, value): used by all variance utilities, AIPE planners, effect size point estimates, and CI helpers. The term column names the quantity; value is the numeric value.
  2. Wide form: one row, many columns: used by omega_squared(), eta_squared_partial(), and similar effect size functions that return one row per effect in a factorial design (with columns effect, omega_squared, F_value, df_effect, df_error, N).

Pipeline code can use dplyr::filter(term == "...") for the long form or column-access for the wide form. The two shapes compose: the wide form can always be reshaped to long with tidyr::pivot_longer().

Summary

DMAR’s architecture rests on four building blocks for any quantity \(\theta\): a point estimator with variance, an exact-noncentral CI, an expected-value function for design-stage planning, and an AIPE sample- size planner. The functions in each block compose horizontally (the same \(\theta\) shows up in all four) and vertically (within a block, related transforms like \(d \to \mathrm{CLES} \to U_3 \to \mathrm{NNT}\) are explicit). Knowing which block a given task lives in is half the battle; the function name and help file complete the picture.

mirror server hosted at Truenetwork, Russian Federation.