Cookbook: Incidence (Binary) Outcome, End to End

One complete, runnable script for a binary (“incidence”) outcome — did the event happen or not. Same four steps as every cookbook (design → assign → record → infer); see vignette("cookbook-continuous") for the narrated version of the pattern. Here the response is 0/1, the natural estimand is a log odds ratio (or a risk difference / risk ratio via the g-computation classes), and the inference classes are the InferenceIncid* family.

Setup

EDI is not on CRAN yet, so install.packages("EDI") fails — install from R-universe (fallback: GitHub, subdir = "R/EDI"). Not evaluated here.

install.packages("EDI", repos = c("https://kapelner.r-universe.dev", "https://cloud.r-project.org"))
# or: remotes::install_github("kapelner/EDI", subdir = "R/EDI")
library(EDI)
set.seed(20260916)

n = 80
X = data.frame(
  age    = round(rnorm(n, 50, 10)),
  smoker = rbinom(n, 1, 0.3)
)
true_log_or = 0.9

Fixed design, logistic regression

des = DesignFixedBernoulli$new(n = n, response_type = "incidence", verbose = FALSE)
des$add_all_subjects_to_experiment(X)
des$assign_w_to_all_subjects()
w = des$get_w()

p = plogis(-1.2 + true_log_or * w + 0.03 * (X$age - 50) + 0.6 * X$smoker)
y = rbinom(n, 1, p)
des$add_all_subject_responses(y)

inf = InferenceIncidLogRegr$new(des, verbose = FALSE)
inf$num_cores = 1L
inf$compute_estimate()                         # log odds ratio for treatment
#> [1] 1.499761
inf$compute_asymp_confidence_interval(alpha = 0.05)
#>      2.5%     97.5% 
#> 0.4997003 2.4998222
inf$compute_asymp_two_sided_pval()
#> [1] 0.003289556

Randomization test and bootstrap, as in the continuous cookbook:

inf$set_seed(1)
inf$compute_rand_two_sided_pval(r = 200, show_progress = FALSE)
#> [1] 0.00264537
inf$set_seed(1)
inf$compute_bootstrap_confidence_interval(alpha = 0.05, B = 200, show_progress = FALSE)
#>  2.5% 97.5% 
#>    NA    NA

A risk difference instead of an odds ratio

The g-computation classes estimate a marginal risk difference or risk ratio by standardizing over the covariates — often the estimand a trial actually reports. Same design object, different class:

inf_rd = InferenceIncidGCompRiskDiff$new(des, verbose = FALSE)
inf_rd$num_cores = 1L
inf_rd$compute_estimate()
#> [1] 0.3322325
inf_rd$compute_asymp_confidence_interval(alpha = 0.05)
#>      2.5%     97.5% 
#> 0.1300459 0.5344191

Everything at once

suite = InferenceSuite$new(des)
res = suite$run_all_inference(screen = TRUE, plots = FALSE, num_cores = 1L,
                              methods = c("wald", "score", "lik_ratio"), max_secs_per_class = 15)
#> inference       cov      estimand    est       se        pval       pval method     status 
#> class           mod                                                                        
#> ===========================================================================================
#> Classes 0/25  [                0%                 ] Status: Estimating...Avg Δ                    mean Δ      0.331     0.104     2.19e-03   wald            ok     
#> Classes 1/25  [=              4%                ] Estimated Time Left: 2sAvg Δ Pooled …           mean Δ      0.331     0.103     1.81e-03   wald            ok     
#> Classes 2/25  [==             8%                ] Estimated Time Left: 1sBinom Ident R…  ~.       mean Δ      0.343     0.104     8.85e-04   wald            ok     
#> Classes 3/25  [===            12%               ] Estimated Time Left: 1sBinom Ident R…  ~.       mean Δ      0.343     0.104     1.95e-03   score           ok     
#> Classes 4/25  [=====          16%               ] Estimated Time Left: 1sBinom Ident R…  ~.       mean Δ      0.343     0.104     1.95e-03   lik_ratio       ok     
#> Classes 5/25  [======         20%               ] Estimated Time Left: 1sCMH                      mean Δ      0.331     0.135     1.38e-02   wald            ok     
#> Classes 6/25  [=======        24%               ] Estimated Time Left: 1sExact Zhang              logodds c…  1.45      NA        NA         NA              ok     
#> Classes 7/25  [=========      28%               ] Estimated Time Left: 1sG Comp Risk Δ   ~.       mean Δ      0.332     0.103     1.28e-03   wald            ok     
#> Classes 8/25  [==========     32%               ] Estimated Time Left: 1sG Comp Risk R…  ~.       risk ratio  2.59      0.850     3.72e-03   wald            ok     
#> Classes 9/25  [===========    36%               ] Estimated Time Left: 1sLog Binom       ~.       log risk …  0.941     0.333     5.36e-03   wald            ok     
#> Classes 10/25 [=============  40%               ] Estimated Time Left: 1sLog Binom       ~.       log risk …  0.941     0.333     1.32e-03   score           ok     
#> Classes 11/25 [============== 44%               ] Estimated Time Left: 1sLog Binom       ~.       log risk …  0.941     0.333     2.13e-03   lik_ratio       ok     
#> Classes 12/25 [============== 48%               ] Estimated Time Left: 0sLogist Regr     ~.       logodds m…  1.50      0.510     3.29e-03   wald            ok     
#> Classes 13/25 [============== 52%               ] Estimated Time Left: 0sLogist Regr     ~.       logodds m…  1.50      0.510     2.45e-03   score           ok     
#> Classes 14/25 [============== 56%               ] Estimated Time Left: 0sLogist Regr     ~.       logodds m…  1.50      0.510     2.25e-03   lik_ratio       ok     
#> Classes 15/25 [============== 60%               ] Estimated Time Left: 0sMiettinen Ris…           mean Δ      0.331     0.103     2.26e-03   wald            ok     
#> Classes 16/25 [============== 64% ==            ] Estimated Time Left: 0sModified Pois…  ~.       log risk …  0.951     0.409     1.99e-02   wald            ok     
#> Classes 17/25 [============== 68% ===           ] Estimated Time Left: 0sModified Pois…  ~.       log risk …  0.951     0.409     1.60e-02   score           ok     
#> Classes 18/25 [============== 72% ====          ] Estimated Time Left: 0sModified Pois…  ~.       log risk …  0.951     0.409     1.51e-02   lik_ratio       ok     
#> Classes 19/25 [============== 76% ======        ] Estimated Time Left: 0sNewcombe Risk…           mean Δ      0.331     NA        1.99e-03   wald            ok     
#> Classes 20/25 [============== 80% =======       ] Estimated Time Left: 0sProbit Regr     ~.       probit ma…  0.922     0.305     2.52e-03   wald            ok     
#> Classes 21/25 [============== 84% ========      ] Estimated Time Left: 0sProbit Regr     ~.       probit ma…  0.922     0.305     2.56e-03   score           ok     
#> Classes 22/25 [============== 88% ==========    ] Estimated Time Left: 0sProbit Regr     ~.       probit ma…  0.922     0.305     2.21e-03   lik_ratio       ok     
#> Classes 23/25 [============== 92% ===========   ] Estimated Time Left: 0sRisk Δ          ~.       mean Δ      0.332     0.103     1.24e-03   wald            ok     
#> Classes 24/25 [============== 96% ============  ] Estimated Time Left: 0sWald                     mean Δ      0.331     0.103     1.27e-03   wald            ok     
#> Classes 25/25 [============= 100% ==============] Estimated Time Left: 0s-------------------------------------------------------------------------------------------
#> Status: Completed in 1s.
#> 
#>   Estimand: risk ratio (1 inferences)      : p =      NA
#>   Estimand: logodds marginal (3 inferences): p = 0.00259
#>   Estimand: log risk ratio (6 inferences)  : p = 0.00377
#>   Estimand: mean Δ (11 inferences)         : p = 0.00168
#>   Estimand: probit marginal (3 inferences) : p = 0.00242
#> 
#> Combined evidence against the sharp null across 5 estimands
#> (24 inferences, weighting = uniform within estimand):
#> p = 0.00259

Sequential design: matching on the fly

DesignSeqOneByOneKK14 matches each arrival to an earlier unmatched subject when a close enough match exists. Its matched inference class for a binary outcome, InferenceIncidKKGCompRiskDiff, uses the pair/reservoir structure directly. As on every sequential design whose assignments depend on earlier subjects, the nonparametric bootstrap is not offered; randomization inference replays the design’s own mechanism instead.

des_seq = DesignSeqOneByOneKK14$new(n = n, response_type = "incidence", verbose = FALSE)
for (i in seq_len(n)) {
  w_i = des_seq$add_one_subject_to_experiment_and_assign(X[i, , drop = FALSE])
  p_i = plogis(-1.2 + true_log_or * w_i + 0.03 * (X$age[i] - 50) + 0.6 * X$smoker[i])
  des_seq$add_one_subject_response(i, rbinom(1, 1, p_i))
}

inf_seq = InferenceIncidKKGCompRiskDiff$new(des_seq, verbose = FALSE)
inf_seq$num_cores = 1L
inf_seq$compute_estimate()
#> [1] 0.2046291
inf_seq$compute_asymp_confidence_interval(alpha = 0.05)
#>       2.5%      97.5% 
#> 0.01616501 0.39309318
inf_seq$set_seed(1)
inf_seq$compute_rand_two_sided_pval(r = 200, show_progress = FALSE)
#> [1] 0.1553835

Where to go next

mirror server hosted at Truenetwork, Russian Federation.