Package {EDI}


Type: Package
Title: Experimental Design and Inference
Version: 1.0.2
Description: Implements a comprehensive suite of experimental designs, both fixed (e.g., block, stratified, matched-pair, cluster, factorial, and mixed-integer-programming-based optimal designs) and sequential (including matching-on-the-fly designs, biased coin designs, and covariate-adaptive urn designs that assign treatment one subject at a time while maintaining covariate balance), for continuous, incidence, count, proportion, survival, and ordinal response types. For each design and response type combination, provides the corresponding inference procedures, including exact, asymptotic, distribution-free, and resampling-based (bootstrap, jackknife, and randomization) methods, so that estimation and testing are always matched to how the data were generated. An 'InferenceSuite' facility runs all applicable inference procedures for a given design and response type at once and reports a single Cauchy-combined p-value summarizing their evidence. A built-in simulation framework supports power analysis and operating-characteristic studies across designs, response types, and inference procedures, with optional parallelization via 'mirai'. Missing covariate data is handled automatically via built-in imputation. Core numerical routines are implemented in C++ via 'Rcpp' for speed on large designs and simulation studies. Machine-specific tuning for optimization is included.
License: GPL-3
Encoding: UTF-8
Language: en-US
Depends: R (≥ 3.5.0)
LinkingTo: Rcpp, RcppEigen, RcppNumerical
Imports: R6, checkmate, Rcpp, data.table, survival, missRanger, missForest, numDeriv, digest, methods, MASS, RhpcBLASctl, randomizr
Suggests: anticlust, blockTools, clinfun, coin, dplyr, mirai, aftgee, betareg, copula, doParallel, ggplot2, fixest, gamlss.dist, geepack, glmmTMB, glmnet, icenReg, interval, jsonlite, knitr, lme4, multgee, nbpMatching, nlme, nnet, ordinal, pscl, qpdf, quantreg, R.utils, RcppXPtrUtils, Rfit, rmarkdown, testthat, VGAM, ompr, ompr.roi, ROI.plugin.glpk, sandwich, withr
VignetteBuilder: knitr
Collate: 'design_abstract.R' 'design_blocking_abstract.R' 'design_matching_abstract.R' 'design_fixed_abstract.R' 'design_component_registry.R' 'design_class_factory.R' 'design_fixed_greedy.R' 'helper_optimal_shared.R' 'helper_optimal_milp_solvers.R' 'helper_user_compiled_fn.R' 'helper_optimal_annealing.R' 'design_fixed_greedy_d_optimal.R' 'design_fixed_bernoulli.R' 'design_fixed_binary_match.R' 'design_fixed_blocked_cluster.R' 'design_fixed_blocking.R' 'design_fixed_cluster.R' 'design_fixed_factorial.R' 'design_fixed_ibcrd.R' 'design_fixed_matching_greedy_pair_switching.R' 'design_fixed_optimal.R' 'design_fixed_optimal_blocks.R' 'design_fixed_rerandomization.R' 'design_observational.R' 'design_observational_blocks.R' 'design_observational_matching.R' 'design_seq_one_by_one_abstract.R' 'design_custom_extensions.R' 'design_seq_one_by_one_atkinson.R' 'design_seq_one_by_one_bernoulli.R' 'design_seq_one_by_one_efron.R' 'design_seq_one_by_one_ibcrd.R' 'design_seq_one_by_one_KK14.R' 'design_seq_one_by_one_KK21_stepwise.R' 'design_seq_one_by_one_KK21.R' 'design_seq_one_by_one_pocock_simon.R' 'design_seq_one_by_one_random_block_size.R' 'design_seq_one_by_one_spbr.R' 'design_seq_one_by_one_urn.R' 'design_class_registry.R' 'EDI.R' 'helper_glm_fit.R' 'globals.R' 'comprehensive_slow_paths.R' 'helper_additional_asserts.R' 'contracts_mixins.R' 'helper_robust_sandwich.R' 'inference_all_abstract_mle_or_KM_summary_table.R' 'inference_ext_ci_inversion.R' 'inference_ext_information_matrix.R' 'inference_ext_likelihood_test_memoization.R' 'inference_mixin_off_optimum_likelihood_eval.R' 'inference_all_abstract_asymp_lik.R' 'inference_all_abstract_marginal_estimand.R' 'inference_ext_bartlett_approx.R' 'inference_mixin_cordeiro_ferrari_approx.R' 'inference_mixin_lemonte_gradient_approx.R' 'inference_mixin_kk_gee_shared.R' 'inference_mixin_kk_glmm_shared.R' 'inference_mixin_kk_passthrough.R' 'inference_mixin_kk_passthrough_compound.R' 'helper_bootstrap_ci.R' 'contracts_resampling_draws.R' 'inference_ext_bca_bootstrap_ci.R' 'inference_ext_exchangeable_resampling_units.R' 'inference_ext_minimum_volatility_selector.R' 'inference_ext_m_out_of_n_bootstrap.R' 'inference_ext_prw_subsampling.R' 'inference_all_abstract_non_param_boot.R' 'inference_all_abstract_rand_bootstrap.R' 'inference_all_abstract_rand_bootstrap_ci.R' 'inference_all_abstract_bayesian_bootstrap.R' 'inference_all_abstract_jackknife.R' 'inference_all_abstract_exact.R' 'inference_all_abstract_asymp.R' 'inference_ext_param_bootstrap_estimate.R' 'inference_all_abstract_param_boot.R' 'inference_all_abstract_asymp_lik_std_mod_cache.R' 'inference_all_abstract_count_likelihood.R' 'inference_ext_quantile_rand_ci.R' 'inference_all_abstract_rand_ci.R' 'inference_ext_sequential_mc_pval.R' 'inference_all_abstract_rand.R' 'inference_all_abstract.R' 'inference_all_abstract_KK_passthrough_compound.R' 'inference_all_abstract_quantile_rand_ci.R' 'inference_all_KK_mean_diff_IVWC.R' 'inference_all_KK_quantile_regr_ivwc_abstract.R' 'inference_all_KK_quantile_regr_one_lik_abstract.R' 'inference_all_KK_wilcox_ivwc.R' 'inference_all_average_diff.R' 'inference_all_simple_mean_diff_pooled_var.R' 'inference_all_simple_wilcox.R' 'inference_continuous_KK_bai_abstract.R' 'inference_continuous_KK_glmm.R' 'inference_continuous_KK_ols_ivwc.R' 'inference_continuous_KK_ols_one_lik.R' 'inference_continuous_KK_quantile_regr_ivwc.R' 'inference_continuous_KK_quantile_regr_one_lik.R' 'inference_continuous_KK_robust_regr_ivwc.R' 'inference_continuous_KK_robust_regr_one_lik.R' 'inference_continuous_KK14_bai.R' 'inference_continuous_KK21_bai.R' 'inference_continuous_lin.R' 'inference_continuous_ols.R' 'inference_continuous_quantile_regr.R' 'inference_continuous_robust_regr.R' 'inference_count_composite_likelihood.R' 'inference_count_KK_combined.R' 'inference_count_KK_cond_poisson.R' 'inference_count_KK_gee.R' 'inference_count_negbin.R' 'inference_count_poisson.R' 'inference_count_quasipoisson.R' 'inference_count_robust_poisson.R' 'inference_count_zero_augmented_poisson_abstract.R' 'inference_count_hurdle.R' 'inference_count_zero_inflated.R' 'inference_custom_extensions.R' 'inference_rand_custom.R' 'inference_helpers_zhang.R' 'inference_incidence_binomial_identity.R' 'inference_incidence_cmh.R' 'inference_incidence_exact_binomial.R' 'inference_incidence_exact_zhang.R' 'inference_incidence_extended_robins.R' 'helper_gcomp.R' 'helper_marginal_estimand.R' 'inference_incidence_gcomp_abstract.R' 'inference_incidence_gcomp.R' 'inference_incidence_KK_cond_logit_glmm_abstract.R' 'inference_incidence_KK_cond_logit_glmm.R' 'inference_incidence_KK_cond_logit.R' 'inference_incidence_KK_combined.R' 'inference_incidence_KK_gcomp_abstract.R' 'inference_incidence_KK_marginal_abstract.R' 'inference_incidence_KK_marginal.R' 'inference_incidence_KK_newcombe_ivwc_univ.R' 'inference_incidence_log_binomial.R' 'inference_incidence_logit.R' 'inference_incidence_miettinen_nurminen_univ.R' 'inference_incidence_modified_poisson.R' 'inference_incidence_newcombe_univ.R' 'inference_incidence_probit.R' 'inference_incidence_risk_diff.R' 'inference_incid_wald.R' 'inference_indicidence_exact_fisher.R' 'inference_ordinal_adj_cat_logit.R' 'inference_ordinal_cauchit.R' 'inference_ordinal_cloglog.R' 'inference_ordinal_gcomp.R' 'inference_ordinal_jonckheere_terpstra_test.R' 'inference_ordinal_KK_clmm_abstract.R' 'inference_ordinal_KK_combined.R' 'inference_ordinal_KK_cond_adj_cat_logit.R' 'inference_ordinal_KK_cond_logit_abstract.R' 'inference_ordinal_ordered_probit.R' 'inference_ordinal_paired_sign_test.R' 'inference_ordinal_partial_proportional_odds.R' 'inference_ordinal_proportional_odds.R' 'inference_ordinal_ridit.R' 'inference_ordinal_stereotype_logit.R' 'inference_proportion_beta.R' 'inference_proportion_fractional_logit.R' 'inference_proportion_gcomp.R' 'inference_proportion_KK_combined.R' 'inference_proportion_KK_quantile_regr_ivwc.R' 'inference_proportion_KK_quantile_regr_one_lik.R' 'inference_proportion_quantile_regr.R' 'inference_proportion_zero_one_inflated_beta.R' 'inference_suite.R' 'inference_survival_coxph.R' 'inference_survival_dep_cens_transform.R' 'inference_survival_gehan_wilcox.R' 'inference_survival_GLMM_weibull_frailty_loggamma.R' 'inference_survival_KK_lwa_cox_ivwc_abstract.R' 'inference_survival_KK_lwa_cox_one_lik_abstract.R' 'inference_survival_KK_lwa_cox.R' 'inference_survival_KK_rank_regr_ivwc_abstract.R' 'inference_survival_KK_rank_regr.R' 'inference_survival_KK_strat_cox.R' 'inference_survival_GLMM_weibull_frailty_normal.R' 'inference_survival_KK_weibull_marginal.R' 'inference_survival_km_diff.R' 'inference_survival_log_rank.R' 'inference_survival_rmst.R' 'inference_survival_strat_cox.R' 'inference_survival_weibull.R' 'inference_class_registry.R' 'helper_package_checks.R' 'helper_math.R' 'helper_model_matrix.R' 'helper_response_asserts.R' 'helper_robust_regression.R' 'helper_rcpp_doc_stubs.R' 'helper_matching.R' 'helper_survival_fits.R' 'helper_zoib.R' 'helper_inference_survival_turnbull.R' 'RcppExports.R' 'simulations_framework.R' 'simulation_framework_report.R' 'local_machine_tuning_synthetic_fixtures.R' 'local_machine_tuning_harness.R' 'local_machine_tuning_axes.R' 'local_machine_tuning_persistence.R' 'local_machine_tuning_correctness.R' 'local_machine_tuning.R' 'zzz.R'
URL: https://github.com/kapelner/EDI, https://kapelner.github.io/EDI/, https://pypi.org/project/edi-kernels/
BugReports: https://github.com/kapelner/EDI/issues
Config/roxygen2/version: 8.0.0.9000
NeedsCompilation: yes
Packaged: 2026-09-25 14:50:44 UTC; kapelner
Author: Adam Kapelner ORCID iD [aut, cre, cph]
Maintainer: Adam Kapelner <adam.kapelner@mail.huji.ac.il>
Repository: CRAN
Date/Publication: 2026-10-07 09:00:15 UTC

Experimental Design and Inference

Description

EDI

Details

Provides comprehensive support for many fixed and sequential experimental designs and many infererential methods (parametric, nonparametric, exact) for response types continuous, incidence, count, proportion, survival (with censoring) and ordinal. Supports automatic missing data imputation, parallelization, provides robustness fallbacks and is optimized with C++.

Author(s)

Adam Kapelner kapelner@qc.cuny.edu

References

Adam Kapelner and Abba Krieger A Matching Procedure for Sequential Experiments that Iteratively Learns which Covariates Improve Power, Arxiv 2010.05980

See Also

Useful links:


Abstract Quantile Regression Compound Estimator for KK Matching-on-the-Fly Designs

Description

An abstract base class providing shared quantile regression logic for KK matching-on-the-fly designs. Subclasses override the transform_y_fn private$m field to apply a response transformation before quantile regression (e.g., identity for continuous, qlogis for proportion outcomes).

Usage

.init_kk_quantile_regr_ivwc(
  self,
  private,
  super,
  des_obj,
  model_formula,
  tau,
  transform_y_fn,
  verbose,
  smart_cold_start_default
)

Abstract Quantile Regression Combined-Likelihood Compound Estimator for KK Designs

Description

Fits a single joint quantile regression over all KK design data by stacking matched-pair differences and reservoir observations into one design matrix.

Usage

.init_kk_quantile_regr_one_lik(
  self,
  private,
  super,
  des_obj,
  model_formula,
  tau,
  transform_y_fn,
  verbose
)

Details

Column layout of X_stack: [beta_0 | beta_T | beta_xs (p cols)] Pair rows: [0 | 1 | Xd_k] -> Q_tau(yd_k) = beta_T + Xd_k' beta_xs Reservoir rows: [1 | w_i | X_i] -> Q_tau(y_i) = beta_0 + w_i*beta_T + X_i'*beta_xs Fitting a single rq() on the stacked dataset minimises the combined check-function loss.

Special cases: Pairs only: beta_0 column is all-zero and dropped; layout [beta_T | beta_xs]. Reservoir only: standard quantile regression layout [beta_0 | beta_T | beta_xs].

Standard errors use Powell's "nid" sandwich estimator, falling back to "iid".


Normalize and Validate an Optimizer Algorithm Name for the fast_* C++ Backends

Description

Internal helper shared by the package's fast_* GLM/survival/ordinal fitting wrappers (e.g. fast_logistic_regression, fast_coxph_regression) to resolve a user-supplied optimization_alg argument to one of the fixed set of optimizer names the underlying C++ backends actually implement, applying a model-specific default when none is supplied and rejecting anything else. This centralizes the default/validation logic so each fast_* wrapper does not have to repeat it.

Usage

.normalize_optimizer_algorithm(
  optimization_alg,
  allow_irls = FALSE,
  default = if (allow_irls) "irls" else "lbfgs"
)

Arguments

optimization_alg

Character string (possibly abbreviated) naming the desired optimizer, NULL, or missing entirely; see Details for resolution order.

allow_irls

Logical. Whether "irls" is a valid choice (and the default default) for this model; FALSE restricts the allowed set to c("lbfgs", "newton_raphson").

default

Character string used when optimization_alg is missing or NULL. Defaults to "irls" when allow_irls = TRUE, else "lbfgs".

Details

The three possible optimizer names, when supported by a given model, correspond to distinct fitting algorithms in the C++ backends: "newton_raphson" (full Newton-Raphson using the analytic Hessian), "lbfgs" (limited-memory quasi-Newton, avoiding an explicit Hessian), and "irls" (iteratively reweighted least squares, the classical GLM-fitting algorithm — only meaningful, and only offered, for exponential-family GLMs, hence gated by allow_irls). Which optimizers a given fast_* function actually accepts (and which is its default) varies by model; this function only encodes the generic irls-vs-not-irls split, not per-model specifics.

optimization_alg is matched against the allowed set via match.arg, so unambiguous partial string matches (e.g. "newton") are accepted; an unmatched or ambiguous value raises match.arg's standard error rather than silently falling back to the default. missing(optimization_alg) or an explicit NULL both resolve to default before matching.

Value

A validated, unabbreviated character string: one of "newton_raphson", "lbfgs", or (only when allow_irls = TRUE) "irls".

See Also

match.arg, which performs the validation/partial-matching.


Inference based on Maximum Likelihood for KK designs

Description

Initialize Bai adjusted-t inference for a completed KK continuous-response design, including the optional convex combination of matched-pair and reservoir estimates.

Computes the appropriate estimate for compound mean difference across pairs and reservoir

Computes a 1-alpha level frequentist confidence interval

Here we use the theory that MLE's computed for GLM's are asymptotically normal (except in the case of estimat_type "median difference" where a nonparametric bootstrap confidence interval (see the controlTest::quantileControlTest method) is employed. Hence these confidence intervals are asymptotically valid and thus approximate for any sample size.

Compute the Bai-adjusted two-sided p-value for the treatment effect using the matched-design adjusted statistic. See related InferenceBaiAdjustedTKK14 methods.

Usage

BaiAdjustedTSource

Details

Inference for mean difference. Note that warm starts are disabled for this class as the Bai adjusted t-test is a closed-form estimator and does not benefit from initialization.

This class requires the nbpMatching package, which is listed in Suggests and is not installed automatically with EDI. Install it manually with install.packages("nbpMatching") before using this class.

Value

The setting-appropriate (see description) numeric estimate of the treatment effect

A (1 - alpha)-sized frequentist confidence interval for the treatment effect

The approximate frequentist p-value

Examples


# (loading the nbpMatching package alone takes a few seconds)
if (requireNamespace("nbpMatching", quietly = TRUE)) {
  seq_des = DesignSeqOneByOneKK14$new(n = 20, response_type = "continuous")
  for (i in 1:20) {
    seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
  }
  seq_des$add_all_subject_responses(rnorm(20))
  seq_des_inf = InferenceBaiAdjustedTKK14$new(seq_des)
  seq_des_inf$compute_estimate()
}



# (loading the nbpMatching package alone takes a few seconds)
if (requireNamespace("nbpMatching", quietly = TRUE)) {
  seq_des = DesignSeqOneByOneKK14$new(n = 20, response_type = "continuous")
  for (i in 1:20) {
    seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
  }
  seq_des$add_all_subject_responses(rnorm(20))
  seq_des_inf = InferenceBaiAdjustedTKK14$new(seq_des)
  seq_des_inf$compute_asymp_confidence_interval()
}



# (loading the nbpMatching package alone takes a few seconds)
if (requireNamespace("nbpMatching", quietly = TRUE)) {
  seq_des = DesignSeqOneByOneKK14$new(n = 20, response_type = "continuous")
  for (i in 1:20) {
    seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
  }
  seq_des$add_all_subject_responses(rnorm(20))
  seq_des_inf = InferenceBaiAdjustedTKK14$new(seq_des)
  seq_des_inf$compute_asymp_two_sided_pval()
}



Robust-Regression IVWC Compound Inference for KK Designs

Description

Initialize KK inverse-variance combined robust-regression inference and prepare the matched/reservoir components used by InferenceContinKKRobustRegrIVWC.

Point estimate of the treatment effect combining a matched-pair robust regression (MASS::rlm on within-pair differences) and a reservoir robust regression (treatment vs. control), combined by inverse-variance weighting: the two component estimates \hat\beta_{T,\text{matched}} and \hat\beta_{T,\text{reservoir}} (each with its own estimated variance) are combined as \hat\beta_T = \sum_k w_k \hat\beta_{T,k} / \sum_k w_k with w_k = 1/\widehat{\mathrm{Var}}(\hat\beta_{T,k}); a component is dropped from the combination if it is not estimable (e.g. too few observations). See Inference for the general estimate contract.

Wald confidence interval for the inverse-variance-combined treatment effect: \hat\beta_T \pm t_{1-\alpha/2,\,df}\cdot \hat{se}(\hat\beta_T), using the combined variance 1/\sum_k w_k from compute_estimate()'s inverse-variance weighting. likelihood_tier = "quasi" for this class (M/MM-estimator objective, not a normalized likelihood), so only the Wald testing type is supported. See InferenceAsymp for the shared contract.

Two-sided Wald p-value for H_0: \beta_T = \code{delta} vs. H_1: \beta_T \neq \code{delta}, using the inverse-variance-combined estimate and standard error from compute_estimate(). See InferenceAsymp for the shared Wald/asymptotic semantics.

Duplicate the robust-regression inference object while preserving the selected match-specific formulas and clearing cached fit results; see Inference for the common duplication contract.

Usage

ContinKKRobustRegrIVWCSource

Details

Fits a variance-weighted compound estimator for KK matching-on-the-fly designs with continuous responses using robust linear regression ('MASS::rlm') for the matched-pair and reservoir components separately.

Model. Two robust M/MM-estimator regressions are fit independently: one on the within-pair outcome differences for matched pairs (with covariate differences as predictors, no intercept), and one on the reservoir subjects (raw covariates plus a treatment indicator). Each yields a treatment-effect estimate \hat\beta_{T,k} and estimated variance. The two are combined by inverse-variance weighting into a single \hat\beta_T, the same combination rule used by InferenceContinKKOLSIVWC but with robust rather than OLS component fits. A component missing enough data (e.g. no matched pairs) is dropped from the combination.

Robust fitting. Uses MASS::rlm (or an internal Rcpp IRLS kernel when use_rcpp = TRUE, the default) with method = "M" or "MM"; "MM" (the default) uses an LQS-based high-breakdown start, while "M" can optionally warm-start from OLS (start_with_ols = TRUE). likelihood_tier = "quasi": the M/MM objective is not a normalized likelihood, so only Wald-type asymptotic inference is available (no score/gradient/likelihood-ratio testing types).

Assumptions. Continuous response; independent matched pairs and/or independent reservoir subjects; no censoring; a KK matching-on-the-fly design. Robust regression down-weights outlying residuals, trading some efficiency under exactly-Gaussian errors for resistance to heavy tails and contamination.

Value

Numeric scalar treatment-effect estimate on the outcome's natural scale.

A length-2 numeric vector c(lower, upper), or NA bounds if nonestimable.

Numeric scalar p-value in [0, 1], or NA_real_ if nonestimable.

References

Kapelner, A. and Krieger, A. M. (2014). Matching on-the-fly: Sequential allocation with higher power and efficiency. Biometrics, 70(2), 378-388. doi:10.1111/biom.12148. (KK14 in REFERENCES.md.)

See Also

Analogous Python API for robust linear models: statsmodels RLM. Robust regression (orientation).

Legacy class. Not fully tested in comprehensive_tests.R.


Count Composite Likelihood Inference Base

Description

Computes the treatment estimate.

Usage

CountCompositeLikelihoodSource

Details

Shared branch for count models whose reported estimator is robust or quasi-likelihood based.


Conditional-Poisson Inference for KK Designs with Combined Likelihood

Description

Initialize conditional-Poisson one-likelihood inference for KK count designs and prepare the combined likelihood used by InferenceCountKKCondPoissonOneLik.

Compute the conditional-Poisson one-likelihood treatment estimate by fitting the combined matched/reservoir likelihood and caching the treatment log-rate coefficient for related likelihood-test methods.

Recomputes the combined conditional-Poisson estimate under Bayesian-bootstrap weights.

Uses the shared asymptotic confidence-interval contract; see InferenceAsymp.

Uses the shared asymptotic two-sided p-value contract; see InferenceAsymp.

Uses the shared Wald confidence-interval contract; see InferenceAsymp.

Computes a Wald two-sided p-value.

Computes a design-adjusted score confidence interval.

Computes a design-adjusted likelihood-ratio confidence interval.

Computes a design-adjusted gradient confidence interval.

Computes a design-adjusted score p-value.

Computes a design-adjusted likelihood-ratio p-value.

Computes a design-adjusted gradient p-value.

Usage

CountKKCondPoissonOneLikLikelihoodSource

KK Hurdle Poisson Combined-Likelihood Inference for Count Responses

Description

Initialize KK hurdle-Poisson one-likelihood inference for count responses and prepare the combined matched/reservoir likelihood. See InferenceCountKKHurdlePoissonOneLik and InferenceParamBootstrap for related likelihood and bootstrap methods.

Compute the one-likelihood hurdle-Poisson treatment-effect estimate by fitting the combined count likelihood and caching the treatment log-rate coefficient for related p-value and interval methods.

Recomputes the combined hurdle-Poisson estimate under Bayesian-bootstrap weights.

Compute the configured asymptotic confidence interval for the one-likelihood hurdle-Poisson treatment coefficient, delegating to Wald, score, likelihood-ratio, or gradient paths as documented in InferenceCountLikelihood.

Computes a design-conservative score confidence interval.

Computes a design-conservative likelihood-ratio confidence interval.

Computes a design-conservative gradient confidence interval.

Compute the configured asymptotic two-sided p-value for the one-likelihood hurdle-Poisson treatment coefficient, delegating to Wald, score, likelihood-ratio, or gradient paths as documented in InferenceCountLikelihood.

Computes a design-conservative score p-value.

Computes a design-conservative likelihood-ratio p-value.

Computes a design-conservative gradient p-value.

Compute the Wald confidence interval for the one-likelihood hurdle-Poisson treatment coefficient, falling back to bootstrap when the model standard error is unavailable. See InferenceAsymp.

Compute the Wald two-sided p-value for the one-likelihood hurdle-Poisson treatment coefficient, falling back to the Bayesian-bootstrap p-value when the model standard error is unavailable. See InferenceAsymp.

Usage

CountKKHurdlePoissonOneLikLikelihoodSource

An Abstract Experimental Design

Description

Internal method. An abstract R6 Class encapsulating the data and functionality for an experimental design. This class takes care of data storage and response handling.

Details

Throughout the package, treatment assignment vectors w use the \{0, 1\} encoding: 1 indicates a treated subject and 0 a control subject. All public methods that return or accept w (e.g. get_w(), draw_ws_according_to_design()) use this convention. A handful of variance estimators (e.g. InferenceIncidCMH, InferenceIncidExtendedRobins) recode to a signed \{-1,+1\} contrast internally where their formulas require it; that recoding is local to those classes and does not affect this public convention.

Saving and loading

Design (and its DesignSeqOneByOne subclasses) is the unit of persistence for a trial. Persist a des_obj with base R's saveRDS()/readRDS() – there is no dedicated save_edi_design()/load_edi_design() wrapper, and none is planned: the audit behind this section found nothing that needs transformation on load beyond what is documented here. Inference* objects are disposable, cheaply reconstructed from a Design object on demand (see each class's $new()), and must never be saveRDS()'d directly – nothing currently prevents it (they serialize "successfully" like any R6 object), but the result is a frozen snapshot a user could easily mistake for something that stays live against the design, and re-running inference from a reloaded Design is both cheap and the only tested path.

Worked example (mirrors the round-trip tests in R/EDI/tests/testthat/test-save-load-design.R):

des_obj = DesignSeqOneByOneBernoulli$new(n = 20, response_type = "continuous")
for (i in 1:10) {
  des_obj$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
  des_obj$add_one_subject_response(i, y = rnorm(1))
}
saveRDS(des_obj, "trial.rds", version = 2)

# ...new R session...
des_obj = readRDS("trial.rds")
for (i in 11:20) {
  des_obj$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
  des_obj$add_one_subject_response(i, y = rnorm(1))
}
inf_obj = InferenceContinOLS$new(des_obj) # reconstructed fresh, never persisted
inf_obj$compute_estimate()

Passing version = 2 to saveRDS() is recommended, matching the one existing internal precedent for RDS serialization in this package (SimulationFramework's replication cache); it is not required for a same-R-version round trip.

Version stamp. Every Design object records the package version it was constructed under (get_edi_version_created()). This is stamped once at construction and is not refreshed by readRDS() – it reflects the version that originally built the object, not whatever version is currently loaded. The first "resume the trial" call after a reload (draw_ws_according_to_design() for fixed designs, add_one_subject_to_experiment_and_assign() for sequential designs) compares the stamped version's major component against the currently loaded package's major component and emits a one-time warning() on a mismatch; minor/patch differences are silent, since most field additions are additive under this class's lock_objects = FALSE R6 fields and do not warrant nagging on every routine upgrade. Objects saved before this field existed self-initialize it to the currently loaded version the first time it is read, rather than erroring on the missing field.

RNG/reproducibility caveat. private$seed is consumed only once, inside maybe_set_seed() at construction time, and is not re-applied on readRDS(). Continuing to enroll subjects after a reload therefore draws from whatever the global .Random.seed happens to be in the new session, not a deterministic continuation of the original stream. This is almost certainly the right behavior for a real trial (bit-for-bit-reproducible continuation across a process restart is not a property a production trial should have), but it means a same-seed reload-and-continue is not expected to reproduce the same draws as an uninterrupted run with that seed – do not rely on that for testing.

Known non-serializable case. A DesignFixedOptimal constructed with objective = "custom" from a raw RcppXPtrUtils::cppXPtr() external pointer (rather than a C++ source string) cannot be safely reloaded: compiled function pointers do not survive a saveRDS()/readRDS() round trip, and there is no retained source to recompile from. This is detected on first use after reload and raises a clear error rather than failing silently; supply custom_objective as a C++ source string instead of a pre-built cppXPtr() object if you need this design to survive a save/reload cycle – that form recompiles itself automatically the first time it is used post-reload. Every other audited private cache on Design and its components (all_subject_data_cache, permutations_cache, lin_centered_covariates, matching/blocking/cluster component state such as m, xm_structural, boot_pair_rows) was traced to its originating C++ return type and confirmed to hold only plain R matrices/vectors/lists, not external pointers or other non-serializable values.

Active bindings

num_cores

Current number of cores in the global budget.

Methods

Public methods


Design$is_blocking_design()

Check whether this design currently has blocking structure.

The base implementation returns FALSE. Designs that compose BlockingStructure override this method with the structural check.

Usage
Design$is_blocking_design()
Returns

FALSE for designs without BlockingStructure.


Design$is_matching_design()

Check whether this design currently has matching structure.

The base implementation returns FALSE. Designs that compose MatchingStructure override this method with the structural check.

Usage
Design$is_matching_design()
Returns

FALSE for designs without MatchingStructure.


Design$is_a_kk_matching_capable()

Characterization: is this a KK matching-on-the-fly-capable design (sequential KK or its fixed binary-match equivalent)? Default FALSE; overridden to TRUE on DesignSeqOneByOneKK14 and DesignFixedBinaryMatch.

Usage
Design$is_a_kk_matching_capable()

Design$is_a_cluster_capable()

Characterization: is this a cluster-structured design? Default FALSE; overridden to TRUE on DesignFixedCluster and DesignFixedBlockedCluster.

Usage
Design$is_a_cluster_capable()

Design$is_a_bernoulli_capable()

Characterization: is this a Bernoulli-randomized design? Default FALSE; overridden to TRUE on DesignSeqOneByOneBernoulli and DesignFixedBernoulli.

Usage
Design$is_a_bernoulli_capable()

Design$new()

Initialize an experimental design

Usage
Design$new(
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = FALSE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  ordinal_levels = NULL,
  seed = NULL
)
Arguments
response_type

"continuous", "incidence", "proportion", "count", "survival", or "ordinal".

prob_T

Probability of treatment assignment.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size (if fixed).

verbose

Flag for verbosity.

missingness_method

How to handle missing values in covariates when building the model matrix for inference. One of:

"impute" (default)

Missing values are filled in using random-forest imputation (missRanger, falling back to missForest on failure). The response vector is included as an auxiliary predictor when available. This preserves all covariates and all subjects but introduces imputed values that influence inference.

"drop_column"

Any covariate column that contains at least one missing value is dropped entirely from the model matrix before inference. No values are invented; the remaining complete columns are used as-is. This is conservative but transparent.

"error"

An error is thrown as soon as any missing value is detected in the covariate matrix. Use this when you want to guarantee that inference runs on exactly the data you supplied, with no silent modification.

design_formula

A formula object used to create the design matrix from covariates. Default is ~ ..

ordinal_levels

If the response type is "ordinal", the labels for the levels.

seed

Integer seed for reproducibility.

Returns

A new 'Design' object


Design$add_one_subject_response()

For CARA designs, add a single subject response.

Usage
Design$add_one_subject_response(t, y = NULL, y_L = NULL, y_R = NULL)
Arguments
t

The subject index.

y

The exact response value. Supply this XOR both y_L and y_R – never together, never just one of the two.

y_L

For a censored survival response, the lower bound of the event-time interval. Right-censored: the last known event-free time (pair with y_R = Inf). Left-censored: 0, which must be stated explicitly rather than defaulted. Interval-censored: the interval's lower bound. Storage accepts any well-formed left-/interval-censored value; whether a given Inference class can actually consume it depends on that class (most survival Inference classes still only accept exact/right-censored data and will reject construction with a clear error otherwise – see individual class docs).

y_R

For a censored survival response, the upper bound of the event-time interval. Right-censored: Inf. Left-/ interval-censored: the confirmed-by time / interval upper bound.


Design$add_all_subject_responses()

For non-CARA designs, add all subject responses.

Usage
Design$add_all_subject_responses(ys = NULL, y_Ls = NULL, y_Rs = NULL)
Arguments
ys

The exact responses as a numeric vector, NA for any subject whose response is censored (supply y_Ls/y_Rs for those instead).

y_Ls

The censored-response lower bounds, NA for any subject with an exact response in ys. Right-censored: the last known event-free time (pair with y_Rs = Inf). Left-censored: 0, stated explicitly. Interval-censored: the interval's lower bound. Storage accepts any well-formed left-/interval-censored value; whether a given Inference class can actually consume it depends on that class (most survival Inference classes still only accept exact/ right-censored data and will reject construction with a clear error otherwise – see individual class docs).

y_Rs

The censored-response upper bounds, NA for any subject with an exact response in ys. Right-censored: Inf. Left-/interval-censored: the confirmed-by time / interval upper bound.


Design$overwrite_all_subject_assignments()

For analysis on already-completed experimental data

Usage
Design$overwrite_all_subject_assignments(w)
Arguments
w

A {0,1} vector of subject assignments (1 = treated, 0 = control).


Design$is_fixed_sample_size()

Check if this design was initialized with a fixed sample size n

Usage
Design$is_fixed_sample_size()
Returns

TRUE if fixed.


Design$assert_all_subjects_arrived()

Asserts if all subjects arrived.

Usage
Design$assert_all_subjects_arrived()

Design$assert_all_responses_recorded()

Asserts if all responses are recorded.

Usage
Design$assert_all_responses_recorded()

Design$check_experiment_completed()

Checks if the experiment is completed.

Usage
Design$check_experiment_completed()
Returns

TRUE if experiment is complete, FALSE otherwise.


Design$assert_even_allocation()

Checks if the experiment has a 50-50 allocation.

Usage
Design$assert_even_allocation()

Design$assert_fixed_sample()

Checks if the experiment has a fixed sample size.

Usage
Design$assert_fixed_sample()

Design$any_censoring()

Checks if the experiment has any censored responses

Usage
Design$any_censoring()
Returns

TRUE if any censored.


Design$has_general_censoring()

Checks if the experiment has any left- or interval-censored survival responses – i.e. any subject whose y_R is finite (right-censored subjects have y_R = Inf, which is excluded). Most survival Inference classes cannot yet consume this shape of data (see get_effective_time()/get_effective_dead()); this is the check Inference$initialize() uses to reject construction cleanly for those classes.

Usage
Design$has_general_censoring()
Returns

TRUE if any subject is left- or interval-censored.


Design$get_t()

Get t

Usage
Design$get_t()
Returns

The current number of subjects.


Design$get_X_raw()

Get raw X information

Usage
Design$get_X_raw()
Returns

A data frame of subject data.


Design$get_X_imp()

Get imputed X information

Usage
Design$get_X_imp()
Returns

Same as Xraw except with imputations.


Design$get_X()

Get X matrix

Usage
Design$get_X()
Returns

A numeric matrix of subject data.


Design$get_y()

Get y

Usage
Design$get_y()
Returns

A numeric vector of subject responses.


Design$get_y_original()

Get y_original

Usage
Design$get_y_original()
Returns

A numeric vector of the original subject responses.


Design$get_w()

Get w

Usage
Design$get_w()
Returns

A {0,1} vector of subject assignments (1 = treated, 0 = control).


Design$draw_ws_according_to_design()

Draw treatment assignment vectors according to the design.

Usage
Design$draw_ws_according_to_design(r = 1L)
Arguments
r

Number of vectors to draw. Default is 1.

Returns

A matrix of size n x r with {0,1} entries (1 = treated, 0 = control).


Design$capabilities()

Returns the capabilities this design instance exposes (see fix_design_hierarchy.md, "Capability Model").

Deliberately instance-level, not a class-registry read (fix_design_hierarchy.md, TODO-28): is_blocking_design()/is_matching_design() depend on real construction-time state (e.g. private$m/private$blocking_capable), not just which components a class composes – DesignFixediBCRD constructed with an unknown n, for instance, composes BlockingStructure but is not blocking-capable for that particular instance. A class-registry-only answer (this function briefly unioned in get_effective_design_capabilities(), a purely class-level, component- composition-based check) would silently report "blocking" for every instance of such a class regardless of its actual construction state – confirmed as a real, reproducible false positive during this TODO's implementation, not a hypothetical. get_effective_design_capabilities()/ design_class_registry.R's direct_components still exist and are correct – they're the right tool for a generator-only query with no instance in hand (see design_class_generator_supports_batch_w_pregeneration()), just not for this instance-level method.

Usage
Design$capabilities()
Returns

A character vector of capability names.


Design$supports()

Returns whether this design object supports a capability. See capabilities().

Usage
Design$supports(capability)
Arguments
capability

A capability name, e.g. "blocking", "matching", or "batch_w_pregeneration".

Returns

TRUE if the capability is present, FALSE otherwise.


Design$applicable_inference_class_names()

Returns the sorted character vector of concrete, exported Inference class names legal for this design object under default constructor arguments, derived purely from this design's own normalized metadata (response type, KK-matching capability, blocking, and both censoring axes) filtered through the registry's compatibility predicates – the same normalization and predicate logic InferenceSuite uses for discovery (see normalize_inference_design_metadata() and is_inference_class_compatible_with_design_metadata() in inference_suite.R). No candidate class is constructed to determine applicability, so this has no side effects and cannot be influenced by a constructor failure or a missing optional package (see unavailable_inference_classes_due_to_missing_packages() for that case, reported separately). A class whose censoring tolerance depends on non-default constructor arguments (e.g. InferenceSurvivalCoxPHRegr only tolerates general censoring with testing_type = "wald") is listed here when its default configuration is compatible; a construction-time error for an incompatible non-default argument combination remains the documented behavior of that class's initialize().

Usage
Design$applicable_inference_class_names()
Returns

A sorted character vector of applicable Inference class names.


Design$unavailable_inference_classes_due_to_missing_packages()

Companion to applicable_inference_class_names(): returns the subset of otherwise design-compatible Inference classes that are excluded solely because a registered required_packages entry is not installed, as a named list (class name -> character vector of missing package names) – kept separate from plain design incompatibility so callers can tell "not applicable to this design" apart from "applicable, but an optional dependency isn't installed."

Usage
Design$unavailable_inference_classes_due_to_missing_packages()
Returns

A named list, class name -> missing package names; empty list if none.


Design$incompatible_inference_classes_due_to_design_structure()

Companion to applicable_inference_class_names(): returns the subset of otherwise design-compatible Inference classes that are excluded because they declared a design_compatibility_reason predicate (a design-*structure* requirement, e.g. even treatment allocation or equal block sizes, beyond what response type/KK/blocking/censoring metadata alone can express) and this design object fails it, as a named list (class name -> one-line reason string) – kept separate from plain design incompatibility and from a missing package for the same reason unavailable_inference_classes_due_to_missing_packages() is kept separate: so callers can tell exactly why a class is missing from applicable_inference_class_names() instead of only discovering it as a construction-time error.

Usage
Design$incompatible_inference_classes_due_to_design_structure()
Returns

A named list, class name -> reason string; empty list if none.


Design$randomization_family()

Returns this design object's registry-backed randomization family (see fix_design_hierarchy.md, "Class Metadata"), e.g. "kk14", "bernoulli", "rerandomization". Replaces class-identity (inherits()/is()) dispatch at call sites that need to distinguish design variants (see "Class-Identity Dispatch Replacement"). Returns NA_character_ if the class is not registered or is one of the unsplit/timing-root abstract bases.

Usage
Design$randomization_family()
Returns

A single character string (or NA_character_).


Design$supports_resampling()

Check if the design supports resampling at all – FALSE only for the abstract timing-family bases themselves (DesignFixed, DesignSeqOneByOne, and their custom-extension abstract bases) instantiated directly; TRUE for every concrete subclass, including ObservationalDesign. This is the general check for resampling methods that never need the design's own randomization mechanism – plain nonparametric bootstrap, Bayesian bootstrap, m-out-of-n bootstrap, PRW subsampling – which only resample already-observed units/rows and their fixed, observed assignment, so they remain valid and available even for a design with no randomization mechanism at all (see ObservationalDesign's class documentation: "resampling subjects with their observed, fixed assignment does not require a known randomization probability"). Contrast with supports_randomization_draw()/ supports_resampling_replay() below, which gate the narrower set of methods that actually do need to invoke the design's mechanism (a plain randomization test/CI, or a bootstrap randomization test that re-randomizes resampled data) and are therefore FALSE for ObservationalDesign specifically – see fix_design_hierarchy.md, "Observational Design Migration" for the live bug that split fixes.

Usage
Design$supports_resampling()
Returns

TRUE if supported.


Design$supports_randomization_draw()

Check if this design can draw a fresh treatment assignment from its own randomization mechanism – the eligibility condition for permutation-style randomization tests/CIs (compute_rand_two_sided_pval() and friends), which redraw w directly. FALSE for the abstract timing-family bases themselves (same as supports_resampling()) and, unlike supports_resampling(), also FALSE for ObservationalDesign (no draw mechanism at all – w is supplied by the user, so there is nothing to redraw); TRUE for every other concrete subclass. See supports_resampling()'s documentation for why this is a narrower, separate capability rather than reusing that one, and "Observational Design Migration" for the live bug this fixes (ObservationalDesign previously answered the old, unsplit supports_resampling() TRUE, silently passing the randomization-test eligibility assert before failing later and deeper, inside draw_ws_raw()'s throwing stub).

Usage
Design$supports_randomization_draw()
Returns

TRUE if a fresh randomization draw is supported.


Design$supports_resampling_replay()

Check if this design's mechanism can be faithfully replayed against resampled data – the eligibility condition specifically for the bootstrap randomization test (BRT), which resamples units and then re-randomizes each resample using the design's own mechanism (see inference_all_abstract_rand_bootstrap.R's repeated draw_ws_according_to_design() calls). Not the eligibility condition for plain nonparametric/Bayesian/m-out-of-n/PRW-subsampling bootstrap – those never redraw w at all (they resample already-observed units and their fixed, observed assignment) and are gated by the broader supports_resampling() instead, which stays TRUE for ObservationalDesign. FALSE for the same abstract timing-family bases as supports_randomization_draw() and for ObservationalDesign (no randomization mechanism to replay); TRUE for every other concrete subclass. See supports_randomization_draw()'s documentation for why this is a separate capability rather than the same flag reused.

Usage
Design$supports_resampling_replay()
Returns

TRUE if bootstrap-randomization-test-style replay is supported.


Design$prepare_for_resampling_replay()

Hook invoked by the bootstrap-randomization-test machinery on a design object whose assignment mechanism is about to be replayed against resampled data (once per replicate draw site, ahead of draw_ws_according_to_design(1L)). The base implementation is a no-op; designs whose replay is a full re-optimization (DesignFixedOptimal) override it to switch to their per-replicate solver profile (solver_args$brt_*). Idempotent.

Usage
Design$prepare_for_resampling_replay()
Returns

invisible(NULL).


Design$warm_all_subject_data_cache()

Warm the per-subject assignment-data cache, when this design uses covariates. This is an internal optimization hook for randomization inference; it keeps cache mutation inside the Design object instead of exposing its private environment to callers.

Usage
Design$warm_all_subject_data_cache()
Returns

TRUE invisibly when a cache warm-up was attempted, or FALSE invisibly when the design does not use covariates.


Design$get_n()

Get n, the sample size

Usage
Design$get_n()
Returns

The number of subjects.


Design$get_y_L()

Get y_L

Usage
Design$get_y_L()
Returns

A numeric vector of censored-response lower bounds (NA for exact-response subjects).


Design$get_y_R()

Get y_R

Usage
Design$get_y_R()
Returns

A numeric vector of censored-response upper bounds (NA for exact-response subjects).


Design$get_effective_time()

Get the effective response time per subject: the exact value y where recorded, or the lower bound y_L for a censored subject. This reconstructs "the one informative number" every response type other than left-/interval-censored survival data has always had, for code that needs a single numeric value per subject rather than the y/y_L/ y_R triple directly.

Usage
Design$get_effective_time()
Returns

A numeric vector, one value per subject.


Design$get_effective_dead()

Get the effective event indicator per subject: 1 for an exact response, 0 for a censored one. This reconstructs today's dead semantics for right-censored survival data (and is trivially all-1 for every other response type, which never has censoring). It is only valid for exact/right-censored data – a left- or interval-censored subject also returns 0 here, which is not meaningful right-censoring status, so callers must confirm (e.g. via any_censoring() plus their own censoring-shape checks) that no such rows are present before relying on this value.

Usage
Design$get_effective_dead()
Returns

An integer vector, one value per subject.


Design$get_prob_T()

Get probability of treatment

Usage
Design$get_prob_T()
Returns

The specified probability.


Design$get_response_type()

Get response type

Usage
Design$get_response_type()
Returns

The specified response type.


Design$get_response_type_original()

Get the original response type

Usage
Design$get_response_type_original()
Returns

The original specified response type.


Design$get_ordinal_levels()

Get ordinal levels

Usage
Design$get_ordinal_levels()
Returns

The levels of the ordinal response.


Design$get_original_ordinal_levels()

Get original ordinal levels

Usage
Design$get_original_ordinal_levels()
Returns

The labels for the levels of the original ordinal response.


Design$get_missingness_method()

Get the missingness method

Usage
Design$get_missingness_method()
Returns

The missingness handling method: "impute", "drop_column", or "error".


Design$get_edi_version_created()

Get the EDI package version this object was created under.

Stamped once, at construction time, from utils::packageVersion("EDI"); never re-stamped on readRDS() reload, so it reflects the version that originally built the object rather than whatever version is currently loaded. Objects saved before this field existed self-initialize it to the currently loaded version the first time it is read (there is no way to recover the true original version for those objects), rather than erroring on the missing field.

Usage
Design$get_edi_version_created()
Returns

A character string, e.g. "1.0.0".


Design$transform_y()

Transform the response vector y

Usage
Design$transform_y(
  transform_fun,
  transformed_response_type,
  ordinal_levels = NULL
)
Arguments
transform_fun

A function that takes y_original and returns a new y.

transformed_response_type

The response type of the transformed y.

ordinal_levels

If the transformed response type is "ordinal", the labels for the levels.


Design$get_design_formula()

Get the model formula

Usage
Design$get_design_formula()
Returns

The model formula.


Design$duplicate()

Duplicate this design object

Usage
Design$duplicate(verbose = FALSE)
Arguments
verbose

A flag for verbosity.

Returns

A new 'Design' object with the same data


Design$clone()

The objects of this class are cloneable with this method.

Usage
Design$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

# Design is abstract and cannot be instantiated directly; construct a
# concrete subclass instead, e.g.:
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))

Internal base for user-defined sequential-design extensions

Description

DesignCustomSequential is intentionally not exported. Subclasses implement assignment_rule() and return a scalar 0/1 assignment for the current subject. EDI handles subject storage, responses, and redraws through DesignSeqOneByOne.

Super classes

Design -> DesignSeqOneByOne -> DesignCustomSequential

Methods

Public methods

+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design

DesignCustomSequential$assignment_rule()

User-defined assignment rule.

Usage
DesignCustomSequential$assignment_rule()
Returns

A binary treatment assignment.


DesignCustomSequential$assign_wt()

Standard internal assignment entry point.

Usage
DesignCustomSequential$assign_wt()
Returns

A binary treatment assignment.


DesignCustomSequential$clone()

The objects of this class are cloneable with this method.

Usage
DesignCustomSequential$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


A Fixed Design

Description

An abstract R6 Class encapsulating the data and functionality for a fixed experimental design. This class takes care of whole-experiment randomization.

Super class

Design -> DesignFixed

Methods

Public methods

+ inherited public methods from Design

DesignFixed$new()

Initialize a fixed experimental design

Usage
DesignFixed$new(
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL,
  ...
)
Arguments
response_type

"continuous", "incidence", "proportion", "count", "survival", or "ordinal".

prob_T

Probability of treatment assignment.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

A flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

...

Extra arguments passed to the Design superclass.

Returns

A new 'DesignFixed' object


DesignFixed$assign_w_to_all_subjects()

Assign treatment to all subjects in the fixed experiment.

Usage
DesignFixed$assign_w_to_all_subjects(w_precomputed = NULL)
Arguments
w_precomputed

Optional {0,1} numeric vector of length n. If supplied the allocation is used directly and draw_ws_according_to_design is not called (avoids e.g. the Java round-trip for DesignFixedGreedy).


DesignFixed$add_all_subjects_to_experiment()

Add all subjects' covariates to a fixed design at once.

Usage
DesignFixed$add_all_subjects_to_experiment(X_all)
Arguments
X_all

A data frame containing the full covariate matrix.

Returns

Invisibly returns the design object.


DesignFixed$add_all_subject_responses()

Add all subject responses for a fixed design.

Usage
DesignFixed$add_all_subject_responses(ys = NULL, y_Ls = NULL, y_Rs = NULL)
Arguments
ys

The exact responses as a numeric vector, NA for any subject whose response is censored (supply y_Ls/y_Rs for those instead).

y_Ls

The censored-response lower bounds, NA for any subject with an exact response in ys. Right-censored: the last known event-free time (pair with y_Rs = Inf). Left-censored: 0, stated explicitly. Interval-censored: the interval's lower bound. Storage accepts any well-formed left-/interval-censored value; whether a given Inference class can actually consume it depends on that class (most survival Inference classes still only accept exact/ right-censored data and will reject construction with a clear error otherwise – see individual class docs).

y_Rs

The censored-response upper bounds, NA for any subject with an exact response in ys. Right-censored: Inf. Left-/interval-censored: the confirmed-by time / interval upper bound.


DesignFixed$overwrite_all_subject_assignments()

Overwrite all subject assignments for a fixed design.

Usage
DesignFixed$overwrite_all_subject_assignments(w)
Arguments
w

A {0,1} vector of subject assignments (1 = treated, 0 = control).


DesignFixed$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixed$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

# DesignFixed is abstract and cannot be instantiated directly; construct a
# concrete subclass instead, e.g.:
des = DesignFixedBernoulli$new(n = 10, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(10)))
des$assign_w_to_all_subjects()

A Fixed-Sample-Size Bernoulli (Independent-Coin-Flip) Randomized Design

Description

A fixed-sample-size DesignFixed in which each subject's treatment assignment w_i is drawn independently as w_i \stackrel{iid}{\sim} \mathrm{Bernoulli}(p), i = 1, \dots, n, where p is prob_T. This is the classical Bernoulli (independent-coin-flip) randomized design: unlike DesignFixediBCRD (complete randomization), which fixes the number of treated subjects at exactly \mathrm{round}(np), a Bernoulli design leaves the realized number of treated subjects n_T = \sum_i w_i \sim \mathrm{Binomial}(n, p) random; the trade-off is independence across subjects (useful for some asymptotic/martingale arguments) at the cost of not guaranteeing exact balance, which can matter for small n or for inference procedures (e.g. exact permutation tests over a fixed number of treated) that assume a fixed n_T.

Draw mechanism. draw_ws_raw(r) delegates to generate_permutations_bernoulli_cpp(), which fills an n \times r matrix of independent \mathrm{Bernoulli}(p) draws (one column per requested replicate, via a Mersenne Twister RNG seeded once per call from R's RNG state), so r replicate allocation vectors are generated with a single C++ call rather than r separate calls into R's own random-number generation. assign_w_to_all_subjects() draws a single such allocation (r = 1) and applies it to all subjects at once.

No exchange/balance search. Because subjects are treated independently, there is no optimization step analogous to DesignFixedGreedyDOptimal: covariates, if supplied, do not influence the assignment probabilities or realized allocation at all.

Super classes

Design -> DesignFixed -> DesignFixedBernoulli

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixedBernoulli$is_a_bernoulli_capable()

Characterization: this design draws each subject's treatment assignment as an independent \mathrm{Bernoulli}(p) coin flip (see class documentation), so it is Bernoulli-capable by construction.

Usage
DesignFixedBernoulli$is_a_bernoulli_capable()
Returns

Always TRUE for this class.


DesignFixedBernoulli$new()

Initialize a fixed Bernoulli (independent-coin-flip) experimental design. Unlike DesignFixediBCRD, the realized number of treated subjects is not fixed at n * prob_T; it is random (\mathrm{Binomial}(n, prob\_T)) because each subject's assignment is an independent coin flip (see class documentation).

Usage
DesignFixedBernoulli$new(
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

"continuous", "incidence", "proportion", "count", "survival", or "ordinal".

prob_T

Per-subject probability p that a given subject is assigned to treatment; need not be 0.5 (unlike DesignFixedGreedyDOptimal).

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

A flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignFixedBernoulli' object


DesignFixedBernoulli$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixedBernoulli$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Neyman, J. (1923, transl. 1990). "On the Application of Probability Theory to Agricultural Experiments." Statistical Science, 5(4), 465-472, for the potential-outcomes framework under which Bernoulli and complete randomization are compared; see also randomized experiment for orientation on Bernoulli vs. complete (restricted) randomization.

Examples

des = DesignFixedBernoulli$new(n = 10, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(10)))
des$assign_w_to_all_subjects()

A Fixed, Non-Bipartite-Matched-Pair Design with Within-Pair Randomization

Description

A fixed-sample-size DesignFixed that (1) partitions the n subjects into n/2 disjoint matched pairs by solving a non-bipartite (optimal) pairwise-matching problem on a covariate distance matrix, minimizing the total within-pair distance across all pairs, then (2) randomizes treatment within each pair independently: for pair k, one of its two members is assigned to treatment and the other to control with probability 1/2 each, independently across pairs. This is the classical matched-pair randomized design (a special case of blocking with block size 2), which guarantees exact covariate balance within every pair (up to the matching algorithm's distance metric) while preserving randomization-based inference validity, in contrast to DesignFixedGreedyDOptimal, which balance the covariates in aggregate via an optimality criterion but do not guarantee pairwise closeness of any two subjects.

Matching algorithm. Pairing is computed once (lazily, on first call to draw_ws_raw()/assign_w_to_all_subjects(), via private$ensure_matching_structure_computed()) by compute_binary_match_structure(), which forms an n \times n pairwise distance matrix — squared Mahalanobis distance if mahal_match = TRUE (the default; using the sample covariance of the covariates, ridge-regularized if singular), or squared Euclidean distance otherwise — and solves the minimum-weight non-bipartite perfect matching on that distance matrix via nonbimatch (nbpMatching, a Suggests-only dependency; loading is deferred until matching is actually needed, so pre-computed w vectors injected via m never require it). For a single covariate (p = 1), pairing instead reduces to simply sorting subjects by that covariate and pairing consecutive subjects, since the non-bipartite matching problem is trivial in one dimension. The resulting pairing is cached in private$bms/private$m for the lifetime of the design object (or until explicitly reset via set_m()) and is not recomputed per draw.

Within-pair randomization. Given the fixed pairing, each replicate allocation (see draw_binary_match_assignments_cpp()) independently flips, for every pair, which of its two members is treated (an independent fair coin flip per pair per replicate, using a splitmix64-seeded Mersenne Twister per replicate column for reproducible parallel draws); this guarantees exactly n/2 treated subjects overall (only prob_T = 0.5 is supported; the constructor errors otherwise). Passing a pre-computed m to the constructor supplies the matched-pair structure directly (each pair ID occurring in exactly 2 rows), bypassing the matching computation entirely while keeping the same within-pair randomization.

No-covariate fallback. If no covariates are available at draw time (private$m is NULL, e.g. matching hasn't run and no explicit m was supplied), draw_ws_raw() falls back to an unmatched balanced complete randomization (a uniformly random permutation of n/2 ones and n/2 zeros), since there is no covariate information to match on.

Batch pregeneration. draw_binary_match_assignments_cpp()'s output is trusted unvalidated – it guarantees exactly n x r valid \{0,1\} columns with n/2 treated subjects per column by construction (see fix_design_hierarchy.md, "AllocationMatrixValidation"). supports_batch_w_pregeneration() returns TRUE so that the calling framework generates all replicate w vectors for a simulation cell in one batch (amortizing the one-time nbpMatching matching cost across all replicates of that cell) rather than recomputing the matching structure per replicate.

Super classes

Design -> DesignFixed -> DesignFixedBinaryMatch

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixedBinaryMatch$is_a_kk_matching_capable()

Characterization: this design computes its matched-pair structure on the fly from covariates (see class documentation), so it is KK matching-capable by construction.

Usage
DesignFixedBinaryMatch$is_a_kk_matching_capable()
Returns

Always TRUE for this class.


DesignFixedBinaryMatch$supports_batch_w_pregeneration()

Returns TRUE so the calling framework pre-generates all replicate w vectors for a simulation cell in one batch, paying the one-time nbpMatching non-bipartite matching cost once per cell and reusing the resulting pairing across all replicates, rather than recomputing it per replicate.

Usage
DesignFixedBinaryMatch$supports_batch_w_pregeneration()
Returns

Always TRUE for this class.


DesignFixedBinaryMatch$new()

Initialize a binary (non-bipartite) matched-pair fixed experimental design. Matching itself is deferred until the first draw (see class documentation); this constructor only records configuration and, if m is supplied, installs the explicit pairing immediately.

Usage
DesignFixedBinaryMatch$new(
  response_type,
  prob_T = 0.5,
  mahal_match = TRUE,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  m = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

The data type of response values.

prob_T

The probability of the treatment assignment. Must be 0.5, since within-pair randomization only supports an even 1-treated/1-control split per pair.

mahal_match

Match using squared Mahalanobis distance (accounting for covariate correlation/scale) if TRUE (default), else squared Euclidean distance on the raw covariate matrix.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

m

Optional integer vector of explicit matched-pair identifiers, one per subject. If supplied, 'n' must also be supplied, 'length(m)' must equal 'n', all values must be positive, and each pair ID must occur exactly twice. This bypasses the package-computed matching step.

verbose

Flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignFixedBinaryMatch' object


DesignFixedBinaryMatch$assign_w_to_all_subjects()

Assign treatment to all subjects (see DesignFixed$assign_w_to_all_subjects() for the general contract). Before delegating, this override ensures the matched-pair structure is computed (private$ensure_matching_structure_computed()) even when w_precomputed is supplied and draw_ws_according_to_design() is therefore never called — downstream code (e.g. blocked/matched-pair inference) still needs private$m to be populated regardless of how w was obtained.

Usage
DesignFixedBinaryMatch$assign_w_to_all_subjects(w_precomputed = NULL)
Arguments
w_precomputed

Optional {0,1} numeric vector of length n. If supplied, it is used directly as the treatment allocation instead of drawing a fresh within-pair-randomized allocation.


DesignFixedBinaryMatch$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixedBinaryMatch$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Greevy, R., Lu, B., Silber, J. H., and Rosenbaum, P. (2004). "Optimal multivariate matching before randomization." Biostatistics, 5(2), 263-275, doi:10.1093/biostatistics/5.2.263, for optimal non-bipartite matched-pair designs prior to randomization. See also matched pair and Mahalanobis distance for orientation.

Examples

des = DesignFixedBinaryMatch$new(n = 10, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(10)))
des$assign_w_to_all_subjects()

A Fixed, Blocked-and-Clustered Randomized Design

Description

A fixed-sample-size DesignFixed in which the unit of randomization is the cluster, not the individual subject: within each block (stratum, formed from strata_cols), whole clusters (identified by cluster_col) are jointly randomized to treatment or control, so all subjects in the same cluster always receive the same assignment. This is the design used when individual-level randomization is infeasible or invalid (e.g. clusters are classrooms, clinics, or households where within-cluster interference/spillover would violate SUTVA under individual randomization), combined with blocking to improve precision by comparing clusters only to other clusters in the same stratum.

Randomization mechanism. Blocking keys are computed per subject via private$get_strata_keys() (shared with other blocking-structure designs): categorical columns in strata_cols are used as-is, continuous columns are discretized into preferred_num_bins_for_continuous_covariate quantile-based bins, and multiple strata_cols are combined into a single composite block key. Within each resulting block, whole clusters (by cluster_col) are randomized to treatment with probability prob_T via block_and_cluster_ra (randomizr), which performs blocked-and-clustered complete random assignment: within each block, clusters (not subjects) are permuted so that, subject to rounding, the target proportion prob_T of clusters in that block is treated, and every subject in a treated cluster receives w = 1. r independent replicate allocation columns are generated via replicate() (one randomizr call per replicate; there is no batch/vectorized draw path for this design, unlike DesignFixedBinaryMatch).

Cluster-aware bootstrap. draw_bootstrap_indices() overrides the default subject-level bootstrap to resample at the cluster level via resample_group_rows_cpp(): with bootstrap_type = "within_blocks" (the default when bootstrap_type is NULL), clusters are resampled with replacement within each block, preserving the block structure; otherwise, whole blocks (strata) are themselves resampled with replacement. This mirrors the standard cluster-robust bootstrap principle that resampling must occur at the level of the randomization unit (clusters), not individual subjects, to yield a valid variance/interval estimate under cluster-correlated outcomes.

Super classes

Design -> DesignFixed -> DesignFixedBlockedCluster

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixedBlockedCluster$is_a_cluster_capable()

Characterization: this design randomizes whole clusters (see class documentation), so it is cluster-structured by construction.

Usage
DesignFixedBlockedCluster$is_a_cluster_capable()
Returns

Always TRUE for this class.


DesignFixedBlockedCluster$new()

Initialize a blocked and cluster randomized fixed experimental design.

Usage
DesignFixedBlockedCluster$new(
  strata_cols,
  cluster_col,
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  preferred_num_bins_for_continuous_covariate = 2,
  num_bins_for_continuous_covariate = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
strata_cols

A character vector of column names to use for stratification (blocks).

cluster_col

The column name in the data that identifies the cluster for each subject.

response_type

The data type of response values.

prob_T

The target probability that a given cluster within a block is assigned to treatment (subjects inherit their cluster's assignment).

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

preferred_num_bins_for_continuous_covariate

The number of quantile bins to use for continuous strata. Default is 2.

num_bins_for_continuous_covariate

Deprecated alias for 'preferred_num_bins_for_continuous_covariate'.

verbose

Flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignFixedBlockedCluster' object


DesignFixedBlockedCluster$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixedBlockedCluster$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Middleton, J. A., and Aronow, P. M. (2015). "Unbiased estimation of the average treatment effect in cluster-randomized experiments." Statistics, Politics and Policy, 6(1-2), 39-75, doi:10.1515/spp-2013-0002, for blocked/clustered randomized-assignment inference; see also the randomizr package vignette for the assignment-generation conventions this design relies on, and cluster randomized controlled trial for orientation.

Examples

des = DesignFixedBlockedCluster$new(n = 20, response_type = 'continuous',
  strata_cols = 'x2', cluster_col = 'cl')
X = data.frame(x1 = rnorm(20), x2 = factor(rep(1:2, each = 10)), cl = factor(rep(1:10, each = 2)))
des$add_all_subjects_to_experiment(X)
des$assign_w_to_all_subjects()

A Fixed, Stratified-Block Randomized Design

Description

A fixed-sample-size DesignFixed that first partitions subjects into blocks (strata) formed from covariates, then randomizes treatment independently within each block at probability prob_T (via block_ra when randomizr is installed, else an internal generate_permutations_blocking_cpp() fallback). Blocking on a covariate removes its between-block variation from the treatment-effect comparison (comparisons are always within-block), improving precision relative to unblocked randomization whenever the blocking covariate(s) are prognostic of the outcome, at the cost of requiring the analysis to account for the blocking structure (e.g. via a block/stratum fixed effect or a CMH-type test). This differs from DesignFixedBlockedCluster, which randomizes whole clusters of subjects together within each block rather than subjects individually.

Block construction. Blocking keys are computed by private$get_strata_keys() (shared across blocking-structure designs): each column in strata_cols contributes a categorical key (continuous columns are discretized into preferred_num_bins_for_continuous_covariate quantile bins), and multiple columns are combined into one composite block key per subject; if strata_cols is NULL, all available covariate columns are used. B_target caps the number of resulting blocks by greedily adding strata_cols in order only while the running block count stays at or below the target (earlier columns take priority); exact_num_blocks = TRUE instead hard-fails if the greedy construction does not land on exactly B_target blocks. equal_block_sizes = TRUE (the default) additionally requires every block to have the same subject count, checked once at construction (via n %% B_target) if n and B_target are both already known, and again once covariates arrive; some downstream inference classes (InferenceIncidCMH, InferenceIncidExtendedRobins) require equal block sizes unconditionally, regardless of this flag. An explicit m (one block ID per subject) bypasses covariate-derived block construction entirely.

Within-block randomization and bootstrap. Within each block, treatment is assigned independently via block_ra's complete random assignment (subject to rounding, prob_T of each block's subjects are treated); the internal C++ fallback (generate_permutations_blocking_cpp()) is used only if randomizr is not installed. draw_bootstrap_indices() resamples within each block by default (bootstrap_type = "within_blocks" or NULL, via stratified_bootstrap_indices_cpp()), or resamples whole blocks with replacement otherwise (via resample_group_rows_cpp()) — mirroring the block structure in the resampling scheme, analogous to the cluster-level bootstrap in DesignFixedBlockedCluster.

Super classes

Design -> DesignFixed -> DesignFixedBlocking

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixedBlocking$new()

Initialize a fixed stratified-block randomized experimental design. Block construction and validation follow the rules described in the class documentation; see the parameter descriptions below for the greedy B_target/exact_num_blocks/equal_block_sizes contract.

Usage
DesignFixedBlocking$new(
  strata_cols = NULL,
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  preferred_num_bins_for_continuous_covariate = 2,
  B_target = NULL,
  exact_num_blocks = FALSE,
  equal_block_sizes = TRUE,
  m = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
strata_cols

A character vector of column names to use for stratification. If 'NULL' (the default), all available covariate columns are used.

response_type

"continuous", "incidence", "proportion", "count", "survival", or "ordinal".

prob_T

Probability of treatment assignment.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

preferred_num_bins_for_continuous_covariate

The number of quantile bins to use for continuous strata. Default is 2.

B_target

The target number of blocks. Columns from 'strata_cols' are added greedily in order, each column being included only if it does not push the total number of unique blocks beyond this target. For categorical covariates their natural levels are used; for continuous covariates 'preferred_num_bins_for_continuous_covariate' quantile bins are used. Earlier columns are always preferred over later ones. When 'n' is known at construction time, the default is the largest divisor of 'n' that is at most 'floor(sqrt(n))' (so the default always satisfies 'equal_block_sizes = TRUE'); if 'n' is not yet known, it is resolved to 'floor(sqrt(n))' when subjects are added. Set 'B_target = NULL' to use all columns unconditionally. An explicitly supplied 'B_target' that does not divide 'n' still errors immediately when 'equal_block_sizes = TRUE'. Set 'exact_num_blocks = TRUE' to hard fail if the final key construction does not produce exactly 'B_target' blocks.

exact_num_blocks

Whether to require the greedy key construction to produce exactly 'B_target' blocks. Default 'FALSE'.

equal_block_sizes

Whether to require all blocks to have the same number of subjects. Default 'TRUE'. When 'TRUE' and both 'n' and 'B_target' are known at construction time, an error is raised immediately if 'n' is not divisible by 'B_target'. A second check fires when subjects are added: if the covariate-based strata produce unequal block counts the design errors at that point. Set to 'FALSE' to allow unequal blocks (note that 'InferenceIncidCMH' and 'InferenceIncidExtendedRobins' still require equal block sizes regardless).

m

Optional integer vector of explicit block identifiers, one per subject. If supplied, 'n' must also be supplied and 'length(m)' must equal 'n'. The constructor then records this blocking structure immediately via 'set_m()', bypassing covariate-derived strata construction.

verbose

A flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignFixedBlocking' object


DesignFixedBlocking$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixedBlocking$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Fisher, R. A. (1935). The Design of Experiments. Oliver and Boyd, for the original rationale for blocking in randomized experiments; Cochran, W. G., and Cox, G. M. (1957). Experimental Designs (2nd ed.), Wiley, for stratified (randomized block) design theory. See also randomized block design for orientation.

Examples

des = DesignFixedBlocking$new(n = 20, response_type = 'continuous',
  strata_cols = 'x2', equal_block_sizes = FALSE)
X = data.frame(x1 = rnorm(20), x2 = factor(rep(1:2, 10)))
des$add_all_subjects_to_experiment(X)
des$assign_w_to_all_subjects()

A Fixed, Unblocked Cluster Randomized Design

Description

A fixed-sample-size DesignFixed in which whole clusters of subjects (identified by cluster_col), rather than individual subjects, are the unit of randomization: every subject in a given cluster always receives the same treatment assignment. This is the unblocked analog of DesignFixedBlockedCluster — there is no stratification step here, so clusters are randomized to treatment as a single pool rather than within strata. Cluster-level randomization is required whenever individual-level randomization would create within-cluster interference/spillover that violates SUTVA (e.g. clusters are classrooms, clinics, villages, or households), at the cost of an effective sample size driven by the number of clusters, not subjects, and a corresponding need for cluster-aware inference.

Randomization mechanism. draw_ws_raw(r) extracts each subject's cluster ID from cluster_col (erroring if any are missing) and calls cluster_ra (randomizr) once per replicate, which performs complete random assignment at the cluster level: subject to rounding, prob_T of clusters are assigned to treatment, and all subjects sharing a cluster inherit that cluster's assignment. r independent replicate columns are generated via replicate() (one randomizr call per replicate).

Cluster-aware bootstrap. draw_bootstrap_indices() overrides the default subject-level bootstrap to resample whole clusters with replacement (via resample_group_rows_cpp()) rather than individual rows, since outcomes are correlated within a cluster (shared assignment plus, typically, shared context) and the exchangeable resampling unit for a valid bootstrap variance/interval estimate is therefore the cluster, not the subject.

Super classes

Design -> DesignFixed -> DesignFixedCluster

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixedCluster$is_a_cluster_capable()

Characterization: this design randomizes whole clusters (see class documentation), so it is cluster-structured by construction.

Usage
DesignFixedCluster$is_a_cluster_capable()
Returns

Always TRUE for this class.


DesignFixedCluster$new()

Initialize a cluster randomized fixed experimental design (no blocking/stratification; see DesignFixedBlockedCluster if stratification is also needed).

Usage
DesignFixedCluster$new(
  cluster_col,
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
cluster_col

The column name in the data that identifies the cluster for each subject.

response_type

The data type of response values.

prob_T

The target probability that a given cluster is assigned to treatment (subjects inherit their cluster's assignment).

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

Flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignFixedCluster' object


DesignFixedCluster$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixedCluster$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Middleton, J. A., and Aronow, P. M. (2015). "Unbiased estimation of the average treatment effect in cluster-randomized experiments." Statistics, Politics and Policy, 6(1-2), 39-75, doi:10.1515/spp-2013-0002. See also cluster randomized controlled trial for orientation, and DesignFixedBlockedCluster for the blocked variant of this design.

Examples

des = DesignFixedCluster$new(n = 20, response_type = 'continuous', cluster_col = 'cl')
X = data.frame(x = rnorm(20), cl = factor(rep(1:5, each = 4)))
des$add_all_subjects_to_experiment(X)
des$assign_w_to_all_subjects()

Internal base for user-defined fixed-design extensions

Description

DesignFixedCustom is intentionally not exported. Subclasses implement draw_assignments(r = 1) and return an n x r 0/1 assignment matrix. EDI handles subject storage, responses, and validation through DesignFixed.

Super classes

Design -> DesignFixed -> DesignFixedCustom

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixedCustom$draw_assignments()

Draw assignments from the custom design.

Usage
DesignFixedCustom$draw_assignments(r = 1)
Arguments
r

Number of assignment vectors to draw.

Returns

An n x r matrix of 0/1 assignments.


DesignFixedCustom$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixedCustom$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


A Fixed, Balanced Two-Arm Factorial Design

Description

A fixed-sample-size DesignFixed for factorial treatment structures: subjects are assigned to one of the cells of a factorial combination of one or more named factors (e.g. list(treatment = 2), or, once multi-arm support lands, list(drug = 2, dose = 2) for a 2 \times 2 design), with assignment counts balanced as evenly as possible across cells within each replicate draw.

Currently restricted to exactly two total factor-level combinations (i.e. two arms), e.g. a single two-level factor — the product of levels across all factors entries must equal exactly 2; the constructor errors otherwise. In this two-arm regime, the design reduces to a balanced complete randomization between cell 1 (w = 0) and cell 2 (w = 1), and w follows the same {0,1} internal / {-1,+1} public convention as every other Design subclass, so DesignFixedFactorial inherits assign_w_to_all_subjects(), draw_ws_according_to_design(), and get_w() unmodified from DesignFixed/Design and works unmodified with every Inference class; only draw_ws_raw() (the low-level allocation-vector generator) and get_w_factorial() (an additional factor-level accessor, see below) are specific to this class. Support for more than two combinations (true multi-factor, multi-arm designs) is tracked separately — see package_metadata/new_feature_plans/multi_arm_designs.md.

Allocation generation. draw_ws_raw(r) builds a base allocation vector by repeating the sequence of cell indices 0:(num_combinations - 1) out to length n (so cells are as close to equally represented as possible, off by at most one subject when n is not a multiple of the number of cells), then independently permutes ( sample) that base vector once per replicate column to produce r balanced-but-randomized allocations.

Super classes

Design -> DesignFixed -> DesignFixedFactorial

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixedFactorial$new()

Initialize a factorial fixed experimental design. The product of levels implied by factors must currently equal exactly 2 (see class documentation); any other total raises an error.

Usage
DesignFixedFactorial$new(
  factors,
  response_type,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
factors

A list where names are factor names and values are number of levels (e.g. list(treatment = 2)). The product of levels across all factors must currently equal exactly 2 (two-arm only).

response_type

The data type of response values.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

Flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignFixedFactorial' object


DesignFixedFactorial$get_w_factorial()

Decode each subject's scalar cell index (private$w, in 0:(num_combinations - 1)) back into its per-factor level assignments, using the same expand.grid enumeration (private$combinations) established at construction. This is the inverse of the encoding draw_ws_raw() produces, and is the only way to recover individual factor levels once support for more than two total combinations lands, since get_w() (inherited, see class documentation) only ever returns the scalar 0/1 (or -1/+1) cell index.

Usage
DesignFixedFactorial$get_w_factorial()
Returns

A data frame with n rows and one column per entry of factors, giving each subject's level (an integer in 1:levels) for that factor; NULL if treatment has not yet been assigned to all subjects (i.e. private$w is empty or contains NA).


DesignFixedFactorial$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixedFactorial$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

des = DesignFixedFactorial$new(n = 12, response_type = 'continuous', factors = list(treatment = 2))
des$add_all_subjects_to_experiment(data.frame(x=1:12))
des$assign_w_to_all_subjects()

A Fixed, Covariate-Balanced Design via Greedy Pairwise-Swap Search

Description

A fixed-sample-size DesignFixed that searches, among balanced (n/2-treated) allocations, for one that directly minimizes a covariate imbalance criterion f(d), d = M(2w - 1), via a native C++ (RcppEigen + OpenMP) greedy pairwise-swap search (greedy_design_search_cpp()). Unlike DesignFixedGreedyDOptimal, which optimizes a model-based information-matrix criterion (D-/A-optimality) implied by an assumed linear model, this design optimizes a direct covariate-distance criterion between the treated and control group means/covariance, with no linear-model assumption: the two supported objectives are

Search algorithm. Starting from a random balanced (Fisher-Yates) allocation, the search runs in one of two modes selected by n_iter: Inf (default) runs exhaustive best-improvement search — each round scans every (treated, control) pair, applies the single globally best improving swap, and repeats until no swap improves f(d), guaranteeing convergence to a strict local optimum; a positive integer instead runs exactly that many stochastic steps, each picking a uniformly random (treated, control) pair and accepting the swap only if it improves f(d) (with patience-based early stopping). r independent design searches (one per requested replicate) run in parallel via OpenMP, each with its own std::mt19937 generator. This search's randomization is reproducible via the constructor's seed argument: per-thread RNGs are seeded from R's own RNG state (GetRNGstate()/unif_rand()) before the parallel region begins, so private$maybe_set_seed() does govern the resulting allocation, independent of the number of OpenMP threads used.

Constraints and fallbacks. Only exactly balanced allocation (prob_T = 0.5, n even) is supported; the constructor errors otherwise. If no covariates are available, the search degenerates to pure balanced Fisher-Yates randomization (no swap search, since there is nothing to balance on). greedy_design_search_cpp()'s output is trusted unvalidated – it guarantees exactly n x r valid \{0,1\} columns with n/2 treated subjects per column by construction (see fix_design_hierarchy.md, "AllocationMatrixValidation").

Super classes

Design -> DesignFixed -> DesignFixedGreedy

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixedGreedy$new()

Initialize a greedy pairwise-swap search fixed experimental design. Only prob_T = 0.5 is supported (see class documentation).

Usage
DesignFixedGreedy$new(
  response_type,
  prob_T = 0.5,
  objective = "mahal_dist",
  n_iter = Inf,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

The data type of response values.

prob_T

The probability of the treatment assignment. Must be 0.5.

objective

The covariate-imbalance objective to minimize: either "mahal_dist" (default, squared Mahalanobis distance between treated and control covariate means) or "abs_sum_diff" (sum of absolute standardized mean differences); see class documentation for the exact criteria.

n_iter

Number of swap iterations. Inf (default) uses exhaustive best-improvement search guaranteed to reach a strict local optimum. A positive integer runs that many stochastic random-pair iterations with patience-based early stopping.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

Flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility. This design's search is reproducible via seed (see class documentation).

Returns

A new 'DesignFixedGreedy' object


DesignFixedGreedy$supports_batch_w_pregeneration()

Returns TRUE so the calling framework pre-generates all replicate w vectors for a simulation cell in a single batched call to greedy_design_search_cpp() (which parallelizes the r independent searches over OpenMP threads internally), rather than issuing r separate single-replicate C++ calls.

Usage
DesignFixedGreedy$supports_batch_w_pregeneration()
Returns

Always TRUE for this class.


DesignFixedGreedy$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixedGreedy$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Krieger, A. M., Azriel, D., and Kapelner, A. (2019). "Nearly random designs with greatly improved balance." Biometrika, 106(3), 695-701, doi:10.1093/biomet/asz026, for the greedy-swap balance-optimization approach this class implements. See also Mahalanobis distance for orientation on the default objective.

Examples

des = DesignFixedGreedy$new(n = 10, response_type = 'continuous')

A Fixed, Model-Based Optimal Design via Greedy Pairwise-Exchange Search

Description

A fixed-sample-size DesignFixed that searches, among allocations with exactly n_T = \mathrm{round}(n \cdot \mathrm{prob}_T) treated subjects, for allocations optimizing a model-based information-matrix criterion implied by the linear model y = \beta_T w + Z_0 \gamma + \epsilon with Z_0 = [1\ X], via a native C++ greedy pairwise-exchange (Fedorov/ DETMAX-style) local search. This class is the merger of the former DesignFixedDOptimal and DesignFixedAOptimal classes; the criterion is selected by the objective and interest constructor arguments.

The optimality-criterion family and its argument mapping. Write M(w) = [w\ Z_0]^\top [w\ Z_0] for the information (moment) matrix, P = Z_0 (Z_0^\top Z_0)^{-1} Z_0^\top, and s(w) = n_T - w^\top P w (the treatment-coefficient information given the covariate block). The classical criteria map to constructor arguments as follows:

D_M – full-matrix determinant optimality (maximize \det M(w), i.e. |M|)

objective = "D", interest = "all". By the Schur-complement identity \det M(w) = \det(Z_0^\top Z_0) \cdot s(w) (with w^\top w = n_T fixed and Z_0 not depending on w), the covariate block factors out, so D_M reduces to maximizing s(w).

D_s – subset determinant optimality (minimize \det(K^\top M(w)^{-1} K) for a coordinate-selection K: the treatment coefficient plus a chosen covariate subset)

the default objective = "D", interest = "treatment" is D_s with the interest set = {treatment}; interest = ~ x1 + x2 (a one-sided formula) or interest = c("x1", "x2") (model-matrix column names) selects treatment + those covariates. Because w enters only the treatment row/column of M(w), every such D_s criterion factorizes as \det(V_{SS})/s(w) with \det(V_{SS}) constant in w – so all determinant-type settings (D_M and every D_s) select identical allocations and share the same search kernel. For the single treatment contrast, D_s-, c-, and per-parameter A-optimality coincide as well, which is why objective = "A", interest = "treatment" is silently equivalent to the default (allowed by design; no message is emitted).

D_A – general contrast optimality (minimize \det(A^\top M(w)^{-1} A) for an arbitrary contrast matrix A)

interest = <contrast matrix> – arrives with Stage 2 of the merge plan (the generalized-criterion kernel) and currently raises an informative error, as do interest sets excluding the treatment coefficient.

D_B (and A_B) – Bayesian optimality (criteria computed on the posterior information M(w) + R)

prior_precision = a scalar \tau or a matrix R_0, combined with either objective; see the Bayesian section below for exactly which coefficients a scalar \tau penalizes.

A – trace optimality (minimize \mathrm{tr}(K^\top M(w)^{-1} K))

objective = "A" with interest = "all" (all parameters: objective (w^\top H w + 1)/s(w), H = Z_0 (Z_0^\top Z_0)^{-2} Z_0^\top), or with interest = formula/names (A_s: same kernel with the subset-restricted H_S; see below). Unlike the determinant family, trace criteria over different interest sets generally select different allocations.

Bayesian variants. Supplying prior_precision replaces (Z_0^\top Z_0)^{-1} with the ridge-regularized (Z_0^\top Z_0 + R_0)^{-1} in the construction of P (and H), yielding Bayesian D_B/A_B-optimality. A scalar \tau penalizes the covariate coefficients only – the treatment coefficient and the intercept are unpenalized (R_0 = \tau \cdot \mathrm{diag}(0, 1, \ldots, 1) over Z_0's columns) – and, when standardize_covariates = TRUE (the default), the covariates are centered and scaled to unit variance first so \tau is interpretable per standardized coefficient. A full matrix prior_precision is used as R_0 verbatim (dimensions (1+p) \times (1+p) over [\mathrm{intercept}, \mathrm{covariates}] of the design's model matrix; standardize_covariates is ignored).

Search algorithm. For each of the r requested allocations independently: start from a uniformly random balanced-count allocation (a BCRD draw with exactly n_T treated), then repeatedly apply the single best improving treated/control pairwise exchange until no exchange improves the criterion (a strict local optimum). The returned allocations therefore form a restricted-randomization distribution over locally optimal allocations, which is what makes randomization inference possible for this design. The search is reproducible via the constructor's seed argument: the C++ kernels seed a local generator from R's own RNG stream, so a fixed seed yields identical draws (this corrects the former classes' documentation, which predated the RNG migration).

Covariate-subset criteria (D_s/A_s) via interest = a formula or names. interest also accepts a one-sided formula (e.g. ~ x1 + x2) or a character vector of model-matrix column names, meaning the treatment coefficient plus the named covariate coefficients (the treatment is always in the interest set; the intercept never is). Both reduce to the existing kernels with no new machinery: under objective = "D", because w only enters the treatment row/column of M(w), the subset determinant factorizes as \det(K^\top M(w)^{-1} K) = \det(V_{SS}) / s(w) with \det(V_{SS}) constant in w – so subset-D selects allocations identical to the default treatment-focused criterion (allowed silently, like objective = "A", interest = "treatment"); under objective = "A", the subset trace criterion is (w^\top H_S w + 1) / s(w) with H_S = (Z_0 V S)(Z_0 V S)^\top built from the selected columns – the same trace kernel with a subset-restricted H. Formula terms are expanded against the design's model matrix, so factor covariates must be referred to by their expanded model-matrix column names. Note that restricting the design's model matrix itself via design_formula also changes the default covariate set downstream inference adjusts for (Inference$initialize() inherits the design's formula), whereas interest affects the allocation criterion only. General contrast matrices (D_A), and interest sets excluding the treatment coefficient, arrive with Stage 2 of the merge plan (the generalized-criterion kernel; see package_metadata/finished_features/fix_design_hierarchy.md).

Constraints and fallbacks. prob_T may be any value in (0, 1) for which 1 \le \mathrm{round}(n \cdot \mathrm{prob}_T) \le n - 1. If no covariates are available, the search degenerates to pure random allocation with n_T treated (there is no criterion to optimize).

Super classes

Design -> DesignFixed -> DesignFixedGreedyDOptimal

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixedGreedyDOptimal$new()

Initialize a model-based optimal-search fixed experimental design. Covariates, if any, are supplied later via add_all_subjects_to_experiment(); the optimality search itself does not run until assign_w_to_all_subjects() (or draw_ws_according_to_design()) is called.

Usage
DesignFixedGreedyDOptimal$new(
  response_type,
  prob_T = 0.5,
  objective = "D",
  interest = "treatment",
  prior_precision = NULL,
  standardize_covariates = TRUE,
  n_iter = Inf,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

"continuous", "incidence", "proportion", "count", "survival", or "ordinal". Determines only which downstream inference/response machinery this design is paired with; it does not affect the optimality search itself.

prob_T

Probability of treatment assignment, in (0, 1). The search fixes the treated count at \mathrm{round}(n \cdot prob_T).

objective

The optimality criterion: "D" (default, determinant) or "A" (trace). See the class documentation for the exact criteria and for why objective = "A" with interest = "treatment" is equivalent to the default.

interest

Which parameters the criterion targets: "treatment" (default), "all", a one-sided formula (e.g. ~ x1 + x2), a single formula string (e.g. "x1 * x2 + x7", promoted to ~ x1 * x2 + x7), or a character vector of model-matrix column names – all but "all" meaning the treatment coefficient plus the named covariate coefficients (D_s/A_s; see class documentation, including why subset-D selects the same allocations as the default). Formula terms (including interactions like x1:x2) must correspond to columns of the design's model matrix: to target an interaction coefficient, the interaction must be in design_formula too – you cannot be "interested in" a coefficient the working model does not contain. Contrast matrices (general D_A) arrive with Stage 2 of the merge plan and currently raise an error.

prior_precision

NULL (default, non-Bayesian), a single positive scalar \tau (ridge prior precision on the covariate coefficients only; treatment and intercept unpenalized), or a full (1+p) \times (1+p) symmetric prior-precision matrix R_0 over [\mathrm{intercept}, \mathrm{covariates}].

standardize_covariates

If TRUE (default) and prior_precision is a scalar, covariates are centered and scaled to unit variance before the penalized criterion matrices are built. Ignored otherwise.

n_iter

Number of exchange iterations. Inf (default) runs the exhaustive best-improvement search to a strict local optimum. Finite values (the stochastic swap mode shared with DesignFixedGreedy) arrive with the Stage-2 shared search engine and currently raise an error.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

Sample size (if fixed).

verbose

Flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility. Unlike the former DesignFixedDOptimal/DesignFixedAOptimal documentation claimed, the optimality search is reproducible via seed (see class documentation).

Returns

A new 'DesignFixedGreedyDOptimal' object


DesignFixedGreedyDOptimal$get_objective()

The optimality criterion this design was constructed with.

Usage
DesignFixedGreedyDOptimal$get_objective()
Returns

"D" or "A".


DesignFixedGreedyDOptimal$get_interest()

The parameter-interest setting this design was constructed with.

Usage
DesignFixedGreedyDOptimal$get_interest()
Returns

"treatment" or "all".


DesignFixedGreedyDOptimal$get_prior_precision()

The Bayesian prior precision this design was constructed with.

Usage
DesignFixedGreedyDOptimal$get_prior_precision()
Returns

NULL, a positive scalar, or a symmetric matrix.


DesignFixedGreedyDOptimal$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixedGreedyDOptimal$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Atkinson, A. C., Donev, A. N., and Tobias, R. D. (2007). Optimum Experimental Designs, with SAS. Oxford University Press, for the D-/A-optimality criteria and exchange algorithms for constrained design search. See also optimal design for orientation.

Examples

des = DesignFixedGreedyDOptimal$new(n = 10, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(10)))
des$assign_w_to_all_subjects()

A Fixed, Matched-Pair Design with Greedy Which-Member-Treated Optimization

Description

A fixed-sample-size DesignFixed that combines DesignFixedBinaryMatch's non-bipartite matched-pair structure with DesignFixedGreedy's greedy imbalance-minimization search, restricted so every move respects the pairing: subjects are first paired by covariate closeness (as in DesignFixedBinaryMatch), guaranteeing exactly one treated and one control subject per pair; then, rather than assigning within-pair treatment status by a coin flip, the greedy search (greedy_design_search_cpp(), pair-constrained mode) chooses which member of each pair is treated so as to directly minimize the same aggregate covariate-imbalance objective as DesignFixedGreedy (squared Mahalanobis distance or sum of absolute standardized mean differences between the treated and control group means) across the whole sample, not just within each pair. This targets both close within-pair matches (from the matching step) and low aggregate covariate imbalance (from the greedy refinement) simultaneously — a strictly more constrained search than plain DesignFixedGreedy, since only the 2^{n/2} which-member-treated assignments consistent with the fixed pairing are considered, rather than all \binom{n}{n/2} balanced allocations.

Search algorithm. Pairing is computed by compute_binary_match_structure() exactly as in DesignFixedBinaryMatch (Mahalanobis or Euclidean distance per objective), lazily on first draw and cached in private$bms. Given the pairing, each replicate search initializes with a random coin flip per pair (which member starts treated), then in exhaustive mode (n_iter = Inf, default) repeatedly finds and applies the single pair-flip that most decreases the imbalance objective, stopping at a strict local optimum (or runs exactly n_iter random-pair stochastic flip-if-improving steps otherwise, with patience-based early stopping) — the same two search modes as DesignFixedGreedy, but with moves restricted to "flip which side of a given pair is treated" rather than "swap any treated/control pair of subjects." Random initialization and swap selection are seeded from R's own RNG (via greedy_design_search_cpp()'s per-thread seeding), so seed does govern reproducibility here.

Pair-preserving bootstrap. draw_bootstrap_indices() resamples whole matched pairs (via draw_matching_bootstrap_sample_cpp()) rather than individual subjects, since the greedy search only ever flips which member of a pair is treated (never crosses pairs), so w always has exactly one treated subject per pair — the pair, not the subject, is the exchangeable resampling unit.

Constraints. Only prob_T = 0.5 is supported (the constructor errors otherwise), and n must be divisible by 4 (draw_ws_raw() errors otherwise); n/2 matched pairs are formed regardless of parity, but the additional divisible-by-4 requirement is enforced by this class specifically (unlike DesignFixedBinaryMatch, which only requires even n).

Super classes

Design -> DesignFixed -> DesignFixedMatchingGreedyPairSwitching

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixedMatchingGreedyPairSwitching$new()

Initialize a fixed design that performs binary matching followed by greedy which-member-treated optimization (see class documentation). Only prob_T = 0.5 is supported, and n must be divisible by 4.

Usage
DesignFixedMatchingGreedyPairSwitching$new(
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n,
  verbose = FALSE,
  objective = "mahal_dist",
  n_iter = Inf,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

The data type of response values.

prob_T

The probability of treatment assignment. Must be 0.5.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size; must be divisible by 4.

verbose

A flag for verbosity.

objective

The covariate-imbalance objective to minimize when choosing which pair member is treated: either "mahal_dist" (default, squared Mahalanobis distance between treated/control means, also used as the matching distance) or "abs_sum_diff" (sum of absolute standardized mean differences); see class documentation for the exact criteria.

n_iter

Number of swap iterations. Inf (default) uses exhaustive best-improvement search guaranteed to reach a strict local optimum. A positive integer runs that many stochastic random-pair iterations with patience-based early stopping.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new DesignFixedMatchingGreedyPairSwitching object.


DesignFixedMatchingGreedyPairSwitching$supports_batch_w_pregeneration()

Returns TRUE so the calling framework pre-generates all replicate w vectors for a simulation cell in one batched call to greedy_design_search_cpp(), paying the one-time nbpMatching pairing cost once per cell (cached in private$bms) and reusing it across all replicates and the OpenMP-parallelized greedy searches, rather than recomputing the pairing per replicate.

Usage
DesignFixedMatchingGreedyPairSwitching$supports_batch_w_pregeneration()
Returns

Always TRUE for this class.


DesignFixedMatchingGreedyPairSwitching$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixedMatchingGreedyPairSwitching$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Krieger, A. M., Azriel, D., and Kapelner, A. (2019). "Nearly random designs with greatly improved balance." Biometrika, 106(3), 695-701, doi:10.1093/biomet/asz026; Greevy, R., Lu, B., Silber, J. H., and Rosenbaum, P. (2004). "Optimal multivariate matching before randomization." Biostatistics, 5(2), 263-275, doi:10.1093/biostatistics/5.2.263, for the matched-pair design this class refines.

Examples

des = DesignFixedMatchingGreedyPairSwitching$new(n = 10, response_type = 'continuous')

A Fixed, Deterministic Single-Allocation Optimal Design

Description

A fixed-sample-size DesignFixed that computes exactly one allocation w^* – the minimizer of a chosen covariate-imbalance or information objective over all allocations with n_T = \mathrm{round}(n \cdot \mathrm{prob}_T) treated subjects – by numerical optimization, rather than drawing from a restricted-randomization distribution the way DesignFixedGreedy/ DesignFixedGreedyDOptimal do.

The objective family and its argument mapping. Write Z_0 = [1\ X], P = Z_0 (Z_0^\top Z_0)^{-1} Z_0^\top, and s(w) = n_T - w^\top P w. The objectives map to constructor arguments and solved forms as follows:

"D" (determinant / D_M, D_s) and "A" with interest = "treatment"

maximize s(w), solved as the binary quadratic program \min_w w^\top P w – identical criteria and interest/prior_precision semantics to DesignFixedGreedyDOptimal (the same shared construction machinery is used, so the two classes optimize literally the same matrices).

"A" with interest = "all" or a covariate subset (A, A_s)

minimize (w^\top H w + 1)/s(w) with the sibling class's H/H_S, solved exactly via Dinkelbach's algorithm (Dinkelbach 1967) over product-linearized MILP subproblems.

Bayesian D_B/A_B

prior_precision = a scalar \tau (covariates only; intercept and treatment unpenalized) or a full matrix R_0, replacing (Z_0^\top Z_0)^{-1} with the ridge-regularized inverse – identical to the sibling class.

"mahal_dist" / "abs_sum_diff"

DesignFixedGreedy's covariate-imbalance criteria, definitionally identical (column-centered X, the same standardization and singular-covariance fallback), translated to exactly solvable forms: the Mahalanobis criterion is the pure quadratic w^\top Q w with Q = 4 X \Sigma^{-1} X^\top / n^2, and the absolute-sum criterion is an l1 objective solved by the standard linear MILP.

"custom"

a user-compiled black box under the user_compiled_fns.h calling convention (double f(const Eigen::MatrixXd& X, const Eigen::VectorXd& w), minimized), supplied via custom_objective; always solved by the annealing path (no structure to linearize).

Solvers and certificates. solver = "auto" (default) uses the exact "ompr" MILP path (optimum_certificate = "global", a certified global optimum) wherever tractable – always for "abs_sum_diff" (pure linear MILP); up to solver_args$linearization_max_n (default 20, set by a GLPK benchmark: the product linearization adds n(n-1)/2 auxiliaries and branch-and-bound cost climbs steeply past n \approx 20) for the quadratic and Dinkelbach criteria – and the native simulated-annealing solver beyond it, or always for "custom". The annealing solver is a formal method, not a heuristic: Metropolis acceptance over treated/control swaps with a configurable cooling schedule, for which Hajek (1988) proves convergence in probability to the global optimum under a slow-enough (logarithmic) schedule; the practical geometric schedule used by default is asymptotically motivated only, so its certificate is always "annealing_converged", never "global". solver = "ompr"/"annealing" force a path. Commercial backends extend the exact range via solver_args$roi_solver; see that parameter's wiring guides.

Inference. There is no usable randomization distribution conditional on the observed data (given X there is exactly one w^* up to the mirror coin), so permutation-style randomization tests/CIs are unavailable (supports_randomization_draw() is FALSE); the bootstrap randomization test IS available (the mechanism – "optimize this dataset" – is replayed on each resampled covariate matrix), as is all model-based and plain-resampling inference.

The mirror coin. At prob_T = 0.5, whenever the mirror 1 - w^* is a verified co-optimum (checked numerically by evaluating the objective, never by a symmetry table), a fair seeded coin picks between w^* and its mirror (mirror_coin = TRUE, the default). This restores exact treated/control label symmetry – and estimator unbiasedness – at zero cost to balance. A mirror that evaluates strictly better than the solver's answer raises an error (it would be a solver bug).

Super classes

Design -> DesignFixed -> DesignFixedOptimal

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixedOptimal$new()

Initialize a deterministic single-allocation optimal fixed experimental design. The optimization itself does not run until assign_w_to_all_subjects() (or draw_ws_according_to_design(r = 1)) is called.

Usage
DesignFixedOptimal$new(
  response_type,
  prob_T = 0.5,
  objective = "D",
  interest = "treatment",
  prior_precision = NULL,
  standardize_covariates = TRUE,
  custom_objective = NULL,
  solver = "auto",
  solver_args = list(),
  mirror_coin = TRUE,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

The data type of response values.

prob_T

Probability of treatment assignment, in (0, 1); the solve fixes the treated count at \mathrm{round}(n \cdot prob_T).

objective

"D" (default), "A", "mahal_dist", "abs_sum_diff", or "custom"; see the class documentation.

interest

For objective = "D"/"A" only: "treatment" (default), "all", a one-sided formula, a formula string, or model-matrix column names – identical semantics to DesignFixedGreedyDOptimal.

prior_precision

For objective = "D"/"A" only: NULL (default), a positive scalar \tau, or a symmetric prior-precision matrix R_0 – identical semantics to DesignFixedGreedyDOptimal.

standardize_covariates

If TRUE (default) and prior_precision is a scalar, covariates are standardized before the penalized criterion matrices are built. ("mahal_dist"/ "abs_sum_diff" standardize internally per their definitions regardless.)

custom_objective

Required iff objective = "custom" (and forbidden otherwise): an RcppXPtrUtils::cppXPtr() external pointer or a C++ source string under the user_compiled_fns.h calling convention (double f(const Eigen::MatrixXd& X, const Eigen::VectorXd& w); X is the design's model matrix, w a candidate 0/1 allocation; the returned value is minimized). A source string is compiled through the same cppXPtr mechanism and retained so parallel workers can recompile locally. A plain R function is not accepted, and cannot be: the annealing solver evaluates the objective once per candidate swap – typically thousands of times per chain, times n_chains, and again per BRT replicate – and an R-call round-trip on every one of those evaluations is orders of magnitude too slow to be usable, not merely slower. Since "custom" is always solved by annealing (never the MILP path), there is no lower-frequency code path where an R closure would be merely inconvenient; the restriction is a hard performance requirement of this objective's only execution path. See the class examples for a worked cppXPtr() construction. Save/ reload note: a raw cppXPtr() object does not survive saveRDS()/readRDS() (see Design's "Saving and loading" section); pass a C++ source string instead if this design needs to be reloadable.

solver

"auto" (default), "ompr", or "annealing".

solver_args

A named list of solver tuning arguments. Supported: roi_solver ("glpk"/"gurobi"/"cplex" – a closed set; arbitrary ROI plugin names are rejected), linearization_max_n, max_dinkelbach_iter, n_chains, max_iter, initial_temp, cooling_rate, and (consumed by the BRT replicate path) brt_max_iter, brt_n_chains, brt_solver.

Wiring up Gurobi (roi_solver = "gurobi"): (1) obtain a Gurobi license (free academic licenses are available) and install the Gurobi Optimizer itself – this sets up GUROBI_HOME and the license file, entirely outside this package's control; (2) install Gurobi's own R package, which is not on CRAN – it ships inside the Gurobi installation: R CMD INSTALL "$GUROBI_HOME/R/gurobi_<version>_R_<Rmajor.minor>.tar.gz" (exact filename depends on your Gurobi version and platform); (3) install.packages("ROI.plugin.gurobi") from CRAN; (4) verify "gurobi" %in% ROI::ROI_registered_solvers() after loading the plugin; (5) pass solver_args = list(roi_solver = "gurobi").

Wiring up CPLEX (roi_solver = "cplex"): (1) obtain an IBM CPLEX license (free academic licenses are available) and install IBM ILOG CPLEX Optimization Studio; (2) install Rcplex (CRAN) – unlike the Gurobi bridge, it compiles from source against your local CPLEX SDK and must be pointed at your CPLEX version's include/lib directories at install time; follow Rcplex's own INSTALL instructions for your CPLEX version rather than a fixed command, since the flags change across CPLEX releases; (3) install.packages("ROI.plugin.cplex") from CRAN; (4) verify "cplex" %in% ROI::ROI_registered_solvers(); (5) pass solver_args = list(roi_solver = "cplex").

ROI.plugin.gurobi/ROI.plugin.cplex/Rcplex are deliberately never listed in this package's Suggests: declaring them would misrepresent the dependency as something install.packages() could satisfy, when the vendor installation/license underneath cannot be. Availability is checked lazily at solve time; if the plugin loads but the solve fails, the likely cause is a missing vendor installation or license.

mirror_coin

If TRUE (default), flip a fair seeded coin between w^* and a verified co-optimal mirror 1 - w^* after every solve (only possible at prob_T = 0.5); see the class documentation.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

Sample size (if fixed).

verbose

Flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility (consumed by the annealing solver and the mirror coin; MILP solves are deterministic up to the seeded label flip).

Returns

A new 'DesignFixedOptimal' object


DesignFixedOptimal$get_objective()

The objective this design was constructed with.

Usage
DesignFixedOptimal$get_objective()
Returns

One of "D", "A", "mahal_dist", "abs_sum_diff", "custom".


DesignFixedOptimal$get_interest()

The parameter-interest setting ("D"/"A" only).

Usage
DesignFixedOptimal$get_interest()
Returns

The interest construction argument.


DesignFixedOptimal$get_prior_precision()

The Bayesian prior precision this design was constructed with.

Usage
DesignFixedOptimal$get_prior_precision()
Returns

NULL, a positive scalar, or a symmetric matrix.


DesignFixedOptimal$get_solver()

The solver setting this design was constructed with.

Usage
DesignFixedOptimal$get_solver()
Returns

"auto", "ompr", or "annealing".


DesignFixedOptimal$get_mirror_coin()

The mirror-coin setting this design was constructed with.

Usage
DesignFixedOptimal$get_mirror_coin()
Returns

TRUE or FALSE.


DesignFixedOptimal$get_optimization_diagnostics()

Diagnostics cached by the most recent solve: the solver used, optimum_certificate ("global" for exact "ompr" solves, "annealing_converged" otherwise), the achieved objective value, mirror-coin outcome (mirror_feasible/mirror_tied/mirror_flipped), elapsed time, and the solver's own detail fields.

Usage
DesignFixedOptimal$get_optimization_diagnostics()
Returns

A named list, or NULL if no solve has run yet.


DesignFixedOptimal$supports_randomization_draw()

Characterization: FALSE – given the observed data there is exactly one w^* (up to the vacuous 2-atom mirror pair), so there is no randomization distribution to draw from and permutation-style randomization tests/CIs are unavailable. The bootstrap randomization test remains available via supports_resampling_replay() (the deterministic mechanism is replayed on each resample).

Usage
DesignFixedOptimal$supports_randomization_draw()
Returns

Always FALSE for this class.


DesignFixedOptimal$prepare_for_resampling_replay()

BRT replicate-mode switch (called by the bootstrap-randomization-test machinery ahead of each replayed draw; see Design$prepare_for_resampling_replay()). Subsequent solves use the per-replicate solver profile: solver_args$brt_solver (default "annealing" with the reduced brt_max_iter/brt_n_chains schedule – replicate assignments need to be faithful applications of the mechanism, not individually re-verified to the observed solve's convergence standard; "ompr" buys exact per-replicate solves at the user's expense). Idempotent; the mirror coin still applies per replicate (the BRT replays the coin-inclusive mechanism).

Usage
DesignFixedOptimal$prepare_for_resampling_replay()
Returns

invisible(NULL).


DesignFixedOptimal$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixedOptimal$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Dinkelbach, W. (1967). On nonlinear fractional programming. Management Science 13(7):492-498, for the exact A-optimality reduction. Hajek, B. (1988). Cooling schedules for optimal annealing. Mathematics of Operations Research 13(2):311-329, for the annealing solver's formal convergence property. Atkinson, A. C., Donev, A. N., and Tobias, R. D. (2007). Optimum Experimental Designs, with SAS. Oxford University Press, for the D-/A-optimality criteria.

Examples

# objective = "mahal_dist" (the default MILP path) needs a MILP solver --
# ompr + ompr.roi + a ROI plugin (ROI.plugin.glpk by default), all Suggests,
# not installed automatically with EDI. Confirmed 2026-09-25: ungated, this
# example errored outright on R CMD check's --no-suggests leg (and would do
# the same for any real user without the optional MILP-solver stack), not
# just a CI-specific issue.
if (requireNamespace("ompr", quietly = TRUE) &&
    requireNamespace("ompr.roi", quietly = TRUE) &&
    requireNamespace("ROI.plugin.glpk", quietly = TRUE)) {
  des = DesignFixedOptimal$new(n = 14, response_type = 'continuous', objective = "mahal_dist")
  des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(14)))
  des$assign_w_to_all_subjects()
  des$get_optimization_diagnostics()
}


# A custom compiled objective (the user_compiled_fns.h calling convention),
# built with RcppXPtrUtils::cppXPtr() -- here, squared imbalance of the
# centered covariate sums. Compiling it needs a C++ toolchain and takes
# several seconds:
if (requireNamespace("RcppXPtrUtils", quietly = TRUE)) {
  fobj = RcppXPtrUtils::cppXPtr(
    "double f(const Eigen::MatrixXd& X, const Eigen::VectorXd& w) {
      Eigen::RowVectorXd mu = X.colwise().mean();
      Eigen::MatrixXd Xc = X.rowwise() - mu;
      Eigen::VectorXd s = 2.0 * w - Eigen::VectorXd::Ones(X.rows());
      return (Xc.transpose() * s).squaredNorm();
    }", depends = "RcppEigen")
  des2 = DesignFixedOptimal$new(n = 14, response_type = 'continuous',
    objective = "custom", custom_objective = fobj)
  des2$add_all_subjects_to_experiment(data.frame(x1 = rnorm(14)))
  des2$assign_w_to_all_subjects()
}


A Fixed, Covariate-Homogeneous-Block Randomized Design

Description

A fixed-sample-size DesignFixed that first partitions subjects into B approximately equal-sized, covariate-homogeneous blocks by (approximately or exactly) minimizing total within-block pairwise covariate distance \sum_{k=1}^{B} \sum_{i, j \in \text{block } k, i < j} D(x_i, x_j), then randomizes treatment independently within each resulting block at probability prob_T (via block_ra, the same mechanism as DesignFixedBlocking). Unlike DesignFixedBlocking, which forms blocks from user-specified column-wise strata (categorical levels / quantile bins), here blocks are formed directly from a multivariate distance/clustering criterion over all covariates jointly, so this design generalizes matched-pair designs (DesignFixedBinaryMatch) from block size 2 to arbitrary block size B. When B is omitted and n is known at initialization, the default is floor(sqrt(n)), truncated below at 1 (the usual heuristic block count for balancing within-block homogeneity against within-block sample size).

Block-formation algorithms. method selects among three block construction strategies with different exactness/scalability trade-offs (see the new() argument documentation for details): "K-way" (default, balanced k-means-style anticlustering via anticlust, fast and empirically close to optimal), "greedy" (nearest-neighbor greedy matching via blockTools, fast even for large n), and "ompr" (an exact mixed-integer program solved with GLPK via ompr/ompr.roi, globally optimal but scaling as O(n^2 B) in the number of decision variables — practical only for small n). Block membership is computed lazily (on first draw, via get_or_compute_block_ids()) and cached in private$block_ids for reuse across replicates.

Distance specification ("ompr" only). dist selects the pairwise distance D(x_i, x_j) the exact solver minimizes: "euclidean", "sum_abs_diff" (sum of absolute coordinate differences), "mahal" (default, Mahalanobis distance accounting for covariate correlation/scale), or a user-supplied distance function. The "K-way" and "greedy" methods use their own respective packages' built-in distance conventions and do not consult dist.

No-covariate fallback. If the covariate matrix has zero columns, block membership is assigned by simple round-robin (rep(seq_len(B), length.out = n)) rather than by any of the three clustering algorithms, since there is no covariate information to cluster on.

Solver backend (method = "ompr" only). roi_solver selects the MILP backend ompr.roi dispatches to, a closed set c("glpk", "gurobi", "cplex") (default "glpk") – validated against this set, not passed through to ROI::ROI_registered_solvers() unchecked, so a typo or an unsupported solver name fails fast with a clear message rather than an opaque ompr error three layers down. Gurobi and CPLEX are supported because they extend which block-formation problems stay practical, not just which are expressible: GLPK's branch-and-bound is single-threaded with no commercial-grade presolve/cutting-plane machinery, while Gurobi/CPLEX are typically an order of magnitude faster on the same MILP and solve in parallel, meaningfully extending the n for which the exact "ompr" method stays practical. Scope is deliberately closed to these two for now – not because other ROI plugins (CBC, SYMPHONY, ...) wouldn't work mechanically (the dispatch is generic), but because Gurobi and CPLEX are the two most widely used commercial solvers and the only ones worth a maintained step-by-step guide at this point; extending the closed set is a small, low-risk addition later if a real need for a third backend appears, not a reason to leave the set open-ended now.

Wiring up Gurobi:

  1. Obtain a Gurobi license (a free academic license is available from Gurobi for non-commercial use) and install the Gurobi Optimizer itself. This sets up GUROBI_HOME and the license file (gurobi.lic, discoverable via the GRB_LICENSE_FILE environment variable or Gurobi's default search path) – entirely outside this package's control or dependency graph.

  2. Install Gurobi's own R package. Not available via CRAN – it ships inside the Gurobi installation itself: R CMD INSTALL "$GUROBI_HOME/R/gurobi_<version>_R_<Rmajor.minor>.tar.gz" (exact filename/path depends on your Gurobi version and platform; see the R/ subdirectory of your Gurobi install). This is the vendor interface ROI.plugin.gurobi wraps – required even though it's not what you call directly.

  3. Install the ROI bridge package from CRAN: install.packages("ROI.plugin.gurobi"). This package is on CRAN (it only depends on ROI + the gurobi R package from step 2 being present at load time) and is the only new artifact this class's own dependency graph ever touches.

  4. Verify: after library(ROI.plugin.gurobi), "gurobi" %in% ROI::ROI_registered_solvers() should be TRUE.

  5. Pass roi_solver = "gurobi" to the constructor.

Wiring up CPLEX:

  1. Obtain an IBM CPLEX license (a free academic license is available from IBM) and install IBM ILOG CPLEX Optimization Studio.

  2. Install Rcplex (CRAN), CPLEX's R interface. Unlike ROI.plugin.gurobi, Rcplex is a source package that compiles against your local CPLEX installation – it needs to be pointed at your CPLEX SDK's include/lib directories at install time (typically via configure.args to install.packages(), naming your CPLEX version's cplex/include/cplex/lib/<platform> paths). The exact flag names and paths are CPLEX-version- and platform-specific – follow Rcplex's own INSTALL/README instructions for your installed CPLEX version rather than a fixed command copied from here, since this changes across CPLEX releases.

  3. Install the ROI bridge package from CRAN: install.packages("ROI.plugin.cplex") (depends on Rcplex from step 2 being present and working).

  4. Verify: after library(ROI.plugin.cplex), "cplex" %in% ROI::ROI_registered_solvers() should be TRUE.

  5. Pass roi_solver = "cplex" to the constructor.

Dependency-graph consequence: ROI.plugin.gurobi/ROI.plugin.cplex (and Rcplex) are never added to Suggests – they're free/CRAN- available themselves, but declaring them would misrepresent the dependency as something install.packages("EDI", dependencies = TRUE) could satisfy, when the vendor package/license underneath cannot be. The lazy-check pattern already used for ompr/ompr.roi/ROI.plugin.glpk extends naturally: check requireNamespace("ROI.plugin.gurobi"/"ROI.plugin.cplex") at solve time (not at package load or class-definition time) for whichever roi_solver was requested, and error informatively – naming the missing package and, if that's present but the solve still fails, noting the likely cause is a missing vendor license/installation, not something this class can diagnose further.

Bootstrap. draw_bootstrap_indices() resamples within blocks by default (bootstrap_type = "within_blocks" or NULL, via stratified_bootstrap_indices_cpp()) or resamples whole blocks otherwise (via resample_group_rows_cpp()), mirroring the block structure in the resampling scheme, as in DesignFixedBlocking.

Super classes

Design -> DesignFixed -> DesignFixedOptimalBlocks

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixedOptimalBlocks$supports_batch_w_pregeneration()

Returns TRUE so the calling framework pre-generates all replicate w vectors for a simulation cell in one batch, paying the one-time block-formation cost (K-way anticlustering, greedy matching, or the exact ompr/GLPK solve) once per cell and reusing the resulting block assignment across replicates, rather than recomputing it per replicate.

Usage
DesignFixedOptimalBlocks$supports_batch_w_pregeneration()
Returns

Always TRUE for this class.


DesignFixedOptimalBlocks$new()

Initialize a fixed optimal-blocks design. Block formation itself is deferred until the first draw (see class documentation); this constructor only validates and records configuration, including checking that B (or its floor(sqrt(n)) default) admits a feasible block-size partition of n when n is already known.

Usage
DesignFixedOptimalBlocks$new(
  B = NULL,
  method = "K-way",
  dist = "mahal",
  roi_solver = "glpk",
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
B

Number of blocks to form. If omitted and n is supplied, defaults to floor(sqrt(n)), with a minimum of 1.

method

Algorithm used to partition subjects into blocks.

"K-way" (default)

Balanced k-means anticlustering via anticlust::balanced_clustering. Requires the anticlust package. Produces well-spread blocks and is significantly faster than "greedy" (e.g., ~10x faster for n=200, p=10, B=10) while achieving a better within-block distance objective (e.g., ~4% lower).

"greedy"

Greedy nearest-neighbour matching via blockTools::block. Requires the blockTools package. Fast even for large n.

"ompr"

Exact mixed-integer programme solved with GLPK via ompr. Globally optimal but scales as O(n^2 B) in variables and is only practical for small n.

dist

Distance specification used only when method = "ompr". Either a function or one of "euclidean", "sum_abs_diff", or "mahal". Default is "mahal".

roi_solver

MILP backend used only when method = "ompr". A closed set c("glpk", "gurobi", "cplex"), default "glpk". See the class documentation's "Solver backend" and wiring-guide sections for the Gurobi/CPLEX setup steps.

response_type

The response type for the design.

prob_T

Treatment assignment probability within each block.

include_is_missing_as_a_new_feature

Whether to include missingness indicators.

n

Planned sample size.

verbose

Whether to print progress messages.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new DesignFixedOptimalBlocks object.


DesignFixedOptimalBlocks$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixedOptimalBlocks$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Higgins, M. J., Sävje, F., and Sekhon, J. S. (2016). "Improving massive experiments with threshold blocking." Proceedings of the National Academy of Sciences, 113(27), 7369-7376, doi:10.1073/pnas.1510504113, for optimal/near-optimal covariate-based blocking prior to randomization. See also randomized block design for orientation and Mahalanobis distance for the default "ompr" distance.

Examples

des = DesignFixedOptimalBlocks$new(n = 9, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x = rnorm(9)))
des$assign_w_to_all_subjects()

A Fixed Rerandomization Design (Rejection-Sampled on Covariate Balance)

Description

A fixed-sample-size DesignFixed implementing rerandomization (Morgan and Rubin, 2012): candidate allocations w are drawn from the design's base randomization law (balanced complete randomization when prob_T = 0.5, i.i.d. \mathrm{Bernoulli}(prob\_T) draws otherwise — see generate_one_rerandomized_w()) and only accepted if a covariate imbalance criterion M(w) between the treated and control groups falls below a threshold a (obj_val_cutoff), i.e. the accepted allocations are drawn from the base randomization distribution truncated to \{w : M(w) \le a\}. Unlike DesignFixedGreedy/ DesignFixedGreedyDOptimal, which search for a single well-balanced allocation, rerandomization instead filters the design's own randomization distribution, which is what makes it directly compatible with Fisherian/randomization-based inference (the accepted allocations remain a well-defined — if truncated — randomization distribution, valid for randomization tests/CIs restricted to that truncated support).

Two mutually exclusive acceptance modes are supported (specifying both errors): obj_val_cutoff accepts/rejects each draw against a fixed threshold a; prop_acceptable instead draws r / prop\_acceptable candidates and keeps the r with the lowest M(w) (an empirical top-quantile acceptance region, equivalent in the large-draw limit to an implied cutoff at the prop_acceptable quantile of M(w)'s distribution under the base randomization law).

Objective M(w). objective = "mahal_dist" (default) uses the squared Mahalanobis distance between treated and control covariate means, M(w) = (\bar X_T - \bar X_C)^\top S^{-1} (\bar X_T - \bar X_C), where S is the sample covariance of all covariates (ridge-regularized by 10^{-6}I if |\det S| < 10^{-10}), computed once and cached in private$S_inv. objective = "abs_sum_diff" instead uses the sum of absolute mean differences, M(w) = \sum_j |\bar X_{T,j} - \bar X_{C,j}|, with no correlation adjustment. Only these two objectives are supported; any other value errors at draw time.

Fast path (native C++). When prob_T = 0.5 and n is even, candidate generation and filtering run via a parallel C++ rejection sampler (rerandomization_search_cpp()), which internally works with a rescaled objective f_{\mathrm{cpp}}: f_{\mathrm{cpp}} = M(w)/4 for "mahal_dist" and f_{\mathrm{cpp}} = M(w)/2 (on a GED-standardized scale) for "abs_sum_diff"; the user-facing obj_val_cutoff is converted to this internal scale before being passed to C++, so the accepted-allocation semantics are unaffected, but this rescaling is a backend implementation detail worth knowing when comparing C++-path and pure-R-path acceptance rates for the same nominal cutoff. The sampler draws up to max(r * 1000, 100000) candidates internally; if fewer than r allocations are accepted within that budget (an overly tight cutoff), this errors naming how many were actually found – loosen obj_val_cutoff or use prop_acceptable instead. (Earlier versions silently recycled the accepted set to pad out to r, duplicating some draws; fixed, since that meant some "independent" replicates were literal duplicates of an accepted allocation.)

Seed reproducibility and multi-core parallelism. The C++ fast path's rejection sampler is a genuine work-stealing search: with more than one core (set_num_cores/a fork cluster/mirai daemons; the package default is a single core), threads race via atomic operations for both which candidate draws to try next and which output column an accepted draw claims, so which per-thread-seeded RNG stream ends up producing a given replicate – and in what order – depends on real-time OS scheduling, not just seed. With the default single core, draws are exactly seed-reproducible; this is not guaranteed once more than one core is in use. Contrast with DesignFixedGreedy/ DesignFixedBinaryMatch, whose C++ kernels use static (not work-stealing) thread scheduling and remain seed-reproducible regardless of core count.

prop_acceptable path. Uses complete_randomization_forced_balanced_cpp() (balanced case) or complete_randomization_imbalanced_cpp() (prob_T != 0.5) to draw n_{\mathrm{draw}} = \mathrm{round}(r / prop\_acceptable) candidate allocations in one batched call, computes M(w) for all of them via compute_objective_vals_cpp(), and keeps the r with smallest M(w).

Pure-R fallback (unbalanced or odd n, obj_val_cutoff mode only). Draws one candidate at a time via generate_one_rerandomized_w() in an unbounded repeat loop that accepts the first candidate with M(w) \le a. Unlike the C++ fast path, this fallback has no draw-count safety limit: if obj_val_cutoff is set tight enough that acceptance probability under the base randomization law is extremely small for this n/ covariate structure, this loop can run for a very long time (in principle indefinitely) before finding an acceptable draw.

Super classes

Design -> DesignFixed -> DesignFixedRerandomization

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixedRerandomization$new()

Initialize a rerandomization fixed experimental design. Exactly one of obj_val_cutoff/prop_acceptable may be specified (or neither, which accepts every candidate, i.e. no filtering); supplying both raises an error. See class documentation for the exact acceptance semantics of each mode and the covariate-imbalance objective.

Usage
DesignFixedRerandomization$new(
  response_type,
  prob_T = 0.5,
  obj_val_cutoff = NULL,
  prop_acceptable = NULL,
  objective = "mahal_dist",
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

The data type of response values.

prob_T

The probability of the treatment assignment.

obj_val_cutoff

The maximum allowable objective value a; a candidate allocation is accepted iff M(w) \le a. Cannot be specified together with prop_acceptable.

prop_acceptable

The proportion of randomizations to accept (draws r/prop_acceptable total, returns r lowest). Cannot be specified together with obj_val_cutoff.

objective

The covariate-imbalance objective M(w) to filter on: either "mahal_dist" (default, squared Mahalanobis distance) or "abs_sum_diff" (sum of absolute mean differences); see class documentation for the exact formulas.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

Flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignFixedRerandomization' object


DesignFixedRerandomization$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixedRerandomization$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Morgan, K. L., and Rubin, D. B. (2012). "Rerandomization to improve covariate balance in experiments." The Annals of Statistics, 40(2), 1263-1282, doi:10.1214/12-AOS1008, for the rerandomization framework and its randomization-inference validity. See also Mahalanobis distance for the default objective.

Examples

des = DesignFixedRerandomization$new(n = 10, response_type = 'continuous')

A Fixed, Individually Balanced Completely Randomized Design (iBCRD)

Description

A fixed-sample-size DesignFixed implementing the individually balanced complete randomized design (iBCRD): the number of treated subjects is fixed at exactly n_T = \mathrm{round}(n \cdot prob\_T), and each allocation w with exactly n_T ones is drawn uniformly at random from the \binom{n}{n_T} possible such allocations (via a Fisher-Yates shuffle of a base vector with n_T ones and n - n_T zeros). This is the classical "complete randomization" reference design of randomization inference: unlike DesignFixedBernoulli (independent per-subject coin flips, random n_T), n_T is fixed here, which is what makes exact permutation/randomization tests over the \binom{n}{n_T} allocations well-defined; unlike DesignFixedGreedyDOptimal/ DesignFixedGreedy, no covariate information is used to select among those allocations — every one of the \binom{n}{n_T} allocations is equally likely.

Draw mechanism. draw_ws_raw(r) delegates to generate_permutations_ibcrd_cpp(), which builds one base allocation vector (n_T ones followed by n - n_T zeros) and independently shuffles (Fisher-Yates via std::shuffle) a fresh copy of it per replicate column, seeded from R's own RNG stream (so seed does govern reproducibility here, unlike the A-/D-optimal exchange searches). assign_w_to_all_subjects() draws one such allocation (r = 1) and applies it to all subjects.

Single implicit block. The constructor sets private$m to a constant vector of 1s (a single block containing every subject) once n is known, so that shared blocking/matching machinery that expects a block-membership vector treats the whole sample as one block by default.

Super classes

Design -> DesignFixed -> DesignFixediBCRD

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

DesignFixediBCRD$new()

Initialize a fixed individually balanced completely randomized experimental design (see class documentation for the exact randomization law).

Usage
DesignFixediBCRD$new(
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

"continuous", "incidence", "proportion", "count", "survival", or "ordinal".

prob_T

Target probability of treatment assignment; the realized number of treated subjects is fixed at round(n * prob_T) for every draw (unlike DesignFixedBernoulli, where it is random).

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

A flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignFixediBCRD' object


DesignFixediBCRD$clone()

The objects of this class are cloneable with this method.

Usage
DesignFixediBCRD$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Fisher, R. A. (1935). The Design of Experiments. Oliver and Boyd, for complete randomization as the canonical reference design of randomization inference. See also randomized experiment for orientation on complete vs. Bernoulli randomization.

Examples

des = DesignFixediBCRD$new(n = 10, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(10)))
des$assign_w_to_all_subjects()

Sequential One-by-One Experimental Design

Description

Abstract R6 class encapsulating data and functionality for a sequential one-by- one experimental design.

Sample size and stopping

Subjects are assigned one at a time, but this class does not implement interim outcome monitoring or an outcome-dependent stopping rule. For the usual fixed-sample analysis, specify the target sample size n before enrollment and stop after exactly n subjects; the caller is responsible for ending enrollment at that point. With n = NULL, the class leaves the final sample size unspecified and does not determine when enrollment ends. Inference methods that assume a fixed sample size require the final size to be chosen independently of accumulating outcomes.

Super class

Design -> DesignSeqOneByOne

Methods

Public methods

+ inherited public methods from Design

DesignSeqOneByOne$new()

Initialize a sequential one-by-one design.

Usage
DesignSeqOneByOne$new(
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL,
  ...
)
Arguments
response_type

The data type of response values.

prob_T

The probability of the treatment assignment.

include_is_missing_as_a_new_feature

If missing data is present, include a dummy variable for it.

n

The prespecified target sample size for fixed-sample analysis. If NULL, the final sample size is left to the caller; the class does not provide a stopping rule.

verbose

Whether to print progress messages.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

...

Extra arguments passed to the Design superclass.


DesignSeqOneByOne$add_one_subject()

Add subject-specific measurements for the next subject entrant.

Usage
DesignSeqOneByOne$add_one_subject(x_new, allow_new_cols = TRUE)
Arguments
x_new

A data frame with one row representing the new subject's covariates.

allow_new_cols

Allow new features in the new subject's covariates.


DesignSeqOneByOne$add_one_subject_to_experiment_and_assign()

Adds a subject and assigns treatment.

Usage
DesignSeqOneByOne$add_one_subject_to_experiment_and_assign(x_new)
Arguments
x_new

A data frame with one row representing the new subject's covariates.

Returns

The treatment assignment as {0,1} (1 = treated, 0 = control).


DesignSeqOneByOne$assign_wt()

Assigns treatment to the current subject.

Usage
DesignSeqOneByOne$assign_wt()
Returns

The treatment assignment (0 or 1).


DesignSeqOneByOne$print_current_subject_assignment()

Prints the current subject's assignment.

Usage
DesignSeqOneByOne$print_current_subject_assignment()

DesignSeqOneByOne$clone()

The objects of this class are cloneable with this method.

Usage
DesignSeqOneByOne$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

# DesignSeqOneByOne is abstract and cannot be instantiated directly;
# construct a concrete subclass instead, e.g.:
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))

Atkinson's (1982) Covariate-Adjusted Biased Coin Sequential Design

Description

A DesignSeqOneByOne that assigns each newly arriving subject's treatment via Atkinson's (1982) D_A-optimum biased coin: a coin whose treatment probability is skewed away from prob_T toward whichever assignment would most improve the current design's efficiency for estimating the treatment effect D_A-optimally, given the covariates observed so far. Compared to a fixed-probability coin (e.g. DesignSeqOneByOneBernoulli), this improves covariate balance/estimation efficiency online without ever fully determinizing the assignment (the coin is always strictly between 0 and 1, so randomization-based inference remains valid), at the cost of requiring a numerically well-conditioned design matrix to compute the bias.

Assignment rule. For subject t, let Z_{t-1} = [w_{1:t-1}, 1, X_{1:t-1}] be the (treatment, intercept, covariates) design matrix accumulated from the first t-1 subjects, and let M = (t-1)(Z_{t-1}^\top Z_{t-1})^{-1}. Writing x_t for the new subject's covariate vector (with a leading 1 for the intercept) and A = M_{[1, 2:]} \cdot x_t (the treatment row of M, projected onto x_t), the treatment probability is

\pi_t = \frac{\big(M_{11}/A + 1\big)^2}{\big(M_{11}/A + 1\big)^2 + 1},

clamped to [0, 1], and subject t is assigned to treatment with probability \pi_t. This is Atkinson's biased-coin formula for D_A-optimal sequential design: the coin biases toward the assignment that would most reduce the variance of the treatment-effect estimate under the linear model implied by Z_t, converging toward more extreme (but never fully deterministic) probabilities as the current covariate imbalance grows in directions that matter for that estimate.

Fallback to a fair(-ish) coin. For the first ncol(private$Xraw) + 3 subjects (too few observations for Z_{t-1}^\top Z_{t-1} to be reliably invertible), and whenever the C++ computation encounters a non-invertible design matrix, a non-finite bias term, or any other numerical failure (caught via tryCatch()), assignment falls back to an unbiased \mathrm{Bernoulli}(prob\_T) draw instead of Atkinson's rule.

Reproducibility. The per-subject C++ draw (atkinson_assign_weight_cpp()) seeds its own generator from R's RNG stream per call, so seed governs reproducibility of the resulting assignment sequence in the usual way.

Super classes

Design -> DesignSeqOneByOne -> DesignSeqOneByOneAtkinson

Methods

Public methods

+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design

DesignSeqOneByOneAtkinson$new()

Initialize an Atkinson (1982) biased-coin sequential experimental design (see class documentation for the assignment rule).

Usage
DesignSeqOneByOneAtkinson$new(
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

The data type of response values.

prob_T

The nominal probability of treatment assignment; used as the fallback coin probability early in the trial and whenever Atkinson's rule cannot be computed (see class documentation), and as the reference probability the biased coin is skewed away from otherwise.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

A flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignSeqOneByOneAtkinson' object


DesignSeqOneByOneAtkinson$assign_wt()

Draw the next subject's treatment assignment via Atkinson's (1982) D_A-optimum biased coin (see class documentation for the exact probability formula), falling back to an unbiased \mathrm{Bernoulli}(prob\_T) draw early in the trial or on numerical failure of the biased-coin computation.

Usage
DesignSeqOneByOneAtkinson$assign_wt()
Returns

The treatment assignment (0 or 1) for the next subject.


DesignSeqOneByOneAtkinson$clone()

The objects of this class are cloneable with this method.

Usage
DesignSeqOneByOneAtkinson$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Atkinson, A. C. (1982). "Optimum biased coin designs for sequential clinical trials with prognostic factors." Biometrika, 69(1), 61-67, doi:10.1093/biomet/69.1.61. See also randomized experiment for orientation on biased-coin sequential designs.

Examples

seq_des = DesignSeqOneByOneAtkinson$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))

A Sequential Bernoulli (Independent-Coin-Flip) Randomized Design

Description

A DesignSeqOneByOne in which each arriving subject's treatment assignment is drawn independently as w_t \stackrel{iid}{\sim} \mathrm{Bernoulli}(prob\_T), with no dependence on covariates or on prior assignments — the direct sequential-enrollment analog of DesignFixedBernoulli. As in the fixed-sample version, the realized number of treated subjects after t arrivals is random (\mathrm{Binomial}(t, prob\_T)), in contrast to sequential designs that actively balance assignment counts or covariates (e.g. DesignSeqOneByOneAtkinson).

Nonparametric bootstrap

The ordinary row bootstrap for this design is supported under a prespecified fixed sample size: choose n before enrollment and stop after exactly n subjects, without using interim outcomes to decide when to stop. EDI does not implement sequential monitoring or check that this stopping condition was followed. In particular, n = NULL does not satisfy the documented fixed-sample justification, even though the bootstrap method is not blocked at runtime. The usual assumptions of iid subjects and potential outcomes, and a regular estimator, also apply.

Super classes

Design -> DesignSeqOneByOne -> DesignSeqOneByOneBernoulli

Methods

Public methods

+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design

DesignSeqOneByOneBernoulli$is_a_bernoulli_capable()

Characterization: this design draws each subject's treatment assignment as an independent \mathrm{Bernoulli}(prob\_T) coin flip (see class documentation), so it is Bernoulli-capable by construction.

Usage
DesignSeqOneByOneBernoulli$is_a_bernoulli_capable()
Returns

Always TRUE for this class.


DesignSeqOneByOneBernoulli$new()

Initialize a Bernoulli (independent-coin-flip) sequential experimental design.

Usage
DesignSeqOneByOneBernoulli$new(
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

The data type of response values which must be one of the following: "continuous", "incidence", "proportion", "count", "survival", "ordinal".

prob_T

The probability of the treatment assignment. This defaults to 0.5.

include_is_missing_as_a_new_feature

If missing data is present in a variable, should we include another dummy variable for its missingness? The default is TRUE.

n

The prespecified sample size for fixed-sample inference. Default is NULL; the nonparametric bootstrap justification above requires a fixed n.

verbose

A flag indicating whether messages should be displayed.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignSeqOneByOneBernoulli' object


DesignSeqOneByOneBernoulli$assign_wt()

Draw the next subject's treatment assignment as a single independent \mathrm{Bernoulli}(prob\_T) coin flip (see class documentation); does not consult covariates or prior assignments.

Usage
DesignSeqOneByOneBernoulli$assign_wt()
Returns

The treatment assignment (0 or 1) for the next subject.


DesignSeqOneByOneBernoulli$clone()

The objects of this class are cloneable with this method.

Usage
DesignSeqOneByOneBernoulli$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))

Efron's (1971) Biased Coin Sequential Design

Description

A DesignSeqOneByOne implementing Efron's (1971) biased coin: no covariates are used, only the running counts of treated (n_T) and control (n_C) subjects assigned so far. If the counts are currently equal, the next subject is assigned by a fair \mathrm{Bernoulli}(0.5) coin; otherwise, the next subject is assigned to the currently under-represented group with probability weighted_coin_prob (> 0.5, e.g. the classical 2/3) and to the over-represented group with probability 1 - weighted_coin_prob. This keeps the running treatment/control counts close to balanced throughout enrollment (unlike DesignSeqOneByOneBernoulli, whose running counts can drift arbitrarily far from balanced) while remaining strictly randomized at every step (the coin is always strictly between 1 - weighted_coin_prob and weighted_coin_prob, never fully deterministic), unlike a purely deterministic alternating allocation. This is a count-balancing design only — it does not use covariates at all, in contrast to DesignSeqOneByOneAtkinson/ DesignSeqOneByOneKK21, which bias the coin toward covariate balance rather than (or in addition to) count balance.

Super classes

Design -> DesignSeqOneByOne -> DesignSeqOneByOneEfron

Methods

Public methods

+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design

DesignSeqOneByOneEfron$new()

Initialize an Efron (1971) biased coin sequential experimental design (see class documentation for the exact assignment rule).

Usage
DesignSeqOneByOneEfron$new(
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  weighted_coin_prob = 2/3,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

"continuous", "incidence", "proportion", "count", "survival", or "ordinal".

prob_T

Nominal probability of treatment assignment; used only as the fair-coin probability when the running treated/control counts are exactly equal (see assign_wt()).

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

A flag for verbosity.

weighted_coin_prob

The probability (> 0.5) of assigning the next subject to whichever of treatment/control currently has fewer subjects, when the running counts are unequal. Default 2/3, the value from Efron (1971).

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignSeqOneByOneEfron' object


DesignSeqOneByOneEfron$assign_wt()

Draw the next subject's treatment assignment via Efron's (1971) biased coin (see class documentation): a fair coin if the running treated/control counts are equal, otherwise a coin biased toward the currently under-represented group at probability weighted_coin_prob.

Usage
DesignSeqOneByOneEfron$assign_wt()
Returns

The treatment assignment (0 or 1) for the next subject.


DesignSeqOneByOneEfron$clone()

The objects of this class are cloneable with this method.

Usage
DesignSeqOneByOneEfron$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Efron, B. (1971). "Forcing a sequential experiment to be balanced." Biometrika, 58(3), 403-417, doi:10.1093/biomet/58.3.403. See also randomized experiment for orientation on biased-coin sequential designs.

Examples

seq_des = DesignSeqOneByOneEfron$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))

Kapelner and Krieger's (2014) Sequential "Matching-on-the-Fly" Design

Description

A DesignSeqOneByOne that matches subjects to each other as they arrive, rather than requiring all subjects up front (as DesignFixedBinaryMatch does): each new subject is either matched to its nearest (Mahalanobis-distance) unmatched prior subject in the "reservoir" — if that distance is small enough to pass a statistical closeness test — and assigned the opposite treatment from its match, or, if no sufficiently close match exists (or matching hasn't started yet), randomized and added to the reservoir for future subjects to potentially match against.

Burn-in. For the first t_0_pct * n subjects (default 35%), or whenever no covariates are yet available, subjects are simply randomized (\mathrm{Bernoulli}(prob\_T)) and placed in the reservoir (private$m set to 0 for that subject) — matching does not begin until enough subjects have accumulated to estimate a stable covariate covariance structure.

Matching test. Once past burn-in, for new subject t with covariate vector x_t, the squared Mahalanobis distance (via compute_proportional_mahal_distances_cpp(), using the sample covariance of all prior subjects' covariates, ridge-regularized by .Machine$double.eps) to every subject currently in the reservoir is computed, and the closest one is a candidate match. The match is accepted only if that squared distance falls below a threshold T^2_{\mathrm{cutoff}} derived from an F critical value,

T^2_{\mathrm{cutoff}} = \frac{p(n-1)}{n-p} \, F_{p,\, t-p}(\lambda),

where p is the rank of the covariate matrix so far, n is the design's target (planned) sample size, t is the number of subjects enrolled so far, and \lambda (lambda, default 0.1) is the F-distribution quantile level — i.e. this is a Hotelling's T^2-type test of whether the candidate pair's covariate difference is small enough to plausibly be exchangeable "noise" rather than a meaningful covariate mismatch; lambda controls how strict that test is (smaller lambda accepts fewer, closer matches). If accepted, both subjects are recorded as a new match (private$m), and the new subject receives the opposite treatment of its match, guaranteeing exactly one treated and one control per matched pair — the same guarantee DesignFixedBinaryMatch provides, but formed incrementally rather than all at once. If rejected (or the reservoir is empty), the subject is randomized and added to the reservoir instead.

Lifecycle note. The morrison and p constructor arguments are currently recorded on the object but not consulted anywhere in the matching or assignment logic in this version of the class; treat them as reserved for future use rather than as active configuration.

Super classes

Design -> DesignSeqOneByOne -> DesignSeqOneByOneKK14

Methods

Public methods

+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design

DesignSeqOneByOneKK14$is_a_kk_matching_capable()

Characterization: this design matches subjects to each other incrementally as they arrive (see class documentation), so it is KK matching-on-the-fly-capable by construction.

Usage
DesignSeqOneByOneKK14$is_a_kk_matching_capable()
Returns

Always TRUE for this class.


DesignSeqOneByOneKK14$new()

Initialize a KK14 sequential matching-on-the-fly experimental design (see class documentation for the burn-in and matching-test rules).

Usage
DesignSeqOneByOneKK14$new(
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  lambda = NULL,
  t_0_pct = NULL,
  morrison = FALSE,
  p = NULL,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

"continuous", "incidence", "proportion", "count", "survival", or "ordinal".

prob_T

Probability of treatment assignment used for burn-in/ unmatched (reservoir) subjects; matched subjects instead always receive the opposite assignment of their match (see class documentation).

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

A flag for verbosity.

lambda

The F-distribution quantile level controlling how strict the matching-acceptance test is (default 0.1; smaller values accept fewer, closer matches). See class documentation for the exact threshold formula.

t_0_pct

The fraction of n subjects to randomize into the reservoir before matching begins (default 0.35).

morrison

Currently unused by this class's matching/assignment logic; reserved for future use.

p

Currently unused by this class's matching/assignment logic; reserved for future use.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignSeqOneByOneKK14' object


DesignSeqOneByOneKK14$assign_wt()

Draw the next subject's treatment assignment via KK14 matching-on-the-fly (see class documentation): during burn-in, or if no sufficiently close reservoir match exists, randomize and add the subject to the reservoir; otherwise match to the nearest reservoir subject and assign the opposite treatment.

Usage
DesignSeqOneByOneKK14$assign_wt()
Returns

The treatment assignment (0 or 1) for the next subject.


DesignSeqOneByOneKK14$clone()

The objects of this class are cloneable with this method.

Usage
DesignSeqOneByOneKK14$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kapelner, A., and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148. See also Hotelling's T-squared distribution for the matching-test statistic, and Mahalanobis distance.

Examples

seq_des = DesignSeqOneByOneKK14$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))

Kapelner and Krieger's (2021) Outcome-Weighted Sequential Matching-on-the-Fly Design

Description

A DesignSeqOneByOneKK14 extension that replaces KK14's unweighted Mahalanobis matching distance with a response-weighted squared distance: at each assignment, a per-covariate weight vector is re-estimated from the responses observed so far (weighting each covariate by the estimated strength of its association with the response, e.g. an absolute standardized regression coefficient), and the new subject is matched to its nearest reservoir subject under that weighted distance rather than the raw (unweighted) Mahalanobis distance KK14 uses. Weighting the match by outcome association means matching effort is spent preferentially on prognostic covariates (those that actually explain outcome variance) rather than equally on every covariate, which is intended to improve estimator efficiency beyond what outcome-agnostic matching achieves. morrison = TRUE additionally switches to Morrison and Owen's (2025) alternative calibration of the matching-acceptance threshold (differing in the fixed- vs. variable-n settings) and removes KK14's burn-in wait before matching begins.

Weight estimation. compute_weights() dispatches on response_type to a corresponding kk21_*_weights_cpp() backend that fits a per-covariate simple regression of the response on that covariate (given all responses observed so far) and returns the absolute t-statistic (coefficient over its standard error) as that covariate's weight: OLS for "continuous", logistic for "incidence", negative-binomial (or, if count_use_speedup = TRUE, OLS on log(y + 1)) for "count", beta regression (or OLS on the logit scale if proportion_use_speedup = TRUE) for "proportion", Weibull/lognormal/ log-logistic AFT (or OLS on log(y) if survival_use_speedup_for_no_censoring = TRUE and there is no censoring yet) for "survival", and proportional-odds (or OLS on the numeric-coerced level if ordinal_use_speedup = TRUE) for "ordinal". Weights are normalized to sum to 1 and cached per-iteration in private$iteration_weights (retrievable via get_iteration_weights()); the *_use_speedup flags trade weight accuracy for speed by substituting a fast continuous-regression proxy for the response-type-appropriate GLM/AFT/ordinal fit on every single assignment call.

Weighted matching test. The weighted squared distance from the new subject to every reservoir subject is computed (compute_weighted_sqd_distances_cpp()), and the match is accepted only if the minimum weighted distance falls below the private$compute_lambda() quantile of a bootstrapped reference distribution of weighted pairwise distances among subjects enrolled so far (compute_bootstrapped_weighted_sqd_distances_cpp(), num_boot resamples) — a nonparametric, simulation-based acceptance threshold, in contrast to KK14's closed-form F-distribution threshold. As in KK14, an accepted match receives the opposite treatment of its match; a rejected (or empty-reservoir) draw is randomized and added to the reservoir.

Fallback to KK14. Before enough responses have accumulated to fit the weight-estimation regressions reliably (fewer than 2 * (ncol(X) + 2) non-missing responses — two observations per regression parameter, at minimum ncol(X) + 2), assign_wt() falls back to the inherited (unweighted) KK14 assignment rule entirely, via super$assign_wt().

Super classes

Design -> DesignSeqOneByOne -> DesignSeqOneByOneKK14 -> DesignSeqOneByOneKK21

Methods

Public methods

+ inherited public methods from DesignSeqOneByOneKK14
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design

DesignSeqOneByOneKK21$new()

Initialize a matching-on-the-fly sequential experimental design which matches based on Kapelner and Krieger (2021) with option to use matching parameters of Morrison and Owen (2025)

Usage
DesignSeqOneByOneKK21$new(
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  lambda = NULL,
  t_0_pct = NULL,
  morrison = FALSE,
  p = NULL,
  num_boot = NULL,
  count_use_speedup = TRUE,
  proportion_use_speedup = TRUE,
  survival_use_speedup_for_no_censoring = TRUE,
  ordinal_use_speedup = TRUE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL,
  ...
)
Arguments
response_type

The data type of response values which must be one of the following: "continuous", "incidence", "proportion", "count", "survival". This package will enforce that all added responses via the add_one_subject_response method will be of the appropriate type.

prob_T

The probability of the treatment assignment. This defaults to 0.5.

include_is_missing_as_a_new_feature

If missing data is present in a variable, should we include another dummy variable for its missingness in addition to imputing its value? If the feature is type factor, instead of creating a new column, we allow missingness to be its own level. The default is TRUE.

n

The sample size (if fixed). Default is NULL for not fixed.

verbose

A flag indicating whether messages should be displayed to the user. Default is FALSE.

lambda

The quantile cutoff of the subject distance distribution for determining matches. If unspecified and morrison = FALSE, default is 10%.

t_0_pct

The percentage of total sample size n where matching begins. If unspecified and morrison = FALSE, default is 35%.

morrison

Default is FALSE which implies matching via the KK14 algorithm using lambda and t_0_pct matching. If TRUE, we use Morrison and Owen (2025)'s formula for lambda which differs in the fixed n versus variable n settings and matching begins immediately with no wait for a certain reservoir size like in KK14.

p

The number of covariate features. Must be specified when morrison = TRUE otherwise do not specify this argument.

num_boot

the number of bootstrap samples taken to approximate the subject-distance distribution. Default is 500.

count_use_speedup

Should we speed up the estimation of the weights in the response = count case via a continuous regression on log(y + 1). instead of a negative binomial regression each time? This is at the expense of the weights being less accurate. Default is TRUE.

proportion_use_speedup

Should we speed up the estimation of the weights in the response = proportion case via a continuous regression on log(y / (1 - y)) instead of a beta regression each time? This is at the expense of the weights being less accurate. Default is TRUE.

survival_use_speedup_for_no_censoring

Should we speed up the estimation of the weights in the response = survival case via a continuous regression on log(y) instead of a Weibull AFT regression each time, but only when there is no censoring in the data collected so far? This is at the expense of the weights being less accurate when censoring is present. Default is TRUE.

ordinal_use_speedup

Should we speed up the estimation of the weights in the response = ordinal case via a continuous regression on the ordinal levels coerced to numeric. instead of a proportional odds model each time? This is at the expense of the weights being less accurate. Default is TRUE.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

...

Extra arguments passed to the DesignSeqOneByOneKK14 superclass.

Returns

A new 'DesignSeqOneByOneKK21' object

Examples
seq_des = DesignSeqOneByOneKK21$new(n = 6, response_type = "continuous")
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x = rnorm(1)))

DesignSeqOneByOneKK21$get_iteration_weights()

Retrieve the full history of normalized covariate weight vectors computed by compute_weights() across every assignment call so far (see class documentation), keyed by subject index t, for inspecting how the outcome-informed weighting evolved as data accrued.

Usage
DesignSeqOneByOneKK21$get_iteration_weights()
Returns

A list of numeric weight vectors (one per assignment call at which weights were computed), each summing to 1 and named by covariate.


DesignSeqOneByOneKK21$get_covariate_weights()

Retrieve the normalized covariate weight vector from the most recent assignment call (see class documentation for how weights are estimated).

Usage
DesignSeqOneByOneKK21$get_covariate_weights()
Returns

A numeric vector of weights, one per covariate, summing to 1 and named by covariate; NULL if weights have not yet been computed (e.g. still in the KK14 fallback regime).


DesignSeqOneByOneKK21$assign_wt()

Draw the next subject's treatment assignment via the KK21 outcome-weighted matching-on-the-fly rule (see class documentation): falls back to unweighted KK14 matching if too few responses have accumulated to estimate covariate weights, otherwise re-estimates weights, matches to the nearest reservoir subject under the weighted distance if it clears the bootstrapped acceptance threshold (assigning the opposite treatment), or randomizes into the reservoir otherwise.

Usage
DesignSeqOneByOneKK21$assign_wt()
Returns

The treatment assignment (0 or 1) for the next subject.


DesignSeqOneByOneKK21$clone()

The objects of this class are cloneable with this method.

Usage
DesignSeqOneByOneKK21$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kapelner, A., and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the base matching-on-the-fly algorithm this class extends with outcome-weighted distances (Kapelner and Krieger, 2021); see also Morrison, T., and Owen, A. B. (2025) for the alternative morrison = TRUE threshold calibration referenced by the morrison argument.

Examples

seq_des = DesignSeqOneByOneKK21$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))

## ------------------------------------------------
## Method `DesignSeqOneByOneKK21$new()`
## ------------------------------------------------


seq_des = DesignSeqOneByOneKK21$new(n = 6, response_type = "continuous")
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x = rnorm(1)))



Stepwise Variant of the KK21 Outcome-Weighted Sequential Matching Design

Description

A DesignSeqOneByOneKK21 variant that computes its per-covariate matching weights via forward stepwise selection (compute_weights_KK21stepwise()) instead of KK21's independent marginal-association regressions: covariates are added to a growing "selected" set one at a time, at each step choosing whichever remaining covariate has the largest absolute association statistic conditional on (i.e. in a model that also includes) the covariates already selected and the treatment-assignment column, rather than each covariate's association with the response considered in isolation. This targets the case where covariates are mutually correlated: KK21's marginal weights can assign similar high weight to several collinear prognostic covariates (effectively double-counting the same information), whereas the stepwise conditional weights down-weight a covariate once its explanatory content is already captured by previously selected covariates.

Weight computation. For each response type, a family-appropriate model (OLS/logistic/negative-binomial/beta/AFT survival/proportional-odds, matching the same response-type dispatch and *_use_speedup fast-path conventions as DesignSeqOneByOneKK21) is repeatedly refit, each time regressing the response on one candidate remaining covariate plus all previously selected covariates plus the treatment column ws; the candidate with the largest absolute association statistic is selected next and assigned that statistic as its weight, then removed from the candidate pool, and the process repeats until every covariate has been assigned a weight (an O(p^2) number of model fits per assignment call, for p covariates). If a candidate's model fit fails to converge (e.g. perfect separation or rank deficiency) partway through, the remaining not-yet-selected covariates' weights are left NA internally and then replaced with 0 (excluding them from the weighted matching distance) rather than propagating the failure.

Everything else is inherited from KK21. Burn-in fallback to KK14, the bootstrapped acceptance-threshold test, matched-pair assignment, and the morrison/lambda/t_0_pct matching-schedule options are all unchanged from DesignSeqOneByOneKK21; only how the covariate weight vector is computed differs.

Super classes

Design -> DesignSeqOneByOne -> DesignSeqOneByOneKK14 -> DesignSeqOneByOneKK21 -> DesignSeqOneByOneKK21stepwise

Methods

Public methods

+ inherited public methods from DesignSeqOneByOneKK21
+ inherited public methods from DesignSeqOneByOneKK14
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design

DesignSeqOneByOneKK21stepwise$new()

Initialize a matching-on-the-fly sequential experimental design whose covariate matching weights are computed via forward stepwise selection (see class documentation), based on the stepwise version of Kapelner and Krieger (2021) with option to use matching parameters of Morrison and Owen (2025)

Usage
DesignSeqOneByOneKK21stepwise$new(
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  lambda = NULL,
  t_0_pct = NULL,
  morrison = FALSE,
  p = NULL,
  num_boot = NULL,
  count_use_speedup = TRUE,
  proportion_use_speedup = TRUE,
  survival_use_speedup_for_no_censoring = TRUE,
  ordinal_use_speedup = TRUE,
  missingness_method = "impute",
  design_formula = ~.,
  ...
)
Arguments
response_type

The data type of response values which must be one of the following: "continuous", "incidence", "proportion", "count", "survival". This package will enforce that all added responses via the add_one_subject_response method will be of the appropriate type.

prob_T

The probability of the treatment assignment. This defaults to 0.5.

include_is_missing_as_a_new_feature

If missing data is present in a variable, should we include another dummy variable for its missingness in addition to imputing its value? If the feature is type factor, instead of creating a new column, we allow missingness to be its own level. The default is TRUE.

n

The sample size (if fixed). Default is NULL for not fixed.

verbose

A flag indicating whether messages should be displayed to the user. Default is FALSE.

lambda

The quantile cutoff of the subject distance distribution for determining matches. If unspecified and morrison = FALSE, default is 10%.

t_0_pct

The percentage of total sample size n where matching begins. If unspecified and morrison = FALSE, default is 35%.

morrison

Default is FALSE which implies matching via the KK14 algorithm using lambda and t_0_pct matching. If TRUE, we use Morrison and Owen (2025)'s formula for lambda which differs in the fixed n versus variable n settings and matching begins immediately with no wait for a certain reservoir size like in KK14.

p

The number of covariate features. Must be specified when morrison = TRUE otherwise do not specify this argument.

num_boot

the number of bootstrap samples taken to approximate the subject-distance distribution. Default is 500.

count_use_speedup

Should we speed up the estimation of the weights in the response = count case via a continuous regression on log(y + 1). instead of a negative binomial regression each time? This is at the expense of the weights being less accurate. Default is TRUE.

proportion_use_speedup

Should we speed up the estimation of the weights in the response = proportion case via a continuous regression on log(y / (1 - y)) instead of a beta regression each time? This is at the expense of the weights being less accurate. Default is TRUE.

survival_use_speedup_for_no_censoring

Should we speed up the estimation of the weights in the response = survival case via a continuous regression on log(y) instead of a Weibull AFT regression each time, but only when there is no censoring in the data collected so far? This is at the expense of the weights being less accurate when censoring is present. Default is TRUE.

ordinal_use_speedup

Should we speed up the estimation of the weights in the response = ordinal case via a continuous regression on the ordinal levels coerced to numeric. instead of a proportional odds model each time? This is at the expense of the weights being less accurate. Default is TRUE.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

...

Extra arguments passed to the DesignSeqOneByOneKK21 superclass.

Returns

A new 'DesignSeqOneByOneKK21stepwise' object

Examples
seq_des = DesignSeqOneByOneKK21stepwise$new(n = 6, response_type = "continuous")

DesignSeqOneByOneKK21stepwise$clone()

The objects of this class are cloneable with this method.

Usage
DesignSeqOneByOneKK21stepwise$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kapelner, A., and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the base matching-on-the-fly algorithm; the outcome-weighted extension follows Kapelner and Krieger (2021), with this class using a forward-stepwise (rather than marginal) weight-estimation scheme. See also Morrison, T., and Owen, A. B. (2025) for the alternative morrison = TRUE threshold calibration referenced by the morrison argument.

Examples

seq_des = DesignSeqOneByOneKK21stepwise$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))

## ------------------------------------------------
## Method `DesignSeqOneByOneKK21stepwise$new()`
## ------------------------------------------------

seq_des = DesignSeqOneByOneKK21stepwise$new(n = 6, response_type = "continuous")


Pocock and Simon's (1975) Minimization Sequential Design

Description

A DesignSeqOneByOne implementing Pocock and Simon's minimization method: for each categorical covariate in strata_cols, the design tracks a running treated/control count per covariate level (private$counts), and for a new subject computes, for each candidate treatment arm k \in \{0, 1\}, a weighted total imbalance

G_k = \sum_{j} weights_j \cdot \mathrm{Var}\big(\text{counts at subject's level of covariate } j, \text{ after hypothetically assigning arm } k\big),

where the variance is taken across the two treatment arms' hypothetical counts at that covariate level (so G_k is large when arm k would leave the subject's covariate-level counts unbalanced, summed with weights across covariates). The subject is then assigned to whichever arm minimizes G_k with probability p_best (and to the other arm with probability 1 - p_best), or — if the two arms are exactly tied — via a plain \mathrm{Bernoulli}(prob\_T) draw. Unlike DesignSeqOneByOneAtkinson/ DesignSeqOneByOneKK14, which use continuous covariate distances, minimization operates on categorical/discretized strata and balances marginal covariate-level counts directly rather than a multivariate distance or matched-pair structure.

Level bookkeeping. private$ensure_factor_metadata() maintains a mapping from each observed level of each strata_cols column to a row index in private$counts (an (total levels across all covariates) x 2 matrix of running treated/control counts), growing both the level map and counts as new levels are encountered; missing values are treated as their own level ("NA").

Non-resampling bootstrap. draw_bootstrap_indices() always performs a plain i.i.d. nonparametric bootstrap over subjects (sample_int_replace_cpp()), since minimization's adaptive assignment process has no simple exchangeable resampling unit to preserve (each subject's assignment probability depends on the full sequence of covariate levels and assignments that preceded it).

Super classes

Design -> DesignSeqOneByOne -> DesignSeqOneByOnePocockSimon

Methods

Public methods

+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design

DesignSeqOneByOnePocockSimon$new()

Initialize a Pocock and Simon (1975) minimization sequential experimental design (see class documentation for the exact imbalance criterion and assignment rule).

Usage
DesignSeqOneByOnePocockSimon$new(
  strata_cols,
  weights = NULL,
  p_best = 0.8,
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
strata_cols

The names of the covariates to be used for minimization. These must be factor or categorical variables.

weights

A numeric vector of per-covariate weights weights_j in the imbalance criterion G_k (see class documentation), one per entry of strata_cols, in the same order. Defaults to 1 for all (equal-weighted covariates).

p_best

The probability of assigning the treatment arm that minimizes G_k (see class documentation); the complementary arm is assigned with probability 1 - p_best. Defaults to 0.8 (an 80/20 biased coin favoring the balancing arm, rather than a fully deterministic minimization rule).

response_type

The data type of response values.

prob_T

The probability of the treatment assignment used only when the two arms' imbalance is exactly tied (see class documentation).

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

Flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignSeqOneByOnePocockSimon' object


DesignSeqOneByOnePocockSimon$assign_wt()

Draw the next subject's treatment assignment via Pocock and Simon minimization (see class documentation for the exact imbalance criterion G_k and the p_best/prob_T assignment rule), and update the running per-covariate-level treated/control counts in-place to reflect this assignment.

Usage
DesignSeqOneByOnePocockSimon$assign_wt()
Returns

The treatment assignment (0 or 1) for the next subject.


DesignSeqOneByOnePocockSimon$clone()

The objects of this class are cloneable with this method.

Usage
DesignSeqOneByOnePocockSimon$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Pocock, S. J., and Simon, R. (1975). "Sequential treatment assignment with balancing for prognostic factors in the controlled clinical trial." Biometrics, 31(1), 103-115, doi:10.2307/2529712. See also minimisation (clinical trials) for orientation.

Examples

seq_des = DesignSeqOneByOnePocockSimon$new(strata_cols = 'x1', n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = factor(1, levels=1:2)))

A Sequential Permuted-Block Design with Randomly Varying Block Sizes

Description

A DesignSeqOneByOne implementing permuted-block randomization with randomly varying block size: subjects are assigned from a queue of pre-shuffled treatment labels (a "block"), refilled with a fresh sampled block whenever it empties. Each new block's size is itself drawn uniformly at random from block_sizes (rather than being fixed), and each block internally contains exactly round(block_size * prob_T) treated and block_size - round(block_size * prob_T) control labels in random order. Randomizing the block size (rather than using a single fixed block length, as in classical permuted-block designs) is a standard clinical-trials safeguard against selection bias: with a fixed, known block size, unblinded staff could predict the last assignment(s) in a block from the ones already observed, whereas an unpredictable block size makes this much harder while still guaranteeing near-perfect treatment/control balance throughout enrollment (balance is exact at every block boundary and never worse than one full block's imbalance in between). If strata_cols is supplied, a separate independent sequence of blocks is maintained per stratum (one queue per distinct combination of strata_cols values), so balance holds within each stratum, not just overall.

Block-size / prob_T compatibility. Every entry of block_sizes must yield an integer number of treated subjects when multiplied by prob_T (checked at construction: abs(bs * prob_T - round(bs * prob_T)) <= 1e-10 for every bs); a block size that would require a fractional number of treated subjects is rejected.

Per-stratum queues. private$strata_states is a hashed environment mapping each stratum key (or the literal key "overall" when strata_cols is NULL) to the vector of not-yet-used assignments remaining in that stratum's current block; assign_wt() pops the next assignment from the relevant queue, refilling it with a freshly drawn block (random size, randomly ordered) whenever it is empty.

Bootstrap. draw_bootstrap_indices() resamples within strata (via stratified_bootstrap_indices_cpp()) when strata_cols is supplied, or performs a plain i.i.d. nonparametric bootstrap over subjects otherwise.

Super classes

Design -> DesignSeqOneByOne -> DesignSeqOneByOneRandomBlockSize

Methods

Public methods

+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design

DesignSeqOneByOneRandomBlockSize$new()

Initialize a sequential permuted-block experimental design with randomly varying block size (see class documentation for the exact block-refill rule and its selection-bias rationale).

Usage
DesignSeqOneByOneRandomBlockSize$new(
  strata_cols = NULL,
  block_sizes = c(4, 6, 8),
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
strata_cols

A character vector of column names to use for stratification. If NULL, simple blocking is used.

block_sizes

A vector of positive integers representing the possible block sizes to choose from. Each must be a multiple of the inverse of prob_T to ensure integer treatment/control counts.

response_type

The data type of response values which must be one of the following: "continuous", "incidence", "proportion", "count", "survival", "ordinal".

prob_T

The probability of the treatment assignment. This defaults to 0.5.

include_is_missing_as_a_new_feature

If missing data is present in a variable, should we include another dummy variable for its missingness? Default is TRUE.

n

The sample size (if fixed). Default is NULL for not fixed.

verbose

A flag indicating whether messages should be displayed. Default is FALSE.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignSeqOneByOneRandomBlockSize' object


DesignSeqOneByOneRandomBlockSize$assign_wt()

Pop the next treatment assignment from the current subject's stratum block queue (see class documentation), refilling that queue with a freshly drawn random-size, randomly-ordered block first if it is empty.

Usage
DesignSeqOneByOneRandomBlockSize$assign_wt()
Returns

The treatment assignment (0 or 1) for the next subject.


DesignSeqOneByOneRandomBlockSize$clone()

The objects of this class are cloneable with this method.

Usage
DesignSeqOneByOneRandomBlockSize$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Efron, B. (1971). "Forcing a sequential experiment to be balanced." Biometrika, 58(3), 403-417, doi:10.1093/biomet/58.3.403, for sequential balanced-block randomization background. See also block randomisation for orientation on permuted-block designs and the selection-bias rationale for varying block size.

Examples

seq_des = DesignSeqOneByOneRandomBlockSize$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))

A Stratified Permuted-Block Sequential Design (SPBR) with Fixed Block Size

Description

A DesignSeqOneByOne implementing classical stratified permuted-block randomization: subjects are assigned from a per-stratum queue of pre-shuffled treatment labels (a "block" of block_size labels, containing exactly round(block_size * prob_T) treated and the remainder control, in random order), refilled with a fresh randomly-ordered block of the same fixed size whenever a stratum's queue empties. This is the fixed-block-size, mandatory-stratification counterpart of DesignSeqOneByOneRandomBlockSize (which varies block size across draws and makes stratification optional): here, strata_cols is required, and a single fixed block_size is used for every stratum and every block, guaranteeing exact treatment/control balance within each stratum at every block boundary.

Block-size / prob_T compatibility. block_size must yield an integer number of treated subjects: the constructor errors unless abs(block_size * prob_T - round(block_size * prob_T)) <= 1e-10.

Per-stratum queues. As in DesignSeqOneByOneRandomBlockSize, private$strata_states is a hashed environment mapping each stratum key (concatenated strata_cols values, "NA" for missing) to the vector of not-yet-used assignments remaining in that stratum's current block; assign_wt() pops the next assignment, refilling with a fresh block when empty. draw_bootstrap_indices() resamples within strata by default (bootstrap_type = "within_blocks" or NULL) or resamples whole strata otherwise.

Super classes

Design -> DesignSeqOneByOne -> DesignSeqOneByOneSPBR

Methods

Public methods

+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design

DesignSeqOneByOneSPBR$new()

Initialize a stratified permuted-block sequential experimental design with fixed block size (see class documentation for the exact block-refill rule).

Usage
DesignSeqOneByOneSPBR$new(
  strata_cols,
  block_size = 4,
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
strata_cols

A character vector of column names to use for stratification.

block_size

The size of the permuted blocks (fixed; see class documentation for its compatibility requirement with prob_T).

response_type

"continuous", "incidence", "proportion", "count", "survival", or "ordinal".

prob_T

Probability of treatment assignment.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

A flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignSeqOneByOneSPBR' object


DesignSeqOneByOneSPBR$assign_wt()

Pop the next treatment assignment from the current subject's stratum block queue (see class documentation), refilling that queue with a freshly drawn fixed-size, randomly-ordered block first if it is empty.

Usage
DesignSeqOneByOneSPBR$assign_wt()
Returns

The treatment assignment (0 or 1) for the next subject.


DesignSeqOneByOneSPBR$clone()

The objects of this class are cloneable with this method.

Usage
DesignSeqOneByOneSPBR$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Zelen, M. (1974). "The randomization and stratification of patients to clinical trials." Journal of Chronic Diseases, 27(7-8), 365-375, doi:10.1016/0021-9681(74)90015-0, for stratified permuted-block randomization. See also block randomisation for orientation, and DesignSeqOneByOneRandomBlockSize for the randomly-varying-block-size variant.

Examples

seq_des = DesignSeqOneByOneSPBR$new(strata_cols = 'x1', n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = factor(1, levels=1:2)))

Wei's (1977, 1978) Adaptive Urn Sequential Design, UD(\alpha, \beta)

Description

A DesignSeqOneByOne implementing Wei's adaptive biased-coin urn design UD(\alpha, \beta): conceptually, an urn starts with \alpha balls of each type (treatment and control), and each assignment is drawn proportionally to the current ball counts, then \beta balls of the opposite type to whatever was drawn are added back to the urn (so drawing treatment adds \beta control balls, and vice versa), pushing subsequent draws toward the under-represented arm. No covariates are used; only the running treated/control counts n_T, n_C matter, via the closed-form assignment probability

\Pr(w_t = 1) = \frac{\alpha + \beta \, n_C}{2\alpha + \beta (n_T + n_C)}.

Like DesignSeqOneByOneEfron, this design balances running assignment counts online while remaining strictly randomized (the probability is always strictly between 0 and 1 for finite \alpha, \beta > 0); unlike Efron's design (which only distinguishes "balanced" vs. "imbalanced" and applies a single fixed weighted_coin_prob in the imbalanced case), the urn design's bias toward the under-represented arm scales continuously and smoothly with the current degree of imbalance, tuned by the ratio \beta/\alpha: larger \beta/\alpha yields stronger balancing pressure, and \beta = 0 recovers a fixed \mathrm{Bernoulli}(0.5) coin (no adaptation at all).

Super classes

Design -> DesignSeqOneByOne -> DesignSeqOneByOneUrn

Methods

Public methods

+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design

DesignSeqOneByOneUrn$new()

Initialize Wei's UD(\alpha, \beta) adaptive urn sequential experimental design (see class documentation for the exact assignment-probability formula).

Usage
DesignSeqOneByOneUrn$new(
  alpha = 1,
  beta = 1,
  response_type,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
alpha

The initial number of balls of each type (Treatment/Control) in the conceptual urn; larger alpha relative to beta weakens the balancing effect (assignment probabilities stay closer to 0.5 for longer).

beta

The number of balls of the opposite type added to the urn after each assignment; beta = 0 recovers an unbiased \mathrm{Bernoulli}(0.5) coin (no balancing).

response_type

The data type of response values.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

A flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignSeqOneByOneUrn' object


DesignSeqOneByOneUrn$assign_wt()

Draw the next subject's treatment assignment from Wei's UD(\alpha, \beta) urn probability (see class documentation), computed from the running treated/control counts.

Usage
DesignSeqOneByOneUrn$assign_wt()
Returns

The treatment assignment (0 or 1) for the next subject.


DesignSeqOneByOneUrn$clone()

The objects of this class are cloneable with this method.

Usage
DesignSeqOneByOneUrn$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Wei, L. J. (1977). "A class of designs for sequential clinical trials." Journal of the American Statistical Association, 72(358), 382-386, doi:10.1080/01621459.1977.10481006; Wei, L. J. (1978). "The adaptive biased coin design for sequential experiments." The Annals of Statistics, 6(1), 92-100, doi:10.1214/aos/1176344068. See also randomized experiment for orientation on adaptive biased-coin sequential designs.

Examples

seq_des = DesignSeqOneByOneUrn$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))

A Sequential Design Guaranteeing Exact Terminal Balance (Random Allocation Rule)

Description

A DesignSeqOneByOne implementing the "random allocation rule" (a sequential realization of complete randomization): treatment is assigned to each arriving subject with probability equal to the fraction of remaining treatment slots among all remaining slots, \Pr(w_t = 1) = n_{T,\mathrm{rem}} / (n_{T,\mathrm{rem}} + n_{C,\mathrm{rem}}), where n_{T,\mathrm{rem}} = \mathrm{round}(n \cdot prob\_T) - n_T and n_{C,\mathrm{rem}} = (n - \mathrm{round}(n \cdot prob\_T)) - n_C are the treatment/control slots not yet used, given the running counts n_T, n_C. This guarantees the realized sequence, once all n subjects have arrived, has exactly \mathrm{round}(n \cdot prob\_T) treated subjects — the same terminal allocation-count guarantee as DesignFixediBCRD's complete randomization, but realized online as subjects arrive one at a time rather than all at once, and with every prefix of the sequence itself drawn from the correct conditional (hypergeometric) distribution given the slots used so far. If a slot type is exhausted (n_{T,\mathrm{rem}} \le 0 or n_{C,\mathrm{rem}} \le 0), the remaining subjects are deterministically assigned to whichever type still has open slots.

No target n: falls back to Bernoulli. If n was not supplied at construction (private$n is NULL), there is no terminal target to balance toward, so assign_wt() falls back to an unbiased \mathrm{Bernoulli}(prob\_T) draw for every subject instead (equivalent to DesignSeqOneByOneBernoulli).

Single implicit block. add_one_subject_to_experiment_and_assign() overrides the inherited method only to additionally set private$m to a constant vector of 1s (a single block containing every subject enrolled so far) after each assignment, mirroring the fixed-sample DesignFixediBCRD's single-block convention for shared blocking/matching machinery.

Super classes

Design -> DesignSeqOneByOne -> DesignSeqOneByOneiBCRD

Methods

Public methods

+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design

DesignSeqOneByOneiBCRD$new()

Initialize a sequential design targeting exact terminal treatment/control balance (see class documentation for the assignment rule and the no-fixed-n fallback).

Usage
DesignSeqOneByOneiBCRD$new(
  response_type,
  prob_T = 0.5,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

"continuous", "incidence", "proportion", "count", "survival", or "ordinal".

prob_T

Target probability of treatment assignment; the terminal number of treated subjects is fixed at round(n * prob_T) when n is known (see class documentation).

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The planned (target) sample size; if NULL, there is no terminal balance target and assignment falls back to an unbiased Bernoulli coin (see class documentation).

verbose

A flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'DesignSeqOneByOneiBCRD' object


DesignSeqOneByOneiBCRD$add_one_subject_to_experiment_and_assign()

Add one subject to the experiment and assign treatment via assign_wt() (delegating to the inherited DesignSeqOneByOne$add_one_subject_to_experiment_and_assign()), then set private$m to a single-block vector of 1s covering every subject enrolled so far (see class documentation), overwritten on every call rather than only once all subjects have arrived.

Usage
DesignSeqOneByOneiBCRD$add_one_subject_to_experiment_and_assign(x_new)
Arguments
x_new

A data frame with one row representing the new subject's covariates.

Returns

The treatment assignment (0 or 1) for the newly added subject.


DesignSeqOneByOneiBCRD$assign_wt()

Draw the next subject's treatment assignment via the random allocation rule (see class documentation): with probability equal to the fraction of remaining treatment slots among all remaining slots, or a deterministic assignment if one slot type is exhausted; falls back to an unbiased Bernoulli coin if no fixed n was supplied.

Usage
DesignSeqOneByOneiBCRD$assign_wt()
Returns

The treatment assignment (0 or 1) for the next subject.


DesignSeqOneByOneiBCRD$clone()

The objects of this class are cloneable with this method.

Usage
DesignSeqOneByOneiBCRD$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Rosenberger, W. F., and Lachin, J. M. (2016). Randomization in Clinical Trials: Theory and Practice (2nd ed.), Wiley, for the random allocation rule as a sequential implementation of complete randomization. See also DesignFixediBCRD for the fixed-sample (all-at-once) version of the same terminal randomization law.

Examples

seq_des = DesignSeqOneByOneiBCRD$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))

Comprehensive-test slow-path registry

Description

Performance-based exclusions used by EDI's comprehensive test harness. These rules describe paths that are implemented but intentionally omitted from routine exhaustive execution because their observed runtime is too high. They do not change an inference class's public capabilities and must not be interpreted as "not implemented" declarations.

Usage

EDI_COMPREHENSIVE_SLOW_PATHS

Format

A named list. 'exact_operations' contains keys of the form 'response_type||InferenceClass||operation', optionally suffixed with '||model_formula' (e.g. '||~1') to restrict the entry to that one formula – an operation with no formula suffix is skipped for every formula, matching the pre-2026-09-08 behavior. 'model_formula' is matched as the deparsed formula text ('"~1"', '"~."', ...). Every other element contains formula-, dataset-, and design-independent concrete inference-class names for the named slow-path family.

Details

Rule-name prefixes use the harness vocabulary: 'boot' is ordinary nonparametric bootstrap, 'bbt' is Bayesian bootstrap, 'brt' is bootstrap randomization, 'pboot'/'param_bootstrap' are parametric bootstrap, and 'rand' is randomization inference. Suffixes identify the affected CI, p-value, or typed variant. A class can appear in more than one category.

'InferenceSuite$run_all_inference()' also omits matching class/method/type combinations when its 'methods' argument is left at the default 'NULL'. Supplying 'methods' explicitly opts into the requested paths even when they appear in this registry. Because of that coupling, an 'exact_operations' entry must never name the base 'compute_estimate' operation itself (only a CI/p-value sub-method, e.g. 'compute_rand_two_sided_pval') – doing so doesn't just skip one slow sub-computation, it silently drops the whole class from ‘run_all_inference()'’s default output (found 2026-08-26: a pre-existing '"survival||InferenceSurvivalCoxPHRegr||compute_estimate"' entry, harmless while this registry was internal-only, started doing exactly that the moment 'run_all_inference()' began consulting it – 'compute_estimate()' for that class takes ~0.09s, nowhere near "too slow"; removed).

See Also

[InferenceSuite]


Split from a single "RR" tag into "RR" (raw risk ratio) vs. "log_risk_ratio" – per user decision, 2026-08-23: 'InferenceIncidGCompRiskRatio'/ 'InferenceIncidKKGCompRiskRatio' compute the raw ratio 'risk1/risk0' directly (a nonlinear function of an underlying logistic model's coefficients) and only touch log space as a delta-method device to get a positive-respecting CI/SE, exponentiating back before returning – "RR" is their natural, directly-computed scale. 'InferenceIncidModifiedPoisson'/ 'InferenceIncidLogBinomial'/'InferenceIncidKKModifiedPoisson' instead fit a genuine log-link regression model (Zou's modified-Poisson working likelihood, or a log-link binomial GLM) whose own coefficient *is* log(RR) by construction, with a directly-computed (non-delta-method) coefficient SE – log(RR) is their natural scale, and "RR" is the derived quantity ('exp(coefficient)'). Both groups previously shared one "RR" tag, which put a log10 x-axis and null-reference line at 1 on what were actually already-log-scale estimates/CIs for the second group – a genuine scale mismatch, not just a display nicety.

Description

Split from a single "RR" tag into "RR" (raw risk ratio) vs. "log_risk_ratio" – per user decision, 2026-08-23: 'InferenceIncidGCompRiskRatio'/ 'InferenceIncidKKGCompRiskRatio' compute the raw ratio 'risk1/risk0' directly (a nonlinear function of an underlying logistic model's coefficients) and only touch log space as a delta-method device to get a positive-respecting CI/SE, exponentiating back before returning – "RR" is their natural, directly-computed scale. 'InferenceIncidModifiedPoisson'/ 'InferenceIncidLogBinomial'/'InferenceIncidKKModifiedPoisson' instead fit a genuine log-link regression model (Zou's modified-Poisson working likelihood, or a log-link binomial GLM) whose own coefficient *is* log(RR) by construction, with a directly-computed (non-delta-method) coefficient SE – log(RR) is their natural scale, and "RR" is the derived quantity ('exp(coefficient)'). Both groups previously shared one "RR" tag, which put a log10 x-axis and null-reference line at 1 on what were actually already-log-scale estimates/CIs for the second group – a genuine scale mismatch, not just a display nicety.

Usage

EDI_INFERENCE_ESTIMAND_TAGS

Exact binomial incidence component source

Description

Source list for the exact-binomial incidence component.

Usage

ExactBinomialIncidenceSource

Exact Fisher incidence component source

Description

Source list for the exact Fisher incidence component.

Usage

ExactFisherIncidenceSource

Exact Zhang incidence component source

Description

Source list for the exact Zhang incidence component.

Usage

ExactZhangIncidenceSource

KK conditional-logit IVWC component source

Description

Initialize conditional-logistic IVWC inference for KK binary responses and prepare separate matched-pair and reservoir likelihood components used by InferenceIncidKKCondLogitIVWC.

Computes the class-specific treatment-effect estimate; see Inference.

Uses the shared asymptotic confidence-interval contract; see InferenceAsymp.

Uses the shared asymptotic two-sided p-value contract; see InferenceAsymp.

Usage

IncidKKCondLogitIVWCSource

Details

Source list for the KK conditional-logit inverse-variance-weighted-combination (IVWC) incidence component.


Conditional Logistic Combined-Likelihood Inference for KK Designs with Binary Responses

Description

Initialize conditional-logistic combined-likelihood inference for KK binary responses and prepare matched-pair conditional-logit plus reservoir Bernoulli likelihood components. See InferenceAsympLik for shared likelihood-test methods.

Computes the class-specific treatment-effect estimate; see Inference.

Recomputes the combined conditional-logistic estimate under Bayesian-bootstrap weights.

Uses the shared asymptotic confidence-interval contract; see InferenceAsymp.

Uses the shared asymptotic two-sided p-value contract; see InferenceAsymp.

Usage

IncidKKCondLogitOneLikLikelihoodSource

Details

Fits a single joint likelihood over all KK design data for incidence responses. The matched-pair component uses the conditional logistic likelihood, and the reservoir component uses the standard Bernoulli log-likelihood.


Identity-link binomial regression component source

Description

Initialize inference for the identity-link binomial risk- difference model P(Y_i = 1) = \beta_0 + \beta_T W_i + X_i^\top \gamma; see InferenceIncidBinomialIdentityRiskDiff for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Refits the identity-link binomial model with subject/block-level weights applied to the fitting log-likelihood (Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights()) via fast_identity_binomial_regression_weighted_cpp, and returns the reweighted risk-difference estimate \hat\beta_T^{(w)}. Uses the same QR column-dropping hardening and fit-reasonableness check as compute_estimate(); a hardened-but-still-unreasonable fit is cached as nonestimable and returns NA.

Likelihood-ratio confidence interval for \beta_T by test inversion (find the set of delta not rejected at level alpha by the likelihood-ratio test); see InferenceAsympLik for the shared inversion contract. Falls back to a nonestimable result (NA bounds) if the underlying root-finding fails.

Usage

IncidenceBinomialIdentityLikelihoodSource

Details

Source list for the IncidenceBinomialIdentityLikelihood component composed by InferenceIncidBinomialIdentityRiskDiff.


KK incidence g-computation component source

Description

Initialize KK marginal g-computation inference for a completed incidence design; prepares the KK match structure used by the cluster-robust sandwich covariance.

Usage

IncidenceKKGComputationSource

Details

Source list for the IncidenceKKGComputation component shared by the KK g-computation incidence classes.


Log-binomial likelihood component source

Description

Initialize inference for the log-link binomial risk-ratio model \log P(Y_i = 1) = \beta_0 + \beta_T W_i + X_i^\top \gamma; see InferenceIncidLogBinomial for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Refits the log-binomial model with subject/block-level weights applied to the fitting log-likelihood (Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights()) via fast_log_binomial_regression_weighted_cpp, and returns the reweighted log-risk-ratio estimate \hat\beta_T^{(w)}. Uses the same QR column-dropping hardening and fit-reasonableness check as compute_estimate(); a hardened-but-still-unreasonable fit is cached as nonestimable and returns NA.

Score confidence interval for \beta_T by test inversion of the score test (find the set of delta not rejected at level alpha); see InferenceAsympLik for the shared inversion contract. Falls back to a nonestimable result (NA bounds) if the underlying root-finding fails or degenerates.

Gradient confidence interval for \beta_T by test inversion of the gradient test; see InferenceAsympLik for the shared inversion contract. Falls back to a nonestimable result (NA bounds) if the underlying root-finding fails or degenerates.

Usage

IncidenceLogBinomialLikelihoodSource

Details

Source list for the IncidenceLogBinomialLikelihood component composed by InferenceIncidLogBinomial.


Logistic likelihood component source

Description

Initialize inference for the logistic regression model \mathrm{logit}(P(Y_i = 1)) = \beta_0 + \beta_T W_i + X_i^\top \gamma; see InferenceIncidLogRegr for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Refits the logistic model with subject/block-level weights applied to the fitting log-likelihood (Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights()) via fast_logistic_regression_weighted_cpp, and returns the reweighted log-odds-ratio estimate \hat\beta_T^{(w)}. Uses the same QR column-dropping hardening and fit-reasonableness check as compute_estimate(); a hardened-but-still-unreasonable fit (e.g. near-perfect separation under the resampled weights) is cached as nonestimable and returns NA.

Usage

IncidenceLogisticLikelihoodSource

Details

Source list for the IncidenceLogisticLikelihood component composed by InferenceIncidLogRegr.


Modified-Poisson likelihood component source

Description

Initialize inference for the modified Poisson model \log E[Y_i \mid w_i, x_i] = \beta_0 + \beta_T w_i + x_i^\top \gamma; see InferenceIncidModifiedPoisson for the model form and the non-robust-SE caveat. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Fits the modified Poisson model by maximizing the Poisson working log-likelihood on the binary response and returns the log-risk-ratio estimate \hat\beta_T.

Wald confidence interval for \beta_T using the model-based (non-robust) Poisson-working-likelihood standard error; see InferenceIncidModifiedPoisson's non-robust-SE caveat and InferenceAsymp for the shared Wald contract.

Two-sided Wald test of H_0: \beta_T = \code{delta} using the model-based (non-robust) Poisson-working-likelihood standard error; see InferenceIncidModifiedPoisson's non-robust-SE caveat.

Refits the modified Poisson model with subject/block-level weights applied to the working log-likelihood (Bayesian-bootstrap or nonparametric-bootstrap draw weights) via fast_poisson_regression_weighted_cpp, and returns the reweighted log-risk-ratio estimate \hat\beta_T^{(w)}. Uses the same QR column-dropping hardening and fit-reasonableness check as compute_estimate(); a hardened-but-still-unreasonable fit is cached as nonestimable and returns NA.

Usage

IncidenceModifiedPoissonLikelihoodSource

Details

Source list for the IncidenceModifiedPoissonLikelihood component composed by InferenceIncidModifiedPoisson.


Probit likelihood component source

Description

Initialize inference for the probit regression model \Phi^{-1}(P(Y_i = 1)) = \beta_0 + \beta_T W_i + X_i^\top \gamma; see InferenceIncidProbitRegr for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Refits the probit model with subject/block-level weights applied to the fitting log-likelihood (Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights()) via fast_probit_regression_weighted_cpp, and returns the reweighted estimate \hat\beta_T^{(w)} on the latent standard-normal-index scale. Uses the same QR column-dropping hardening and fit-reasonableness check as compute_estimate(); a hardened-but-still-unreasonable fit is cached as nonestimable and returns NA.

Usage

IncidenceProbitLikelihoodSource

Details

Source list for the IncidenceProbitLikelihood component composed by InferenceIncidProbitRegr.


Inference for A Sequential Design

Description

An abstract R6 Class that estimates, tests and provides intervals for a treatment effect in a completed design. This class takes a completed Design object as an input where this object contains data for a fully completed experiment (i.e. all treatment assignments were allocated and all responses were collected).

Active bindings

num_cores

Current number of cores for this inference object. Defaults to the global budget unless overridden on the object.

Methods

Public methods


Inference$new()

Initialize an estimation and test object after the design is completed.

Usage
Inference$new(
  des_obj,
  verbose = FALSE,
  harden = TRUE,
  model_formula = NULL,
  smart_cold_start_default = NULL,
  seed = NULL
)
Arguments
des_obj

A completed Design object whose entire n subjects are assigned and response y is recorded within.

verbose

Whether to print progress messages.

harden

Whether to apply robustness measures (default TRUE). When TRUE, the inference methods employ defensive strategies including QR-based rank reduction of the design matrix, progressive correlation-threshold dropping, and fallback fits (e.g.\ robust survival regression, treatment-only models) to avoid crashes on ill-conditioned data. When FALSE, the vanilla algorithm runs on the full design matrix as supplied; any rank deficiency or convergence failure will surface as an error rather than being silently worked around. Set to FALSE when you want to verify that the raw model converges without intervention.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

smart_cold_start_default

Whether to use smart cold start values by default for likelihood-based models. Explicit starts always override this object-level policy. NULL (default) consults the global cold-start dispatch policy.

seed

Integer seed for reproducibility.


Inference$capabilities()

Returns the effective metadata-backed capabilities for this inference object.

Usage
Inference$capabilities()
Returns

A character vector of capability names.


Inference$supports()

Returns whether this inference object supports a metadata-backed capability.

Usage
Inference$supports(capability)
Arguments
capability

Capability name or names.

Returns

A logical vector aligned with capability.


Inference$compute_exact_two_sided_pval_for_treatment_effect()

Computes an exact two-sided p-value. Subclasses that support exact inference override this; inference objects that do not support exact methods throw an error.

Usage
Inference$compute_exact_two_sided_pval_for_treatment_effect(...)
Arguments
...

Other arguments passed to the method.


Inference$compute_exact_confidence_interval()

Computes an exact confidence interval. Subclasses that support exact inference override this; inference objects that do not support exact methods throw an error.

Usage
Inference$compute_exact_confidence_interval(...)
Arguments
...

Other arguments passed to the method.


Inference$compute_asymp_two_sided_pval()

Computes an asymptotic two-sided p-value. Subclasses that support asymptotic inference override this; inference objects that do not support asymptotic methods throw an error.

Usage
Inference$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect.


Inference$compute_asymp_confidence_interval()

Computes an asymptotic confidence interval. Subclasses that support asymptotic inference override this; inference objects that do not support asymptotic methods throw an error.

Usage
Inference$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level.


Inference$compute_estimate()

Computes the treatment-effect estimate. Concrete subclasses implement the model-specific estimator, such as a fitted regression coefficient, maximum-likelihood parameter, estimating-equation solution, mean or risk contrast, survival contrast, or rank statistic. Interval, p-value, bootstrap, jackknife, and randomization methods use this method as the canonical point-estimate contract.

Usage
Inference$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.

Returns

A numeric treatment estimate.


Inference$is_nonestimable()

Returns whether the most recent inference attempt explicitly marked the result as non-estimable.

Usage
Inference$is_nonestimable(type = c("any", "estimate", "se"))
Arguments
type

Which stage to query: "any", "estimate", or "se".

Returns

A logical scalar.


Inference$get_nonestimable_reason()

Returns the reason recorded for the most recent explicit non-estimability.

Usage
Inference$get_nonestimable_reason()
Returns

A character scalar or NULL.


Inference$get_nonestimable_stage()

Returns the stage recorded for the most recent explicit non-estimability.

Usage
Inference$get_nonestimable_stage()
Returns

A character scalar or NULL.


Inference$duplicate()

Duplicate this inference object

Usage
Inference$duplicate(verbose = FALSE, make_fork_cluster = FALSE)
Arguments
verbose

A flag indicating whether messages should be displayed.

make_fork_cluster

Whether the duplicate should be allowed to create a fork cluster. Default FALSE.

Returns

A new Inference object with the same data


Inference$get_response()

Return the response vector used by inference extension classes.

This accessor is part of the supported extension contract for user-defined R6 inference classes. Prefer this method over direct access to private fields.

Usage
Inference$get_response()
Returns

A numeric response vector.


Inference$get_treatment()

Return the treatment-assignment vector used by inference extension classes.

This accessor is part of the supported extension contract for user-defined R6 inference classes. Treatment is encoded as 0/1.

Usage
Inference$get_treatment()
Returns

A numeric or integer 0/1 treatment vector.


Inference$get_covariates()

Return the processed covariate matrix used by inference extension classes.

This accessor returns the design object's model-matrix covariates, after the package's missingness handling and encoding. It may be NULL if no covariates are available.

Usage
Inference$get_covariates()
Returns

A numeric matrix of covariates or NULL.


Inference$get_analysis_data()

Return a data frame with response, treatment, censoring status, and covariates.

This accessor is the preferred data interface for user-defined R6 inference classes. It avoids reliance on private implementation fields. The returned data frame always contains y, w, and dead; covariate columns are appended when available.

Usage
Inference$get_analysis_data()
Returns

A data frame suitable for user-defined model fitting.


Inference$get_design_object()

Return the completed design object backing this inference object.

This accessor is part of the supported extension contract. Extension classes should use this method instead of private$des_obj.

Usage
Inference$get_design_object()
Returns

The completed Design object.


Inference$get_response_type()

Return the response type for the backing design.

Usage
Inference$get_response_type()
Returns

A character scalar such as "continuous", "incidence", "proportion", "count", "survival", or "ordinal".


Inference$get_model_formula()

Return the model formula used for covariate adjustment.

Usage
Inference$get_model_formula()
Returns

A formula object or NULL.


Inference$set_optimization_alg()

Set the optimizer used by likelihood-based inference implementations.

Usage
Inference$set_optimization_alg(
  optimization_alg = NULL,
  allow_irls = private$optimization_alg_allow_irls,
  default = private$optimization_alg_default
)
Arguments
optimization_alg

The optimizer name. Valid values are configured by the concrete inference class.

allow_irls

Whether to allow IRLS (Iteratively Reweighted Least Squares) as a fallback or primary optimization algorithm.

default

The default optimizer to use if none is specified.

Returns

Invisibly returns self.


Inference$get_optimization_alg()

Return the optimizer used by likelihood-based inference implementations.

Usage
Inference$get_optimization_alg()

Inference$set_seed()

Set the seed for reproducibility.

Usage
Inference$set_seed(seed)
Arguments
seed

Integer seed for reproducibility.


Inference$clone()

The objects of this class are cloneable with this method.

Usage
Inference$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Abstract Conditional Logistic GLMM Inference

Description

Fits one likelihood with a conditional-logistic contribution from discordant matched pairs and a random-intercept logistic GLMM contribution from concordant matched pairs and reservoir subjects.

Super class

Inference -> InferenceAbstractKKCondLogitGLMM

Methods

Public methods

+ inherited public methods from Inference

InferenceAbstractKKCondLogitGLMM$new()

Initialize KK conditional-logit GLMM incidence inference, validate the binary response, and prepare the matched-pair conditional likelihood and reservoir mixed-model components. See InferenceAbstractKKCondLogitGLMM.

Usage
InferenceAbstractKKCondLogitGLMM$new(
  des_obj,
  model_formula = NULL,
  max_abs_reasonable_coef = 50,
  max_abs_reasonable_se = 10,
  max_abs_log_sigma = 8,
  verbose = FALSE,
  smart_cold_start_default = NULL,
  optimization_alg = NULL
)
Arguments
des_obj

A completed Design object with an incidence or proportion response.

model_formula

Optional formula for covariate adjustment.

max_abs_reasonable_coef

Cap for reasonable coefficient estimates.

max_abs_reasonable_se

Cap for reasonable treatment standard errors.

max_abs_log_sigma

Cap for reasonable log random effect variance.

verbose

Logical. Whether to print progress messages.

smart_cold_start_default

Logical. Whether to use smart starting values for the optimizer.

optimization_alg

Character. Optimization algorithm (default "lbfgs").


InferenceAbstractKKCondLogitGLMM$compute_estimate()

Computes the class-specific treatment-effect estimate; see Inference.

Usage
InferenceAbstractKKCondLogitGLMM$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

Logical. If TRUE, skip variance component calculations.


InferenceAbstractKKCondLogitGLMM$compute_estimate_with_bootstrap_weights()

Recomputes the class-specific treatment estimate for a bootstrap sample; see InferenceNonParamBootstrap.

Usage
InferenceAbstractKKCondLogitGLMM$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Numeric vector. Row weights for bootstrap.

estimate_only

Logical. If TRUE, skip variance component calculations.


InferenceAbstractKKCondLogitGLMM$compute_asymp_confidence_interval()

Uses the shared asymptotic confidence-interval contract; see InferenceAsymp.

Usage
InferenceAbstractKKCondLogitGLMM$compute_asymp_confidence_interval(
  alpha = 0.05
)
Arguments
alpha

Numeric. Significance level (default 0.05).


InferenceAbstractKKCondLogitGLMM$compute_asymp_two_sided_pval()

Uses the shared asymptotic two-sided p-value contract; see InferenceAsymp.

Usage
InferenceAbstractKKCondLogitGLMM$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Numeric. Null treatment effect value (default 0).


InferenceAbstractKKCondLogitGLMM$clone()

The objects of this class are cloneable with this method.

Usage
InferenceAbstractKKCondLogitGLMM$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Abstract class for all-subject marginal incidence inference in KK designs

Description

Abstract class for all-subject marginal incidence inference in KK designs

Super class

Inference -> InferenceAbstractKKMarginalIncid

Methods

Public methods

+ inherited public methods from Inference

InferenceAbstractKKMarginalIncid$new()

Initialize the shared KK marginal-incidence inference base, validate the binary matched/reservoir design, and prepare caches used by InferenceAbstractKKMarginalIncid.

Usage
InferenceAbstractKKMarginalIncid$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

A flag indicating whether messages should be displayed.

smart_cold_start_default

Whether to use smart cold start values.


InferenceAbstractKKMarginalIncid$clone()

The objects of this class are cloneable with this method.

Usage
InferenceAbstractKKMarginalIncid$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Abstract class for all-subject modified-Poisson inference in KK designs

Description

Abstract class for all-subject modified-Poisson inference in KK designs

Super classes

Inference -> InferenceAbstractKKMarginalIncid -> InferenceAbstractKKModifiedPoisson

Methods

Public methods

+ inherited public methods from InferenceAbstractKKMarginalIncid
+ inherited public methods from Inference

InferenceAbstractKKModifiedPoisson$compute_estimate()

Compute the KK marginal incidence treatment-effect estimate using the class-specific marginal risk-difference or risk-ratio estimator and cache it for related InferenceAbstractKKMarginalIncid methods.

Usage
InferenceAbstractKKModifiedPoisson$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

Logical. If TRUE, skip variance component calculations.


InferenceAbstractKKModifiedPoisson$compute_estimate_with_bootstrap_weights()

Recomputes the KK marginal incidence estimate under Bayesian-bootstrap weights.

Usage
InferenceAbstractKKModifiedPoisson$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Numeric vector. Row weights for bootstrap.

estimate_only

Logical. If TRUE, skip variance component calculations.


InferenceAbstractKKModifiedPoisson$compute_asymp_confidence_interval()

Compute the KK marginal incidence asymptotic confidence interval using the cached marginal effect and design-aware standard error. See InferenceAsymp.

Usage
InferenceAbstractKKModifiedPoisson$compute_asymp_confidence_interval(
  alpha = 0.05
)
Arguments
alpha

Numeric. Significance level (default 0.05).


InferenceAbstractKKModifiedPoisson$compute_asymp_two_sided_pval()

Compute the KK marginal incidence asymptotic two-sided p-value using the cached marginal effect and design-aware standard error. See InferenceAsymp.

Usage
InferenceAbstractKKModifiedPoisson$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Numeric. Null treatment effect value (default 0).


InferenceAbstractKKModifiedPoisson$clone()

The objects of this class are cloneable with this method.

Usage
InferenceAbstractKKModifiedPoisson$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Abstract class for ordinal CLMM-based Inference in KK designs

Description

Abstract class for ordinal CLMM-based Inference in KK designs

Super class

Inference -> InferenceAbstractKKOrdinalCLMM

Methods

Public methods

+ inherited public methods from Inference

InferenceAbstractKKOrdinalCLMM$new()

Initialize KK cumulative-link mixed-model inference for ordinal responses, validate the matched design, and prepare the ordinal likelihood used by InferenceAbstractKKOrdinalCLMM.

Usage
InferenceAbstractKKOrdinalCLMM$new(
  des_obj,
  model_formula = NULL,
  use_rcpp = TRUE,
  verbose = FALSE,
  harden = TRUE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object.

model_formula

Optional formula for covariate adjustment.

use_rcpp

Logical. If TRUE (default), use the internal Rcpp implementation (no external packages required). Set FALSE to fall back to ordinal::clmm.

verbose

A flag indicating whether messages should be displayed.

harden

Whether to apply robustness measures.

smart_cold_start_default

Whether to use smart cold start values.


InferenceAbstractKKOrdinalCLMM$compute_estimate()

Compute the ordinal CLMM treatment-effect estimate by fitting the cumulative-link mixed model and caching the treatment coefficient for related InferenceAsymp methods.

Usage
InferenceAbstractKKOrdinalCLMM$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

Logical. If TRUE, skip variance component calculations.


InferenceAbstractKKOrdinalCLMM$compute_estimate_with_bootstrap_weights()

Recomputes the KK ordinal CLMM treatment estimate under Bayesian-bootstrap weights.

Usage
InferenceAbstractKKOrdinalCLMM$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Numeric vector. Row weights for bootstrap.

estimate_only

Logical. If TRUE, skip variance component calculations.


InferenceAbstractKKOrdinalCLMM$compute_asymp_confidence_interval()

Compute the ordinal CLMM asymptotic confidence interval for the treatment coefficient using the fitted-model standard error. See InferenceAsymp.

Usage
InferenceAbstractKKOrdinalCLMM$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Numeric. Significance level (default 0.05).


InferenceAbstractKKOrdinalCLMM$compute_asymp_two_sided_pval()

Compute the ordinal CLMM asymptotic two-sided p-value for the treatment coefficient using the fitted-model standard error. See InferenceAsymp.

Usage
InferenceAbstractKKOrdinalCLMM$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Numeric. Null treatment effect value (default 0).


InferenceAbstractKKOrdinalCLMM$clone()

The objects of this class are cloneable with this method.

Usage
InferenceAbstractKKOrdinalCLMM$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Abstract mixin: Zhang combined randomisation CI for quantile regression

Description

Provides compute_rand_confidence_interval() via Zhang's combined test-inversion method for both Bernoulli (m = 0, all subjects in the reservoir) and KK matching-on-the-fly designs (m > 0).

Super classes

Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap -> InferenceJackknife -> InferenceAsymp -> InferenceMLEorKMSummaryTable -> InferenceAsympLik -> InferenceKKPassThroughCompoundNoParamBootstrap -> InferenceAbstractQuantileRandCI

Methods

Public methods

+ inherited public methods from InferenceKKPassThroughCompoundNoParamBootstrap
  • InferenceKKPassThroughCompoundNoParamBootstrap$approximate_bootstrap_distribution_beta_hat_T()
  • InferenceKKPassThroughCompoundNoParamBootstrap$compute_estimate_with_bootstrap_weights()
+ inherited public methods from InferenceAsympLik
+ inherited public methods from InferenceMLEorKMSummaryTable
+ inherited public methods from InferenceAsymp
+ inherited public methods from InferenceJackknife
+ inherited public methods from InferenceBayesianBootstrap
+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference

InferenceAbstractQuantileRandCI$new()

Initialize the inference object.

Usage
InferenceAbstractQuantileRandCI$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A DesignSeqOneByOne object.

model_formula

Optional formula for covariate adjustment.

verbose

Whether to print messages.

smart_cold_start_default

Whether to use smart cold start values.


InferenceAbstractQuantileRandCI$clone()

The objects of this class are cloneable with this method.

Usage
InferenceAbstractQuantileRandCI$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Mean-Difference IVWC Inference for KK Matching-on-the-Fly Designs

Description

Fits a compound (inverse-variance-weighted combination, "IVWC") mean-difference estimator of the treatment effect for continuous responses under a DesignSeqOneByOne-family KK matching-on-the-fly design (see DesignSeqOneByOneKK14 and DesignSeqOneByOneKK21). Such a design produces two structurally different kinds of subjects: subjects successfully matched into pairs during the sequential design, and unmatched "reservoir" subjects randomized independently. This estimator combines both:

\hat\beta_T = w^* \bar d + (1 - w^*)\, \bar r, \qquad w^* = \frac{\widehat{\mathrm{Var}}(\bar r)}{\widehat{\mathrm{Var}}(\bar r) + \widehat{\mathrm{Var}}(\bar d)},

where \bar d is the mean within-pair (treated minus control) difference among matched subjects and \bar r is the treated-minus-control difference in means among reservoir subjects, weighted inversely by their estimated variances (see $compute_asymp_confidence_interval() for the full variance formula and the fallback behavior when only one of the two sub-estimates is usable). Inference is Wald-only: this class has no likelihood tier (likelihood_tier = "none") and provides asymptotic Wald, randomization, and bootstrap (including Bayesian bootstrap) confidence intervals and p-values, but no score/likelihood-ratio/gradient tests.

Initialize KK IVWC mean-difference inference.

Computes the compound IVWC (inverse-variance-weighted compound) mean-difference point estimate \hat\beta_T: the inverse-variance-weighted combination w^* \bar d + (1-w^*)\, \bar r of the matched-pair mean within-pair difference \bar d and the reservoir treated-minus-control mean difference \bar r, falling back to whichever of the two is usable if the other is not (see $compute_asymp_confidence_interval() for the full weighting formula and usability conditions).

Computes a 1-\alpha level frequentist confidence interval for the compound IVWC (inverse-variance-weighted compound) mean-difference estimator \hat\beta_T.

Computes a two-sided Wald p-value for the compound IVWC mean-difference estimator \hat\beta_T testing H_0: \beta_T = \code{delta}, using the same asymptotically-normal point estimate and standard error (z = (\hat\beta_T - \code{delta})/\widehat{\mathrm{SE}}(\hat\beta_T)) that $compute_asymp_confidence_interval() inverts to form its interval — see that method's documentation for the full inverse-variance-weighted combination formula. This class has no likelihood tier (likelihood_tier = "none"), so no score, likelihood-ratio, or gradient test is available here; this is a plain Wald test, not a likelihood-backed one.

Details

The point estimate combines two sub-estimates depending on which are usable: the mean within-pair difference among matched subjects, \bar d, with estimated variance \widehat{\mathrm{Var}}(\bar d), and the treated-minus-control difference in means among reservoir (unmatched) subjects, \bar r, with estimated variance \widehat{\mathrm{Var}}(\bar r). When both are usable (at least 2 matched pairs and at least 2 treated/2 control reservoir subjects, with finite positive variance estimates), they are combined by classical inverse-variance weighting,

\hat\beta_T = w^* \bar d + (1 - w^*)\, \bar r, \qquad w^* = \frac{\widehat{\mathrm{Var}}(\bar r)}{\widehat{\mathrm{Var}}(\bar r) + \widehat{\mathrm{Var}}(\bar d)},

with combined variance the standard inverse-variance-pooled form \widehat{\mathrm{Var}}(\hat\beta_T) = \left(\widehat{\mathrm{Var}}(\bar r)^{-1} + \widehat{\mathrm{Var}}(\bar d)^{-1}\right)^{-1} = \widehat{\mathrm{Var}}(\bar r)\,\widehat{\mathrm{Var}}(\bar d) \big/ \left(\widehat{\mathrm{Var}}(\bar r) + \widehat{\mathrm{Var}}(\bar d)\right). If only one of the two sub-estimates is usable (e.g. the reservoir is empty or degenerate, or no pairs matched), \hat\beta_T and its variance fall back to that sub-estimate alone. The compound estimator is treated as asymptotically normal, so the interval is \hat\beta_T \pm z_{1-\alpha/2}\sqrt{\widehat{\mathrm{Var}}(\hat\beta_T)} (or a t-based critical value, depending on private$compute_z_or_t_ci_from_s_and_df's degrees-of-freedom resolution).

Value

The setting-appropriate (see description) numeric estimate of the treatment effect

A (1 - alpha)-sized frequentist confidence interval for the treatment effect

The approximate frequentist p-value

Legacy status

Legacy class. Not fully tested in comprehensive_tests.R; prefer a more actively maintained KK continuous-response inference class (e.g. InferenceContinKKOLSIVWC) for new analyses unless this specific unadjusted mean-difference estimator is required.

Super class

Inference -> InferenceAllKKMeanDiffIVWC

Methods

Public methods

+ inherited public methods from Inference

InferenceAllKKMeanDiffIVWC$new()

Usage
InferenceAllKKMeanDiffIVWC$new(
  des_obj,
  verbose = FALSE,
  harden = TRUE,
  model_formula = NULL,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A KK matching-on-the-fly design object.

verbose

Whether to print progress messages.

harden

Whether to use hardened model-matrix fitting.

model_formula

Optional formula for covariate adjustment.

smart_cold_start_default

Whether to use smart cold start values.


InferenceAllKKMeanDiffIVWC$compute_estimate()

Usage
InferenceAllKKMeanDiffIVWC$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, compute only the point estimate \hat\beta_T and skip the variance-component computations needed for confidence intervals or p-values (faster when only the point estimate is needed).


InferenceAllKKMeanDiffIVWC$compute_asymp_confidence_interval()

Usage
InferenceAllKKMeanDiffIVWC$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.


InferenceAllKKMeanDiffIVWC$compute_asymp_two_sided_pval()

Usage
InferenceAllKKMeanDiffIVWC$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null difference to test against. For any treatment effect at all this is set to zero (the default).


InferenceAllKKMeanDiffIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferenceAllKKMeanDiffIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kapelner, A., and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the KK matching-on-the-fly design this estimator targets, and for the inverse-variance combination of matched-pair and reservoir estimates.

Examples

seq_des = DesignSeqOneByOneKK14$new(n = 6, response_type = "continuous")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10])
seq_des$add_all_subject_responses(c(4.71, 1.23, 4.78, 6.11, 5.95, 8.43))

seq_des_inf = InferenceAllKKMeanDiffIVWC$
  new(seq_des)
seq_des_inf$compute_estimate()
seq_des_inf$compute_asymp_confidence_interval()
seq_des_inf$compute_asymp_two_sided_pval()

Non-parametric Wilcoxon-based Compound Inference for KK Matching-on-the-Fly Designs

Description

Fits a non-parametric, rank-based compound (inverse-variance-weighted, IVWC) estimator of the treatment effect under a DesignSeqOneByOne-family KK matching-on-the-fly design (see DesignSeqOneByOneKK14). For matched pairs, the sub-estimate \hat\beta_m is the Hodges-Lehmann estimate from a Wilcoxon signed-rank test on the within-pair differences (the median of the Walsh averages (d_i+d_j)/2); for reservoir (unmatched) subjects, \hat\beta_r is the Hodges-Lehmann estimate from a Wilcoxon rank-sum (Mann-Whitney U) test on treated-vs-control reservoir responses (the median of all pairwise differences). The two are combined by classical inverse-variance weighting,

\hat\beta_T = w^* \hat\beta_m + (1-w^*)\, \hat\beta_r, \qquad w^* = \frac{\widehat{\mathrm{Var}}(\hat\beta_r)}{\widehat{\mathrm{Var}}(\hat\beta_r) + \widehat{\mathrm{Var}}(\hat\beta_m)},

with variance the standard inverse-variance-pooled form (see $compute_estimate()'s method-level documentation for the full formula and fallback behavior when only one sub-estimate is usable). Because it is built from Hodges-Lehmann/Wilcoxon estimators rather than sample means, this method is robust to outliers and does not assume a specific parametric distribution for the response — but it does not currently support censored survival data or incidence (binary) responses (see $initialize()), and its jackknife methods all report explicit non-estimability rather than computing a (statistically unreliable) delete-1 jackknife of the Hodges-Lehmann functional.

Override to avoid O(n^2) per-resample HL computation during the bootstrap warm-start inside compute_rand_confidence_interval. The asymptotic MLE CI is a perfectly adequate starting bound for the bisection and is computed in O(1).

Initialize KK inverse-variance combined Wilcoxon inference and prepare matched/reservoir rank-based components used by InferenceAllKKWilcoxIVWC. Requires des_obj to be a KK matching-on-the-fly-capable design (des_obj$is_a_kk_matching_capable()); errors otherwise. Also rejects response_type = "incidence" (a message-and-error recommends a compound mean-difference or conditional-logistic estimator instead — rank-based methods are not well suited to binary outcomes) and rejects censored survival data (recommends a restricted-mean or Cox-based method instead, since this estimator has no censoring handling). Legal response_type values are "continuous", "count", "proportion", "survival" (uncensored only), and "ordinal".

Returns the estimated treatment effect: an inverse-variance-weighted compound (IVWC) of two Hodges-Lehmann median-shift estimates.

Computes a 1-\alpha level confidence interval for the compound Hodges-Lehmann treatment effect estimator. Although each sub-estimate is itself derived from a non-parametric rank test, the inverse-variance-weighted combination \hat\beta_T (see $compute_estimate() for the full formula) is treated as asymptotically normal, so the interval is \hat\beta_T \pm z_{1-\alpha/2}\sqrt{\widehat{\mathrm{Var}}(\hat\beta_T)} (or a t-based critical value, depending on private$compute_z_or_t_ci_from_s_and_df's degrees-of-freedom resolution).

Compute the KK Wilcoxon compound two-sided p-value testing H_0: \beta_T = \code{delta}, from the same asymptotically-normal compound estimate/variance (z = (\hat\beta_T - \code{delta})/\widehat{\mathrm{SE}}(\hat\beta_T)) that $compute_asymp_confidence_interval() inverts to form its interval — see that method's documentation, and $compute_estimate(), for the compound Hodges-Lehmann estimator's full formula. Only delta = 0 is currently supported: a non-zero null shift raises an error (when assertions are enabled) rather than testing it, because the underlying Wilcoxon tests' null-shift handling has not been extended to the compound combined estimator. See related simple Wilcoxon behavior in InferenceAllSimpleWilcox.

Reports the jackknife point-estimate as explicitly non-estimable for this compound Hodges-Lehmann estimator, rather than computing a leave-one-out jackknife. Deletion-based (jackknife) resampling of a Hodges-Lehmann/Wilcoxon-derived statistic is known to behave poorly — the median-of-Walsh-averages functional is not smooth enough for the delete-1 jackknife's linear-approximation machinery to be reliable at the small matched-pair/reservoir sample sizes typical of KK designs, and combining two already-jackknife-unstable sub-estimates compounds the problem. This method exists purely to record that unavailability (via private$cache_nonestimable_estimate()) rather than silently returning a misleading number; see InferenceJackknife for the shared jackknife contract this method participates in.

Reports the jackknife bias-correction estimate as non-estimable for this Wilcoxon compound estimator, for the same reason as $compute_jackknife_estimate() (the Hodges-Lehmann functional is not smooth enough for the delete-1 jackknife); see InferenceJackknife for the shared jackknife contract.

Reports the jackknife standard error as non-estimable for this Wilcoxon compound estimator, for the same reason as $compute_jackknife_estimate(); see InferenceJackknife for the shared jackknife contract.

Reports the jackknife-Wald p-value as non-estimable here, for the same reason as $compute_jackknife_estimate(); see InferenceJackknife.

Reports the jackknife-Wald confidence interval as non-estimable here, for the same reason as $compute_jackknife_estimate(); see InferenceJackknife.

Details

For matched pairs, \hat\beta_m is the Hodges-Lehmann estimate from a Wilcoxon signed-rank test on the within-pair differences (stats::wilcox.test(diffs, conf.int = TRUE)'s estimate, the median of the Walsh averages (d_i + d_j)/2), with variance estimated as the sample variance of those Walsh averages divided by the number of pairs m. For reservoir (unmatched) subjects, \hat\beta_r is the Hodges-Lehmann estimate from a Wilcoxon rank-sum test between treated and control reservoir responses (median of all pairwise differences y_{T,i} - y_{C,j}), with an analogous pairwise-difference-variance-based estimate. When both sub-estimates are usable, the compound estimate is the inverse-variance-weighted combination

\hat\beta_T = w^* \hat\beta_m + (1-w^*)\, \hat\beta_r, \qquad w^* = \frac{\widehat{\mathrm{Var}}(\hat\beta_r)}{\widehat{\mathrm{Var}}(\hat\beta_r) + \widehat{\mathrm{Var}}(\hat\beta_m)},

with combined variance \widehat{\mathrm{Var}}(\hat\beta_r)\, \widehat{\mathrm{Var}}(\hat\beta_m) / (\widehat{\mathrm{Var}}(\hat\beta_r) + \widehat{\mathrm{Var}}(\hat\beta_m)) — the same combination scheme as InferenceAllKKMeanDiffIVWC, but applied to rank-based rather than mean-based sub-estimates. If only one sub-estimate is usable (e.g. no matched pairs, or a degenerate reservoir), \hat\beta_T falls back to that sub-estimate alone.

Super class

Inference -> InferenceAllKKWilcoxIVWC

Methods

Public methods

+ inherited public methods from Inference

InferenceAllKKWilcoxIVWC$compute_bootstrap_confidence_interval()

Usage
InferenceAllKKWilcoxIVWC$compute_bootstrap_confidence_interval(
  alpha = 0.05,
  ...
)
Arguments
alpha

The confidence level. Default is 0.05.

alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.

alpha

Significance level. Default 0.05.

...

Additional arguments passed to super.


InferenceAllKKWilcoxIVWC$new()

Usage
InferenceAllKKWilcoxIVWC$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A DesignSeqOneByOne object (must be a KK design).

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.


InferenceAllKKWilcoxIVWC$compute_estimate()

Usage
InferenceAllKKWilcoxIVWC$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.


InferenceAllKKWilcoxIVWC$compute_asymp_confidence_interval()

Usage
InferenceAllKKWilcoxIVWC$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level. Default is 0.05.

alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.

alpha

Significance level. Default 0.05.


InferenceAllKKWilcoxIVWC$compute_asymp_two_sided_pval()

Usage
InferenceAllKKWilcoxIVWC$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null difference to test against. For any treatment effect at all this is set to zero (the default).

delta

Null treatment-effect value. Default 0.


InferenceAllKKWilcoxIVWC$compute_jackknife_estimate()

Usage
InferenceAllKKWilcoxIVWC$compute_jackknife_estimate(unit = "auto")
Arguments
unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceAllKKWilcoxIVWC$compute_jackknife_bias_estimate()

Usage
InferenceAllKKWilcoxIVWC$compute_jackknife_bias_estimate(unit = "auto")
Arguments
unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceAllKKWilcoxIVWC$compute_jackknife_std_error()

Usage
InferenceAllKKWilcoxIVWC$compute_jackknife_std_error(unit = "auto")
Arguments
unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceAllKKWilcoxIVWC$compute_jackknife_wald_two_sided_pval()

Usage
InferenceAllKKWilcoxIVWC$compute_jackknife_wald_two_sided_pval(
  delta = 0,
  unit = "auto"
)
Arguments
delta

The null difference to test against. For any treatment effect at all this is set to zero (the default).

delta

Null treatment-effect value. Default 0.

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceAllKKWilcoxIVWC$compute_jackknife_wald_confidence_interval()

Usage
InferenceAllKKWilcoxIVWC$compute_jackknife_wald_confidence_interval(
  alpha = 0.05,
  unit = "auto"
)
Arguments
alpha

The confidence level. Default is 0.05.

alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.

alpha

Significance level. Default 0.05.

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceAllKKWilcoxIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferenceAllKKWilcoxIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Hodges, J. L., and Lehmann, E. L. (1963). "Estimates of Location Based on Rank Tests." The Annals of Mathematical Statistics, 34(2), 598-611, doi:10.1214/aoms/1177704172, for the Hodges-Lehmann estimator underlying both sub-estimates; Kapelner, A., and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the KK matching-on-the-fly design and the inverse-variance combination of matched-pair and reservoir estimates.

Legacy class. Not fully tested in comprehensive_tests.R.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceAllKKWilcoxIVWC$new(seq_des)
inf$compute_estimate()


Simple Mean-Difference Inference for Continuous Responses

Description

Fits the simplest possible treatment-effect estimator for a continuous response: the unadjusted difference in sample means between the treated and control arms, \hat\beta_T = \bar y_T - \bar y_C, with no covariate adjustment. Inference is by Welch's unequal-variance t-test: standard error \sqrt{s_T^2/n_T + s_C^2/n_C} (per-arm sample variances, not pooled) with Satterthwaite-Welch degrees of freedom — see $compute_asymp_confidence_interval() for the exact formula. This class has no likelihood tier (likelihood_tier = "none") and provides asymptotic Wald, randomization, and bootstrap (including Bayesian bootstrap) confidence intervals and p-values. Warm starts are disabled for this class, since the simple mean difference is a closed-form estimator (no iterative fit to warm-start).

Initialize a simple mean-difference inference object.

Computes a 1-\alpha level confidence interval for the simple (unadjusted) mean-difference treatment effect \hat\beta_T = \bar y_T - \bar y_C, using Welch's unequal-variance formula: standard error \widehat{\mathrm{SE}}(\hat\beta_T) = \sqrt{s_T^2/n_T + s_C^2/n_C} (sample variances s_T^2, s_C^2 computed separately per arm, not pooled) with Satterthwaite-Welch degrees of freedom \mathrm{df} = (s_T^2/n_T + s_C^2/n_C)^2 \big/ \left(\frac{(s_T^2/n_T)^2}{n_T-1} + \frac{(s_C^2/n_C)^2}{n_C-1}\right); the interval is \hat\beta_T \pm t_{\mathrm{df}, 1-\alpha/2}\,\widehat{\mathrm{SE}}(\hat\beta_T). Requires at least 2 observations per arm; otherwise the standard error and interval are NA. See InferenceAsymp for the shared asymptotic confidence-interval contract this delegates to.

Computes a two-sided Welch's t-test p-value testing H_0: \beta_T = \code{delta}, from the same Welch unequal-variance standard error and Satterthwaite-Welch degrees of freedom used by $compute_asymp_confidence_interval() — see that method's documentation for the full formula. See InferenceAsymp for the shared asymptotic two-sided p-value contract this delegates to.

Computes the simple (unadjusted) mean-difference point estimate \hat\beta_T = \bar y_T - \bar y_C, the difference in sample means between the treated and control arms. NA if either arm has zero observations. See InferenceMLEorKMSummaryTable for the shared estimate-contract this participates in.

Recomputes the simple mean-difference estimate under subject/block bootstrap weights (used by the Bayesian bootstrap and related weighted-resampling machinery — see InferenceNonParamBootstrap). The weighted point estimate is \hat\beta_T = \bar y_T^w - \bar y_C^w, weighted arm means \bar y_T^w = \sum_i r_i y_i \mathbb{1}[w_i=1] / \sum_i r_i \mathbb{1}[w_i=1] (and analogously for control), where r_i are the expanded row weights. Unless estimate_only = TRUE, the standard error uses a weighted, effective-sample-size Welch formula: n_{\mathrm{eff}} = (\sum r_i)^2 / \sum r_i^2 (the usual Kish effective-sample-size correction for unequal weights) in place of the raw n in both the per-arm weighted variance denominator and the Satterthwaite-Welch degrees-of-freedom formula (see $compute_asymp_confidence_interval() for the unweighted version of the same formula). Rows with non-finite or non-positive weight, or a non-finite response, are dropped before computing; if no rows survive, returns NA with all cached variance components set to NA.

Value

A two-sided p-value.

The setting-appropriate (see description) numeric estimate of the treatment effect

Super class

Inference -> InferenceAllSimpleAverageDiff

Methods

Public methods

+ inherited public methods from Inference

InferenceAllSimpleAverageDiff$compute_rand_two_sided_pval()

Uses the randomization-CI layer's two-sided p-value contract (InferenceRandCI's version, not InferenceRand's): for incidence responses this dispatches to the Zhang exact randomization test where applicable rather than refusing outright, matching this class's pre-migration old-ladder behavior (see InferenceIncidRiskDiff's identical rationale). Previously bound to InferenceRand's version instead, which silently regressed Zhang dispatch after migration – see inference_all_abstract_rand_ci.R's compute_rand_two_sided_pval for why it's now safe to splice this in outside the old inheritance chain.

Usage
InferenceAllSimpleAverageDiff$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  type = NULL,
  args_for_type = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors.

delta

Null treatment effect value.

delta

Null treatment effect value.

transform_responses

Response transformation to apply during the test. For survival responses the default "log" multiplies the recorded times of the units treated under each reference allocation by e^\delta, event and censoring times alike, with censoring indicators unchanged – the rank-based AFT residual construction (Tsiatis 1990; Wei, Ying and Lin 1990; Jin, Lin, Wei and Ying 2003); see compute_rand_confidence_interval() for the assumptions.

na.rm

Whether to remove non-finite simulated statistics.

show_progress

Whether to show progress.

permutations

Optional pre-generated assignment draws.

type

Optional incidence-specific exact randomization type.

args_for_type

Optional arguments keyed by type.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceAllSimpleAverageDiff$new()

Usage
InferenceAllSimpleAverageDiff$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  max_resample_attempts = 50L,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A DesignSeqOneByOne object whose entire n subjects are assigned and response y is recorded within.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages. Default FALSE.

max_resample_attempts

Maximum number of times a single bootstrap replicate may be redrawn when the drawn sample fails validity screening. If all attempts fail the replicate is recorded as NA, silently reducing the effective B. Must be a positive integer. Default 50L.

smart_cold_start_default

Whether to use smart cold start values.


InferenceAllSimpleAverageDiff$compute_asymp_confidence_interval()

Usage
InferenceAllSimpleAverageDiff$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Confidence level.


InferenceAllSimpleAverageDiff$compute_asymp_two_sided_pval()

Usage
InferenceAllSimpleAverageDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect value.

delta

Null treatment effect value.


InferenceAllSimpleAverageDiff$compute_estimate()

Usage
InferenceAllSimpleAverageDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.

estimate_only

If TRUE, skip variance calculations.


InferenceAllSimpleAverageDiff$compute_estimate_with_bootstrap_weights()

Usage
InferenceAllSimpleAverageDiff$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Row weights for the bootstrap sample.

estimate_only

If TRUE, skip variance component calculations.

estimate_only

If TRUE, skip variance calculations.


InferenceAllSimpleAverageDiff$clone()

The objects of this class are cloneable with this method.

Usage
InferenceAllSimpleAverageDiff$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Welch, B. L. (1947). "The Generalization of 'Student's' Problem when Several Different Population Variances are Involved." Biometrika, 34(1-2), 28-35, doi:10.1093/biomet/34.1-2.28, for the unequal-variance t-test and its Satterthwaite-Welch degrees-of-freedom approximation used here.

Examples

seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "continuous")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10])
seq_des$add_all_subject_responses(c(4.71, 1.23, 4.78, 6.11, 5.95, 8.43))

seq_des_inf = InferenceAllSimpleAverageDiff$new(seq_des)
seq_des_inf$compute_estimate()
seq_des_inf$compute_asymp_confidence_interval()
seq_des_inf$compute_asymp_two_sided_pval()

Simple Mean-Difference Inference with Pooled Variance

Description

Fits the same unadjusted mean-difference point estimate as InferenceAllSimpleAverageDiff, \hat\beta_T = \bar y_T - \bar y_C, but performs inference via the classical pooled equal-variance Student's t-test instead of Welch's unequal-variance version: pooled variance s_p^2 = \left((n_T-1)s_T^2 + (n_C-1)s_C^2\right)/(n_T+n_C-2), standard error s_p\sqrt{1/n_T + 1/n_C}, and exact degrees of freedom n_T+n_C-2 — see $compute_asymp_confidence_interval() for the full formula. This assumes the two arms have equal population variance; prefer InferenceAllSimpleAverageDiff when that assumption is doubtful, since the pooled estimator's nominal coverage degrades under heteroskedasticity with unequal arm sizes. This class does not support censored survival data (enforced at construction). This class has no likelihood tier (likelihood_tier = "none") and provides asymptotic Wald, randomization, and bootstrap (including Bayesian bootstrap) confidence intervals and p-values. Warm starts are disabled for this class, since the simple mean difference is a closed-form estimator (no iterative fit to warm-start).

Initialize simple pooled-variance mean-difference inference for continuous responses and prepare the pooled standard-error calculation used by InferenceAllSimpleMeanDiffPooledVar. Disables warm starts (closed-form estimator) and asserts des_obj has no censored observations (unsupported by this class).

Computes a 1-\alpha level confidence interval for the simple (unadjusted) mean-difference treatment effect \hat\beta_T = \bar y_T - \bar y_C, using the classical pooled equal-variance Student's t-test formula (unlike InferenceAllSimpleAverageDiff's Welch unequal-variance version): the pooled variance estimate s_p^2 = \left((n_T-1)s_T^2 + (n_C-1)s_C^2\right) / (n_T+n_C-2) gives standard error \widehat{\mathrm{SE}}(\hat\beta_T) = s_p\sqrt{1/n_T + 1/n_C} with exact degrees of freedom n_T + n_C - 2; the interval is \hat\beta_T \pm t_{\mathrm{df}, 1-\alpha/2}\, \widehat{\mathrm{SE}}(\hat\beta_T). Assumes equal population variances in the two arms — use InferenceAllSimpleAverageDiff instead when that assumption is doubtful. Requires at least 2 observations per arm; otherwise returns c(NA, NA). See InferenceAsymp for the shared asymptotic confidence-interval contract this participates in.

Computes a two-sided pooled-variance Student's t-test p-value testing H_0: \beta_T = \code{delta}, from the same pooled standard error and exact n_T+n_C-2 degrees of freedom used by $compute_asymp_confidence_interval() — see that method's documentation for the full formula. See InferenceAsymp for the shared asymptotic two-sided p-value contract this participates in.

Value

A new InferenceAllSimpleMeanDiffPooledVar object.

Super class

Inference -> InferenceAllSimpleMeanDiffPooledVar

Methods

Public methods

+ inherited public methods from Inference

InferenceAllSimpleMeanDiffPooledVar$new()

Usage
InferenceAllSimpleMeanDiffPooledVar$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed design object.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.


InferenceAllSimpleMeanDiffPooledVar$compute_asymp_confidence_interval()

Usage
InferenceAllSimpleMeanDiffPooledVar$compute_asymp_confidence_interval(
  alpha = 0.05
)
Arguments
alpha

Confidence level.


InferenceAllSimpleMeanDiffPooledVar$compute_asymp_two_sided_pval()

Usage
InferenceAllSimpleMeanDiffPooledVar$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect value.


InferenceAllSimpleMeanDiffPooledVar$clone()

The objects of this class are cloneable with this method.

Usage
InferenceAllSimpleMeanDiffPooledVar$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Student [Gosset, W. S.] (1908). "The Probable Error of a Mean." Biometrika, 6(1), 1-25, doi:10.1093/biomet/6.1.1, for the pooled-variance two-sample t-test used here.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceAllSimpleMeanDiffPooledVar$new(seq_des)
inf$compute_estimate()


Simple Wilcoxon Rank-Sum (Hodges-Lehmann) Inference

Description

Fits a non-parametric treatment-effect estimator based on the two-sample Wilcoxon rank-sum test: the point estimate is the Hodges-Lehmann location-shift estimate (the median of all pairwise treatment-minus-control differences y_{T,i} - y_{C,j}), and both the confidence interval and two-sided p-value are the standard rank-based Wilcoxon quantities from stats::wilcox.test() (normal approximation with continuity correction), not Wald intervals/tests built around the point estimate and a separately estimated standard error. Robust to outliers and does not assume normality or equal arm variances. Not supported for incidence (binary) responses (the Hodges-Lehmann estimator degenerates on 0/1 data — use InferenceAllSimpleAverageDiff or a conditional-logistic estimator instead) or censored survival data (use InferenceSurvivalGehanWilcox instead). This class has no likelihood tier (likelihood_tier = "none") and does not support the Bayesian bootstrap; its jackknife methods all report explicit non-estimability rather than computing a (statistically unreliable) delete-1 jackknife of the Hodges-Lehmann functional.

Initialize simple Wilcoxon inference and prepare the rank-based treatment statistic used by InferenceAllSimpleWilcox. Rejects response_type = "incidence" (Hodges-Lehmann degenerates on binary data) and rejects censored survival data at construction; see the class-level documentation for recommended alternatives in both cases. Legal response_type values are "continuous", "count", "proportion", "survival" (uncensored only), and "ordinal".

Returns the Hodges-Lehmann estimate of location shift: the median of all pairwise treatment-minus-control differences y_{T,i} - y_{C,j} (via wilcox_hl_point_estimate_cpp()), the standard point estimate associated with the Wilcoxon rank-sum test. Robust to outliers and does not assume normality or equal variances.

Wilcoxon rank-sum test two-sided p-value testing H_0: \beta_T = \code{delta} (via stats::wilcox.test(yT, yC - delta, exact = FALSE)$p.value, the normal approximation with continuity correction) — a genuine rank-based test, not a Wald test built from the Hodges-Lehmann estimate and its standard error, despite living alongside $compute_asymp_confidence_interval() in this class's "asymptotic" method family. For delta != 0, the control arm's values are shifted by delta before testing, so the test checks whether y_T and y_C + \code{delta} come from the same distribution.

Returns the Hodges-Lehmann confidence interval directly from stats::wilcox.test(yT, yC, conf.int = TRUE, exact = FALSE, conf.level = 1 - alpha) — the standard nonparametric interval associated with the Wilcoxon rank-sum test, based on inverting the rank-sum test statistic rather than a Wald normal-approximation interval around $compute_estimate()'s point estimate (though the two coincide asymptotically).

Delegates to the genuine rank-based $compute_asymp_two_sided_pval() rather than the generic Wald-component z/t formula.

Fixed 2026-09-06: this class did not override compute_wald_two_sided_pval, so it fell through to the composed Wald component's generic (estimate - delta) / se formula built from compute_estimate() (the Hodges-Lehmann median-of-pairwise- differences) and get_standard_error(). On heavily tied, small-integer count/ordinal data the Hodges-Lehmann estimate lands on exactly 0 far more often than a continuous estimator would, so the Wald statistic came out exactly 0/se = 0 regardless of se, forcing p = 1 deterministically (observed: pinned at 1 in ~75-98 compute_asymp_two_sided_pval() does not have this failure mode.

Delegates to the genuine rank-based $compute_asymp_confidence_interval() rather than the generic Wald normal-approximation interval, for the same reason as compute_wald_two_sided_pval above.

Fixed 2026-09-06: the generic Wald component's normal-approximation interval is built from get_standard_error(), which this class derives by back-solving se = (ci[2]-ci[1]) / (2*1.96) from stats::wilcox.test()'s own asymptotic CI width. Under heavy ties, that root search can converge to a numerically near-zero-width interval as a search artifact, not a real sampling-uncertainty statement; that spurious near-zero SE then produced a near-[0,0] Wald interval (observed in over 1,200 rows of comprehensive-results data). The rank-based compute_asymp_confidence_interval() inverts the rank-sum test directly and does not go through this derived SE at all.

Reports the jackknife point-estimate as explicitly non-estimable for this Hodges-Lehmann estimator, rather than computing a leave-one-out jackknife: the median-of-pairwise-differences functional is not smooth enough for the delete-1 jackknife's linear-approximation machinery to be reliable. This method exists purely to record that unavailability (via private$cache_nonestimable_estimate()) rather than silently returning a misleading number; see InferenceJackknife for the shared jackknife contract this method participates in.

Reports the jackknife bias-correction estimate as non-estimable for this simple Wilcoxon estimator, for the same reason as $compute_jackknife_estimate() (the Hodges-Lehmann functional is not smooth enough for the delete-1 jackknife); see InferenceJackknife for the shared jackknife contract.

Reports the jackknife standard error as non-estimable for this simple Wilcoxon estimator, for the same reason as $compute_jackknife_estimate(); see InferenceJackknife for the shared jackknife contract.

Reports the jackknife-Wald p-value as non-estimable here, for the same reason as $compute_jackknife_estimate(); see InferenceJackknife.

Reports the jackknife-Wald confidence interval as non-estimable here, for the same reason as $compute_jackknife_estimate(); see InferenceJackknife.

Super class

Inference -> InferenceAllSimpleWilcox

Methods

Public methods

+ inherited public methods from Inference

InferenceAllSimpleWilcox$new()

Usage
InferenceAllSimpleWilcox$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  max_resample_attempts = 50L,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed DesignSeqOneByOne object.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages. Default FALSE.

max_resample_attempts

Maximum number of times a single bootstrap replicate may be redrawn when the drawn sample fails validity screening. If all attempts fail the replicate is recorded as NA, silently reducing the effective B. Must be a positive integer. Default 50L.

smart_cold_start_default

Flag for consistent API.


InferenceAllSimpleWilcox$compute_estimate()

Usage
InferenceAllSimpleWilcox$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.


InferenceAllSimpleWilcox$compute_asymp_two_sided_pval()

Usage
InferenceAllSimpleWilcox$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect. Default 0.

delta

Null treatment effect. Default 0.

delta

Null treatment-effect value. Default 0.


InferenceAllSimpleWilcox$compute_asymp_confidence_interval()

Usage
InferenceAllSimpleWilcox$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.


InferenceAllSimpleWilcox$compute_wald_two_sided_pval()

Usage
InferenceAllSimpleWilcox$compute_wald_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect. Default 0.

delta

Null treatment effect. Default 0.

delta

Null treatment-effect value. Default 0.


InferenceAllSimpleWilcox$compute_wald_confidence_interval()

Usage
InferenceAllSimpleWilcox$compute_wald_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.


InferenceAllSimpleWilcox$compute_jackknife_estimate()

Usage
InferenceAllSimpleWilcox$compute_jackknife_estimate(unit = "auto")
Arguments
unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceAllSimpleWilcox$compute_jackknife_bias_estimate()

Usage
InferenceAllSimpleWilcox$compute_jackknife_bias_estimate(unit = "auto")
Arguments
unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceAllSimpleWilcox$compute_jackknife_std_error()

Usage
InferenceAllSimpleWilcox$compute_jackknife_std_error(unit = "auto")
Arguments
unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceAllSimpleWilcox$compute_jackknife_wald_two_sided_pval()

Usage
InferenceAllSimpleWilcox$compute_jackknife_wald_two_sided_pval(
  delta = 0,
  unit = "auto"
)
Arguments
delta

Null treatment effect. Default 0.

delta

Null treatment effect. Default 0.

delta

Null treatment-effect value. Default 0.

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceAllSimpleWilcox$compute_jackknife_wald_confidence_interval()

Usage
InferenceAllSimpleWilcox$compute_jackknife_wald_confidence_interval(
  alpha = 0.05,
  unit = "auto"
)
Arguments
alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceAllSimpleWilcox$clone()

The objects of this class are cloneable with this method.

Usage
InferenceAllSimpleWilcox$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Hodges, J. L., and Lehmann, E. L. (1963). "Estimates of Location Based on Rank Tests." The Annals of Mathematical Statistics, 34(2), 598-611, doi:10.1214/aoms/1177704172, for the Hodges-Lehmann estimator; Wilcoxon, F. (1945). "Individual Comparisons by Ranking Methods." Biometrics Bulletin, 1(6), 80-83, doi:10.2307/3001968, for the underlying rank-sum test.

Examples

seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "continuous")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10])
seq_des$add_all_subject_responses(c(4.71, 1.23, 4.78, 6.11, 5.95, 8.43))

seq_des_inf = InferenceAllSimpleWilcox$new(seq_des)
seq_des_inf$compute_estimate()

Asymptotic Inference

Description

Abstract class for asymptotic inference.

Super classes

Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap -> InferenceJackknife -> InferenceAsymp

Methods

Public methods

+ inherited public methods from InferenceJackknife
+ inherited public methods from InferenceBayesianBootstrap
+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference

InferenceAsymp$compute_asymp_confidence_interval()

Computes an asymptotic confidence interval for the treatment effect using the configured large-sample test. For the default Wald path, the method first calls compute_estimate(), retrieves the class-specific standard error, and forms a normal or t interval around the estimate. Likelihood-backed subclasses may override the dispatch; see InferenceAsympLik.

Usage
InferenceAsymp$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level 1 - alpha. Default 0.05.

Returns

A confidence interval.


InferenceAsymp$compute_asymp_two_sided_pval()

Computes an asymptotic two-sided p-value for the treatment effect using the configured large-sample test. For the default Wald path, the method compares compute_estimate() to the null value delta using the class-specific standard error and a normal or t reference distribution. Likelihood-backed subclasses may override the dispatch; see InferenceAsympLik.

Usage
InferenceAsymp$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect to test against. Default 0.

Returns

The asymptotic p-value.


InferenceAsymp$get_supported_testing_types()

Gets the asymptotic testing methods supported by this inference object.

Usage
InferenceAsymp$get_supported_testing_types()

InferenceAsymp$set_testing_type()

Sets the asymptotic testing method used by p-values and CIs. This base (Wald-only) implementation accepts only "wald" and rejects everything else with a clear message; likelihood-tier classes override this with a richer version supporting score/gradient/lik_ratio testing types (see InferenceAsympLik). Without this base method, a Wald-only class (one composing only the Wald component, e.g. a robust-sandwich or Bai-adjusted-t estimator) has no set_testing_type() at all, so calling it fails with an opaque "attempt to apply non-function" instead of a clear rejection.

Usage
InferenceAsymp$set_testing_type(testing_type = "wald")
Arguments
testing_type

One of "wald" for this base implementation (likelihood-tier subclasses accept more values).

Returns

The inference object, invisibly.


InferenceAsymp$compute_wald_two_sided_pval()

Computes the Wald two-sided p-value regardless of configured testing type. This directly uses the treatment estimate, its standard error, and the available degrees of freedom; compare with compute_asymp_two_sided_pval() for configured-test dispatch.

Usage
InferenceAsymp$compute_wald_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect.


InferenceAsymp$compute_wald_confidence_interval()

Computes the Wald confidence interval regardless of configured testing type. This directly uses the treatment estimate, its standard error, and the available degrees of freedom; compare with compute_asymp_confidence_interval() for configured-test dispatch.

Usage
InferenceAsymp$compute_wald_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level. Default 0.05.


InferenceAsymp$compute_estimate()

Abstract method to compute the treatment-effect estimate. Concrete subclasses implement the model-specific calculation, such as an MLE coefficient, estimating-equation coefficient, standardized contrast, or rank/statistic-based treatment effect. Related p-value and interval methods call this method before using class-specific uncertainty estimates.

Usage
InferenceAsymp$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.

Returns

A scalar treatment estimate.


InferenceAsymp$get_mod()

Returns the model object from the last call that produced the treatment estimate and SE. Calls compute_estimate() first if needed.

Usage
InferenceAsymp$get_mod()
Returns

The cached model object (type depends on the concrete class).


InferenceAsymp$get_summary()

Prints a summary of the model from the last call that produced the treatment estimate and SE.

Usage
InferenceAsymp$get_summary()

InferenceAsymp$clone()

The objects of this class are cloneable with this method.

Usage
InferenceAsymp$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Likelihood-Backed Asymptotic Inference

Description

Intermediate base class for asymptotic inference families that expose likelihood / partial-likelihood / working-likelihood test paths in addition to Wald inference. The term "likelihood" is used broadly: subclasses may be backed by a true full likelihood, a partial likelihood (e.g. Cox PH), a quasi-likelihood (e.g. GEE, quasi-Poisson), or a composite/combined likelihood. Classes requiring a full generative likelihood — i.e. those supporting parametric-bootstrap LR calibration — inherit instead from InferenceParamBootstrap.

Super classes

Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap -> InferenceJackknife -> InferenceAsymp -> InferenceMLEorKMSummaryTable -> InferenceAsympLik

Methods

Public methods

+ inherited public methods from InferenceMLEorKMSummaryTable
+ inherited public methods from InferenceAsymp
+ inherited public methods from InferenceJackknife
+ inherited public methods from InferenceBayesianBootstrap
+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference

InferenceAsympLik$compute_asymp_confidence_interval()

Computes an asymptotic confidence interval for the treatment effect using the configured likelihood-backed test. Wald intervals use the fitted estimate and standard error; score, likelihood-ratio, gradient, and Bartlett-corrected likelihood-ratio intervals are obtained by inverting the corresponding test. For purely Wald asymptotic dispatch, see InferenceAsymp.

Usage
InferenceAsympLik$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level 1 - alpha. Default 0.05.

Returns

A confidence interval.


InferenceAsympLik$compute_asymp_two_sided_pval()

Computes an asymptotic two-sided p-value for the treatment effect using the configured likelihood-backed test. Depending on testing_type, this evaluates a Wald, score, likelihood-ratio, gradient, or Bartlett-corrected likelihood-ratio statistic under the null value delta. For count-specific likelihood families, see InferenceCountLikelihood.

Usage
InferenceAsympLik$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect to test against. Default 0.

Returns

The asymptotic p-value.


InferenceAsympLik$set_testing_type()

Sets the asymptotic testing method used by p-values and CIs.

Usage
InferenceAsympLik$set_testing_type(
  testing_type = c("wald", "score", "gradient", "lik_ratio", "lik_ratio_bartlett_approx",
    "lik_ratio_bartlett_exact")
)
Arguments
testing_type

One of "wald", "score", "gradient", "lik_ratio", "lik_ratio_bartlett_approx", or "lik_ratio_bartlett_exact".

Returns

The inference object, invisibly.


InferenceAsympLik$set_information_preference()

Sets the information matrix preference used by score-test dispatch.

Usage
InferenceAsympLik$set_information_preference(
  information_preference = c("auto", "fisher", "observed")
)
Arguments
information_preference

One of "auto", "fisher", or "observed".

Returns

The inference object, invisibly.


InferenceAsympLik$get_testing_type()

Gets the asymptotic testing method used by p-values and CIs.

Usage
InferenceAsympLik$get_testing_type()

InferenceAsympLik$get_information_preference()

Gets the score-test information matrix preference.

Usage
InferenceAsympLik$get_information_preference()

InferenceAsympLik$get_information_source_used()

Gets the actual information source used by the most recent information-backed computation.

Usage
InferenceAsympLik$get_information_source_used()

InferenceAsympLik$get_supported_testing_types()

Gets the asymptotic testing methods supported by this inference object.

Usage
InferenceAsympLik$get_supported_testing_types()

InferenceAsympLik$get_supported_information_preferences()

Gets the score-test information matrix preferences supported by this inference object.

Usage
InferenceAsympLik$get_supported_information_preferences()

InferenceAsympLik$compute_score_two_sided_pval()

Computes the score two-sided p-value regardless of configured testing type. The score test evaluates the null-restricted fit and uses the configured information matrix preference; compare with compute_asymp_two_sided_pval() for configured-test dispatch.

Usage
InferenceAsympLik$compute_score_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect.


InferenceAsympLik$compute_score_confidence_interval()

Computes the score confidence interval regardless of configured testing type by inverting score-test p-values over candidate treatment effects. Compare with compute_asymp_confidence_interval() for configured-test dispatch.

Usage
InferenceAsympLik$compute_score_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level. Default 0.05.


InferenceAsympLik$compute_lik_ratio_two_sided_pval()

Computes the likelihood-ratio two-sided p-value regardless of configured testing type. This compares unrestricted and null-restricted fits at delta; subclasses may use a full, partial, quasi-, or composite likelihood as described in InferenceAsympLik.

Usage
InferenceAsympLik$compute_lik_ratio_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect.


InferenceAsympLik$compute_lik_ratio_confidence_interval()

Computes the likelihood-ratio confidence interval regardless of configured testing type by inverting likelihood-ratio p-values over candidate treatment effects. Compare with score and gradient intervals in this class.

Usage
InferenceAsympLik$compute_lik_ratio_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level. Default 0.05.


InferenceAsympLik$compute_lik_ratio_bartlett_approx_two_sided_pval()

Computes the approximate (Monte-Carlo) Bartlett-corrected likelihood-ratio two-sided p-value regardless of configured testing type. Returns NA_real_ for subclasses that do not implement an approximate Bartlett correction factor.

The approximate Bartlett factor is estimated by Monte Carlo (e.g. the generic InferenceParamBootstrap factor): B datasets are simulated under the null-restricted fit at delta and refit to approximate E[LR | H0], the quantity a classical analytic Bartlett correction targets exactly. The Monte-Carlo draws are seeded from this object's own seed (see set_seed()), so repeated calls at the same delta with the same B are reproducible; there is no separate seed argument here.

See compute_lik_ratio_bartlett_exact_two_sided_pval() for the closed-form analytic counterpart (no simulation, no B).

Usage
InferenceAsympLik$compute_lik_ratio_bartlett_approx_two_sided_pval(
  delta = 0,
  B = 99
)
Arguments
delta

Null treatment effect. Default 0.

B

Number of Monte-Carlo replicates used to estimate the Bartlett factor. Default 99.


InferenceAsympLik$compute_lik_ratio_bartlett_approx_confidence_interval()

Computes the approximate (Monte-Carlo) Bartlett-corrected likelihood-ratio confidence interval regardless of configured testing type. Returns c(NA_real_, NA_real_) for subclasses that do not implement an approximate Bartlett correction factor.

See compute_lik_ratio_bartlett_approx_two_sided_pval() for what B controls and how the Monte-Carlo seed is inherited from this object's own seed. Each p-value evaluation during the confidence-interval search re-simulates B replicates, so this can be substantially more expensive than the p-value alone.

Usage
InferenceAsympLik$compute_lik_ratio_bartlett_approx_confidence_interval(
  alpha = 0.05,
  B = 99
)
Arguments
alpha

Significance level. Default 0.05.

B

Number of Monte-Carlo replicates used to estimate the Bartlett factor. Default 99.


InferenceAsympLik$compute_lik_ratio_bartlett_exact_two_sided_pval()

Computes the exact (closed-form analytic) Bartlett-corrected likelihood-ratio two-sided p-value regardless of configured testing type. Returns NA_real_ for subclasses that do not implement an exact, bespoke analytic Bartlett correction factor. Concrete families opt in through get_bartlett_factor_exact().

Unlike compute_lik_ratio_bartlett_approx_two_sided_pval(), this path involves no simulation and no Monte-Carlo replicate count.

Usage
InferenceAsympLik$compute_lik_ratio_bartlett_exact_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect. Default 0.


InferenceAsympLik$compute_lik_ratio_bartlett_exact_confidence_interval()

Computes the exact (closed-form analytic) Bartlett-corrected likelihood-ratio confidence interval regardless of configured testing type. Returns c(NA_real_, NA_real_) for subclasses that do not implement an exact, bespoke analytic Bartlett correction factor.

Usage
InferenceAsympLik$compute_lik_ratio_bartlett_exact_confidence_interval(
  alpha = 0.05
)
Arguments
alpha

Significance level. Default 0.05.


InferenceAsympLik$compute_lik_ratio_bartlett_two_sided_pval()

Computes "the best available" Bartlett-corrected likelihood-ratio two-sided p-value regardless of configured testing type: uses the exact (closed-form analytic) factor if this class implements one, otherwise falls back to the approximate (Monte-Carlo) factor. Errors if the class supports neither (see supports_bartlett_likelihood_ratio_exact()/ supports_bartlett_likelihood_ratio_approx()).

This is a convenience entry point for callers who want a Bartlett-corrected p-value without caring which mechanism produced it. Because exact and approximate factors are computed differently (deterministic closed form vs. seeded Monte-Carlo simulation), the same call can silently start returning different numeric results on a future package version once a family gains an exact implementation where previously only the approximate path existed. Callers who need results stable across package versions (e.g. for reproducibility or regression tests) should call compute_lik_ratio_bartlett_approx_two_sided_pval() or compute_lik_ratio_bartlett_exact_two_sided_pval() directly instead.

Usage
InferenceAsympLik$compute_lik_ratio_bartlett_two_sided_pval(delta = 0, B = 99)
Arguments
delta

Null treatment effect. Default 0.

B

Number of Monte-Carlo replicates, used only when the exact factor is unavailable and the approximate factor is used instead. If explicitly supplied but the exact factor is used (so B has no effect), a warning is issued; B left at its default is silently ignored in that case. Default 99.


InferenceAsympLik$compute_lik_ratio_bartlett_confidence_interval()

Computes "the best available" Bartlett-corrected likelihood-ratio confidence interval regardless of configured testing type: uses the exact (closed-form analytic) factor if this class implements one, otherwise falls back to the approximate (Monte-Carlo) factor. Errors if the class supports neither.

See compute_lik_ratio_bartlett_two_sided_pval() for the exact-over-approx selection rule, the B-ignored warning behavior, and why callers who need version-to-version reproducibility should prefer the explicit _approx/_exact methods instead.

Usage
InferenceAsympLik$compute_lik_ratio_bartlett_confidence_interval(
  alpha = 0.05,
  B = 99
)
Arguments
alpha

Significance level. Default 0.05.

B

Number of Monte-Carlo replicates, used only when the exact factor is unavailable and the approximate factor is used instead. Default 99.


InferenceAsympLik$compute_gradient_two_sided_pval()

Computes the gradient two-sided p-value regardless of configured testing type.

Usage
InferenceAsympLik$compute_gradient_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect.


InferenceAsympLik$compute_gradient_confidence_interval()

Computes the gradient confidence interval regardless of configured testing type.

Usage
InferenceAsympLik$compute_gradient_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level. Default 0.05.


InferenceAsympLik$clone()

The objects of this class are cloneable with this method.

Usage
InferenceAsympLik$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Bai Adjusted-t Mean-Difference Inference for KK14 Designs

Description

Continuous-response mean-difference inference for designs assigned by DesignSeqOneByOneKK14 (the Kapelner-Krieger 2014 sequential matching-on-the-fly design). The point estimate and its variance are the closed-form Bai-adjusted-t combination of the matched-pairs mean difference and the unmatched-reservoir mean difference, inverse-variance-weighted when both are usable; the full formula, pair-distance definition, and confidence-interval/p-value construction are shared with InferenceBaiAdjustedTKK21. The two leaves differ only in how pair distance is defined during matching: this class (KK14) uses the plain squared Euclidean distance \sum_j (x_{1j} - x_{2j})^2 between candidate subjects' covariate vectors, unlike KK21's covariate-weighted distance. Because the estimator is closed-form, initialization does not use warm starts (there is no iterative fit to warm-start).

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

Legacy class. Not fully tested in comprehensive_tests.R.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceBaiAdjustedTKK14

Methods

Public methods

+ inherited public methods from Inference

InferenceBaiAdjustedTKK14$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceBaiAdjustedTKK14$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceBaiAdjustedTKK14$supports_rand_pval_for_incidence()

Usage
InferenceBaiAdjustedTKK14$supports_rand_pval_for_incidence()

InferenceBaiAdjustedTKK14$compute_rand_two_sided_pval()

Usage
InferenceBaiAdjustedTKK14$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceBaiAdjustedTKK14$clone()

The objects of this class are cloneable with this method.

Usage
InferenceBaiAdjustedTKK14$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceBaiAdjustedTKK14$new(seq_des)
inf$compute_estimate()


Bai Adjusted-t Mean-Difference Inference for KK21 Designs

Description

Continuous-response mean-difference inference for designs assigned by DesignSeqOneByOneKK21 (the Kapelner-Krieger 2021 sequential matching-on-the-fly design with covariate-weighted matching). The point estimate and its variance are the closed-form Bai-adjusted-t combination of the matched-pairs mean difference and the unmatched-reservoir mean difference, inverse-variance-weighted when both are usable; see InferenceBaiAdjustedTKK14 for the full formula, the pair-distance definition, and the confidence-interval/p-value construction shared with this class. The two leaves differ only in how pair distance is defined during matching: this class (KK21) uses the design's covariate-weighted squared distance \sum_j w_j (x_{1j} - x_{2j})^2, where w_j are the design's covariate_weights (see DesignSeqOneByOneKK21), unlike KK14's unweighted distance. Because the estimator is closed-form, initialization does not use warm starts (there is no iterative fit to warm-start).

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

Legacy class. Not fully tested in comprehensive_tests.R.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceBaiAdjustedTKK21

Methods

Public methods

+ inherited public methods from Inference

InferenceBaiAdjustedTKK21$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceBaiAdjustedTKK21$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceBaiAdjustedTKK21$supports_rand_pval_for_incidence()

Usage
InferenceBaiAdjustedTKK21$supports_rand_pval_for_incidence()

InferenceBaiAdjustedTKK21$compute_rand_two_sided_pval()

Usage
InferenceBaiAdjustedTKK21$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceBaiAdjustedTKK21$clone()

The objects of this class are cloneable with this method.

Usage
InferenceBaiAdjustedTKK21$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceBaiAdjustedTKK21$new(seq_des)
inf$compute_estimate()


Bayesian Bootstrap-capable Inference

Description

Abstract class for Dirichlet-weight Bayesian bootstrap inference layered on top of the existing nonparametric bootstrap infrastructure.

Super classes

Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap

Methods

Public methods

+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference

InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_pval_types()

Returns the type values compute_bayesian_bootstrap_two_sided_pval() accepts.

Usage
InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_pval_types()

InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_ci_types()

Returns the type values compute_bayesian_bootstrap_confidence_interval() accepts.

Usage
InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_ci_types()

InferenceBayesianBootstrap$compute_estimate_with_bootstrap_weights()

Recomputes the treatment estimate under Bayesian-bootstrap subject-, block-, cluster-, or matched-set weights.

This is an abstract hook implemented by concrete inference families that support weighted re-estimation.

Usage
InferenceBayesianBootstrap$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Numeric Bayesian-bootstrap weights at the design's exchangeable resampling unit. For ordinary designs these are subject-level weights. For blocking, clustering, or matching designs these may instead be block-, cluster-, pair-, or matched-set-level weights, depending on weighting_unit_type.

estimate_only

If TRUE, compute only the point estimate for the weighted replicate.

Returns

A numeric treatment-effect estimate for the weighted replicate.


InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T()

Creates the Bayesian-bootstrap distribution of the treatment estimate using Dirichlet weights.

Usage
InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T(
  B = 501,
  show_progress = TRUE,
  debug = FALSE,
  weighting_unit_type = NULL
)
Arguments
B

Number of Bayesian-bootstrap replicates. The default is 501.

show_progress

A flag indicating whether a progress bar should be displayed.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

weighting_unit_type

Optional Bayesian-bootstrap weighting-unit scheme. Legal public values are:

NULL

Use the design's default weighting-unit logic. For ordinary non-blocking designs this is the usual subject-level Bayesian bootstrap. For certain blocking designs, NULL maps to the same behavior as "within_blocks".

"within_blocks"

Only legal for blocking-style designs that support block-aware weighting: DesignFixedBlocking, DesignFixedOptimalBlocks, DesignSeqOneByOneSPBR, and DesignFixedBlockedCluster. Draws Dirichlet weights on observational units within each observed block/stratum. For blocked cluster designs this means cluster-within-stratum weights.

"resample_blocks"

Only legal for the same blocking-style designs as "within_blocks". Draws Dirichlet weights on whole observed blocks/strata rather than on units within each block.

Any non-NULL value is rejected for designs outside that blocking family.

Returns

When debug = FALSE (default), a numeric vector of length B containing the Bayesian-bootstrap estimates. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.


InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval()

Computes a Bayesian-bootstrap-based two-sided p-value for the treatment effect.

Usage
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval(
  delta = 0,
  B = 501,
  type = NULL,
  na.rm = FALSE,
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  weighting_unit_type = NULL
)
Arguments
delta

Null hypothesis value. Default 0.

B

Number of Bayesian-bootstrap replicates. Default 501.

type

Type of Bayesian-bootstrap p-value. Supported values are "percentile" (default), "symmetric", "wald", "studentized" / "bootstrap-t" (pivots by replicate SE from compute_estimate_with_bootstrap_weights(..., estimate_only = FALSE)), and "bca" (bias-corrected and accelerated via leave-one-unit-out Bayesian jackknife).

na.rm

If TRUE, discard non-finite bootstrap replicates before computing the p-value. Otherwise, any non-finite replicate returns NA.

show_progress

A flag indicating whether a progress bar should be displayed.

min_number_usable_samples

Minimum number of finite Bayesian-bootstrap replicates required after filtering. Default 5.

weighting_unit_type

Optional Bayesian-bootstrap weighting-unit scheme. See InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T().

Returns

A numeric two-sided p-value, or NA_real_ if too few usable replicates remain or the estimate is non-finite.


InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval()

Computes a Bayesian-bootstrap confidence interval for the treatment effect.

Usage
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 501,
  type = NULL,
  na.rm = TRUE,
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  weighting_unit_type = NULL
)
Arguments
alpha

Significance level. Default 0.05.

B

Number of Bayesian-bootstrap replicates. Default 501.

type

Type of Bayesian-bootstrap interval. Supported values are "percentile" (default), "basic", "wald", "studentized" / "bootstrap-t" (pivots by replicate SE from compute_estimate_with_bootstrap_weights(..., estimate_only = FALSE)), and "bca" (bias-corrected and accelerated via leave-one-unit-out Bayesian jackknife).

na.rm

If TRUE, discard non-finite bootstrap replicates before constructing the interval.

show_progress

A flag indicating whether a progress bar should be displayed.

min_number_usable_samples

Minimum number of finite Bayesian-bootstrap replicates required after filtering. Default 5.

weighting_unit_type

Optional Bayesian-bootstrap weighting-unit scheme. See InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T().

Returns

A length-2 numeric confidence interval. Returns c(NA_real_, NA_real_) when the estimate is non-finite or too few usable replicates remain.


InferenceBayesianBootstrap$clone()

The objects of this class are cloneable with this method.

Usage
InferenceBayesianBootstrap$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Linear Mixed Model Inference for KK Designs with Continuous Response

Description

Fits a linear mixed model for continuous responses under a KK matching-on-the-fly design. The matched-pair strata enter as a subject-level random intercept (1 | group_id), accounting for within-pair correlation.

When use_rcpp = TRUE (default) the likelihood is maximised by an internal Rcpp/L-BFGS routine that requires no external packages. Set use_rcpp = FALSE to fall back to glmmTMB.

The treatment coefficient \beta_T is on the response's natural (untransformed) scale — a mean difference, not a ratio or log-scale effect. likelihood_tier = "full": likelihood-ratio, score, and Wald tests are all available when the model converges (see $get_likelihood_test_spec() inherited from the shared count/GLMM likelihood plumbing). Validity requires the random-intercept-per-pair structure to correctly capture the design's matching dependence and the usual linear mixed model assumptions (conditional normality of responses and pair effects, correctly specified fixed-effects formula).

Super class

Inference -> InferenceContinKKGLMM

Methods

Public methods

+ inherited public methods from Inference

InferenceContinKKGLMM$new()

Initialize inference for the linear mixed model Y_i = \beta_0 + \beta_T W_i + X_i^\top \gamma + b_{g(i)} + \epsilon_i, b_g \sim N(0, \sigma_b^2), \epsilon_i \sim N(0, \sigma_e^2), where g(i) is subject i's matched-pair group id, W_i is the treatment indicator, X_i are covariates, and \beta_T is the treatment effect (mean difference on the response's natural scale). The random intercept b_g absorbs the within-pair correlation induced by matching, so \beta_T's standard error correctly reflects the design.

Usage
InferenceContinKKGLMM$new(
  des_obj,
  model_formula = NULL,
  use_rcpp = TRUE,
  use_gls_fast_path = TRUE,
  use_gls_fast_path_bootstrap = FALSE,
  verbose = FALSE,
  smart_cold_start_default = NULL,
  optimization_alg = NULL
)
Arguments
des_obj

A completed Design object with a continuous response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

use_rcpp

Logical. If TRUE (default), use the optimised Rcpp Gaussian LMM implementation (no external package required). If FALSE, use glmmTMB.

use_gls_fast_path

Logical. If TRUE (default), use a fast GLS estimator (no optimisation) for estimate_only calls during randomisation inference once variance components are cached from a prior full fit. Statistically exact: fixing VC at the null-fit MLE and permuting only the treatment assignment gives a valid permutation test by exchangeability. Set FALSE to always run full L-BFGS.

use_gls_fast_path_bootstrap

Logical. If TRUE, also use the fast GLS estimator for non-studentised bootstrap draws (estimate_only = TRUE weighted calls). Asymptotically valid by the plug-in principle (VC orthogonal to beta_T in the Fisher information), but not exact in finite samples. Default FALSE.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.

optimization_alg

The optimization algorithm to use. Default is dispatched via policy.


InferenceContinKKGLMM$compute_estimate()

Fits the linear mixed model by maximum likelihood and returns \hat\beta_T. When use_rcpp = TRUE (the default), the log-likelihood \ell(\beta, \sigma_b^2, \sigma_e^2) is maximized directly by an internal Rcpp/L-BFGS routine over \beta and the log-variance components; otherwise glmmTMB performs the fit. If use_gls_fast_path = TRUE and variance components are already cached from a prior full fit (used during randomization inference, where only the treatment column changes across permutations), estimate_only = TRUE calls instead solve the generalized least-squares problem at the cached variance components — exact under exchangeability of the permuted treatment assignment, since fixing the variance components at their null-fit MLE and permuting only W preserves the permutation test's validity.

Usage
InferenceContinKKGLMM$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations (standard error, degrees of freedom) needed for confidence intervals or p-values; only \hat\beta_T is returned.


InferenceContinKKGLMM$compute_estimate_with_bootstrap_weights()

Refits the linear mixed model with subject/block-level weights applied to each row's contribution to the likelihood (Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded from subject/block level to individual rows via private$expand_subject_or_block_weights_to_row_weights()), and returns the reweighted estimate \hat\beta_T^{(w)}. When weights are effectively constant, this collapses to the unweighted compute_estimate() call (returns df = Inf to signal a degenerate/skipped bootstrap replicate rather than refitting). When use_rcpp = TRUE, a weighted Rcpp fast path is tried first via private$weighted_rcpp_estimate(); otherwise private$compute_weighted_glmm_bootstrap_estimate() refits via glmmTMB-based machinery.

Usage
InferenceContinKKGLMM$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Numeric vector of nonnegative weights, one per matched-pair group or reservoir subject (bootstrap draw weights), expanded to per-row weights before fitting.

estimate_only

Logical. If TRUE, skip standard-error computation.

Returns

The reweighted treatment estimate \hat\beta_T^{(w)}.


InferenceContinKKGLMM$compute_asymp_confidence_interval()

Computes a 1-\alpha Wald confidence interval for \beta_T from the fitted-model standard error, using a normal or t critical value depending on the resolved degrees of freedom (see private$compute_z_or_t_ci_from_s_and_df).

Usage
InferenceContinKKGLMM$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level of the interval is 1 - \code{alpha}. Default 0.05.


InferenceContinKKGLMM$compute_asymp_two_sided_pval()

Computes a two-sided Wald p-value for H_0: \beta_T = \code{delta} using the fitted-model estimate and standard error (same statistic that $compute_asymp_confidence_interval() inverts).

Usage
InferenceContinKKGLMM$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null treatment-effect value. Default 0.


InferenceContinKKGLMM$clone()

The objects of this class are cloneable with this method.

Usage
InferenceContinKKGLMM$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

See Also

Comparable Python API: statsmodels MixedLM. See also: Mixed model (Wikipedia).

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinKKGLMM$new(seq_des)
inf$compute_estimate()


OLS IVWC Compound Inference for KK Designs

Description

Fits a variance-weighted compound estimator for KK matching-on-the-fly designs with continuous responses using OLS regression for matched-pair differences and reservoir outcomes, with the treatment indicator and, optionally, all recorded covariates as predictors. Note that warm starts are disabled for this class as OLS is a closed-form estimator and does not benefit from initialization.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

The point estimate \hat\beta_T is the inverse-variance-weighted combination of an OLS fit on matched-pair within-pair differences and an OLS fit on reservoir (unmatched) subjects' outcomes, falling back to whichever sub-fit is usable if the other is not — the same compound combination rule used by InferenceAllKKMeanDiffIVWC, generalized here to allow covariate adjustment via model_formula. likelihood_tier = "none": this is an estimating-equation (least squares) estimator, not a fitted likelihood, so only Wald-type asymptotic inference is available (no likelihood-ratio or score test).

Legacy class. Not fully tested in comprehensive_tests.R.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceContinKKOLSIVWC

Methods

Public methods

+ inherited public methods from Inference

InferenceContinKKOLSIVWC$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceContinKKOLSIVWC$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceContinKKOLSIVWC$supports_rand_pval_for_incidence()

Usage
InferenceContinKKOLSIVWC$supports_rand_pval_for_incidence()

InferenceContinKKOLSIVWC$compute_rand_two_sided_pval()

Usage
InferenceContinKKOLSIVWC$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceContinKKOLSIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferenceContinKKOLSIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinKKOLSIVWC$new(seq_des)
inf$compute_estimate()


OLS Combined-Likelihood Inference for KK Designs

Description

Fits a single stacked OLS regression over matched-pair differences and reservoir observations for KK matching-on-the-fly designs with continuous responses, using the treatment indicator and, optionally, all recorded covariates as predictors. Note that warm starts are disabled for this class as OLS is a closed-form estimator and does not benefit from initialization.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

Model. Let m be the number of matched pairs and n_R = n_{RT} + n_{RC} the number of unmatched reservoir subjects. The design matrix stacks two blocks: m matched-pair difference rows (each row's response is the within-pair outcome difference, coded with an implicit unit treatment column and covariate differences X_{d}), and n_R reservoir rows (raw covariates plus a treatment/matching-status indicator column). The stacked regression is fit by ordinary least squares (lm.fit), and \hat\beta_T is the coefficient on the treatment/matched-difference column, i.e. an additive mean-difference estimand on the outcome's natural scale. If only matched pairs exist, j_treat = 1; if only reservoir data exist, j_treat = 2; if both exist, the combined design uses j_treat = 2. If neither matched pairs nor a treatment-and-control-populated reservoir exist, the estimate is marked nonestimable ("no_usable_matched_or_reservoir_data").

Variance. Standard errors use the HC2 heteroskedasticity-consistent sandwich estimator (ols_hc2_post_fit_cpp), not the classical OLS variance, so the Wald confidence interval/p-value are robust to heteroskedasticity across the matched/reservoir blocks.

Likelihood tier. likelihood_tier = "full": this is a genuine Gaussian likelihood (not a quasi-likelihood or partial likelihood), so score, gradient, and likelihood-ratio testing types are available in addition to Wald, and an exact (not higher-order-accurate) Bartlett correction reproduces base R's lm() classical partial F-test exactly under the classical homoskedastic-Gaussian-errors assumption (a stronger assumption than the HC2-robust Wald path uses, so the two paths need not agree numerically).

Assumptions. Continuous response; independent matched pairs and/or independent reservoir subjects; no censoring (assertNoCensoring() is enforced); a KK matching-on-the-fly design (DesignSeqOneByOneKK14 or subclass).

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceContinKKOLSOneLik

Methods

Public methods

+ inherited public methods from Inference

InferenceContinKKOLSOneLik$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceContinKKOLSOneLik$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceContinKKOLSOneLik$supports_rand_pval_for_incidence()

Usage
InferenceContinKKOLSOneLik$supports_rand_pval_for_incidence()

InferenceContinKKOLSOneLik$compute_rand_two_sided_pval()

Usage
InferenceContinKKOLSOneLik$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceContinKKOLSOneLik$clone()

The objects of this class are cloneable with this method.

Usage
InferenceContinKKOLSOneLik$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kapelner, A. and Krieger, A. M. (2014). Matching on-the-fly: Sequential allocation with higher power and efficiency. Biometrics, 70(2), 378-388. doi:10.1111/biom.12148. (KK14 in REFERENCES.md.)

See Also

InferenceContinKKOLSIVWC for the inverse-variance-weighted-combination alternative to this one-likelihood combined-fit approach; analogous Python API: statsmodels GLM/OLS.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinKKOLSOneLik$new(seq_des)
inf$compute_estimate()


Quantile Regression Compound Estimator for KK Matching-on-the-Fly Designs

Description

A variance-weighted compound quantile regression estimator for KK matching-on-the-fly designs with continuous responses. The estimator combines:

  1. Quantile regression on within-pair differences (matched pairs)

  2. Quantile regression on reservoir subjects (treatment vs control)

using the same variance-weighted combination logic as the OLS compound estimator.

Default quantile: tau = 0.5 (median regression). At tau = 0.5 this estimates the median treatment effect, which is the canonical nonparametric location estimator and is more robust to outliers and heavy-tailed response distributions than the OLS mean-based estimator. To target a different quantile of the treatment effect distribution — for example the 25th or 75th percentile — pass tau = 0.25 or tau = 0.75 to the constructor:

  inf = InferenceContinKKQuantileRegrIVWC$
  new(seq_des, tau = 0.75)

Any value strictly between 0 and 1 is accepted.

Standard errors use Powell's "nid" sandwich estimator (non-iid), which is more robust than the "iid" (constant-density) assumption; the implementation falls back to "iid" on failure. Asymptotic z-based inference is used throughout.

The randomization-based confidence interval is inherited from the base class and is valid for location-shift models at all quantiles: shifting y by delta maps the tau-th quantile treatment effect to delta under the null.

This class requires the quantreg package, which is listed in Suggests and is not installed automatically with EDI. Install quantreg before using this class.

Legacy class. Not fully tested in comprehensive_tests.R.

Super class

Inference -> InferenceContinKKQuantileRegrIVWC

Methods

Public methods

+ inherited public methods from Inference

InferenceContinKKQuantileRegrIVWC$new()

Initialize continuous-response KK IVWC quantile-regression inference; see InferenceContinKKQuantileRegrIVWC.

Usage
InferenceContinKKQuantileRegrIVWC$new(
  des_obj,
  model_formula = NULL,
  tau = 0.5,
  verbose = FALSE
)
Arguments
des_obj

A DesignSeqOneByOne object whose entire n subjects are assigned and response y is recorded within.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

tau

The quantile level for regression, strictly between 0 and 1. The default tau = 0.5 estimates the median treatment effect. Pass a different value (e.g. tau = 0.25 or tau = 0.75) to target the corresponding percentile of the treatment effect distribution.

verbose

A flag indicating whether messages should be displayed to the user. Default is FALSE.

Examples
set.seed(1)
x_dat <- data.frame(
  x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
  x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneKK14$new(n = nrow(x_dat), response_type = "continuous", verbose =
FALSE)
for (i in seq_len(nrow(x_dat))) {
  seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(c(1.2, 0.9, 1.5, 1.8, 2.1, 1.7, 2.6, 2.2))
infer <- InferenceContinKKQuantileRegrIVWC$new(seq_des, verbose = FALSE)
infer

InferenceContinKKQuantileRegrIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferenceContinKKQuantileRegrIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinKKQuantileRegrIVWC$new(seq_des)
inf$compute_estimate()


## ------------------------------------------------
## Method `InferenceContinKKQuantileRegrIVWC$new()`
## ------------------------------------------------

set.seed(1)
x_dat <- data.frame(
  x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
  x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneKK14$new(n = nrow(x_dat), response_type = "continuous", verbose =
FALSE)
for (i in seq_len(nrow(x_dat))) {
  seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(c(1.2, 0.9, 1.5, 1.8, 2.1, 1.7, 2.6, 2.2))
infer <- InferenceContinKKQuantileRegrIVWC$new(seq_des, verbose = FALSE)
infer


Quantile Regression Combined-Likelihood Compound Estimator for KK Designs (Continuous)

Description

Fits the combined stacked quantile regression (matched-pair differences + reservoir) using the treatment indicator and all recorded covariates for continuous responses. Minimises the joint check-function loss over both data sources simultaneously. Inference is based on the stacked combined-likelihood quantile-regression fit.

Model. Analogous to InferenceContinKKOLSOneLik's single stacked design (matched-pair difference rows plus reservoir rows fit jointly), but the objective is the quantreg check-function loss \rho_\tau(u) = u(\tau - \mathbf{1}_{u<0}) rather than squared error, so the estimand \beta_T is the treatment effect on the \tau-th quantile of the response, not the mean. tau = 0.5 (default) targets the median treatment effect, which is more robust to outliers and heavy-tailed responses than the OLS mean-based estimator; any value strictly between 0 and 1 is accepted. likelihood_tier = "none": no likelihood-based (score/gradient/lik_ratio) testing types, only Wald.

Assumptions. Continuous response; independent matched pairs and/or independent reservoir subjects; no censoring; a KK matching-on-the-fly design. Requires the quantreg package (listed in Suggests, not installed automatically).

Super class

Inference -> InferenceContinKKQuantileRegrOneLik

Methods

Public methods

+ inherited public methods from Inference

InferenceContinKKQuantileRegrOneLik$new()

Initialize continuous-response KK combined-likelihood quantile-regression inference. The shared stacked matched-pair/reservoir quantile-regression fit and its tau semantics are documented on InferenceContinKKQuantileRegrOneLik and on the shared compute_estimate() method it inherits from the KKQuantileRegrOneLik component.

Usage
InferenceContinKKQuantileRegrOneLik$new(
  des_obj,
  model_formula = NULL,
  tau = 0.5,
  verbose = FALSE
)
Arguments
des_obj

A DesignSeqOneByOne object whose entire n subjects are assigned and response y is recorded within.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

tau

The quantile level for regression, strictly between 0 and 1. Default is 0.5.

verbose

Whether to print progress messages.

Examples
set.seed(1)
x_dat <- data.frame(
  x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
  x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneKK14$new(n = nrow(x_dat), response_type = "continuous", verbose =
FALSE)
for (i in seq_len(nrow(x_dat))) {
  seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(c(1.2, 0.9, 1.5, 1.8, 2.1, 1.7, 2.6, 2.2))
infer <- InferenceContinKKQuantileRegrOneLik$new(seq_des, verbose
= FALSE)
infer

InferenceContinKKQuantileRegrOneLik$clone()

The objects of this class are cloneable with this method.

Usage
InferenceContinKKQuantileRegrOneLik$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kapelner, A. and Krieger, A. M. (2014). Matching on-the-fly: Sequential allocation with higher power and efficiency. Biometrics, 70(2), 378-388. doi:10.1111/biom.12148. (KK14 in REFERENCES.md.) Koenker, R. (2005). Quantile Regression. Cambridge University Press.

See Also

InferenceContinKKQuantileRegrIVWC for the inverse-variance-weighted-combination alternative to this one-likelihood combined-fit approach. Quantile regression (orientation).

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinKKQuantileRegrOneLik$new(seq_des)
inf$compute_estimate()


## ------------------------------------------------
## Method `InferenceContinKKQuantileRegrOneLik$new()`
## ------------------------------------------------

set.seed(1)
x_dat <- data.frame(
  x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
  x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneKK14$new(n = nrow(x_dat), response_type = "continuous", verbose =
FALSE)
for (i in seq_len(nrow(x_dat))) {
  seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(c(1.2, 0.9, 1.5, 1.8, 2.1, 1.7, 2.6, 2.2))
infer <- InferenceContinKKQuantileRegrOneLik$new(seq_des, verbose
= FALSE)
infer


Robust-Regression IVWC Compound Inference for KK Designs

Description

Fits a variance-weighted compound estimator for KK matching-on-the-fly designs with continuous responses using robust regression for matched-pair differences and reservoir outcomes, with treatment and, optionally, all recorded covariates as predictors.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceContinKKRobustRegrIVWC

Methods

Public methods

+ inherited public methods from Inference

InferenceContinKKRobustRegrIVWC$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceContinKKRobustRegrIVWC$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceContinKKRobustRegrIVWC$supports_rand_pval_for_incidence()

Usage
InferenceContinKKRobustRegrIVWC$supports_rand_pval_for_incidence()

InferenceContinKKRobustRegrIVWC$compute_rand_two_sided_pval()

Usage
InferenceContinKKRobustRegrIVWC$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceContinKKRobustRegrIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferenceContinKKRobustRegrIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinKKRobustRegrIVWC$new(seq_des)
inf$compute_estimate()


Robust-Regression Combined-Likelihood Inference for KK Designs

Description

Fits a single stacked robust regression over matched-pair differences and reservoir observations for KK matching-on-the-fly designs with continuous responses.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

Model. Analogous to InferenceContinKKOLSOneLik's single stacked design (matched-pair difference rows plus reservoir rows fit in one regression, treatment coefficient \beta_T), but fit with a robust M/MM-estimator (MASS::rlm, or an internal Rcpp IRLS kernel when use_rcpp = TRUE) instead of ordinary least squares. "MM" (the default method) starts from an LQS-based high-breakdown fit; "M" can warm-start from OLS (start_with_ols = TRUE).

Likelihood tier. likelihood_tier = "quasi": the robust objective is not a normalized likelihood, so only Wald-type asymptotic inference is available (compute_wald_confidence_interval()/ compute_wald_two_sided_pval(), aliased by the standard compute_asymp_* names) — no score/gradient/likelihood-ratio testing types, unlike the OLS one-likelihood sibling.

Assumptions. Continuous response; independent matched pairs and/or independent reservoir subjects; no censoring; a KK matching-on-the-fly design. Robust regression trades some efficiency under exactly-Gaussian errors for resistance to outliers and heavy tails.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceContinKKRobustRegrOneLik

Methods

Public methods

+ inherited public methods from Inference

InferenceContinKKRobustRegrOneLik$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceContinKKRobustRegrOneLik$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceContinKKRobustRegrOneLik$supports_rand_pval_for_incidence()

Usage
InferenceContinKKRobustRegrOneLik$supports_rand_pval_for_incidence()

InferenceContinKKRobustRegrOneLik$compute_rand_two_sided_pval()

Usage
InferenceContinKKRobustRegrOneLik$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceContinKKRobustRegrOneLik$clone()

The objects of this class are cloneable with this method.

Usage
InferenceContinKKRobustRegrOneLik$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kapelner, A. and Krieger, A. M. (2014). Matching on-the-fly: Sequential allocation with higher power and efficiency. Biometrics, 70(2), 378-388. doi:10.1111/biom.12148. (KK14 in REFERENCES.md.)

See Also

InferenceContinKKRobustRegrIVWC for the inverse-variance-weighted-combination alternative to this one-likelihood combined-fit approach. Analogous Python API: statsmodels RLM.


Lin (2013) Covariate-Adjusted OLS Inference for Continuous Responses

Description

Fits Lin's (2013) covariate-adjusted linear estimator for continuous responses: OLS of Y_i on [1, W_i, X_i^c, W_i X_i^c], where X_i^c = X_i - \bar X are covariates centered at their sample means and W_i X_i^c are treatment-by-centered-covariate interactions (omitted when there are no covariates, reducing to plain OLS with \hat\beta_T the simple mean difference). Centering makes \hat\beta_T interpretable as the average treatment effect regardless of whether interactions are included, following the design-based reinterpretation of Freedman's critique of ANCOVA in randomized experiments. Standard errors use the HC2 heteroskedasticity-consistent (Huber-White-type) covariance estimator (ols_hc2_post_fit_cpp), not classical OLS SEs assuming homoskedasticity — this is the estimator Lin (2013) recommends since it remains conservative under treatment-effect heterogeneity, unlike the classical or HC0 sandwich variants. likelihood_tier = "full": Wald, score, gradient, and likelihood-ratio tests are all available, with the parametric-likelihood bootstrap using the OLS Gaussian-errors model (Y_i \mid x_i \sim N(x_i^\top \beta, \sigma^2)) as the generative null even though the design-based HC2 standard error does not itself assume homoskedastic Gaussian errors — the likelihood-ratio/score/gradient machinery is a secondary, model-based inference path alongside the primary HC2-Wald and randomization paths. Validity of \hat\beta_T as an average-treatment-effect estimator relies on randomization (of W), not on any particular outcome model; the working linear model need not be correctly specified.

Super class

Inference -> InferenceContinLin

Methods

Public methods

+ inherited public methods from Inference

InferenceContinLin$new()

Initialize inference for Lin's (2013) covariate-adjusted OLS estimator (intercept, treatment, centered covariates, and treatment-by-centered-covariate interactions); see InferenceContinLin for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferenceContinLin$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  harden = TRUE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with a continuous response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

harden

Flag for consistent API.

smart_cold_start_default

Flag for consistent API.


InferenceContinLin$compute_estimate()

Fits Lin's covariate-adjusted OLS model (stats's lm.fit on the centered-covariate design matrix) and returns \hat\beta_T, the estimated average treatment effect. A design matrix that is rank-deficient or has fewer usable rows than columns, or a fit with non-finite coefficients, is cached as nonestimable rather than returned.

Usage
InferenceContinLin$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip HC2 variance computation and cache only the point estimate; used by randomization and bootstrap resampling paths.


InferenceContinLin$compute_estimate_with_bootstrap_weights()

Refits Lin's model with subject/block-level weights applied to a weighted least-squares fit (stats::lm.wfit) — Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights() — and returns the reweighted estimate \hat\beta_T^{(w)}. When estimate_only = FALSE, also computes a weighted-residual variance estimate (not the HC2 estimator used by compute_asymp_confidence_interval()) for internal bootstrap diagnostics. Rows with non-finite or non-positive weight, or non-finite response, are dropped from the weighted fit; if no rows remain, or the fit fails, the estimate is NA.

Usage
InferenceContinLin$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip variance calculations.


InferenceContinLin$compute_asymp_confidence_interval()

Wald confidence interval for \beta_T using the HC2 heteroskedasticity-robust standard error (ols_hc2_post_fit_cpp); see InferenceAsymp for the shared t/z interval contract. Fits the model first if not already cached.

Usage
InferenceContinLin$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level. The default is 0.05.


InferenceContinLin$compute_asymp_two_sided_pval()

Two-sided Wald test of H_0: \beta_T = \code{delta} using the HC2 heteroskedasticity-robust standard error; see InferenceAsymp for the shared t/z test contract. Fits the model first if not already cached.

Usage
InferenceContinLin$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null treatment effect. Defaults to 0.


InferenceContinLin$clone()

The objects of this class are cloneable with this method.

Usage
InferenceContinLin$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Lin, W. (2013). "Agnostic notes on regression adjustments to experimental data: Reexamining Freedman's critique." The Annals of Applied Statistics, 7(1), 295-318, doi:10.1214/12-AOAS583.

See Also

InferenceContinOLS for the uncentered, non-interacted OLS estimator this class generalizes.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinLin$new(seq_des)
inf$compute_estimate()


OLS Inference for Continuous Responses

Description

Fits an ordinary least squares regression for continuous responses: Y_i = \beta_0 + \beta_T W_i + X_i^\top \gamma + \epsilon_i, using the treatment indicator and, optionally, all recorded covariates as predictors (uncentered, no treatment-covariate interactions — see InferenceContinLin for the centered-covariate, interacted variant). \hat\beta_T is a mean difference on the response's natural scale. likelihood_tier = "full": Wald, score, gradient, and likelihood-ratio tests are all available (fast_ols_cpp/fast_ols_with_var_cpp), with the parametric-likelihood bootstrap using the OLS Gaussian-errors model as the generative null. Standard errors are the closed-form OLS variance under homoskedastic errors (unlike InferenceContinLin's HC2 heteroskedasticity-robust SE). Warm starts are disabled (fit_warm_start_enabled = FALSE set at construction) because OLS is a closed-form estimator and gains nothing from an iterative optimizer's warm-started initial values. Validity requires the usual OLS assumptions: correctly specified linear predictor, and (for the asymptotic/likelihood inference path specifically) homoskedastic, approximately normal errors; the randomization-inference path relies only on randomization of W.

Super class

Inference -> InferenceContinOLS

Methods

Public methods

+ inherited public methods from Inference

InferenceContinOLS$new()

Initialize inference for the OLS model Y_i = \beta_0 + \beta_T W_i + X_i^\top \gamma + \epsilon_i; see InferenceContinOLS for the model form. Disables warm-started optimizer initial values (not applicable to this closed-form estimator). Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferenceContinOLS$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  max_resample_attempts = 50L,
  harden = TRUE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with a continuous response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

max_resample_attempts

Maximum number of times a single bootstrap replicate may be redrawn when the drawn sample fails validity screening. Default 50L.

harden

Whether to apply robustness measures.

smart_cold_start_default

Flag for consistent API.


InferenceContinOLS$compute_estimate()

Fits the OLS model (fast_ols_cpp/fast_ols_with_var_cpp) and returns \hat\beta_T. If not hardened (private$harden == FALSE), fits directly on the full design matrix; otherwise uses QR column-dropping hardening to handle rank-deficient designs. A fit with a non-finite treatment coefficient is cached as nonestimable rather than returned.

Usage
InferenceContinOLS$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.


InferenceContinOLS$compute_estimate_with_bootstrap_weights()

Refits the OLS model with subject/block-level weights applied to a weighted least-squares fit (stats::lm.wfit) — Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights() — and returns the reweighted estimate \hat\beta_T^{(w)}. When estimate_only = FALSE, also computes a weighted-residual variance estimate for internal bootstrap diagnostics. Rows with non-finite or non-positive weight, or non-finite response, are dropped from the weighted fit; if no rows remain, or the fit fails, the estimate is NA.

Usage
InferenceContinOLS$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip variance calculations.


InferenceContinOLS$compute_asymp_confidence_interval()

Uses the shared asymptotic confidence-interval contract; see InferenceAsymp.

Usage
InferenceContinOLS$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.


InferenceContinOLS$compute_asymp_two_sided_pval()

Uses the shared asymptotic two-sided p-value contract; see InferenceAsymp.

Usage
InferenceContinOLS$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null difference to test against. Default is zero.


InferenceContinOLS$clone()

The objects of this class are cloneable with this method.

Usage
InferenceContinOLS$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Rosenbaum, P. R. (2002). Observational Studies (2nd ed.). Springer, for the OLS mean-difference estimator's design-based justification under randomization.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinOLS$new(seq_des)
inf$compute_estimate()


Quantile Regression Inference for Continuous Responses

Description

Fits a linear quantile regression, Q_\tau(Y_i \mid x_i) = x_i^\top\beta_\tau, for a continuous response, estimating \beta_\tau by minimizing the asymmetric ("pinball" / "check") loss

\hat\beta_\tau = \operatorname*{arg\,min}_\beta \sum_{i=1}^n \rho_\tau(y_i - x_i^\top\beta), \qquad \rho_\tau(u) = u\left(\tau - \mathbb{1}[u < 0]\right),

via quantreg's simplex method (quantreg::rq/rq.fit(..., method = "br")). The treatment coefficient is the estimated shift in the \tau-th conditional quantile of the response attributable to treatment, holding any other covariates in model_formula fixed; by default tau = 0.5, so this is median regression (robust to outliers and distributional skew relative to mean-based estimators, at the cost of losing the mean-shift interpretation away from \tau = 0.5).

Standard errors use quantreg's Powell-style "nid" (non-i.i.d., kernel-based sparsity/local-density estimator) sandwich covariance when available, with fallback to the simpler "iid" estimator if needed. Inference (confidence intervals, p-values) is based on the resulting asymptotic normal approximation, not an exact finite-sample distribution.

This class requires the quantreg package, which is listed under Suggests and is not installed automatically with EDI. Install quantreg manually before use.

Super class

Inference -> InferenceContinQuantileRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceContinQuantileRegr$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize a quantile-regression inference object for a completed design with a continuous, uncensored response. Requires the quantreg package to be installed.

Usage
InferenceContinQuantileRegr$new(
  des_obj,
  model_formula = NULL,
  tau = 0.5,
  verbose = FALSE
)
Arguments
des_obj

A completed Design object with a continuous response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

tau

The quantile \tau \in (0, 1) to estimate (default 0.5, i.e. median regression).

verbose

Whether to print progress messages. Default FALSE.


InferenceContinQuantileRegr$compute_estimate()

Computes the treatment coefficient \hat\beta_{T,\tau} from a check-loss quantile regression fit at quantile tau (see class documentation for the full model). Rank-deficient covariate columns are dropped before fitting (see private$reduce_design_matrix_for_quantile()); returns NA if the reduced design has no usable treatment column or too few residual degrees of freedom.

Usage
InferenceContinQuantileRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.


InferenceContinQuantileRegr$compute_estimate_with_bootstrap_weights()

Recomputes the quantile-regression treatment estimate under subject/block bootstrap weights (quantreg::rq(..., weights = row_weights)), used by the Bayesian bootstrap and related weighted-resampling machinery; see InferenceBayesianBootstrap. Unlike $compute_estimate(), this always reduces the design matrix from scratch (reuse_factorizations = FALSE) rather than reusing a cached rank-reduction from a prior warm-started fit.

Usage
InferenceContinQuantileRegr$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip variance calculations.


InferenceContinQuantileRegr$compute_asymp_confidence_interval()

Computes a 1-\alpha level confidence interval for the quantile-regression treatment coefficient \hat\beta_{T,\tau}, using quantreg's Powell-style "nid" asymptotic standard error (falling back to "iid" if unavailable — see class documentation) and residual degrees of freedom n - p. See InferenceAsymp for the shared asymptotic confidence-interval contract this delegates to.

Usage
InferenceContinQuantileRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.


InferenceContinQuantileRegr$compute_asymp_two_sided_pval()

Computes a two-sided Wald p-value testing H_0: \beta_{T,\tau} = \code{delta}, from the same quantreg sandwich standard error and degrees of freedom used by $compute_asymp_confidence_interval(). See InferenceAsymp for the shared asymptotic two-sided p-value contract this delegates to.

Usage
InferenceContinQuantileRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null difference to test against. Default is zero.


InferenceContinQuantileRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceContinQuantileRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Koenker, R., and Bassett, G. (1978). "Regression Quantiles." Econometrica, 46(1), 33-50, doi:10.2307/1913643, for the check-loss quantile regression estimator; Koenker, R. (2005). Quantile Regression, Cambridge University Press, for the Powell-style sandwich standard error estimators used here.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinQuantileRegr$new(seq_des)
inf$compute_estimate()


Robust (M/MM-Estimator) Regression Inference for Continuous Responses

Description

Fits a robust linear regression — by default via this package's fast_robust_regression_cpp C++ backend (use_rcpp = TRUE; see that page for the full M/MM-estimator model, weight functions, and asymptotic variance formula), or via MASS::rlm when use_rcpp = FALSE — for a continuous response using the treatment indicator and, optionally, all recorded covariates as predictors. This provides a Huber/MM-style robustness upgrade over ordinary least squares when outcomes are heavy-tailed or outlier-prone, down-weighting large residuals rather than letting them dominate the fit the way squared-error loss does.

The method argument is passed through to either backend and may be either "M" (Huber's psi function) or "MM" (Tukey's bisquare weight, the default — higher breakdown point than "M" at some cost in asymptotic efficiency under normality). When use_rcpp = TRUE (the default), coefficient standard errors come from the C++ backend's own M-estimator asymptotic variance (ssq_b_j); when FALSE, they come from the coefficient table returned by summary.rlm(). Either way, confidence intervals and p-values use that standard error with residual degrees of freedom n - p in a normal-theory Wald approximation (not an exact finite-sample distribution).

Super class

Inference -> InferenceContinRobustRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceContinRobustRegr$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize a robust-regression inference object for a completed design with a continuous, uncensored response.

Usage
InferenceContinRobustRegr$new(
  des_obj,
  model_formula = NULL,
  method = "MM",
  use_rcpp = TRUE,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with a continuous response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

method

Robust estimation method, "M" (Huber) or "MM" (Tukey bisquare, higher breakdown point). Default "MM".

use_rcpp

Whether to use the fast_robust_regression_cpp C++ backend (TRUE, default) instead of MASS::rlm (FALSE).

verbose

Whether to print progress messages. Default FALSE.

smart_cold_start_default

Whether to use smart starting values for the optimizer.


InferenceContinRobustRegr$compute_estimate()

Computes the robust-regression treatment coefficient \hat\beta_T from an M/MM-estimator fit (see class documentation for the full model and use_rcpp backend choice). Rank-deficient covariate columns are dropped before fitting via private$fit_with_hardened_qr_column_dropping().

Usage
InferenceContinRobustRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.


InferenceContinRobustRegr$compute_estimate_with_bootstrap_weights()

Recomputes the robust-regression treatment estimate under subject/block bootstrap weights, used by the Bayesian bootstrap and related weighted-resampling machinery; see InferenceBayesianBootstrap. When use_rcpp = TRUE, reproduces MASS::rlm's default wt.method = "inv.var" weighting exactly by pre-multiplying X and y by \sqrt{\text{weight}} and running unweighted M-estimation on the transformed data via the C++ backend (rather than adding native weight support to that backend); when use_rcpp = FALSE, passes weights directly to MASS::rlm. This variant never populates a standard error or degrees of freedom (both left NA) — it is estimate-only by construction, matching the estimate_only default of the fast path it always uses internally.

Usage
InferenceContinRobustRegr$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip variance calculations.


InferenceContinRobustRegr$compute_asymp_confidence_interval()

Computes a 1-\alpha level confidence interval for the robust-regression treatment coefficient \hat\beta_T, using the M/MM-estimator's asymptotic standard error (see class documentation for its source depending on use_rcpp) and residual degrees of freedom n - p in a normal-theory Wald approximation. See InferenceAsymp for the shared asymptotic confidence-interval contract this delegates to.

Usage
InferenceContinRobustRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.


InferenceContinRobustRegr$compute_asymp_two_sided_pval()

Computes a two-sided Wald p-value testing H_0: \beta_T = \code{delta}, from the same M/MM-estimator standard error and residual degrees of freedom used by $compute_asymp_confidence_interval(). See InferenceAsymp for the shared asymptotic two-sided p-value contract this delegates to.

Usage
InferenceContinRobustRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null difference to test against. Default is zero.


InferenceContinRobustRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceContinRobustRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinRobustRegr$new(seq_des)
inf$compute_estimate()


Hurdle Negative Binomial Regression Inference for Count Responses

Description

Fits a hurdle negative binomial regression for count responses: a binary hurdle submodel P(Y_i > 0) = \mathrm{logit}^{-1}(X_i^{h\top} \gamma^h) (fit jointly with the count submodel) crossed with a zero-truncated negative-binomial count submodel for Y_i \mid Y_i > 0: \log E[Y_i \mid Y_i > 0, w_i, x_i] = \beta_0 + \beta_T w_i + x_i^\top \gamma, \mathrm{Var}(Y_i \mid Y_i > 0) = \mu_i + \mu_i^2 / \theta (fast_hurdle_negbin_cpp/ fast_hurdle_negbin_with_var_cpp). The hurdle and count submodels may use different covariate formulas (model_formula/model_formula_hurdle). The reported treatment effect is the coefficient from the conditional (truncated, Y > 0) count component, on the log-rate scale, conditional on clearing the hurdle: it is not the effect on the unconditional mean E[Y], which also depends on how treatment shifts the hurdle-crossing probability. A marginal (unconditional-mean) estimand is not yet implemented for this class (see marginal_estimand_report.md). likelihood_tier = "full": Wald, gradient, and (bootstrap-calibrated) likelihood-ratio tests are available for the count submodel's treatment coefficient; a plain score test is not exposed. Jackknife inference is not supported: delete-one refits of this two-part model with a jointly-estimated dispersion parameter are numerically unstable, so compute_jackknife_estimate() and related methods report explicit non-estimability rather than attempting delete-one refits.

Super class

Inference -> InferenceCountHurdleNegBin

Methods

Public methods

+ inherited public methods from Inference

InferenceCountHurdleNegBin$new()

Initialize inference for the two-part hurdle negative binomial model (binary hurdle submodel plus zero-truncated negative-binomial count submodel); see InferenceCountHurdleNegBin for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferenceCountHurdleNegBin$new(
  des_obj,
  model_formula = NULL,
  model_formula_hurdle = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object.

model_formula

Optional formula for covariate adjustment.

model_formula_hurdle

Formula for the hurdle submodel. If NULL (default), it uses the same formula as model_formula.

verbose

A flag indicating whether messages should be displayed.

smart_cold_start_default

Whether to use smart cold start values.


InferenceCountHurdleNegBin$compute_asymp_confidence_interval()

Compute the hurdle negative-binomial asymptotic confidence interval for the treatment coefficient, using the shared count-likelihood semantics documented in InferenceCountLikelihood.

Usage
InferenceCountHurdleNegBin$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The significance level (default 0.05).


InferenceCountHurdleNegBin$compute_asymp_two_sided_pval()

Compute the hurdle negative-binomial asymptotic two-sided p-value for the treatment coefficient, falling back through the shared count-likelihood machinery when needed; see InferenceCountLikelihood.

Usage
InferenceCountHurdleNegBin$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null treatment effect (default 0).


InferenceCountHurdleNegBin$compute_gradient_two_sided_pval()

Gradient test of H_0: \beta_T = \code{delta} on the truncated count submodel's treatment coefficient (a score-test variant using the observed rather than expected information); see InferenceCountLikelihood for the shared likelihood-test dispatch.

Usage
InferenceCountHurdleNegBin$compute_gradient_two_sided_pval(delta = 0)
Arguments
delta

The null treatment effect (default 0).


InferenceCountHurdleNegBin$compute_gradient_confidence_interval()

Compute a hurdle negative-binomial likelihood-based confidence interval by inverting the configured likelihood test. See InferenceCountLikelihood for related score, likelihood-ratio, and gradient methods.

Usage
InferenceCountHurdleNegBin$compute_gradient_confidence_interval(alpha = 0.05)
Arguments
alpha

The significance level (default 0.05).


InferenceCountHurdleNegBin$compute_estimate_with_bootstrap_weights()

Refits the hurdle negative-binomial model with subject/ block-level weights (Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights()) via glmmTMB's glmmTMB(family = truncated_nbinom2()) (not the package's internal C++ solver, which has no weighted variant for this model), and returns the reweighted conditional-count log-rate-ratio estimate \hat\beta_T^{(w)}. Requires the glmmTMB package; errors if unavailable. No standard error is computed (s_beta_hat_T is always NA). A fit that fails, or whose fitted treatment coefficient is missing or non-finite, is cached as nonestimable and returns NA.

Usage
InferenceCountHurdleNegBin$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip variance calculations.


InferenceCountHurdleNegBin$compute_jackknife_estimate()

Hurdle negative-binomial delete-one refits are unstable for jackknife inference; report explicit non-estimability.

Usage
InferenceCountHurdleNegBin$compute_jackknife_estimate(unit = "auto")
Arguments
unit

Deletion unit. Default "auto".


InferenceCountHurdleNegBin$compute_jackknife_bias_estimate()

Report that the jackknife bias estimate is unavailable for hurdle negative-binomial fits because delete-one refits are unstable; see InferenceJackknife for the shared jackknife contract.

Usage
InferenceCountHurdleNegBin$compute_jackknife_bias_estimate(unit = "auto")
Arguments
unit

Deletion unit. Default "auto".


InferenceCountHurdleNegBin$compute_jackknife_std_error()

Report that the jackknife standard error is unavailable for hurdle negative-binomial fits because delete-one refits are unstable; see InferenceJackknife for the shared jackknife contract.

Usage
InferenceCountHurdleNegBin$compute_jackknife_std_error(unit = "auto")
Arguments
unit

Deletion unit. Default "auto".


InferenceCountHurdleNegBin$compute_jackknife_wald_two_sided_pval()

Reports that jackknife-Wald p-values are unavailable here; see InferenceJackknife.

Usage
InferenceCountHurdleNegBin$compute_jackknife_wald_two_sided_pval(
  delta = 0,
  unit = "auto"
)
Arguments
delta

Null treatment-effect value. Default 0.

unit

Deletion unit. Default "auto".


InferenceCountHurdleNegBin$compute_jackknife_wald_confidence_interval()

Reports that jackknife-Wald intervals are unavailable here; see InferenceJackknife.

Usage
InferenceCountHurdleNegBin$compute_jackknife_wald_confidence_interval(
  alpha = 0.05,
  unit = "auto"
)
Arguments
alpha

Significance level. Default 0.05.

unit

Deletion unit. Default "auto".


InferenceCountHurdleNegBin$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCountHurdleNegBin$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Mullahy, J. (1986). "Specification and Testing of Some Modified Count Data Models." Journal of Econometrics, 33(3), 341-365, doi:10.1016/0304-4076(86)90002-3, for the hurdle count-model framework.

See Also

InferenceCountNegBin for the single-part negative binomial model this class's count submodel generalizes to two parts.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountHurdleNegBin$new(seq_des, model_formula = ~ x1)
inf$compute_estimate()


Hurdle Poisson Regression Inference for Count Responses

Description

Fits a hurdle Poisson regression for count responses: a binary hurdle submodel P(Y_i > 0) = \mathrm{logit}^{-1}(X_i^{h\top} \gamma^h) (fit jointly with the count submodel) crossed with a zero-truncated Poisson count submodel for Y_i \mid Y_i > 0: \log E[Y_i \mid Y_i > 0, w_i, x_i] = \beta_0 + \beta_T w_i + x_i^\top \gamma. The hurdle and count submodels may use different covariate formulas (model_formula/model_formula_hurdle). The reported treatment effect is the coefficient from the conditional (truncated, Y > 0) count component, on the log-rate scale, conditional on clearing the hurdle: it is not the effect on the unconditional mean E[Y], which also depends on how treatment shifts the hurdle-crossing probability, under the default estimand = "conditional". likelihood_tier = "full": Wald, gradient, and (bootstrap-calibrated) likelihood-ratio tests are available for the count submodel's treatment coefficient under that estimand; a plain score test is not exposed. Jackknife inference is not supported: delete-one refits of this two-part model are numerically unstable, so compute_jackknife_estimate() and related methods report explicit non-estimability rather than attempting delete-one refits. Unlike InferenceCountHurdleNegBin, the count submodel here assumes Poisson (equidispersion) conditional on clearing the hurdle, with no separate dispersion parameter.

Marginal (unconditional-mean) estimand. Via set_estimand(), this class also supports estimand = "marginal_mean_diff" and "marginal_ratio": the g-computation average, over the empirical covariate distribution, of the model-implied unconditional mean E[Y_i \mid w_i, x_i] = (1 - \pi(x_i)) \cdot \lambda(x_i) / (1 - e^{-\lambda(x_i)}) — the hurdle-crossing probability times the zero-truncated Poisson mean, E[Y \mid Y>0] = \lambda / (1 - e^{-\lambda}) (exact for Poisson: truncating at 0 changes the normalizing constant but not the rate parameter \lambda; see Cameron and Trivedi, Regression Analysis of Count Data, ch. 4.2) — at w_i = 1 vs. w_i = 0. A pure post-fit transform of the same maximum-likelihood fit (no refit), with a delta-method standard error against the sandwich-robust covariance matrix already used for this class's conditional Wald inference. Only "wald"-type inference is available under a marginal estimand.

Super classes

Inference -> InferenceCountZeroAugmentedPoissonAbstract -> InferenceCountHurdlePoisson

Methods

Public methods

+ inherited public methods from InferenceCountZeroAugmentedPoissonAbstract
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_bayesian_bootstrap_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_bootstrap_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_jackknife_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_rand_bootstrap_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_randomization_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_subsampling_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_bayesian_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_bayesian_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_estimate_with_bootstrap_weights()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_gradient_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_gradient_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_bias_estimate()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_estimate()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_std_error()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_wald_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_wald_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_approx_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_approx_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_exact_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_exact_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_m_out_of_n_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_m_out_of_n_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_estimate()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_rand_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_rand_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_rand_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_rand_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_score_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_score_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_sensitivity()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_wald_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_wald_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$get_information_preference()
  • InferenceCountZeroAugmentedPoissonAbstract$get_information_source_used()
  • InferenceCountZeroAugmentedPoissonAbstract$get_last_param_bootstrap_diagnostics()
  • InferenceCountZeroAugmentedPoissonAbstract$get_last_param_bootstrap_estimate_diagnostics()
  • InferenceCountZeroAugmentedPoissonAbstract$get_mod()
  • InferenceCountZeroAugmentedPoissonAbstract$get_summary()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_bayesian_bootstrap_ci_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_bayesian_bootstrap_pval_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_bootstrap_ci_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_bootstrap_pval_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_information_preferences()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_rand_bootstrap_ci_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_rand_bootstrap_pval_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_testing_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_testing_type()
  • InferenceCountZeroAugmentedPoissonAbstract$select_optimal_b_subsampling()
  • InferenceCountZeroAugmentedPoissonAbstract$select_optimal_m_out_of_n_bootstrap()
  • InferenceCountZeroAugmentedPoissonAbstract$set_information_preference()
  • InferenceCountZeroAugmentedPoissonAbstract$set_testing_type()
  • InferenceCountZeroAugmentedPoissonAbstract$supports_rand_pval_for_incidence()
+ inherited public methods from Inference

InferenceCountHurdlePoisson$new()

Initialize inference for the hurdle Poisson model (binary hurdle submodel plus zero-truncated Poisson count submodel); see InferenceCountHurdlePoisson for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferenceCountHurdlePoisson$new(
  des_obj,
  model_formula = NULL,
  model_formula_hurdle = NULL,
  use_rcpp = TRUE,
  verbose = FALSE,
  smart_cold_start_default = NULL,
  optimization_alg = NULL
)
Arguments
des_obj

A completed Design object with a count response.

model_formula

Optional formula for the count submodel.

model_formula_hurdle

Formula for the hurdle submodel. If NULL (default), it uses the same formula as model_formula.

use_rcpp

Logical. If TRUE (default), use the internal Rcpp implementation. If FALSE, use glmmTMB.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.

optimization_alg

Optimization algorithm. Default is dispatched via policy.


InferenceCountHurdlePoisson$compute_estimate()

Fits the hurdle Poisson model. Under the default estimand = "conditional", returns \hat\beta_T, the treatment log-rate coefficient from the zero-truncated count submodel (conditional on clearing the hurdle). Under estimand = "marginal_mean_diff" or "marginal_ratio" (set via set_estimand()), returns the g-computation marginal mean difference or log-scale marginal ratio of the unconditional mean instead — see the class-level @details for the formula. A pure post-fit transform of the same cached fit, no refit.

Usage
InferenceCountHurdlePoisson$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.


InferenceCountHurdlePoisson$compute_asymp_confidence_interval()

Asymptotic confidence interval. Under the conditional estimand, delegates to the shared zero-augmented count-model Wald/ bootstrap-fallback contract; under a marginal estimand, the delta-method interval computed by compute_estimate(). Calls self$compute_estimate() first (not private$shared() directly) so the estimand-aware cache is always current regardless of call order.

Usage
InferenceCountHurdlePoisson$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The significance level (default 0.05).


InferenceCountHurdlePoisson$compute_asymp_two_sided_pval()

Asymptotic two-sided p-value, dispatched exactly as compute_asymp_confidence_interval(); see that method's description for the marginal-estimand path.

Usage
InferenceCountHurdlePoisson$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null treatment effect under the current estimand (default 0).


InferenceCountHurdlePoisson$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCountHurdlePoisson$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Mullahy, J. (1986). "Specification and Testing of Some Modified Count Data Models." Journal of Econometrics, 33(3), 341-365, doi:10.1016/0304-4076(86)90002-3, for the hurdle count-model framework.

See Also

InferenceCountPoisson for the single-part Poisson model this class's count submodel generalizes to two parts; InferenceCountHurdleNegBin for the overdispersion-robust negative-binomial variant (does not support a marginal estimand — the mean-function derivation here is Poisson-specific).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountHurdlePoisson$new(seq_des)
inf$compute_estimate()


One-Likelihood Conditional-Poisson Inference for KK Count Designs

Description

Estimates a treatment log-rate-ratio \beta_T for count outcomes collected under a KK matching-on-the-fly design (DesignSeqOneByOneKK14 or subclass) by maximizing a single combined likelihood that couples a conditional (within-matched-pair, intercept-free) Poisson likelihood for matched subjects with an ordinary Poisson likelihood for reservoir subjects, sharing one treatment coefficient across both pieces. This is the "one-likelihood" alternative to the inverse-variance-weighted combination (...IVWC pattern used elsewhere in the KK family): rather than fitting matched and reservoir models separately and pooling by inverse-variance weights, the treatment coefficient here is estimated jointly from the full combined log-likelihood, and its standard error, score, likelihood-ratio, and gradient statistics are all "design-conservative" – each is the pointwise-wider of the model-based asymptotic quantity and a design-based quantity computed by treating the estimate as a plug-in statistic under InferenceAsymp's z/t machinery, so inference never overstates precision relative to the design alone.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

Estimand. \beta_T, the treatment coefficient in a log-linear (Poisson) mean model E[Y \mid w, x] = \exp(\beta_0 + \beta_T w + x\beta), interpreted as a log rate ratio (equivalently, \exp(\hat\beta_T) is the treatment-vs-control incidence rate ratio).

Model. Matched subjects contribute a conditional-Poisson term that eliminates the pair-specific nuisance intercept by conditioning on the pair total count (removing the need to estimate one intercept per pair); reservoir subjects contribute an ordinary Poisson log-likelihood with a single shared intercept. Both pieces are summed into one combined negative log-likelihood and maximized jointly in (\beta_0, \beta_T, \beta) (see get_cpoisson_combined_hessian_cpp and fast_cpoisson_combined_with_var_cpp for the backend fitting contract). likelihood_tier = "full", so likelihood-ratio, score, and gradient tests and a parametric likelihood bootstrap are all available in addition to the design-conservative Wald path.

Assumptions. Independence of counts across matched pairs and reservoir subjects given covariates; correct log-linear mean specification; a KK matching-on-the-fly design supplying the matched/reservoir partition. No response censoring is supported (checked at construction via assertNoCensoring()).

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceCountKKCondPoissonOneLik

Methods

Public methods

+ inherited public methods from Inference

InferenceCountKKCondPoissonOneLik$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceCountKKCondPoissonOneLik$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceCountKKCondPoissonOneLik$supports_rand_pval_for_incidence()

Usage
InferenceCountKKCondPoissonOneLik$supports_rand_pval_for_incidence()

InferenceCountKKCondPoissonOneLik$compute_rand_two_sided_pval()

Usage
InferenceCountKKCondPoissonOneLik$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceCountKKCondPoissonOneLik$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCountKKCondPoissonOneLik$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kapelner, A. and Krieger, A. (2014). "Matching on-the-fly: A group sequential covariate balanced randomization procedure." arXiv preprint arXiv:1305.6259. (KK14 in REFERENCES.md.)

See Also

Analogous Python API for count models: statsmodels discrete models (ConditionalPoisson, Poisson). Poisson regression (orientation).


GLMM Inference for KK Designs with Count Response

Description

Fits a Poisson GLMM for count responses under a KK matching-on-the-fly design. The random intercept per matched pair is integrated out via Gauss-Hermite quadrature.

When use_rcpp = TRUE (default) the likelihood is maximised by an internal Rcpp routine. Set use_rcpp = FALSE to fall back to glmmTMB.

Model. Y_{ij} \mid b_i \sim \mathrm{Poisson}(\mu_{ij}) with \log \mu_{ij} = X_{ij}'\beta + \beta_T \cdot W_{ij} + b_i, where i indexes matched pairs, j \in \{1, 2\} the two subjects within a pair, W_{ij} is the treatment indicator, and b_i \sim \mathcal{N}(0, \sigma_b^2) is a pair-level random intercept absorbing within-pair correlation induced by matching. \beta_T is a log-rate (log relative risk) treatment effect: \exp(\hat\beta_T) is the estimated rate ratio. The random effect is integrated out of the marginal likelihood by adaptive Gauss-Hermite quadrature rather than a Laplace approximation.

Likelihood tier. likelihood_tier = "full": both Wald (model-based standard error) and likelihood-ratio testing types are available. Because the GLMM likelihood alone does not encode the KK design's matched-pair randomization structure, the likelihood-ratio CI/p-value are conservatively widened/calibrated against the design-aware Wald result (see compute_lik_ratio_confidence_interval()/ compute_lik_ratio_two_sided_pval()) so the model-based test is never anti-conservative relative to the design.

Assumptions. Count response modeled as conditionally Poisson given the random intercept (equidispersion conditional on b_i); pair-level random effects independent across pairs; a KK matching-on-the-fly design.

Super class

Inference -> InferenceCountKKGLMM

Methods

Public methods

+ inherited public methods from Inference

InferenceCountKKGLMM$new()

Initialize a KK Poisson-GLMM inference object for a matched-pair KK design with a count response and prepare the matched-pair random-intercept likelihood machinery; see the class topic for the model.

Usage
InferenceCountKKGLMM$new(
  des_obj,
  model_formula = NULL,
  use_rcpp = TRUE,
  optimization_alg = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed KK matching-on-the-fly Design object (DesignSeqOneByOneKK14 or subclass) with a count response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used.

use_rcpp

Logical. If TRUE (default), maximize the Gauss-Hermite-quadrature marginal likelihood with the internal Rcpp Poisson-GLMM routine; if FALSE, fall back to glmmTMB.

optimization_alg

Optimization algorithm passed to the likelihood maximizer. If NULL (default), an algorithm is dispatched via the package's optimizer policy.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart starting values for the optimizer.


InferenceCountKKGLMM$compute_estimate()

Point estimate of the treatment log-rate coefficient \beta_T from a Poisson GLMM with a matched-pair random intercept, fit by maximizing the Gauss-Hermite-quadrature-integrated marginal likelihood (internal Rcpp routine when use_rcpp = TRUE, else glmmTMB). See the class topic for the model form.

Usage
InferenceCountKKGLMM$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance-component calculations.

Returns

Numeric scalar: the treatment coefficient on the log-rate (link) scale, i.e. \exp(\hat\beta_T) is a rate ratio.


InferenceCountKKGLMM$compute_estimate_with_bootstrap_weights()

Recomputes the KK Poisson-GLMM treatment estimate under nonparametric/ Bayesian-bootstrap subject-or-block weights, refitting the weighted GLMM (compute_weighted_glmm_bootstrap_estimate()). Standard error, degrees of freedom, and the cached summary table are cleared/set to NA/Inf/ NULL since only the point estimate is meaningful under resampling weights. Falls back to the unweighted point estimate when the weights are effectively constant.

Usage
InferenceCountKKGLMM$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Numeric vector of nonnegative bootstrap replicate weights, one per subject or per matched block (KK match structure).

estimate_only

If TRUE, compute only the weighted point estimate (this method never computes a weighted standard error regardless of this argument).

Returns

Numeric scalar treatment-effect estimate (log-rate scale) under the given weights.


InferenceCountKKGLMM$compute_wald_confidence_interval()

Wald confidence interval for the treatment log-rate coefficient: \hat\beta_T \pm t_{1-\alpha/2,\,df}\cdot \hat{se}(\hat\beta_T), using the model-based GLMM standard error. See InferenceAsymp for the shared contract.

Usage
InferenceCountKKGLMM$compute_wald_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.

Returns

A length-2 numeric vector c(lower, upper) on the log-rate scale.


InferenceCountKKGLMM$compute_wald_two_sided_pval()

Two-sided Wald p-value for H_0: \beta_T = \code{delta} vs. H_1: \beta_T \neq \code{delta}, using the model-based GLMM standard error.

Usage
InferenceCountKKGLMM$compute_wald_two_sided_pval(delta = 0)
Arguments
delta

The null value of \beta_T to test against; 0 (the default) tests for any treatment effect at all.

Returns

Numeric scalar p-value in [0, 1].


InferenceCountKKGLMM$compute_lik_ratio_confidence_interval()

Likelihood-ratio confidence interval for the treatment log-rate coefficient, inverting the GLMM's profile likelihood-ratio test against \chi^2_1. Because the GLMM likelihood does not itself account for the KK matched-pair design's randomization structure, this interval is conservatively widened to be at least as wide as the design-aware Wald interval (compute_wald_confidence_interval()) via .conservative_kk_onelik_ci() – guarding against the model-based interval being anti-conservative relative to the design.

Usage
InferenceCountKKGLMM$compute_lik_ratio_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.

Returns

A length-2 numeric vector c(lower, upper) on the log-rate scale.


InferenceCountKKGLMM$compute_lik_ratio_two_sided_pval()

Two-sided likelihood-ratio p-value for H_0: \beta_T = \code{delta}, from the GLMM's profile likelihood-ratio test referred to \chi^2_1. As with compute_lik_ratio_confidence_interval(), this is conservatively calibrated (via .conservative_kk_onelik_pval()) against the design-aware Wald p-value so the model-based test cannot be anti-conservative relative to the KK matched-pair design.

Usage
InferenceCountKKGLMM$compute_lik_ratio_two_sided_pval(delta = 0)
Arguments
delta

The null value of \beta_T to test against; 0 (the default) tests for any treatment effect at all.

Returns

Numeric scalar p-value in [0, 1].


InferenceCountKKGLMM$compute_asymp_confidence_interval()

Asymptotic confidence interval, dispatching to compute_wald_confidence_interval() or compute_lik_ratio_confidence_interval() depending on self$get_testing_type() (defaults to Wald if the testing type is neither). See InferenceAsymp for the shared testing-type dispatch contract.

Usage
InferenceCountKKGLMM$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.

Returns

A length-2 numeric vector c(lower, upper) on the log-rate scale.


InferenceCountKKGLMM$compute_asymp_two_sided_pval()

Asymptotic two-sided p-value, dispatching to compute_wald_two_sided_pval() or compute_lik_ratio_two_sided_pval() depending on self$get_testing_type() (defaults to Wald if the testing type is neither). See InferenceAsymp for the shared testing-type dispatch contract.

Usage
InferenceCountKKGLMM$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null value of \beta_T to test against; 0 (the default) tests for any treatment effect at all.

Returns

Numeric scalar p-value in [0, 1].


InferenceCountKKGLMM$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCountKKGLMM$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kapelner, A. and Krieger, A. M. (2014). Matching on-the-fly: Sequential allocation with higher power and efficiency. Biometrics, 70(2), 378-388. doi:10.1111/biom.12148. (KK14 in REFERENCES.md.)

See Also

Analogous Python API for Poisson/count GLMs: statsmodels discrete models. Generalized linear model and Gauss-Hermite quadrature (orientation).

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'count')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountKKGLMM$new(seq_des)
inf$compute_estimate()


KK Hurdle Poisson IVWC Inference for Count Responses

Description

Inverse-variance weighted combined inference for count responses under a KK matching-on-the-fly design. The matched-pair component is fit with a hurdle-Poisson mixed model using pair random intercepts, and the reservoir component is fit with an ordinary Poisson log-link regression. The reported treatment effect is on the log-rate scale.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceCountKKHurdlePoissonIVWC

Methods

Public methods

+ inherited public methods from Inference

InferenceCountKKHurdlePoissonIVWC$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceCountKKHurdlePoissonIVWC$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceCountKKHurdlePoissonIVWC$supports_rand_pval_for_incidence()

Usage
InferenceCountKKHurdlePoissonIVWC$supports_rand_pval_for_incidence()

InferenceCountKKHurdlePoissonIVWC$compute_rand_two_sided_pval()

Usage
InferenceCountKKHurdlePoissonIVWC$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceCountKKHurdlePoissonIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCountKKHurdlePoissonIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


KK Hurdle-Poisson Combined-Likelihood Inference for Count Responses

Description

Fits a two-part hurdle-Poisson model to a KK matching-on-the-fly count design by maximizing a single combined likelihood over the matched pairs and reservoir subjects jointly, rather than fitting matched and reservoir submodels separately and combining them afterward (contrast with the inverse-variance-weighted-combination sibling, InferenceCountKKHurdlePoissonIVWC).

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

Model. For subject i with count response Y_i \ge 0, a hurdle model factors the likelihood into (1) a binary zero-vs-positive part P(Y_i = 0) = 1 - \pi_i, \mathrm{logit}(\pi_i) = X_i'\gamma, and (2) a zero-truncated Poisson part for the positive counts, Y_i \mid Y_i > 0 \sim \text{Poisson}_{+}(\lambda_i), \log \lambda_i = X_i'\beta + \beta_T W_i, where W_i is the treatment indicator and \beta_T is the treatment log-rate coefficient for the positive-count submodel (the estimand returned by compute_estimate()). Unlike a standard hurdle model fit by maximum likelihood on i.i.d. rows, this class's negative log-likelihood combines the matched-pair rows and reservoir rows of a KK design into one objective (see private$fit_combined_hurdle()), so the fitted \beta_T and its curvature already reflect the design's matched/reservoir structure rather than treating all subjects as exchangeable.

Likelihood tier. likelihood_tier = "full": Wald, score, likelihood-ratio, and gradient testing types are all available (see get_testing_type()). Because a hurdle-Poisson combined likelihood does not by itself encode the KK design's finite-sample matched-pair randomization distribution, the score/likelihood-ratio/gradient confidence intervals and p-values are computed twice — once from this class's design-aware asymptotic variance (the same Wald-type calculation used by compute_wald_confidence_interval()) and once from the generic likelihood-based calculation inherited from InferenceAsympLik — and the wider interval / larger p-value of the two is returned (see .conservative_kk_onelik_ci()/.conservative_kk_onelik_pval()), so the model-based test is never anti-conservative relative to the design. The Wald confidence interval and p-value fall back to the BayesianBootstrap component's bootstrap distribution when the model-based standard error is unavailable or non-finite (e.g. a boundary/separation fit).

Assumptions. Independence across matched pairs and reservoir subjects conditional on covariates; correct specification of the logistic hurdle and log-linear positive-count submodels; a KK matching-on-the-fly design (DesignSeqOneByOneKK14 or subclass) supplying the matched/reservoir partition. No response censoring is supported (checked at construction via assertNoCensoring()).

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceCountKKHurdlePoissonOneLik

Methods

Public methods

+ inherited public methods from Inference

InferenceCountKKHurdlePoissonOneLik$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceCountKKHurdlePoissonOneLik$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceCountKKHurdlePoissonOneLik$supports_rand_pval_for_incidence()

Usage
InferenceCountKKHurdlePoissonOneLik$supports_rand_pval_for_incidence()

InferenceCountKKHurdlePoissonOneLik$compute_rand_two_sided_pval()

Usage
InferenceCountKKHurdlePoissonOneLik$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceCountKKHurdlePoissonOneLik$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCountKKHurdlePoissonOneLik$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Mullahy, J. (1986). "Specification and Testing of Some Modified Count Data Models." Journal of Econometrics, 33(3), 341-365. doi:10.1016/0304-4076(86)90002-3. (Mullahy1986 in REFERENCES.md.)

See Also

Analogous Python API for hurdle/zero-truncated count models: statsmodels discrete models (HurdleCountModel). Poisson regression (orientation).


Count-Specific Likelihood Inference

Description

Component source (CountLikelihoodPlumbingSource) for the count-based likelihood families (Poisson, Negative Binomial, Zero-Inflated, Hurdle): centralizes count-specific parameter packing, warm starts, and likelihood dispatch. Composed by every count class through the registered CountLikelihoodPlumbing component.

Computes the treatment-effect estimate using the underlying count likelihood model. Concrete subclasses fit a Poisson, negative binomial, zero-inflated, hurdle, or combined likelihood model and cache the treatment coefficient for related InferenceCountLikelihood p-value and confidence-interval methods.

Usage

CountLikelihoodPlumbingSource

Negative Binomial Regression Inference for Count Responses

Description

Fits a negative binomial regression for count responses: Y_i \mid w_i, x_i \sim \mathrm{NegBin}(\mu_i, \theta), \log \mu_i = \beta_0 + \beta_T w_i + x_i^\top \gamma, \mathrm{Var}(Y_i) = \mu_i + \mu_i^2 / \theta, jointly maximizing over the regression coefficients and the dispersion parameter \theta (fast_neg_bin_cpp/fast_neg_bin_with_var_cpp). \hat\beta_T is a log-rate-ratio: \exp(\hat\beta_T) is the estimated rate ratio. Unlike InferenceCountPoisson, the negative-binomial model allows overdispersion (\mathrm{Var}(Y_i) > E[Y_i]) via \theta; smaller \theta indicates more overdispersion, and the model converges to Poisson as \theta \to \infty. likelihood_tier = "full": Wald, score, gradient, and likelihood-ratio tests are all available, plus parametric-likelihood bootstrap calibration of the likelihood-ratio test (simulating new responses from \mathrm{NegBin}(\hat\mu_i, \hat\theta) under the null). Jackknife inference is not supported: delete-one refits of a jointly-estimated dispersion parameter are numerically unstable, so compute_jackknife_estimate() and related methods report explicit non-estimability rather than attempting delete-one refits. Validity requires the negative-binomial mean-variance relationship to hold and the usual correctly-specified-linear-predictor-on-the-log-scale assumption.

Super class

Inference -> InferenceCountNegBin

Methods

Public methods

+ inherited public methods from Inference

InferenceCountNegBin$new()

Initialize inference for the negative binomial regression model Y_i \mid w_i, x_i \sim \mathrm{NegBin}(\mu_i, \theta), \log \mu_i = \beta_0 + \beta_T w_i + x_i^\top \gamma; see InferenceCountNegBin for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferenceCountNegBin$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL,
  optimization_alg = NULL
)
Arguments
des_obj

A completed Design object with a count response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart optimizer start values by default.

optimization_alg

Optimization algorithm to use. Default is dispatched via policy.


InferenceCountNegBin$compute_estimate_with_bootstrap_weights()

Refits the negative binomial model with subject/block-level weights applied to the fitting log-likelihood (Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights()) via fast_neg_bin_weighted_cpp, and returns the reweighted log-rate-ratio estimate \hat\beta_T^{(w)}. If the weighted negative-binomial fit fails to converge, falls back to a weighted Poisson GLM (stats::glm(family = poisson())) as an estimating-equation-consistent point estimate of the same mean structure (this fallback does not itself estimate \theta, so no standard error is computed in that path); no standard error is computed in either path (s_beta_hat_T is always NA).

Usage
InferenceCountNegBin$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip variance calculations.


InferenceCountNegBin$compute_jackknife_estimate()

Negative-binomial delete-one refits are unstable for jackknife inference; report explicit non-estimability.

Usage
InferenceCountNegBin$compute_jackknife_estimate(unit = "auto")
Arguments
unit

Deletion unit. Default "auto".


InferenceCountNegBin$compute_jackknife_bias_estimate()

Report that the jackknife bias estimate is unavailable for negative-binomial fits when delete-one refits are not stable; see InferenceJackknife for the shared jackknife contract.

Usage
InferenceCountNegBin$compute_jackknife_bias_estimate(unit = "auto")
Arguments
unit

Deletion unit. Default "auto".


InferenceCountNegBin$compute_jackknife_std_error()

Report that the jackknife standard error is unavailable for negative-binomial fits when delete-one refits are not stable; see InferenceJackknife for the shared jackknife contract.

Usage
InferenceCountNegBin$compute_jackknife_std_error(unit = "auto")
Arguments
unit

Deletion unit. Default "auto".


InferenceCountNegBin$compute_jackknife_wald_two_sided_pval()

Reports that jackknife-Wald p-values are unavailable here; see InferenceJackknife.

Usage
InferenceCountNegBin$compute_jackknife_wald_two_sided_pval(
  delta = 0,
  unit = "auto"
)
Arguments
delta

Null treatment-effect value. Default 0.

unit

Deletion unit. Default "auto".


InferenceCountNegBin$compute_jackknife_wald_confidence_interval()

Reports that jackknife-Wald intervals are unavailable here; see InferenceJackknife.

Usage
InferenceCountNegBin$compute_jackknife_wald_confidence_interval(
  alpha = 0.05,
  unit = "auto"
)
Arguments
alpha

Significance level. Default 0.05.

unit

Deletion unit. Default "auto".


InferenceCountNegBin$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCountNegBin$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Cameron, A. C., and Trivedi, P. K. (2013). Regression Analysis of Count Data (2nd ed.). Cambridge University Press, for the negative binomial regression model and its maximum-likelihood theory.

See Also

Comparable Python API: statsmodels discrete models (NegativeBinomial). See also: Negative binomial distribution (Wikipedia).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountNegBin$new(seq_des)
inf$compute_estimate()


Poisson Regression Inference for Count Responses

Description

Fits a Poisson log-link regression for count responses: Y_i \mid w_i, x_i \sim \mathrm{Poisson}(\mu_i), \log \mu_i = \beta_0 + \beta_T w_i + x_i^\top \gamma (fast_poisson_regression_cpp/ fast_poisson_regression_with_var_cpp). \hat\beta_T is a log-rate-ratio: \exp(\hat\beta_T) is the estimated rate ratio. likelihood_tier = "full": Wald, score, gradient, and likelihood-ratio tests are all available, plus parametric-likelihood bootstrap calibration of the likelihood-ratio test (simulating new Poisson responses under the null).

Design-conservative testing. Every asymptotic/likelihood test method on this class does not report the raw model-based result directly. Instead, it also computes a design-based jackknife-Wald test (compute_jackknife_wald_two_sided_pval()/ compute_jackknife_wald_confidence_interval(), which do not assume the Poisson mean-variance relationship) and combines the two conservatively: p-values report \max of the model-based and design-based p-values, and confidence intervals report the union of the model-based and design-based intervals. This guards against the model-based test being anti-conservative when the Poisson equidispersion assumption (\mathrm{Var}(Y_i) = E[Y_i]) fails — a real risk for count data, which is frequently overdispersed (see InferenceCountNegBin for a model that estimates dispersion directly instead). If either component is unavailable, the available one is used alone; if neither is available, the result is NA. Validity requires the usual correctly-specified linear predictor on the log scale; unlike the raw Poisson likelihood alone, this class's actual reported inference degrades gracefully (rather than becoming anti-conservative) under mean-variance misspecification.

Estimand. Composes MarginalEstimand (set_estimand()/get_estimand()/get_supported_estimands()). Under the default estimand = "conditional", \hat\beta_T is the log-rate-ratio above. Under estimand = "marginal_mean_diff", the reported quantity is instead the g-computation marginal rate difference \frac{1}{n}\sum_i \{\exp(\hat\beta_0 + \hat\beta_T + X_i^\top \hat\gamma) - \exp(\hat\beta_0 + X_i^\top \hat\gamma)\}. Under estimand = "marginal_ratio", the log of the corresponding marginal rate ratio — which is numerically identical to the conditional \hat\beta_T for this family: because the log link is linear in w_i with no treatment-by-covariate interaction term, every subject's treated-vs-control mean ratio is \exp(\hat\beta_0 + \hat\beta_T + X_i^\top\hat\gamma) / \exp(\hat\beta_0 + X_i^\top\hat\gamma) = \exp(\hat\beta_T) exactly, so averaging over subjects before or after taking the ratio makes no difference. "marginal_ratio" is provided for estimand-API consistency with the other model families, not because it differs numerically from "conditional" here; "marginal_mean_diff" is the estimand where g-computation actually changes the reported number for a Poisson GLM, since a difference (unlike a ratio) does not collapse under a nonlinear (log) mean function. Because there is no latent submodel for this family (unlike e.g. InferenceCountZeroInflatedPoisson's excess-zero mixture), the marginal mean function is exactly the model's own fitted mean; no separate standardization step beyond the g-computation average is needed. Standard errors under a marginal estimand use the delta method against the model's coefficient covariance (degrees of freedom Inf), including in the design-conservative union/max combination above (the design-based jackknife-Wald component also refits under the active estimand); testing_type is restricted to "wald" whenever the estimand is non-conditional. The underlying model fit is identical regardless of estimand — switching estimand is a pure post-fit transform, never a refit.

Super class

Inference -> InferenceCountPoisson

Methods

Public methods

+ inherited public methods from Inference

InferenceCountPoisson$new()

Initialize inference for the Poisson regression model Y_i \mid w_i, x_i \sim \mathrm{Poisson}(\mu_i), \log \mu_i = \beta_0 + \beta_T w_i + x_i^\top \gamma; see InferenceCountPoisson for the model form and the design-conservative testing mechanism. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferenceCountPoisson$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL,
  harden = TRUE
)
Arguments
des_obj

A completed Design object with a count response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values by default.

harden

Whether to apply robustness measures.


InferenceCountPoisson$compute_estimate()

Fits the Poisson regression model by maximum likelihood. Under the default estimand = "conditional", returns \hat\beta_T, the treatment log-rate-ratio. Under estimand = "marginal_mean_diff"/"marginal_ratio" (set via set_estimand()), returns the g-computation marginal rate difference/log-rate-ratio instead — see the class-level @details for the formula. The underlying model fit is identical either way (a pure post-fit transform of the same cached fit, no refit).

Usage
InferenceCountPoisson$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.


InferenceCountPoisson$compute_asymp_confidence_interval()

Design-conservative confidence interval for \beta_T using whichever test type is configured (private$testing_type: "wald", "score", "gradient", or "lik_ratio"); see InferenceCountPoisson for the union-with-jackknife-Wald combination rule.

Usage
InferenceCountPoisson$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level. Default 0.05.


InferenceCountPoisson$compute_asymp_two_sided_pval()

Design-conservative two-sided p-value for H_0: \beta_T = \code{delta} using whichever test type is configured (private$testing_type); see InferenceCountPoisson for the max-with-jackknife-Wald combination rule.

Usage
InferenceCountPoisson$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect. Default 0.


InferenceCountPoisson$compute_wald_confidence_interval()

Wald confidence interval for \beta_T using the fitted Poisson model's Fisher-information-based standard error, unioned with the design-based jackknife-Wald interval; see InferenceCountPoisson for the combination rule and InferenceAsymp for the underlying Wald contract.

Usage
InferenceCountPoisson$compute_wald_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level. Default 0.05.


InferenceCountPoisson$compute_wald_two_sided_pval()

Wald test of H_0: \beta_T = \code{delta} using the fitted Poisson model's Fisher-information-based standard error, taking the max with the design-based jackknife-Wald p-value; see InferenceCountPoisson for the combination rule.

Usage
InferenceCountPoisson$compute_wald_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect. Default 0.


InferenceCountPoisson$compute_score_confidence_interval()

Score-test confidence interval for \beta_T (inverting the Poisson score test at each candidate null, no full re-fit needed at the observed information), unioned with the design-based jackknife-Wald interval; see InferenceCountPoisson for the combination rule.

Usage
InferenceCountPoisson$compute_score_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level. Default 0.05.


InferenceCountPoisson$compute_score_two_sided_pval()

Score test of H_0: \beta_T = \code{delta}, taking the max with the design-based jackknife-Wald p-value; see InferenceCountPoisson for the combination rule.

Usage
InferenceCountPoisson$compute_score_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect. Default 0.


InferenceCountPoisson$compute_lik_ratio_confidence_interval()

Likelihood-ratio-test confidence interval for \beta_T (test inversion, requiring a null refit at each candidate value), unioned with the design-based jackknife-Wald interval; see InferenceCountPoisson for the combination rule.

Usage
InferenceCountPoisson$compute_lik_ratio_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level. Default 0.05.


InferenceCountPoisson$compute_lik_ratio_two_sided_pval()

Likelihood-ratio test of H_0: \beta_T = \code{delta}, taking the max with the design-based jackknife-Wald p-value; see InferenceCountPoisson for the combination rule.

Usage
InferenceCountPoisson$compute_lik_ratio_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect. Default 0.


InferenceCountPoisson$compute_gradient_confidence_interval()

Gradient-test confidence interval for \beta_T (a score-test variant using the observed rather than expected information), unioned with the design-based jackknife-Wald interval; see InferenceCountPoisson for the combination rule.

Usage
InferenceCountPoisson$compute_gradient_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level. Default 0.05.


InferenceCountPoisson$compute_gradient_two_sided_pval()

Gradient test of H_0: \beta_T = \code{delta}, taking the max with the design-based jackknife-Wald p-value; see InferenceCountPoisson for the combination rule.

Usage
InferenceCountPoisson$compute_gradient_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect. Default 0.


InferenceCountPoisson$compute_lik_ratio_bootstrap_two_sided_pval()

Parametric-likelihood-bootstrap-calibrated likelihood-ratio test of H_0: \beta_T = \code{delta} (simulating new Poisson responses from the null-constrained fit to calibrate the LR statistic's null distribution), taking the max with the design-based jackknife-Wald p-value; see InferenceCountPoisson for the combination rule.

Usage
InferenceCountPoisson$compute_lik_ratio_bootstrap_two_sided_pval(
  delta = 0,
  B = 199,
  show_progress = FALSE,
  min_number_usable_samples = 5L,
  max_attempts_per_replicate = 2L
)
Arguments
delta

Null treatment effect. Default 0.

B

Number of bootstrap replicates.

show_progress

Whether to show progress.

min_number_usable_samples

Minimum usable bootstrap samples.

max_attempts_per_replicate

Maximum attempts per replicate.


InferenceCountPoisson$compute_lik_ratio_bootstrap_confidence_interval()

Parametric-likelihood-bootstrap-calibrated likelihood-ratio confidence interval for \beta_T (test inversion using the bootstrap-calibrated null distribution), unioned with the design-based jackknife-Wald interval; see InferenceCountPoisson for the combination rule.

Usage
InferenceCountPoisson$compute_lik_ratio_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 199,
  show_progress = FALSE,
  min_number_usable_samples = 5L,
  max_attempts_per_replicate = 2L,
  root_tolerance = NULL,
  max_root_iterations = 8L
)
Arguments
alpha

Significance level. Default 0.05.

B

Number of bootstrap replicates.

show_progress

Whether to show progress.

min_number_usable_samples

Minimum usable bootstrap samples.

max_attempts_per_replicate

Maximum attempts per replicate.

root_tolerance

Root tolerance.

max_root_iterations

Maximum root iterations.


InferenceCountPoisson$compute_estimate_with_bootstrap_weights()

Refits the Poisson model with subject/block-level weights applied to the fitting log-likelihood (Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights()) via fast_poisson_regression_weighted_cpp, and returns the reweighted log-rate-ratio estimate \hat\beta_T^{(w)}. Uses the same QR column-dropping hardening as the unweighted fit; a hardened fit with a non-finite treatment coefficient is cached as nonestimable and returns NA. The weighted refit always targets the conditional treatment coefficient, whatever the active estimand (it is not estimand-aware).

Usage
InferenceCountPoisson$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Row weights for the bootstrap sample.

estimate_only

If TRUE, skip variance calculations.


InferenceCountPoisson$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCountPoisson$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Cameron, A. C., and Trivedi, P. K. (2013). Regression Analysis of Count Data (2nd ed.). Cambridge University Press, for the Poisson regression model and its maximum-likelihood theory.

See Also

Comparable Python API: statsmodels discrete models (Poisson). See also: Poisson regression (Wikipedia).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountPoisson$new(seq_des)
inf$compute_estimate()


inf$set_seed(1)
inf$compute_lik_ratio_bootstrap_two_sided_pval(delta = 0, B = 9, show_progress = FALSE)


GEE Inference for KK Designs with Count Response

Description

Fits a Generalized Estimating Equations (GEE) model with a Poisson family and log link, \log E[Y_i \mid x_i] = x_i^\top\beta, for count responses under a KK matching-on-the-fly design, using an exchangeable working correlation structure where each cluster is either a matched pair (2 members) or a reservoir singleton (1 member) — see $compute_estimate()'s method-level documentation for the full fitting contract (internal Rcpp solver vs. geepack fallback, hardening/retry behavior). GEE is used here purely to fit one marginal model jointly across matched-pair and reservoir subjects while accounting for the within-pair correlation the matching induces, not as a longitudinal/repeated-measures tool. Inference is quasi-likelihood/estimating-equation based (likelihood_tier = "quasi"): standard errors are GEE sandwich (robust) standard errors, not model-likelihood-based.

Super class

Inference -> InferenceCountPoissonKKGEE

Methods

Public methods

+ inherited public methods from Inference

InferenceCountPoissonKKGEE$new()

Initialize KK count-response GEE inference, validate the matched/reservoir design, and prepare the exchangeable-working-correlation Poisson (log-link) GEE fitting machinery used by InferenceCountPoissonKKGEE.

Usage
InferenceCountPoissonKKGEE$new(
  des_obj,
  model_formula = NULL,
  use_rcpp = TRUE,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with a count response.

model_formula

Optional formula for covariate adjustment.

use_rcpp

Whether to use the internal Rcpp GEE solver (TRUE, default) with automatic fallback to geepack::geeglm on failure, or always use geepack::geeglm directly (FALSE).

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.


InferenceCountPoissonKKGEE$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCountPoissonKKGEE$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Liang, K.-Y., and Zeger, S. L. (1986). "Longitudinal Data Analysis Using Generalized Linear Models." Biometrika, 73(1), 13-22, doi:10.1093/biomet/73.1.13, for the GEE estimating-equation framework and sandwich variance estimator used here.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'count')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountPoissonKKGEE$new(seq_des)
inf$compute_estimate()


Quasi-Poisson Regression Inference for Count Responses

Description

Fits a Poisson log-link mean model, \log E[Y_i \mid x_i] = x_i^\top\beta, for count responses using the treatment indicator and, optionally, all recorded covariates as predictors, via fast_quasipoisson_regression_with_var_cpp — see that page for the full model and the Pearson-dispersion-scaled ("quasi-Poisson") variance formula, \widehat{\mathrm{Var}}(\hat\beta_k) = \hat\phi\,[(X^\top \hat{W}X)^{-1}]_{kk}, which corrects standard errors for overdispersion (\mathrm{Var}(Y_i) > E[Y_i]) relative to the strict Poisson assumption without changing the point estimate \hat\beta. This class has no likelihood-ratio/score/gradient testing capability (likelihood_tier = "quasi"): the dispersion-scaled quasi-likelihood is not a normalized model likelihood, so only Wald inference is available. Rank-deficient covariate columns are dropped automatically before fitting (via private$fit_with_hardened_qr_column_dropping()).

Estimand. Composes MarginalEstimand (set_estimand()/get_estimand()/get_supported_estimands()). Under the default estimand = "conditional", \hat\beta_T is the treatment log-rate-ratio. Under "marginal_mean_diff" it is the g-computed difference in the average fitted count under treatment vs. control, \frac{1}{n}\sum_i \{\exp(x_{i1}^\top\hat\beta) - \exp(x_{i0}^\top\hat\beta)\}, with every subject plugged in at treatment 1 and 0. "marginal_ratio" is the log of the corresponding ratio, which for this log-link family equals the conditional \hat\beta_T exactly (there is no treatment-by-covariate term), so it is offered for estimand-API consistency; its delta-method SE equals the conditional SE. Under a marginal estimand the standard error is the delta-method SE against the dispersion-scaled coefficient covariance \hat\phi (X^\top \hat W X)^{-1}, i.e. the plain-Poisson marginal SE inflated by \sqrt{\hat\phi}, with a normal reference (degrees of freedom Inf). Switching the estimand is a pure post-fit transform of the cached fit, never a refit. The Bayesian-bootstrap weighted refit is not estimand-aware: it always targets the conditional coefficient.

Super class

Inference -> InferenceCountQuasiPoisson

Methods

Public methods

+ inherited public methods from Inference

InferenceCountQuasiPoisson$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize a quasi-Poisson regression inference object for a completed design with a count, uncensored response.

Usage
InferenceCountQuasiPoisson$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL,
  harden = TRUE
)
Arguments
des_obj

A completed Design object with a count response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.

harden

Whether to apply robustness measures.


InferenceCountQuasiPoisson$compute_estimate()

Computes the quasi-Poisson point estimate via fast_quasipoisson_regression_with_var_cpp (see class documentation for the full model). Under the default estimand = "conditional" this is the treatment coefficient \hat\beta_T; under "marginal_mean_diff" or "marginal_ratio" (set via set_estimand()) it is the g-computed marginal mean difference / log ratio, a pure post-fit transform of the same cached fit (no refit). Rank-deficient covariate columns are dropped before fitting.

Usage
InferenceCountQuasiPoisson$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance calculations.


InferenceCountQuasiPoisson$compute_estimate_with_bootstrap_weights()

Recomputes the Poisson-mean-model treatment estimate under subject/block bootstrap weights (via fast_poisson_regression_weighted_cpp), used by the Bayesian bootstrap and related weighted-resampling machinery; see InferenceBayesianBootstrap. When estimate_only = FALSE, also computes a weighted Pearson-dispersion-scaled standard error (fixed 2026-09-07 – previously always NA regardless of estimate_only, which starved the Bayesian-bootstrap studentized/BCa variants of a per-replicate SE and left them NA on the large majority of calls). The weighted refit always targets the conditional treatment coefficient, whatever the active estimand (it is not estimand-aware).

Usage
InferenceCountQuasiPoisson$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip the dispersion-correction computation.


InferenceCountQuasiPoisson$compute_asymp_confidence_interval()

Computes a 1-\alpha level Wald confidence interval for the active estimand. Under the default conditional estimand this is the quasi-Poisson treatment coefficient \hat\beta_T with the Pearson-dispersion-scaled standard error from fast_quasipoisson_regression_with_var_cpp (see class documentation); under a marginal estimand it is the g-computed functional with its delta-method standard error. Both use a normal reference. See InferenceAsymp for the shared asymptotic confidence-interval contract this delegates to.

Usage
InferenceCountQuasiPoisson$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Confidence level.


InferenceCountQuasiPoisson$compute_asymp_two_sided_pval()

Computes a two-sided Wald p-value testing H_0: \beta_T = \code{delta} under the conditional estimand, or that the active marginal functional equals delta under a marginal estimand, from the same standard error used by $compute_asymp_confidence_interval() (normal reference). See InferenceAsymp for the shared asymptotic two-sided p-value contract this delegates to.

Usage
InferenceCountQuasiPoisson$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect value.


InferenceCountQuasiPoisson$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCountQuasiPoisson$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountQuasiPoisson$new(seq_des)
inf$compute_estimate()


Robust (Sandwich-Variance) Poisson Regression Inference for Count Responses

Description

Fits the same Poisson log-link mean model as InferenceCountPoisson (point estimate via fast_poisson_regression_cpp, maximum likelihood), but computes standard errors via a Huber-White (Eicker-Huber-White) sandwich estimator instead of the model-based Poisson Fisher information or the quasi-Poisson dispersion scaling used by InferenceCountQuasiPoisson: \widehat{\mathrm{Var}}(\hat\beta) = B\,M\,B, with "bread" B = (X^\top \hat W X)^{-1} (the Poisson Fisher information at \hat\beta) and "meat" M = X^\top \mathrm{diag}((y_i-\hat\mu_i)^2) X (the empirical score outer product), via robust_sandwich_variance_from_xtwx(). This is robust to arbitrary mean-variance misspecification (not just proportional overdispersion), at the cost of somewhat higher variance in the SE estimate itself for small samples. This class has no likelihood-ratio/ score/gradient testing capability (likelihood_tier = "quasi"): only Wald inference is available. Rank-deficient covariate columns are dropped automatically before fitting.

Super class

Inference -> InferenceCountRobustPoisson

Methods

Public methods

+ inherited public methods from Inference

InferenceCountRobustPoisson$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize a robust (sandwich-variance) Poisson regression inference object for a completed design with a count, uncensored response.

Usage
InferenceCountRobustPoisson$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with a count response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart starting values for the optimizer.


InferenceCountRobustPoisson$compute_estimate()

Computes the Poisson treatment coefficient \hat\beta_T via fast_poisson_regression_cpp (see class documentation for the sandwich-variance model). Rank-deficient covariate columns are dropped before fitting.

Usage
InferenceCountRobustPoisson$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance calculations.


InferenceCountRobustPoisson$compute_estimate_with_bootstrap_weights()

Recomputes the Poisson-mean-model treatment estimate under subject/block bootstrap weights (via fast_poisson_regression_weighted_cpp), used by the Bayesian bootstrap and related weighted-resampling machinery; see InferenceBayesianBootstrap. When estimate_only = FALSE, also computes a weighted Huber-White sandwich standard error (fixed 2026-09-07 – previously always NA regardless of estimate_only, which starved the Bayesian-bootstrap studentized/BCa variants of a per-replicate SE and left them NA on the large majority of calls).

Usage
InferenceCountRobustPoisson$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip the sandwich-variance computation.


InferenceCountRobustPoisson$compute_asymp_confidence_interval()

Computes a 1-\alpha level confidence interval for the robust Poisson treatment coefficient \hat\beta_T, using the Huber-White sandwich standard error (see class documentation). See InferenceAsymp for the shared asymptotic confidence-interval contract this delegates to.

Usage
InferenceCountRobustPoisson$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Confidence level.


InferenceCountRobustPoisson$compute_asymp_two_sided_pval()

Computes a two-sided Wald p-value testing H_0: \beta_T = \code{delta}, from the same Huber-White sandwich standard error used by $compute_asymp_confidence_interval(). See InferenceAsymp for the shared asymptotic two-sided p-value contract this delegates to.

Usage
InferenceCountRobustPoisson$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect value.


InferenceCountRobustPoisson$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCountRobustPoisson$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountRobustPoisson$new(seq_des)
inf$compute_estimate()


Zero-Inflated Negative Binomial Regression Inference for Count Responses

Description

Fits a zero-inflated negative binomial regression for count responses: a binary excess-zero submodel P(\text{structural zero}_i) = \mathrm{logit}^{-1}(X_i^{h\top} \gamma^h) mixed with a (non-truncated) negative-binomial count submodel \log E[Y_i \mid \text{not structural zero}, w_i, x_i] = \beta_0 + \beta_T w_i + x_i^\top \gamma, \mathrm{Var}(Y_i \mid \text{not structural zero}) = \mu_i + \mu_i^2 / \theta. Unlike a hurdle model, zero counts can arise from either the structural-zero mechanism or from an ordinary negative-binomial draw of 0. The hurdle and count submodels may use different covariate formulas (model_formula/model_formula_zero). The reported treatment effect is the coefficient from the conditional count component, on the log-rate scale, conditional on the response coming from the count process, not the excess-zero-inflation mechanism: it is not the effect on the unconditional mean E[Y], which also depends on how treatment shifts the excess-zero probability. A marginal (unconditional-mean) estimand is not yet implemented for this class (see marginal_estimand_report.md). likelihood_tier = "full": Wald, gradient, score, and (bootstrap-calibrated) likelihood-ratio tests are all available for the count submodel's treatment coefficient (unlike the Poisson variant, this class's private get_supported_testing_types_impl() includes "score"). Jackknife inference is not supported: delete-one refits of this two-part mixture model with a jointly-estimated dispersion parameter are numerically unstable, so compute_jackknife_estimate() and related methods report explicit non-estimability rather than attempting delete-one refits.

Super classes

Inference -> InferenceCountZeroAugmentedPoissonAbstract -> InferenceCountZeroInflatedNegBin

Methods

Public methods

+ inherited public methods from InferenceCountZeroAugmentedPoissonAbstract
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_bayesian_bootstrap_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_bootstrap_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_jackknife_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_rand_bootstrap_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_randomization_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_subsampling_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_asymp_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_asymp_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_bayesian_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_bayesian_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_estimate()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_estimate_with_bootstrap_weights()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_gradient_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_gradient_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_bias_estimate()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_estimate()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_std_error()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_wald_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_wald_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_approx_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_approx_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_exact_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_exact_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_m_out_of_n_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_m_out_of_n_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_estimate()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_rand_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_rand_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_rand_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_rand_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_score_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_score_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_sensitivity()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_wald_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_wald_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$get_information_preference()
  • InferenceCountZeroAugmentedPoissonAbstract$get_information_source_used()
  • InferenceCountZeroAugmentedPoissonAbstract$get_last_param_bootstrap_diagnostics()
  • InferenceCountZeroAugmentedPoissonAbstract$get_last_param_bootstrap_estimate_diagnostics()
  • InferenceCountZeroAugmentedPoissonAbstract$get_mod()
  • InferenceCountZeroAugmentedPoissonAbstract$get_summary()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_bayesian_bootstrap_ci_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_bayesian_bootstrap_pval_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_bootstrap_ci_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_bootstrap_pval_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_information_preferences()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_rand_bootstrap_ci_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_rand_bootstrap_pval_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_testing_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_testing_type()
  • InferenceCountZeroAugmentedPoissonAbstract$select_optimal_b_subsampling()
  • InferenceCountZeroAugmentedPoissonAbstract$select_optimal_m_out_of_n_bootstrap()
  • InferenceCountZeroAugmentedPoissonAbstract$set_information_preference()
  • InferenceCountZeroAugmentedPoissonAbstract$set_testing_type()
  • InferenceCountZeroAugmentedPoissonAbstract$supports_rand_pval_for_incidence()
+ inherited public methods from Inference

InferenceCountZeroInflatedNegBin$new()

Initialize inference for the zero-inflated negative binomial model (binary excess-zero submodel mixed with a negative-binomial count submodel); see InferenceCountZeroInflatedNegBin for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferenceCountZeroInflatedNegBin$new(
  des_obj,
  model_formula = NULL,
  model_formula_zero = NULL,
  use_rcpp = TRUE,
  verbose = FALSE,
  optimization_alg = NULL
)
Arguments
des_obj

A completed Design object with a count response.

model_formula

Optional formula for covariate adjustment.

model_formula_zero

Formula for the zero-inflation submodel. If NULL (default), it uses the same formula as model_formula.

use_rcpp

Logical. If TRUE (default), use our internal Rcpp implementation. If FALSE, use glmmTMB.

verbose

Whether to print progress messages.

optimization_alg

Optimization algorithm. Default is dispatched via policy.


InferenceCountZeroInflatedNegBin$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCountZeroInflatedNegBin$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Lambert, D. (1992). "Zero-Inflated Poisson Regression, with an Application to Defects in Manufacturing." Technometrics, 34(1), 1-14, doi:10.2307/1269547, for the zero-inflated count-model framework.

See Also

InferenceCountNegBin for the single-part negative binomial model this class's count submodel generalizes; InferenceCountZeroInflatedPoisson for the Poisson (equidispersed) variant.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountZeroInflatedNegBin$new(seq_des, model_formula = ~ x1)
inf$compute_estimate()


Zero-Inflated Poisson Regression Inference for Count Responses

Description

Fits a zero-inflated Poisson regression for count responses: a binary excess-zero submodel P(\text{structural zero}_i) = \mathrm{logit}^{-1}(X_i^{h\top} \gamma^h) mixed with a (non-truncated) Poisson count submodel \log E[Y_i \mid \text{not structural zero}, w_i, x_i] = \beta_0 + \beta_T w_i + x_i^\top \gamma. Unlike a hurdle model, zero counts can arise from either the structural-zero mechanism or from an ordinary Poisson draw of 0, so the two mixture components are not identified by disjoint support. The hurdle and count submodels may use different covariate formulas (model_formula/model_formula_zero). The reported treatment effect is the coefficient from the conditional count component, on the log-rate scale, conditional on the response coming from the count process, not the excess-zero-inflation mechanism: it is not the effect on the unconditional mean E[Y], which also depends on how treatment shifts the excess-zero probability, under the default estimand = "conditional". likelihood_tier = "full": Wald, gradient, and (bootstrap-calibrated) likelihood-ratio tests are available for the count submodel's treatment coefficient under that estimand; a plain score test is not exposed. Jackknife inference is not supported: delete-one refits of this two-part mixture model are numerically unstable, so compute_jackknife_estimate() and related methods report explicit non-estimability rather than attempting delete-one refits.

Marginal (unconditional-mean) estimand. Via set_estimand(), this class also supports estimand = "marginal_mean_diff" and "marginal_ratio": the g-computation average, over the empirical covariate distribution, of the model-implied unconditional mean E[Y_i \mid w_i, x_i] = (1 - \pi(x_i)) \lambda(x_i) (the untruncated Poisson mean weighted by the non-structural-zero probability) at w_i = 1 vs. w_i = 0 — a mean difference or, on the log scale, a mean ratio. This is a pure post-fit transform of the same maximum- likelihood fit (no refit), with a delta-method standard error computed against the sandwich-robust covariance matrix already used for this class's conditional Wald inference. Only "wald"-type inference is available under a marginal estimand (no likelihood-ratio/score/gradient test, since the marginal quantity is a functional of the fitted parameters, not itself a likelihood).

Super classes

Inference -> InferenceCountZeroAugmentedPoissonAbstract -> InferenceCountZeroInflatedPoisson

Methods

Public methods

+ inherited public methods from InferenceCountZeroAugmentedPoissonAbstract
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_bayesian_bootstrap_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_bootstrap_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_jackknife_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_rand_bootstrap_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_randomization_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$approximate_subsampling_distribution_beta_hat_T()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_bayesian_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_bayesian_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_estimate_with_bootstrap_weights()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_gradient_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_gradient_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_bias_estimate()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_estimate()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_std_error()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_wald_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_wald_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_approx_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_approx_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_exact_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_exact_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_m_out_of_n_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_m_out_of_n_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_estimate()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_rand_bootstrap_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_rand_bootstrap_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_rand_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_rand_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_score_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_score_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_sensitivity()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_wald_confidence_interval()
  • InferenceCountZeroAugmentedPoissonAbstract$compute_wald_two_sided_pval()
  • InferenceCountZeroAugmentedPoissonAbstract$get_information_preference()
  • InferenceCountZeroAugmentedPoissonAbstract$get_information_source_used()
  • InferenceCountZeroAugmentedPoissonAbstract$get_last_param_bootstrap_diagnostics()
  • InferenceCountZeroAugmentedPoissonAbstract$get_last_param_bootstrap_estimate_diagnostics()
  • InferenceCountZeroAugmentedPoissonAbstract$get_mod()
  • InferenceCountZeroAugmentedPoissonAbstract$get_summary()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_bayesian_bootstrap_ci_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_bayesian_bootstrap_pval_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_bootstrap_ci_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_bootstrap_pval_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_information_preferences()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_rand_bootstrap_ci_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_rand_bootstrap_pval_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_supported_testing_types()
  • InferenceCountZeroAugmentedPoissonAbstract$get_testing_type()
  • InferenceCountZeroAugmentedPoissonAbstract$select_optimal_b_subsampling()
  • InferenceCountZeroAugmentedPoissonAbstract$select_optimal_m_out_of_n_bootstrap()
  • InferenceCountZeroAugmentedPoissonAbstract$set_information_preference()
  • InferenceCountZeroAugmentedPoissonAbstract$set_testing_type()
  • InferenceCountZeroAugmentedPoissonAbstract$supports_rand_pval_for_incidence()
+ inherited public methods from Inference

InferenceCountZeroInflatedPoisson$new()

Initialize inference for the zero-inflated Poisson model (binary excess-zero submodel mixed with a Poisson count submodel); see InferenceCountZeroInflatedPoisson for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferenceCountZeroInflatedPoisson$new(
  des_obj,
  model_formula = NULL,
  model_formula_zero = NULL,
  use_rcpp = TRUE,
  verbose = FALSE,
  optimization_alg = NULL
)
Arguments
des_obj

A completed Design object with a count response.

model_formula

Optional formula for the count submodel.

model_formula_zero

Formula for the zero-inflation submodel. If NULL (default), it uses the same formula as model_formula.

use_rcpp

Logical. If TRUE (default), use our internal Rcpp implementation. If FALSE, use glmmTMB.

verbose

Whether to print progress messages.

optimization_alg

Optimization algorithm. Default is dispatched via policy.


InferenceCountZeroInflatedPoisson$compute_estimate()

Fits the zero-inflated Poisson model. Under the default estimand = "conditional", returns \hat\beta_T, the treatment log-rate coefficient from the conditional count submodel (see the class-level caveat that this is conditional on the response coming from the count process, not an unconditional-mean effect). Under estimand = "marginal_mean_diff" or "marginal_ratio" (set via set_estimand()), returns the g-computation marginal mean difference or log-scale marginal ratio of the unconditional mean E[Y \mid w, x] = (1-\pi(x))\lambda(x) instead — a pure post-fit transform of the same cached fit, no refit.

Usage
InferenceCountZeroInflatedPoisson$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.


InferenceCountZeroInflatedPoisson$compute_asymp_confidence_interval()

Asymptotic confidence interval. Under the conditional estimand, delegates to the shared zero-augmented count-model Wald/ bootstrap-fallback contract; under a marginal estimand, the delta-method interval computed by compute_estimate(). Calls self$compute_estimate() first (not private$shared() directly) so the estimand-aware cache is always current regardless of call order.

Usage
InferenceCountZeroInflatedPoisson$compute_asymp_confidence_interval(
  alpha = 0.05
)
Arguments
alpha

The significance level (default 0.05).


InferenceCountZeroInflatedPoisson$compute_asymp_two_sided_pval()

Asymptotic two-sided p-value, dispatched exactly as compute_asymp_confidence_interval(); see that method's description for the marginal-estimand path.

Usage
InferenceCountZeroInflatedPoisson$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null treatment effect under the current estimand (default 0).


InferenceCountZeroInflatedPoisson$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCountZeroInflatedPoisson$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Lambert, D. (1992). "Zero-Inflated Poisson Regression, with an Application to Defects in Manufacturing." Technometrics, 34(1), 1-14, doi:10.2307/1269547, for the zero-inflated count-model framework.

See Also

InferenceCountPoisson for the single-part Poisson model this class's count submodel generalizes; InferenceCountHurdlePoisson for the related hurdle (disjoint-support) variant; InferenceCountZeroInflatedNegBin for the overdispersion-robust negative-binomial variant (does not support a marginal estimand — the mean-function derivation here is Poisson-specific).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountZeroInflatedPoisson$new(seq_des)
inf$compute_estimate()


Internal base for user-defined asymptotic inference extensions

Description

InferenceCustomAsymp is intentionally not exported. Extension packages may retrieve it with getFromNamespace("InferenceCustomAsymp", "EDI") while this API is experimental.

Subclasses implement a public fit(estimate_only = FALSE) method and return a named list with the custom-fit result contract:

estimate

Required numeric scalar treatment-effect estimate.

se

Optional numeric scalar standard error. Required for Wald confidence intervals and asymptotic p-values unless estimate_only is TRUE.

df

Optional numeric scalar degrees of freedom. Use NA_real_ for z inference.

model

Optional fitted model object retained for get_mod() and get_summary().

nonestimable_reason

Optional character scalar. When supplied with a non-finite estimate or standard error, EDI records the result as explicitly non-estimable.

Subclasses should use public accessors such as get_analysis_data(), get_response(), get_treatment(), and get_covariates() rather than EDI private fields.

Value

A two-sided p-value.

Super class

Inference -> InferenceCustomAsymp

Methods

Public methods

+ inherited public methods from Inference

InferenceCustomAsymp$compute_rand_two_sided_pval()

Computes a randomization two-sided p-value. Delegates to the 'RandomizationCI'-provided dispatch (Zhang incidence support, type/args_for_type) since 'NonparametricBootstrap' pulls in both 'RandomizationTest' and 'RandomizationCI', and the two provide conflicting 'compute_rand_two_sided_pval' implementations.

Usage
InferenceCustomAsymp$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  type = NULL,
  args_for_type = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors.

delta

Null treatment effect value.

transform_responses

Response transformation to apply during the test. For survival responses the default "log" multiplies the recorded times of the units treated under each reference allocation by e^\delta, event and censoring times alike, with censoring indicators unchanged – the rank-based AFT residual construction (Tsiatis 1990; Wei, Ying and Lin 1990; Jin, Lin, Wei and Ying 2003); see compute_rand_confidence_interval() for the assumptions.

na.rm

Whether to remove non-finite simulated statistics.

show_progress

Whether to show progress.

permutations

Optional pre-generated assignment draws.

type

Optional incidence-specific exact randomization type.

args_for_type

Optional arguments keyed by type.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceCustomAsymp$fit()

Calls the user-defined fit callback for this custom inference path; see InferenceCustomAsymp.

Usage
InferenceCustomAsymp$fit(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance calculations.

Returns

A list with fit results.


InferenceCustomAsymp$compute_estimate()

Compute the treatment-effect estimate by delegating to the user-supplied custom estimator. See InferenceCustomRand, InferenceCustomAsymp, and InferenceCustomBoot for related extension classes.

Usage
InferenceCustomAsymp$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance calculations.

Returns

The treatment estimate.


InferenceCustomAsymp$compute_asymp_confidence_interval()

Compute asymptotic confidence interval.

Usage
InferenceCustomAsymp$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level.

Returns

Confidence interval.


InferenceCustomAsymp$compute_asymp_two_sided_pval()

Compute asymptotic p-value.

Usage
InferenceCustomAsymp$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect.

Returns

P-value.


InferenceCustomAsymp$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCustomAsymp$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Internal base for user-defined bootstrap inference extensions

Description

This class uses the same fit() result contract as InferenceCustomAsymp, but only promises estimate/bootstrap behavior.

Value

A two-sided p-value.

Super class

Inference -> InferenceCustomBoot

Methods

Public methods

+ inherited public methods from Inference

InferenceCustomBoot$compute_rand_two_sided_pval()

Computes a randomization two-sided p-value. Delegates to the 'RandomizationCI'-provided dispatch (Zhang incidence support, type/args_for_type) since 'NonparametricBootstrap' pulls in both 'RandomizationTest' and 'RandomizationCI', and the two provide conflicting 'compute_rand_two_sided_pval' implementations.

Usage
InferenceCustomBoot$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  type = NULL,
  args_for_type = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors.

delta

Null treatment effect value.

transform_responses

Response transformation to apply during the test. For survival responses the default "log" multiplies the recorded times of the units treated under each reference allocation by e^\delta, event and censoring times alike, with censoring indicators unchanged – the rank-based AFT residual construction (Tsiatis 1990; Wei, Ying and Lin 1990; Jin, Lin, Wei and Ying 2003); see compute_rand_confidence_interval() for the assumptions.

na.rm

Whether to remove non-finite simulated statistics.

show_progress

Whether to show progress.

permutations

Optional pre-generated assignment draws.

type

Optional incidence-specific exact randomization type.

args_for_type

Optional arguments keyed by type.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceCustomBoot$fit()

Calls the user-defined fit callback for this custom inference path; see InferenceCustomAsymp.

Usage
InferenceCustomBoot$fit(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance calculations.

Returns

A list with fit results.


InferenceCustomBoot$compute_estimate()

Compute the treatment-effect estimate by delegating to the user-supplied custom asymptotic estimator. See InferenceCustomAsymp and InferenceAsymp.

Usage
InferenceCustomBoot$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance calculations.

Returns

The treatment estimate.


InferenceCustomBoot$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCustomBoot$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Internal base for user-defined randomization inference extensions

Description

This class uses the same fit() result contract as InferenceCustomAsymp, but only promises estimate/randomization/ randomization-CI behavior.

Value

A two-sided p-value.

Super class

Inference -> InferenceCustomRand

Methods

Public methods

+ inherited public methods from Inference

InferenceCustomRand$compute_rand_two_sided_pval()

Computes a randomization two-sided p-value. Delegates to the 'RandomizationCI'-provided dispatch (Zhang incidence support, type/args_for_type) since 'RandomizationCI' pulls in 'RandomizationTest', and the two provide conflicting 'compute_rand_two_sided_pval' implementations – the same conflict 'InferenceCustomAsymp' and 'InferenceCustomBoot' resolve the same way.

Usage
InferenceCustomRand$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  type = NULL,
  args_for_type = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors.

delta

Null treatment effect value.

transform_responses

Response transformation to apply during the test. For survival responses the default "log" multiplies the recorded times of the units treated under each reference allocation by e^\delta, event and censoring times alike, with censoring indicators unchanged – the rank-based AFT residual construction (Tsiatis 1990; Wei, Ying and Lin 1990; Jin, Lin, Wei and Ying 2003); see compute_rand_confidence_interval() for the assumptions.

na.rm

Whether to remove non-finite simulated statistics.

show_progress

Whether to show progress.

permutations

Optional pre-generated assignment draws.

type

Optional incidence-specific exact randomization type.

args_for_type

Optional arguments keyed by type.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceCustomRand$fit()

Calls the user-defined fit callback for this custom inference path; see InferenceCustomAsymp.

Usage
InferenceCustomRand$fit(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance calculations.

Returns

A list with fit results.


InferenceCustomRand$compute_estimate()

Compute the treatment-effect estimate by delegating to the user-supplied custom bootstrap estimator. See InferenceCustomBoot and InferenceNonParamBootstrap.

Usage
InferenceCustomRand$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance calculations.

Returns

The treatment estimate.


InferenceCustomRand$clone()

The objects of this class are cloneable with this method.

Usage
InferenceCustomRand$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Binomial Identity Risk Difference Inference for Incidence Responses

Description

Fits a binomial regression with the identity link for binary (incidence) responses: P(Y_i = 1) = \beta_0 + \beta_T W_i + X_i^\top \gamma, where W_i is the treatment indicator and X_i are optional recorded covariates, by maximum likelihood (fast_identity_binomial_regression_cpp/ fast_identity_binomial_regression_weighted_cpp). Because the link is the identity rather than the logit, \hat\beta_T is directly a risk difference on the probability scale, not a log-odds-ratio — the class name and estimand differ from InferenceIncidLogRegr for exactly this reason. likelihood_tier = "full": likelihood-ratio, score, gradient, and Wald tests are all available when the model converges, plus parametric-likelihood-bootstrap calibration of the likelihood-ratio test. Because the identity link does not constrain fitted probabilities to [0,1], fits are hardened by QR column-dropping and rejected as nonestimable when the fitted linear predictor produces implausible coefficients (see private$is_identity_binomial_fit_reasonable()); this is a real practical limitation of the identity link relative to logit/probit, not a bug. Validity requires the additive risk-difference model to be correctly specified over the covariate range actually observed (an identity-link fit can be well-behaved in-sample yet imply out-of-range probabilities for other covariate values).

Estimand. Composes MarginalEstimand (set_estimand()/get_estimand()/get_supported_estimands()). estimand = "marginal_mean_diff" is supported for API consistency with the other GLM families, but it is algebraically identical to the default estimand = "conditional" for this class: the g-computation marginal risk difference is \frac{1}{n}\sum_i \{(\hat\beta_0 + \hat\beta_T + X_i^\top \hat\gamma) - (\hat\beta_0 + X_i^\top \hat\gamma)\}, which simplifies to exactly \hat\beta_T for every subject (not merely on average) because the identity link is linear in W_i with no treatment-by-covariate interaction term — the per-subject treated-minus-control difference \hat\beta_T does not depend on X_i at all, so standardizing over the covariate distribution changes nothing. Contrast with InferenceCountPoisson's "marginal_ratio" (also a collapsing case, for the same no-interaction reason) and "marginal_mean_diff" (which does not collapse, since a difference does not distribute through the nonlinear log-link mean). This collapsing is a genuine property of the identity-link model, not a wiring bug — it is documented here so a user comparing estimands for this class is not surprised the two never differ. Standard errors are computed independently for each estimand (the marginal path uses the delta method against the model's coefficient covariance; the conditional path uses the model information matrix), so while the two point estimates coincide exactly, their standard errors may differ slightly by construction even though both are asymptotically valid.

Super class

Inference -> InferenceIncidBinomialIdentityRiskDiff

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidBinomialIdentityRiskDiff$compute_estimate()

Fits the identity-link binomial regression model by maximum likelihood. Under the default estimand = "conditional", returns \hat\beta_T, the risk difference coefficient. Under estimand = "marginal_mean_diff" (set via set_estimand()), returns the g-computation marginal risk difference — see the class-level @details for why this is algebraically identical to the conditional estimate for this family. The underlying model fit is identical either way (a pure post-fit transform of the same cached fit, no refit).

Usage
InferenceIncidBinomialIdentityRiskDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.


InferenceIncidBinomialIdentityRiskDiff$compute_asymp_confidence_interval()

Wald confidence interval, dispatched by testing_type for the conditional estimand; under a marginal estimand testing_type is always "wald" (the only value set_estimand() permits there). Calls self$compute_estimate() first (not private$shared() directly) so the estimand-aware cache is always current regardless of call order.

Usage
InferenceIncidBinomialIdentityRiskDiff$compute_asymp_confidence_interval(
  alpha = 0.05
)
Arguments
alpha

Two-sided miscoverage rate; the returned interval targets 1 - alpha coverage.


InferenceIncidBinomialIdentityRiskDiff$compute_asymp_two_sided_pval()

Wald two-sided p-value, dispatched by testing_type exactly as compute_asymp_confidence_interval(); see that method's description for the marginal-estimand always-Wald note.

Usage
InferenceIncidBinomialIdentityRiskDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment-effect value under the current estimand (both scales coincide for this family — see the class-level @details).


InferenceIncidBinomialIdentityRiskDiff$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidBinomialIdentityRiskDiff$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

McCullagh, P., and Nelder, J. A. (1989). Generalized Linear Models (2nd ed.). Chapman and Hall/CRC, for the binomial GLM family and identity-link risk-difference parameterization.

See Also

InferenceIncidLogRegr (logit link, log-odds-ratio estimand), InferenceIncidLogBinomial (log link, log-risk-ratio estimand) for alternative link/estimand choices on the same response type. Comparable Python API: statsmodels GLM (family=Binomial(link=identity())). See also: Generalized linear model (Wikipedia).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidBinomialIdentityRiskDiff$new(seq_des)
inf$compute_estimate()


CMH Blocked Incidence Inference

Description

Unadjusted blocked-design incidence inference using the simple mean-difference point estimate with a randomization-based standard error.

Legacy inference class. This class is retained for backwards compatibility and is not comprehensively tested by the package comprehensive-test harness.

Internally, this class recodes treatment assignments to w_i \in \{-1, +1\} (the package-wide convention is \{0,1\}; see Design). For a balanced design the treatment-effect estimator is \hat\tau = (2/n)\,\mathbf{y}'\mathbf{w}, and since E_w[\mathbf{y}'\mathbf{w}] = 0 for any balanced randomization the standard error is

SE(\hat\tau) = \frac{2}{n}\sqrt{\frac{\sum_k (\mathbf{y}'\mathbf{w}_k)^2}{K}}

where K draws \mathbf{w}_1,\ldots,\mathbf{w}_K come from the design's reference distribution. Centering at the known zero mean (rather than the sample mean) makes the denominator K rather than K-1.

For blocking designs the expectation is evaluated exactly:

SE(\hat\tau) = \frac{2}{n}\sqrt{\sum_b \frac{n_{1b}\,n_{0b}}{n_B - 1}}

where n_{1b}, n_{0b} are the numbers of positive and negative responses in block b and n_B is the (common) block size. This equals 2\sqrt{V_{\rm CMH}} where V_{\rm CMH} is the CMH variance from Azriel et al. (2026), Equation 3.

For non-blocking designs, the "balanced design" precondition above requires the observed treatment allocation to be exactly balanced (n_T = n_C), not merely drawn from a prob\_T = 0.5 mechanism – e.g. plain Bernoulli randomization has prob\_T = 0.5 but does not guarantee an exactly balanced realized allocation. A warning (not an error) is issued once, the first time the standard error is actually computed (i.e. on the first confidence-interval / p-value / standard-error request, not at construction or for estimate-only use), when this is violated – erroring would make this class unusable with Bernoulli-style non-blocking designs entirely; the warning tells the caller the reported standard error may be miscalibrated.

Super class

Inference -> InferenceIncidCMH

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidCMH$compute_asymp_confidence_interval()

Uses the randomization-CI layer's two-sided p-value contract (InferenceRandCI's version, not InferenceRand's): for incidence responses this dispatches to the Zhang exact randomization test where applicable rather than refusing outright, matching this class's pre-migration old-ladder behavior (it inherited from InferenceAllSimpleAverageDiff, whose own pin was already corrected to InferenceRandCI – see that file's identical rationale). This class independently composes the same components rather than truly inheriting InferenceAllSimpleAverageDiff, so it had its own stale copy of the old InferenceRand pin, which silently regressed Zhang dispatch for the non-blocking balanced-design path – found via test-incid-cmh-extended-robins-migration-golden.R's randomization_pval case going from '"ok"' to '"unsupported"'.

Wald confidence interval for the balanced-design/CMH risk-difference estimate \hat\tau, using the randomization-based (blocking-design: exact CMH variance formula; non-blocking design: Monte Carlo over se_est_num_vectors design draws) standard error documented in the class @details. See InferenceAsymp for the shared Wald contract.

Usage
InferenceIncidCMH$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.

Returns

A length-2 numeric vector c(lower, upper) on the risk-difference scale.


InferenceIncidCMH$compute_asymp_two_sided_pval()

Two-sided Wald p-value for H_0: \tau = \code{delta} vs. H_1: \tau \neq \code{delta}, using the same randomization-based standard error as compute_asymp_confidence_interval().

Usage
InferenceIncidCMH$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null value of \tau to test against; 0 (the default) tests for any treatment effect at all.

Returns

Numeric scalar p-value in [0, 1].


InferenceIncidCMH$new()

Initialize Cochran-Mantel-Haenszel incidence inference, validate the stratified binary-response design, and prepare the stratum-adjusted test used by InferenceIncidCMH.

Usage
InferenceIncidCMH$new(
  des_obj,
  model_formula = NULL,
  se_est_num_vectors = 5000L,
  verbose = FALSE
)
Arguments
des_obj

A completed design object.

model_formula

Optional formula for covariate adjustment.

se_est_num_vectors

For non-block designs, the number of randomization vectors drawn from the design to estimate the standard error. Default 1000L.

verbose

Logical. Whether to print progress messages.

Returns

A new InferenceIncidCMH object.


InferenceIncidCMH$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidCMH$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

seq_des = DesignSeqOneByOneRandomBlockSize$new(n = 20, response_type = 'incidence',
  strata_cols = 'x1')
for (i in 1:20) {
  seq_des$add_one_subject_to_experiment_and_assign(
    data.frame(x1 = factor(rep(1:2, 10)[i], levels=1:2)))
}
seq_des$add_all_subject_responses(rbinom(20, 1, 0.5))
inf = InferenceIncidCMH$new(seq_des)
inf$compute_estimate()

Exact Binomial (McNemar-Type) Incidence Inference for Matched-Pair Designs

Description

Performs exact matched-pair inference for binary (incidence) outcomes using only discordant matched pairs — pairs where the treated and control member's outcomes differ — the same reduction classical McNemar's test makes. Writing d_+ for the count of discordant pairs where the treated subject had the event and the control did not, and d_- for the reverse, the point estimate is the Haldane-Anscombe continuity-corrected log odds ratio \log\left((d_+ + 0.5)/(d_- + 0.5)\right); the confidence interval inverts the exact (Clopper-Pearson) binomial confidence interval for d_+ / (d_+ + d_-) against 1/2 (via stats::binom.test) onto the log-odds scale; and the two-sided p-value is an exact binomial test of d_+ vs. d_- (via zhang_exact_binom_pval_cpp) against a null log odds ratio. This class is available for DesignFixedBinaryMatch and KK matching-on-the-fly designs. For KK designs, only the matched-pair data are used and the reservoir is ignored. If there are no matched pairs, or no discordant pairs, the relevant quantities are reported as non-estimable rather than as NaN/Inf.

Initialize exact matched-pair binomial inference for incidence outcomes. Requires des_obj to be DesignFixedBinaryMatch or a KK matching-on-the-fly-capable design; errors otherwise. Requires an uncensored incidence response.

Computes the Haldane-Anscombe continuity-corrected matched-pair log odds ratio \log\left((d_+ + 0.5)/(d_- + 0.5)\right) from the discordant matched-pair counts (see class documentation for the full model). NA if there are no matched pairs.

Value

A new InferenceIncidExactBinomial object.

The treatment estimate.

Super class

Inference -> InferenceIncidExactBinomial

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidExactBinomial$new()

Usage
InferenceIncidExactBinomial$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed design object.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values by default.


InferenceIncidExactBinomial$compute_estimate()

Usage
InferenceIncidExactBinomial$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

Ignored for this estimator (the exact statistic is always cheap to compute; there is no separate variance step to skip).


InferenceIncidExactBinomial$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidExactBinomial$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidExactBinomial$new(seq_des)
inf$compute_estimate()


Exact Fisher (Conditional Hypergeometric) Incidence Inference

Description

Performs exact conditional inference for binary (incidence) outcomes via Fisher's exact test on one or more 2x2 (treated/control by case/noncase) tables. When the design provides no stratification structure (e.g. an unstructured or iBCRD design), a single overall 2x2 table is built and fisher.test is used directly, giving the conditional MLE odds ratio and its exact confidence interval/p-value. When the design has blocking structure (DesignFixedBlocking, DesignSeqOneByOneSPBR, DesignSeqOneByOneRandomBlockSize), a separate 2x2 table is built per block-defining covariate stratum. When the design has matched-pair structure (KK matching-on-the-fly designs), each matched pair becomes its own 2x2 table, with any reservoir (unmatched) subjects pooled into one additional stratum table. In either stratified case, mantelhaen.test (exact conditional test) is used instead, giving the common odds ratio across strata; stratified inference only supports testing/estimating against a null odds ratio of 1 (log odds ratio 0) — a non-zero null shift is rejected with an error. Strata with no cases or no noncases in either arm are dropped before analysis; if no informative strata remain, this errors rather than returning a degenerate result.

Initialize exact Fisher inference for incidence outcomes. Requires an uncensored incidence response; the design's structure (unstructured, blocked, or matched) determines the stratification used at estimation time (see class documentation).

Computes the log of the (conditional MLE, or common-odds-ratio if stratified) odds ratio from fisher.test or mantelhaen.test (see class documentation for which applies and why).

Value

A new InferenceIncidExactFisher object.

The treatment estimate.

Super class

Inference -> InferenceIncidExactFisher

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidExactFisher$new()

Usage
InferenceIncidExactFisher$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed design object.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values by default.


InferenceIncidExactFisher$compute_estimate()

Usage
InferenceIncidExactFisher$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

Ignored for this estimator (the exact statistic is always cheap to compute; there is no separate variance step to skip).


InferenceIncidExactFisher$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidExactFisher$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Fisher, R. A. (1935). "The Logic of Inductive Inference." Journal of the Royal Statistical Society, 98(1), 39-82, doi:10.2307/2342435, for the exact conditional test underlying fisher.test; Mantel, N., and Haenszel, W. (1959). "Statistical Aspects of the Analysis of Data from Retrospective Studies of Disease." Journal of the National Cancer Institute, 22(4), 719-748, for the stratified common-odds-ratio test used when the design provides multiple strata.

Examples

des = DesignFixediBCRD$new(n = 20, response_type = 'incidence')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(20)))
des$assign_w_to_all_subjects()
des$add_all_subject_responses(rbinom(20, 1, 0.5))
inf = InferenceIncidExactFisher$new(des)
inf$compute_estimate()
inf$compute_exact_two_sided_pval_for_treatment_effect()

Exact Zhang Combined-Test Incidence Inference

Description

Performs exact inference for a binary (incidence) outcome that combines two exact component tests when the design has both matched-pair and reservoir (unmatched) subjects — an internal-to-this-package method (not drawn from external literature) analogous in spirit to InferenceIncidExactBinomial (matched pairs) and InferenceIncidExactFisher (unmatched 2x2 table), fused into one combined exact test rather than a Wald-style variance combination. The point estimate is always the Haldane-Anscombe continuity-corrected log odds ratio \log\left((n_{11} + 0.5)(n_{00} + 0.5) / \left((n_{10}+0.5)(n_{01}+0.5)\right)\right) from the pooled 2\times2 table across all subjects (matched and reservoir together). For p-values and confidence intervals, the two subsets are tested separately (an exact matched-pairs binomial test, as in InferenceIncidExactBinomial, on discordant pairs; an exact Fisher test on the reservoir 2\times2 table, as in InferenceIncidExactFisher), and their p-values are combined via combination_method: "Fisher" (default; -2(\log p_M + \log p_R) \sim \chi^2_4 under independence), "Stouffer" (averaged z-scores), or "min_p" (Šidák-style 1-(1-\min(p_M,p_R))^2). If only one of the two subsets is informative (e.g. a pure-Bernoulli design with no matching, or no discordant pairs), the combined p-value degenerates to that one component's p-value. Confidence intervals are obtained by numerically inverting (bisection) the combined p-value as a function of the hypothesized log odds ratio, starting from a normal-approximation (Haldane-Anscombe MLE) interval as the search bracket. Requires a Bernoulli-capable or matching-capable design.

Initialize exact Zhang combined-test incidence inference. Requires des_obj to be Bernoulli-capable or matching-capable, an uncensored incidence response.

Computes the Haldane-Anscombe continuity-corrected log odds ratio \log\left((n_{11}+0.5)(n_{00}+0.5) / \left((n_{10}+0.5)(n_{01}+0.5)\right)\right) from the pooled 2\times2 table across all subjects (matched and reservoir combined) — see class documentation for the full combined-test model.

Computes an exact confidence interval for the log odds ratio by bisection-inverting the combined matched-pairs + Fisher-exact p-value (see class documentation for the full combination methodology).

Value

A new InferenceIncidExactZhang object.

The treatment estimate.

A confidence interval.

Super class

Inference -> InferenceIncidExactZhang

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidExactZhang$new()

Usage
InferenceIncidExactZhang$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed design object.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values by default.


InferenceIncidExactZhang$compute_estimate()

Usage
InferenceIncidExactZhang$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

Ignored for this estimator (the exact statistic is always cheap to compute; there is no separate variance step to skip).


InferenceIncidExactZhang$compute_exact_confidence_interval()

Usage
InferenceIncidExactZhang$compute_exact_confidence_interval(
  alpha = 0.05,
  pval_epsilon = 0.005,
  type = NULL,
  args_for_type = NULL
)
Arguments
alpha

Significance level.

pval_epsilon

Bisection tolerance for the inversion routine.

type

Exact inference type; only "Zhang" (the default) is supported.

args_for_type

Optional arguments keyed by exact type; recognizes combination_method ("Fisher" (default), "Stouffer", or "min_p") inside the "Zhang" entry.


InferenceIncidExactZhang$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidExactZhang$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

seq_des = DesignSeqOneByOneKK14$new(n = 20, response_type = 'incidence')
for (i in 1:20) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(20, 1, 0.5))
inf = InferenceIncidExactZhang$new(seq_des)
inf$compute_estimate()
inf$compute_exact_two_sided_pval_for_treatment_effect()

Extended Robins Blocked Incidence Inference

Description

Unadjusted blocked-design incidence inference using the simple mean-difference point estimate with a block-stratified standard error.

Legacy inference class. This class is retained for backwards compatibility and is not comprehensively tested by the package comprehensive-test harness.

Super class

Inference -> InferenceIncidExtendedRobins

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidExtendedRobins$compute_asymp_confidence_interval()

Uses the randomization-CI layer's two-sided p-value contract (InferenceRandCI's version, not InferenceRand's): for incidence responses this dispatches to the Zhang exact randomization test where applicable rather than refusing outright, matching this class's pre-migration old-ladder behavior (it inherited from InferenceAllSimpleAverageDiff, whose own pin was already corrected to InferenceRandCI – see that file's identical rationale). This class independently composes the same components rather than truly inheriting InferenceAllSimpleAverageDiff, so it had its own stale copy of the old InferenceRand pin (same bug as InferenceIncidWald/InferenceIncidCMH, fixed alongside them even though this class's own golden test's design doesn't happen to trigger the Zhang-eligible path that would have caught it).

Uses the shared asymptotic confidence-interval contract; see InferenceAsymp.

Usage
InferenceIncidExtendedRobins$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Numeric. Significance level (default 0.05).


InferenceIncidExtendedRobins$compute_asymp_two_sided_pval()

Uses the shared asymptotic two-sided p-value contract; see InferenceAsymp.

Usage
InferenceIncidExtendedRobins$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Numeric. Null treatment effect value (default 0).


InferenceIncidExtendedRobins$new()

Initialize Extended Robins blocked-design incidence inference.

Usage
InferenceIncidExtendedRobins$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE
)
Arguments
des_obj

A completed design object.

model_formula

Optional formula for covariate adjustment.

verbose

Logical. Whether to print progress messages.

Returns

A new InferenceIncidExtendedRobins object.


InferenceIncidExtendedRobins$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidExtendedRobins$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

seq_des = DesignSeqOneByOneRandomBlockSize$new(n = 20, response_type = 'incidence',
  strata_cols = 'x1')
for (i in 1:20) {
  seq_des$add_one_subject_to_experiment_and_assign(
    data.frame(x1 = factor(rep(1:2, 10)[i], levels=1:2)))
}
seq_des$add_all_subject_responses(rbinom(20, 1, 0.5))
inf = InferenceIncidExtendedRobins$new(seq_des)
inf$compute_estimate()

G-Computation Risk-Difference Inference for Binary Responses

Description

Fits a logistic working model, \mathrm{logit}\,\Pr(Y_i=1\mid x_i) = x_i^\top\hat\beta, for an incidence outcome using treatment and, optionally, all recorded covariates, then estimates the marginal (standardized) risk difference \mathrm{RD} = \overline{\mathrm{risk}}_1 - \overline{\mathrm{risk}}_0 by G-computation: setting every subject's treatment indicator to 1 (respectively 0) while holding their other observed covariates fixed, averaging the model-implied risk over the empirical covariate distribution under each counterfactual, and differencing — see gcomp_logistic_point_estimate_cpp for the exact standardization formula. Inference is nonparametric-bootstrap/randomization/jackknife-based (likelihood_tier = "none"): no closed-form asymptotic standard error is used.

Uses the shared nonparametric bootstrap distribution contract; see InferenceNonParamBootstrap.

Computes a bootstrap confidence interval for the treatment effect.

Computes a bootstrap two-sided p-value for the treatment effect.

Computes a Bayesian-bootstrap two-sided p-value for the treatment effect.

Computes a Bayesian-bootstrap confidence interval for the treatment effect.

Computes a jackknife-Wald two-sided p-value for the treatment effect.

Computes a jackknife-Wald confidence interval for the treatment effect.

Computes a PRW subsampling two-sided p-value for the treatment effect.

Computes a PRW subsampling confidence interval for the treatment effect.

Computes an m-out-of-n bootstrap two-sided p-value for the treatment effect.

Computes an m-out-of-n bootstrap confidence interval for the treatment effect.

Value

A numeric vector of bootstrap estimates.

Super class

Inference -> InferenceIncidGCompRiskDiff

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidGCompRiskDiff$approximate_bootstrap_distribution_beta_hat_T()

Usage
InferenceIncidGCompRiskDiff$approximate_bootstrap_distribution_beta_hat_T(
  B = 501,
  show_progress = TRUE,
  debug = FALSE,
  bootstrap_type = NULL
)
Arguments
B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

debug

Whether to return diagnostics.

bootstrap_type

Optional resampling scheme.

bootstrap_type

Optional empirical-resampling scheme.

bootstrap_type

Optional empirical-resampling scheme.


InferenceIncidGCompRiskDiff$compute_bootstrap_confidence_interval()

Usage
InferenceIncidGCompRiskDiff$compute_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 501,
  type = NULL,
  na.rm = TRUE,
  show_progress = TRUE,
  min_number_usable_samples = 5L
)
Arguments
alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.


InferenceIncidGCompRiskDiff$compute_bootstrap_two_sided_pval()

Usage
InferenceIncidGCompRiskDiff$compute_bootstrap_two_sided_pval(
  delta = NULL,
  B = 501,
  type = "symmetric",
  na.rm = FALSE,
  show_progress = TRUE,
  min_number_usable_samples = 5L
)
Arguments
delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.


InferenceIncidGCompRiskDiff$compute_bayesian_bootstrap_two_sided_pval()

Usage
InferenceIncidGCompRiskDiff$compute_bayesian_bootstrap_two_sided_pval(
  delta = NULL,
  B = 501,
  type = NULL,
  na.rm = FALSE,
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  weighting_unit_type = NULL
)
Arguments
delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

weighting_unit_type

Optional resampling unit override.

weighting_unit_type

Optional resampling unit override.


InferenceIncidGCompRiskDiff$compute_bayesian_bootstrap_confidence_interval()

Usage
InferenceIncidGCompRiskDiff$compute_bayesian_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 501,
  type = NULL,
  na.rm = TRUE,
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  weighting_unit_type = NULL
)
Arguments
alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

weighting_unit_type

Optional resampling unit override.

weighting_unit_type

Optional resampling unit override.


InferenceIncidGCompRiskDiff$compute_jackknife_wald_two_sided_pval()

Usage
InferenceIncidGCompRiskDiff$compute_jackknife_wald_two_sided_pval(
  delta = NULL,
  unit = "auto"
)
Arguments
delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceIncidGCompRiskDiff$compute_jackknife_wald_confidence_interval()

Usage
InferenceIncidGCompRiskDiff$compute_jackknife_wald_confidence_interval(
  alpha = 0.05,
  unit = "auto"
)
Arguments
alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceIncidGCompRiskDiff$compute_subsampling_two_sided_pval()

Usage
InferenceIncidGCompRiskDiff$compute_subsampling_two_sided_pval(
  delta = NULL,
  B = 501,
  b = NULL,
  type = "centered",
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  subsampling_type = NULL
)
Arguments
delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

b

Subsample size. See InferenceNonParamBootstrap$compute_subsampling_two_sided_pval.

b

Subsample size. See InferenceNonParamBootstrap$compute_subsampling_confidence_interval.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

subsampling_type

Optional empirical-resampling scheme.

subsampling_type

Optional empirical-resampling scheme.


InferenceIncidGCompRiskDiff$compute_subsampling_confidence_interval()

Usage
InferenceIncidGCompRiskDiff$compute_subsampling_confidence_interval(
  alpha = 0.05,
  B = 501,
  b = NULL,
  type = "basic",
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  subsampling_type = NULL
)
Arguments
alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

b

Subsample size. See InferenceNonParamBootstrap$compute_subsampling_two_sided_pval.

b

Subsample size. See InferenceNonParamBootstrap$compute_subsampling_confidence_interval.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

subsampling_type

Optional empirical-resampling scheme.

subsampling_type

Optional empirical-resampling scheme.


InferenceIncidGCompRiskDiff$compute_m_out_of_n_bootstrap_two_sided_pval()

Usage
InferenceIncidGCompRiskDiff$compute_m_out_of_n_bootstrap_two_sided_pval(
  delta = NULL,
  B = 501,
  m = NULL,
  type = "centered",
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  bootstrap_type = NULL
)
Arguments
delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

m

Resample size. See InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval.

m

Resample size. See InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

bootstrap_type

Optional resampling scheme.

bootstrap_type

Optional empirical-resampling scheme.

bootstrap_type

Optional empirical-resampling scheme.


InferenceIncidGCompRiskDiff$compute_m_out_of_n_bootstrap_confidence_interval()

Usage
InferenceIncidGCompRiskDiff$compute_m_out_of_n_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 501,
  m = NULL,
  type = "basic",
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  bootstrap_type = NULL
)
Arguments
alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

m

Resample size. See InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval.

m

Resample size. See InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

bootstrap_type

Optional resampling scheme.

bootstrap_type

Optional empirical-resampling scheme.

bootstrap_type

Optional empirical-resampling scheme.


InferenceIncidGCompRiskDiff$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidGCompRiskDiff$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

See Also

InferenceIncidGCompRiskRatio for the risk-ratio version of this same standardized logistic working model.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidGCompRiskDiff$new(seq_des)
inf$compute_estimate()


G-Computation Risk-Ratio Inference for Binary Responses

Description

Fits a logistic working model, \mathrm{logit}\,\Pr(Y_i=1\mid x_i) = x_i^\top\hat\beta, for an incidence outcome using treatment and, optionally, all recorded covariates, then estimates the marginal (standardized) risk ratio \mathrm{RR} = \overline{\mathrm{risk}}_1 / \overline{\mathrm{risk}}_0 by G-computation: setting every subject's treatment indicator to 1 (respectively 0) while holding their other observed covariates fixed, averaging the model-implied risk over the empirical covariate distribution under each counterfactual, and taking the ratio — see gcomp_logistic_point_estimate_cpp for the exact standardization formula (mean1/mean0). Bootstrap/jackknife inference on this estimand is generally done on the log risk-ratio scale internally (see $compute_bootstrap_confidence_interval(), $compute_bayesian_bootstrap_confidence_interval(), and the jackknife-Wald methods, whose "basic"/"wald" interval types route through log-scale-specific helpers for this estimand), then back-transformed, since ratio estimators are typically closer to normally distributed on the log scale. Inference is nonparametric-bootstrap/ randomization/jackknife-based (likelihood_tier = "none"): no closed-form asymptotic standard error is used.

Uses the shared nonparametric bootstrap distribution contract; see InferenceNonParamBootstrap.

Computes a bootstrap confidence interval for the treatment effect.

Computes a bootstrap two-sided p-value for the treatment effect.

Computes a Bayesian-bootstrap two-sided p-value for the treatment effect.

Computes a Bayesian-bootstrap confidence interval for the treatment effect.

Computes a jackknife-Wald two-sided p-value for the treatment effect.

Computes a jackknife-Wald confidence interval for the treatment effect.

Computes a PRW subsampling two-sided p-value for the treatment effect.

Computes a PRW subsampling confidence interval for the treatment effect.

Computes an m-out-of-n bootstrap two-sided p-value for the treatment effect.

Computes an m-out-of-n bootstrap confidence interval for the treatment effect.

Value

A numeric vector of bootstrap estimates.

Super class

Inference -> InferenceIncidGCompRiskRatio

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidGCompRiskRatio$approximate_bootstrap_distribution_beta_hat_T()

Usage
InferenceIncidGCompRiskRatio$approximate_bootstrap_distribution_beta_hat_T(
  B = 501,
  show_progress = TRUE,
  debug = FALSE,
  bootstrap_type = NULL
)
Arguments
B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

debug

Whether to return diagnostics.

bootstrap_type

Optional resampling scheme.

bootstrap_type

Optional empirical-resampling scheme.

bootstrap_type

Optional empirical-resampling scheme.


InferenceIncidGCompRiskRatio$compute_bootstrap_confidence_interval()

Usage
InferenceIncidGCompRiskRatio$compute_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 501,
  type = NULL,
  na.rm = TRUE,
  show_progress = TRUE,
  min_number_usable_samples = 5L
)
Arguments
alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.


InferenceIncidGCompRiskRatio$compute_bootstrap_two_sided_pval()

Usage
InferenceIncidGCompRiskRatio$compute_bootstrap_two_sided_pval(
  delta = NULL,
  B = 501,
  type = "symmetric",
  na.rm = FALSE,
  show_progress = TRUE,
  min_number_usable_samples = 5L
)
Arguments
delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.


InferenceIncidGCompRiskRatio$compute_bayesian_bootstrap_two_sided_pval()

Usage
InferenceIncidGCompRiskRatio$compute_bayesian_bootstrap_two_sided_pval(
  delta = NULL,
  B = 501,
  type = NULL,
  na.rm = FALSE,
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  weighting_unit_type = NULL
)
Arguments
delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

weighting_unit_type

Optional resampling unit override.

weighting_unit_type

Optional resampling unit override.


InferenceIncidGCompRiskRatio$compute_bayesian_bootstrap_confidence_interval()

Usage
InferenceIncidGCompRiskRatio$compute_bayesian_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 501,
  type = NULL,
  na.rm = TRUE,
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  weighting_unit_type = NULL
)
Arguments
alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

na.rm

Whether to remove non-finite bootstrap replicates.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

weighting_unit_type

Optional resampling unit override.

weighting_unit_type

Optional resampling unit override.


InferenceIncidGCompRiskRatio$compute_jackknife_wald_two_sided_pval()

Usage
InferenceIncidGCompRiskRatio$compute_jackknife_wald_two_sided_pval(
  delta = NULL,
  unit = "auto"
)
Arguments
delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceIncidGCompRiskRatio$compute_jackknife_wald_confidence_interval()

Usage
InferenceIncidGCompRiskRatio$compute_jackknife_wald_confidence_interval(
  alpha = 0.05,
  unit = "auto"
)
Arguments
alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

unit

Deletion unit. Default "auto".

unit

Deletion unit. Default "auto".


InferenceIncidGCompRiskRatio$compute_subsampling_two_sided_pval()

Usage
InferenceIncidGCompRiskRatio$compute_subsampling_two_sided_pval(
  delta = NULL,
  B = 501,
  b = NULL,
  type = "centered",
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  subsampling_type = NULL
)
Arguments
delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

b

Subsample size. See InferenceNonParamBootstrap$compute_subsampling_two_sided_pval.

b

Subsample size. See InferenceNonParamBootstrap$compute_subsampling_confidence_interval.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

subsampling_type

Optional empirical-resampling scheme.

subsampling_type

Optional empirical-resampling scheme.


InferenceIncidGCompRiskRatio$compute_subsampling_confidence_interval()

Usage
InferenceIncidGCompRiskRatio$compute_subsampling_confidence_interval(
  alpha = 0.05,
  B = 501,
  b = NULL,
  type = "basic",
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  subsampling_type = NULL
)
Arguments
alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

b

Subsample size. See InferenceNonParamBootstrap$compute_subsampling_two_sided_pval.

b

Subsample size. See InferenceNonParamBootstrap$compute_subsampling_confidence_interval.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

subsampling_type

Optional empirical-resampling scheme.

subsampling_type

Optional empirical-resampling scheme.


InferenceIncidGCompRiskRatio$compute_m_out_of_n_bootstrap_two_sided_pval()

Usage
InferenceIncidGCompRiskRatio$compute_m_out_of_n_bootstrap_two_sided_pval(
  delta = NULL,
  B = 501,
  m = NULL,
  type = "centered",
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  bootstrap_type = NULL
)
Arguments
delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

The null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

delta

Null treatment effect. Defaults to 0 for RD and 1 for RR.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

m

Resample size. See InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval.

m

Resample size. See InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

bootstrap_type

Optional resampling scheme.

bootstrap_type

Optional empirical-resampling scheme.

bootstrap_type

Optional empirical-resampling scheme.


InferenceIncidGCompRiskRatio$compute_m_out_of_n_bootstrap_confidence_interval()

Usage
InferenceIncidGCompRiskRatio$compute_m_out_of_n_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 501,
  m = NULL,
  type = "basic",
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  bootstrap_type = NULL
)
Arguments
alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

alpha

Significance level. Default 0.05.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of Bayesian-bootstrap samples.

B

Number of subsamples.

B

Number of subsamples.

B

Number of resamples.

B

Number of resamples.

m

Resample size. See InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval.

m

Resample size. See InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval.

type

Bootstrap CI type. See InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

type

Bayesian-bootstrap p-value type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.

type

Bayesian-bootstrap CI type. See InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.

type

P-value type.

type

Confidence-interval type.

type

P-value type.

type

Confidence-interval type.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite bootstrap samples required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite subsampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

min_number_usable_samples

Minimum number of finite resampled estimates required.

bootstrap_type

Optional resampling scheme.

bootstrap_type

Optional empirical-resampling scheme.

bootstrap_type

Optional empirical-resampling scheme.


InferenceIncidGCompRiskRatio$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidGCompRiskRatio$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

See Also

InferenceIncidGCompRiskDiff for the risk-difference version of this same standardized logistic working model.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidGCompRiskRatio$new(seq_des)
inf$compute_estimate()


Conditional Logistic Plus GLMM IVWC Inference for KK Designs

Description

Fits a combined conditional-logit-plus-random-intercept-GLMM likelihood for incidence responses under a KK matching-on-the-fly design, where reservoir (unmatched) subjects are excluded from the GLMM component (private$combine_reservoir_into_glmm() == FALSE): only concordant matched pairs contribute their random-intercept GLMM likelihood alongside the discordant-pair conditional-logit term, both sharing a single treatment coefficient \beta_T. See InferencePropKKGLMM for the full model form (conditional-logit-on-discordant plus random-intercept-GLMM, jointly maximized) and InferenceAbstractKKCondLogitGLMM for the shared fitting/caching contract. Contrast with the sibling InferenceIncidKKCondLogitGLMMOneLik, which instead includes reservoir subjects in the GLMM component (combine_reservoir_into_glmm() == TRUE) — this class's naming ("IVWC") reflects that reservoir information, when used, is intended to be combined with this fit's estimate via inverse-variance weighting rather than folded into the same likelihood.

Details

Legacy class. Not fully tested in comprehensive_tests.R.

Super classes

Inference -> InferenceAbstractKKCondLogitGLMM -> InferenceIncidKKCondLogitGLMMIVWC

Methods

Public methods

+ inherited public methods from InferenceAbstractKKCondLogitGLMM
+ inherited public methods from Inference

InferenceIncidKKCondLogitGLMMIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidKKCondLogitGLMMIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKCondLogitGLMMIVWC$new(seq_des)
inf$compute_estimate()

Conditional Logistic Plus GLMM Combined-Likelihood Inference for KK Designs

Description

Fits a combined conditional-logit-plus-random-intercept-GLMM likelihood for incidence responses under a KK matching-on-the-fly design, where reservoir (unmatched) subjects are included in the GLMM component (private$combine_reservoir_into_glmm() == TRUE), so all subjects (discordant matched pairs, concordant matched pairs, and reservoir) enter one joint likelihood with a single treatment coefficient \beta_T. See InferencePropKKGLMM for the full model form (conditional-logit-on-discordant plus random-intercept-GLMM, jointly maximized) and InferenceAbstractKKCondLogitGLMM for the shared fitting/caching contract. Contrast with the sibling InferenceIncidKKCondLogitGLMMIVWC, which excludes reservoir subjects from the GLMM component.

Super classes

Inference -> InferenceAbstractKKCondLogitGLMM -> InferenceIncidKKCondLogitGLMMOneLik

Methods

Public methods

+ inherited public methods from InferenceAbstractKKCondLogitGLMM
+ inherited public methods from Inference

InferenceIncidKKCondLogitGLMMOneLik$new()

Initialize inference for the combined conditional-logit (discordant matched pairs) plus random-intercept-GLMM (concordant pairs and reservoir subjects) incidence model; see InferenceIncidKKCondLogitGLMMOneLik for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferenceIncidKKCondLogitGLMMOneLik$new(
  des_obj,
  model_formula = NULL,
  max_abs_reasonable_coef = 50,
  max_abs_reasonable_se = 1.25,
  max_abs_log_sigma = 8,
  verbose = FALSE,
  smart_cold_start_default = NULL,
  optimization_alg = NULL
)
Arguments
des_obj

A completed Design object with an incidence response.

model_formula

Optional formula for covariate adjustment.

max_abs_reasonable_coef

Cap for reasonable coefficient estimates.

max_abs_reasonable_se

Cap for reasonable treatment standard errors.

max_abs_log_sigma

Cap for reasonable log random effect variance.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart optimizer start values.

optimization_alg

Character. Optimization algorithm (default "lbfgs").


InferenceIncidKKCondLogitGLMMOneLik$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidKKCondLogitGLMMOneLik$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKCondLogitGLMMOneLik$new(seq_des)
inf$compute_estimate()


Conditional Logistic IVWC Inference (KK Designs, Binary Response)

Description

Inverse-variance-weighted combination (IVWC) of two independently fit conditional-likelihood pieces for KK matched-pair-plus-reservoir binary designs: matched pairs are analyzed with exact conditional logistic regression (conditional_logit_fit_matched_pairs(), which conditions out the pair-specific nuisance intercept and estimates only the treatment log-odds-ratio \beta_T from discordant pairs, or the joint clogit-style likelihood when covariates are present), and reservoir subjects are analyzed with ordinary logistic regression (conditional_logit_fit_reservoir()). If \hat\beta_m, \hat\sigma^2_m and \hat\beta_r, \hat\sigma^2_r are the matched-pair and reservoir estimates and their variances, the combined estimate is the variance-weighted average

\hat\beta_T = w^\star \hat\beta_m + (1-w^\star) \hat\beta_r, \quad w^\star = \frac{\hat\sigma^2_r}{\hat\sigma^2_r + \hat\sigma^2_m},

with combined variance \hat\sigma^2_m \hat\sigma^2_r / (\hat\sigma^2_m + \hat\sigma^2_r). This is the classical fixed-effects inverse-variance meta-analysis pooling formula (see Cochrane Handbook / DerSimonian-Laird), applied here to combine the two conditionally-independent likelihood contributions of a KK design rather than to pool separate studies. When only one of the two components is estimable the combined estimate falls back to that component alone. Contrast this with InferenceIncidKKCondLogitOneLik, which instead fits a single joint likelihood over both pieces (see that class's documentation) – likelihood_tier = "partial" here reflects that the matched-pair piece is a genuine conditional (partial) likelihood, but the two-piece combination itself is a closed-form Wald/meta-analytic step, not a further likelihood evaluation.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Details

Legacy class. Not fully tested in comprehensive_tests.R.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Super class

Inference -> InferenceIncidKKCondLogitIVWC

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidKKCondLogitIVWC$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceIncidKKCondLogitIVWC$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

delta

The null difference. Default 0.

transform_responses

Type of transformation. Default "none".

show_progress

Show progress bar. Default TRUE.

permutations

Pre-computed permutations. Default NULL.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceIncidKKCondLogitIVWC$supports_rand_pval_for_incidence()

Usage
InferenceIncidKKCondLogitIVWC$supports_rand_pval_for_incidence()

InferenceIncidKKCondLogitIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidKKCondLogitIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Fleiss, J.L., Levin, B., Paik, M.C. (2003). Statistical Methods for Rates and Proportions, 3rd ed. Wiley. (conditional logistic regression for matched pairs)

See Also

InferenceIncidKKCondLogitOneLik for the one-likelihood alternative combining strategy.


One-Likelihood Conditional-Logistic Inference for KK Binary Designs

Description

Estimates a treatment log-odds-ratio \beta_T for binary (incidence) outcomes collected under a KK matching-on-the-fly design (DesignSeqOneByOneKK14 or subclass) by maximizing one combined likelihood that couples a conditional-logistic (intercept-free, within-matched-pair) likelihood for matched subjects with an ordinary logistic likelihood for reservoir subjects, sharing a single treatment coefficient across both pieces. This is the "one-likelihood" counterpart to InferenceIncidKKCondLogitIVWC, which instead fits the matched and reservoir pieces separately and pools them by inverse-variance weighting; here the treatment coefficient is a single joint MLE, and likelihood_tier = "full" exposes likelihood-ratio, score, and gradient inference plus a parametric likelihood bootstrap in addition to Wald.

Estimand. \beta_T, the treatment coefficient of a logistic mean model \mathrm{logit}(P(Y=1 \mid w,x)) = \beta_0 + \beta_T w + x\beta; \exp(\hat\beta_T) is the treatment-vs-control odds ratio.

Model. Matched pairs contribute McFadden-style conditional logistic likelihood terms that condition away the pair-specific nuisance intercept (see build_matching_combined_clogit_design_cpp/ collect_discordant_pairs_cpp); reservoir subjects contribute an ordinary logistic likelihood with one shared intercept. The combined negative log-likelihood is minimized jointly in (\beta_0, \beta_T, \beta) via fast_logistic_regression_cpp/ fast_logistic_regression_with_var_cpp. When get_testing_type() != "wald", asymptotic CI/p-value calls are routed through InferenceAsympLik's generic score/likelihood-ratio/gradient dispatch instead of the design's own Wald machinery.

Assumptions. Independence across matched pairs and reservoir subjects given covariates; correct logistic mean specification; a KK matching-on-the-fly design supplying the matched/reservoir partition. No response censoring is supported (checked at construction via assertNoCensoring()).

Super class

Inference -> InferenceIncidKKCondLogitOneLik

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidKKCondLogitOneLik$compute_rand_two_sided_pval()

Computes a randomization-based two-sided p-value for the treatment effect, preflighting the observed combined-likelihood treatment statistic (see the class-header note above) before delegating to InferenceRandCI's Zhang-dispatch-aware implementation.

Usage
InferenceIncidKKCondLogitOneLik$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  type = NULL,
  args_for_type = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization (permutation) draws.

delta

The null treatment effect. Default 0.

transform_responses

Optional response transform applied before the randomization statistic is computed. Default "none".

na.rm

Whether to remove non-finite permutation replicates.

show_progress

Whether to show a progress bar.

permutations

Optional pre-generated permutation matrix/list to reuse instead of drawing new permutations.

type

Optional randomization-statistic type override.

args_for_type

Optional list of extra arguments for type.

zero_one_logit_clamp

Clamp applied to responses at the 0/1 boundary before a logit-scale transform, to avoid infinite values. Default .Machine$double.eps.


InferenceIncidKKCondLogitOneLik$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidKKCondLogitOneLik$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kapelner, A., and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388. doi:10.1111/biom.12148. (KK14 in REFERENCES.md.)

See Also

Analogous Python API for conditional logistic regression: statsmodels discrete models (ConditionalLogit). Logistic regression (orientation).


G-Computation Risk-Difference Inference for KK Designs with Binary Responses

Description

Fits an all-subject logistic working model \mathrm{logit}\,P(Y=1\mid w,x) = \beta_0 + \beta_T w + \beta_X^\top x for a KK incidence outcome using treatment w and, optionally, all recorded covariates x, then estimates the marginal (standardized, g-computation) risk difference \hat\theta = n^{-1}\sum_i \{\hat p(1, x_i) - \hat p(0, x_i)\} by averaging the fitted-model predicted risks under all-treated and all-control assignments over the empirical covariate distribution (Robins 1986). Matched pairs are treated as clusters and reservoir subjects are treated as singletons when computing the sandwich covariance of the standardized estimator (the delta-method variance of the empirical mean of the two counterfactual-risk contrasts, not the naive logistic-regression coefficient variance).

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Details

This estimator has likelihood_tier = "none": the fitted logistic model is a working model for standardization only, and reported inference is sandwich/bootstrap-based, not likelihood-based. compute_estimate() fits the model and returns \hat\theta on the risk-difference (probability) scale; compute_estimate_with_bootstrap_weights() refits under Bayesian-bootstrap subject weights for compute_bayesian_bootstrap_confidence_interval(). Jackknife deletes one cluster (matched pair or singleton reservoir subject) at a time. If the working model fails to converge or the design has no treatment-arm variation, the estimate is marked non-estimable via is_nonestimable().

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Super class

Inference -> InferenceIncidKKGCompRiskDiff

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidKKGCompRiskDiff$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceIncidKKGCompRiskDiff$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

delta

The null difference. Default 0.

transform_responses

Type of transformation. Default "none".

show_progress

Show progress bar. Default TRUE.

permutations

Pre-computed permutations. Default NULL.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceIncidKKGCompRiskDiff$supports_rand_pval_for_incidence()

Usage
InferenceIncidKKGCompRiskDiff$supports_rand_pval_for_incidence()

InferenceIncidKKGCompRiskDiff$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidKKGCompRiskDiff$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Robins, J. (1986). A new approach to causal inference in mortality studies with a sustained exposure period. Mathematical Modelling, 7(9-12), 1393-1512. doi:10.1016/0270-0255(86)90088-6

See Also

InferenceIncidKKGCompRiskRatio for the risk-ratio analog on the same standardization machinery.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKGCompRiskDiff$new(seq_des)
inf$compute_estimate()


G-Computation Risk-Ratio Inference for KK Designs with Binary Responses

Description

Fits the same all-subject logistic working model as InferenceIncidKKGCompRiskDiff for a KK incidence outcome using treatment and, optionally, all recorded covariates, then estimates the marginal (standardized, g-computation) risk ratio \hat\theta = \left(n^{-1}\sum_i \hat p(1, X_i)\right) / \left(n^{-1}\sum_i \hat p(0, X_i)\right) by averaging fitted-model predicted risks under all-treated and all-control assignments over the empirical covariate distribution (Robins 1986). Matched pairs are treated as clusters and reservoir subjects are treated as singletons when computing the sandwich covariance; the delta method is applied on the log-risk-ratio scale to keep the reported ratio and its confidence interval positive, then back-transformed for reporting.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Details

This estimator has likelihood_tier = "none". If the working model fails to converge, has no treatment-arm variation, or the all-control standardized risk is zero (undefined ratio), the estimate is marked non-estimable via is_nonestimable().

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Super class

Inference -> InferenceIncidKKGCompRiskRatio

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidKKGCompRiskRatio$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceIncidKKGCompRiskRatio$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

delta

The null difference. Default 0.

transform_responses

Type of transformation. Default "none".

show_progress

Show progress bar. Default TRUE.

permutations

Pre-computed permutations. Default NULL.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceIncidKKGCompRiskRatio$supports_rand_pval_for_incidence()

Usage
InferenceIncidKKGCompRiskRatio$supports_rand_pval_for_incidence()

InferenceIncidKKGCompRiskRatio$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidKKGCompRiskRatio$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Robins, J. (1986). A new approach to causal inference in mortality studies with a sustained exposure period. Mathematical Modelling, 7(9-12), 1393-1512. doi:10.1016/0270-0255(86)90088-6

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKGCompRiskRatio$new(seq_des)
inf$compute_estimate()


GEE Inference for KK Designs with Binary Response

Description

Fits a Generalized Estimating Equations (GEE) model with a binomial family and logit link, \mathrm{logit}\,\Pr(Y_i = 1 \mid x_i) = x_i^\top\beta, for binary (incidence) responses under a KK matching-on-the-fly design, using an exchangeable working correlation structure where each cluster is either a matched pair (2 members) or a reservoir singleton (1 member) — see $compute_estimate()'s method-level documentation for the full fitting contract (internal Rcpp solver vs. geepack fallback, hardening/retry behavior). GEE is used here purely to fit one marginal model jointly across matched-pair and reservoir subjects while accounting for the within-pair correlation the matching induces, not as a longitudinal/repeated-measures tool. Inference is quasi-likelihood/estimating-equation based (likelihood_tier = "quasi"): standard errors are GEE sandwich (robust) standard errors, not model-likelihood-based.

Super class

Inference -> InferenceIncidKKGEE

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidKKGEE$new()

Initialize KK binary-response GEE inference, validate the matched/reservoir design, and prepare the exchangeable-working-correlation binomial (logit-link) GEE fitting machinery used by InferenceIncidKKGEE.

Usage
InferenceIncidKKGEE$new(
  des_obj,
  model_formula = NULL,
  use_rcpp = TRUE,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with an incidence response.

model_formula

Optional formula for covariate adjustment.

use_rcpp

Whether to use the internal Rcpp GEE solver (TRUE, default) with automatic fallback to geepack::geeglm on failure, or always use geepack::geeglm directly (FALSE).

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.


InferenceIncidKKGEE$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidKKGEE$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Liang, K.-Y., and Zeger, S. L. (1986). "Longitudinal Data Analysis Using Generalized Linear Models." Biometrika, 73(1), 13-22, doi:10.1093/biomet/73.1.13, for the GEE estimating-equation framework and sandwich variance estimator used here.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKGEE$new(seq_des)
inf$compute_estimate()


Modified-Poisson Inference for KK Designs with Binary Responses

Description

Fits Zou's (2004) modified-Poisson working model for binary incidence outcomes under a KK matching-on-the-fly design: a log-link Poisson model \log E[Y_i \mid w_i, x_i] = \beta_0 + \beta_T w_i + x_i^\top \gamma is fit to the binary (0/1) response by ordinary Poisson maximum likelihood (a working, misspecified likelihood — the true response is Bernoulli, not Poisson), and the coefficient standard errors are corrected by a cluster-robust sandwich covariance rather than the (invalid, for a misspecified likelihood) model-based Poisson information. Matched pairs are treated as clusters (2 members) and reservoir subjects as singleton clusters when computing the sandwich covariance, so the matched-pair correlation induced by the design is accounted for even though the modified-Poisson working model itself does not encode it directly. \exp(\hat\beta_T) is the estimated risk ratio, directly interpretable unlike a logistic regression's odds ratio (which only approximates the risk ratio when the outcome is rare). likelihood_tier = "none" (the sandwich-corrected inference is not a normalized model likelihood): only Wald inference is exposed. See InferenceAbstractKKMarginalIncid for the shared marginal-incidence fitting contract.

Super classes

Inference -> InferenceAbstractKKMarginalIncid -> InferenceAbstractKKModifiedPoisson -> InferenceIncidKKModifiedPoisson

Methods

Public methods

+ inherited public methods from InferenceAbstractKKModifiedPoisson
+ inherited public methods from InferenceAbstractKKMarginalIncid
+ inherited public methods from Inference

InferenceIncidKKModifiedPoisson$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidKKModifiedPoisson$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Zou, G. (2004). "A Modified Poisson Regression Approach to Prospective Studies with Binary Data." American Journal of Epidemiology, 159(7), 702-706, doi:10.1093/aje/kwh090; Kapelner, A. and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the KK matching-on-the-fly design this class is built for.

See Also

InferenceIncidModifiedPoisson for the non-KK analog.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKModifiedPoisson$new(seq_des)
inf$compute_estimate()


KK Newcombe Risk-Difference IVWC Inference for Binary Responses

Description

Initialize KK Newcombe risk-difference IVWC inference and prepare matched/reservoir paired-binomial components used by InferenceIncidKKNewcombeRiskDiff.

Computes the compound Newcombe risk-difference point estimate \hat\theta = w_1 \hat\theta_1 + w_2 \hat\theta_2 on the risk-difference (probability) scale, where \hat\theta_1 is the paired-Newcombe discordant-pair estimate from matched pairs and \hat\theta_2 is the independent-Newcombe estimate from reservoir subjects, combined by inverse-variance weighting w_j \propto 1/\widehat{\mathrm{Var}}(\hat\theta_j) (falls back to the single available component when one has zero subjects). Caches intermediate match/reservoir statistics for reuse by compute_asymp_confidence_interval() and compute_asymp_two_sided_pval().

Usage

KKNewcombeRiskDiffIVWCSource

Details

Implements a compound Newcombe risk-difference estimator for KK designs. This class pools information from matched pairs (using the Paired Newcombe method) and the reservoir (using the Independent Newcombe method) via inverse-variance weighted combination (IVWC).

The matched-pair component applies the paired Newcombe (Method-10-style Wilson-score) interval to the discordant pairs to estimate the treatment effect and its variance (Newcombe 1998). The reservoir component applies the independent-samples Newcombe interval, treating unmatched subjects as two independent binomial samples. The two component estimates \hat\theta_1, \hat\theta_2 are combined by inverse-variance weighting, \hat\theta = w_1 \hat\theta_1 + w_2 \hat\theta_2, w_j = (1/\hat V_j) / \sum_k (1/\hat V_k), the standard IVWC framework used throughout the package's KK inference classes. likelihood_tier = "none": this is a closed-form Wilson-score-type estimator, not a fitted likelihood model, so no likelihood-ratio or parametric-bootstrap methods are exposed. If a design has no matched pairs or no reservoir subjects, the single available component is used directly rather than combined.

References

Newcombe, R. G. (1998). Interval Estimation for the Difference Between Independent Proportions: Comparison of Eleven Methods. Statistics in Medicine, 17(8), 873-890. doi:10.1002/(SICI)1097-0258(19980430)17:8<873::AID-SIM779>3.0.CO;2-I

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKNewcombeRiskDiff$new(seq_des)
inf$compute_estimate()


Log-Binomial Regression Inference for Incidence Responses

Description

Fits a binomial regression with the log link for binary (incidence) responses: \log P(Y_i = 1) = \beta_0 + \beta_T W_i + X_i^\top \gamma, where W_i is the treatment indicator and X_i are optional recorded covariates, by maximum likelihood (fast_log_binomial_regression_cpp/ fast_log_binomial_regression_weighted_cpp). \hat\beta_T is a log risk ratio: \exp(\hat\beta_T) is the estimated treatment risk ratio (relative risk) directly, unlike the log-odds-ratio from InferenceIncidLogRegr's logit link. likelihood_tier = "full": Wald, score, gradient, and likelihood-ratio tests are all available when the model converges, plus parametric-likelihood-bootstrap calibration of the likelihood-ratio test. Because the log link does not constrain fitted probabilities to [0,1] (only to [0,\infty)), fits are hardened by QR column-dropping and a coefficient-magnitude cap (max_abs_reasonable_coef) and rejected as nonestimable when the fit is implausible — the same practical limitation as the identity-link sibling InferenceIncidBinomialIdentityRiskDiff, here applying to the upper rather than both tails of the probability scale. Validity requires the multiplicative log-linear risk model to be correctly specified over the covariate range observed.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Super class

Inference -> InferenceIncidLogBinomial

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidLogBinomial$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceIncidLogBinomial$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

delta

The null difference. Default 0.

transform_responses

Type of transformation. Default "none".

show_progress

Show progress bar. Default TRUE.

permutations

Pre-computed permutations. Default NULL.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceIncidLogBinomial$supports_rand_pval_for_incidence()

Usage
InferenceIncidLogBinomial$supports_rand_pval_for_incidence()

InferenceIncidLogBinomial$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidLogBinomial$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

McCullagh, P., and Nelder, J. A. (1989). Generalized Linear Models (2nd ed.). Chapman and Hall/CRC, for the binomial GLM family and log-link relative-risk parameterization.

See Also

InferenceIncidLogRegr (logit link, log-odds-ratio estimand), InferenceIncidBinomialIdentityRiskDiff (identity link, risk-difference estimand) for alternative link/estimand choices on the same response type. Comparable Python API: statsmodels GLM (family=Binomial(link=log())). See also: Generalized linear model (Wikipedia).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidLogBinomial$new(seq_des)
inf$compute_estimate()


Logistic Regression Inference for Incidence Responses

Description

Fits a logistic regression model for binary (incidence) responses: \mathrm{logit}(P(Y_i = 1)) = \beta_0 + \beta_T W_i + X_i^\top \gamma, where W_i is the treatment indicator and X_i are optional recorded covariates, by maximum likelihood (fast_logistic_regression_cpp/ fast_logistic_regression_weighted_cpp). \hat\beta_T is a log-odds-ratio: \exp(\hat\beta_T) is the estimated treatment odds ratio. likelihood_tier = "full": Wald, score, gradient, and likelihood-ratio tests are all available (via the shared StandardModelCache model-caching contract), plus parametric-likelihood-bootstrap calibration of the likelihood-ratio test. A fit whose coefficients exceed max_abs_reasonable_coef in magnitude (a proxy for near-perfect separation) is cached as nonestimable rather than returned. Validity requires the usual logistic-regression assumptions: correctly specified linear predictor on the logit scale, independence across subjects conditional on covariates, and no perfect/quasi-complete separation.

Estimand. Composes MarginalEstimand (set_estimand()/get_estimand()/get_supported_estimands()). Under the default estimand = "conditional", \hat\beta_T is the log-odds-ratio above. Under estimand = "marginal_mean_diff", the reported quantity is instead the g-computation marginal risk difference \frac{1}{n}\sum_i \{\mathrm{plogis}(\hat\beta_0 + \hat\beta_T + X_i^\top \hat\gamma) - \mathrm{plogis}(\hat\beta_0 + X_i^\top \hat\gamma)\} — every subject's covariates plugged in once under treatment and once under control, averaged over the empirical covariate distribution. Under estimand = "marginal_ratio", the log of the corresponding marginal risk ratio. Because there is no latent submodel for this family (unlike e.g. InferencePropZeroOneInflatedBetaRegr's zero/one-inflation mixture), the marginal mean function is exactly the model's own fitted mean; no separate standardization step beyond the g-computation average is needed. Standard errors under a marginal estimand use the delta method against the model's coefficient covariance (degrees of freedom Inf); testing_type is restricted to "wald" whenever the estimand is non-conditional (set_testing_type() errors otherwise). The underlying model fit is identical regardless of estimand — switching estimand is a pure post-fit transform, never a refit.

Super class

Inference -> InferenceIncidLogRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidLogRegr$compute_estimate()

Fits the logistic regression model by maximum likelihood. Under the default estimand = "conditional", returns \hat\beta_T, the treatment log-odds-ratio. Under estimand = "marginal_mean_diff"/"marginal_ratio" (set via set_estimand()), returns the g-computation marginal risk difference/log-risk-ratio instead — see the class-level @details for the formula. The underlying model fit is identical either way (a pure post-fit transform of the same cached fit, no refit).

Usage
InferenceIncidLogRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.


InferenceIncidLogRegr$compute_asymp_confidence_interval()

Wald confidence interval, dispatched by testing_type for the conditional estimand (score/gradient/ likelihood-ratio/Bartlett available); under a marginal estimand testing_type is always "wald" (the only value set_estimand() permits there), so this always resolves to the delta-method interval. Calls self$compute_estimate() first (not private$shared() directly) so the estimand-aware cache is always current regardless of call order.

Usage
InferenceIncidLogRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Two-sided miscoverage rate; the returned interval targets 1 - alpha coverage.


InferenceIncidLogRegr$compute_asymp_two_sided_pval()

Wald two-sided p-value, dispatched by testing_type exactly as compute_asymp_confidence_interval(); see that method's description for the marginal-estimand always-Wald note.

Usage
InferenceIncidLogRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment-effect value under the current estimand (conditional log-odds-ratio, or marginal risk difference/ log-risk-ratio).


InferenceIncidLogRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidLogRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

McCullagh, P., and Nelder, J. A. (1989). Generalized Linear Models (2nd ed.). Chapman and Hall/CRC, for the logistic regression model and its maximum-likelihood theory.

See Also

Comparable Python API: statsmodels GLM. See also: Logistic regression (Wikipedia).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidLogRegr$new(seq_des)
inf$compute_estimate()


inf$set_seed(1)
inf$compute_lik_ratio_bootstrap_two_sided_pval(delta = 0, B = 9, show_progress = FALSE)


Miettinen-Nurminen Risk-Difference Inference for Binary Responses

Description

Fits the classical Miettinen-Nurminen score method for the risk difference in a two-arm binary trial. The point estimate is the observed risk difference, while confidence intervals and p-values are obtained by inverting the constrained score test under the null p_T - p_C = \delta.

This class is intentionally unadjusted. It operates on the 2 \times 2 table induced by treatment assignment and incidence response (counts x_T, x_C of events among n_T, n_C treated/control subjects), and is therefore the natural classical binary-endpoint complement to the regression-based incidence methods already in the package. The point estimate is the plain risk difference \hat p_T - \hat p_C; the Miettinen-Nurminen confidence interval and p-value invert a restricted-maximum-likelihood score test of H_0: p_T - p_C = \delta (via mn_ci_cpp/mn_pvalue_cpp) — under the null, the two arms' event probabilities are re-estimated subject to the constraint \hat p_T - \hat p_C = \delta, and the resulting score statistic is compared to its asymptotic normal distribution, with a small-sample bias correction factor (n_T+n_C)/(n_T+n_C-1) applied to the naive Wald variance used elsewhere (e.g. in $compute_estimate_with_bootstrap_weights(), which never gets this correction since it skips the score-test path entirely). This score-based interval generally has better small-sample coverage than the naive normal-approximation Wald interval on the risk difference.

Super class

Inference -> InferenceIncidMiettinenNurminenRiskDiff

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidMiettinenNurminenRiskDiff$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize a Miettinen-Nurminen risk-difference inference object for a completed design with an incidence response.

Usage
InferenceIncidMiettinenNurminenRiskDiff$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE
)
Arguments
des_obj

A completed DesignSeqOneByOne object with an incidence response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 20, response_type = "incidence")
for (i in 1:20) {
	x_i = data.frame(x1 = rnorm(1), x2 = rnorm(1))
	w_i = seq_des$add_one_subject_to_experiment_and_assign(x_i)
	p_i = plogis(-0.8 + 0.5 * w_i)
	seq_des$add_one_subject_response(i, rbinom(1, 1, p_i))
}
seq_des_inf = InferenceIncidMiettinenNurminenRiskDiff$new(seq_des)
seq_des_inf$compute_estimate()

InferenceIncidMiettinenNurminenRiskDiff$compute_estimate()

Computes the observed (unadjusted) risk-difference estimate \hat p_T - \hat p_C (see class documentation for the full Miettinen-Nurminen inference model). NA if either arm is empty.

Usage
InferenceIncidMiettinenNurminenRiskDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.


InferenceIncidMiettinenNurminenRiskDiff$compute_estimate_with_bootstrap_weights()

Recomputes the risk-difference estimate under subject/block bootstrap weights: the weighted event proportions \hat p_T^w = \sum_i r_i y_i \mathbb{1}[w_i=1] / \sum_i r_i \mathbb{1}[w_i=1] (and analogously for control), differenced. Used by the Bayesian bootstrap and related weighted-resampling machinery; see InferenceBayesianBootstrap. Always leaves the standard error and degrees of freedom unavailable (NA) regardless of estimate_only — this weighted path never computes the Miettinen-Nurminen score-based variance.

Usage
InferenceIncidMiettinenNurminenRiskDiff$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

Present for interface parity; this method never computes variance components regardless of its value.


InferenceIncidMiettinenNurminenRiskDiff$compute_asymp_confidence_interval()

Computes a 1-\alpha Miettinen-Nurminen restricted-MLE score confidence interval for the risk difference (see class documentation for the full method), by bisection-inverting the score test (mn_ci_cpp) to pval_epsilon tolerance.

Usage
InferenceIncidMiettinenNurminenRiskDiff$compute_asymp_confidence_interval(
  alpha = 0.05,
  pval_epsilon = 1e-07
)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.

pval_epsilon

Bisection tolerance for CI bounds.


InferenceIncidMiettinenNurminenRiskDiff$compute_asymp_two_sided_pval()

Computes a two-sided Miettinen-Nurminen restricted-MLE score p-value (mn_pvalue_cpp) testing H_0: p_T - p_C = \code{delta} (see class documentation for the full method).

Usage
InferenceIncidMiettinenNurminenRiskDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null treatment effect on the risk-difference scale.


InferenceIncidMiettinenNurminenRiskDiff$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidMiettinenNurminenRiskDiff$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Miettinen, O., and Nurminen, M. (1985). "Comparative Analysis of Two Rates." Statistics in Medicine, 4(2), 213-226, doi:10.1002/sim.4780040211, for the restricted-maximum-likelihood score method used here.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidMiettinenNurminenRiskDiff$new(seq_des)
inf$compute_estimate()


## ------------------------------------------------
## Method `InferenceIncidMiettinenNurminenRiskDiff$new()`
## ------------------------------------------------

seq_des = DesignSeqOneByOneBernoulli$new(n = 20, response_type = "incidence")
for (i in 1:20) {
	x_i = data.frame(x1 = rnorm(1), x2 = rnorm(1))
	w_i = seq_des$add_one_subject_to_experiment_and_assign(x_i)
	p_i = plogis(-0.8 + 0.5 * w_i)
	seq_des$add_one_subject_response(i, rbinom(1, 1, p_i))
}
seq_des_inf = InferenceIncidMiettinenNurminenRiskDiff$new(seq_des)
seq_des_inf$compute_estimate()

Modified Poisson Regression Inference for Incidence Responses

Description

Fits Zou's (2004) modified Poisson regression for binary (incidence) responses: \log E[Y_i \mid w_i, x_i] = \beta_0 + \beta_T w_i + x_i^\top \gamma, fit by maximizing the ordinary Poisson log-likelihood treating the binary Y_i as if it were Poisson-distributed (a valid estimating equation for the conditional mean regardless of the true outcome distribution, exactly as InferencePropFractionalLogit's quasi-binomial fit is for fractional responses). \hat\beta_T is a log risk ratio: \exp(\hat\beta_T) is the estimated treatment relative risk, the same estimand as InferenceIncidLogBinomial's log-binomial model, but modified Poisson never produces a fit failure from the [0,1]-probability constraint that a genuine binomial log-link model can hit. Caveat: this implementation's standard error comes from the ordinary (model-based) Poisson Fisher information (fast_poisson_regression_with_var_cpp's ssq_b_j), not a robust/sandwich correction — Zou's (2004) original proposal specifically pairs the misspecified Poisson working model with a robust sandwich variance estimator to obtain valid standard errors under the resulting overdispersion; users needing the fully robust modified-Poisson variance should treat this class's standard errors/CIs/p-values as approximate. likelihood_tier = "full" metadata is set for component-composition purposes, but private$supports_likelihood_tests() is hard FALSE — only Wald inference is exposed (get_supported_testing_types_impl() returns "wald" only), not likelihood-ratio/score/gradient tests. Fits with implausible coefficients or fitted linear predictors (checked via private$is_modified_poisson_fit_reasonable(), capped by max_abs_reasonable_coef/max_abs_reasonable_linear_predictor) are cached as nonestimable rather than returned.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Super class

Inference -> InferenceIncidModifiedPoisson

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidModifiedPoisson$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceIncidModifiedPoisson$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

delta

The null difference. Default 0.

transform_responses

Type of transformation. Default "none".

show_progress

Show progress bar. Default TRUE.

permutations

Pre-computed permutations. Default NULL.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceIncidModifiedPoisson$supports_rand_pval_for_incidence()

Usage
InferenceIncidModifiedPoisson$supports_rand_pval_for_incidence()

InferenceIncidModifiedPoisson$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidModifiedPoisson$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Zou, G. (2004). "A Modified Poisson Regression Approach to Prospective Studies with Binary Data." American Journal of Epidemiology, 159(7), 702-706, doi:10.1093/aje/kwh090.

See Also

InferenceIncidLogBinomial for the genuine log-binomial alternative with the same log-risk-ratio estimand. Comparable Python API: statsmodels discrete models (Poisson family on binary data). See also: Poisson regression (Wikipedia).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidModifiedPoisson$new(seq_des)
inf$compute_estimate()


Newcombe Risk-Difference Inference for Binary Responses

Description

Fits the Newcombe hybrid score method (Method 10) for the risk difference in a two-arm binary trial. This method constructs a confidence interval for the difference between two independent proportions by combining Wilson score intervals for each group.

This class is unadjusted and assumes independent samples (e.g. from a Bernoulli design). It ignores any matched-pair structure if present; for matched data, use InferenceIncidKKNewcombeRiskDiff. The point estimate is the plain risk difference \hat p_T - \hat p_C. The confidence interval (newcombe_independent_ci_cpp) is Newcombe's "Method 10" hybrid score interval: separate Wilson score intervals [\ell_T, u_T] and [\ell_C, u_C] are computed for each arm's proportion individually, then combined into a difference interval via [\hat p_T - \hat p_C - \sqrt{(\hat p_T - \ell_T)^2 + (u_C - \hat p_C)^2},\ \hat p_T - \hat p_C + \sqrt{(u_T - \hat p_T)^2 + (\hat p_C - \ell_C)^2}] — this avoids the boundary/coverage problems of the naive Wald interval on a risk difference while remaining closed-form (no iterative score-test inversion, unlike the Miettinen-Nurminen method in InferenceIncidMiettinenNurminenRiskDiff). The two-sided p-value has no closed form here: it is obtained by numerically inverting the confidence interval (bisection via stats::uniroot) to find the significance level at which delta falls exactly on the interval boundary.

Super class

Inference -> InferenceIncidNewcombeRiskDiff

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidNewcombeRiskDiff$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize a Newcombe risk-difference inference object for a completed design with an uncensored incidence response.

Usage
InferenceIncidNewcombeRiskDiff$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE
)
Arguments
des_obj

A completed DesignSeqOneByOne object with an incidence response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.


InferenceIncidNewcombeRiskDiff$compute_estimate()

Computes the observed (unadjusted) risk-difference estimate \hat p_T - \hat p_C (see class documentation for the full Newcombe interval method).

Usage
InferenceIncidNewcombeRiskDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.


InferenceIncidNewcombeRiskDiff$compute_estimate_with_bootstrap_weights()

Recomputes the risk-difference estimate under subject/block bootstrap weights: the weighted event proportions \hat p_T^w = \sum_i r_i y_i \mathbb{1}[w_i=1] / \sum_i r_i \mathbb{1}[w_i=1] (and analogously for control), differenced. Used by the Bayesian bootstrap and related weighted-resampling machinery; see InferenceBayesianBootstrap. Always leaves the standard error and degrees of freedom unavailable (NA) regardless of estimate_only — the Newcombe interval method has no separate variance quantity to compute on this path.

Usage
InferenceIncidNewcombeRiskDiff$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

Present for interface parity; this method never computes variance components regardless of its value.


InferenceIncidNewcombeRiskDiff$compute_asymp_confidence_interval()

Computes a 1-\alpha Newcombe hybrid Wilson-score confidence interval for the risk difference (see class documentation for the full formula), via newcombe_independent_ci_cpp.

Usage
InferenceIncidNewcombeRiskDiff$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The significance level.


InferenceIncidNewcombeRiskDiff$compute_asymp_two_sided_pval()

Computes a two-sided p-value testing H_0: p_T - p_C = \code{delta} by numerically finding (stats::uniroot) the significance level \alpha at which delta falls exactly on the boundary of the Newcombe confidence interval (see class documentation) — there is no closed-form p-value for this method. Returns 1 if no root is found in (10^{-10}, 1-10^{-10}) (interpreted as delta being far inside the interval at every plausible \alpha).

Usage
InferenceIncidNewcombeRiskDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null risk difference.


InferenceIncidNewcombeRiskDiff$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidNewcombeRiskDiff$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Newcombe, R. G. (1998). "Interval Estimation for the Difference Between Independent Proportions: Comparison of Eleven Methods." Statistics in Medicine, 17(8), 873-890, doi:10.1002/(SICI)1097-0258(19980430)17:8<873::AID-SIM779>3.0.CO;2-I, for "Method 10", the hybrid Wilson-score interval used here.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidNewcombeRiskDiff$new(seq_des)
inf$compute_estimate()


Probit Regression Inference for Incidence Responses

Description

Fits a probit regression model for binary (incidence) responses: \Phi^{-1}(P(Y_i = 1)) = \beta_0 + \beta_T W_i + X_i^\top \gamma, where \Phi is the standard normal CDF, W_i is the treatment indicator, and X_i are optional recorded covariates, by maximum likelihood (fast_probit_regression_cpp/ fast_probit_regression_weighted_cpp). Unlike InferenceIncidLogRegr's logit link, \hat\beta_T here is not an odds-ratio scale parameter: it is the treatment's additive effect on the latent standard-normal index underlying the binary outcome. likelihood_tier = "full": Wald, score, gradient, and likelihood-ratio tests are all available when the model converges, plus parametric-likelihood-bootstrap calibration of the likelihood-ratio test. A fit whose coefficients exceed max_abs_reasonable_coef in magnitude (a proxy for near-perfect separation) is cached as nonestimable rather than returned. Validity requires the usual probit assumptions: correctly specified linear predictor on the latent-normal scale, independence across subjects conditional on covariates, and no perfect/quasi-complete separation.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Super class

Inference -> InferenceIncidProbitRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidProbitRegr$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceIncidProbitRegr$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

delta

The null difference. Default 0.

transform_responses

Type of transformation. Default "none".

show_progress

Show progress bar. Default TRUE.

permutations

Pre-computed permutations. Default NULL.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceIncidProbitRegr$supports_rand_pval_for_incidence()

Usage
InferenceIncidProbitRegr$supports_rand_pval_for_incidence()

InferenceIncidProbitRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidProbitRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

McCullagh, P., and Nelder, J. A. (1989). Generalized Linear Models (2nd ed.). Chapman and Hall/CRC, for the binomial GLM family and probit link.

See Also

InferenceIncidLogRegr for the logit-link alternative with a log-odds-ratio estimand. Comparable Python API: statsmodels GLM (family=Binomial(link=probit())). See also: Probit model (Wikipedia).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidProbitRegr$new(seq_des)
inf$compute_estimate()


Risk Difference Inference for Incidence Responses

Description

Fits a linear probability model E[Y \mid w, x] = \beta_0 + \beta_T w + \beta_X^\top x via ordinary least squares for binary (incidence) responses Y \in \{0,1\}, using the treatment indicator w and, optionally, all recorded covariates x as predictors. \hat\beta_T is reported directly as the risk-difference estimate: because Y is 0/1, the OLS fit coincides with a saturated/linear model for the conditional risk P(Y=1\mid w,x), so the coefficient on w is already on the risk-difference (probability) scale with no back-transformation needed. This is a misspecified working model for a binary response (heteroskedastic, errors not Gaussian: \mathrm{Var}(Y \mid w,x) = P(w,x)(1-P(w,x)) varies by treatment arm and covariates, so a single pooled residual variance is the wrong variance model), so likelihood_tier = "none": standard errors and the Wald CI use the Huber-White (HC0) sandwich variance of \hat\beta_T – (X'X)^{-1} X'\,\mathrm{diag}(e_i^2)\,X\, (X'X)^{-1}, from the OLS residuals e_i – not the classical homoskedastic OLS variance and not a binomial likelihood, and no likelihood-ratio or parametric-bootstrap methods are exposed.

Super class

Inference -> InferenceIncidRiskDiff

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidRiskDiff$new()

Uses the randomization-CI layer's two-sided p-value contract (InferenceRandCI's version, not InferenceRand's): for incidence responses it dispatches to the Zhang exact randomization test rather than refusing outright, matching this class's pre-migration old-ladder behavior. This deliberately differs from the InferenceAllSimpleAverageDiff-family precedent of pinning InferenceRand's version, which would have regressed the working Zhang dispatch this class had on the old ladder.

Initialize a risk-difference inference object.

Usage
InferenceIncidRiskDiff$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with an incidence response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.

r

Number of randomization vectors. @param delta Null difference.

transform_responses

Transformation. @param na.rm Remove NAs.

show_progress

Show progress. @param permutations Pre-computed permutations.

type

Optional exact-inference type for incidence dispatch.

args_for_type

Optional arguments for type.

zero_one_logit_clamp

Clamp for exact 0/1 values when logging.


InferenceIncidRiskDiff$compute_estimate()

Fits the OLS linear-probability model and returns the risk-difference point estimate \hat\beta_T, the coefficient on treatment. On a hardened design (private$harden), or when estimate_only = FALSE, delegates to fast_ols_with_var_cpp() via the shared model cache so the variance is available for later confidence-interval/p-value calls without refitting.

Usage
InferenceIncidRiskDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip the variance/degrees-of-freedom computation and return only \hat\beta_T (cheaper for simulation/randomization callers that never request inference).


InferenceIncidRiskDiff$compute_asymp_confidence_interval()

Wald confidence interval for the risk difference, \hat\beta_T \pm z_{1-\alpha/2}\, \hat s(\hat\beta_T), using the Huber-White (HC0) sandwich standard error from the cached model fit (see the class-level documentation) and a normal-quantile multiplier – fast_ols_with_var_cpp() doesn't report residual degrees of freedom, so private$compute_z_or_t_ci_from_s_and_df() always falls back to its z branch here, never a t reference, regardless of n. Interval bounds are not clamped to [-1, 1]; a linear-probability model can produce out-of-range endpoints near the boundary of the covariate space.

Usage
InferenceIncidRiskDiff$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Two-sided miscoverage rate; the returned interval has nominal coverage 1-\alpha.


InferenceIncidRiskDiff$compute_asymp_two_sided_pval()

Two-sided Wald test of H_0: \beta_T = \delta vs. H_1: \beta_T \neq \delta, via the Wald statistic (\hat\beta_T - \delta) / \hat s(\hat\beta_T) (\hat s the Huber-White sandwich SE) referred to a standard normal distribution – same z, not t, reference as $compute_asymp_confidence_interval(), for the same reason (no residual degrees of freedom are ever cached for this class).

Usage
InferenceIncidRiskDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null risk-difference value under H_0 (default 0, no treatment effect).


InferenceIncidRiskDiff$compute_estimate_with_bootstrap_weights()

Refits the linear-probability model by weighted least squares (stats::lm.wfit()) under subject/block resampling weights (nonparametric-bootstrap replicate weights or Bayesian-bootstrap Dirichlet weights, both expanded to row weights via expand_subject_or_block_weights_to_row_weights()) and returns the re-estimated treatment coefficient. Rows with non-finite or non-positive weight are excluded; if too few positive-weight rows remain to identify the design, the replicate estimate is NA_real_.

Usage
InferenceIncidRiskDiff$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Numeric weights, one per subject or per resampling block (matched pair/cluster), as produced by the bootstrap or Bayesian-bootstrap resampling machinery.

estimate_only

Accepted for interface compatibility; standard errors are never computed for a single bootstrap replicate regardless of this flag (only the point estimate is used to build the bootstrap distribution).


InferenceIncidRiskDiff$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidRiskDiff$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidRiskDiff$new(seq_des)
inf$compute_estimate()


Wald Incidence Inference

Description

Unadjusted incidence inference using the empirical risk difference \hat\theta = \bar Y_T - \bar Y_C (sample proportions in the treatment and control arms) together with the standard unpooled Wald standard error \hat s(\hat\theta) = \sqrt{\bar Y_T(1-\bar Y_T)/n_T + \bar Y_C(1-\bar Y_C)/n_C} and normal-approximation confidence interval / hypothesis test \hat\theta \pm z_{1-\alpha/2}\,\hat s(\hat\theta). This is the classical two-proportion Wald interval (e.g. Wald 1943; see InferenceIncidNewcombeRiskDiff for a small-sample-robust alternative). likelihood_tier = "none": no likelihood is fit, so no likelihood-ratio or parametric-bootstrap methods are exposed; the reservoir/covariate structure is ignored, unlike InferenceIncidRiskDiff's covariate- adjusted linear-probability model. Non-estimable if either arm has zero subjects.

Super class

Inference -> InferenceIncidWald

Methods

Public methods

+ inherited public methods from Inference

InferenceIncidWald$new()

Uses the randomization-CI layer's two-sided p-value contract (InferenceRandCI's version, not InferenceRand's): for incidence responses this dispatches to the Zhang exact randomization test where applicable rather than refusing outright, matching this class's pre-migration old-ladder behavior (it inherited from InferenceAllSimpleAverageDiff, whose own pin was already corrected to InferenceRandCI – see that file's identical rationale). This class independently composes the same components rather than truly inheriting InferenceAllSimpleAverageDiff, so it had its own stale copy of the old InferenceRand pin, which silently regressed Zhang dispatch ('compute_rand_two_sided_pval()' started throwing "Randomization tests are not supported for incidence" for the same designs the old ladder handled correctly) – found via test-incid-wald-migration-golden.R's randomization_pval case going from '"ok"' to '"unsupported"'.

Pins asymptotic CI dispatch to the composed Wald component's implementation (InferenceAsymp), which uses private$get_standard_error() (this class's two-proportion Wald SE) and private$get_degrees_of_freedom(). Without this explicit pin, the SimpleMeanDifference component – composed after Wald in this class's components list – silently wins the assembly-order collision and dispatches its own Welch's t-test on raw y instead, making the documented Wald formula dead code. See class documentation.

Pins asymptotic p-value dispatch to the composed Wald component's implementation; see $compute_asymp_confidence_interval() for the rationale.

Initialize Wald risk-difference incidence inference and prepare the treatment/control binomial summaries used by InferenceIncidWald and related InferenceIncidRiskDiff methods.

Usage
InferenceIncidWald$new(des_obj, model_formula = NULL, verbose = FALSE)
Arguments
des_obj

A completed design object.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.

delta

Null treatment effect value.

Returns

A new InferenceIncidWald object.


InferenceIncidWald$clone()

The objects of this class are cloneable with this method.

Usage
InferenceIncidWald$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidWald$new(seq_des)
inf$compute_estimate()


Jackknife-based Inference

Description

Abstract class for delete-1 jackknife estimate correction and jackknife-Wald inference layered on top of bootstrap-capable inference classes.

Super classes

Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap -> InferenceJackknife

Methods

Public methods

+ inherited public methods from InferenceBayesianBootstrap
+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference

InferenceJackknife$approximate_jackknife_distribution_beta_hat_T()

Returns the leave-one-out jackknife estimate distribution.

Usage
InferenceJackknife$approximate_jackknife_distribution_beta_hat_T(unit = "auto")
Arguments
unit

Deletion unit. Default '\"auto\"', which chooses a design-aware unit automatically.

Returns

A numeric vector of jackknife replicate estimates.


InferenceJackknife$compute_jackknife_estimate()

Computes the delete-1 jackknife bias-corrected treatment estimate.

For blocking designs, this uses leave-one-block-out deletion units. For matching designs, it uses leave-match-out deletion units. For KK designs, it uses leave-match-out for matched pairs and leave-one-out for reservoir subjects.

Usage
InferenceJackknife$compute_jackknife_estimate(unit = "auto")
Arguments
unit

Deletion unit. Default '\"auto\"', which chooses a design-aware unit automatically.

Returns

A numeric jackknife bias-corrected treatment estimate.


InferenceJackknife$compute_jackknife_bias_estimate()

Computes the jackknife bias estimate.

Usage
InferenceJackknife$compute_jackknife_bias_estimate(unit = "auto")
Arguments
unit

Deletion unit. Default '\"auto\"'.

Returns

A numeric jackknife bias estimate.


InferenceJackknife$compute_jackknife_std_error()

Computes the delete-1 jackknife standard error.

For blocking designs, this uses leave-one-block-out deletion units. For matching designs, it uses leave-match-out deletion units. For KK designs, it uses leave-match-out for matched pairs and leave-one-out for reservoir subjects.

Usage
InferenceJackknife$compute_jackknife_std_error(unit = "auto")
Arguments
unit

Deletion unit. Default '\"auto\"', which chooses a design-aware unit automatically.

Returns

A numeric jackknife standard error.


InferenceJackknife$compute_jackknife_wald_two_sided_pval()

Computes a two-sided Wald p-value using the jackknife estimate and jackknife standard error.

For blocking designs, this uses leave-one-block-out deletion units. For matching designs, it uses leave-match-out deletion units. For KK designs, it uses leave-match-out for matched pairs and leave-one-out for reservoir subjects.

Usage
InferenceJackknife$compute_jackknife_wald_two_sided_pval(
  delta = 0,
  unit = "auto"
)
Arguments
delta

Null treatment-effect value. Default 0.

unit

Deletion unit. Default '\"auto\"', which chooses a design-aware unit automatically.

Returns

A two-sided jackknife-Wald p-value.


InferenceJackknife$compute_jackknife_wald_confidence_interval()

Computes a normal-approximation confidence interval using the jackknife estimate and jackknife standard error.

For blocking designs, this uses leave-one-block-out deletion units. For matching designs, it uses leave-match-out deletion units. For KK designs, it uses leave-match-out for matched pairs and leave-one-out for reservoir subjects.

Usage
InferenceJackknife$compute_jackknife_wald_confidence_interval(
  alpha = 0.05,
  unit = "auto"
)
Arguments
alpha

Significance level. Default 0.05.

unit

Deletion unit. Default '\"auto\"', which chooses a design-aware unit automatically.

Returns

A jackknife-Wald confidence interval.


InferenceJackknife$clone()

The objects of this class are cloneable with this method.

Usage
InferenceJackknife$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Internal Base Class for KK Matching-on-the-Fly Designs

Description

Internal method. An abstract R6 class that provides relevant methods when the designs are KK matching-on-the-fly.

Initialize

Recomputes the class-specific treatment estimate under bootstrap weights; see InferenceBayesianBootstrap.

Usage

kk_passthrough_compound_host_public

Inference for A Sequential Design

Description

An abstract R6 Class that provides asymptotic tests and intervals for a treatment effect in a sequential design where the common denominator is a summary table from a glm.

Super classes

Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap -> InferenceJackknife -> InferenceAsymp -> InferenceMLEorKMSummaryTable

Methods

Public methods

+ inherited public methods from InferenceAsymp
+ inherited public methods from InferenceJackknife
+ inherited public methods from InferenceBayesianBootstrap
+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference

InferenceMLEorKMSummaryTable$compute_estimate()

Computes the appropriate estimate for mean difference

Usage
InferenceMLEorKMSummaryTable$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.

Returns

The setting-appropriate (see description) numeric estimate of the treatment effect

Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "continuous")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10])
seq_des$add_all_subject_responses(c(4.71, 1.23, 4.78, 6.11, 5.95, 8.43))

seq_des_inf = InferenceContinOLS$new(seq_des)
seq_des_inf$compute_estimate()

InferenceMLEorKMSummaryTable$compute_asymp_confidence_interval()

Computes a 1-alpha level frequentist confidence interval

Usage
InferenceMLEorKMSummaryTable$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.

Returns

A (1 - alpha)-sized frequentist confidence interval for the treatment effect


InferenceMLEorKMSummaryTable$compute_asymp_two_sided_pval()

Compute a two-sided p-value for model-summary-table inference by using the cached treatment estimate and standard error from the fitted model or Kaplan-Meier summary. See InferenceMLEorKMSummaryTable and InferenceAsymp.

Usage
InferenceMLEorKMSummaryTable$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null difference to test against. For any treatment effect at all this is set to zero (the default).

Returns

The approximate frequentist p-value


InferenceMLEorKMSummaryTable$clone()

The objects of this class are cloneable with this method.

Usage
InferenceMLEorKMSummaryTable$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


## ------------------------------------------------
## Method `InferenceMLEorKMSummaryTable$compute_estimate()`
## ------------------------------------------------

seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "continuous")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10])
seq_des$add_all_subject_responses(c(4.71, 1.23, 4.78, 6.11, 5.95, 8.43))

seq_des_inf = InferenceContinOLS$new(seq_des)
seq_des_inf$compute_estimate()


Marginal vs. Conditional Estimand Switch

Description

Component scaffold providing set_estimand()/ get_estimand()/get_supported_estimands() for classes whose reported treatment coefficient is conditional on a latent mixture component – e.g. the interior beta submodel in zero/one-inflated beta regression, or the count-process submodel in zero-augmented/hurdle Poisson – rather than the unconditional response mean E[Y]. See marginal_estimand_report.md for the full design discussion.

Mirrors the testing_type switch (InferenceAsympLik) on its own, orthogonal axis: default "conditional" (today's behavior, fully backward compatible – every class that does not compose this component is implicitly conditional-only), with "marginal_mean_diff" and "marginal_ratio" available to classes that declare support for them via their own get_supported_estimands_impl() override (the same override pattern already used for get_supported_testing_types_impl()).

Scope note (2026-08-18): this component provides only the get/set/supported-values switch and the cache-key helper. It does not itself compute any marginal estimate – no class currently composes it. Wiring a class's own compute_estimate() to consult self$get_estimand() and, for a marginal estimand, call into a family-specific model-implied mean function plus a shared g-computation-average/delta-method-gradient helper, is marginal_estimand_report.md → TODO-4/5/9 – deliberately deferred until their target classes (still on the legacy deep-hierarchy ladder as of this writing) migrate to the shallow hierarchy under fix_inference_hierarchy.md's Full-Likelihood Estimators remainder. compute_estimate() itself stays 100 percent class-owned either way – this component never overrides or wraps it, so no allowed_host_overrides declaration is needed.

Methods

Public methods


InferenceMarginalEstimand$set_estimand()

Sets the target estimand for this inference object.

Usage
InferenceMarginalEstimand$set_estimand(estimand)
Arguments
estimand

One of get_supported_estimands(). Accepts the canonical values ("conditional", "marginal_mean_diff", "marginal_ratio") case-insensitively.

Details

If this object also composes LikelihoodTests (checked via self$supports("likelihood_tests"), the sanctioned capability query – see marginal_estimand_report.md → TODO-6), switching to a non-"conditional" estimand shrinks the set of supported testing types to "wald" only. If the currently configured testing_type is no longer in that shrunk set, this errors loudly and leaves the estimand unchanged, rather than silently leaving the object in an inconsistent state – the same guarantee holds regardless of which of set_testing_type()/ set_estimand() is called first.

Returns

The inference object, invisibly.


InferenceMarginalEstimand$get_estimand()

Gets the current target estimand.

Usage
InferenceMarginalEstimand$get_estimand()
Returns

A character scalar, one of get_supported_estimands().


InferenceMarginalEstimand$get_supported_estimands()

Gets the estimands supported by this inference object.

Usage
InferenceMarginalEstimand$get_supported_estimands()
Returns

A character vector. Always includes "conditional". Canonicalizes a requested estimand value, rejecting anything not in the fixed set of recognized spellings – unlike 'get_supported_estimands_impl()' below (host-overridable, varies by class), this recognizes syntax, not per-class support. Default: every class implicitly supports only the conditional estimand until it declares otherwise. Concrete classes override this private method (via ‘define_inference_class()'’s 'overrides' argument, the same pattern 'get_supported_testing_types_impl()' already uses) once they wire a marginal mean function. Cache-key fragment for the current estimand, generalizing ‘likelihood_test_delta_key()'’s testing_type/delta pattern to this orthogonal axis. Any cache keyed partly by estimand should prefix or combine with this so a cache entry built under one estimand is never reused under another.


InferenceMarginalEstimand$clone()

The objects of this class are cloneable with this method.

Usage
InferenceMarginalEstimand$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Bootstrap-based Inference

Description

Abstract class for bootstrap-based inference.

The default m = NULL rule is a cheap deterministic intermediate sequence: m \to \infty and m / n \to 0, as required by the standard m-out-of-n bootstrap asymptotic setup (Bickel, Gotze, and van Zwet; Bickel and Sakov). The exponent 0.7 is a pragmatic interior point in (0, 1); it is not a silver-bullet optimal choice. Use select_optimal_m_out_of_n_bootstrap() for data-adaptive minimum-volatility selection.

The default m = NULL follows the intermediate-sequence convention from the m-out-of-n bootstrap literature: m \to \infty and m / n \to 0. The deterministic exponent 0.7 is a first-pass default; for unstable paths prefer the minimum-volatility selector.

The NULL default is grounded in the standard m-out-of-n asymptotic condition m \to \infty and m / n \to 0. The minimum-volatility selector is available when a fixed deterministic exponent is too brittle for a specific estimator/design path.

This implements the same minimum-volatility idea used in PTE's m-selection workflow: scan admissible intermediate sizes and choose a stable region of the target statistic rather than assuming one exponent is uniformly optimal.

The default b = NULL rule is a cheap deterministic intermediate sequence: b \to \infty and b / n \to 0, as required by the Politis, Romano, and Wolf subsampling framework. The exponent 0.7 is a pragmatic interior point in (0, 1); it is not a universal optimum. Use select_optimal_b_subsampling() for data-adaptive minimum-volatility selection.

The default b = NULL follows the intermediate-sequence convention from the Politis/Romano/Wolf subsampling literature: b \to \infty and b / n \to 0. The deterministic exponent 0.7 is a first-pass default; for unstable paths prefer the minimum-volatility selector.

The NULL default is grounded in the standard Politis/Romano/Wolf asymptotic condition b \to \infty and b / n \to 0. The minimum-volatility selector is available when a fixed deterministic exponent is too brittle for a specific estimator/design path.

This implements the same minimum-volatility idea used in PTE's m-selection workflow, applied to the PRW block/subsample-size choice: scan admissible intermediate sizes and choose a stable region of the target statistic rather than assuming one exponent is uniformly optimal.

Design-specific validity caveats for the nonparametric bootstrap

Nonparametric bootstrap methods are not supported for DesignSeqOneByOne designs and their subclasses, except the concrete DesignSeqOneByOneBernoulli class. The restriction includes m-out-of-n bootstrap and subsampling methods registered under the same capability. Calling restricted methods reports that the method is not supported.

For supported designs, the nonparametric bootstrap resamples experimental units with replacement from their empirical distribution, carrying each unit's realized (x, w, y) into the replicate, and recomputes the estimator. Its validity rests on the resampled units being (approximately) iid draws from the design's unit-level superpopulation. Since covariate-adaptive designs induce dependence among the assignments w_i (and between w and X), the appropriate resampling unit and the fidelity with which the design's dependence is replicated differ by design. In all cases below the inference is asymptotic, never finite-sample exact (for exact finite-sample inference under the design's actual randomization mechanism, use the randomization tests and randomization confidence intervals where their assumptions hold). Calibration depends on the design and estimator; omitted dependence does not in general guarantee conservative inference.

DesignFixedBernoulli and DesignSeqOneByOneBernoulli

Assignments are independent coin flips that do not use covariates or past assignments. If subjects and their potential outcomes arrive iid, with a noninformative sample size, observed rows are iid and ordinary row-level resampling has its usual asymptotic justification for regular estimators. Independent assignments alone do not establish iid rows under time trends, dependent recruitment, or informative stopping.

DesignFixediBCRD

Assignment depends only on the treatment counts (completely randomized / without-replacement urn), inducing negative correlation among the w_i through the fixed-margin constraint. Row-level iid resampling does not replicate this constraint: replicates have a random number of treated subjects. The extra variability is O(1/n), so the bootstrap is conservative by an asymptotically negligible amount.

DesignFixedBlocking

Resampling is within-strata by default (bootstrap_type = "within_blocks"), preserving stratum sizes; the exact within-stratum treatment/control split is not enforced in replicates, so the block-randomization variance reduction is partially unreplicated. Conservative, minor. bootstrap_type = "resample_blocks" instead resamples whole blocks, preserving within-block composition at the price of fewer resampling atoms.

DesignFixedOptimalBlocks

Same within-block resampling caveats as DesignFixedBlocking, plus the blocks themselves are computed from the realized covariate sample: the block structure is a global function of the data that the bootstrap conditions on rather than re-derives. The justification for this conditioning is asymptotic: as n grows the blocking depends on the sample only through the (convergent) empirical distribution of X, so between-block dependence vanishes. Conservative.

DesignFixedCluster

Assignment is at the cluster level and outcomes are correlated within clusters, so whole clusters are resampled with replacement. This is the correct exchangeable unit; with few clusters the bootstrap distribution rests on few resampling atoms and becomes unstable. Asymptotics are in the number of clusters, not the number of subjects.

DesignFixedBlockedCluster

Clusters are resampled within strata, matching both levels of the design's dependence (stratum and cluster). Sound, with the same small-sample caution: few clusters per stratum means few resampling atoms per stratum, and asymptotics are in the number of clusters.

DesignFixedGreedyDOptimal, DesignFixedGreedy, DesignFixedRerandomization

The observed w vector is one draw from a tightly constrained (optimized or acceptance-sampled) set of allocations. Resampled replicates carry per-row assignments whose recombined w vector no longer satisfies the balance constraint, so the bootstrap reflects the variance of unconstrained assignment (cf. Li, Ding & Rubin 2018 for rerandomization). Conservative, moderate-to-large: the stronger the optimization, the greater the over-coverage.

DesignFixedMatchingGreedyPairSwitching

The greedy switching search only ever flips assignments within binary-match pairs, so every pair has exactly one treated subject; the bootstrap resamples intact pairs, preserving the within-pair anticorrelation. Remaining caveats: the pairing is a global function of the sample (conditioned on, justified asymptotically as for the matched designs below), and the greedy choice of which pair member is treated couples the pairs, which resampling does not replicate — the residual effect errs conservative.

DesignFixedBinaryMatch

Matched pairs are resampled intact, preserving the within-pair anticorrelation of w and the pair-level variance reduction. The pairing itself is a global function of the covariate sample (an Abadie & Imbens 2008-type concern): pairs are exchangeable but not exactly independent. Validity is asymptotic — as n grows the pairing depends on the sample only through the empirical distribution of X and between-pair dependence vanishes — and the bootstrap conditions on the realized match structure.

DesignFixedFactorial

Row-level resampling does not replicate the balanced allocation across factor combinations. Conservative, minor.

DesignFixedCustom

Warning: iid row-level resampling is used because the package has no knowledge of the user-supplied assignment mechanism. If that mechanism balances on covariates, the bootstrap is likely conservative; if it induces clustering or other positive dependence, the bootstrap may not even be valid (anti-conservative). Use the randomization-based inference, which draws from the actual custom mechanism, whenever possible.

Super classes

Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap

Methods

Public methods

+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference

InferenceNonParamBootstrap$get_supported_bootstrap_pval_types()

Returns the type values compute_bootstrap_two_sided_pval() accepts.

Usage
InferenceNonParamBootstrap$get_supported_bootstrap_pval_types()

InferenceNonParamBootstrap$get_supported_bootstrap_ci_types()

Returns the type values compute_bootstrap_confidence_interval() accepts.

Usage
InferenceNonParamBootstrap$get_supported_bootstrap_ci_types()

InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()

Creates the m-out-of-n bootstrap distribution of the treatment-effect estimate.

Usage
InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T(
  B = 501,
  m = NULL,
  show_progress = TRUE,
  debug = FALSE,
  bootstrap_type = NULL,
  scaling = "sqrt_n",
  center = "full_estimate"
)
Arguments
B

Number of resamples. Default 501.

m

Number of exchangeable resampling units drawn with replacement. If NULL (default), use the deterministic intermediate-size rule floor(n_units^0.7), where n_units is the number of exchangeable units used by the design (observations, clusters, pairs, or matched sets). The resolved value must satisfy max(5, p_eff + 2) <= m <= floor(n_units / 2).

show_progress

A flag indicating whether a progress bar should be displayed.

debug

If TRUE, return distribution diagnostics in addition to the resampled estimates.

bootstrap_type

Optional empirical-resampling scheme. See approximate_bootstrap_distribution_beta_hat_T().

scaling

Scaling sequence for centered m-out-of-n pivots. The default "sqrt_n" uses sqrt(m) for the m-sample distribution and converts back to the full-sample scale using sqrt(n_units).

center

Centering convention for diagnostics and cache keys.

Returns

A numeric vector of bootstrap estimates, or when debug = TRUE, a diagnostic list.


InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval()

Computes a centered m-out-of-n bootstrap two-sided p-value.

Usage
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval(
  delta = 0,
  B = 501,
  m = NULL,
  type = "centered",
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  bootstrap_type = NULL,
  scaling = "sqrt_n"
)
Arguments
delta

Null treatment effect. Default 0.

B

Number of resamples. Default 501.

m

Number of exchangeable units drawn with replacement. If NULL (default), use floor(n_units^0.7) subject to the validation bounds documented for approximate_m_out_of_n_bootstrap_distribution_beta_hat_T().

type

P-value type. Currently only "centered" is supported.

show_progress

A flag indicating whether a progress bar should be displayed.

min_number_usable_samples

Minimum number of finite resampled estimates required after filtering. Default 5.

bootstrap_type

Optional empirical-resampling scheme.

scaling

Scaling sequence for centered m-out-of-n pivots.

Returns

A numeric two-sided p-value, or NA_real_ if the path is non-estimable.


InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval()

Computes a basic m-out-of-n bootstrap confidence interval.

Usage
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 501,
  m = NULL,
  type = "basic",
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  bootstrap_type = NULL,
  scaling = "sqrt_n"
)
Arguments
alpha

Significance level. Default 0.05.

B

Number of resamples. Default 501.

m

Number of exchangeable units drawn with replacement. If NULL (default), use floor(n_units^0.7) subject to the validation bounds documented for approximate_m_out_of_n_bootstrap_distribution_beta_hat_T().

type

Confidence-interval type. Currently only "basic" is supported.

show_progress

A flag indicating whether a progress bar should be displayed.

min_number_usable_samples

Minimum number of finite resampled estimates required after filtering. Default 5.

bootstrap_type

Optional empirical-resampling scheme.

scaling

Scaling sequence for centered m-out-of-n pivots.

Returns

A length-2 numeric confidence interval, or c(NA_real_, NA_real_) if the path is non-estimable.


InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap()

Selects an m-out-of-n bootstrap size by minimum volatility.

Usage
InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap(
  B = 251,
  alpha = 0.05,
  m_pow_of_n_grid = seq(0.5, 0.9, by = 0.05),
  m_grid = NULL,
  objective = "ci_width",
  target = "ci",
  volatility_window = 3L,
  bootstrap_type = NULL,
  scaling = "sqrt_n",
  show_progress = TRUE,
  min_finite_fraction = 0.8
)
Arguments
B

Number of resamples per candidate size. Default 251.

alpha

Significance level for interval-width objectives.

m_pow_of_n_grid

Candidate exponent grid used when m_grid = NULL. Defaults to seq(0.5, 0.9, by = 0.05).

m_grid

Optional explicit integer candidate sizes.

objective

Selection objective. Currently "ci_width".

target

Target summary. Currently "ci".

volatility_window

Rolling window size used to measure local volatility across candidate sizes.

bootstrap_type

Optional empirical-resampling scheme.

scaling

Scaling sequence for centered m-out-of-n pivots.

show_progress

A flag indicating whether a progress bar should be displayed.

min_finite_fraction

Minimum finite-resample fraction required for a candidate size to be eligible.

Returns

An EDIMOutOfNBootstrapMSelection list with the selected m, mapped exponent, candidate table, status, and reason.


InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T()

Creates the Politis/Romano/Wolf subsampling distribution of the treatment-effect estimate.

Usage
InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T(
  B = 501,
  b = NULL,
  show_progress = TRUE,
  debug = FALSE,
  subsampling_type = NULL,
  scaling = "sqrt_n",
  center = "full_estimate"
)
Arguments
B

Number of subsamples. Default 501.

b

Number of exchangeable units drawn without replacement. If NULL (default), use the deterministic intermediate-size rule floor(n_units^0.7), where n_units is the number of exchangeable units used by the design (observations, clusters, pairs, or matched sets). The resolved value must satisfy max(5, p_eff + 2) <= b <= floor(n_units / 2).

show_progress

A flag indicating whether a progress bar should be displayed.

debug

If TRUE, return distribution diagnostics in addition to the subsampled estimates.

subsampling_type

Optional empirical-resampling scheme. See approximate_bootstrap_distribution_beta_hat_T().

scaling

Scaling sequence for centered subsampling pivots. The default "sqrt_n" uses sqrt(b) for the subsample distribution and converts back to the full-sample scale using sqrt(n_units).

center

Centering convention for diagnostics and cache keys.

Returns

A numeric vector of subsampled estimates, or when debug = TRUE, a diagnostic list.


InferenceNonParamBootstrap$compute_subsampling_two_sided_pval()

Computes a centered PRW subsampling two-sided p-value.

Usage
InferenceNonParamBootstrap$compute_subsampling_two_sided_pval(
  delta = 0,
  B = 501,
  b = NULL,
  type = "centered",
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  subsampling_type = NULL,
  scaling = "sqrt_n"
)
Arguments
delta

Null treatment effect. Default 0.

B

Number of subsamples. Default 501.

b

Number of exchangeable units drawn without replacement. If NULL (default), use floor(n_units^0.7) subject to the validation bounds documented for approximate_subsampling_distribution_beta_hat_T().

type

P-value type. Currently only "centered" is supported.

show_progress

A flag indicating whether a progress bar should be displayed.

min_number_usable_samples

Minimum number of finite subsampled estimates required after filtering. Default 5.

subsampling_type

Optional empirical-resampling scheme.

scaling

Scaling sequence for centered subsampling pivots.

Returns

A numeric two-sided p-value, or NA_real_ if the path is non-estimable.


InferenceNonParamBootstrap$compute_subsampling_confidence_interval()

Computes a basic PRW subsampling confidence interval.

Usage
InferenceNonParamBootstrap$compute_subsampling_confidence_interval(
  alpha = 0.05,
  B = 501,
  b = NULL,
  type = "basic",
  show_progress = TRUE,
  min_number_usable_samples = 5L,
  subsampling_type = NULL,
  scaling = "sqrt_n"
)
Arguments
alpha

Significance level. Default 0.05.

B

Number of subsamples. Default 501.

b

Number of exchangeable units drawn without replacement. If NULL (default), use floor(n_units^0.7) subject to the validation bounds documented for approximate_subsampling_distribution_beta_hat_T().

type

Confidence-interval type. Currently only "basic" is supported.

show_progress

A flag indicating whether a progress bar should be displayed.

min_number_usable_samples

Minimum number of finite subsampled estimates required after filtering. Default 5.

subsampling_type

Optional empirical-resampling scheme.

scaling

Scaling sequence for centered subsampling pivots.

Returns

A length-2 numeric confidence interval, or c(NA_real_, NA_real_) if the path is non-estimable.


InferenceNonParamBootstrap$select_optimal_b_subsampling()

Selects a PRW subsampling size by minimum volatility.

Usage
InferenceNonParamBootstrap$select_optimal_b_subsampling(
  B = 251,
  alpha = 0.05,
  b_pow_of_n_grid = seq(0.5, 0.9, by = 0.05),
  b_grid = NULL,
  objective = "ci_width",
  target = "ci",
  volatility_window = 3L,
  subsampling_type = NULL,
  scaling = "sqrt_n",
  show_progress = TRUE,
  min_finite_fraction = 0.8
)
Arguments
B

Number of subsamples per candidate size. Default 251.

alpha

Significance level for interval-width objectives.

b_pow_of_n_grid

Candidate exponent grid used when b_grid = NULL. Defaults to seq(0.5, 0.9, by = 0.05).

b_grid

Optional explicit integer candidate sizes.

objective

Selection objective. Currently "ci_width".

target

Target summary. Currently "ci".

volatility_window

Rolling window size used to measure local volatility across candidate sizes.

subsampling_type

Optional empirical-resampling scheme.

scaling

Scaling sequence for centered subsampling pivots.

show_progress

A flag indicating whether a progress bar should be displayed.

min_finite_fraction

Minimum finite-subsample fraction required for a candidate size to be eligible.

Returns

An EDISubsamplingBSelection list with the selected b, mapped exponent, candidate table, status, and reason.


InferenceNonParamBootstrap$compute_subsampling_sensitivity()

Computes PRW subsampling sensitivity over candidate sizes.

Usage
InferenceNonParamBootstrap$compute_subsampling_sensitivity(
  B = 251,
  alpha = 0.05,
  b_pow_of_n_grid = seq(0.5, 0.9, by = 0.05),
  b_grid = NULL,
  objective = "ci_width",
  target = "ci",
  volatility_window = 3L,
  subsampling_type = NULL,
  scaling = "sqrt_n",
  show_progress = TRUE,
  min_finite_fraction = 0
)
Arguments
B

Number of subsamples per candidate size. Default 251.

alpha

Significance level for interval-width objectives.

b_pow_of_n_grid

Candidate exponent grid used when b_grid = NULL. Defaults to seq(0.5, 0.9, by = 0.05).

b_grid

Optional explicit integer candidate sizes.

objective

Selection objective. Currently "ci_width".

target

Target summary. Currently "ci".

volatility_window

Rolling window size used to measure local volatility across candidate sizes.

subsampling_type

Optional empirical-resampling scheme.

scaling

Scaling sequence for centered subsampling pivots.

show_progress

A flag indicating whether a progress bar should be displayed.

min_finite_fraction

Minimum finite-subsample fraction required for a candidate size to be eligible. Defaults to 0 for sensitivity scans.

Returns

An EDISubsamplingSensitivity list containing the candidate grid table without selecting a final b.


InferenceNonParamBootstrap$approximate_bootstrap_distribution_beta_hat_T()

Creates the bootstrap distribution of the estimate for the treatment effect. The resampling unit is design-specific (rows, within-strata rows, matched pairs plus reservoir, or clusters); see the class-level section Design-specific validity caveats for the nonparametric bootstrap for the conservativeness and asymptotics of each concrete design.

Usage
InferenceNonParamBootstrap$approximate_bootstrap_distribution_beta_hat_T(
  B = 501,
  show_progress = TRUE,
  debug = FALSE,
  bootstrap_type = NULL
)
Arguments
B

Number of bootstrap samples. The default is 501.

show_progress

A flag indicating whether a progress bar should be displayed.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

bootstrap_type

Optional bootstrap-resampling scheme. Legal public values are:

NULL

Use the design's default row-resampling bootstrap. For ordinary non-blocking designs this is the usual subject-level resample-with-replacement bootstrap. For certain blocking designs, NULL maps to the same behavior as "within_blocks".

"within_blocks"

Only legal for blocking-style designs that support block-aware bootstrap resampling: DesignFixedBlocking, DesignFixedOptimalBlocks, DesignSeqOneByOneSPBR, and DesignFixedBlockedCluster. Resamples observational units within each observed block/stratum. For blocked cluster designs this means resampling clusters within strata.

"resample_blocks"

Only legal for the same blocking-style designs as "within_blocks". Resamples entire observed blocks/strata with replacement rather than resampling units within each block.

Any non-NULL value is rejected for designs outside that blocking family.

Returns

When debug = FALSE (default), a numeric vector of length B containing the bootstrap estimates. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.


InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval()

Computes a bootstrap-based two-sided p-value for the treatment effect. Validity is asymptotic and design-dependent; for most covariate-adaptive designs the p-value errs conservative. See the class-level section Design-specific validity caveats for the nonparametric bootstrap.

Usage
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval(
  delta = 0,
  B = 501,
  type = NULL,
  na.rm = FALSE,
  show_progress = TRUE,
  min_number_usable_samples = 5L
)
Arguments
delta

Null hypothesis value. Default 0.

B

Number of bootstrap samples. Default 501.

type

Bootstrap p-value type. Supported values are "percentile" (default), "symmetric", "studentized", "bootstrap-t", and "bca". "percentile": shifts the bootstrap distribution to be centred at delta and counts the two-tail proportion (Hall 1992). "symmetric": uses |T^* - \bar{T}^*| \ge |t_{\rm obs} - \delta| for a symmetric one-sample test; recommended by Hall & Wilson (1991) when the null distribution may be skewed. This pooled-tail test is offered only as a p-value here, not as a confidence-interval type in compute_bootstrap_confidence_interval: pooling both tails via |\cdot| improves testing power (Hall & Wilson's original use case), but inverting it unstudentized would add no value as an interval. The unstudentized pivot is not asymptotically pivotal, so the resulting interval would have the same first-order O(n^\{-1/2\}) coverage error as "percentile"/ "basic", while forcing symmetric bounds around a possibly skewed bootstrap distribution — strictly worse than "percentile"/"basic" for shape-adaptivity, and strictly worse than "symmetric-percentile-t" for accuracy, since studentizing (not the absolute-value pooling) is what buys the O(n^\{-1\}) improvement. The CI-worthy symmetric variant is therefore "symmetric-percentile-t" (studentized pivot), not a plain "symmetric" CI type. "studentized" / "bootstrap-t": pivots by the per-replicate standard error, giving O(n^{-1}) error versus O(n^{-1/2}) for the percentile method (Hall 1992; Davidson & MacKinnon 1999). "bca": bias-corrected and accelerated p-value via closed-form CI inversion using the jackknife acceleration and bias-correction constants; second-order accurate (Efron 1987; Efron & Tibshirani 1993).

na.rm

Remove non-finite bootstrap replicates. Default FALSE.

show_progress

A flag indicating whether a progress bar should be displayed.

min_number_usable_samples

Minimum number of finite bootstrap samples required after filtering. Default 5. Must be less than or equal to B.

Returns

A bootstrap two-sided p-value.


InferenceNonParamBootstrap$compute_bootstrap_confidence_interval()

Computes a bootstrap-based confidence interval. Coverage is asymptotic and design-dependent; for most covariate-adaptive designs the interval errs conservative (over-coverage). See the class-level section Design-specific validity caveats for the nonparametric bootstrap.

Usage
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 501,
  type = NULL,
  na.rm = TRUE,
  show_progress = TRUE,
  min_number_usable_samples = 5L
)
Arguments
alpha

The confidence level 1 - alpha. Default 0.05.

B

Number of bootstrap samples. Default 501.

type

Bootstrap CI type. Supported values are "percentile", "basic", "studentized", "bootstrap-t", "symmetric-percentile-t", "bca", "prepivoted", "double-bootstrap", "calibrated", and "smoothed". There is no plain "symmetric" CI type (contrast with the "symmetric" p-value type in compute_bootstrap_two_sided_pval): inverting the unstudentized Hall & Wilson pooled-tail statistic would add no value as an interval, since it is not asymptotically pivotal and so has the same first-order O(n^\{-1/2\}) coverage error as "percentile"/"basic", while forcing symmetric bounds around a possibly skewed bootstrap distribution — strictly worse than "percentile"/ "basic" for shape-adaptivity, and strictly worse than "symmetric-percentile-t" for accuracy, since studentizing (not the absolute-value pooling) is what buys the O(n^\{-1\}) improvement. "symmetric-percentile-t" is the CI-worthy symmetric variant.

na.rm

Remove non-finite bootstrap replicates. Default TRUE. Non-finite replicates are always removed internally.

show_progress

Show progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required after filtering. Default 5. Must be less than or equal to B.

Returns

A bootstrap confidence interval.


InferenceNonParamBootstrap$clone()

The objects of this class are cloneable with this method.

Usage
InferenceNonParamBootstrap$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Bickel, P. J., Gotze, F., and van Zwet, W. R. (1997). Resampling fewer than n observations: gains, losses, and remedies for losses. Statistica Sinica.

Bickel, P. J. and Sakov, A. (2008). On the choice of m in the m out of n bootstrap. The Annals of Statistics.

Politis, D. N., Romano, J. P., and Wolf, M. (1999). Subsampling. Springer.


Adjacent Category Logit Regression Inference for Ordinal Responses

Description

Fits an adjacent-category logit regression for ordinal responses (via fast_adjacent_category_logit_cpp — see that page for the full model, an alternative ordinal parameterization to the cumulative-logit proportional-odds model) using the treatment indicator and, optionally, all recorded covariates as predictors. This is a full-likelihood class (likelihood_tier = "full") supporting score, gradient, and likelihood-ratio tests, plus parametric likelihood-ratio bootstrap calibration, in addition to Wald and resampling-based inference. Bayesian-bootstrap inference is temporarily unavailable because the current non-uniform weighted hook fits a cumulative-logit surrogate rather than the adjacent-category likelihood. It will remain disabled until the native weighted adjacent-category backend described in the package implementation plan lands.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceOrdinalAdjCatLogitRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalAdjCatLogitRegr$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceOrdinalAdjCatLogitRegr$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalAdjCatLogitRegr$supports_rand_pval_for_incidence()

Usage
InferenceOrdinalAdjCatLogitRegr$supports_rand_pval_for_incidence()

InferenceOrdinalAdjCatLogitRegr$compute_rand_two_sided_pval()

Usage
InferenceOrdinalAdjCatLogitRegr$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalAdjCatLogitRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalAdjCatLogitRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalAdjCatLogitRegr$new(seq_des)
inf$compute_estimate()


Cauchit Regression Inference for Ordinal Responses

Description

Cauchit-link cumulative-odds ordinal regression: P(Y \le k \mid w, x) = F_{\mathrm{Cauchy}}(\alpha_k - \beta_T w - \beta_X^\top x), where F_{\mathrm{Cauchy}} is the standard Cauchy CDF, \alpha_k are category-specific cutpoints, and \beta_T is the treatment log-odds coefficient on the cauchit scale (proportional-odds-style shift common to all categories). Fit by maximum likelihood. The heavy-tailed Cauchy link is markedly less sensitive to outlying/extreme response categories than the logit or probit link, at the cost of a less familiar effect-size interpretation. likelihood_tier = "full": exposes likelihood-ratio, score, gradient, and parametric-likelihood-bootstrap inference in addition to the Wald/asymptotic and Bayesian-bootstrap paths.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceOrdinalCauchitRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalCauchitRegr$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceOrdinalCauchitRegr$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalCauchitRegr$supports_rand_pval_for_incidence()

Usage
InferenceOrdinalCauchitRegr$supports_rand_pval_for_incidence()

InferenceOrdinalCauchitRegr$compute_rand_two_sided_pval()

Usage
InferenceOrdinalCauchitRegr$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalCauchitRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalCauchitRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Agresti, A. (2010). Analysis of Ordinal Categorical Data (2nd ed.). Wiley. Ch. 3-4 (cumulative link models).

See Also

https://en.wikipedia.org/wiki/Ordinal_regression, https://www.statsmodels.org/stable/discretemod.html for an analogous Python cumulative-link API.


Cumulative Cloglog Inference for Ordinal Responses

Description

Complementary log-log cumulative-odds ordinal regression: P(Y \le k \mid w, x) = 1 - \exp\{-\exp(\alpha_k - \beta_T w - \beta_X^\top x)\}, where \alpha_k are category-specific cutpoints and \beta_T is the treatment coefficient on the cloglog scale. Fit by maximum likelihood. The cloglog link is asymmetric (unlike logit/probit) and is the natural ordinal generalization of a proportional-hazards/grouped survival-time model, so it is preferred when the underlying process is plausibly a discretized time-to-event or extreme-value mechanism. likelihood_tier = "full": exposes likelihood-ratio, score, gradient, and parametric-likelihood-bootstrap inference in addition to Wald/asymptotic and Bayesian-bootstrap paths.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceOrdinalCloglogRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalCloglogRegr$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceOrdinalCloglogRegr$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalCloglogRegr$supports_rand_pval_for_incidence()

Usage
InferenceOrdinalCloglogRegr$supports_rand_pval_for_incidence()

InferenceOrdinalCloglogRegr$compute_rand_two_sided_pval()

Usage
InferenceOrdinalCloglogRegr$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalCloglogRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalCloglogRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Agresti, A. (2010). Analysis of Ordinal Categorical Data (2nd ed.). Wiley. Ch. 3-4 (cumulative link models); McCullagh, P. (1980). "Regression Models for Ordinal Data." JRSS-B, 42(2), 109-142.

See Also

https://en.wikipedia.org/wiki/Ordinal_regression


Continuation Ratio Regression Inference for Ordinal Responses

Description

Fits a conditional (stratified) continuation-ratio logit model for ordinal responses: for cut j = 1, \dots, K-1, among subjects who have reached at least category j,

\log\frac{\Pr(Y_i > j \mid Y_i \ge j)}{\Pr(Y_i = j \mid Y_i \ge j)} = \alpha_j + \beta_T W_i + X_i^\top \gamma,

a discrete-time-hazard-model analog for ordinal data, with a treatment coefficient \beta_T constrained equal across all cuts. \exp(\hat\beta_T) is the common "continue vs. stop here" odds ratio: a positive \beta_T means treatment pushes subjects toward higher categories of Y, matching the sign convention of every other ordinal estimator in the package. Fitting proceeds by expand_continuation_ratio_data_cpp's stacked-binary expansion followed by conditional logistic regression on the expanded data. likelihood_tier = "full": likelihood-ratio, score, gradient, and Wald tests are all available when the model converges, plus parametric-likelihood-bootstrap calibration of the likelihood-ratio test. Validity requires the continuation-ratio proportionality assumption (a common \beta_T across all K-1 cuts).

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceOrdinalContRatioRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalContRatioRegr$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceOrdinalContRatioRegr$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalContRatioRegr$supports_rand_pval_for_incidence()

Usage
InferenceOrdinalContRatioRegr$supports_rand_pval_for_incidence()

InferenceOrdinalContRatioRegr$compute_rand_two_sided_pval()

Usage
InferenceOrdinalContRatioRegr$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalContRatioRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalContRatioRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Agresti, A. (2010). Analysis of Ordinal Categorical Data (2nd ed.). Wiley, for the continuation-ratio model family.

See Also

InferenceOrdinalStereotypeLogitRegr and InferenceOrdinalKKCondAdjCatLogitRegr for related ordinal-logit expansions. See also: Ordinal regression (Wikipedia).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalContRatioRegr$new(seq_des)
inf$compute_estimate()


G-Computation Mean-Difference Inference for Ordinal Responses

Description

Fits a proportional-odds working model for an ordinal outcome using treatment and, optionally, all recorded covariates (fast_ordinal_regression_with_var_cpp), then estimates the marginal difference in expected ordinal category score by G-computation — see gcomp_ordinal_proportional_odds_post_fit_cpp for the exact standardization formula (mean1 - mean0). Standard errors are obtained by the delta method: a central finite-difference gradient of the mean-difference functional with respect to the fitted [\alpha, \beta] parameters, propagated through the model's fitted variance-covariance matrix, \widehat{\mathrm{Var}}(\widehat{\mathrm{md}}) = \nabla^\top \widehat{\mathrm{Var}}(\hat\theta) \nabla. If that delta-method standard error is unavailable or non-finite, the Wald-style methods ($compute_asymp_confidence_interval(), $compute_asymp_two_sided_pval(), $compute_wald_confidence_interval(), $compute_wald_two_sided_pval()) all silently fall back to a nonparametric bootstrap interval/p-value instead (with a warning), rather than returning NA.

Super class

Inference -> InferenceOrdinalGCompMeanDiff

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalGCompMeanDiff$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize the ordinal g-computation (G-Comp) inference object for a completed design with an ordinal, uncensored response.

Usage
InferenceOrdinalGCompMeanDiff$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed DesignSeqOneByOne object with an ordinal response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values by default.


InferenceOrdinalGCompMeanDiff$compute_estimate()

Computes the G-computation standardized mean-difference treatment-effect estimate (see class documentation for the full proportional-odds-based standardization).

Usage
InferenceOrdinalGCompMeanDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.


InferenceOrdinalGCompMeanDiff$compute_estimate_with_bootstrap_weights()

Recomputes the G-computation mean-difference estimate under subject/block bootstrap weights (via fast_ordinal_regression_weighted_cpp plus gcomp_ordinal_proportional_odds_post_fit_cpp), used by the Bayesian bootstrap and related weighted-resampling machinery; see InferenceNonParamBootstrap. Runs side-effect free: the ordinary (unweighted) cached fit, warm-start state, and rank-reduced column selection are saved before the weighted refit and restored afterward (on.exit), so a weighted bootstrap replicate cannot corrupt the class's own point estimate or subsequent fits.

Usage
InferenceOrdinalGCompMeanDiff$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Row weights for the bootstrap sample.

estimate_only

If TRUE, skip variance calculations.


InferenceOrdinalGCompMeanDiff$compute_asymp_confidence_interval()

Computes a 1-\alpha confidence interval for the G-Comp mean difference using the delta-method standard error (see class documentation), or falls back (with a warning) to a nonparametric bootstrap interval if that standard error is unavailable. Identical to $compute_wald_confidence_interval().

Usage
InferenceOrdinalGCompMeanDiff$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The significance level (default 0.05).


InferenceOrdinalGCompMeanDiff$compute_asymp_two_sided_pval()

Computes a two-sided Wald p-value testing H_0: \mathrm{md} = \code{delta} using the delta-method standard error (see class documentation), or falls back (with a warning) to a nonparametric bootstrap p-value if that standard error is unavailable. Identical to $compute_wald_two_sided_pval().

Usage
InferenceOrdinalGCompMeanDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null treatment effect (default 0).


InferenceOrdinalGCompMeanDiff$compute_wald_confidence_interval()

Identical to $compute_asymp_confidence_interval() (both compute the same delta-method-based Wald interval, with the same bootstrap fallback); provided as an explicit alias for callers that want to name the Wald method directly rather than via the generic "asymptotic" dispatch.

Usage
InferenceOrdinalGCompMeanDiff$compute_wald_confidence_interval(alpha = 0.05)
Arguments
alpha

The significance level (default 0.05).


InferenceOrdinalGCompMeanDiff$compute_wald_two_sided_pval()

Identical to $compute_asymp_two_sided_pval() (both compute the same delta-method-based Wald p-value, with the same bootstrap fallback); provided as an explicit alias for callers that want to name the Wald method directly rather than via the generic "asymptotic" dispatch.

Usage
InferenceOrdinalGCompMeanDiff$compute_wald_two_sided_pval(delta = 0)
Arguments
delta

The null treatment effect (default 0).


InferenceOrdinalGCompMeanDiff$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalGCompMeanDiff$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalGCompMeanDiff$new(seq_des)
inf$compute_estimate()


Jonckheere-Terpstra (JT) Test for Ordinal Responses

Description

Two-arm Jonckheere-Terpstra (JT) rank test for an ordinal response — for two groups, this reduces to the Mann-Whitney U statistic. The point estimate is the stochastic superiority probability, centered at 0 under the null: \hat\beta_T = \widehat{\Pr}(Y_T > Y_C) + \tfrac{1}{2}\widehat{\Pr}(Y_T = Y_C) - \tfrac{1}{2}, computed from category counts as U/(n_T n_C) - 1/2. Asymptotic inference ($compute_asymp_confidence_interval(), $compute_asymp_two_sided_pval()) uses the classical null variance of the Mann-Whitney U statistic, \mathrm{Var}(U) = n_T n_C (n_T+n_C+1)/12 (no tie correction), matching clinfun::jonckheere.test()'s normal approximation. This class also provides an exact, permutation-distribution-based two-sided p-value via $compute_exact_two_sided_pval_for_treatment_effect() (exact_jonckheere_terpstra_pval_cpp), which does not rely on the normal approximation.

Super class

Inference -> InferenceOrdinalJonckheereTerpstraTest

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalJonckheereTerpstraTest$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize the JT test object for a completed design with an ordinal, uncensored response.

Usage
InferenceOrdinalJonckheereTerpstraTest$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE
)
Arguments
des_obj

A completed DesignSeqOneByOne object.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress.


InferenceOrdinalJonckheereTerpstraTest$compute_estimate()

Returns the estimated treatment effect: the stochastic superiority measure \widehat{\Pr}(Y_T > Y_C) + \tfrac12\widehat{\Pr}(Y_T=Y_C) - \tfrac12, computed from the Mann-Whitney U statistic (see class documentation).

Usage
InferenceOrdinalJonckheereTerpstraTest$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.


InferenceOrdinalJonckheereTerpstraTest$compute_estimate_with_bootstrap_weights()

Recomputes the JT superiority estimate under subject/block bootstrap weights: the weighted version of the same stochastic superiority quantity, \sum_{i,j} w_i w_j\left(\mathbb{1}[y_{T,i} > y_{C,j}] + \tfrac12\mathbb{1}[y_{T,i}=y_{C,j}]\right) \big/ \sum_{i,j} w_i w_j - \tfrac12, used by the Bayesian bootstrap and related weighted-resampling machinery. Always leaves the standard error unavailable (NA) regardless of estimate_only — this weighted path never computes the null-variance approximation.

Usage
InferenceOrdinalJonckheereTerpstraTest$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject/block level.

estimate_only

Present for interface parity; this method never computes variance components regardless of its value.


InferenceOrdinalJonckheereTerpstraTest$compute_exact_two_sided_pval_for_treatment_effect()

Returns the exact, permutation-distribution-based two-sided p-value (exact_jonckheere_terpstra_pval_cpp) — unlike $compute_asymp_two_sided_pval(), this does not rely on the normal approximation to the Mann-Whitney U null distribution.

Usage
InferenceOrdinalJonckheereTerpstraTest$compute_exact_two_sided_pval_for_treatment_effect(
  
)

InferenceOrdinalJonckheereTerpstraTest$compute_asymp_confidence_interval()

Computes the asymptotic normal confidence interval, using the same Mann-Whitney U null-variance approximation (n_T n_C(n_T+n_C+1)/12, no tie correction) as clinfun::jonckheere.test(); see class documentation.

Usage
InferenceOrdinalJonckheereTerpstraTest$compute_asymp_confidence_interval(
  alpha = 0.05
)
Arguments
alpha

The significance level (default 0.05).


InferenceOrdinalJonckheereTerpstraTest$compute_asymp_two_sided_pval()

Computes the asymptotic normal two-sided p-value, using the same Z-approximation as clinfun::jonckheere.test(); see class documentation and $compute_exact_two_sided_pval_for_treatment_effect() for the exact (non-approximate) alternative.

Usage
InferenceOrdinalJonckheereTerpstraTest$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null treatment effect (default 0).


InferenceOrdinalJonckheereTerpstraTest$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalJonckheereTerpstraTest$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Jonckheere, A. R. (1954). "A Distribution-Free k-Sample Test Against Ordered Alternatives." Biometrika, 41(1-2), 133-145, doi:10.1093/biomet/41.1-2.133; Terpstra, T. J. (1952). "The Asymptotic Normality and Consistency of Kendall's Test Against Trend, When Ties Are Present in One Ranking." Indagationes Mathematicae, 14, 327-333.

Examples

set.seed(1)
x_dat <- data.frame(
  x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
  x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneBernoulli$new(n = nrow(x_dat), response_type = "ordinal",
  verbose = FALSE)
for (i in seq_len(nrow(x_dat))) {
  seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(as.integer(c(1, 2, 2, 3, 3, 4, 4, 5)))
infer <- InferenceOrdinalJonckheereTerpstraTest$
  new(seq_des, verbose = FALSE)
infer


Ordinal KK CLMM (Proportional Odds / logit link)

Description

Cumulative-link random-intercept mixed model for ordinal responses under a KK matching-on-the-fly design, using the logit link (proportional odds): \mathrm{logit}(P(Y_i \le k)) = \alpha_k - (\beta_T W_i + X_i^\top \gamma) - b_{g(i)}, b_g \sim N(0, \sigma_b^2), where g(i) is subject i's matched-pair group id (reservoir subjects get singleton groups). \exp(\hat\beta_T) is the (conditional-on-b_g) treatment odds ratio. See InferenceAbstractKKOrdinalCLMM for the shared model-fitting, caching, and likelihood-tier contract common to all four link-function siblings (...Probit, ...Cauchit, ...Cloglog); this class supplies only the link-function choice (private$clmm_link() == "logit").

Super classes

Inference -> InferenceAbstractKKOrdinalCLMM -> InferenceOrdinalKKCLMM

Methods

Public methods

+ inherited public methods from InferenceAbstractKKOrdinalCLMM
+ inherited public methods from Inference

InferenceOrdinalKKCLMM$new()

Initialize the logit-link ordinal KK CLMM subclass; see the shared ordinal mixed-model contract in InferenceAbstractKKOrdinalCLMM.

Usage
InferenceOrdinalKKCLMM$new(
  des_obj,
  model_formula = NULL,
  use_rcpp = TRUE,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object.

model_formula

Optional formula for covariate adjustment.

use_rcpp

Use internal Rcpp implementation (default TRUE).

verbose

Print messages?

smart_cold_start_default

Use smart cold start values?


InferenceOrdinalKKCLMM$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalKKCLMM$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKCLMM$new(seq_des)
inf$compute_estimate()


Ordinal KK CLMM (Cauchit link)

Description

Cumulative-link random-intercept mixed model for ordinal responses under a KK matching-on-the-fly design, using the cauchit (inverse-Cauchy-CDF) link: \tan(\pi (P(Y_i \le k) - 1/2)) = \alpha_k - (\beta_T W_i + X_i^\top \gamma) - b_{g(i)}, b_g \sim N(0, \sigma_b^2), where g(i) is subject i's matched-pair group id. The cauchit link's heavy-tailed latent distribution makes it more robust than logit/probit to a small number of subjects near the response's extreme categories. See InferenceAbstractKKOrdinalCLMM for the shared model-fitting, caching, and likelihood-tier contract common to all four link-function siblings; this class supplies only the link-function choice (private$clmm_link() == "cauchit").

Super classes

Inference -> InferenceAbstractKKOrdinalCLMM -> InferenceOrdinalKKCLMMCauchit

Methods

Public methods

+ inherited public methods from InferenceAbstractKKOrdinalCLMM
+ inherited public methods from Inference

InferenceOrdinalKKCLMMCauchit$new()

Initialize the cauchit-link ordinal KK CLMM subclass; see the shared ordinal mixed-model contract in InferenceAbstractKKOrdinalCLMM.

Usage
InferenceOrdinalKKCLMMCauchit$new(
  des_obj,
  model_formula = NULL,
  use_rcpp = TRUE,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object.

model_formula

Optional formula for covariate adjustment.

use_rcpp

Use internal Rcpp implementation (default TRUE).

verbose

Print messages?

smart_cold_start_default

Use smart cold start values?


InferenceOrdinalKKCLMMCauchit$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalKKCLMMCauchit$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKCLMMCauchit$new(seq_des)
inf$compute_estimate()


Ordinal KK CLMM (Complementary log-log link)

Description

Cumulative-link random-intercept mixed model for ordinal responses under a KK matching-on-the-fly design, using the complementary log-log link: \log(-\log(1 - P(Y_i \le k))) = \alpha_k - (\beta_T W_i + X_i^\top \gamma) - b_{g(i)}, b_g \sim N(0, \sigma_b^2), where g(i) is subject i's matched-pair group id. Unlike the symmetric logit/probit/ cauchit links, the cloglog link is asymmetric, making it appropriate when the ordinal categories arise from an underlying continuous-time proportional-hazards process discretized into intervals. See InferenceAbstractKKOrdinalCLMM for the shared model-fitting, caching, and likelihood-tier contract common to all four link-function siblings; this class supplies only the link-function choice (private$clmm_link() == "cloglog").

Super classes

Inference -> InferenceAbstractKKOrdinalCLMM -> InferenceOrdinalKKCLMMCloglog

Methods

Public methods

+ inherited public methods from InferenceAbstractKKOrdinalCLMM
+ inherited public methods from Inference

InferenceOrdinalKKCLMMCloglog$new()

Initialize the cloglog-link ordinal KK CLMM subclass; see the shared ordinal mixed-model contract in InferenceAbstractKKOrdinalCLMM.

Usage
InferenceOrdinalKKCLMMCloglog$new(
  des_obj,
  model_formula = NULL,
  use_rcpp = TRUE,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object.

model_formula

Optional formula for covariate adjustment.

use_rcpp

Use internal Rcpp implementation (default TRUE).

verbose

Print messages?

smart_cold_start_default

Use smart cold start values?


InferenceOrdinalKKCLMMCloglog$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalKKCLMMCloglog$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKCLMMCloglog$new(seq_des)
inf$compute_estimate()


Ordinal KK CLMM (Probit link)

Description

Cumulative-link random-intercept mixed model for ordinal responses under a KK matching-on-the-fly design, using the probit link: \Phi^{-1}(P(Y_i \le k)) = \alpha_k - (\beta_T W_i + X_i^\top \gamma) - b_{g(i)}, b_g \sim N(0, \sigma_b^2), where \Phi is the standard normal CDF and g(i) is subject i's matched-pair group id. Unlike the logit-link sibling, \hat\beta_T here is not an odds-ratio scale parameter; it is the treatment's effect on the latent standard-normal index underlying the ordinal categories. See InferenceAbstractKKOrdinalCLMM for the shared model-fitting, caching, and likelihood-tier contract common to all four link-function siblings; this class supplies only the link-function choice (private$clmm_link() == "probit").

Super classes

Inference -> InferenceAbstractKKOrdinalCLMM -> InferenceOrdinalKKCLMMProbit

Methods

Public methods

+ inherited public methods from InferenceAbstractKKOrdinalCLMM
+ inherited public methods from Inference

InferenceOrdinalKKCLMMProbit$new()

Initialize the probit-link ordinal KK CLMM subclass; see the shared ordinal mixed-model contract in InferenceAbstractKKOrdinalCLMM.

Usage
InferenceOrdinalKKCLMMProbit$new(
  des_obj,
  model_formula = NULL,
  use_rcpp = TRUE,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object.

model_formula

Optional formula for covariate adjustment.

use_rcpp

Use internal Rcpp implementation (default TRUE).

verbose

Print messages?

smart_cold_start_default

Use smart cold start values?


InferenceOrdinalKKCLMMProbit$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalKKCLMMProbit$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKCLMMProbit$new(seq_des)
inf$compute_estimate()


Adjacent Category Logit Inference for KK Matching-on-the-fly Designs

Description

Fits a conditional (stratified) adjacent-category logit model for ordinal responses under a KK matching-on-the-fly design:

\log\frac{\Pr(Y_i = j+1 \mid Y_i \in \{j, j+1\})}{\Pr(Y_i = j \mid Y_i \in \{j, j+1\})} = \alpha_j + \beta_T W_i + X_i^\top \gamma,

for adjacent category comparisons j = 1, \dots, K-1, with cut-specific intercepts \alpha_j and a treatment coefficient \beta_T constrained equal across all cuts (the parallel/proportional adjacent-category assumption). \exp(\hat\beta_T) is the common adjacent-category odds ratio. Fitting proceeds by expand_adjacent_category_data_cpp's stacked-binary expansion (each subject contributes a 0/1 row per adjacent cut they border, stratified by matched pair) followed by conditional logistic regression on the expanded data — the matched-pair identity becomes the conditioning stratum, so the pair's shared nuisance intercept is conditioned out exactly as in a single binary conditional-logit KK model, and reservoir (unmatched) subjects each form their own singleton stratum. likelihood_tier = "partial" (a conditional/partial likelihood, matched-set effects are profiled out rather than estimated); supports_likelihood_tests() is hard FALSE — only Wald inference is exposed, not likelihood-ratio, score, or gradient tests. Validity requires the adjacent-category proportionality assumption (a common \beta_T across all K-1 cuts) in addition to the usual conditional-logit exchangeability-within-strata assumption induced by the KK design.

Super class

Inference -> InferenceOrdinalKKCondAdjCatLogitRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalKKCondAdjCatLogitRegr$new()

Initialize inference for the conditional adjacent-category logit model \log(\Pr(Y_i = j+1 \mid Y_i \in \{j,j+1\}) / \Pr(Y_i = j \mid Y_i \in \{j,j+1\})) = \alpha_j + \beta_T W_i + X_i^\top \gamma and prepare KK matched-pair structure for the stratified conditional-logit fit. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferenceOrdinalKKCondAdjCatLogitRegr$new(
  des_obj,
  verbose = FALSE,
  harden = TRUE,
  model_formula = NULL,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed KK DesignSeqOneByOne object with an ordinal response.

verbose

Flag for progress messages.

harden

Whether to apply robustness measures.

model_formula

Optional formula for covariate adjustment.

smart_cold_start_default

Whether to use smart cold start values by default.


InferenceOrdinalKKCondAdjCatLogitRegr$compute_estimate()

Fits the conditional adjacent-category logit model via stacked-binary expansion (expand_adjacent_category_data_cpp) plus conditional logistic regression, and returns the shared log-odds-ratio estimate \hat\beta_T.

Usage
InferenceOrdinalKKCondAdjCatLogitRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.


InferenceOrdinalKKCondAdjCatLogitRegr$compute_estimate_with_bootstrap_weights()

Recomputes the treatment estimate under subject/block-level bootstrap weights (Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights()). When weights are effectively constant, this collapses to the unweighted compute_estimate() call. Otherwise, rather than refitting the full expanded conditional-logit model under weights, it calls weighted_ordinal_bootstrap_surrogate_fit() — a fast weighted ordinal-logistic surrogate fit on the raw (unexpanded) design matrix — as an approximation to the weighted adjacent-category likelihood; this trades exact reweighted refitting for speed across many bootstrap replicates. No standard error is computed (s_beta_hat_T is always NA); the surrogate returns NA if the fit fails.

Usage
InferenceOrdinalKKCondAdjCatLogitRegr$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Subject-, block-, cluster-, or matched-set bootstrap weights.

estimate_only

If TRUE, compute only the weighted point estimate.


InferenceOrdinalKKCondAdjCatLogitRegr$compute_asymp_confidence_interval()

Wald confidence interval for the shared adjacent-category log-odds-ratio \beta_T, using the conditional-logit model's standard error; see InferenceAsymp for the shared Wald contract. Fits the model first if not already cached.

Usage
InferenceOrdinalKKCondAdjCatLogitRegr$compute_asymp_confidence_interval(
  alpha = 0.05
)
Arguments
alpha

Two-sided miscoverage rate; the returned interval targets 1 - alpha coverage.


InferenceOrdinalKKCondAdjCatLogitRegr$compute_asymp_two_sided_pval()

Return the adjacent-category conditional-logit asymptotic p-value for the treatment coefficient, using the shared Wald semantics documented in InferenceAsymp.

Usage
InferenceOrdinalKKCondAdjCatLogitRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null hypothesis treatment effect.


InferenceOrdinalKKCondAdjCatLogitRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalKKCondAdjCatLogitRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Agresti, A. (2010). Analysis of Ordinal Categorical Data (2nd ed.). Wiley, for the adjacent-category logit model family; Kapelner, A. and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the KK matching-on-the-fly design this class is built for.

See Also

InferenceOrdinalAdjCatLogitRegr for the non-KK analog. See also: Ordinal regression (Wikipedia).

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKCondAdjCatLogitRegr$new(seq_des)
inf$compute_estimate()


GEE Inference for KK Designs with Ordinal Response

Description

Fits a proportional-odds local-odds-ratio Generalized Estimating Equations model, via multgee::ordLORgee, for ordinal responses under a KK matching-on-the-fly design, using the treatment indicator and, optionally, all recorded covariates as predictors. Each GEE cluster is either a matched pair (2 members) or a reservoir singleton (1 member) — GEE is used here purely to fit one marginal cumulative-logit model jointly across matched-pair and reservoir subjects while accounting for the within-pair correlation the matching induces, not as a longitudinal/repeated-measures tool. Unlike the other ⁠Inference*KKGEE⁠ classes in this family (continuous/count/incidence/proportion, which use an internal Rcpp solver or geepack::geeglm with an exchangeable working correlation), this ordinal class always requires the multgee package and has no use_rcpp option. The raw multgee treatment coefficient is negated when reported so that, consistently with EDI's other ordinal estimators, a positive estimate means movement toward higher response categories. Inference is quasi-likelihood/ estimating-equation based (likelihood_tier = "quasi"): standard errors are GEE sandwich (robust) standard errors, not model-likelihood-based. Bayesian-bootstrap inference is temporarily unavailable because multgee::ordLORgee does not accept the non-uniform observation weights needed to refit the same clustered estimator. It will remain disabled until the weighted ordinal-GEE implementation planned for v1.1.0 is complete.

Super class

Inference -> InferenceOrdinalKKGEE

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalKKGEE$new()

Initialize KK ordinal GEE inference, validate the ordinal matched/reservoir design, and prepare the multgee::ordLORgee proportional-odds local-odds-ratio GEE fitting machinery used by InferenceOrdinalKKGEE. Requires the multgee package; errors at construction if it is not installed.

Usage
InferenceOrdinalKKGEE$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with an ordinal response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.


InferenceOrdinalKKGEE$compute_estimate_with_bootstrap_weights()

Recomputes the KK ordinal treatment estimate under subject/block bootstrap weights, used by the Bayesian bootstrap and related weighted-resampling machinery. If the supplied weights are all (numerically) equal, this short-circuits to the unweighted $compute_estimate(estimate_only = TRUE) (the multgee proportional-odds GEE fit) rather than refitting. Otherwise, since multgee::ordLORgee does not support observation weights, this falls back to a different, approximating model: a plain (non-GEE, no matched-pair clustering) weighted proportional-odds ordinal logistic regression via fast_ordinal_regression_weighted_cpp, treating the coefficient on the first predictor column as the treatment effect. This always leaves the standard error and degrees of freedom unavailable (s_beta_hat_T = NA, df = Inf) regardless of estimate_only, since it is a point-estimate-only fallback path.

Usage
InferenceOrdinalKKGEE$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Subject-, block-, cluster-, or matched-set bootstrap weights.

estimate_only

If TRUE, compute only the weighted point estimate. Has no effect on the weighted (non-uniform-weight) fallback path, which never computes a standard error regardless.


InferenceOrdinalKKGEE$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalKKGEE$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Touloumis, A. (2015). "R Package multgee: A Generalized Estimating Equations Solver for Multinomial Responses." Journal of Statistical Software, 64(8), 1-14, doi:10.18637/jss.v064.i08, for the local-odds-ratio GEE solver used here; Liang, K.-Y., and Zeger, S. L. (1986). "Longitudinal Data Analysis Using Generalized Linear Models." Biometrika, 73(1), 13-22, doi:10.1093/biomet/73.1.13, for the underlying GEE estimating-equation framework.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKGEE$new(seq_des)
inf$compute_estimate()


GLMM Inference for KK Designs with Ordinal Response

Description

Fits a cumulative-logit random-intercept mixed model (proportional odds) for ordinal responses under a KK matching-on-the-fly design: \mathrm{logit}(P(Y_i \le k \mid w_i, x_i, b_{g(i)})) = \alpha_k - (\beta_T w_i + x_i^\top \gamma) - b_{g(i)}, for cutpoints \alpha_1 < \cdots < \alpha_{K-1}, treatment indicator w_i, covariates x_i, and a matched-pair random intercept b_g \sim N(0, \sigma_b^2) that is integrated out of the marginal likelihood (either by adaptive Gauss-Hermite quadrature when use_rcpp = TRUE, the default; see fast_ordinal_glmm_cpp for the quadrature order and optimizer details, or by glmmTMB's Laplace approximation when use_rcpp = FALSE). g(i) is subject i's matched-pair group id; reservoir (unmatched) subjects each get their own singleton group, contributing no within-group correlation but still entering the joint likelihood. The treatment coefficient \beta_T is a log-odds-ratio: \exp(\beta_T) is the (conditional-on-b_g) odds ratio of being at or above any given response category. likelihood_tier = "full": likelihood-ratio, score, Wald, and gradient tests are available when use_rcpp = TRUE and the fit converges; use_rcpp = FALSE disables likelihood-test support (private$supports_likelihood_tests() returns FALSE) because glmmTMB's Laplace-approximate likelihood is not wired into this package's score/gradient/LR machinery. Validity requires the random-intercept structure to correctly capture the design's matching dependence, proportional odds (the treatment/covariate effect is constant across cutpoints), and correct specification of the fixed-effects formula.

This differs from the GEE sibling InferenceOrdinalKKGEE (documented above) in estimand and inference basis: the GLMM's \beta_T is a subject-specific (conditional) log-odds-ratio with model-likelihood-based inference, while the GEE's is a population-averaged (marginal) log-odds-ratio with sandwich-based inference; the two need not numerically agree even on the same data, and the correct choice depends on whether a subject-specific/conditional or population-averaged/marginal treatment effect is of interest.

Super class

Inference -> InferenceOrdinalKKGLMM

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalKKGLMM$new()

Initialize inference for the cumulative-logit random-intercept mixed model \mathrm{logit}(P(Y_i \le k)) = \alpha_k - (\beta_T W_i + X_i^\top \gamma) - b_{g(i)}, b_g \sim N(0, \sigma_b^2), where g(i) is subject i's matched-pair group id (reservoir subjects get singleton groups). Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferenceOrdinalKKGLMM$new(
  des_obj,
  model_formula = NULL,
  use_rcpp = TRUE,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with an ordinal response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

use_rcpp

Logical. If TRUE (default), use internal Rcpp.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.


InferenceOrdinalKKGLMM$compute_estimate()

Fits the cumulative-logit random-intercept mixed model by (adaptive-Gauss-Hermite- or Laplace-)approximate maximum likelihood and returns \hat\beta_T, the estimated treatment log-odds-ratio, conditional on the matched-pair random intercept. Caches the fitted model object, full parameter vector, and (when estimate_only = FALSE) the standard error and degrees of freedom for reuse by compute_asymp_confidence_interval(), compute_asymp_two_sided_pval(), and likelihood-test methods; a fit that fails the kernel's projected-gradient convergence check, produces non-finite parameters, reaches the upper random-effect variance boundary, exceeds private$max_abs_reasonable_coef, or lacks a finite positive treatment-coefficient variance is cached as nonestimable rather than returned. A valid near-zero random-effect variance boundary is accepted using conditional fixed-effect information. The native optimizer retains multistart L-BFGS for basin selection and, only when its finite selected point fails the projected-score tolerance, applies a damped-Newton polish using the numerical Hessian. The polished point is retained only if it remains finite and does not increase the negative log-likelihood; at a valid lower variance boundary, the KKT-satisfied variance coordinate is excluded from that Newton system.

Usage
InferenceOrdinalKKGLMM$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip standard-error/variance-component computation and cache only the point estimate; used by randomization and bootstrap resampling paths where only \hat\beta_T is needed per replicate.


InferenceOrdinalKKGLMM$compute_asymp_confidence_interval()

Wald confidence interval for \beta_T using the fitted model's standard error and degrees of freedom; see InferenceAsymp for the shared \hat\beta_T \pm t_{\alpha/2, df} \cdot \widehat{se}(\hat\beta_T) (or z-based when df = Inf) contract. Fits the model first if not already cached.

Usage
InferenceOrdinalKKGLMM$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Two-sided miscoverage rate; the returned interval targets 1 - alpha coverage.


InferenceOrdinalKKGLMM$compute_asymp_two_sided_pval()

Two-sided Wald test of H_0: \beta_T = \code{delta} against H_1: \beta_T \ne \code{delta}, using the fitted model's standard error and degrees of freedom; see InferenceAsymp for the shared t/z test contract. Fits the model first if not already cached.

Usage
InferenceOrdinalKKGLMM$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Treatment log-odds-ratio value under the null hypothesis.


InferenceOrdinalKKGLMM$compute_estimate_with_bootstrap_weights()

Refits the mixed model with subject/block-level weights applied to each row's contribution to the marginal likelihood (Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded from subject/block level to individual rows via private$expand_subject_or_block_weights_to_row_weights()) and returns the reweighted estimate \hat\beta_T^{(w)}. Uses fast_ordinal_regression_weighted_cpp — an ordinary (non-mixed-effects) weighted cumulative-logit fit, not a reweighted GLMM refit — as a fast approximation to the weighted marginal likelihood; this trades exact random-effects refitting for speed across many bootstrap replicates. When weights are effectively constant, this collapses to the unweighted compute_estimate() call (returns df = Inf to signal a degenerate/skipped bootstrap replicate rather than refitting). Rows with non-finite or non-positive weight, or non-finite response, are dropped from the weighted fit; if no rows remain, the estimate is NA.

Usage
InferenceOrdinalKKGLMM$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Subject-, block-, cluster-, or matched-set bootstrap weights.

estimate_only

If TRUE, compute only the weighted point estimate.


InferenceOrdinalKKGLMM$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalKKGLMM$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Hedeker, D., and Gibbons, R. D. (1994). "A Random-Effects Ordinal Regression Model for Multilevel Analysis." Biometrics, 50(4), 933-944, doi:10.2307/2533433, for the random-effects cumulative-logit model; Pinheiro, J. C., and Bates, D. M. (1995). "Approximations to the Log-Likelihood Function in the Nonlinear Mixed-Effects Model." Journal of Computational and Graphical Statistics, 4(1), 12-35, doi:10.1080/10618600.1995.10474663, for the adaptive Gauss-Hermite quadrature approximation used to integrate out the random intercept.

See Also

Comparable Python API: statsmodels MixedLM (continuous analog; no ordinal-GLMM in statsmodels). See also: Ordinal regression and Mixed model (Wikipedia).

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKGLMM$new(seq_des)
inf$compute_estimate()


Ordered Probit Regression Inference for Ordinal Responses

Description

Fits a cumulative-probit ("ordered probit") model for ordinal responses: \Phi^{-1}(P(Y_i \le k)) = \alpha_k - (\beta_T W_i + X_i^\top \gamma), for cutpoints \alpha_1 < \cdots < \alpha_{K-1}, where \Phi is the standard normal CDF, W_i is the treatment indicator, and X_i are optional recorded covariates, by maximum likelihood (fast_ordinal_probit_regression_cpp/ fast_ordinal_probit_regression_with_var_cpp). As with binary probit regression, \hat\beta_T is not an odds-ratio-scale parameter: it is the treatment's effect on the latent standard-normal index underlying the ordinal categories. likelihood_tier = "full": likelihood-ratio, score, gradient, and Wald tests are all available when the model converges, plus parametric-likelihood-bootstrap calibration of the likelihood-ratio test. Validity requires the proportional/parallel cutpoints assumption (a single \beta_T shared across all cutpoints) in addition to the usual latent-normal-index assumption.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceOrdinalOrderedProbitRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalOrderedProbitRegr$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceOrdinalOrderedProbitRegr$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalOrderedProbitRegr$supports_rand_pval_for_incidence()

Usage
InferenceOrdinalOrderedProbitRegr$supports_rand_pval_for_incidence()

InferenceOrdinalOrderedProbitRegr$compute_rand_two_sided_pval()

Usage
InferenceOrdinalOrderedProbitRegr$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalOrderedProbitRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalOrderedProbitRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

McCullagh, P. (1980). "Regression Models for Ordinal Data." Journal of the Royal Statistical Society, Series B, 42(2), 109-142, doi:10.1111/j.2517-6161.1980.tb01109.x, for the cumulative-link ordinal model family this class's probit link instantiates.

See Also

InferenceOrdinalCauchitRegr, InferenceOrdinalCloglogRegr for other cumulative-link function choices on the same ordinal model family. See also: Ordinal regression and Probit model (Wikipedia).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalOrderedProbitRegr$new(seq_des)
inf$compute_estimate()


Paired Sign Test Inference for KK Designs with Ordinal Response

Description

Fits the classical paired sign test for ordinal responses under a KK matching-on-the-fly design. For each matched pair i with treated member response Y_{i,T} and control member response Y_{i,C}, only the sign of the within-pair difference Y_{i,T} - Y_{i,C} is used; tied pairs (Y_{i,T} = Y_{i,C}) are dropped from the effective sample. The estimand is \theta = P(Y_T > Y_C \mid \text{pair untied}), and the reported treatment effect is \hat\beta_T = \hat p - 0.5, where \hat p is the sample proportion of untied pairs favoring treatment; \beta_T = 0 corresponds to \theta = 0.5 (no directional preference). The standard error is the usual binomial-proportion formula \sqrt{\hat p (1 - \hat p) / n_{\text{eff}}}, where n_{\text{eff}} is the number of untied pairs. Reservoir (unmatched) subjects are not included — this is a purely within-pair test, unlike the IVWC-style classes elsewhere in the KK family that combine matched-pair and reservoir information. likelihood_tier = "none" (supports_likelihood_tests() is hard FALSE): only Wald inference on the proportion scale is exposed. Bootstrap and jackknife are deliberately unsupported and throw explicit errors (see approximate_bootstrap_distribution_beta_hat_T() and approximate_jackknife_distribution_beta_hat_T()), since subject-level resampling or deletion would violate the matched-pair design's dependence structure; randomization inference (compute_rand_two_sided_pval()) remains available since it permutes treatment assignment within the design's own randomization mechanism rather than resampling subjects. Requires a KK matching-on-the-fly design (DesignSeqOneByOneKK14/KK21) or DesignFixedBinaryMatch; a design with no discordant (untied) pairs is cached as nonestimable for the standard error (point estimate 0) or fully nonestimable, per harden.

Super class

Inference -> InferenceOrdinalPairedSignTest

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalPairedSignTest$new()

Uses the shared randomization-test two-sided p-value contract; see InferenceRand. Pinned from plain InferenceRand (not InferenceRandCI) per the established ordinal-class precedent – Zhang dispatch is incidence-only.

Initialize inference for the paired sign test on within-pair response differences Y_{i,T} - Y_{i,C}; see InferenceOrdinalPairedSignTest for the model form. Requires a KK matching-on-the-fly design (DesignSeqOneByOneKK14/KK21) or DesignFixedBinaryMatch. Does not compute the sign-test statistic; that is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferenceOrdinalPairedSignTest$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed KK matching-on-the-fly design object.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.

r

Number of randomization draws.

delta

Null treatment effect.

transform_responses

Optional response transformation.

na.rm

Whether to drop non-finite draws.

show_progress

Whether to show a progress bar.

permutations

Optional pre-computed permutations.

zero_one_logit_clamp

Clamp for logit transforms.


InferenceOrdinalPairedSignTest$compute_estimate()

Computes the pair-sign counts (pos/neg, ties dropped) from the design's matched-pair structure and returns \hat\beta_T = \hat p - 0.5, where \hat p is the proportion of untied pairs favoring treatment. If every pair is tied, the estimate is 0 (no directional preference) and the fit is cached as standard-error-nonestimable (or fully nonestimable when harden = FALSE).

Usage
InferenceOrdinalPairedSignTest$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip the standard-error computation and cache only the point estimate.


InferenceOrdinalPairedSignTest$compute_estimate_with_bootstrap_weights()

Recomputes \hat\beta_T under subject/block-level bootstrap weights (Bayesian-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights()): for each matched pair, a weighted vote is cast toward whichever member has the higher response, using the mean bootstrap weight of the pair's two rows; \hat\beta_T^{(w)} is the weighted proportion of treatment-favoring pairs minus 0.5. No standard error is computed (s_beta_hat_T is always NA). Pairs with no discordant (untied) weighted votes are cached as nonestimable.

Usage
InferenceOrdinalPairedSignTest$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Subject-, block-, cluster-, or matched-set bootstrap weights.

estimate_only

If TRUE, compute only the weighted point estimate.


InferenceOrdinalPairedSignTest$compute_asymp_confidence_interval()

Wald confidence interval for \beta_T = \theta - 0.5 (equivalently, for \theta = P(Y_T > Y_C \mid \text{pair untied})), using the binomial-proportion standard error; see InferenceAsymp for the shared Wald contract. Fits (computes the pair-sign counts) first if not already cached.

Usage
InferenceOrdinalPairedSignTest$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Two-sided miscoverage rate; the returned interval targets 1 - alpha coverage.


InferenceOrdinalPairedSignTest$compute_asymp_two_sided_pval()

Two-sided Wald test of H_0: \theta = 0.5 (equal chance of favoring treatment vs. control among untied pairs) against H_1: \theta \ne 0.5, using the binomial-proportion standard error; see InferenceAsymp for the shared Wald contract. Only delta = 0 is supported (the sign test's null is fixed at no directional preference; a non-zero delta throws).

Usage
InferenceOrdinalPairedSignTest$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null value for \beta_T; must be 0.


InferenceOrdinalPairedSignTest$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalPairedSignTest$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Dixon, W. J., and Mood, A. M. (1946). "The Statistical Sign Test." Journal of the American Statistical Association, 41(236), 557-566, doi:10.2307/2280577, for the classical paired sign test; Kapelner, A. and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the KK matching-on-the-fly design this class is built for.

See Also

Sign test (Wikipedia).

Examples

set.seed(1)
x_dat <- data.frame(
  x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
  x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneKK21$new(n = nrow(x_dat), response_type = "ordinal",
verbose = FALSE)
for (i in seq_len(nrow(x_dat))) {
  seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(as.integer(c(1, 2, 2, 3, 3, 4, 4, 5)))
infer <- InferenceOrdinalPairedSignTest$
  new(seq_des, verbose = FALSE)
infer


Partial Proportional-Odds Regression Inference for Ordinal Responses

Description

Fits a partial proportional-odds cumulative-logit model for an ordinal response: a subset of covariates named in nonparallel are allowed a separate coefficient at each cumulative threshold (relaxing the proportional-odds/parallel-lines assumption for exactly those covariates), while every other covariate — including the treatment indicator, which is always fit as a parallel (proportional) term regardless of nonparallel — keeps one shared coefficient across all thresholds. The reported treatment effect is therefore always a single proportional (threshold-invariant) log-odds shift, even when other covariates' effects are allowed to vary by threshold. When nonparallel is empty, fitting uses this package's fast Rcpp full-proportional-odds solver (fast_ordinal_regression_with_var_cpp); otherwise it falls back, in order, to VGAM::vglm(family = VGAM::cumulative(parallel = ...)), ordinal::clm(nominal = ...), and (only when nonparallel is empty and the earlier fast/VGAM/clm attempts failed) MASS::polr. Each fallback requires its corresponding package to be installed; unavailable packages are silently skipped in favor of the next fallback.

Super class

Inference -> InferenceOrdinalPartialProportionalOddsRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalPartialProportionalOddsRegr$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize partial proportional-odds ordinal regression inference for a completed design with an ordinal, uncensored response.

Usage
InferenceOrdinalPartialProportionalOddsRegr$new(
  des_obj,
  verbose = FALSE,
  harden = TRUE,
  model_formula = NULL,
  nonparallel = character(0),
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed DesignSeqOneByOne object with an ordinal response.

verbose

Whether to print progress messages.

harden

Whether to apply robustness measures.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

nonparallel

Names of covariates (not including "treatment", which is always fit as a parallel/proportional term) allowed a separate coefficient at each cumulative threshold, relaxing the proportional-odds assumption for those covariates specifically.

smart_cold_start_default

Whether to use smart cold start values by default.


InferenceOrdinalPartialProportionalOddsRegr$compute_estimate()

Retrieves the estimated (always-parallel) treatment log-odds shift from the partial proportional-odds fit (see class documentation for the fitting backend cascade).

Usage
InferenceOrdinalPartialProportionalOddsRegr$compute_estimate(
  estimate_only = FALSE
)
Arguments
estimate_only

If TRUE, skip variance component calculations.

Returns

The estimated treatment effect.


InferenceOrdinalPartialProportionalOddsRegr$compute_estimate_with_bootstrap_weights()

Recomputes the partial-proportional-odds treatment estimate under subject/block bootstrap weights, used by the Bayesian bootstrap and related weighted-resampling machinery. If the weights are effectively constant, short-circuits to the unweighted $compute_estimate(estimate_only = TRUE). Otherwise refits with weights via the same backend cascade as the unweighted fit (VGAM/ordinal/MASS::polr, each weighted), and if all of those fail, falls back further to a plain weighted binary-logistic surrogate fit (weighted_ordinal_bootstrap_surrogate_fit(..., method = "logistic")) that does not model the ordinal structure at all. Never computes a standard error on any weighted path (s_beta_hat_T is always NA), regardless of estimate_only.

Usage
InferenceOrdinalPartialProportionalOddsRegr$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Subject-, block-, cluster-, or matched-set bootstrap weights.

estimate_only

Present for interface parity; this method never computes variance components regardless of its value.


InferenceOrdinalPartialProportionalOddsRegr$compute_asymp_confidence_interval()

Computes a Wald-style confidence interval for the treatment log-odds shift, using the model-based standard error from whichever backend (fast Rcpp solver, VGAM, ordinal, or MASS::polr) successfully fit the unweighted model (see class documentation). If that standard error is unavailable (NA, non-finite, or 0) — e.g. because the fit succeeded via a fallback path that doesn't report one — the interval is explicitly marked non-estimable (c(NA, NA)) when private$harden is TRUE, or raises an error otherwise, rather than silently returning a misleading result. Identical to $compute_wald_confidence_interval().

Usage
InferenceOrdinalPartialProportionalOddsRegr$compute_asymp_confidence_interval(
  alpha = 0.05
)
Arguments
alpha

Significance level for the interval.

Returns

A confidence interval for the treatment effect.


InferenceOrdinalPartialProportionalOddsRegr$compute_asymp_two_sided_pval()

Computes a Wald-style two-sided p-value testing H_0: \beta_T = \code{delta}, using the same model-based standard error as $compute_asymp_confidence_interval(); if unavailable, marked non-estimable (NA) or an error is raised, per private$harden — see that method's documentation. Identical to $compute_wald_two_sided_pval().

Usage
InferenceOrdinalPartialProportionalOddsRegr$compute_asymp_two_sided_pval(
  delta = 0
)
Arguments
delta

Null treatment effect to test.

Returns

A two-sided p-value.


InferenceOrdinalPartialProportionalOddsRegr$compute_wald_confidence_interval()

Identical to $compute_asymp_confidence_interval(); provided as an explicit alias for callers that want to name the Wald method directly rather than via the generic "asymptotic" dispatch.

Usage
InferenceOrdinalPartialProportionalOddsRegr$compute_wald_confidence_interval(
  alpha = 0.05
)
Arguments
alpha

Significance level for the interval.

Returns

A confidence interval for the treatment effect.


InferenceOrdinalPartialProportionalOddsRegr$compute_wald_two_sided_pval()

Identical to $compute_asymp_two_sided_pval(); provided as an explicit alias for callers that want to name the Wald method directly rather than via the generic "asymptotic" dispatch.

Usage
InferenceOrdinalPartialProportionalOddsRegr$compute_wald_two_sided_pval(
  delta = 0
)
Arguments
delta

Null treatment effect to test.

Returns

A two-sided p-value.


InferenceOrdinalPartialProportionalOddsRegr$benchmark_asymp_two_sided_pval_breakdown()

Diagnostic helper for performance investigation: runs the same computation as $compute_asymp_two_sided_pval() (fit the partial proportional-odds model requiring a standard error, cache the estimate/SE/df, compute the two-sided Wald p-value) but separately times each of the three stages — model fit, cache materialization, and final p-value arithmetic — via proc.time(). If the fit fails or has no usable standard error, returns immediately with only fit_time populated and every other timing/result field NA.

Usage
InferenceOrdinalPartialProportionalOddsRegr$benchmark_asymp_two_sided_pval_breakdown(
  delta = 0
)
Arguments
delta

Null treatment effect to test.

Returns

A named list: fit_time, cache_time, pval_math_time, total_time (all in seconds), pval, beta_hat_T, and s_beta_hat_T.


InferenceOrdinalPartialProportionalOddsRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalPartialProportionalOddsRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Peterson, B., and Harrell, F. E. (1990). "Partial Proportional Odds Models for Ordinal Response Variables." Journal of the Royal Statistical Society, Series C (Applied Statistics), 39(2), 205-217, doi:10.2307/2347760, for the partial (non-parallel-covariate) proportional-odds model fit here. McCullagh, P. (1980). "Regression Models for Ordinal Data." Journal of the Royal Statistical Society, Series B, 42(2), 109-142, doi:10.1111/j.2517-6161.1980.tb01109.x, for the full proportional-odds model this generalizes (see InferenceOrdinalPropOddsRegr).


Proportional Odds Regression Inference for Ordinal Responses

Description

Fits a proportional-odds (cumulative-logit) regression, via fast_ordinal_regression_with_var_cpp (see that page for the full model), for ordinal responses using the treatment indicator and, optionally, all recorded covariates as predictors. This is a full-likelihood class (likelihood_tier = "full") supporting score, gradient, and likelihood-ratio tests, plus parametric likelihood-ratio bootstrap calibration, in addition to Wald and resampling-based inference.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceOrdinalPropOddsRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalPropOddsRegr$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceOrdinalPropOddsRegr$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalPropOddsRegr$supports_rand_pval_for_incidence()

Usage
InferenceOrdinalPropOddsRegr$supports_rand_pval_for_incidence()

InferenceOrdinalPropOddsRegr$compute_rand_two_sided_pval()

Usage
InferenceOrdinalPropOddsRegr$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalPropOddsRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalPropOddsRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

McCullagh, P. (1980). "Regression Models for Ordinal Data." Journal of the Royal Statistical Society, Series B, 42(2), 109-142, doi:10.1111/j.2517-6161.1980.tb01109.x, for the proportional-odds cumulative-logit model fit here.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalPropOddsRegr$new(seq_des)
inf$compute_estimate()


inf$set_seed(1)
inf$compute_lik_ratio_bootstrap_two_sided_pval(delta = 0, B = 9, show_progress = FALSE)


Ridit Analysis for Ordinal Responses

Description

Performs Ridit analysis (Relative to an Identified Distribution unit) for comparing two groups on an ordinal scale. Every subject's category k is converted to a ridit score — its relative rank position within the reference distribution's empirical CDF: r_k = F_{\mathrm{ref}}(k{-}1) + \tfrac12 f_{\mathrm{ref}}(k), where F_{\mathrm{ref}} and f_{\mathrm{ref}} are the reference group's empirical cumulative and point probabilities. The reference distribution — controlled by the reference constructor argument — may be the control arm (default), the treatment arm, or the pooled sample. The treatment effect is the mean ridit score among treated subjects minus 0.5 (the value it would take under the null of no group difference, since a group's own ridit scores against itself as reference always average to 0.5); this mean ridit score also has a direct interpretation as (an estimate of) the probability that a randomly selected treated subject's outcome exceeds a randomly selected reference-distribution subject's outcome (a Mann-Whitney-type stochastic superiority probability), similar in spirit to InferenceOrdinalJonckheereTerpstraTest's superiority measure but referenced against a chosen distribution rather than always symmetric between the two arms. Standard errors and p-values come from fast_ridit_analysis_cpp's asymptotic formula, not a resampling approximation.

Super class

Inference -> InferenceOrdinalRidit

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalRidit$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize a Ridit analysis inference object for a completed design with an ordinal, uncensored response.

Usage
InferenceOrdinalRidit$new(
  des_obj,
  model_formula = NULL,
  reference = "control",
  verbose = FALSE,
  max_resample_attempts = 50L
)
Arguments
des_obj

A DesignSeqOneByOne object whose entire n subjects are assigned and response y is recorded within.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

reference

The group to use as the "Identified Distribution" (reference). Must be one of "control", "treatment", or "pooled". Default is "control".

verbose

A flag indicating whether messages should be displayed.

max_resample_attempts

Maximum number of times a single bootstrap replicate may be redrawn when the drawn sample fails validity screening. If all attempts fail the replicate is recorded as NA, silently reducing the effective B. Must be a positive integer. Default 50L.


InferenceOrdinalRidit$compute_estimate()

Returns the estimated treatment effect: the mean ridit score among treated subjects minus 0.5 (see class documentation for the full ridit-score definition and reference-distribution choice).

Usage
InferenceOrdinalRidit$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.

Returns

The numeric estimate.


InferenceOrdinalRidit$compute_estimate_with_bootstrap_weights()

Recomputes the ridit treatment estimate under subject/block bootstrap weights: the reference distribution's category proportions and every subject's ridit score are recomputed using the weights, then the weighted mean ridit score among treated subjects (minus 0.5) is returned. Used by the Bayesian bootstrap and related weighted-resampling machinery. Always leaves the standard error and degrees of freedom unavailable (NA) regardless of estimate_only — this weighted path never computes the asymptotic variance.

Usage
InferenceOrdinalRidit$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Subject-, block-, cluster-, or matched-set bootstrap weights.

estimate_only

Present for interface parity; this method never computes variance components regardless of its value.


InferenceOrdinalRidit$get_mean_ridit_treatment()

Returns the mean ridit score among treated subjects (not centered — this is the raw mean, unlike $compute_estimate() which subtracts 0.5).

Usage
InferenceOrdinalRidit$get_mean_ridit_treatment()
Returns

The numeric Mean Ridit.


InferenceOrdinalRidit$get_ridit_scores()

Returns each subject's individual ridit score (see class documentation for the ridit-score formula), in subject order.

Usage
InferenceOrdinalRidit$get_ridit_scores()
Returns

A numeric vector of scores.


InferenceOrdinalRidit$compute_asymp_confidence_interval()

Computes the asymptotic confidence interval for the treatment effect (mean ridit - 0.5), using fast_ridit_analysis_cpp's asymptotic standard error.

Usage
InferenceOrdinalRidit$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level.

Returns

A numeric vector of length 2.


InferenceOrdinalRidit$compute_asymp_two_sided_pval()

Computes a two-sided Wald p-value testing H_0: \text{mean ridit} - 0.5 = \code{delta} (i.e. delta = 0 tests the null of no group difference, mean ridit = 0.5), using fast_ridit_analysis_cpp's asymptotic standard error.

Usage
InferenceOrdinalRidit$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null value (centered at 0, so delta=0 means Ridit=0.5).

Returns

The p-value.


InferenceOrdinalRidit$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalRidit$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Bross, I. D. J. (1958). "How to Use Ridit Analysis." Biometrics, 14(1), 18-38, doi:10.2307/2527727, for the ridit transformation and its interpretation used here.

Examples

set.seed(1)
x_dat <- data.frame(
  x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
  x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneBernoulli$new(n = nrow(x_dat), response_type = "ordinal",
  verbose = FALSE)
for (i in seq_len(nrow(x_dat))) {
  seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(as.integer(c(1, 2, 2, 3, 3, 4, 4, 5)))
infer <- InferenceOrdinalRidit$
  new(seq_des, verbose = FALSE)
infer


Stereotype Logit Regression Inference for Ordinal Responses

Description

Fits Anderson's (1984) stereotype logit model for ordinal responses (see fast_stereotype_logit_cpp for the full reduced-rank multinomial-softmax formula and reparameterization): a single linear predictor \eta_i = \beta_T W_i + X_i^\top \gamma is scaled by a category-specific score \phi_k \in [0,1] (jointly estimated, monotone in k) in a softmax over all K categories, rather than assuming a single proportional/parallel effect across cuts as InferenceOrdinalContRatioRegr/ InferenceOrdinalKKCondAdjCatLogitRegr do. This makes the stereotype model a genuinely more flexible (multinomial-logit-like, reduced-rank) alternative to the standard proportional-odds/adjacent-category/continuation-ratio ordinal families, at the cost of a less directly interpretable treatment coefficient (\beta_T enters multiplicatively through the \phi_k scores rather than as a single additive log-odds-ratio). likelihood_tier = "full": likelihood-ratio, score, gradient, and Wald tests are all available when the model converges, plus parametric-likelihood-bootstrap calibration of the likelihood-ratio test.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

compute_lik_ratio_two_sided_pval() can be anti-conservative at small n. The stereotype model's category-score parameters \phi_k are a Davies (1977) non-regular case, not identified when \beta_T is at/near zero – exactly the neighborhood every null hypothesis test sits in – so the LR statistic's null distribution need not be the chi-square(1) this method assumes. The size of the problem depends on the constrained refit: before 2026-09-21 the delta-constrained refit started cold and could stall in a much worse local optimum (e.g. neg-log-likelihood 131.3 vs the full fit's 118.9 at \delta equal to the MLE itself), which inflated the LR statistic and produced ~18-23% Type-I error at n = 100; it also collapsed the bootstrap and inverted-LR confidence intervals to zero width. The refit now also starts from the full-fit parameters and keeps the better fit. Measured after that change (400 simulated null datasets, 5 response categories, a continuous covariate): about 6.5% Type-I error and 0.93 confidence-interval coverage at n = 100, with no zero-width intervals, and about 13% Type-I error at n = 50 with 3 categories (169 converged fits). Part of the small-sample excess comes from the likelihood being multimodal: in a few percent of n = 50 fits the reported point estimate sits at a local optimum whose likelihood is below that of another optimum, which makes the LR statistic at the estimate positive rather than zero. For small samples, prefer the better-calibrated alternatives on this class: compute_lik_ratio_bootstrap_two_sided_pval() (parametric-bootstrap calibration via the class's own null simulation; ~6-7% Type-I error) and the cheaper compute_lik_ratio_bartlett_two_sided_pval() (Monte-Carlo Bartlett correction; ~4% Type-I error); those two figures predate the refit change and were not re-measured. Both are roughly 20-40x slower per call than the raw chi-square test (a fresh bootstrap/Monte-Carlo refit at every delta candidate), which is why they are excluded from this package's own routine comprehensive test suite (see comprehensive_slow_paths.R). compute_lik_ratio_two_sided_pval() itself is left unchanged (not silently recalibrated) to avoid an undocumented behavior change to an existing method's contract. Bayesian-bootstrap inference is temporarily unavailable because the current non-uniform weighted hook fits a cumulative-logit surrogate rather than the stereotype likelihood. It will remain disabled until the native weighted stereotype-logit backend described in the package implementation plan lands.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceOrdinalStereotypeLogitRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceOrdinalStereotypeLogitRegr$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceOrdinalStereotypeLogitRegr$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalStereotypeLogitRegr$supports_rand_pval_for_incidence()

Usage
InferenceOrdinalStereotypeLogitRegr$supports_rand_pval_for_incidence()

InferenceOrdinalStereotypeLogitRegr$compute_rand_two_sided_pval()

Usage
InferenceOrdinalStereotypeLogitRegr$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceOrdinalStereotypeLogitRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceOrdinalStereotypeLogitRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Anderson, J. A. (1984). "Regression and Ordered Categorical Variables." Journal of the Royal Statistical Society, Series B, 46(1), 1-30, doi:10.1111/j.2517-6161.1984.tb01276.x, for the stereotype logit model.

See Also

InferenceOrdinalContRatioRegr for a proportional (non-reduced-rank) ordinal alternative. See also: Ordinal regression (Wikipedia).


Parametric-Bootstrap-Capable Likelihood Inference

Description

Intermediate abstract base for the subset of likelihood-backed inference families that are plausible targets for parametric null-bootstrap likelihood-ratio calibration.

This class sits between InferenceAsympLik and InferenceAsympLikStdModCache in the hierarchy. Families with highly bespoke partial-likelihood, quadrature, frailty, copula, or custom combined-likelihood geometry remain direct children of InferenceAsympLik and do not pass through here.

The only operational user-facing parametric-bootstrap LR methods on this surface are compute_lik_ratio_bootstrap_two_sided_pval(...) and compute_lik_ratio_bootstrap_confidence_interval(...). Diagnostic accessors such as get_last_param_bootstrap_diagnostics() are supplementary and not alternative execution entry points.

Parametric-bootstrap LR calibration is available only for concrete classes that inherit from InferenceParamBootstrap and whose private method supports_lik_ratio_param_bootstrap() returns TRUE. Families that are intentionally unsupported are kept off this branch entirely.

Super classes

Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap -> InferenceJackknife -> InferenceAsymp -> InferenceMLEorKMSummaryTable -> InferenceAsympLik -> InferenceParamBootstrap

Methods

Public methods

+ inherited public methods from InferenceAsympLik
+ inherited public methods from InferenceMLEorKMSummaryTable
+ inherited public methods from InferenceAsymp
+ inherited public methods from InferenceJackknife
+ inherited public methods from InferenceBayesianBootstrap
+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference

InferenceParamBootstrap$get_last_param_bootstrap_diagnostics()

Returns diagnostics from the most recent parametric-bootstrap LR run.

Usage
InferenceParamBootstrap$get_last_param_bootstrap_diagnostics()
Returns

A list of diagnostics, or NULL if no parametric-bootstrap LR run has been executed.


InferenceParamBootstrap$compute_lik_ratio_bootstrap_two_sided_pval()

Bootstrap-calibrated likelihood-ratio two-sided p-value.

Fits the null model at delta, simulates B datasets from that fitted null, refits unrestricted and null models on each, and returns the empirical tail probability of the observed LR statistic. This is the primary user-facing entry point for bootstrap-calibrated likelihood-ratio p-values.

This method is available only for classes whose private supports_lik_ratio_param_bootstrap() method returns TRUE. For unsupported classes it errors immediately rather than silently falling back to another procedure.

The standard user-facing arguments are delta = 0, B = 199, and show_progress = FALSE. The remaining arguments control replicate-quality thresholds and retry behavior.

Runtime cost is roughly one unrestricted fit plus B simulated unrestricted/null refit pairs, so this is typically much more expensive than the asymptotic LR p-value.

Usage
InferenceParamBootstrap$compute_lik_ratio_bootstrap_two_sided_pval(
  delta = 0,
  B = 199,
  show_progress = FALSE,
  min_number_usable_samples = 5L,
  max_attempts_per_replicate = 2L
)
Arguments
delta

Null treatment effect. Default 0.

B

Number of bootstrap replicates. Default 199.

show_progress

Logical; show a progress bar. Default FALSE.

min_number_usable_samples

Minimum number of usable bootstrap replicates required to return a finite p-value. Default 5L.

max_attempts_per_replicate

Maximum number of simulation/refit retries per bootstrap replicate. Default 2L.

Returns

A scalar p-value, or NA_real_ if the computation fails.


InferenceParamBootstrap$compute_lik_ratio_bootstrap_confidence_interval()

Bootstrap-calibrated likelihood-ratio confidence interval.

Inverts compute_lik_ratio_bootstrap_two_sided_pval via a bracket-and-bisect search seeded with the Wald interval. Each p-value evaluation costs B bootstrap refits, so this method is substantially more expensive than the p-value alone. This is the primary user-facing entry point for bootstrap-calibrated likelihood-ratio confidence intervals.

This method is available only for classes whose private supports_lik_ratio_param_bootstrap_confidence_interval() method returns TRUE.

The standard user-facing arguments are B = 199 and show_progress = FALSE. Runtime cost is high because each confidence-interval bound requires repeated bootstrap p-value evaluations.

Usage
InferenceParamBootstrap$compute_lik_ratio_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 199,
  show_progress = FALSE,
  min_number_usable_samples = 5L,
  max_attempts_per_replicate = 2L,
  root_tolerance = NULL,
  max_root_iterations = 8L
)
Arguments
alpha

Significance level. Default 0.05.

B

Bootstrap replicates per p-value evaluation. Default 199.

show_progress

Logical; show a progress bar. Default FALSE.

min_number_usable_samples

Minimum number of usable bootstrap replicates required within each p-value evaluation. Default 5L.

max_attempts_per_replicate

Maximum number of simulation/refit retries per bootstrap replicate. Default 2L.

root_tolerance

Effect-scale tolerance for the inversion root. If NULL, a tolerance proportional to the Wald standard error is used. The bootstrap p-value being inverted is Monte Carlo-discrete, so solving to machine precision is not meaningful.

max_root_iterations

Maximum number of bisection iterations per bound during interval inversion. Use 0L to return the first finite outer bracket. Default 8L.

Returns

Named two-element numeric vector with the confidence-interval bounds.


InferenceParamBootstrap$get_last_param_bootstrap_estimate_diagnostics()

Returns diagnostics from the most recent compute_param_bootstrap_estimate() run.

Usage
InferenceParamBootstrap$get_last_param_bootstrap_estimate_diagnostics()
Returns

A list of diagnostics, or NULL if no parametric-bootstrap estimate bias correction has been run.


InferenceParamBootstrap$compute_param_bootstrap_estimate()

Parametric-bootstrap bias-corrected point estimate.

Simulates B datasets from the model at the unrestricted fit (not a null-restricted fit, unlike the LR-bootstrap methods on this class), refits each unrestricted, and returns 2 * theta_hat - mean(theta_hat_star) – the standard single-level parametric-bootstrap bias correction (Efron & Tibshirani), the Monte-Carlo analog of an analytic Cox-Snell (1968) first-order bias correction. Reuses the same simulate_under_lik_null() contract every supports_lik_ratio_param_bootstrap() == TRUE family already implements, anchoring the simulation at the unrestricted fit instead of a null-restricted one.

This method is available only for classes whose private supports_param_bootstrap_estimate() method returns TRUE (by default, this delegates to supports_lik_ratio_param_bootstrap()).

Usage
InferenceParamBootstrap$compute_param_bootstrap_estimate(
  B = 199,
  show_progress = FALSE,
  min_number_usable_samples = 5L,
  max_attempts_per_replicate = 2L
)
Arguments
B

Number of bootstrap replicates. Default 199.

show_progress

Logical; show a progress bar. Default FALSE.

min_number_usable_samples

Minimum number of usable bootstrap replicates required to return a finite estimate. Default 5L.

max_attempts_per_replicate

Maximum number of simulation/refit retries per bootstrap replicate. Default 2L.

Returns

A scalar bias-corrected estimate, or NA_real_ if the computation fails.


InferenceParamBootstrap$compute_param_bootstrap_confidence_interval()

Parametric-bootstrap "basic" (reflected) confidence interval.

Uses the same kind of simulate-at-full_fit replicates as compute_param_bootstrap_estimate() – run once, not re-simulated per candidate null – and reflects their empirical quantiles around the raw estimate: CI = [2*theta_hat - q_(1-alpha/2)(theta_star), 2*theta_hat - q_(alpha/2)(theta_star)]. This is the standard "basic" bootstrap interval (Efron & Tibshirani): the bias-aware companion to compute_param_bootstrap_estimate() – unlike a naive percentile interval, it reflects around the same point the estimate's own bias correction targets.

Unlike compute_lik_ratio_bootstrap_confidence_interval(), which must re-simulate B fresh replicates at every candidate null value visited during bisection (because that method's replicates are anchored at a null-restricted fit that changes with the candidate value), this method's replicates are anchored at the single unrestricted fit and therefore do not need to be regenerated per alpha or inverted against any particular null – one batch of B replicates supports the whole interval.

This method is available only for classes whose private supports_param_bootstrap_estimate() method returns TRUE.

Usage
InferenceParamBootstrap$compute_param_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 199,
  show_progress = FALSE,
  min_number_usable_samples = 5L,
  max_attempts_per_replicate = 2L
)
Arguments
alpha

Significance level. Default 0.05.

B

Number of bootstrap replicates. Default 199.

show_progress

Logical; show a progress bar. Default FALSE.

min_number_usable_samples

Minimum number of usable bootstrap replicates required to return a finite interval. Default 5L.

max_attempts_per_replicate

Maximum number of simulation/refit retries per bootstrap replicate. Default 2L.

Returns

Named two-element numeric vector with the confidence-interval bounds.


InferenceParamBootstrap$compute_param_bootstrap_pval()

Parametric-bootstrap two-sided p-value for H0: theta = delta, obtained by inverting the same "basic" reflected-quantile construction as compute_param_bootstrap_confidence_interval() rather than by simulating fresh replicates under each candidate delta (contrast compute_lik_ratio_bootstrap_two_sided_pval(), which does resimulate per delta because its replicates are anchored at a null-restricted fit that changes with delta).

Reflects delta through the raw estimate, t = 2*theta_hat - delta, and reports twice the smaller empirical tail of the replicate distribution beyond t (with the same +1 continuity correction used by compute_lik_ratio_bootstrap_two_sided_pval()). Because the replicate batch does not depend on delta, this is a direct lookup against one batch of B replicates – valid for any delta with no additional simulation.

As with the confidence interval, this test implicitly assumes the shape of theta_hat's sampling distribution near delta resembles its shape at the unrestricted fit – an approximation that degrades the further delta is from the observed estimate, and is why this p-value should be treated as a cheap default rather than a replacement for compute_lik_ratio_bootstrap_two_sided_pval() when accuracy near a specific null matters.

Usage
InferenceParamBootstrap$compute_param_bootstrap_pval(
  delta,
  B = 199,
  show_progress = FALSE,
  min_number_usable_samples = 5L,
  max_attempts_per_replicate = 2L
)
Arguments
delta

Null value of the coefficient of interest.

B

Number of bootstrap replicates. Default 199.

show_progress

Logical; show a progress bar. Default FALSE.

min_number_usable_samples

Minimum number of usable bootstrap replicates required to return a finite p-value. Default 5L.

max_attempts_per_replicate

Maximum number of simulation/refit retries per bootstrap replicate. Default 2L.

Returns

A scalar p-value, or NA_real_ if the computation fails. Shared replicate-running core for anything built on B simulated null datasets via simulate_under_lik_null() (currently: compute_lik_ratio_bootstrap_two_sided_pval() and, via InferenceExtBartlettApprox, get_bartlett_factor_approx()). Handles seeding, multi-core parallelism, the reusable-worker-state optimization, and deterministic-mode thread budgeting uniformly, so every caller gets the same performance characteristics for free. Returns a list with $results (one per-replicate result object per B, each carrying at least $lr), $used_worker_path, and $used_deterministic_mode.

Callers are responsible for setting private$active_resampling_operation themselves (not done here, to avoid a nested caller clobbering an already-active outer flag via a premature on.exit reset). Simulate a bootstrap dataset under the fitted null likelihood and return a minimal spec list for refitting.

Must be overridden by families that set supports_lik_ratio_param_bootstrap() to TRUE. The returned list must contain at least: - full_fit: unrestricted fit on the simulated data - fit_null: function(delta, start) returning a constrained fit - neg_loglik: function(fit) returning the neg-log-likelihood


InferenceParamBootstrap$clone()

The objects of this class are cloneable with this method.

Usage
InferenceParamBootstrap$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Beta Regression Inference for Proportion Responses

Description

Fits Ferrari and Cribari-Neto's (2004) beta regression for proportion responses Y_i \in (0, 1): \mathrm{logit}(E[Y_i \mid w_i, x_i]) = \beta_0 + \beta_T w_i + x_i^\top \gamma, Y_i \mid w_i, x_i \sim \mathrm{Beta}(\mu_i \phi, (1-\mu_i)\phi) for fitted mean \mu_i and a single (constant, not covariate-dependent) precision parameter \phi, by maximum likelihood (fast_beta_regression_cpp/ fast_beta_regression_weighted_cpp). \hat\beta_T is a log-odds-ratio on the conditional-mean scale: \exp(\hat\beta_T) is the odds ratio for the expected proportion. Unlike InferencePropFractionalLogit's quasi-likelihood (which specifies only the conditional mean), beta regression also specifies the conditional variance/shape via \phi — a correctly specified beta model yields a fully efficient likelihood-based fit and genuine likelihood-ratio/score/gradient tests, at the cost of requiring the beta-distribution shape assumption to actually hold. likelihood_tier = "full": likelihood-ratio, score, gradient, and Wald tests are all available when the model converges, plus parametric-likelihood-bootstrap calibration of the likelihood-ratio test. Y_i values of exactly 0 or 1 are not supported by the beta density and are handled by sanitize_beta_response()'s boundary adjustment before fitting.

Estimand. Composes MarginalEstimand (set_estimand()/get_estimand()/get_supported_estimands()). Under the default estimand = "conditional", \hat\beta_T is the log-odds-ratio above. Under estimand = "marginal_mean_diff", the reported quantity is instead the g-computation marginal mean difference \frac{1}{n}\sum_i \{\mathrm{plogis}(\hat\beta_0 + \hat\beta_T + X_i^\top \hat\gamma) - \mathrm{plogis}(\hat\beta_0 + X_i^\top \hat\gamma)\} (the precision parameter \phi does not enter the mean, so it plays no role in this functional). Only "marginal_mean_diff" is supported — a ratio of two mean proportions, both bounded in [0,1], is not the standard estimand for a beta-regression treatment effect the way a rate ratio is for count data. Because there is no latent submodel for this family (unlike e.g. InferencePropZeroOneInflatedBetaRegr's zero/one-inflation mixture), the marginal mean function is exactly the model's own fitted mean; no separate standardization step beyond the g-computation average is needed. Standard errors under the marginal estimand use the delta method against the mean-submodel coefficient covariance (degrees of freedom Inf); testing_type is restricted to "wald" whenever the estimand is non-conditional. The underlying model fit is identical regardless of estimand — switching estimand is a pure post-fit transform, never a refit.

Super class

Inference -> InferencePropBetaRegr

Methods

Public methods

+ inherited public methods from Inference

InferencePropBetaRegr$new()

Initialize inference for the beta regression model \mathrm{logit}(E[Y_i \mid w_i, x_i]) = \beta_0 + \beta_T w_i + x_i^\top \gamma, Y_i \sim \mathrm{Beta}(\mu_i \phi, (1-\mu_i) \phi); see InferencePropBetaRegr for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferencePropBetaRegr$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL,
  optimization_alg = NULL
)
Arguments
des_obj

A completed Design object with a proportion response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values by default.

optimization_alg

Character scalar specifying the optimization algorithm. Default is dispatched via policy.


InferencePropBetaRegr$compute_estimate()

Fits the beta regression model by maximum likelihood (jointly estimating the mean coefficients and the precision parameter \phi). Under the default estimand = "conditional", returns the log-odds-ratio estimate \hat\beta_T on the conditional-mean scale. Under estimand = "marginal_mean_diff" (set via set_estimand()), returns the g-computation marginal mean difference instead — see the class-level @details for the formula. The underlying model fit is identical either way (a pure post-fit transform of the same cached fit, no refit).

Usage
InferencePropBetaRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.


InferencePropBetaRegr$compute_asymp_confidence_interval()

Wald confidence interval, dispatched by testing_type for the conditional estimand (score/gradient/ likelihood-ratio/Bartlett available; see InferenceAsympLik); under a marginal estimand testing_type is always "wald" (the only value set_estimand() permits there), so this always resolves to the delta-method interval. Calls self$compute_estimate() first (not private$shared() directly) so the estimand-aware cache is always current regardless of call order.

Usage
InferencePropBetaRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Two-sided miscoverage rate; the returned interval targets 1 - alpha coverage.


InferencePropBetaRegr$compute_asymp_two_sided_pval()

Wald two-sided p-value, dispatched by testing_type exactly as compute_asymp_confidence_interval(); see that method's description for the marginal-estimand always-Wald note.

Usage
InferencePropBetaRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment-effect value under the current estimand (conditional log-odds-ratio, or marginal mean difference).


InferencePropBetaRegr$compute_estimate_with_bootstrap_weights()

Refits the beta model with subject/block-level weights applied to the fitting log-likelihood (Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights()) via fast_beta_regression_weighted_cpp, and returns the reweighted estimate \hat\beta_T^{(w)}. Uses the same QR column-dropping hardening as compute_estimate(); a hardened-but-still-unreasonable fit is cached as nonestimable.

Usage
InferencePropBetaRegr$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip variance calculations.


InferencePropBetaRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferencePropBetaRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Ferrari, S., and Cribari-Neto, F. (2004). "Beta regression for modelling rates and proportions." Journal of Applied Statistics, 31(7), 799-815, doi:10.1080/0266476042000214501.

See Also

InferencePropFractionalLogit for a quasi-likelihood proportion model that specifies only the conditional mean. Comparable Python API: no direct beta-regression equivalent in statsmodels; see statsmodels GLM for the general exponential-family GLM framework. See also: Beta distribution (Wikipedia).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropBetaRegr$new(seq_des)
inf$compute_estimate()


Fractional Logit Inference for Proportion Responses

Description

Fits Papke and Wooldridge's (1996) fractional logistic (quasi-binomial) regression for proportion responses Y_i \in [0, 1] (not restricted to \{0, 1\}): E[Y_i \mid w_i, x_i] = \mathrm{logit}^{-1}(\beta_0 + \beta_T w_i + x_i^\top \gamma), fit by maximizing the Bernoulli quasi-log-likelihood \sum_i \{Y_i \log \mu_i + (1 - Y_i) \log(1 - \mu_i)\} treating Y_i as if it were binary (a valid estimating equation for the conditional mean even though Y_i is fractional — the Bernoulli log-likelihood's score is unbiased for the true mean regardless of the actual distribution of Y_i on [0,1]). \hat\beta_T is a log-odds-ratio on the conditional-mean scale: \exp(\hat\beta_T) is the odds ratio for the expected proportion. Standard errors use the model-based (non-robust/non-sandwich) Fisher information from this quasi-likelihood, scaled by an estimated quasi-binomial dispersion parameter \hat\phi (Papke & Wooldridge's own prescription; fixed 2026-09-06 – the unscaled Bernoulli-based variance systematically overstates \mathrm{Var}(\hat\beta_T) for a genuinely fractional response, since Bernoulli is the maximum-variance distribution on [0,1] for a given mean); only Wald inference is exposed (private$supports_likelihood_tests() is hard FALSE here even though likelihood_tier = "full" metadata is set for component-composition purposes — this class deliberately does not compose ParametricLikelihoodBootstrap, so no likelihood-ratio/score/gradient test surface is exposed). Validity requires that the conditional mean is correctly specified on the logit scale; unlike beta regression, no assumption is made about the conditional variance or shape of Y_i's distribution.

Super class

Inference -> InferencePropFractionalLogit

Methods

Public methods

+ inherited public methods from Inference

InferencePropFractionalLogit$new()

Initialize inference for the fractional logit model E[Y_i \mid w_i, x_i] = \mathrm{logit}^{-1}(\beta_0 + \beta_T w_i + x_i^\top \gamma); see InferencePropFractionalLogit for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferencePropFractionalLogit$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  harden = TRUE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with a proportion response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

harden

Whether to apply robustness measures.

smart_cold_start_default

Whether to use smart cold start values.


InferencePropFractionalLogit$compute_estimate()

Fits the fractional logit model by maximizing the Bernoulli quasi-log-likelihood on the fractional response and returns the log-odds-ratio estimate \hat\beta_T. When estimate_only = TRUE and hardening is disabled (harden = FALSE), uses a fast path via base R's glm.fit(family = quasibinomial()) instead of the package's own fitting routine; otherwise dispatches through the shared hardened-fit path.

Usage
InferencePropFractionalLogit$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations; when combined with harden = FALSE, also switches to the quasibinomial() fast path.


InferencePropFractionalLogit$compute_estimate_with_bootstrap_weights()

Refits the fractional logit model with subject/block-level weights applied to the fitting quasi-log-likelihood (Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights()), and returns the reweighted estimate \hat\beta_T^{(w)}. Uses the same QR column-dropping hardening as compute_estimate()'s hardened path; a hardened-but-still-unreasonable fit is cached as nonestimable.

Usage
InferencePropFractionalLogit$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip variance calculations.


InferencePropFractionalLogit$clone()

The objects of this class are cloneable with this method.

Usage
InferencePropFractionalLogit$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Papke, L. E., and Wooldridge, J. M. (1996). "Econometric Methods for Fractional Response Variables with an Application to 401(K) Plan Participation Rates." Journal of Applied Econometrics, 11(6), 619-632, doi:10.1002/(SICI)1099-1255(199611)11:6<619::AID-JAE418>3.0.CO;2-1.

See Also

InferencePropBetaRegr for a proportion model that also specifies the conditional variance/shape. Comparable Python API: statsmodels GLM (family=Binomial() on fractional response data). See also: Logistic regression (Wikipedia).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropFractionalLogit$new(seq_des)
inf$compute_estimate()


G-Computation Mean-Difference Inference for Proportion Responses

Description

Fits a fractional-logit working model, \mathrm{logit}\,E[Y_i \mid x_i] = x_i^\top\hat\beta (via fast_logistic_regression_cpp, treating the continuous-in-(0,1) proportion as a quasi-binomial mean model — the same mean-model idea as InferencePropFractionalLogit), for a proportion outcome using treatment and, optionally, all recorded covariates, then estimates the marginal mean difference by G-computation: standardizing predicted mean proportions under all-treated and all-control assignments over the empirical covariate distribution (see gcomp_fractional_logit_point_estimate_cpp for the exact standardization formula). Inference uses a Huber-White sandwich-robust covariance for the regression coefficients and the delta method (analytic gradient of the standardized mean-difference functional with respect to \hat\beta, by default) to propagate that covariance onto the mean-difference scale, \widehat{\mathrm{Var}}(\widehat{\mathrm{md}}) = \nabla^\top \widehat{\mathrm{Var}}(\hat\beta) \nabla.

The implementation is optimized for resampling-based inference. It utilizes a fast C++ IRLS solver for the underlying fractional logit regression. During resampling draws, it bypasses the calculation of the sandwich covariance matrix and delta-method standard errors, providing a significant speedup when computing bootstrap or randomization distributions.

Variance fallback cascade. If the primary analytic-gradient/sandwich-covariance variance is non-finite (e.g. near-boundary fitted probabilities), up to eight progressively more conservative fallback strategies are tried in order (see the variance_fallback_methods constructor argument for the full list and their individual definitions): stabilized (PSD-projected) sandwich covariance, model-based (Fisher information) covariance, finite-difference gradients in place of the analytic delta-method gradient, and combinations with progressively stronger probability clipping. The first strategy in the ordered list that yields a finite, positive variance is used; an empty variance_fallback_methods vector always returns NA variance rather than erroring.

Super class

Inference -> InferencePropGCompMeanDiff

Methods

Public methods

+ inherited public methods from Inference

InferencePropGCompMeanDiff$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize the g-computation inference object.

Usage
InferencePropGCompMeanDiff$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  prob_clip_eps = 1e-06,
  prob_clip_strong_eps = 1e-04,
  max_resample_attempts = 50L,
  smart_cold_start_default = NULL,
  harden = TRUE,
  variance_fallback_methods = c("robust", "stabilized_robust", "model_based",
    "stabilized_robust_fd", "model_based_fd", "stabilized_robust_strong_clip",
    "model_based_strong_clip", "model_based_fd_strong_clip")
)
Arguments
des_obj

A completed DesignSeqOneByOne object with a proportion response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

Whether to print progress messages.

prob_clip_eps

Primary probability clamp applied to fitted values during model-based variance computation. Predicted probabilities are clipped to [prob_clip_eps, 1 - prob_clip_eps] before computing IWLS weights. Must be in [0, 0.5). Default 1e-6.

prob_clip_strong_eps

Stronger clamp used as a fallback when the primary variance strategy fails. Predicted probabilities and gradients are clipped to [prob_clip_strong_eps, 1 - prob_clip_strong_eps]. Must be in [0, 0.5) and should be \ge prob_clip_eps. Default 1e-4.

max_resample_attempts

Maximum number of times a single bootstrap replicate may be redrawn when the drawn sample fails validity screening (e.g. near-perfect separation, too few observations per arm, excessive boundary mass). If all attempts fail the replicate is recorded as NA, silently reducing the effective B. Can be overridden per-call in approximate_bootstrap_distribution_beta_hat_T. Must be a positive integer. Default 50L.

smart_cold_start_default

Whether to use smart cold start values.

harden

Whether to apply robustness measures.

variance_fallback_methods

Ordered character vector of variance strategies to attempt in sequence. Each name corresponds to a (gradient, covariance-matrix) pair; the first strategy that yields a finite, positive variance is used. Allowed values (in their default order) are:

"robust"

Analytic delta-method gradient with the sandwich (HC) covariance.

"stabilized_robust"

Same gradient; covariance projected to the nearest PSD matrix.

"model_based"

Same gradient; Fisher-information (IWLS) covariance. Fitted probabilities are first clipped to [prob_clip_eps, 1 - prob_clip_eps], then the binomial variance weight w_i = \hat\mu_i(1-\hat\mu_i) is further capped at 0.25. The cap is the global maximum of p(1-p), attained at p = 0.5; it prevents a near-boundary fitted probability that slips through the clip from inflating the information matrix, which is mathematically correct because no Bernoulli variance can exceed 0.25.

"stabilized_robust_fd"

Finite-difference (central-difference) gradient; stabilized sandwich covariance. The step size for coefficient j is h_j = \varepsilon^{1/3}(|\hat\beta_j| + 1), where \varepsilon = .Machine$double.eps. The cube-root of machine epsilon is the theoretically optimal step that balances truncation error (O(h^2) for central differences) against floating-point cancellation (O(\varepsilon / h)), giving a total error of O(\varepsilon^{2/3}).

"model_based_fd"

Finite-difference gradient (same step rule as "stabilized_robust_fd"); model-based covariance with the 0.25 weight cap.

"stabilized_robust_strong_clip"

Strong-clipped analytic gradient; stabilized sandwich covariance.

"model_based_strong_clip"

Strong-clipped analytic gradient; model-based covariance (strong-clipped) with the 0.25 weight cap.

"model_based_fd_strong_clip"

Strong-clipped finite-difference gradient (same step rule); model-based covariance (strong-clipped) with the 0.25 weight cap.

Pass a shorter vector or a single string to restrict which strategies are tried. An empty vector always returns NA variance.


InferencePropGCompMeanDiff$compute_estimate()

Computes the g-computation treatment-effect estimate.

Usage
InferencePropGCompMeanDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.


InferencePropGCompMeanDiff$get_standard_error()

Returns the standard error of the g-computation mean-difference estimate (NA if it is unavailable).

Usage
InferencePropGCompMeanDiff$get_standard_error()
Returns

A single numeric standard error, or NA_real_.


InferencePropGCompMeanDiff$compute_estimate_with_bootstrap_weights()

Recomputes the g-computation mean-difference estimate under the supplied subject- or block-level bootstrap weights and caches it; used by the Bayesian-bootstrap and resampling paths.

Usage
InferencePropGCompMeanDiff$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Numeric vector of bootstrap weights, one per subject (or per block when the design is blocked).

estimate_only

If TRUE, only the point estimate is required (no variance).

Returns

The weighted mean-difference estimate, or NA_real_ if the weighted fit is unusable.


InferencePropGCompMeanDiff$compute_asymp_confidence_interval()

Computes a 1 - alpha confidence interval.

Usage
InferencePropGCompMeanDiff$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha.


InferencePropGCompMeanDiff$compute_asymp_two_sided_pval()

Uses the shared asymptotic two-sided p-value contract; see InferenceAsymp.

Usage
InferencePropGCompMeanDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null mean difference. Defaults to 0.


InferencePropGCompMeanDiff$compute_wald_two_sided_pval()

Computes a Wald two-sided p-value for the treatment effect.

Usage
InferencePropGCompMeanDiff$compute_wald_two_sided_pval(delta = 0)
Arguments
delta

The null mean difference. Defaults to 0.


InferencePropGCompMeanDiff$compute_wald_confidence_interval()

Computes a Wald confidence interval for the treatment effect.

Usage
InferencePropGCompMeanDiff$compute_wald_confidence_interval(alpha = 0.05)
Arguments
alpha

The significance level. Default 0.05.


InferencePropGCompMeanDiff$compute_bootstrap_two_sided_pval()

Computes a bootstrap two-sided p-value for the treatment effect.

Usage
InferencePropGCompMeanDiff$compute_bootstrap_two_sided_pval(
  delta = 0,
  B = 501,
  type = "symmetric",
  na.rm = FALSE,
  boundary_tol = 0.02,
  max_boundary_mass = 0.95,
  sep_tol = 0.02,
  min_group_n = 5L,
  show_progress = TRUE,
  min_number_usable_samples = 5L
)
Arguments
delta

The null mean difference. Defaults to 0.

B

Number of bootstrap samples.

type

Bootstrap p-value type. See InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.

na.rm

Whether to remove non-finite bootstrap replicates.

boundary_tol

Resample screening threshold for boundary mass near 0/1.

max_boundary_mass

Reject a resample when at least this fraction is near the boundary.

sep_tol

Separation tolerance used to reject nearly perfectly separated resamples.

min_group_n

Minimum number of observations required in each treatment arm.

show_progress

Whether to show a progress bar.

min_number_usable_samples

Minimum number of finite bootstrap samples required.


InferencePropGCompMeanDiff$compute_bootstrap_confidence_interval()

Generic (non-screening-modified) bootstrap two-sided p-value, aliased directly from 'InferenceNonParamBootstrap' so 'compute_bootstrap_two_sided_pval()' can dispatch to it without relying on 'super$', which does not resolve under flat component composition.

Computes a bootstrap confidence interval.

Usage
InferencePropGCompMeanDiff$compute_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 501,
  type = NULL,
  na.rm = TRUE,
  show_progress = TRUE,
  boundary_tol = 0.02,
  max_boundary_mass = 0.95,
  sep_tol = 0.02,
  min_group_n = 5L,
  min_number_usable_samples = 5L
)
Arguments
alpha

The confidence level 1 - alpha.

B

Number of bootstrap samples.

type

Bootstrap CI type.

na.rm

Whether to remove non-finite bootstrap replicates.

show_progress

Whether to show bootstrap progress.

boundary_tol

Resample screening threshold for boundary mass near 0/1.

max_boundary_mass

Reject a resample when at least this fraction is near the boundary.

sep_tol

Separation tolerance used to reject nearly perfectly separated resamples.

min_group_n

Minimum number of observations required in each treatment arm.

min_number_usable_samples

Minimum number of finite bootstrap samples required.


InferencePropGCompMeanDiff$approximate_bootstrap_distribution_beta_hat_T()

Generic (non-screening-modified) bootstrap confidence interval; see 'compute_bootstrap_two_sided_pval_generic'.

Abbreviated bootstrap sampler that reuses a bootstrap worker.

Usage
InferencePropGCompMeanDiff$approximate_bootstrap_distribution_beta_hat_T(
  B = 501,
  show_progress = TRUE,
  max_resample_attempts = NULL,
  boundary_tol = 0.02,
  max_boundary_mass = 0.95,
  sep_tol = 0.02,
  min_group_n = 5L,
  debug = FALSE
)
Arguments
B

The number of bootstrap samples (default 501).

show_progress

Whether to show a progress bar.

max_resample_attempts

Maximum redraw attempts per bootstrap replicate before the replicate is recorded as NA. NULL (default) uses the value set at construction time.

boundary_tol

Resample screening threshold for boundary mass near 0/1.

max_boundary_mass

Reject a resample when at least this fraction is near the boundary.

sep_tol

Separation tolerance used to reject nearly perfectly separated resamples.

min_group_n

Minimum number of observations required in each treatment arm.

debug

If TRUE, return per-replicate diagnostics (values, errors, warnings) instead of just the bootstrap values.


InferencePropGCompMeanDiff$clone()

The objects of this class are cloneable with this method.

Usage
InferencePropGCompMeanDiff$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropGCompMeanDiff$new(seq_des)
inf$compute_estimate()


GEE Inference for KK Designs with Proportion Response

Description

Fits a Generalized Estimating Equations model with a binomial (quasi-likelihood, fractional-response) family and logit link, \mathrm{logit}\,E[Y_i \mid x_i] = x_i^\top\beta, for proportion (continuous values in ⁠(0, 1)⁠) responses under a KK matching-on-the-fly design — the same fractional-logit mean-model idea as InferencePropFractionalLogit, extended to jointly account for matched-pair and reservoir clustering via GEE. Each GEE cluster is either a matched pair (2 members) or a reservoir singleton (1 member), with an exchangeable working correlation structure — see $compute_estimate()'s method-level documentation for the full fitting contract (internal Rcpp solver vs. geepack fallback, hardening/retry behavior). Inference is quasi-likelihood/estimating-equation based (likelihood_tier = "quasi"): standard errors are GEE sandwich (robust) standard errors, not model-likelihood-based.

Super class

Inference -> InferencePropKKGEE

Methods

Public methods

+ inherited public methods from Inference

InferencePropKKGEE$new()

Initialize KK proportion-response GEE inference, validate the matched/reservoir design, and prepare the exchangeable-working-correlation fractional-logit GEE fitting machinery used by InferencePropKKGEE.

Usage
InferencePropKKGEE$new(
  des_obj,
  model_formula = NULL,
  use_rcpp = TRUE,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with a proportion response.

model_formula

Optional formula for covariate adjustment.

use_rcpp

Whether to use the internal Rcpp GEE solver (TRUE, default) with automatic fallback to geepack::geeglm on failure, or always use geepack::geeglm directly (FALSE, requires geepack to be installed).

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.


InferencePropKKGEE$clone()

The objects of this class are cloneable with this method.

Usage
InferencePropKKGEE$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Liang, K.-Y., and Zeger, S. L. (1986). "Longitudinal Data Analysis Using Generalized Linear Models." Biometrika, 73(1), 13-22, doi:10.1093/biomet/73.1.13, for the GEE estimating-equation framework and sandwich variance estimator used here.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropKKGEE$new(seq_des)
inf$compute_estimate()


KK GLMM Inference for Proportion Responses

Description

Fits a combined conditional-logit-plus-random-intercept-GLMM likelihood for proportion responses under a KK matching-on-the-fly design. Matched pairs with a discordant pair-difference are handled by a conditional (fixed pair-effect) logistic term, while concordant/reservoir subjects are handled by a random-intercept logistic mixed model with intercept b_g \sim N(0, \sigma_b^2) per matched-set/reservoir group g; both terms share the same treatment coefficient \beta_T, jointly maximized by the internal fast_clogit_plus_glmm_cpp routine. This combines the design-exact conditional-logit treatment of matched pairs (no nuisance pair-intercept to estimate) with a GLMM's ability to still contribute information from concordant pairs and reservoir subjects, which a pure conditional-logit-on-discordant-pairs-only approach would discard. \exp(\hat\beta_T) is the common treatment odds ratio. likelihood_tier = "full": likelihood-ratio, score, and Wald tests are all available when the model converges. See InferenceAbstractKKCondLogitGLMM for the shared model-fitting and caching contract used by this class's incidence-response siblings (InferenceIncidKKCondLogitGLMMIVWC, InferenceIncidKKCondLogitGLMMOneLik).

Super classes

Inference -> InferenceAbstractKKCondLogitGLMM -> InferencePropKKGLMM

Methods

Public methods

+ inherited public methods from InferenceAbstractKKCondLogitGLMM
+ inherited public methods from Inference

InferencePropKKGLMM$new()

Initialize inference for the combined conditional-logit (discordant matched pairs) plus random-intercept-GLMM (concordant pairs/reservoir) proportion model; see InferencePropKKGLMM for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferencePropKKGLMM$new(
  des_obj,
  model_formula = NULL,
  max_abs_reasonable_coef = 10000,
  max_abs_log_sigma = 8,
  verbose = FALSE,
  smart_cold_start_default = NULL,
  optimization_alg = NULL
)
Arguments
des_obj

A completed Design object with a proportion response.

model_formula

Optional formula for covariate adjustment.

max_abs_reasonable_coef

Cap for reasonable coefficient estimates.

max_abs_log_sigma

Cap for reasonable log random effect variance.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.

optimization_alg

Character. Optimization algorithm (default "lbfgs").


InferencePropKKGLMM$clone()

The objects of this class are cloneable with this method.

Usage
InferencePropKKGLMM$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kapelner, A. and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the KK matching-on-the-fly design this class is built for; Breslow, N. E., and Clayton, D. G. (1993). "Approximate Inference in Generalized Linear Mixed Models." Journal of the American Statistical Association, 88(421), 9-25, doi:10.2307/2290687, for the GLMM likelihood framework combined with the conditional-logit term here.


Quantile Regression Compound Estimator for KK Matching-on-the-Fly Designs (Proportion Outcomes)

Description

A variance-weighted compound quantile regression estimator for KK matching-on-the-fly designs with proportion responses. Inference is performed on the logit (log-odds) scale: responses y \in (0,1) are transformed via \text{logit}(y) = \log(y/(1-y)) before quantile regression.

The estimator combines:

  1. Quantile regression on logit-scale within-pair differences \text{logit}(y_T) - \text{logit}(y_C) (matched pairs)

  2. Quantile regression of \text{logit}(y) on treatment and covariates (reservoir)

using the same variance-weighted combination logic as the OLS compound estimator.

The estimated treatment effect is a log-odds-ratio shift at quantile tau. At beta_T = 1 (one log-odds-ratio unit of treatment effect), the population treatment effect on the logit scale is exactly 1, so no skip_ci is needed.

Default quantile: tau = 0.5 (median regression). To target a different quantile — for example the 25th or 75th percentile — pass tau = 0.25 or tau = 0.75 to the constructor:

  inf = InferencePropKKQuantileRegrIVWC$
  new(seq_des, tau = 0.75)

Any value strictly between 0 and 1 is accepted.

Standard errors use Powell's "nid" sandwich estimator (non-iid), falling back to "iid" on failure. Asymptotic z-based inference is used throughout.

This class requires the quantreg package, which is listed in Suggests and is not installed automatically with EDI. Install quantreg before using this class.

Legacy class. Not fully tested in comprehensive_tests.R.

Super class

Inference -> InferencePropKKQuantileRegrIVWC

Methods

Public methods

+ inherited public methods from Inference

InferencePropKKQuantileRegrIVWC$new()

Initialize proportion-response KK IVWC quantile-regression inference on the logit response scale; see InferencePropKKQuantileRegrIVWC.

Usage
InferencePropKKQuantileRegrIVWC$new(
  des_obj,
  model_formula = NULL,
  tau = 0.5,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A DesignSeqOneByOne object whose entire n subjects are assigned and response y is recorded within.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

tau

The quantile level for regression on the logit scale, strictly between 0 and 1. The default tau = 0.5 estimates the median log-odds-ratio treatment effect. Pass a different value (e.g. tau = 0.25 or tau = 0.75) to target a different percentile of the treatment effect distribution.

verbose

A flag indicating whether messages should be displayed to the user. Default is FALSE.

smart_cold_start_default

Whether to use smart cold start values.


InferencePropKKQuantileRegrIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferencePropKKQuantileRegrIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropKKQuantileRegrIVWC$new(seq_des)
inf$compute_estimate()


Quantile Regression Combined-Likelihood Compound Estimator for KK Designs (Proportion)

Description

Fits the combined stacked quantile regression (matched-pair differences + reservoir) using the treatment indicator and all recorded covariates for proportion responses. Responses y \in (0,1) are transformed via \mathrm{logit}(y) = \log(y/(1-y)) before regression; the estimated treatment effect \hat\beta_T is a log-odds-ratio shift at quantile tau of the logit-transformed response. Minimizes the joint check-function (pinball) loss \rho_\tau(u) = u(\tau - \mathbb{1}\{u<0\}) over both data sources simultaneously in one quantreg fit, unlike the IVWC sibling, which fits matched-pair and reservoir quantile regressions separately and pools them by inverse-variance weighting. Standard errors use Powell's sandwich estimator. likelihood_tier = "none": quantile regression minimizes an asymmetric-loss objective, not a proper likelihood, so no likelihood-ratio or parametric-bootstrap methods are exposed. Requires the quantreg package.

Super class

Inference -> InferencePropKKQuantileRegrOneLik

Methods

Public methods

+ inherited public methods from Inference

InferencePropKKQuantileRegrOneLik$new()

Initialize proportion-response KK combined-likelihood quantile-regression inference. Responses are fitted on the logit scale; the shared stacked quantile-regression fit is documented in InferenceContinKKQuantileRegrOneLik (this class's continuous-response sibling, sharing the same KKQuantileRegrOneLik component).

Usage
InferencePropKKQuantileRegrOneLik$new(
  des_obj,
  model_formula = NULL,
  tau = 0.5,
  verbose = FALSE
)
Arguments
des_obj

A DesignSeqOneByOne object whose entire n subjects are assigned and response y is recorded within.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

tau

The quantile level on the logit scale, strictly between 0 and 1. Default is 0.5.

verbose

Whether to print progress messages.


InferencePropKKQuantileRegrOneLik$clone()

The objects of this class are cloneable with this method.

Usage
InferencePropKKQuantileRegrOneLik$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Koenker, R. (2005). Quantile Regression. Cambridge University Press. doi:10.1017/CBO9780511754098

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropKKQuantileRegrOneLik$new(seq_des)
inf$compute_estimate()


Quantile Regression Inference for Proportion Responses

Description

Fits a quantile regression for proportion responses (constrained to (0, 1)) using the treatment indicator and, optionally, all recorded covariates as predictors. Inference is performed on the logit (log-odds) scale: responses y \in (0,1) are transformed via \text{logit}(y) = \log(y/(1-y)) before quantile regression, so the estimated treatment effect is a log-odds-ratio shift at quantile tau; by default tau = 0.5, so this is a median log-odds-ratio shift.

Fitting is via rq (method "br", the Barrodale-Roberts simplex algorithm, for the point estimate; the default Frisch-Newton-adjacent interior-point path for the variance-computing fit) on logit(y) ~ w + covariates with no intercept column (the design matrix already carries one). Standard errors use quantreg's Powell (1991) kernel sandwich "nid" estimator (heteroskedasticity- and design-robust, valid under non-i.i.d. errors) when available, falling back to the i.i.d.-errors "iid" estimator if "nid" extraction fails; inference on the resulting standard error uses the asymptotic normal (Wald) approximation, not a resampling-based reference distribution, for the asymptotic CI/p-value paths. compute_asymp_confidence_interval/compute_asymp_two_sided_pval use the fit's residual degrees of freedom n - p in a t-reference (via compute_z_or_t_ci_from_s_and_df) rather than a plain normal reference, so the interval/test remain slightly conservative in small samples relative to a bare Wald z.

This class requires the quantreg package, which is listed under Suggests and is not installed automatically with EDI. Install quantreg manually before use. Only uncensored proportion responses are supported (checked via assertNoCensoring at construction).

Super class

Inference -> InferencePropQuantileRegr

Methods

Public methods

+ inherited public methods from Inference

InferencePropQuantileRegr$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize a quantile-regression inference object for a completed design with a proportion response.

Usage
InferencePropQuantileRegr$new(
  des_obj,
  model_formula = NULL,
  tau = 0.5,
  verbose = FALSE
)
Arguments
des_obj

A completed Design object with a proportion response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

tau

The quantile to estimate (default 0.5).. Default 0.5.

verbose

Whether to print progress messages.. Default FALSE.


InferencePropQuantileRegr$compute_estimate()

Computes the fitted treatment coefficient of the tau-quantile regression of logit(y) on the treatment indicator (plus any adjustment covariates) — a log-odds-ratio shift at quantile tau of the proportion response, not a difference in means or in the raw-scale quantile. Caches beta_hat_T (and, unless estimate_only, the standard error and residual degrees of freedom) so repeated calls are cheap; returns NA_real_ if the reduced design matrix is degenerate (fewer usable rows than columns) or the quantreg fit fails/errors.

Usage
InferencePropQuantileRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.


InferencePropQuantileRegr$compute_estimate_with_bootstrap_weights()

Recomputes compute_estimate's treatment log-odds-ratio-shift coefficient with each subject's (or block's) contribution to the tau-quantile fit reweighted by subject_or_block_weights (expanded to per-row weights and passed as quantreg::rq(..., weights = ...)), for the Bayesian bootstrap contract; see InferenceBayesianBootstrap. Writes into the same beta_hat_T/s_beta_hat_T/df cache fields that compute_estimate reads from — a call to this method overwrites the cached original-data estimate with the bootstrap-reweighted one, so a subsequent compute_estimate() call will return the bootstrap replicate's value from cache rather than recomputing on the original data, until the cache is reset by whatever higher-level bootstrap driver owns this object's lifecycle. Returns NA_real_ under the same degenerate-design/fit-failure conditions as compute_estimate.

Usage
InferencePropQuantileRegr$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip variance calculations.


InferencePropQuantileRegr$compute_asymp_confidence_interval()

Uses the shared asymptotic confidence-interval contract; see InferenceAsymp.

Usage
InferencePropQuantileRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.


InferencePropQuantileRegr$compute_asymp_two_sided_pval()

Uses the shared asymptotic two-sided p-value contract; see InferenceAsymp.

Usage
InferencePropQuantileRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null difference to test against. Default is zero.


InferencePropQuantileRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferencePropQuantileRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Koenker, R. and Bassett, G. (1978). "Regression Quantiles." Econometrica, 46(1), 33-50, doi:10.2307/1913643, for quantile regression itself. Powell, J. L. (1991). "Estimation of Monotonic Regression Models under Quantile Restrictions," in Nonparametric and Semiparametric Methods in Econometrics and Statistics, Cambridge University Press, for the "nid" sandwich standard error.

See Also

InferenceContinQuantileRegr for the untransformed (continuous-scale) analogue of this class.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropQuantileRegr$new(seq_des)
inf$compute_estimate()


Zero/One-Inflated Beta Inference for Proportion Responses

Description

Internal class for non-KK zero/one-inflated beta regression models. The response is modeled as a three-component mixture with point masses at 0 and 1 plus a beta-distributed interior component on (0, 1). The reported treatment effect is the treatment coefficient from the beta mean submodel, on the logit scale, conditional on the response falling strictly inside (0, 1): it is not, and should not be read as, the effect on the unconditional mean E[Y], which also depends on how treatment shifts the zero/one inflation probabilities. A design where treatment moves mass between the point masses and the interior, with no shift in the interior beta mean, will report a null treatment coefficient here even though E[Y] changed.

The beta mean submodel uses treatment alone in the univariate class and treatment plus covariates in the multivariate class. The zero/one inflation submodels use model_formula_zero_one, which defaults to ~ . so that treatment plus all available covariates enter those auxiliary pieces. See the class-level description above for what the reported coefficient does and does not represent.

Marginal estimand. This class composes MarginalEstimand (set_estimand()/get_estimand()/get_supported_estimands()); in addition to the default "conditional" estimand described above, it supports "marginal_mean_diff": the model-implied unconditional mean recombines all three mixture components, E[Y \mid x, w] = \pi_1(x, w) \cdot 1 + (1 - \pi_0(x, w) - \pi_1(x, w)) \cdot \mathrm{logit}^{-1}(x^\top \beta + \beta_T w) (the zero mass contributes nothing), where \pi_0/\pi_1 are the normalized zero/one-inflation mixture probabilities. The reported treatment effect under "marginal_mean_diff" is the g-computation average \hat\tau = n^{-1} \sum_i [\hat E(Y \mid x_i, w=1) - \hat E(Y \mid x_i, w=0)], on the response's natural [0,1] scale (not the log-odds scale of the conditional estimand). Standard errors are delta-method, against the joint covariance of [\beta, \log\phi, \gamma_0, \gamma_1] already returned by fast_zero_one_inflated_beta_cpp, using a numerical (central-difference) gradient of \hat\tau — see marginal_estimand_report.md → TODO-4 for why analytic differentiation was not used. This is a pure post-fit transform of the same cached maximum-likelihood fit (no refit), so only Wald-via-delta-method inference is available under a marginal estimand (no likelihood-ratio/ score/gradient test — see set_estimand()'s testing-type interaction).

Super class

Inference -> InferencePropZeroOneInflatedBetaRegr

Methods

Public methods

+ inherited public methods from Inference

InferencePropZeroOneInflatedBetaRegr$new()

Initialize inference for the three-component zero/one-inflated beta mixture model; see InferencePropZeroOneInflatedBetaRegr for the model form and the important caveat that the reported treatment coefficient is conditional on the interior (0,1) component, not an unconditional-mean effect. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Usage
InferencePropZeroOneInflatedBetaRegr$new(
  des_obj,
  model_formula = NULL,
  model_formula_zero_one = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with a proportion response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

model_formula_zero_one

Formula for the zero/one inflation submodels. Defaults to ~ ., meaning treatment plus all available covariates.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.


InferencePropZeroOneInflatedBetaRegr$compute_estimate()

Fits the zero/one-inflated beta mixture model by maximum likelihood (jointly the beta mean submodel, the zero/one inflation submodels, and the beta precision). Under the default estimand = "conditional", returns \hat\beta_T, the treatment log-odds-ratio from the beta mean submodel conditional on the interior (0,1) component — see InferencePropZeroOneInflatedBetaRegr's estimand caveat. Under estimand = "marginal_mean_diff" (set via set_estimand()), returns the g-computation marginal mean difference instead — see the class-level @details for the formula. The underlying model fit is identical either way (a pure post-fit transform of the same cached fit, no refit).

Usage
InferencePropZeroOneInflatedBetaRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.


InferencePropZeroOneInflatedBetaRegr$compute_asymp_confidence_interval()

Wald confidence interval, dispatched by testing_type for the conditional estimand (score/gradient/ likelihood-ratio/Bartlett available; see InferenceAsympLik); under a marginal estimand testing_type is always "wald" (the only value set_estimand() permits there), so this always resolves to the delta-method interval. Calls self$compute_estimate() first (not private$shared() directly) so the estimand-aware cache is always current regardless of call order.

Usage
InferencePropZeroOneInflatedBetaRegr$compute_asymp_confidence_interval(
  alpha = 0.05
)
Arguments
alpha

Two-sided miscoverage rate; the returned interval targets 1 - alpha coverage.


InferencePropZeroOneInflatedBetaRegr$compute_asymp_two_sided_pval()

Wald two-sided p-value, dispatched by testing_type exactly as compute_asymp_confidence_interval(); see that method's description for the marginal-estimand always-Wald note.

Usage
InferencePropZeroOneInflatedBetaRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment-effect value under the current estimand (conditional log-odds-ratio, or marginal mean difference).


InferencePropZeroOneInflatedBetaRegr$compute_estimate_with_bootstrap_weights()

Refits the zero/one-inflated beta model with subject/block-level weights applied to the fitting log-likelihood (Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded to row level via private$expand_subject_or_block_weights_to_row_weights()), and returns the reweighted conditional log-odds-ratio estimate \hat\beta_T^{(w)}.

Usage
InferencePropZeroOneInflatedBetaRegr$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip variance calculations.


InferencePropZeroOneInflatedBetaRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferencePropZeroOneInflatedBetaRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Ospina, R., and Ferrari, S. L. P. (2010). "Inflated beta distributions." Statistical Papers, 51(1), 111-126, doi:10.1007/s00362-008-0125-4, for the zero/one-inflated beta mixture density; Ferrari, S., and Cribari-Neto, F. (2004). "Beta regression for modelling rates and proportions." Journal of Applied Statistics, 31(7), 799-815, doi:10.1080/0266476042000214501, for the interior beta-regression submodel.

See Also

InferencePropBetaRegr for the plain (non-inflated) beta regression model.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropZeroOneInflatedBetaRegr$new(seq_des)
inf$compute_estimate()


Randomization-based Inference

Description

Abstract class for randomization-based inference.

Super class

Inference -> InferenceRand

Methods

Public methods

+ inherited public methods from Inference

InferenceRand$approximate_randomization_distribution_beta_hat_T()

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Usage
InferenceRand$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

delta

The null difference. Default 0.

transform_responses

Type of transformation. Default "none".

show_progress

Show progress bar. Default TRUE.

permutations

Pre-computed permutations. Default NULL.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

Returns

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.


InferenceRand$supports_rand_pval_for_incidence()

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Usage
InferenceRand$supports_rand_pval_for_incidence()
Returns

A single logical.


InferenceRand$compute_rand_two_sided_pval()

Computes a randomization-based p-value.

Usage
InferenceRand$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors.

delta

Null difference.

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

Returns

Randomization p-value.


InferenceRand$clone()

The objects of this class are cloneable with this method.

Usage
InferenceRand$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Bootstrap Randomization Test Inference

Description

Each of the B null draws is generated by (1) resampling n subject rows with replacement from the observed data and (2) drawing one fresh assignment vector w from the actual experimental design run on the resampled covariates. The test statistic computed on each such draw forms a null distribution against which the observed statistic (actual data, actual w) is compared.

Details

Abstract class implementing the bootstrap randomization test (BRT), a hybrid of the nonparametric bootstrap and the randomization test of Fisher's sharp null. This construction is inspired by Kallus, N. (2018), "Optimal a priori balance in the design of controlled experiments," Journal of the Royal Statistical Society: Series B, 80(1), 85-112, Section 5 ("Algorithms for inference"), Algorithm 4 – see @references below for the full attribution and what EDI's implementation adds beyond it.

Motivating heuristic: the sharp null makes the science table fully known. Under Fisher's sharp null H_0: y_i(0) = y_i(1) = y_i for all i, the observed outcome is the outcome regardless of assignment, so each observed row (x_i, y_i) is a complete description of that subject: it can be paired with any w and the outcome is still correct. Resampling rows and drawing a fresh w therefore simulates an entire new experiment — new subjects from \hat{F}_n (the empirical distribution of subjects), new assignment from the true, known assignment mechanism — under the assumption that outcomes do not respond to treatment. Both simulation ingredients are faithful to the real data-generating process under H_0: the assignment mechanism is exact (drawn from the actual design, including sequential covariate-dependent designs, since the design depends only on the covariates), and the subject distribution is approximately right by the usual bootstrap argument \hat{F}_n \to F. This is a motivating heuristic, not a proof — see "Known limitations and open theoretical question" below for exactly where it stops short of establishing asymptotic validity, and why. If H_0 is false, the construction still generates the null distribution (it forcibly treats y as invariant to w) while the observed statistic drifts into the tail — that asymmetry is the intended source of power, independent of the open validity question below.

Why this outgrew its original motivation. This construction was originally implemented for a narrow reason: it works even when w is deterministic or near-deterministic, exactly the degenerate case Kallus (2018) needed it for (see @references). It turns out to be useful well beyond that: (1) it is design-agnostic for free, since it calls draw_ws_according_to_design() on the resampled data rather than enumerating a permutation space – for the matching-on-the-fly designs (DesignSeqOneByOneKK14/KK21/KK21stepwise), that permutation space depends on the whole sequential arrival/matching history and has no closed form worth enumerating, so a classic randomization test would need bespoke combinatorics this construction avoids entirely; (2) it generalizes the inferential target from finite-population (conditional on exactly these n subjects) to superpopulation, independent of the degenerate-design motivation; (3) it inherits the rest of the inference machinery (CI inversion, the estimand/testing-type axes) automatically by living in the same class hierarchy as everything else, rather than being a bolted-on special case usable only for the one scenario that motivated it.

How it differs from the pure randomization test. The classic randomization test conditions on the realized sample (the "science table") and is exact for any n. The BRT is unconditional: it targets the distribution of the statistic over both subject sampling and randomization, and the null tested is the compound "sharp null AND subjects are i.i.d. draws from F". Exactness is traded for a population-level (superpopulation) interpretation. Equivalently, the BRT is a parametric bootstrap test whose "parameter" is the pair (F, design): the design part is known exactly, F is plugged in via \hat{F}_n, and the sharp null is precisely what makes \hat{F}_n estimable from one arm-agnostic dataset. The pure randomization test is the special case where F is conditioned away.

Known limitations and open theoretical question: asymptotic validity is not proven, for two distinct, identifiable reasons. Kallus (2018) himself states that asymptotic validity of this bootstrap test is "believed" rather than established, and explicitly calls it an open question. EDI's implementation does not resolve that question; the two specific gaps are:

  1. Glivenko-Cantelli is not bootstrap consistency. The motivating heuristic above leans on \hat{F}_n \to F to justify substituting a bootstrap resample for a genuinely fresh i.i.d. draw from F. But \hat{F}_n \to F (uniform convergence of the empirical CDF, a Glivenko-Cantelli statement) is a claim about the empirical distribution itself; it is not the same as consistency of the bootstrap, i.e. convergence of the sampling distribution of the test statistic computed under resampling from \hat{F}_n to the statistic's true sampling distribution under F. That second, load-bearing claim is well known to require statistic-specific regularity conditions (smoothness/Hadamard differentiability of the statistic as a functional of the empirical process, adequate moment conditions, etc. — the classical Bickel-Freedman-type distinction) and is known to fail for some statistics even when \hat{F}_n \to F holds trivially (e.g. non-smooth statistics, extreme-value/max-type statistics, heavy-tailed distributions without enough moments, non-regular estimators). No such regularity condition is verified, case by case, for the estimators this class is composed into.

  2. The naive version of this concern is already mitigated by design; a narrower residual question remains open. A naive i.i.d. row bootstrap would create exact duplicate rows that never occur in real data, and for a design whose assignment mechanism depends on the joint covariate configuration of the whole sample — exactly what EDI's matching-on-the-fly designs (DesignSeqOneByOneKK14/KK21/KK21stepwise) and fixed matched-pair designs (DesignFixedBinaryMatch) do — those artificial exact ties would fabricate zero-distance ("perfect") matches that a genuinely fresh sample from F would essentially never produce. This package does not resample rows i.i.d. for matching-capable designs. DesignMatchingAbstract$draw_bootstrap_indices() (design_matching_abstract.R), inherited by every matching-capable design, dispatches to a pair-aware resampler (private$draw_matching_bootstrap_indices()) that resamples reservoir subjects i.i.d. and matched pairs as intact units — preserving each pair's true, historically realized within-pair covariate distance rather than fabricating an artificial zero-distance one. This is the correct fix for the naive-duplicate-row concern, and it is already active for every BRT draw on a matching-capable design (InferenceRandBootstrap inherits bootstrap_sample_indices() from the same chain, so no separate wiring is needed). What remains open, narrower than the naive concern above: reservoir subjects are still resampled i.i.d. independently of the intact pairs, so a single reservoir subject can still appear more than once in one bootstrap draw; whether two duplicate copies of the same original reservoir subject can subsequently be matched to each other by the re-run sequential matching algorithm (fabricating a same-subject zero-distance pair as a second-order effect, distinct from the naive first-order concern this mitigation addresses) has not been analyzed in this package. This is a narrower, unquantified residual question, not a demonstrated bias.

Status: point 1 above is a documented open theoretical question, not a settled result; point 2's naive form is mitigated by the pair-aware resampler described above, with only the narrower residual question left open. Neither this package nor Kallus (2018) supplies a general proof of asymptotic validity covering point 1; both explicitly flag it as unresolved rather than claiming a proof. Treat the p-values and confidence intervals from this class as resting on a well-motivated but unproven asymptotic argument. A rigorous asymptotic proof (or a demonstrated counterexample) remains future work.

Other implementation notes. (1) For sequential designs the bootstrap sample needs an arrival order; the i.i.d. resampling order plays that role, matching the i.i.d.-arrivals assumption of the KK designs. (2) Studentization is not needed for the motivating heuristic above but, as with any bootstrap test, an asymptotically pivotal statistic improves the level's rate of convergence where the construction is valid.

Users do not instantiate this class directly: every concrete inference class in the package inherits from it, so its methods (compute_rand_bootstrap_two_sided_pval, approximate_rand_bootstrap_distribution_beta_hat_T) are available on any inference object. See InferenceRandBootstrapCI for the companion confidence interval.

Super classes

Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap

Methods

Public methods

+ inherited public methods from InferenceNonParamBootstrap
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference

InferenceRandBootstrap$get_supported_rand_bootstrap_pval_types()

Returns the type values compute_rand_bootstrap_two_sided_pval() accepts.

Usage
InferenceRandBootstrap$get_supported_rand_bootstrap_pval_types()

InferenceRandBootstrap$approximate_rand_bootstrap_distribution_beta_hat_T()

Computes the bootstrap randomization null distribution of the test statistic under Fisher's sharp null (shifted by delta): each draw resamples subject rows with replacement and draws one fresh assignment vector from the design.

Usage
InferenceRandBootstrap$approximate_rand_bootstrap_distribution_beta_hat_T(
  B = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  debug = FALSE,
  bootstrap_type = NULL,
  rand_bootstrap_draws = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
B

Number of bootstrap randomization draws. Default 501.

delta

The null treatment effect (on the transform_responses scale). Default 0.

transform_responses

Type of response transformation used to impose the sharp null shift for nonzero delta. Default "none".

show_progress

A flag indicating whether a progress bar should be displayed.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

bootstrap_type

Optional bootstrap-resampling scheme; see approximate_bootstrap_distribution_beta_hat_T for legal values. Default NULL.

rand_bootstrap_draws

Optional pre-generated draws (as returned by the private method generate_rand_bootstrap_draws) enabling common random numbers across calls with different delta values (used by the CI inversion). Default NULL.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging.

Returns

When debug = FALSE (default), a numeric vector of length B containing the null-distribution draws. When debug = TRUE, a list with: values, errors, warnings, num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.


InferenceRandBootstrap$compute_rand_bootstrap_two_sided_pval()

Computes a bootstrap randomization two-sided p-value for Fisher's sharp null (shifted by delta): the observed statistic (actual data, actual w) is compared against the null distribution generated by resampling rows and drawing fresh assignments from the design.

Usage
InferenceRandBootstrap$compute_rand_bootstrap_two_sided_pval(
  B = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  bootstrap_type = NULL,
  rand_bootstrap_draws = NULL,
  zero_one_logit_clamp = .Machine$double.eps,
  type = "percentile"
)
Arguments
B

Number of bootstrap randomization draws. Default 501.

delta

The null treatment effect. Default 0.

transform_responses

Type of response transformation for the sharp null shift. The default "none" resolves by response type (logit for proportion, log for count and survival, identity otherwise), matching compute_rand_two_sided_pval.

na.rm

Remove non-finite null draws. Default TRUE.

show_progress

A flag indicating whether a progress bar should be displayed.

bootstrap_type

Optional bootstrap-resampling scheme; see approximate_bootstrap_distribution_beta_hat_T for legal values. Default NULL.

rand_bootstrap_draws

Optional pre-generated draws for common random numbers across delta values. Default NULL.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging.

type

Test statistic type. "percentile" (default) uses the raw estimator as the BRT test statistic, giving an asymptotic p-value that inherits the unconditional superpopulation validity of the BRT. "studentized" divides each null draw's signed deviation from the null by its per-draw standard error: p = 2\min(P(z^0_b \ge z), P(z^0_b \le z)) where z^0_b = (t^0_b - \delta)/\hat{s}^0_b; yields asymmetric CI under inversion. "symmetric-percentile-t" uses the absolute pivot p = P(|t^0_b - \delta|/\hat{s}^0_b \ge |t - \delta|/\hat{s}); CI is symmetric. Both SE-based types require the class to expose s_beta_hat_T; fall back to "percentile" if the SE is unavailable. "smoothed" adds kernel noise \varepsilon_b \sim N(0, \hat{\sigma}/\sqrt{n}) to each resampled draw before imposing the null shift, reducing discreteness in the null distribution. Only meaningful for continuous responses. For count responses the noisy draw is rounded and floored at zero so it stays on the non-negative integer support the Poisson-family likelihoods require; at the default bandwidth this makes the smoothing nearly a no-op for low counts.

Theoretical justification. Order-statistic/rank-based estimators (e.g. the Hodges-Lehmann pseudo-median) take only finitely many values, so their bootstrap/ randomization null distribution is a step function; this both coarsens p-values and destabilizes the uniroot-based delta search in InferenceRandBootstrapCI's compute_rand_bootstrap_confidence_interval, which assumes an approximately continuous, monotone p-value curve. Restoring continuity by convolving the resampling distribution with a shrinking-bandwidth kernel is the classical "smoothed bootstrap" device: Silverman (1981, "Density ratios, empirical likelihood and cot death", Applied Statistics 30(2):142-145) for kernel smoothing of a resampled empirical distribution; Silverman & Young (1987, "The bootstrap: To smooth or not to smooth?", Biometrika 74(3):469-479) for applying that smoothing directly to the bootstrap resampling scheme; and Hall, DiCiccio & Romano (1989, "On smoothing and the bootstrap", Annals of Statistics 17(2):692-704) for the bandwidth conditions under which smoothing improves, rather than degrades, coverage accuracy.

Implementation caveat (ad hoc, not a certified instance of the above). The bandwidth used here, \hat{\sigma}/\sqrt{n} (the raw SE-of-the-mean scale of the response), is a pragmatic engineering choice, not one derived from or validated against the bandwidth-selection results in the sources above. Those references smooth the resampling distribution itself with a bandwidth chosen to trade off bias against variance (often shrinking slower than n^{-1/2}, e.g. n^{-1/5}-type KDE rates); here, noise is instead added directly to each already-resampled response at a fixed n^{-1/2} rate. The bandwidth is not exposed as a parameter, has no zero-noise escape hatch (the only way to disable smoothing is to pick a different type), and its coverage behavior has not been validated by simulation in this package. Treat it as a discreteness patch that is qualitatively motivated by the literature above, not a certified implementation of it.

Performance. Every class with a C++ compute_fast_rand_bootstrap_distr kernel that operates on real-valued responses (Wilcox HL, simple mean difference, OLS, robust regression, CoxPH, Weibull marginal, log-rank, RMST, KM-diff) accepts the smoothing noise directly in its kernel. This only speeds up InferenceRandBootstrapCI's compute_rand_bootstrap_confidence_interval(type = "smoothed"), since CI inversion pre-materializes one set of fresh assignments up front (common random numbers reused across every delta evaluated during root-finding) and the fast kernel can engage on each evaluation; measured on InferenceAllSimpleWilcox, n = 30, B = 99: the forced-slow-fallback CI took 25.5s, the fast-kernel CI took 0.5s (about 50x). A standalone compute_rand_bootstrap_two_sided_pval(type = "smoothed") call (no CI inversion) is not accelerated by this: it deliberately draws the fresh assignment lazily per replicate (materialize_w = FALSE), so rand_bootstrap_draw_matrices() cannot build the matrices the fast kernels need and the R-level fallback still runs regardless of this fix. The two ordinal classes (InferenceOrdinalRidit, InferenceOrdinalJonckheereTerpstraTest) still use the slower R-level fallback in every case, because adding continuous Gaussian noise to integer category codes is not statistically meaningful — see the response-type caveat above.

Returns

A two-sided p-value.


InferenceRandBootstrap$clone()

The objects of this class are cloneable with this method.

Usage
InferenceRandBootstrap$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kallus, N. (2018), "Optimal a priori balance in the design of controlled experiments," Journal of the Royal Statistical Society: Series B, 80(1), 85-112, Section 5 ("Algorithms for inference"), Algorithm 4 – the nearest prior appearance of this exact construction (resample subjects, then redraw one fresh assignment from the design on the resampled covariates), introduced there for the degenerate case of his a priori balancing designs (PSODs), where a design can admit only one or very few distinct treatment permutations, so the classic Fisher randomization test has no power (it always returns p-value 1). Kallus himself cites Efron, B. and Tibshirani, R. (1993), An Introduction to the Bootstrap, Chapman & Hall, for the bootstrap ingredient, and Good, P. (2005), Permutation, Parametric and Bootstrap Tests of Hypotheses, Springer, for the test/confidence-interval duality used to invert his Algorithm 4 into intervals – and states asymptotic validity of the bootstrap test as an open question, not a proven result (see "Known limitations and open theoretical question" above, which this package inherits and extends with the matching-design- specific failure mode). EDI's contribution beyond Algorithm 4 is generalizing the construction to arbitrary designs (fixed and sequential, including the matching-on- the-fly family) and response types, and pairing it with a confidence-interval inversion (InferenceRandBootstrapCI); it does not resolve the open validity question.

Examples

seq_des = DesignSeqOneByOneKK14$new(n = 20, response_type = "continuous")
for (i in 1:20) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(20))
seq_des_inf = InferenceAllSimpleAverageDiff$new(seq_des)
seq_des_inf$compute_rand_bootstrap_two_sided_pval(B = 101)

Bootstrap Randomization Confidence Intervals

Description

The CI is the set of delta values whose bootstrap randomization p-value exceeds alpha. All p-value evaluations across candidate delta values share one set of pre-generated draws (resampled row indices and fresh design assignments) — common random numbers — so the p-value is a deterministic, near-monotone function of delta and the bound search is stable. See InferenceRandBootstrap for the statistical justification of the test being inverted; the resulting interval inherits its unconditional, superpopulation interpretation and asymptotic validity.

Users do not instantiate this class directly: every concrete inference class in the package inherits from it, so compute_rand_bootstrap_confidence_interval is available on any inference object.

Details

Abstract class implementing confidence intervals by inverting the bootstrap randomization test of InferenceRandBootstrap over the null effect delta.

Super classes

Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI

Methods

Public methods

+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference

InferenceRandBootstrapCI$get_supported_rand_bootstrap_ci_types()

Returns the type values compute_rand_bootstrap_confidence_interval() accepts.

Usage
InferenceRandBootstrapCI$get_supported_rand_bootstrap_ci_types()

InferenceRandBootstrapCI$compute_rand_bootstrap_confidence_interval()

Computes a confidence interval by inverting the bootstrap randomization test over the null effect delta. For statistics that are affine in the additive sharp-null shift (e.g. the simple mean difference and the OLS treatment coefficient), the inversion is performed in closed form from the breakpoints of the p-value step function — exact given the draws and requiring no bisection. Otherwise, the generic bisection search is used. When the p-value does not drop below alpha / 2 anywhere within the search radius, a conservative bound is returned at the search boundary rather than NA; each such event emits a message() and increments the private field rand_bootstrap_ci_conservative_count.

Usage
InferenceRandBootstrapCI$compute_rand_bootstrap_confidence_interval(
  alpha = 0.05,
  B = 501,
  pval_epsilon = 0.005,
  show_progress = TRUE,
  max_expansions = 7L,
  bootstrap_type = NULL,
  zero_one_logit_clamp = .Machine$double.eps,
  type = "percentile"
)
Arguments
alpha

The confidence level 1 - alpha. Default 0.05.

B

Number of bootstrap randomization draws. Default 501.

pval_epsilon

Bisection tolerance (on both the delta bracket width and the p-value span). Default 0.005.

show_progress

A flag indicating whether progress should be displayed.

max_expansions

Maximum number of bound-doubling expansions when the seed interval does not bracket the target p-value. Default 7.

bootstrap_type

Optional bootstrap-resampling scheme; see approximate_bootstrap_distribution_beta_hat_T for legal values. Default NULL.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging.

type

CI type. "percentile" (default) inverts the raw BRT p-value via bisection (or closed-form for affine statistics). "studentized" inverts the signed-pivot BRT p-value p(\delta) = 2\min(P(z^0_b \ge z), P(z^0_b \le z)) where z^0_b = (t^0_b(\delta) - \delta)/\hat{s}^0_b; yields asymmetric CI. "symmetric-percentile-t" inverts the absolute-pivot version p(\delta) = P(|t^0_b - \delta|/\hat{s}^0_b \ge |t - \delta|/\hat{s}); yields a CI symmetric around the observed estimate. Both SE-based types pre-compute \hat{s}^0_b once at \delta = 0 and reuse it across bisection steps. Yield O(n^{-1}) coverage error versus O(n^{-1/2}) for "percentile" when the pivot is asymptotically normal. Require the class to expose a standard error; return NA bounds in harden mode if unavailable. "smoothed" adds per-draw kernel noise \varepsilon_b \sim N(0, \hat{\sigma}/\sqrt{n}) to the resampled responses before imposing the null shift, reducing discreteness. Only meaningful for continuous responses; count responses are rounded and floored at zero after the noise so the Poisson-family refits stay on the integer support. See InferenceRandBootstrap's compute_rand_bootstrap_two_sided_pval for the theoretical justification (Silverman 1981; Silverman & Young 1987; Hall, DiCiccio & Romano 1989), an explicit caveat that the bandwidth used here (\hat{\sigma}/\sqrt{n}, fixed, unexposed, with no zero-noise escape hatch) is a pragmatic ad hoc choice rather than one derived from or validated against those sources, and measured performance (this CI, unlike a standalone smoothed p-value, is accelerated by the noise-aware fast kernels: about 50x on n = 30, B = 99).

Returns

A bootstrap randomization confidence interval. The interval lives on the response-transformation scale used by the test (identity for continuous, logit for proportion, log for count and survival). Bounds may be conservative (wider than necessary) when the p-value inversion cannot be completed within the search radius.


InferenceRandBootstrapCI$clone()

The objects of this class are cloneable with this method.

Usage
InferenceRandBootstrapCI$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

seq_des = DesignSeqOneByOneKK14$new(n = 20, response_type = "continuous")
for (i in 1:20) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(20))
seq_des_inf = InferenceAllSimpleAverageDiff$new(seq_des)
seq_des_inf$compute_rand_bootstrap_confidence_interval(alpha = 0.05, B = 101)

Randomization-based Confidence Intervals

Description

Abstract class for randomization-based confidence interval inference.

Super classes

Inference -> InferenceRand -> InferenceRandCI

Methods

Public methods

+ inherited public methods from InferenceRand
+ inherited public methods from Inference

InferenceRandCI$compute_rand_two_sided_pval()

Compute a randomization-based two-sided p-value for the treatment effect.

Usage
InferenceRandCI$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  type = NULL,
  args_for_type = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors.

delta

Null treatment effect value.

transform_responses

Response transformation to apply during the test. For survival responses the default "log" multiplies the recorded times of the units treated under each reference allocation by e^\delta, event and censoring times alike, with censoring indicators unchanged – the rank-based AFT residual construction (Tsiatis 1990; Wei, Ying and Lin 1990; Jin, Lin, Wei and Ying 2003); see compute_rand_confidence_interval() for the assumptions.

na.rm

Whether to remove non-finite simulated statistics.

show_progress

Whether to show progress.

permutations

Optional pre-generated assignment draws.

type

Optional incidence-specific exact randomization type.

args_for_type

Optional arguments keyed by type.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

Returns

A two-sided p-value.


InferenceRandCI$compute_rand_confidence_interval()

Computes a randomization-based confidence interval by inverting the randomization test: the interval is the set of \delta for which the two-sided randomization p-value of the sharp null "treatment effect = \delta" is at least alpha. For each candidate \delta the control potential outcomes are first imputed under that sharp null by removing the hypothesised effect from the units that were actually treated; then, for each reference allocation w_b, the effect is applied to the units treated under w_b — on the response type's scale (additive for continuous responses, multiplicative e^\delta for counts and survival times, a logit shift for proportions) — the estimator is recomputed, and the observed estimate is compared with the resulting reference distribution. This is the impute-then-permute construction of Rosenbaum (2002, ch. 2) and Imbens and Rubin (2015, ch. 5): y_{sim} = y - \delta w_{obs} + \delta w_b on the additive scale.

Survival responses. The sharp null is an accelerated-failure-time (AFT) effect: for units treated under w_b the recorded time is multiplied by e^\delta — both event times and censoring times — and the censoring indicator is carried over unchanged. Equivalently, the test is run on the residual times \log y_i - \delta w_i of every observation, censored or not. This is the residual construction underlying rank-based inference for the AFT model: Tsiatis (1990) shows that linear rank statistics computed on these residuals have mean zero at the true \delta under independent censoring, even though the residual censoring distribution then depends on treatment; Wei, Ying and Lin (1990) invert exactly this family of tests to obtain confidence intervals for AFT regression coefficients; Jin, Lin, Wei and Ying (2003) give the modern estimation and inference machinery for the same model. The alternative of rescaling only the event times and re-deriving the censoring indicator is not identifiable, because a unit's censoring time is unobserved whenever its event was. Two assumptions therefore apply: (i) independent censoring (C \perp T \mid w) for the asymptotic validity of the inverted test, and (ii) for the finite-sample exactness of the permutation version specifically, that censoring times are on the same accelerated clock as event times (C_i(1) = e^\delta C_i(0)), which is plausible for health-driven dropout and does not hold for calendar-time administrative censoring; under administrative censoring the interval is asymptotically, not exactly, valid. Because \delta is a log time-ratio, this construction is coherent only for classes whose estimand is on that scale (the Weibull AFT, marginal Weibull, Weibull-frailty, and rank-regression classes); classes whose estimand is a log hazard ratio cannot invert an AFT shift without a parametric link between the two scales and refuse this method (the six classes in EDI_LOG_HAZARD_RATIO_INFERENCE_CLASSES — the Cox family — which also do not advertise the randomization_ci capability, so InferenceSuite never offers it for them; their randomization p-value and randomization-bootstrap CI are unaffected).

Usage
InferenceRandCI$compute_rand_confidence_interval(
  alpha = 0.05,
  r = 501,
  pval_epsilon = 0.005,
  show_progress = TRUE,
  type = NULL,
  args_for_type = NULL,
  ci_search_control = NULL
)
Arguments
alpha

Significance level.

r

Number of randomization vectors.

pval_epsilon

Bisection tolerance.

show_progress

Show progress.

type

Optional incidence-specific exact randomization type.

args_for_type

Optional arguments keyed by type.

ci_search_control

Optional control list for randomization-CI search. Supported entries are fallback, seed, max_radius_se_mult (default 25), max_radius_scale_mult (default 6), max_expansions (default 7), seed_boot_B, Monte Carlo settings mc_enable, mc_batch_size, mc_min_draws, and mc_conf_level, midpoint-cache settings pval_cache_enable and pval_cache_resolution, and model-fit reuse settings fit_warm_start_enable and fit_reuse_factorizations. Set mc_enable = FALSE to force full enumeration of all requested randomization draws. The search radius is max(max_radius_se_mult * se_guess, max_radius_scale_mult * sd(y)). When the randomization p-value does not drop below alpha/2 anywhere within the search radius (e.g. the test has low power or the design has few unique permutations), a conservative CI bound is returned at the search boundary rather than NA. This guarantees a valid (though possibly wide) interval. Each such event emits a message() and increments the private field rand_ci_conservative_count for monitoring. high_precision_confirm (default TRUE) re-checks each converged bound with one full-enumeration (no early stopping) p-value evaluation and, only if that disagrees with the early-stopped bisection's conclusion, spends a short additional full-precision re-bisection to correct it – a final high-precision confirmation pass that catches sequential-Monte-Carlo early-stopping noise the cheap bisection has no way to notice on its own; see high_precision_confirm_and_refine_ci_bound()'s own comment for the full rationale and R/package_metadata/new_feature_plans/ for the deeper, not-yet-implemented redesign (anytime-valid confidence sequences) this pass is a pragmatic stopgap for.

Returns

Randomization CI. Bounds may be conservative (wider than necessary) when the p-value inversion cannot be completed within the search radius; see ci_search_control for details.


InferenceRandCI$clone()

The objects of this class are cloneable with this method.

Usage
InferenceRandCI$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Tsiatis, A. A. (1990). Estimating regression parameters using linear rank tests for censored data. The Annals of Statistics, 18(1), 354-372, doi:10.1214/aos/1176347504. Wei, L. J., Ying, Z., and Lin, D. Y. (1990). Linear regression analysis of censored survival data based on rank tests. Biometrika, 77(4), 845-851, doi:10.1093/biomet/77.4.845. Jin, Z., Lin, D. Y., Wei, L. J., and Ying, Z. (2003). Rank-based inference for the accelerated failure time model. Biometrika, 90(2), 341-353, doi:10.1093/biomet/90.2.341.


Randomization test/CI on a user-supplied statistic

Description

Runs a randomization test (and, via compute_rand_confidence_ interval(), a randomization confidence interval) for an arbitrary user-supplied statistic, without writing an Inference subclass. The statistic function is called once per permutation as fn(y, w, dead) – plain numeric/integer vectors: the current (possibly permuted) response, treatment assignment, and event indicator. It must return one scalar. For maximum speed, supply custom_randomization_statistic_cpp instead: C++ source defining a function of (NumericVector y, IntegerVector w) or (NumericVector y, IntegerVector w, IntegerVector dead) returning a scalar double, using the same convention as DesignFixedOptimal's custom_objective.

Only a randomization test and randomization confidence interval are available – there is no package point-estimator, Wald path, or bootstrap machinery on this class, so no other action needs disabling.

Super classes

Inference -> InferenceCustomRand -> InferenceRandCustom

Methods

Public methods

+ inherited public methods from InferenceCustomRand
+ inherited public methods from Inference

InferenceRandCustom$new()

Initialize.

Usage
InferenceRandCustom$new(
  des_obj,
  custom_randomization_statistic_function = NULL,
  custom_randomization_statistic_cpp = NULL,
  verbose = FALSE
)
Arguments
des_obj

See class description.

custom_randomization_statistic_function

See class description.

custom_randomization_statistic_cpp

See class description.

verbose

See class description.


InferenceRandCustom$fit()

Calls the user-supplied statistic on the observed data.

Usage
InferenceRandCustom$fit(estimate_only = FALSE)
Arguments
estimate_only

Unused; present for the 'fit()' contract.


InferenceRandCustom$clone()

The objects of this class are cloneable with this method.

Usage
InferenceRandCustom$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Inference Suite: Discover and Bundle Every Applicable Inference Class for a Design

Description

A lightweight coordinator (not itself an Inference subclass, and not part of the Inference R6 hierarchy) that pairs a single completed Design object with the full set of concrete Inference classes compatible with it. On construction, the suite consults package-level inference metadata to discover every exported, non-abstract Inference subclass whose declared response-type, matched-design (KK), blocking, and censoring requirements are all satisfied by des_obj, storing the resulting sorted class-name vector in applicable_design_classes. Because discovery is driven by metadata lookups rather than by actually attempting to construct each candidate class, the applicable list automatically stays current as new inference classes are registered elsewhere in the package, without this class needing any changes, and without risking side effects (e.g. an optional-package load failure inside some class's constructor) from a doomed construction attempt.

Compatibility rules (see is_inference_class_compatible_with_design_metadata(), also used by Design's own applicable_inference_class_names()): a candidate class is excluded if it is abstract or not exported; if it declares no compatible response types, or none match des_obj's response type; if its name contains "KK" (a matched-design-only class) but des_obj does not support KK matching; if the class's requires_blocking_design() is TRUE (currently only InferenceIncidExtendedRobins – InferenceIncidCMH works on both blocking and non-blocking designs via different standard-error estimators, so it is not excluded) but des_obj does not support blocking; or if des_obj has any left-/interval-censored subjects (a finite y_R) but the class's supports_interval_or_left_censored_data() is FALSE – both of the latter two mirror Inference$initialize()'s own construction-time gate exactly, via each class's registered requires_blocking_design/supports_general_censoring metadata (see infer_inference_requires_blocking_design()/ infer_inference_supports_general_censoring() in inference_class_registry.R). Ordinary right-censoring alone never excludes a class. The KK-name-pattern rule is still hardcoded in this class rather than looked up from a central registry.

A class that is design-compatible but whose registered required_packages are not all installed is excluded from applicable_design_classes and reported separately, by class name, in unavailable_due_to_missing_packages – so callers can tell "not applicable to this design" apart from "applicable, but an optional dependency isn't installed" (see the Discovery section of fix_inference_hierarchy.md). Package availability is never a reason a class is treated as design-incompatible.

Construction itself does not compute any estimates, p-values, or confidence intervals – it only discovers and validates which inference classes are applicable and does not eagerly construct any of them. This class's run_all_inference() method does that: it constructs and fits every applicable class and returns a uniform comparison across them (see that method's own documentation for the output schema, the CI/p-value method selection policy, and the screen/html/plots/pdf/ save_results_as_JSON output options). lock_objects = FALSE allows ad hoc fields to be attached to an instance after construction.

Every row this class discovers and fits is a test about the same outcome variable, by construction. response_type is a required, immutable constructor argument on Design (read-only thereafter via get_response_type()), and this class discovers every candidate in applicable_design_classes from one attached Design object's one response_type – there is no code path here that spans two response types in a single instance. This matters beyond bookkeeping: a planned comparison-across-classes feature (a single combined-evidence p-value summarizing every row, via a dependence-robust combination test) relies on every combined class sharing one sharp null of "no treatment effect on this outcome" – which only holds when every test concerns the same outcome variable (combining, say, a survival model's p-value with an unrelated continuous-outcome model's p-value would not be valid, since a real effect on one with none on the other is entirely plausible). That precondition is guaranteed here structurally, not by caller discipline.

Same-Y does not mean same estimand, and that matters for interpretation. Rows discovered here routinely test genuinely different \theta_i on the same outcome – a mean difference, a log-odds ratio under one link, a rank-based effect, a quantile shift – each \theta_i its own parameterization's own null. The combined H_0: \theta_1 = 0 \cap \dots \cap \theta_k = 0 is one coherent claim (Fisher's unit-level sharp null: no treatment effect whatsoever) only under a randomization-based/exact procedure, where every possible summary of "effect" is simultaneously zero by construction. Asymptotic/likelihood-based procedures instead test a weak, population-level null specific to their own parameterization, and weak nulls for genuinely different summaries are not generally nested – treatment can shift a distribution's variance or a high quantile while leaving its mean exactly unchanged, so "mean difference = 0" does not imply "quantile effect = 0" outside location-shift-style models. This is comparatively safe across different link functions for the same latent effect (e.g. cauchit/probit/cloglog/logit on the same binary/ordinal response, which typically share one underlying latent-variable null) and comparatively riskier across different kinds of summary (a mean-difference test combined with a rank-based test combined with a quantile-regression test). Under that weak-null reading, a rejection is more honestly read as "at least one specific summary of this outcome's distribution differs" rather than "there is one coherent nonzero effect" – reinforcing, not loosening, the interpretation caveat below.

Combined Evidence interpretation caveat (read before using combined_evidence$pval): the Cauchy combination test is a union-intersection test of H_0: \theta_1 = 0 \cap \theta_2 = 0 \cap \dots \cap \theta_k = 0 against the alternative that at least one \theta_i \neq 0. A significant combined_evidence$pval is therefore evidence of an effect in at least one of these senses, not evidence for a specific estimate or direction – it does not say which class's estimand is nonzero, nor does a small combined p-value imply every (or even most) constituent p-values were small. Do not report combined_evidence$pval as if it estimated a single effect size, and do not treat it as validating any one class's estimate over another's; its only valid use is as evidence that *some* legitimate way of looking for a treatment effect on this outcome found one. This guarantee assumes every constituent p_i is itself valid (i.e. uniform under its own null) – the Cauchy combination's dependence-robust size control says nothing about which such p_i is small, so a single miscalibrated or misspecified procedure (e.g. an asymptotic approximation that breaks down for this sample) can dominate the combined result the same way it would dominate a plain minimum-p-value test, even when every other procedure shows nothing. A significant combined_evidence$pval is worth cross-checking against the per-estimand CI-forest plot (plots$ci_forest) before trusting it – one outlying interval sitting apart from a cluster of concordant ones is a miscalibration flag, not confirmed evidence.

This same-Y precondition is guaranteed within one InferenceSuite instance structurally (one Design, one response_type), but the architecture cannot stop a caller from manually combining raw pvals pulled from two separate InferenceSuite objects' results_tables outside cct_combine_pvalues() – doing so is outside this function's validity guarantee and is not a supported use of this metric.

Public fields

applicable_design_classes

Character vector of applicable inference class names derived during initialization.

unavailable_due_to_missing_packages

A named list, one entry per otherwise-design-compatible class whose registered required_packages are not all installed: names are class names, values are the character vector of missing package names. These classes are excluded from applicable_design_classes but are reported here separately from plain design/response-type incompatibility (see the class-level docs' "Discovery" rules), so callers can distinguish "not applicable to this design" from "applicable, but an optional dependency isn't installed." Empty named list if every design-compatible class has all required packages available.

Methods

Public methods


InferenceSuite$new()

Discover every Inference class applicable to des_obj (see the class-level documentation for the compatibility rules) and validate any per-class constructor overrides in inference_params, storing applicable_design_classes for later use. This constructor does not instantiate any Inference object itself; callers are expected to construct the specific classes they need (optionally passing the validated inference_params) from the discovered list.

Usage
InferenceSuite$new(des_obj, model_formula = NULL, inference_params = list())
Arguments
des_obj

A completed Design object (validated via is(des_obj, "Design") when assertions are enabled; see toggle_asserts).

model_formula

Accepted for interface/future-extension purposes but currently not used anywhere in this method's body – supplying a non-NULL value has no effect on discovery, validation, or any stored state. Do not rely on this parameter to affect covariate adjustment; a design's own model formula and design matrix are what individual Inference classes actually consult when later constructed from this suite's discovered class list.

inference_params

A named list of lists supplying additional constructor arguments for specific inference classes. Each name must be the name of a concrete Inference subclass that is applicable to des_obj (checked against applicable_design_classes once discovered – an inapplicable class name raises an error); the corresponding list contains keyword arguments (beyond des_obj) forwarded to that class's initialize, and every argument name supplied must match a formal parameter of that class's initialize method (other than des_obj and ...) or an error is raised naming the unknown argument(s) and the valid ones. Defaults to an empty list (no extra arguments for any class).


InferenceSuite$run_all_inference()

Construct and fit every class in applicable_design_classes, and report one uniform comparison row per class – estimate, SE, CI, p-value (each via the highest-priority available method; see inference_suite_inspect.md's "Method Selection Policy"), likelihood tier, estimand (where declared), fit time, captured warnings, and status. Unlike the constructor, this method fits models and is not free; call it explicitly when you want the comparison, not automatically.

A single class's construction or fit failure never aborts the report – it is caught and recorded as that class's status/message (see "Per-Class Failure Isolation" in the design doc).

Side effects (v1.0.0 slice): screen prints each row as its class finishes fitting (computation order), not buffered to the end, with a percent-done/estimated-time-remaining progress bar line underneath each row (the ETA is the mean per-class elapsed time so far times classes remaining), followed by a footer listing classes excluded for missing optional packages. html = TRUE writes a self-contained, timestamped HTML report (the same table plus the same footer) to output_dir and opens it via browseURL; it requires the knitr package. The plots ggplot2 visualizations, and their embedding into this HTML report, are not yet implemented (inference_suite_inspect.md TODO-7); pdf output is not yet a parameter of this method.

Usage
InferenceSuite$run_all_inference(
  screen = TRUE,
  html = FALSE,
  alpha = 0.05,
  save_results_as_JSON = FALSE,
  plots = screen,
  pdf = FALSE,
  classes = NULL,
  exclude_classes = character(),
  max_secs_per_class = NULL,
  num_cores = 1L,
  formulas = NULL,
  methods = NULL,
  basic_bootstrap = FALSE,
  compute_conf_intervals = FALSE,
  output_dir = "~",
  combined_evidence_estimands = NULL,
  combined_evidence_weighting = c("estimand_grouped", "equal", "custom"),
  combined_evidence_weights = NULL
)
Arguments
screen

Print results to the console as each class finishes. At least one of screen/html must be TRUE.

html

Render, save (output_dir, timestamped filename), and auto-open a self-contained HTML report of the results.

alpha

Significance level: confidence intervals are computed at 1 - alpha and alpha is the significance threshold used anywhere the report flags significance. Default 0.05.

save_results_as_JSON

If TRUE, serialize the return object (excluding plot objects) to a timestamped JSON file in output_dir. Requires the optional jsonlite package; if it is not installed, a warning() is issued and this artifact is skipped rather than erroring. Default FALSE.

plots

If TRUE, build and display (on the current graphics device) one visualization per estimand: an annotated confidence-interval forest plot (p-value left of each interval, interval width right of it, class/method label right-aligned on each row, color-keyed to significance at alpha) stacked over its own “Estimates” box-and-whisker subplot – a free x-axis (same label, same log10/linear choice as the forest, but its own limits) summarizing the point estimates, one point per inference class/formula, collapsed over method/type since those share one estimate. That subplot scales with how many distinct estimates there are: none for a single estimate (redundant with the forest's own dot), dots alone for 2-5, and dots over a box-and-whisker for more than 5. Built with ggplot2 and stacked into a single gtable grob (draw with grid::grid.draw()); requires the optional ggplot2 package, if it is not installed, a warning() is issued and plotting is skipped rather than erroring. Defaults to the value of screen.

pdf

If TRUE, save the visualization to one timestamped multi-page PDF file in output_dir (one page per estimand; page height scales with the largest estimand's number of CI rows). Same ggplot2 dependency and missing-package handling as plots. Default FALSE.

classes

Optional character vector restricting which applicable classes to fit – e.g. re-running against only the few classes a user is actually deciding between, without reconstructing the suite. Every name must already be in applicable_design_classes or this errors, naming the unknown name(s) and the valid ones. NULL (default) fits every applicable class.

exclude_classes

Optional character vector of applicable classes to skip, applied after classes. Same validation as classes. Default none.

max_secs_per_class

Optional per-class elapsed-time limit in seconds (via setTimeLimit), after which that class's row gets status = "timeout" instead of hanging the whole report. Protects against one pathological class (e.g. a bootstrap/randomization method with many replicates) blocking every other class. Known limitation: R's time limits are checked at R-level interrupt points, so this reliably cuts off slow R-level work but is not guaranteed to interrupt one very slow single native (C/C++/BLAS) call with no intervening R-level check. NULL (default) means no limit.

num_cores

If greater than 1, fit classes in parallel across this many forked workers (makeForkCluster) – Unix/Linux only; on other platforms this falls back to sequential (num_cores = 1) with a warning(). Screen output changes under parallel execution: fitting is a single blocking call that only returns once every worker has finished, so there is no meaningful per-class ETA to show while running – screen = TRUE instead prints a "fitting N classes across K workers" message up front, then every result row together once complete, then a total-elapsed-time summary line (a deliberate design choice, not a degraded default – see inference_suite_inspect.md's TODO-13). Default 1L (sequential, with the normal incremental streaming/progress bar).

formulas

NULL (default), a single formula (~ .), a single formula string ("~ ."), or a collection of either – including c(~ 1, ~ .), which base R already returns as a plain list of formula objects (formulas have no c() method of their own), or a character vector (c("~ .", "~ age + sex * smoking")). NULL means each class fits once with its own default formula – identical to omitting this argument entirely, since model_formula = NULL at construction already resolves to des_obj$get_design_formula() (default ~ .). When non-NULL, only classes whose constructor syntactically accepts a model_formula argument are fit once per formula in formulas (one results_table row each, disambiguated in results by "<class>[<formula>]" names); classes without a model_formula constructor argument at all still fit exactly once, ignoring formulas. Note this is a syntactic check (does the constructor accept one), not a semantic one (does the fit actually use it) – some classes accept-and-ignore model_formula (e.g. InferenceAllSimpleAverageDiff's unadjusted Welch's t-test); see fix_inference_hierarchy.md's adjusts_for_covariates registry-metadata audit, which makes the cov_model column semantics-aware wherever that audit has landed.

methods

NULL (default), a character vector of method sentinel strings, or (TODO-22) a named list, sentinel to a character vector of requested type values, or NULL, restricting which inference method(s) – and, for the three resampling sentinels marked "typed" below, which resampling/CI- construction type flavor(s) – get fit and reported per class. NULL considers every sentinel in EDI_INFERENCE_SUITE_METHOD_SENTINELS, and for each typed sentinel, every type value that class supports (queried at runtime via its own get_supported_bootstrap_ {pval,ci}_types() / get_supported_bayesian_bootstrap_ {pval,ci}_types() / get_supported_rand_bootstrap_ {pval,ci}_types() accessor – never a hardcoded type table in this package), except class/method/type combinations declared in EDI_COMPREHENSIVE_SLOW_PATHS. Those implemented but prohibitively slow paths are omitted only from this default selection. Supplying methods explicitly opts into the named sentinel/type combinations even when the registry marks them slow. Thus the default remains broad without allowing known multi-minute paths to dominate a routine report; it is not a single "best available" cascade. List-shaped example: methods = list(bootstrap = c("percentile", "bca"), rand_bootstrap = NULL) fits only "bootstrap" (restricted to the "percentile"/"bca" types that class actually supports) and "rand_bootstrap" (every type that class supports); a sentinel present as a list name with value NULL still means "every valid type" for it, exactly like the flat-vector shape – to not fit a sentinel at all, simply don't name it. Requesting a type for a sentinel with no type axis (any sentinel not marked "typed" below) errors. Valid sentinels, each corresponding to one asymptotic/exact/ randomization/resampling inference family a class may (or may not) implement:

"wald"

Asymptotic Wald inference – compute_asymp_confidence_interval()/ compute_asymp_two_sided_pval() (capability "wald"). The standard closed-form normal-approximation CI/test.

"exact"

Exact inference – compute_exact_confidence_interval()/ compute_exact_two_sided_pval_for_treatment_effect() (capability "exact_test"). Finite-sample-exact methods (e.g. Fisher's exact test, exact binomial).

"rand"

Randomization inference – compute_rand_confidence_interval()/ compute_rand_two_sided_pval() (capabilities "randomization_ci"/"randomization_test" – distinct capability names for the CI vs. p-value side, since a class can support one without the other). Design-based inference via re-randomizing the observed treatment assignment.

"rand_bootstrap" (typed)

Randomization-bootstrap inference – compute_rand_bootstrap_confidence_interval()/ compute_rand_bootstrap_two_sided_pval() (capabilities "randomization_bootstrap_ci"/"randomization_bootstrap" – distinct capability names for the CI vs. p-value side, since a class can support one without the other). Resamples under the randomization null rather than the usual iid-resampling bootstrap. type (both sides agree on the same four values, unlike "bootstrap"/"bayes_boot" below): "percentile", "studentized", "symmetric-percentile-t", "smoothed".

"jackknife"

Jackknife-Wald inference – compute_jackknife_wald_confidence_interval()/ compute_jackknife_wald_two_sided_pval() (capability "jackknife"). Leave-one-out variance estimate feeding a Wald-style CI/test.

"score"

Score (Rao) test – compute_score_confidence_interval()/ compute_score_two_sided_pval() (capability "likelihood_tests", one of its three independent sub-procedures).

"lik_ratio"

Likelihood-ratio test – compute_lik_ratio_confidence_interval()/ compute_lik_ratio_two_sided_pval() (capability "likelihood_tests").

"gradient"

Gradient test – compute_gradient_confidence_interval()/ compute_gradient_two_sided_pval() (capability "likelihood_tests").

(The plain, "best available" auto-selecting compute_lik_ratio_bartlett_two_sided_pval()/ compute_lik_ratio_bartlett_confidence_interval() dispatcher is deliberately not its own sentinel – it only picks between the two explicit variants below depending on which this class implements, so it is never a distinct inference procedure.)

"lik_ratio_bartlett_approx"

Bartlett-corrected likelihood-ratio test, Monte-Carlo-approximated correction factor pinned explicitly (for reproducibility) – compute_lik_ratio_bartlett_approx_confidence_interval()/ compute_lik_ratio_bartlett_approx_two_sided_pval() (capability "likelihood_tests"; degrades to NA for classes without an approximate Bartlett factor).

"lik_ratio_bartlett_exact"

Bartlett-corrected likelihood-ratio test, closed-form analytic correction factor pinned explicitly – compute_lik_ratio_bartlett_exact_confidence_interval()/ compute_lik_ratio_bartlett_exact_two_sided_pval() (capability "likelihood_tests"; degrades to NA for classes without an exact Bartlett factor).

"param_boot"

Bootstrap-calibrated likelihood-ratio test – compute_lik_ratio_bootstrap_confidence_interval()/ compute_lik_ratio_bootstrap_two_sided_pval() (capability "parametric_likelihood_bootstrap").

"param_boot_direct"

Direct parametric-bootstrap estimate/CI/pval for the treatment coefficient itself – compute_param_bootstrap_confidence_interval()/ compute_param_bootstrap_pval() (capability "parametric_likelihood_bootstrap"; distinct from "param_boot" above, which is a bootstrap-calibrated likelihood-ratio test, not a direct estimate).

"bayes_boot" (typed)

Bayesian bootstrap inference – compute_bayesian_bootstrap_confidence_interval()/ compute_bayesian_bootstrap_two_sided_pval() (capability "bayesian_bootstrap"). CI-side type: "percentile", "basic", "wald", "studentized", "bootstrap-t", "bca"; pval-side type swaps "basic" for "symmetric" (all others the same).

"bootstrap" (typed)

Nonparametric bootstrap inference – compute_bootstrap_confidence_interval()/ compute_bootstrap_two_sided_pval() (capability "nonparametric_bootstrap"). CI-side type: "percentile", "basic", "studentized", "bootstrap-t", "symmetric-percentile-t", "bca", "prepivoted", "double-bootstrap", "calibrated", "smoothed"; pval-side type is a smaller set – "percentile", "symmetric", "studentized", "bootstrap-t", "bca" – neither "basic" nor the other CI-only variants apply on the pval side.

For the three "typed" sentinels above, an exhaustive type list is documented here for orientation only – the actual set consulted at runtime always comes from that class's own accessor (see the top of this section), so a class need not support every value listed. ("likelihood_ratio"/"estimating_equation_likelihood_ratio" are deliberately not separate sentinels – both capabilities gate the exact same method pair "lik_ratio" above already covers.) For each class, only sentinels the class has any CI or p-value capability for (among the requested methods) get a row; for a typed sentinel, one row per type that class actually supports (intersected with any requested type subset) – a class with zero applicable sentinels, or a typed sentinel with zero resulting types, still gets exactly one row with method/type = NA_character_ (mirrors the pre-methods "no capability" row) rather than being silently dropped. A class contributing more than one applicable-sentinel row is disambiguated in results/results_table by "<class>{<method>}" or "<class>{<method>:<type>}" (or with a "[<formula>]" tag too under simultaneous formulas fan-out) names. Unlike the removed cascade, ci_method/pval_method on a given row now always match that row's own method (or are NA if this class lacks that half of the sentinel's capability, including when type is valid on one side but not the other) – there is no fallback to a different sentinel within one row.

basic_bootstrap

FALSE (default). Convenience flag: when TRUE, restricts every typed sentinel ("bootstrap"/ "bayes_boot"/"rand_bootstrap") to just that class's first (i.e. default) type value instead of fitting every type it supports – "just run the default bootstrap flavor for nonparametric/Bayesian/randomization resampling" without having to spell out methods = list(bootstrap = ..., bayes_boot = ..., rand_bootstrap = ...) by hand. Only takes effect for a typed sentinel the caller didn't already restrict via an explicit methods list entry – an explicit type request there always wins over this flag. No effect on non-typed sentinels ("param_boot"/"param_boot_direct" included – neither has a type axis, so they already run their one procedure).

compute_conf_intervals

FALSE (default). When FALSE, every task's confidence-interval side (ci_a/ci_b/ ci_method) is skipped entirely – only the p-value side runs. Several sentinels' CI search (Bartlett-approx, "rand", "rand_bootstrap", "param_boot") re-invokes the same expensive machinery as its own p-value roughly 15-40 times per bound during root-finding, by far the dominant cost of a full run_all_inference() run for those sentinels; skipping it can cut total runtime dramatically. ci_a/ci_b/ci_method stay present but always NA in results_table (stable schema either way) and are omitted entirely from the live/print/HTML display tables when FALSE. Set TRUE to compute CIs as before.

output_dir

Directory for the html/pdf/ save_results_as_JSON output files. Default "~" (the user's home directory), not the current working directory – these calls are routinely made from inside the package's own source tree (a demo/dev script run from R/EDI/), and a stray timestamped report left in getwd() there is exactly the kind of untracked file that can end up bundled into a source tarball (R CMD build) or committed by accident. Pass an explicit path (e.g. the current directory, or a temp directory) to write elsewhere.

combined_evidence_estimands

NULL (default: include every declared estimand), or a character vector of estimand values to restrict the Combined Evidence p-value/weights to. Validated argument-time against the estimand values actually declared among classes/exclude_classes-filtered candidates.

combined_evidence_weighting

One of "estimand_grouped" (default – w_i = 1 / (G * m_i)), "equal" (flat w_i = 1/k), or "custom" (caller supplies combined_evidence_weights). See inference_suite_inspect.md's TODO-15.

combined_evidence_weights

Named numeric vector (inference_class name -> weight), required when and only when combined_evidence_weighting = "custom". Names must be a subset of the classes being fit; an unnamed usable class defaults to weight 0 (excluded). Need not pre-sum to 1.

Returns

Invisibly, an object of class c("EDIInferenceSuiteResults", "list") with elements results (one named sub-list per class, in computation order), results_table (the same rows as a flat data.frame, sorted/grouped by estimand – NA_character_ last – with a secondary sort by inference_class; includes the weight column driven by combined_evidence_weighting/ combined_evidence_estimands), combined_evidence (list(pval, stat, method = "cauchy_combination", n_classes_used, n_estimand_groups, estimands_used, weighting, weights_used, classes_used) – the Cauchy-combination-test p-value/statistic across all usable rows under the resolved weighting policy; weights_used/classes_used are keyed/valued by each row's results name, not results_table$inference_class directly, since that column can repeat under formulas; pval = stat = NA_real_ if fewer than 2 rows are usable), design, alpha, unavailable_due_to_missing_packages, plots (list(ci_forest); ci_forest is a named list of one gtable grob per estimand – the CI forest stacked over its “Estimates” box-and-whisker subplot; draw with grid::grid.draw() or pass to ggplot2::ggsave() – possibly empty – rather than a single plot, since the visualization is split one-per-estimand, per user request, 2026-08-19; the former separate estimates plot became that subplot, 2026-08-21), files (list(html, pdf, json), each a path or NULL; pdf is one multi-page PDF with one page per estimand), timestamp, total_secs, and edi_version.


InferenceSuite$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSuite$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Madigan, D., Ryan, P. B., and Schuemie, M. (2013), "Does design matter? Systematic evaluation of the impact of analytical choices on effect estimates in observational studies," Therapeutic Advances in Drug Safety, 4(2), 53-62, PMID 25083251 – the motivating finding behind this class's "every legitimate way to look for an effect, compared honestly" default output.

Liu, Y. and Xie, J. (2020), "Cauchy combination test: a powerful test with analytic p-value calculation under arbitrary dependency structures," Journal of the American Statistical Association, 115(529), 393-402 – run_all_inference()'s Combined Evidence Metric.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 20, response_type = "continuous")
for (i in 1:20) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(20))

suite = InferenceSuite$new(seq_des)
suite$applicable_design_classes

# Fit and compare every applicable class:
results = suite$run_all_inference(screen = TRUE)
results$results_table


Cox Proportional Hazards Regression Inference for Survival Responses

Description

Fits a Cox proportional hazards model, \lambda(t \mid x_i) = \lambda_0(t) \exp(x_i^\top\beta), for survival responses using the treatment indicator and, optionally, all recorded covariates as predictors, by maximizing the Breslow-tie-corrected partial likelihood. For exact/right-censored data, fitting uses this package's internal Newton-Raphson C++ solver (fast_coxph_regression_prebuilt_cpp) by default (use_rcpp = TRUE), falling back to survival::coxph.fit()/survival::coxph() if that fails to converge or use_rcpp = FALSE. For left- or interval-censored data (a genuinely different likelihood, with no closed-form partial-likelihood score/information), fitting instead dispatches to icenReg::ic_sp(model = "ph"), a semiparametric NPMLE Cox fit whose standard errors come from icenReg's own internal bootstrap (not a closed-form covariance) — only Wald inference (testing_type = "wald") is supported on that path; score/gradient/likelihood-ratio/Bartlett testing types raise an informative error for such data. This is a partial-likelihood class (likelihood_tier = "partial") supporting score, gradient, and likelihood-ratio tests, plus parametric likelihood-ratio bootstrap calibration (both only for the exact/right-censored path), in addition to Wald and resampling-based inference. Fitted coefficients exceeding a fixed magnitude threshold (20, on the log-hazard-ratio scale) are treated as non-estimable (a numerical-divergence guard) rather than returned.

Super class

Inference -> InferenceSurvivalCoxPHRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalCoxPHRegr$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand. Deliberately pulled from InferenceRand, not InferenceRandCI – despite InferenceRandCI's override having a richer signature (type/args_for_type), its body calls super$...(), which resolves against this class's *actual* R6 superclass (Inference, which has no such method) once the method body is extracted and merged flatly by the component system – not against InferenceRand the way it would inside InferenceRandCI's own real inheritance chain. InferenceRandCI's only other content is an incidence-response special case (Zhang exact test) that never applies to survival data anyway, so the two are behaviorally identical for this class – confirmed by tracing the body, not assumed. Matches the pattern already used by the sibling migrated classes (LogRank/ GehanWilcox/KMDiff/RestrictedMeanDiff).

Initialize a Cox PH inference object for a completed design with a survival response. Unlike most survival inference classes in this package, this one accepts left- and interval-censored data (via an icenReg-backed fallback fit; see class documentation), not only exact/right-censored.

Usage
InferenceSurvivalCoxPHRegr$new(
  des_obj,
  model_formula = NULL,
  use_rcpp = TRUE,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with a survival response.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

use_rcpp

Logical. If TRUE (default), enable internal Rcpp score/information helpers for likelihood inference.

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values by default.


InferenceSurvivalCoxPHRegr$compute_estimate()

Computes the Cox PH treatment coefficient \hat\beta_T (log hazard ratio) — see class documentation for the fitting backend used (partial-likelihood C++/survival solver for exact/ right-censored data; icenReg NPMLE for left-/interval-censored data). NA if the fit fails or the fitted coefficients are numerically extreme.

Usage
InferenceSurvivalCoxPHRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.


InferenceSurvivalCoxPHRegr$compute_asymp_confidence_interval()

Computes an asymptotic confidence interval using the configured likelihood-backed test.

Usage
InferenceSurvivalCoxPHRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level 1 - alpha. Default 0.05.


InferenceSurvivalCoxPHRegr$compute_asymp_two_sided_pval()

Computes an asymptotic two-sided p-value using the configured likelihood-backed test.

Usage
InferenceSurvivalCoxPHRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect to test against. Default 0.


InferenceSurvivalCoxPHRegr$compute_estimate_with_bootstrap_weights()

Recomputes the Cox PH treatment estimate under subject/block bootstrap weights (via weighted_cox_bootstrap_surrogate_fit(), which assumes ordinary right-censoring semantics), used by the Bayesian bootstrap and related weighted-resampling machinery. If the weights are effectively constant, short-circuits to the unweighted $compute_estimate(estimate_only = TRUE). Not supported for left- or interval-censored data — raises an error immediately, since the surrogate weighted fit has no extension for that likelihood. Always leaves the standard error unavailable (NA) regardless of estimate_only — this weighted path never computes a variance.

Usage
InferenceSurvivalCoxPHRegr$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Subject-, block-, cluster-, or matched-set bootstrap weights.

estimate_only

Present for interface parity; this method never computes variance components regardless of its value.


InferenceSurvivalCoxPHRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalCoxPHRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Cox, D. R. (1972). "Regression Models and Life-Tables." Journal of the Royal Statistical Society, Series B, 34(2), 187-220, for the proportional hazards model and partial likelihood; Breslow, N. E. (1974). "Covariance Analysis of Censored Survival Data." Biometrics, 30(1), 89-99, doi:10.2307/2529620, for the tied-event partial-likelihood approximation used for exact/right-censored data.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalCoxPHRegr$new(seq_des)
inf$compute_estimate()


Dependent-Censoring Transformation Inference for Survival Responses

Description

Fits a joint bivariate log-normal transformation model for a latent event time T^E_i and a latent censoring time T^C_i that are allowed to be dependent (a violation of the usual independent-censoring assumption): \log T^E_i = X_i^\top \beta_{\mathrm{event}} + \sigma_{\mathrm{event}} \epsilon^E_i, \log T^C_i = X_i^\top \beta_{\mathrm{cens}} + \sigma_{\mathrm{cens}} \epsilon^C_i, with (\epsilon^E_i, \epsilon^C_i) jointly standard bivariate normal with correlation \rho (estimated via an atanh-reparameterized, clamped nuisance parameter). X_i includes the treatment indicator W_i as its first column, so \hat\beta_T (the first entry of \hat\beta_{\mathrm{event}}) is a log-time-ratio for the event submodel, on the same AFT interpretation scale as InferenceSurvivalWeibullRegr but log-normal rather than Weibull, and jointly modeling the censoring mechanism rather than assuming it independent. This is the correct tool when censoring is suspected to depend on the same latent factors driving the event time (e.g. sicker patients are both more likely to be censored — dropout — and more likely to fail early), a scenario under which ordinary Kaplan-Meier/Cox/AFT methods (which assume independent censoring) are biased. likelihood_tier = "full": likelihood-ratio, score, gradient, and Wald tests are available when the model converges, plus parametric-likelihood-bootstrap calibration of the likelihood-ratio test. Substantial method-support limitations, all deliberate: randomization inference and jackknife bias correction/standard errors are hard-unsupported (each randomization draw would require a full dependent-censoring likelihood refit, too unstable/slow for the comprehensive test suite; jackknife bias correction is unstable for this likelihood on small censored samples) — every jackknife/randomization method returns NA and marks the result nonestimable rather than computing a value. Nonparametric-bootstrap confidence intervals are computed but additionally validated/sanity-checked (excessively wide or zero-excluding-by-construction intervals are treated as unstable and replaced with NA), and Bayesian-bootstrap weighted re-estimation uses a fast Cox-model surrogate fit (weighted_cox_bootstrap_surrogate_fit()) as an approximation rather than a full weighted joint-likelihood refit.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Value

A single logical.

Super class

Inference -> InferenceSurvivalDepCensTransformRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalDepCensTransformRegr$supports_rand_pval_for_incidence()

Usage
InferenceSurvivalDepCensTransformRegr$supports_rand_pval_for_incidence()

InferenceSurvivalDepCensTransformRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalDepCensTransformRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

See Also

InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik for a different (Clayton-copula, KK-design) approach to dependence between two survival-type quantities. Survival analysis (Wikipedia, general orientation; no direct Wikipedia page for dependent-censoring copula/transformation models specifically).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalDepCensTransformRegr$new(seq_des)
inf$compute_estimate()


Clayton Copula / Standard Weibull Compound Inference for KK Designs

Description

This class implements a compound estimator for KK matching-on-the-fly designs with survival responses using a Clayton copula with Weibull AFT margins for matched pairs and a standard Weibull AFT model for the reservoir. The two treatment-effect estimates (on the log-time ratio scale) are combined by inverse-variance weighting.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

Frailty distribution. The Clayton copula for a matched pair, S(t1,t2) = (S1(t1)^-theta + S2(t2)^-theta - 1)^(-1/theta) with Weibull margins S_i, is exactly the closed-form bivariate survival function obtained by multiplying two conditionally-independent Weibull hazards by a shared gamma frailty term Z ~ Gamma(1/theta, 1/theta) and integrating Z out analytically (Clayton 1978; Oakes 1989); theta is the frailty variance / dependence parameter (see ClaytonWeibullLikelihood in fast_survival_models_optim.cpp, which builds the likelihood from the per-subject Weibull cumulative hazards H1, H2). This is the classic textbook Weibull-gamma shared-frailty model, fit here in its closed form (no numerical integration required) rather than as an AFT Gaussian-random-intercept model.

This is a different (and equally standard) frailty assumption from the log-normal-frailty Weibull AFT GLMM implemented by InferenceSurvivalGLMMWeibullFrailtyNormalIVWC / InferenceSurvivalGLMMWeibullFrailtyNormalOneLik, which instead places a Gaussian random intercept on the log-time (AFT) scale and integrates it out by Gauss-Hermite quadrature. Prefer this Clayton-copula class for the classic gamma-frailty / proportional-hazards dependence structure; prefer the Weibull-frailty class for a Gaussian-random-intercept / GLMM-style dependence structure.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$supports_rand_pval_for_incidence()

Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$supports_rand_pval_for_incidence(
  
)

InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$compute_rand_two_sided_pval()

Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Clayton DG (1978). "A Model for Association in Bivariate Life Tables and Its Application in Epidemiological Studies of Familial Tendency in Chronic Disease Incidence." Biometrika, 65(1), 141-151. doi:10.2307/2335289

Oakes D (1989). "Bivariate Survival Models Induced by Frailties." Journal of the American Statistical Association, 84(406), 487-493. doi:10.2307/2289934

Legacy class. Not fully tested in comprehensive_tests.R.

See Also

InferenceSurvivalGLMMWeibullFrailtyNormalIVWC for the corresponding log-normal-frailty IVWC estimator.


One-Likelihood Clayton-Copula Weibull AFT Inference for KK Survival Designs

Description

Estimates a treatment log-time-ratio \beta_T for right-censored survival outcomes collected under a KK matching-on-the-fly design (DesignSeqOneByOneKK14 or subclass) by maximizing a single combined likelihood: matched-pair survival times are modeled with a Weibull accelerated-failure-time (AFT) margin joined by a Clayton copula (dependence parameter \theta) to account for within-pair correlation induced by shared matching covariates, while unmatched reservoir subjects are modeled by the same Weibull AFT margin marginally (no dependence term). All subjects share one treatment coefficient, estimated jointly.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

Estimand. \beta_T, the treatment coefficient of a Weibull AFT model \log T = \beta_0 + \beta_T W + X\beta + \sigma\epsilon with \epsilon extreme-value-distributed; \exp(\hat\beta_T) is the treatment-vs-control survival-time ratio (an acceleration factor). This is a distinct scale from the log hazard ratio reported by Cox-based KK survival classes.

Model. .fit_clayton_weibull_aft() jointly optimizes the AFT regression coefficients, the Weibull shape (\log\sigma), and the Clayton copula dependence parameter (\log\theta) by direct maximum likelihood over the combined matched-pair-copula / reservoir-marginal log-likelihood; right-censoring enters as the usual survival contribution (density for observed failures, survival function for censored times). likelihood_tier = "full", so a parametric likelihood bootstrap (simulate_under_lik_null, which draws new pair times from the fitted Clayton copula and new singleton times from the marginal Weibull) is available alongside Wald inference.

Assumptions. Weibull AFT margin correctly specified; Clayton copula correctly captures within-pair dependence (a positive-dependence, single-parameter Archimedean copula); independent censoring given covariates; a KK matching-on-the-fly design supplying the matched/ reservoir partition.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$supports_rand_pval_for_incidence()

Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$supports_rand_pval_for_incidence(
  
)

InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$compute_rand_two_sided_pval()

Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Clayton, D. G. (1978). "A model for association in bivariate life tables and its application in epidemiological studies of familial tendency in chronic disease incidence." Biometrika, 65(1), 141-151. doi:10.1093/biomet/65.1.141. (Clayton1978 in REFERENCES.md.)

Oakes, D. (1989). "Bivariate survival models induced by frailties." Journal of the American Statistical Association, 84(406), 487-493. doi:10.1080/01621459.1989.10478795. (Oakes1989 in REFERENCES.md.)

See Also

Analogous Python API for AFT/copula survival models: lifelines WeibullAFTFitter, copulas. Copula (probability theory) (orientation).


Weibull Frailty IVWC Inference for KK Designs

Description

Log-normal (Gaussian random-intercept) frailty Weibull AFT estimator; see InferenceSurvivalGLMMWeibullFrailtyNormalIVWC for the frailty-distribution details and contrast with the gamma-frailty InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC (Clayton copula) alternative.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

Legacy class. Not fully tested in comprehensive_tests.R.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceSurvivalGLMMWeibullFrailtyNormalIVWC

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$supports_rand_pval_for_incidence()

Usage
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$supports_rand_pval_for_incidence(
  
)

InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$compute_rand_two_sided_pval()

Usage
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Weibull Frailty Combined-Likelihood Inference for KK Designs

Description

Log-normal (Gaussian random-intercept) frailty Weibull AFT estimator; see InferenceSurvivalGLMMWeibullFrailtyNormalOneLik for the frailty-distribution details and contrast with the gamma-frailty InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik (Clayton copula) alternative.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceSurvivalGLMMWeibullFrailtyNormalOneLik

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$supports_rand_pval_for_incidence()

Usage
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$supports_rand_pval_for_incidence(
  
)

InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$compute_rand_two_sided_pval()

Usage
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.


Gehan-Wilcoxon (Peto-Prentice) Inference for Survival Data with Censoring

Description

Non-parametric inference for survival outcomes supporting censored data, using the Peto-Prentice modification of the Gehan-Wilcoxon test. The treatment effect estimate is the mean difference in Peto-Prentice weighted martingale residuals between the treatment and control groups. Specifically, for each subject the weighted residual is M_i^w = \hat{S}(t_i^-) \cdot M_i, where M_i = \delta_i - \hat\Lambda_0(t_i) is the martingale residual and \hat{S}(t_i^-) is the overall Kaplan-Meier survival estimate just before time t_i. These weights downweight late events, analogously to the Wilcoxon rank-sum test for uncensored data (which also weights early observations more heavily via their larger rank denominator).

The p-value uses survival::survdiff(rho = 1) (Peto-Prentice / Fleming-Harrington p=1, q=0), which is distinct from the log-rank test (rho = 0) used in InferenceSurvivalKMDiff.

Super class

Inference -> InferenceSurvivalGehanWilcox

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalGehanWilcox$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize Gehan-Wilcoxon survival inference and prepare the rank-based treatment statistic used by InferenceSurvivalGehanWilcox.

Usage
InferenceSurvivalGehanWilcox$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE
)
Arguments
des_obj

The design object.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

If TRUE, print additional information.


InferenceSurvivalGehanWilcox$compute_estimate()

Returns the mean difference in Peto-Prentice weighted martingale residuals between the treatment and control groups. Positive values indicate that treatment subjects experienced fewer early events than expected. For left- or interval-censored data, dispatches instead to interval::ictest(..., scores = "wmw") (the Wilcoxon-Mann-Whitney interval-censored generalization of the Peto-Prentice test) and returns its estimate.

Usage
InferenceSurvivalGehanWilcox$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.

Returns

A numeric scalar (the Peto-Prentice weighted score treatment effect estimate).

Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2:10])
seq_des$add_all_subject_responses(
  ys = c(4.71, NA, 4.78, 6.11, NA, 8.43),
  y_Ls = c(NA, 1.23, NA, NA, 5.95, NA),
  y_Rs = c(NA, Inf, NA, NA, Inf, NA)
)

seq_des_inf = InferenceSurvivalGehanWilcox$new(seq_des)
seq_des_inf$compute_estimate()

InferenceSurvivalGehanWilcox$compute_estimate_with_bootstrap_weights()

Recomputes the class-specific treatment estimate under bootstrap weights; see InferenceBayesianBootstrap.

Usage
InferenceSurvivalGehanWilcox$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip variance calculations.


InferenceSurvivalGehanWilcox$compute_asymp_confidence_interval()

Computes a (1 - alpha)-level confidence interval based on the asymptotic normality of the Peto-Prentice weighted martingale residual mean difference. Falls back to bootstrap if the SE is unavailable.

Usage
InferenceSurvivalGehanWilcox$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level. Default is 0.05.

Returns

A numeric vector of length 2: (lower, upper) confidence bounds.

Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2:10])
seq_des$add_all_subject_responses(
  ys = c(4.71, NA, 4.78, 6.11, NA, 8.43),
  y_Ls = c(NA, 1.23, NA, NA, 5.95, NA),
  y_Rs = c(NA, Inf, NA, NA, Inf, NA)
)

seq_des_inf = InferenceSurvivalGehanWilcox$new(seq_des)
seq_des_inf$compute_asymp_confidence_interval()

InferenceSurvivalGehanWilcox$compute_asymp_two_sided_pval()

Computes the Peto-Prentice (Gehan-Wilcoxon) two-sided p-value via survival::survdiff(rho = 1), which puts greater weight on early events relative to the standard log-rank test (rho = 0). For delta != 0, not yet implemented.

Usage
InferenceSurvivalGehanWilcox$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect to test against. Default is 0.

Returns

A p-value in [0, 1].

Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2:10])
seq_des$add_all_subject_responses(
  ys = c(4.71, NA, 4.78, 6.11, NA, 8.43),
  y_Ls = c(NA, 1.23, NA, NA, 5.95, NA),
  y_Rs = c(NA, Inf, NA, NA, Inf, NA)
)

seq_des_inf = InferenceSurvivalGehanWilcox$new(seq_des)
seq_des_inf$compute_asymp_two_sided_pval()

InferenceSurvivalGehanWilcox$compute_rand_confidence_interval()

Randomization confidence intervals are not supported for this class because the Peto-Prentice weighted score scale is not commensurate with the time-ratio null used by the randomization CI bisection algorithm.

Usage
InferenceSurvivalGehanWilcox$compute_rand_confidence_interval(
  alpha = 0.05,
  r = 501,
  pval_epsilon = 0.005,
  show_progress = TRUE,
  ci_search_control = NULL
)
Arguments
alpha

Unused.

r

Unused.

pval_epsilon

Unused.

show_progress

Unused.

ci_search_control

Unused.


InferenceSurvivalGehanWilcox$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalGehanWilcox$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Gehan, E. A. (1965). "A generalized Wilcoxon test for comparing arbitrarily singly-censored samples." Biometrika, 52(1-2), 203-223, doi:10.1093/biomet/52.1-2.203, for the original generalized (Gehan) Wilcoxon test for censored data. Peto, R., and Peto, J. (1972). "Asymptotically Efficient Rank Invariant Test Procedures." Journal of the Royal Statistical Society, Series A, 135(2), 185-207, doi:10.2307/2344317, for the survival-weighted (Peto-Prentice) modification this class implements via \rho=1.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalGehanWilcox$new(seq_des)
inf$compute_estimate()


## ------------------------------------------------
## Method `InferenceSurvivalGehanWilcox$compute_estimate()`
## ------------------------------------------------

seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2:10])
seq_des$add_all_subject_responses(
  ys = c(4.71, NA, 4.78, 6.11, NA, 8.43),
  y_Ls = c(NA, 1.23, NA, NA, 5.95, NA),
  y_Rs = c(NA, Inf, NA, NA, Inf, NA)
)

seq_des_inf = InferenceSurvivalGehanWilcox$new(seq_des)
seq_des_inf$compute_estimate()

## ------------------------------------------------
## Method `InferenceSurvivalGehanWilcox$compute_asymp_confidence_interval()`
## ------------------------------------------------

seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2:10])
seq_des$add_all_subject_responses(
  ys = c(4.71, NA, 4.78, 6.11, NA, 8.43),
  y_Ls = c(NA, 1.23, NA, NA, 5.95, NA),
  y_Rs = c(NA, Inf, NA, NA, Inf, NA)
)

seq_des_inf = InferenceSurvivalGehanWilcox$new(seq_des)
seq_des_inf$compute_asymp_confidence_interval()

## ------------------------------------------------
## Method `InferenceSurvivalGehanWilcox$compute_asymp_two_sided_pval()`
## ------------------------------------------------

seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2:10])
seq_des$add_all_subject_responses(
  ys = c(4.71, NA, 4.78, 6.11, NA, 8.43),
  y_Ls = c(NA, 1.23, NA, NA, 5.95, NA),
  y_Rs = c(NA, Inf, NA, NA, Inf, NA)
)

seq_des_inf = InferenceSurvivalGehanWilcox$new(seq_des)
seq_des_inf$compute_asymp_two_sided_pval()

LWA-style Marginal Cox IVWC Compound Inference for KK Designs

Description

Fits a compound (IVWC) estimator for KK matching-on-the-fly designs with survival responses: matched pairs are analyzed with a marginal Cox model \lambda(t \mid w) = \lambda_0(t)\exp(\beta_T w) whose robust variance uses the Lee-Wei-Amato (1992) cluster-robust sandwich (treating each matched pair as an independent cluster of correlated failure times), while reservoir subjects are analyzed with a standard (independent-subjects) Cox partial likelihood; the two log-hazard-ratio estimates are then combined by inverse-variance weighting. likelihood_tier = "partial" (Cox partial likelihood), but likelihood-ratio/score/gradient tests are not exposed on this IVWC compound (only on the OneLik sibling, which fits one combined partial likelihood across both sources instead of pooling two separate fits).

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

Legacy class. Not fully tested in comprehensive_tests.R.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceSurvivalKKLWACoxPHIVWC

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalKKLWACoxPHIVWC$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceSurvivalKKLWACoxPHIVWC$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalKKLWACoxPHIVWC$supports_rand_pval_for_incidence()

Usage
InferenceSurvivalKKLWACoxPHIVWC$supports_rand_pval_for_incidence()

InferenceSurvivalKKLWACoxPHIVWC$compute_rand_two_sided_pval()

Usage
InferenceSurvivalKKLWACoxPHIVWC$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalKKLWACoxPHIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalKKLWACoxPHIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Lee, E. W., Wei, L. J., and Amato, D. A. (1992). "Cox-Type Regression Analysis for Large Numbers of Small Groups of Correlated Failure Time Observations." In Survival Analysis: State of the Art, 237-247. Springer. doi:10.1007/978-94-015-7983-4_14

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'survival')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalKKLWACoxPHIVWC$new(seq_des)
inf$compute_estimate()


LWA-style Marginal Cox Combined-Likelihood Inference for KK Designs

Description

Fits a single combined Cox partial likelihood \lambda(t \mid w, x) = \lambda_0(t)\exp(\beta_T w + \beta_X^\top x) jointly over matched-pair and reservoir subjects for KK matching-on-the-fly designs with survival responses (a marginal, not stratified, Cox model: matched pairs do not get pair-specific baseline hazards). This is the one-likelihood combined-fit analog of InferenceSurvivalKKLWACoxPHIVWC, which instead fits and pools two separate estimators. likelihood_tier = "partial": exposes likelihood-ratio and parametric-likelihood-bootstrap inference in addition to Wald/asymptotic and Bayesian-bootstrap paths.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceSurvivalKKLWACoxPHOneLik

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalKKLWACoxPHOneLik$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceSurvivalKKLWACoxPHOneLik$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalKKLWACoxPHOneLik$supports_rand_pval_for_incidence()

Usage
InferenceSurvivalKKLWACoxPHOneLik$supports_rand_pval_for_incidence()

InferenceSurvivalKKLWACoxPHOneLik$compute_rand_two_sided_pval()

Usage
InferenceSurvivalKKLWACoxPHOneLik$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalKKLWACoxPHOneLik$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalKKLWACoxPHOneLik$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Cox, D. R. (1972). "Regression Models and Life-Tables." Journal of the Royal Statistical Society, Series B, 34(2), 187-220.

Examples


seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'survival')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalKKLWACoxPHOneLik$new(seq_des)
inf$compute_estimate()


Rank Regression Inference for Survival Responses under KK Designs

Description

Fits a multivariate Gehan-Wilcoxon rank regression for survival outcomes under a KK matching-on-the-fly design. The model adjusts for the treatment indicator and, optionally, all recorded covariates.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

Legacy class. Not fully tested in comprehensive_tests.R.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceSurvivalKKRankRegrIVWC

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalKKRankRegrIVWC$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceSurvivalKKRankRegrIVWC$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalKKRankRegrIVWC$supports_rand_pval_for_incidence()

Usage
InferenceSurvivalKKRankRegrIVWC$supports_rand_pval_for_incidence()

InferenceSurvivalKKRankRegrIVWC$compute_rand_two_sided_pval()

Usage
InferenceSurvivalKKRankRegrIVWC$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalKKRankRegrIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalKKRankRegrIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'survival')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalKKRankRegrIVWC$new(seq_des)
inf$compute_estimate()

Stratified Cox / Standard Cox Compound Inference for KK Designs

Description

This class implements a compound estimator for KK matching-on-the-fly designs with survival responses. For matched pairs, it uses stratified Cox proportional hazards regression (each pair is a stratum). For reservoir subjects, it uses standard Cox regression. The two estimates (both log-hazard ratios) are combined via a variance-weighted linear combination.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Details

Under harden = TRUE, multivariate fits preserve the treatment column and progressively retry reduced covariate sets after QR-based rank reduction and correlation-based pruning. Extreme finite coefficients / standard errors are rejected and treated as non-estimable.

The matched-pair sub-estimate treats each pair as its own stratum (a pair-specific baseline hazard, exactly canceling shared-frailty effects within the pair via Cox's partial likelihood) and is a special case of the Lee-Wei-Amato (1992) large-numbers-of-small-groups stratified Cox approach; the reservoir sub-estimate is a standard unstratified Cox partial-likelihood fit (Cox 1972). The two log-hazard-ratio estimates are combined by inverse-variance weighting, the same rule used throughout the KK IVWC family.

Legacy class. Not fully tested in comprehensive_tests.R.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceSurvivalKKStratCoxPHIVWC

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalKKStratCoxPHIVWC$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceSurvivalKKStratCoxPHIVWC$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalKKStratCoxPHIVWC$supports_rand_pval_for_incidence()

Usage
InferenceSurvivalKKStratCoxPHIVWC$supports_rand_pval_for_incidence()

InferenceSurvivalKKStratCoxPHIVWC$compute_rand_two_sided_pval()

Usage
InferenceSurvivalKKStratCoxPHIVWC$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalKKStratCoxPHIVWC$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalKKStratCoxPHIVWC$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Cox, D. R. (1972). "Regression Models and Life-Tables." Journal of the Royal Statistical Society, Series B, 34(2), 187-220.

Lee, E. W., Wei, L. J., and Amato, D. A. (1992). "Cox-Type Regression Analysis for Large Numbers of Small Groups of Correlated Failure Time Observations." In Survival Analysis: State of the Art, 237-247. Springer. doi:10.1007/978-94-015-7983-4_14


Stratified Cox Combined-Likelihood Compound Inference for KK Designs

Description

Fits a single combined partial likelihood for KK matching-on-the-fly designs with survival responses: matched pairs contribute a stratified Cox term (one stratum per pair, canceling shared-frailty effects within the pair) and reservoir subjects contribute a standard unstratified Cox term, summed into one joint partial log-likelihood and optimized jointly for a single shared treatment coefficient. This differs from the two-stage IVWC sibling InferenceSurvivalKKStratCoxPHIVWC, which fits the matched and reservoir sub-models separately and combines the two log-hazard-ratio estimates by inverse-variance weighting; this class instead estimates one coefficient from the combined likelihood directly, which additionally supports likelihood-ratio tests and parametric likelihood bootstrap (likelihood_tier = "partial").

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Computes a randomization-based p-value.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Randomization p-value.

Super class

Inference -> InferenceSurvivalKKStratCoxPHOneLik

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalKKStratCoxPHOneLik$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceSurvivalKKStratCoxPHOneLik$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalKKStratCoxPHOneLik$supports_rand_pval_for_incidence()

Usage
InferenceSurvivalKKStratCoxPHOneLik$supports_rand_pval_for_incidence()

InferenceSurvivalKKStratCoxPHOneLik$compute_rand_two_sided_pval()

Usage
InferenceSurvivalKKStratCoxPHOneLik$compute_rand_two_sided_pval(
  r = 501,
  delta = 0,
  transform_responses = "none",
  na.rm = TRUE,
  show_progress = TRUE,
  permutations = NULL,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

r

Number of randomization vectors.

delta

The null difference. Default 0.

delta

Null difference.

transform_responses

Type of transformation. Default "none".

transform_responses

Transformation.

na.rm

Remove NAs.

show_progress

Show progress bar. Default TRUE.

show_progress

Show progress.

permutations

Pre-computed permutations. Default NULL.

permutations

Pre-computed permutations.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalKKStratCoxPHOneLik$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalKKStratCoxPHOneLik$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Cox, D. R. (1972). "Regression Models and Life-Tables." Journal of the Royal Statistical Society, Series B, 34(2), 187-220.

Lee, E. W., Wei, L. J., and Amato, D. A. (1992). "Cox-Type Regression Analysis for Large Numbers of Small Groups of Correlated Failure Time Observations." In Survival Analysis: State of the Art, 237-247. Springer. doi:10.1007/978-94-015-7983-4_14


Marginal (Cluster-Robust) Weibull Inference for KK Matched-Pair Survival Designs

Description

Initialize the marginal (cluster-robust) Weibull inference object.

Returns the pooled treatment effect estimate (log-time-ratio scale).

Recomputes the treatment estimate under Bayesian-bootstrap weights.

Computes the asymptotic (cluster-robust) confidence interval.

Computes the asymptotic (cluster-robust) two-sided p-value.

Duplicates this subclass while preserving fit caches; see Inference.

Usage

SurvivalKKWeibullMarginalSource

Details

Fits a single pooled Weibull Accelerated Failure Time (AFT) model across all subjects (treatment plus, optionally, all recorded covariates), ignoring the matched-pair structure in the mean model. Standard errors are computed via a cluster-robust (sandwich) covariance estimator: matched pairs from a KK matching-on-the-fly or binary-match design form size-2 clusters, and unmatched reservoir subjects each form their own singleton cluster.

This is the "marginal" competitor to InferenceSurvivalGLMMWeibullFrailtyNormalOneLik: rather than modeling the within-pair correlation explicitly via a frailty term, it fits an ordinary (working-independence) Weibull AFT model and corrects the treatment-effect standard error post hoc for the within-pair dependence. The model is fit via the package's fast C++ Weibull AFT backend (fast_weibull_regression_general_cpp) and the cluster-robust sandwich is assembled from per-subject dfbeta contributions collapsed within clusters, which is numerically equivalent to survival::survreg(..., cluster = ..., robust = TRUE) (retained as a fallback if the C++ fit fails to converge).

Examples


des = DesignSeqOneByOneKK14$new(n = 20, response_type = 'survival')
for (i in 1:20) {
  des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
des$add_all_subject_responses(rexp(20))
inf = InferenceSurvivalKKWeibullMarginal$new(des)
inf$compute_estimate()


Kaplan-Meier Median-Difference Inference for Survival Responses

Description

Fits a non-parametric treatment-effect estimator for censored survival responses: the difference in Kaplan-Meier median survival times between the treated and control arms, \hat m_T - \hat m_C. Standard errors are obtained by back-calculating from each arm's separate Brookmeyer-Crowley confidence interval for its own median (via survival::survfit's default log-log-transformed interval): \hat\sigma_i = (\mathrm{upper}_i - \mathrm{lower}_i) / (2 z_{\alpha/2}), combined (the two arms are independent by design) as \sqrt{\hat\sigma_T^2 + \hat\sigma_C^2}. When either arm's median is inestimable (its Kaplan-Meier curve never reaches 0.5) or the back-calculated bounds are non-finite, the Wald-style confidence interval and p-value methods ($compute_asymp_confidence_interval(), $compute_asymp_two_sided_pval()) silently fall back to a nonparametric bootstrap instead of returning NA. A convenience method, $compute_asymp_log_rank_two_sided_pval_for_treatment_effect(), is also provided for the log-rank p-value on the same fitted survival curves. For left- or interval-censored data, the point estimate instead comes from a Turnbull NPMLE median contrast (interval::icfit(), via turnbull_npmle_stat_diff()), which has no closed-form standard error — inference on that path relies entirely on the bootstrap fallback described above. Randomization confidence intervals are not supported (the median difference's units are not commensurate with the randomization CI bisection algorithm's transformed-scale null search).

Super class

Inference -> InferenceSurvivalKMDiff

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalKMDiff$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize Kaplan-Meier median-difference survival inference and prepare treatment-group survival curves used by InferenceSurvivalKMDiff.

Usage
InferenceSurvivalKMDiff$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

The design object.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

If TRUE, print additional information.

smart_cold_start_default

Whether to use smart cold start values by default.


InferenceSurvivalKMDiff$compute_estimate()

Computes the class-specific mean or survival contrast; see InferenceMLEorKMSummaryTable.

Usage
InferenceSurvivalKMDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.

Returns

The setting-appropriate (see description) numeric estimate of the treatment effect

Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10])
seq_des$add_all_subject_responses(
  ys = c(4.71, NA, 4.78, 6.11, NA, 8.43),
  y_Ls = c(NA, 1.23, NA, NA, 5.95, NA),
  y_Rs = c(NA, Inf, NA, NA, Inf, NA)
)

seq_des_inf = InferenceSurvivalKMDiff$new(seq_des)
seq_des_inf$compute_estimate()

InferenceSurvivalKMDiff$compute_estimate_with_bootstrap_weights()

Recomputes the class-specific treatment estimate for a bootstrap sample; see InferenceNonParamBootstrap.

Usage
InferenceSurvivalKMDiff$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Row weights for the bootstrap sample.

estimate_only

If TRUE, skip variance calculations.


InferenceSurvivalKMDiff$compute_asymp_confidence_interval()

Computes a (1 - alpha)-level confidence interval for the difference in Kaplan-Meier median survival times (treatment minus control).

The Brookmeyer-Crowley confidence interval is obtained for each group's median separately via survival::survfit (using a log-log transformation of the survival function by default). The per-group SE is back-calculated from the CI half-width as \hat\sigma_i = (\text{upper}_i - \text{lower}_i) / (2 z_{\alpha/2}). The two groups are independent by design, so the SE of the difference is \sqrt{\hat\sigma_T^2 + \hat\sigma_C^2}, and the CI is (\hat{m}_T - \hat{m}_C) \pm z_{\alpha/2} \cdot \sqrt{\hat\sigma_T^2 + \hat\sigma_C^2}.

Falls back to compute_bootstrap_confidence_interval when either group's median is not estimable (i.e., the Kaplan-Meier curve does not reach 0.5) or when the Brookmeyer-Crowley CI bounds are NA.

Usage
InferenceSurvivalKMDiff$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

The significance level; the confidence level is 1 - alpha. Default is 0.05.

Returns

A numeric vector of length 2 giving the (lower, upper) confidence bounds for the difference in median survival times, on the original time scale.


InferenceSurvivalKMDiff$compute_asymp_two_sided_pval()

Computes a Wald-style 2-sided p-value based on the median difference and its back-calculated standard error.

Usage
InferenceSurvivalKMDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null difference to test against. Default is 0.

Returns

The approximate frequentist p-value


InferenceSurvivalKMDiff$compute_asymp_log_rank_two_sided_pval_for_treatment_effect()

Computes a 2-sided p-value via the log rank test

Usage
InferenceSurvivalKMDiff$compute_asymp_log_rank_two_sided_pval_for_treatment_effect(
  delta = 0
)
Arguments
delta

The null difference to test against. For any treatment effect at all this is set to zero (the default).

Returns

The approximate frequentist p-value


InferenceSurvivalKMDiff$compute_rand_confidence_interval()

Uses the shared randomization confidence-interval contract; see InferenceRandCI.

Usage
InferenceSurvivalKMDiff$compute_rand_confidence_interval(
  alpha = 0.05,
  r = 501,
  pval_epsilon = 0.005,
  show_progress = TRUE,
  ci_search_control = NULL
)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.

r

The number of randomization vectors. The default is 501.

pval_epsilon

The bisection algorithm tolerance. The default is 0.005.

show_progress

Show a text progress indicator.

ci_search_control

Unused.

Returns

A 1 - alpha sized frequentist confidence interval


InferenceSurvivalKMDiff$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalKMDiff$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kaplan, E. L., and Meier, P. (1958). "Nonparametric Estimation from Incomplete Observations." Journal of the American Statistical Association, 53(282), 457-481, doi:10.2307/2281868, for the Kaplan-Meier survival curve estimator each arm's median is read from. Brookmeyer, R., and Crowley, J. (1982). "A Confidence Interval for the Median Survival Time." Biometrics, 38(1), 29-41, doi:10.2307/2530286, for the per-arm median confidence interval this class's standard error is back-calculated from. Turnbull, B. W. (1976). "The Empirical Distribution Function with Arbitrarily Grouped, Censored and Truncated Data." Journal of the Royal Statistical Society, Series B, 38(3), 290-295, doi:10.1111/j.2517-6161.1976.tb01597.x, for the NPMLE used on the left-/interval-censored path.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalKMDiff$new(seq_des)
inf$compute_estimate()


## ------------------------------------------------
## Method `InferenceSurvivalKMDiff$compute_estimate()`
## ------------------------------------------------

seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10])
seq_des$add_all_subject_responses(
  ys = c(4.71, NA, 4.78, 6.11, NA, 8.43),
  y_Ls = c(NA, 1.23, NA, NA, 5.95, NA),
  y_Rs = c(NA, Inf, NA, NA, Inf, NA)
)

seq_des_inf = InferenceSurvivalKMDiff$new(seq_des)
seq_des_inf$compute_estimate()


Log-Rank Inference for Survival Data with Censoring

Description

Non-parametric all-subject inference for survival outcomes supporting right censoring, based on the standard two-sample log-rank test. The treatment effect estimate is the difference in mean martingale residuals between the treatment and control groups under the pooled null hazard. The p-value uses the classic log-rank score statistic with its hypergeometric tie-adjusted variance. For left- or interval-censored data, this class dispatches instead to interval::ictest(..., scores = "logrank1") (Sun's-scores interval-censored generalization of the log-rank test) for both the point estimate and testing.

Super class

Inference -> InferenceSurvivalLogRank

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalLogRank$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize log-rank survival inference and prepare the treatment-group survival data used by InferenceSurvivalLogRank.

Usage
InferenceSurvivalLogRank$new(des_obj, model_formula = NULL, verbose = FALSE)
Arguments
des_obj

The design object.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

If TRUE, print additional information.


InferenceSurvivalLogRank$compute_estimate()

Computes the treatment-effect estimate on the martingale-residual mean-difference scale. Under left-/interval-censored data, dispatched instead through interval::ictest()'s Sun's-scores log-rank test (TODO-7, interval_censored_survival_response.md); its estimate field (mean score difference between groups) is on the same "difference of group-mean scores" scale as the right-censored martingale-residual difference this method otherwise returns.

Usage
InferenceSurvivalLogRank$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.


InferenceSurvivalLogRank$compute_estimate_with_bootstrap_weights()

Recomputes the class-specific treatment estimate under bootstrap weights; see InferenceBayesianBootstrap.

Usage
InferenceSurvivalLogRank$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Bootstrap weights at the subject or block level.

estimate_only

If TRUE, skip variance calculations.


InferenceSurvivalLogRank$compute_asymp_confidence_interval()

Computes a (1 - alpha)-level confidence interval based on the asymptotic normality of the martingale-residual mean-difference estimate. Falls back to bootstrap if the estimated standard error is unavailable.

Usage
InferenceSurvivalLogRank$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level.


InferenceSurvivalLogRank$compute_asymp_two_sided_pval()

Computes a Wald-style 2-sided p-value by inverting the confidence interval.

Usage
InferenceSurvivalLogRank$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null difference to test against. Default is 0.

Returns

The approximate frequentist p-value


InferenceSurvivalLogRank$compute_asymp_log_rank_two_sided_pval_for_treatment_effect()

Computes the standard two-sided log-rank p-value for a zero treatment effect. Under left-/interval-censored data, this is interval::ictest()'s own p-value (TODO-7) rather than a re-derived chi-squared statistic.

Usage
InferenceSurvivalLogRank$compute_asymp_log_rank_two_sided_pval_for_treatment_effect(
  delta = 0
)
Arguments
delta

Null treatment effect to test against. Only 0 is supported.


InferenceSurvivalLogRank$compute_rand_confidence_interval()

Randomization confidence intervals are not supported for this class because the martingale-residual score scale is not commensurate with the transformed time-ratio null used by the randomization CI algorithm.

Usage
InferenceSurvivalLogRank$compute_rand_confidence_interval(
  alpha = 0.05,
  r = 501,
  pval_epsilon = 0.005,
  show_progress = TRUE,
  ci_search_control = NULL
)
Arguments
alpha

Unused.

r

Unused.

pval_epsilon

Unused.

show_progress

Unused.

ci_search_control

Unused.


InferenceSurvivalLogRank$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalLogRank$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Mantel, N. (1966). "Evaluation of survival data and two new rank order statistics arising in its consideration." Cancer Chemotherapy Reports, 50(3), 163-170, for the log-rank test. Peto, R., and Peto, J. (1972). "Asymptotically Efficient Rank Invariant Test Procedures." Journal of the Royal Statistical Society, Series A, 135(2), 185-207, doi:10.2307/2344317, for its asymptotic-efficiency properties and the \rho=0 case of the Fleming-Harrington family this class corresponds to (see also InferenceSurvivalGehanWilcox for the \rho=1 member of the same family). Sun, J. (1996). "A non-parametric test for interval-censored failure time data with application to AIDS studies." Statistics in Medicine, 15(13), 1387-1395, for the interval-censored generalization used on the left-/interval-censored path.

Examples

set.seed(1)
x_dat <- data.frame(
  x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
  x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneBernoulli$
  new(
  n = nrow(x_dat),
  response_type = "survival",
  verbose = FALSE
)
for (i in seq_len(nrow(x_dat))) {
  seq_des$
  add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$
  add_all_subject_responses(
  ys = c(1.2, 2.4, NA, 3.1, NA, 4.0, 3.3, NA),
  y_Ls = c(NA, NA, 1.8, NA, 2.7, NA, NA, 4.5),
  y_Rs = c(NA, NA, Inf, NA, Inf, NA, NA, Inf)
)
infer <- InferenceSurvivalLogRank$
  new(
  seq_des,
  verbose = FALSE
)
infer


Restricted Mean Survival Time (RMST) Difference Inference for Survival Responses

Description

Fits a non-parametric treatment-effect estimator for censored survival responses: the difference in restricted mean survival time (RMST) between the treated and control arms, \hat\mu_T(\tau) - \hat\mu_C(\tau), where each arm's RMST is the area under its Kaplan-Meier survival curve up to a truncation horizon \tau (\hat\mu(\tau) = \int_0^\tau \hat S(t)\,dt), computed by trapezoidal integration of the step-function KM curve. The standard error of the difference comes from the Greenwood-type variance of each arm's RMST, combined across the two (independent) arms via get_restricted_mean_se_diff(). When that standard error is unavailable or non-finite, $compute_asymp_confidence_interval() falls back to a nonparametric bootstrap interval rather than returning NA. Randomization confidence intervals are not supported (the RMST-difference units are not commensurate with the randomization CI bisection algorithm's transformed-scale null search).

Super class

Inference -> InferenceSurvivalRestrictedMeanDiff

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalRestrictedMeanDiff$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Initialize restricted-mean-survival-time difference inference and prepare treatment-group survival summaries used by InferenceSurvivalRestrictedMeanDiff.

Usage
InferenceSurvivalRestrictedMeanDiff$new(
  des_obj,
  model_formula = NULL,
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

The design object.

model_formula

Optional formula for covariate adjustment. If NULL (default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.

verbose

If TRUE, print additional information.

smart_cold_start_default

Whether to use smart cold start values by default.


InferenceSurvivalRestrictedMeanDiff$compute_estimate()

Computes the class-specific mean or survival contrast; see InferenceMLEorKMSummaryTable.

Usage
InferenceSurvivalRestrictedMeanDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_only

If TRUE, skip variance component calculations.

Returns

The setting-appropriate (see description) numeric estimate of the treatment effect


InferenceSurvivalRestrictedMeanDiff$compute_estimate_with_bootstrap_weights()

Recomputes the class-specific treatment estimate for a bootstrap sample; see InferenceNonParamBootstrap.

Usage
InferenceSurvivalRestrictedMeanDiff$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Row weights for the bootstrap sample.

estimate_only

If TRUE, skip variance calculations.


InferenceSurvivalRestrictedMeanDiff$compute_asymp_confidence_interval()

Computes a 1-\alpha level Wald confidence interval for the RMST-difference treatment effect \hat\mu_T(\tau) - \hat\mu_C(\tau), using its Greenwood-based standard error (see class documentation). Falls back to a nonparametric bootstrap interval if that standard error is unavailable or non-finite.

Usage
InferenceSurvivalRestrictedMeanDiff$compute_asymp_confidence_interval(
  alpha = 0.05
)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.

Returns

A (1 - alpha)-sized frequentist confidence interval for the treatment effect


InferenceSurvivalRestrictedMeanDiff$compute_asymp_two_sided_pval()

Computes a two-sided Wald p-value testing H_0: \mu_T(\tau) - \mu_C(\tau) = 0 (only delta = 0 is currently supported; a non-zero null raises an error), using the RMST-difference estimate and its Greenwood-based standard error — see class documentation. Falls back to a nonparametric bootstrap p-value if that standard error is unavailable.

Usage
InferenceSurvivalRestrictedMeanDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

The null difference to test against. For any treatment effect at all this is set to zero (the default).

Returns

The approximate frequentist p-value


InferenceSurvivalRestrictedMeanDiff$compute_rand_confidence_interval()

Uses the shared randomization confidence-interval contract; see InferenceRandCI.

Usage
InferenceSurvivalRestrictedMeanDiff$compute_rand_confidence_interval(
  alpha = 0.05,
  r = 501,
  pval_epsilon = 0.005,
  show_progress = TRUE,
  ci_search_control = NULL
)
Arguments
alpha

The confidence level in the computed confidence interval is 1 - alpha. The default is 0.05.

r

The number of randomization vectors. The default is 501.

pval_epsilon

The bisection algorithm tolerance. The default is 0.005.

show_progress

Show a text progress indicator.

ci_search_control

Unused.

Returns

A 1 - alpha sized frequentist confidence interval


InferenceSurvivalRestrictedMeanDiff$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalRestrictedMeanDiff$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Royston, P., and Parmar, M. K. B. (2013). "Restricted mean survival time: an alternative to the hazard ratio for the design and analysis of randomized trials with a time-to-event outcome." BMC Medical Research Methodology, 13, 152, doi:10.1186/1471-2288-13-152, for RMST as a treatment-effect summary. Kaplan, E. L., and Meier, P. (1958). "Nonparametric Estimation from Incomplete Observations." Journal of the American Statistical Association, 53(282), 457-481, doi:10.2307/2281868, for the underlying survival curve estimator each arm's RMST is integrated from.

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalRestrictedMeanDiff$new(seq_des)
inf$compute_estimate()


Stratified Cox PH Inference for Survival Responses

Description

Fits an auto-stratified Cox proportional hazards regression: rather than the plain single-baseline-hazard model of InferenceSurvivalCoxPHRegr, this class allows a separate baseline hazard per stratum, \lambda(t \mid x_i, s_i) = \lambda_{0,s_i}(t) \exp(x_i^\top\beta), relaxing the proportional-hazards assumption across strata while keeping it within each. Stratification variables are chosen automatically from the recorded low-cardinality (categorical-like) covariates (compute_survival_strata_ids_cpp) — no stratification variables are specified explicitly by the caller. If no suitable stratification covariates are found, the fit falls back to the corresponding standard (unstratified) Cox PH model. Fitting uses survival::coxph.fit()/survival::coxph() with strata passed through when applicable. This is a partial-likelihood class (likelihood_tier = "partial") supporting Wald, score, gradient, and likelihood-ratio tests, plus parametric likelihood-ratio bootstrap calibration. Randomization confidence intervals are not supported (the log-hazard-ratio estimator units are not commensurate with the randomization CI bisection algorithm's log-time-ratio/AFT-effect null search).

Super class

Inference -> InferenceSurvivalStratCoxPHRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalStratCoxPHRegr$new()

Uses the shared randomization two-sided p-value contract; see InferenceRand. Pinned from InferenceRand for the same traced reason as InferenceSurvivalCoxPHRegr (see that factory call's comment): InferenceRandCI's richer override calls super$...(), which resolves against Inference once flattened, and its only other behavior is an incidence-only Zhang special case that never applies to survival data.

Initialize stratified Cox proportional-hazards inference and prepare the partial-likelihood fit used by InferenceSurvivalStratCoxPHRegr.

Usage
InferenceSurvivalStratCoxPHRegr$new(
  des_obj,
  model_formula = NULL,
  use_rcpp = TRUE,
  optimization_alg = "lbfgs",
  verbose = FALSE,
  smart_cold_start_default = NULL
)
Arguments
des_obj

A completed Design object with a survival response.

model_formula

Optional formula for covariate adjustment. If NULL (default), covariates from the design object are included. Use ~ 1 for univariate.

use_rcpp

Logical. If TRUE (default), enable internal Rcpp score/information helpers for likelihood inference. Cox optimization uses survival::coxph.fit.

optimization_alg

Optimization algorithm: "newton_raphson" (default) or "lbfgs".

verbose

Whether to print progress messages.

smart_cold_start_default

Whether to use smart cold start values.


InferenceSurvivalStratCoxPHRegr$compute_asymp_confidence_interval()

Computes an asymptotic confidence interval using the configured likelihood-backed test.

Usage
InferenceSurvivalStratCoxPHRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alpha

Significance level 1 - alpha. Default 0.05.


InferenceSurvivalStratCoxPHRegr$compute_asymp_two_sided_pval()

Computes an asymptotic two-sided p-value using the configured likelihood-backed test.

Usage
InferenceSurvivalStratCoxPHRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
delta

Null treatment effect to test against. Default 0.


InferenceSurvivalStratCoxPHRegr$compute_estimate_with_bootstrap_weights()

Recomputes the stratified Cox PH treatment estimate under Bayesian-bootstrap weights.

Usage
InferenceSurvivalStratCoxPHRegr$compute_estimate_with_bootstrap_weights(
  subject_or_block_weights,
  estimate_only = FALSE
)
Arguments
subject_or_block_weights

Subject-, block-, cluster-, or matched-set bootstrap weights.

estimate_only

If TRUE, compute only the weighted point estimate.


InferenceSurvivalStratCoxPHRegr$compute_rand_confidence_interval()

Compute a randomization-based confidence interval for the stratified Cox treatment effect by inverting the class-specific randomization p-value. See InferenceRandCI.

Usage
InferenceSurvivalStratCoxPHRegr$compute_rand_confidence_interval(
  alpha = 0.05,
  r = 501,
  pval_epsilon = 0.005,
  show_progress = TRUE,
  ci_search_control = NULL
)
Arguments
alpha

The significance level (default 0.05).

r

Number of vectors to draw.

pval_epsilon

The bisection convergence tolerance.

show_progress

Whether to show a progress bar.

ci_search_control

Unused.


InferenceSurvivalStratCoxPHRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalStratCoxPHRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalStratCoxPHRegr$new(seq_des)
inf$compute_estimate()

Weibull AFT Inference for Survival Responses

Description

Fits a Weibull Accelerated Failure Time (AFT) model for survival responses: \log T_i = \beta_0 + \beta_T W_i + X_i^\top \gamma + \sigma \epsilon_i, \epsilon_i \sim standard extreme-value (Gumbel-minimum), so that T_i is marginally Weibull-distributed with shape 1/\sigma and treatment-dependent scale, by maximum likelihood (fast_weibull_regression_cpp). \hat\beta_T is a log-time-ratio (log acceleration factor): \exp(\hat\beta_T) is the estimated multiplicative effect of treatment on survival time (an AFT model, not a proportional-hazards model — the Weibull distribution is the one location where AFT and proportional-hazards parameterizations coincide, since \exp(-\beta_T/\sigma) also equals the treatment hazard ratio). likelihood_tier = "full": likelihood-ratio, score, gradient, and Wald tests are all available when the model converges, plus parametric-likelihood-bootstrap calibration of the likelihood-ratio test. Right-censored and interval-censored observations enter the likelihood via their appropriate survival/density contributions. Validity requires the Weibull shape assumption for the (log-)survival-time distribution and, when interpreted causally, the usual design-based/model-based assumptions.

Computes the randomization distribution of the treatment effect estimate under the sharp null.

Whether compute_rand_two_sided_pval() is actually usable on this instance right now – FALSE exactly when it would stop(): an incidence-response instance with no custom randomization statistic and a design not eligible for design-randomization-based incidence inference (see private$should_use_design_randomization_for_incidence()). TRUE for every other case, including every non-incidence response type. Public, self-contained (only reads already-set instance state, no side effects), so InferenceSuite can check this before attempting the sentinel instead of relying on the stop() being silently swallowed into a pval = NA "ok" row – the single source of truth for both this check and compute_rand_two_sided_pval()'s own guard, so the two can never drift apart.

Value

When debug = FALSE (default), a numeric vector of length r. When debug = TRUE, a list with: values, errors (list of character vectors, one per iteration), warnings (list of character vectors, one per iteration), num_errors, num_warnings, prop_iterations_with_errors, prop_iterations_with_warnings, and prop_illegal_values.

A single logical.

Super class

Inference -> InferenceSurvivalWeibullRegr

Methods

Public methods

+ inherited public methods from Inference

InferenceSurvivalWeibullRegr$approximate_randomization_distribution_beta_hat_T()

Usage
InferenceSurvivalWeibullRegr$approximate_randomization_distribution_beta_hat_T(
  r = 501,
  delta = 0,
  transform_responses = "none",
  show_progress = TRUE,
  permutations = NULL,
  debug = FALSE,
  zero_one_logit_clamp = .Machine$double.eps
)
Arguments
r

Number of randomization vectors. Default 501.

delta

The null difference. Default 0.

transform_responses

Type of transformation. Default "none".

show_progress

Show progress bar. Default TRUE.

permutations

Pre-computed permutations. Default NULL.

debug

If TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. Default FALSE.

zero_one_logit_clamp

The clamping amount for exact 0 and 1 values when logging


InferenceSurvivalWeibullRegr$supports_rand_pval_for_incidence()

Usage
InferenceSurvivalWeibullRegr$supports_rand_pval_for_incidence()

InferenceSurvivalWeibullRegr$clone()

The objects of this class are cloneable with this method.

Usage
InferenceSurvivalWeibullRegr$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

References

Kalbfleisch, J. D., and Prentice, R. L. (2002). The Statistical Analysis of Failure Time Data (2nd ed.). Wiley, for the Weibull AFT model and its equivalence to a proportional-hazards model. For the randomization confidence interval (compute_rand_confidence_interval()), which inverts an AFT sharp null by rescaling the recorded times of treated units — event and censoring times alike, censoring indicators unchanged — the residual construction and its validity under independent censoring are from Tsiatis, A. A. (1990). Estimating regression parameters using linear rank tests for censored data. The Annals of Statistics, 18(1), 354-372, doi:10.1214/aos/1176347504; Wei, L. J., Ying, Z., and Lin, D. Y. (1990). Linear regression analysis of censored survival data based on rank tests. Biometrika, 77(4), 845-851, doi:10.1093/biomet/77.4.845; and Jin, Z., Lin, D. Y., Wei, L. J., and Ying, Z. (2003). Rank-based inference for the accelerated failure time model. Biometrika, 90(2), 341-353, doi:10.1093/biomet/90.2.341.

See Also

Comparable Python API: lifelines WeibullAFTFitter. See also: Proportional hazards model (Wikipedia, for the AFT/PH equivalence note).

Examples


seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
  seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalWeibullRegr$new(seq_des)
inf$compute_estimate()


inf$set_seed(1)
inf$compute_lik_ratio_bootstrap_two_sided_pval(delta = 0, B = 9, show_progress = FALSE)


Abstract class for LWA-style Marginal Cox / Standard Cox Compound Inference

Description

Initialize KK LWA Cox IVWC inference and prepare the matched/reservoir marginal Cox partial-likelihood components used by InferenceSurvivalKKLWACoxPHIVWC.

Returns the estimated treatment effect (log-hazard ratio).

Uses the shared asymptotic confidence-interval contract; see InferenceAsymp.

Compute the LWA Cox asymptotic p-value for the treatment log-hazard ratio using the cluster-robust partial-likelihood standard error. See InferenceAsymp.

Usage

KKLWACoxIVWCPartialLikelihoodSource

Details

This class implements a compound estimator for KK matching-on-the-fly designs with survival responses. For matched pairs, it uses a marginal Cox proportional hazards model with Lee-Wei-Amato style cluster-robust variance, treating each pair as a cluster of size two. For reservoir subjects, it uses standard Cox regression. The two estimates (both log-hazard ratios) are combined via a variance-weighted linear combination.

Under harden = TRUE, multivariate component fits preserve the treatment column and retry reduced covariate sets after QR-based rank reduction and correlation-based pruning. Extreme finite coefficients / standard errors are rejected and treated as non-estimable.


Abstract class for LWA-style Marginal Cox Combined-Likelihood Inference

Description

Initialize KK LWA Cox one-likelihood inference and prepare the combined marginal Cox partial-likelihood fit used by InferenceSurvivalKKLWACoxPHOneLik.

Returns the model-specific combined-likelihood treatment estimate; see InferenceAsympLik.

Recomputes the LWA one-likelihood treatment estimate under Bayesian-bootstrap weights.

Compute an asymptotic confidence interval.

Compute an asymptotic two-sided p-value.

Usage

KKLWACoxOneLikPartialLikelihoodSource

Details

Fits a single joint marginal Cox model over all KK design data for survival responses. Matched subjects share their pair ID as a cluster, and reservoir subjects are treated as independent (unique clusters). Standard errors are obtained via the Huber-White cluster-robust sandwich estimator (LWA style).


Abstract base class for KK Wilcoxon-based compound inference

Description

Shared base for all KK Wilcoxon inference classes. Overrides the per-permutation statistic used in randomization tests with standardized Wilcoxon W statistics (O(n log n), conf.int = FALSE), avoiding the O(n^2) Walsh-average computation required by the full Hodges-Lehmann estimate.

Usage

KKWilcoxIVWCSource

A Fixed Observational (Non-Randomized) Design

Description

A fixed-sample-size DesignFixed whose treatment assignment vector w is supplied by the user (e.g. via assign_w_to_all_subjects(w_precomputed = ...) or overwrite_all_subject_assignments()) rather than drawn from any randomization mechanism. Unlike every other DesignFixed subclass, there is no prob_T to specify – treatment was not assigned by the experimenter according to a known probability law, so no such probability exists to declare.

No draw mechanism. draw_ws_according_to_design() (and, transitively, the fallback branch of assign_w_to_all_subjects() that would otherwise call it) always throws. Any inference procedure that must redraw w from the design's own randomization law – randomization tests, randomization confidence intervals, and randomization/assignment bootstrap – therefore throws the same clear error rather than silently fabricating a randomization mechanism that never existed. Procedures that resample subjects instead of redrawing w (plain nonparametric bootstrap, Bayesian bootstrap) are unaffected and remain available, since resampling subjects with their observed, fixed assignment does not require a known randomization probability.

No balance target. assert_even_allocation() is a no-op here (rather than the inherited check against prob_T = 0.5): there is no targeted allocation ratio for an observational design to be out of balance with.

Super classes

Design -> DesignFixed -> ObservationalDesign

Methods

Public methods

+ inherited public methods from DesignFixed
+ inherited public methods from Design

ObservationalDesign$new()

Initialize a fixed observational (non-randomized) design. No treatment vector is drawn or requested here; the constructor only records configuration (covariates, response type, etc.) and internally fixes prob_T = 0.5 purely so the shared Design/DesignFixed machinery has a value to store — it is never used to draw or validate an allocation for this class (see assert_even_allocation() and class documentation). w itself is supplied afterward via assign_w_to_all_subjects(w_precomputed = ...) or overwrite_all_subject_assignments().

Usage
ObservationalDesign$new(
  response_type,
  include_is_missing_as_a_new_feature = TRUE,
  n = NULL,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

"continuous", "incidence", "proportion", "count", "survival", or "ordinal".

include_is_missing_as_a_new_feature

Flag for missingness indicators.

n

The sample size.

verbose

A flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility. Note this design has no randomization mechanism to seed (see class documentation); seed only affects RNG-dependent behavior inherited from Design unrelated to treatment assignment (e.g. imputation, bootstrap resampling).

Returns

A new 'ObservationalDesign' object


ObservationalDesign$assert_even_allocation()

Observational designs have no targeted allocation ratio to check balance against, so this is a no-op rather than an error (contrast with the inherited Design$assert_even_allocation(), which errors if the realized allocation deviates from prob_T = 0.5).

Usage
ObservationalDesign$assert_even_allocation()
Returns

invisible(NULL), always; never errors.


ObservationalDesign$supports_randomization_draw()

Characterization: FALSE – this design has no randomization mechanism to redraw w from (see class documentation). Metadata-declared replacement for the old draw_ws_raw() throwing stub (still present below as a fallback for any caller that reaches it without checking this first – see fix_design_hierarchy.md, "Observational Design Migration").

Usage
ObservationalDesign$supports_randomization_draw()
Returns

Always FALSE for this class.


ObservationalDesign$supports_resampling_replay()

Characterization: FALSE – this design has no randomization mechanism to replay against resampled data (bootstrap randomization test eligibility). Plain nonparametric/Bayesian/m-out-of-n/ PRW-subsampling bootstrap are unaffected (see Design$supports_resampling()'s documentation) and remain available.

Usage
ObservationalDesign$supports_resampling_replay()
Returns

Always FALSE for this class.


ObservationalDesign$clone()

The objects of this class are cloneable with this method.

Usage
ObservationalDesign$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

des = ObservationalDesign$new(n = 10, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(10)))
des$assign_w_to_all_subjects(w_precomputed = rbinom(10, 1, 0.3))

A Fixed Observational (Non-Randomized) Design With Blocks

Description

An ObservationalDesign whose subjects are additionally partitioned into user-supplied blocks (matched sets / strata), via the block-membership vector m. As with ObservationalDesign, there is no randomization mechanism at all – neither the treatment assignment w nor the block membership m is drawn by this class, both are supplied by the user – so draw_ws_according_to_design() still always throws (inherited unchanged from ObservationalDesign).

Why blocks, if there's no randomization to block on? Blocking here is not a randomization restriction (there is none); it is a resampling structure. Supplying m lets bootstrap procedures resample within blocks – exactly as they do for DesignFixedBlocking and DesignFixedOptimalBlocks – which is appropriate when the observational data itself has a matched/stratified/clustered structure (e.g. matched case-control sets, repeated measurements within site) that the bootstrap should respect.

No auto-derived n. Unlike plain ObservationalDesign, n is not a constructor argument here at all – it is always length(m), since a block membership vector with one entry per subject already fixes the sample size.

Super classes

Design -> DesignFixed -> ObservationalDesign -> ObservationalDesignBlocks

Methods

Public methods

+ inherited public methods from ObservationalDesign
+ inherited public methods from DesignFixed
+ inherited public methods from Design

ObservationalDesignBlocks$new()

Initialize a fixed observational (non-randomized) design with user-supplied block membership. Neither treatment w nor block membership m is drawn here — both must be supplied (m at construction, w afterward via assign_w_to_all_subjects(w_precomputed = ...)); see class documentation for why blocks are still useful without a randomization mechanism to block on.

Usage
ObservationalDesignBlocks$new(
  response_type,
  m,
  include_is_missing_as_a_new_feature = TRUE,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

"continuous", "incidence", "proportion", "count", "survival", or "ordinal".

m

A positive-integer vector of block (matched-set/stratum) identifiers, one entry per subject; a block may contain any number of subjects (unlike ObservationalDesignMatching, which fixes block size at exactly 2). n is derived as length(m) and is not a separate argument.

include_is_missing_as_a_new_feature

Flag for missingness indicators.

verbose

A flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'ObservationalDesignBlocks' object


ObservationalDesignBlocks$clone()

The objects of this class are cloneable with this method.

Usage
ObservationalDesignBlocks$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

des = ObservationalDesignBlocks$new(response_type = 'continuous', m = c(1, 1, 2, 2, 3, 3))
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(6)))
des$assign_w_to_all_subjects(w_precomputed = c(1, 0, 1, 0, 1, 0))

A Fixed Observational (Non-Randomized) Matched-Pair Design

Description

An ObservationalDesign whose n subjects are organized into n/2 matched pairs, each of size exactly 2. This is a convenience wrapper: it builds the canonical pair-membership vector m = (1, 1, 2, 2, \dots, n/2, n/2) internally – subject 2k - 1 and subject 2k are pair k – and installs it via set_m(), so the caller need only supply subjects (and, later, responses) in that paired order; there is no separate m argument to get wrong.

Why not just ObservationalDesignBlocks with block size 2? It would look equivalent but silently behave differently: matching is a distinct capability (private$matching_capable, checked via is_matching_design()) from generic blocking, and several Inference components (jackknife, nonparametric bootstrap, Bayesian bootstrap, exchangeable-resampling-unit selection) branch on it to use pair-preserving resampling (MatchingStructure's draw_bootstrap_indices(), via draw_matching_bootstrap_sample_cpp()) instead of generic per-stratum resampling. ObservationalDesignBlocks overrides draw_bootstrap_indices() with the generic stratified version, so subclassing it here would advertise is_matching_design() == TRUE while still running the wrong bootstrap underneath. This class instead extends ObservationalDesign directly – the same relationship DesignFixedBinaryMatch has to DesignFixed – so it inherits MatchingStructure's matched-pair bootstrap machinery unmodified.

Super classes

Design -> DesignFixed -> ObservationalDesign -> ObservationalDesignMatching

Methods

Public methods

+ inherited public methods from ObservationalDesign
+ inherited public methods from DesignFixed
+ inherited public methods from Design

ObservationalDesignMatching$new()

Initialize a fixed observational (non-randomized) design whose subjects are organized into n / 2 matched pairs of size 2 (subjects 2k - 1 and 2k form pair k). Unlike ObservationalDesignBlocks there is no separate m argument — the pair structure is fixed by subject order and installed automatically via set_m(), and private$matching_capable is set so downstream Inference classes use pair-preserving (not generic stratified) resampling (see class documentation). As with every ObservationalDesign, w itself is supplied afterward via assign_w_to_all_subjects(w_precomputed = ...), in the same paired subject order.

Usage
ObservationalDesignMatching$new(
  response_type,
  n,
  include_is_missing_as_a_new_feature = TRUE,
  verbose = FALSE,
  missingness_method = "impute",
  design_formula = ~.,
  seed = NULL
)
Arguments
response_type

"continuous", "incidence", "proportion", "count", "survival", or "ordinal".

n

The sample size; must be even (subjects 2k - 1/2k form pair k, so an odd n would leave one subject unpaired).

include_is_missing_as_a_new_feature

Flag for missingness indicators.

verbose

A flag for verbosity.

missingness_method

How to handle missing values in covariates.

design_formula

A formula object.

seed

Integer seed for reproducibility.

Returns

A new 'ObservationalDesignMatching' object


ObservationalDesignMatching$clone()

The objects of this class are cloneable with this method.

Usage
ObservationalDesignMatching$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples

des = ObservationalDesignMatching$new(response_type = 'continuous', n = 6)
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(6)))
des$assign_w_to_all_subjects(w_precomputed = c(1, 0, 0, 1, 1, 0))

Simulation Framework for Experimental Designs and Inference Methods

Description

An R6 class for benchmarking experimental designs and inference methods by Monte Carlo simulation. Each replication generates synthetic covariates and responses, runs every requested (design, inference) pair, and records point estimates, confidence intervals, and p-values. Raw and aggregated results are available through SimulationFrameworkReport.

Details

Covariates are drawn independently from \mathrm{Uniform}(0, 1).

The continuous base signal is transformed to the scale appropriate for response_type. Treatment effects are applied per-subject: additive on the linear/logit/ordinal scale, log-multiplicative for count and survival.

For each (design, inference) pair the framework runs whichever of the following are supported by the inference class:

Incompatible (design, inference) pairs (e.g.\ a KK-specific inference class with a non-KK design) are silently skipped via tryCatch.

Reported summary metrics include:

Methods

Public methods


SimulationFramework$new()

Create a new SimulationFramework object that stores simulation design settings, response-generation settings, inference methods, and replication controls.

Usage
SimulationFramework$new(
  response_type,
  design_classes_and_params = NULL,
  inference_classes_and_params = NULL,
  n = 100L,
  p = 5L,
  cond_exp_func_model = "linear",
  Nrep_W = 100L,
  Nrep_Y_w = 1L,
  betaT = 1,
  alpha = 0.05,
  B_boot = 201L,
  r_rand = 201L,
  pval_epsilon = 0.02,
  sd_noise = 1,
  n_ordinal_levels = 4L,
  proportion_epsilon = 1e-06,
  phi_proportion = 100,
  k_survival = 2,
  incidence_clamp = 1e-09,
  proportion_clamp = 1e-09,
  count_clamp = 1e-09,
  survival_clamp = 1e-09,
  survival_min_time = 0.1,
  count_min_rate = 0L,
  count_shift = 0,
  norm_sq_beta_vec = 1,
  X_mat = NULL,
  num_cores = 1L,
  seed = NULL,
  cov_draw_method = stats::rnorm,
  cov_draw_method_args = list(mean = 0, sd = 1),
  random_X_draws = TRUE,
  prob_censoring = 0.25,
  custom_replication_data_generator = NULL,
  custom_apply_treatment_and_noise = NULL,
  make_estimand_fn = NULL,
  dgp_params = list(),
  custom_dgp = NULL,
  verbose = TRUE,
  keep_all_intermediate_data = FALSE,
  turn_off_asserts_for_speed = TRUE,
  inference_types_and_params = NULL,
  results_filename = "simulation_framework_results.csv.bz2",
  continue_from_last_result_row = TRUE,
  reuse_cache = TRUE,
  stop_on_error = TRUE,
  save_to_disk_every_n_rep = 25L,
  save_model_control_fits = TRUE
)
Arguments
response_type

(required) Character scalar or vector. The type of outcome variable. One of "continuous", "incidence", "proportion", "count", "survival", "ordinal".

design_classes_and_params

NULL (default) or a list describing design classes and optional constructor parameters. Unnamed R6 class generators use default parameters, for example list(DesignSeqOneByOneKK21, DesignFixedBernoulli). Named entries use the entry name as the design class and the value as the parameter list, for example list(DesignSeqOneByOneUrn = list(alpha = 2, beta = 2)). Duplicate named entries are allowed for repeated designs with different parameters. Each generator must be constructable with only response_type and n plus any extra params supplied in this list. NULL uses the package's standard design set. Designs requiring strata_cols, cluster_col, or factors have sensible defaults auto-injected (first covariate column; second for cluster_col; list(treatment=2) for factors) when not supplied in the parameter list. Example:

design_classes_and_params = list(
  DesignSeqOneByOneKK21 = list(lambda = 0.5, t_0_pct = 0.1),
  DesignSeqOneByOneUrn  = list(alpha  = 2,   beta    = 2),
  DesignFixedBernoulli  # default params
)

Commonly useful design constructor parameters:

lambda

Matching-weight decay for KK14 / KK21 / KK21stepwise.

t_0_pct

Burn-in fraction for KK14 / KK21 / KK21stepwise.

morrison

Logical; Morrison correction for KK14.

alpha, beta

Shape parameters for DesignSeqOneByOneUrn.

preferred_num_bins_for_continuous_covariate

Bin count for DesignFixedBlocking and DesignFixedBlockedCluster.

B_target

Target number of blocks for DesignFixedBlocking.

inference_classes_and_params

NULL (default) or a list describing inference classes and optional constructor parameters. Unnamed R6 class generators use default parameters, for example list(InferenceContinOLS, InferenceContinKKOLSIVWC). Named entries use the entry name as the inference class and the value as the constructor parameter list, for example list(InferenceContinOLS = list(max_resample_attempts = 25L)). Duplicate named entries are allowed for repeated inference classes with different parameters. Supplied parameters must be accepted by the inference class constructor. NULL selects a curated set for the given response_type: several universal classes that work with any design, plus representative KK-specific classes (silently skipped for non-KK designs at runtime).

n

Integer scalar or vector. Sample size per simulation replication. Default 100.

p

Integer scalar or vector. Number of covariates. Must be \ge 5 when cond_exp_func_model = "nonlinear". Default 5.

cond_exp_func_model

Character scalar or vector. How the latent continuous signal is constructed before transformation to the response_type scale.

"linear"

Linear combination X\beta with coefficients evenly spaced from 1 to -1.

"nonlinear"

Friedman (1991) function 10\sin(\pi x_1 x_2)+20(x_3-0.5)^2+10x_4+5x_5; requires p \ge 5.

Default "linear".

Nrep_W

Positive integer. Number of treatment-assignment draws (w-reps). Each w-rep generates a fresh covariate matrix X and a new treatment assignment vector w. Default 100L.

Nrep_Y_w

Positive integer. Number of draws of the response per draw of w, the allocation vector. For each w-rep, Nrep_Y_w independent response vectors y are drawn from the same (X, w). The effective total number of replications recorded is Nrep_W * Nrep_Y_w. Default 1 (standard behaviour: one outcome draw per allocation-vector draw).

betaT

Numeric scalar or vector. True treatment effect added to treated subjects' outcomes. The scale is response-type specific: additive for continuous, proportion, and ordinal; on the logit scale for incidence; log-multiplicative for count and survival. Default 1. Set betaT = 0 to check type-I error.

alpha

Numeric in (0,1). Significance level used for all confidence intervals and for computing power (p < \alpha). Default 0.05.

B_boot

Positive integer. Bootstrap resamples per CI / p-value call. Default 201.

r_rand

Positive integer. Randomisation draws per rand p-value call, and per bisection step of the rand CI. Default 201.

pval_epsilon

Numeric. Bisection convergence tolerance for randomisation-based CIs (compute_rand_confidence_interval). Default 0.02.

sd_noise

Numeric > 0. Standard deviation of independent Gaussian noise added to each subject's outcome. Default 1.

n_ordinal_levels

Positive integer. Number of ordinal categories when response_type = "ordinal". Default 4L.

proportion_epsilon

Numeric scalar. Small value added to proportion base responses to avoid 0 and 1. Default 1e-6.

phi_proportion

Positive numeric scalar. Precision parameter for beta-distributed observed proportion outcomes. The beta mean is y_linear_model[i] + betaT * w[i]. Default 100.

k_survival

Positive numeric scalar. Scale parameter passed to the Weibull draw for observed survival outcomes. Default 2.

incidence_clamp

Numeric scalar in (0, 0.5). Clamp applied to the Bernoulli probability for observed incidence outcomes. Default 1e-9.

proportion_clamp

Numeric scalar in (0, 0.5). Clamp applied to the beta mean for observed proportion outcomes. Default 1e-9.

count_clamp

Positive numeric scalar. Minimum Poisson mean for observed count outcomes. Default 1e-9.

survival_clamp

Positive numeric scalar. Minimum Weibull shape for observed survival outcomes. Default 1e-9.

survival_min_time

Numeric scalar. Minimum survival time and shift for base responses. Default 0.1.

count_min_rate

Integer scalar. Minimum baseline rate for count responses. Default 0L.

count_shift

Numeric scalar. Constant added to counts after zero-centering for base responses. Default 0.

norm_sq_beta_vec

Positive numeric scalar. The desired squared Euclidean norm of the latent linear coefficient vector \beta. The generated vector is scaled to match this norm. Default 1.

X_mat

Numeric matrix of dimensions n x p, or NULL (default). If provided, these fixed covariates are used for every replication. In this case, cov_draw_method must be NULL.

num_cores

Positive integer. Number of worker processes for parallel execution of Monte Carlo replications. Note that when num_cores > 1, parallelization *within* individual inference routines (e.g. bootstrap, randomization) is automatically disabled to prevent thread oversubscription.

Unix/Linux (recommended): A makeForkCluster pool is created once at run() start. Workers inherit all pre-generated design and SE caches via copy-on-write with zero serialization overhead. Parallelism operates at the replication level: num_cores replications run simultaneously, each executing all DGP cells serially. This eliminates the per-batch dispatch overhead that would arise from cycling through cells within every replication, and keeps all cores fully subscribed regardless of the number of DGP cells. For best performance, run on a Unix/Linux machine and set num_cores to the number of physical cores available.

Non-Unix (Windows/macOS): mirai daemons are used when available. Every active (replication, DGP-cell) pair is one work unit and num_cores units are kept in flight continuously, so all cores stay subscribed even when the grid has fewer DGP cells than cores. Cell state is pushed to the daemons once at run() start rather than re-serialized per dispatch. If mirai is not installed, execution falls back to serial with a warning. Default 1.

seed

Integer or NULL (default). Random seed for the entire simulation run.

cov_draw_method

A function used to draw n * p i.i.d. covariate values for every replication. The function must accept the total number of values as its first argument, followed by arguments in cov_draw_method_args. Default stats::rnorm. Must be NULL when X_mat is supplied.

cov_draw_method_args

Named list of additional arguments forwarded to cov_draw_method beyond the sample-size first argument. Default is list(mean = 0, sd = 1).

random_X_draws

Logical. If TRUE (default), a new set of covariates is drawn for every single replication. If FALSE, one set is drawn per (n, p) cell and shared across its replications.

prob_censoring

Numeric in [0,1]. Per-subject independent censoring probability; applied only when response_type = "survival". Default 0.25.

custom_replication_data_generator

Optional function for custom replication data. When supplied, it is called as fn(state, rep) and must return a list containing at least X and y_linear_model. Any additional fields in the returned list (e.g. latent frailty draws) are passed forward as rep_data to custom_apply_treatment_and_noise and the function built by make_estimand_fn.

custom_apply_treatment_and_noise

Optional function for custom response generation. Signature: fn(y_linear_model, w, rep_data, state). w uses {0, 1} encoding, with 1 for treatment and 0 for control. rep_data is the full list returned by custom_replication_data_generator (or NULL for the standard path). Must return a list with components y and dead. Three-argument functions fn(y_linear_model, w, state) are still accepted for backwards compatibility.

make_estimand_fn

Optional factory function for a custom true estimand. Signature: fn(beta_T), returning a function with signature fn(y_linear_model, X, w, rep_data, state). Called once per grid cell with that cell's beta_T so the returned estimand function is always tied to the right effect size (important when betaT is a vector of multiple values). The returned function is invoked once per design class per replication after the design completes, so w (in {0, 1} encoding) and X reflect the realized assignment. Must return a numeric scalar. When supplied, its return value is used as the ground truth for all inference classes (overriding the is_mean_diff gate). Three-argument functions fn(y_linear_model, state) are still accepted for backwards compatibility (they will not receive X or w). Default NULL uses beta_T directly as the ground truth.

dgp_params

Optional named list of DGP configuration values (e.g. list(frailty_dist = "gamma", censoring_rate = 0.8)). Injected into state as state\$dgp_params and accessible in all three custom-DGP hooks. Recommended over using closures to pass DGP parameters.

custom_dgp

Optional function for a fully custom DGP. Signature: fn(n, p, rep, state) returning a list with components X (data.frame, n x p), w (integer vector in {0, 1}, length n), y (numeric, length n), dead (integer {0,1} or NULL for non-survival), true_estimand (numeric scalar, optional). When supplied, the design class acts as a data container only; it does not run its own randomization or matching. Requires a fixed design class (not DesignSeqOneByOne variants). Cannot be combined with custom_replication_data_generator or custom_apply_treatment_and_noise.

verbose

Logical. If TRUE, prints a message for every replication and for every (design, inference) pair that is skipped due to an error. Default TRUE.

keep_all_intermediate_data

Logical. If TRUE, the framework saves the instantiated design and inference objects for every replication. These can be retrieved after the run using $get_all_intermediate_data(). Warning: this can consume a lot of memory for many replications. Default FALSE.

turn_off_asserts_for_speed

Logical. If TRUE (default), all checkmate assertions across the package are globally disabled during the simulation run to improve performance.

inference_types_and_params

NULL (default) or a named list from inference type to a named list of arguments for that type's function invocation. The list names control which inference outputs are computed. Valid names are "asymp_ci", "asymp_pval", "exact_ci", "exact_pval", "boot_ci", "boot_pval", "rand_ci", and "rand_pval". Each value must be a named list whose names are accepted by the corresponding inference function. NULL runs all eight types with default invocation arguments. Example:

inference_types_and_params = list(
  asymp_pval = list(delta = 0),
  boot_ci    = list(B = 99, type = "perc"),
  rand_pval  = list(r = 999, transform_responses = TRUE)
)

When no *_ci type is requested, coverage is omitted from SimulationFrameworkReport$summarize(). When no *_pval type is requested, power is omitted.

results_filename

Character scalar. The filename for the results file. Supported extensions are .csv and .csv.bz2. Default "simulation_framework_results.csv.bz2".

continue_from_last_result_row

Logical. If TRUE (default), the framework loads existing results from results_filename and skips previously completed replications.

reuse_cache

Logical. If TRUE (default), expensive pre-generated design / SE cache objects are loaded from disk when available. If FALSE, these cache objects are regenerated from scratch, but each regenerated object is still saved to disk for later restarts.

stop_on_error

Logical. If TRUE (default), any error raised during a simulation path aborts the run immediately. If FALSE, the framework records the error, skips the failing path, and continues with the remaining replications / design / inference combinations. Use $get_errors() after $run() to inspect the captured errors.

save_to_disk_every_n_rep

Positive integer. Results are flushed to the on-disk staging file only once every this many replications, and always after the final replication. Larger values reduce disk I/O overhead at the cost of losing more progress if the run is interrupted. Default 25L.

save_model_control_fits

Logical. If TRUE (default), after the design/SE cache is built, saves the per-subject model-implied potential outcomes under treatment and control as CSV files in a subfolder named <stem>_response_values/ (where <stem> is results_filename with its .csv/.csv.bz2 extension stripped) next to results_filename. One file is written per unique (response_type, cond_exp_func_model, n, p, betaT) cell. Only meaningful when random_X_draws = FALSE; silently skipped otherwise. Column names depend on response_type:

"continuous", "survival"

columns yt and yc

"incidence", "proportion"

columns pt and pc

"count"

columns rt and rc

Default TRUE.


SimulationFramework$run()

Execute the configured simulation replications, run each requested design and inference method, collect estimates/p-values/CIs and errors, and return a SimulationFrameworkReport.

Usage
SimulationFramework$run()
Returns

The SimulationFramework object itself (invisibly).


SimulationFramework$get_all_intermediate_data()

Retrieve the stored intermediate data (design and inference objects) for every replication. Only available if keep_all_intermediate_data = TRUE was passed to the constructor.

Usage
SimulationFramework$get_all_intermediate_data()
Returns

A nested list containing the intermediate data for each replication, or NULL if not recorded.


SimulationFramework$clear_all_intermediate_data_and_gc()

Release all stored intermediate data and invoke the garbage collector. Useful after inspecting intermediate results to free memory before further processing. Sets the internal store to NULL and calls gc().

Usage
SimulationFramework$clear_all_intermediate_data_and_gc()
Returns

The SimulationFramework object itself (invisibly).


SimulationFramework$clone()

The objects of this class are cloneable with this method.

Usage
SimulationFramework$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


# Simple simulation with two designs and two inference methods.
# n/Nrep_W/num_boot/B_boot/r_rand are kept small so this example runs in a
# few seconds; a real simulation would use much larger values (this
# class's defaults, or larger still) for adequately precise estimates.
sim = SimulationFramework$new(
  response_type = "continuous",
  design_classes_and_params = list(
    DesignSeqOneByOneKK21 = list(lambda = 0.5, num_boot = 50L),
    DesignSeqOneByOneBernoulli = list()
  ),
  inference_classes_and_params = list(
    InferenceContinOLS = list(),
    InferenceContinKKOLSIVWC = list()
  ),
  n = 20, p = 3, Nrep_W = 2L, betaT = 1, B_boot = 50L, r_rand = 101L,
  results_filename = tempfile(fileext = ".csv.bz2"),
  continue_from_last_result_row = FALSE
)
sim$run()
report = SimulationFrameworkReport$new(sim)
report$summarize()


Reporting class for SimulationFramework results

Description

An R6 class for accessing and summarizing the results of a SimulationFramework run. It can be constructed either from a completed SimulationFramework object (via SimulationFrameworkReport$new(sim)) or by loading results from a previously saved CSV / CSV.BZ2 file (via SimulationFrameworkReport$new("path/to/results.csv")).

Details

When constructed from a SimulationFramework object all design/inference parameter metadata is preserved, so $summarize() can annotate each row with human-readable parameter strings. When constructed from a file only the raw results are available; parameter annotation columns will be empty strings.

Methods

Public methods


SimulationFrameworkReport$new()

Create a new SimulationFrameworkReport object that stores simulation results, captured errors, and summary helpers returned by SimulationFramework.

Usage
SimulationFrameworkReport$new(sim_or_filename, alpha = NULL)
Arguments
sim_or_filename

Either a completed SimulationFramework object or a character string giving the path to a .csv or .csv.bz2 results file written by SimulationFramework.

alpha

Numeric in (0,1). Significance level for coverage and power calculations. When sim_or_filename is a SimulationFramework object and alpha is NULL (default), the framework's own alpha is used. When loading from a file, defaults to 0.05.


SimulationFrameworkReport$get_results()

Get the raw per-replication results.

Usage
SimulationFrameworkReport$get_results()
Returns

A data.table with one row per (replication, design, inference class, inference type).


SimulationFrameworkReport$get_errors()

Return all errors captured during the simulation run.

Usage
SimulationFrameworkReport$get_errors()
Returns

A list of named lists, one per captured error. Each element includes the simulation cell metadata, replication number, design / inference path, user-supplied parameters, error stage, and error message. Empty when constructed from a file.


SimulationFrameworkReport$summarize()

Aggregate and summarize simulation results.

Usage
SimulationFrameworkReport$summarize()
Returns

A data.table with one row per unique (response_type, cond_exp_func_model, n, p, betaT, design, inference, inference_type) combination. Columns include MSE, coverage, ci_length, and coverage_pval (when CI types were run; coverage_pval is the exact two-sided binomial test p-value of H0: true coverage = 1 - alpha), power (when betaT != 0 and p-value types were run), size and size_pval (when betaT == 0 and p-value types were run; size_pval is the exact two-sided binomial test p-value of H0: true size = alpha, suitable for multiplicity-corrected calibration checks across settings), and parameter annotation strings.


SimulationFrameworkReport$print()

Print a concise summary of the report.

Usage
SimulationFrameworkReport$print()

SimulationFrameworkReport$clone()

The objects of this class are cloneable with this method.

Usage
SimulationFrameworkReport$clone(deep = FALSE)
Arguments
deep

Whether to make a deep clone.

Examples


sim <- SimulationFramework$new(
  response_type = "continuous",
  design_classes_and_params = list(DesignFixedBernoulli),
  inference_classes_and_params = list(InferenceAllSimpleAverageDiff),
  n = 20L, Nrep_W = 5L, betaT = 1,
  results_filename = tempfile(fileext = ".csv"),
  verbose = FALSE, continue_from_last_result_row = FALSE
)
sim$run()
report <- SimulationFrameworkReport$new(sim)
report$get_results()
report$summarize()


GLM and Kaplan-Meier Inference

Description

Computes the treatment estimate using the underlying model.

Computes an asymptotic confidence interval using the configured test.

Computes an asymptotic two-sided p-value using the configured test.

Usage

StandardModelCacheSource

Details

Abstract class providing MLE/KM-based inference methods for GLM and survival models.

Value

A confidence interval.

The asymptotic p-value.


Dependent-censoring transformation component source

Description

Initialize inference for the bivariate log-normal dependent-censoring transformation model; see InferenceSurvivalDepCensTransformRegr for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Fits the joint bivariate log-normal event/censoring transformation model by maximum likelihood and returns the event-time log-time-ratio estimate \hat\beta_T; see InferenceSurvivalDepCensTransformRegr for the model form.

Recomputes the treatment estimate under subject/block-level Bayesian-bootstrap weights, using weighted_cox_bootstrap_surrogate_fit() — a fast weighted Cox-model surrogate fit — as an approximation to the weighted joint dependent-censoring likelihood, rather than a full weighted refit of the joint bivariate model.

Reports jackknife bias correction as unavailable for this model; leave-one-out bias correction is unstable for the dependent-censoring transformation likelihood on small censored samples.

Reports jackknife bias estimate as unavailable for this model.

Reports jackknife standard error as unavailable for this model.

Reports jackknife Wald two-sided p-value as unavailable for this model.

Reports jackknife Wald confidence interval as unavailable for this model.

Reports randomization inference as unavailable for this model; each randomization draw requires a full dependent-censoring likelihood refit and is not stable enough for the comprehensive suite.

Reports randomization confidence interval as unavailable for this model.

Bootstrap confidence interval, validated for this model.

Basic bootstrap confidence interval, validated for this model.

BCa bootstrap confidence interval, validated for this model.

Studentized bootstrap confidence interval, validated for this model.

Reports the randomization distribution as unavailable for this model.

Wald confidence interval for the event-submodel log-time-ratio \beta_T using the fitted joint model's standard error; see InferenceAsymp for the shared Wald contract. Fits the model first if not already cached.

Two-sided Wald test of H_0: \beta_T = \code{delta} for the event-submodel log-time-ratio, using the fitted joint model's standard error; see InferenceAsymp for the shared Wald contract. Fits the model first if not already cached.

Computes a score two-sided p-value, falling back to the asymptotic test when unavailable.

Computes a likelihood-ratio confidence interval, reporting unstable inversion failures as explicitly non-estimable.

Usage

SurvivalDepCensTransformSource

Details

Source list for the dependent-censoring transformation survival component.


GLMM Weibull log-gamma-frailty IVWC component source

Description

Initialize KK Clayton-copula survival inference and prepare the matched/reservoir likelihood components used by InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC.

Returns the model-specific log-time-ratio treatment estimate; see Inference.

Recomputes the IVWC Clayton-copula treatment estimate under Bayesian-bootstrap weights.

Uses the shared asymptotic confidence-interval contract; see InferenceAsymp.

Compute the Clayton-copula survival asymptotic p-value for the treatment effect using the fitted frailty/dependence model. See InferenceAsymp for shared p-value semantics.

Usage

SurvivalGLMMWeibullFrailtyLoggammaIVWCSource

Details

Source list for the KK inverse-variance-weighted-combination (IVWC) Weibull log-gamma-frailty survival component.


Clayton Copula Combined-Likelihood Inference for KK Designs

Description

Initialize KK Clayton-copula one-likelihood survival inference and prepare the combined likelihood used by InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik.

Returns the model-specific log-time-ratio treatment estimate; see Inference.

Uses the shared asymptotic confidence-interval contract; see InferenceAsymp.

Compute the one-likelihood Clayton-copula survival asymptotic p-value for the treatment effect using the fitted dependence model. See InferenceAsymp.

Duplicates this subclass while preserving fit caches; see Inference.

Usage

SurvivalGLMMWeibullFrailtyLoggammaOneLikSource

Details

Gamma-frailty (Clayton copula) Weibull estimator; see InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC for the frailty-distribution details and contrast with the log-normal-frailty InferenceSurvivalGLMMWeibullFrailtyNormalOneLik alternative.


Abstract class for Weibull Frailty / Standard Weibull Compound Inference

Description

Initialize KK Weibull-frailty IVWC survival inference and prepare matched/reservoir parametric survival components used by InferenceSurvivalGLMMWeibullFrailtyNormalIVWC.

Returns the model-specific log-time-ratio treatment estimate; see Inference.

Uses the shared asymptotic confidence-interval contract; see InferenceAsymp.

Compute the Weibull-frailty asymptotic p-value for the treatment effect using the fitted parametric survival likelihood. See InferenceAsymp.

Usage

SurvivalGLMMWeibullFrailtyNormalIVWCSource

Details

This class implements a compound estimator for KK matching-on-the-fly designs with survival responses using a Weibull AFT GLMM for matched pairs. The matched-pair component uses a shared log-normal random intercept per pair, fitted by the package's native Rcpp likelihood optimizer. The reservoir component uses standard Weibull AFT regression. The two treatment-effect estimates are combined by inverse-variance weighting.

This compound estimator accounts for the dependence within matched pairs by modeling it as a shared frailty.

Frailty distribution. The matched-pair likelihood is an AFT (accelerated failure time) parameterization, log(T) = X beta + u + sigma_eps * epsilon, where epsilon is standard extreme-value (giving Weibull margins) and the pair-shared random intercept u is Gaussian, u ~ N(0, sigma_u^2). On the natural time scale this is a multiplicative log-normal frailty, exp(u). A Gaussian random effect has no closed-form marginal likelihood under a Weibull baseline, so the pair likelihood is evaluated by Gauss-Hermite quadrature (fast_weibull_frailty_cpp) rather than in closed form.

This is a different (and equally standard) frailty assumption from the classic gamma-frailty Weibull model (Clayton 1978; Vaupel, Manton & Stallard 1979; Hougaard 2000), which multiplies the hazard (not the AFT error) by a shared Gamma(1/theta, 1/theta) term and has a closed-form marginal survival function via the frailty's Laplace transform. That model is implemented in this package as the Clayton copula of InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC / InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik (the Clayton copula with Weibull margins is exactly the closed-form bivariate survival function obtained by integrating a shared gamma frailty out of two conditionally-independent Weibull hazards). Prefer this log-normal-frailty class for a Gaussian-random-intercept / GLMM-style dependence structure; prefer the Clayton-copula class for the classic gamma-frailty / proportional-hazards dependence structure.

Univariate (ncol(as.matrix(private$X)) == 0): uses the native Rcpp Weibull frailty likelihood with formula = survival::Surv(y, dead) ~ w and a pair-level random intercept.

Multivariate (ncol(as.matrix(private$X)) > 0): fits the same native Rcpp likelihood with covariate adjustment, dropping rank-deficient columns when needed.

See Also

InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC for the corresponding gamma-frailty (Clayton copula) IVWC estimator.


Weibull Frailty Combined-Likelihood Inference for KK Designs

Description

Initialize the one-likelihood Weibull-frailty inference object.

Usage

SurvivalGLMMWeibullFrailtyNormalOneLikLeafSource

Details

Log-normal (Gaussian random-intercept) frailty Weibull AFT estimator; see InferenceSurvivalGLMMWeibullFrailtyNormalOneLik for the frailty-distribution details and contrast with the gamma-frailty InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik (Clayton copula) alternative.


Abstract class for Weibull Frailty Combined-Likelihood Inference

Description

Initialize KK Weibull-frailty one-likelihood survival inference and prepare the combined parametric survival likelihood used by InferenceSurvivalGLMMWeibullFrailtyNormalOneLik.

Returns the model-specific combined-likelihood treatment estimate; see InferenceAsympLik.

Recomputes the one-likelihood Weibull-frailty treatment estimate under Bayesian-bootstrap weights.

Computes an asymptotic confidence interval for the treatment effect.

Returns a 2-sided p-value for H0: beta_T = delta.

Usage

SurvivalGLMMWeibullFrailtyNormalOneLikSource

Details

One-likelihood (combined matched-pair + reservoir) analog of InferenceSurvivalGLMMWeibullFrailtyNormalIVWC: same Weibull-AFT-with-log-normal-random-intercept (Gaussian, Gauss-Hermite quadrature) frailty assumption for matched pairs, but the matched-pair and reservoir contributions are fit as a single combined likelihood rather than combined by inverse-variance weighting. See InferenceSurvivalGLMMWeibullFrailtyNormalIVWC for the frailty-distribution details and its contrast with the gamma-frailty Clayton copula model implemented by InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik.

See Also

InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik for the corresponding gamma-frailty (Clayton copula) combined-likelihood estimator.


Abstract class for Survival Rank-based Regression (AFT) Compound Inference

Description

Initialize KK survival rank-regression IVWC inference and prepare matched/reservoir Gehan-Wilcoxon rank-regression components used by InferenceSurvivalKKRankRegrIVWC.

Returns the model-specific log-time-ratio treatment estimate; see Inference.

Uses the shared asymptotic confidence-interval contract; see InferenceAsymp.

Compute the survival rank-regression asymptotic p-value for the treatment effect using the fitted rank-regression estimate and standard error. See InferenceAsymp.

Usage

SurvivalKKRankRegrIVWCSource

Details

This class implements a robust compound estimator for KK matching-on-the-fly designs with survival responses using rank-based estimating equations via the aftgee package. For matched pairs, it fits a rank-based AFT model with clustering. For reservoir subjects, it fits a standard rank-based AFT model. The two estimates (both log-time ratios) are combined via a variance-weighted linear combination.

This class requires the aftgee package. Under harden = TRUE, multivariate component fits preserve the treatment column and retry reduced covariate sets after QR-based rank reduction and correlation-based pruning. Extreme finite coefficients / standard errors are rejected and treated as non-estimable.


KK stratified Cox IVWC component source

Description

Initialize KK stratified Cox IVWC inference and prepare the matched/reservoir partial-likelihood components used by InferenceSurvivalKKStratCoxPHIVWC.

Returns the estimated treatment effect (log-hazard ratio).

Uses the shared asymptotic confidence-interval contract; see InferenceAsymp.

Compute the stratified-Cox asymptotic p-value for the treatment log-hazard ratio using the fitted partial-likelihood standard error. See InferenceAsymp.

Usage

SurvivalKKStratCoxIVWCSource

Details

Source list for the KK stratified Cox inverse-variance-weighted-combination (IVWC) survival component.


Stratified Cox Combined-Likelihood Compound Inference for KK Designs

Description

Initialize KK stratified Cox one-likelihood inference and prepare the combined partial-likelihood fit used by InferenceSurvivalKKStratCoxPHOneLik.

Returns the model-specific combined-likelihood treatment estimate; see InferenceAsympLik.

Recomputes the one-likelihood stratified Cox treatment estimate under Bayesian-bootstrap weights.

Computes an asymptotic confidence interval for the treatment effect.

Returns a 2-sided p-value for H0: beta_T = delta.

Usage

SurvivalKKStratCoxOneLikPartialLikelihoodSource

Weibull likelihood component source

Description

Initialize inference for the Weibull AFT model \log T_i = \beta_0 + \beta_T W_i + X_i^\top \gamma + \sigma \epsilon_i, \epsilon_i \sim standard extreme-value (so T_i is marginally Weibull); see InferenceSurvivalWeibullRegr for the model form. Does not fit the model; the fit is deferred to the first call to compute_estimate() or a method that requires it.

Fits the Weibull AFT model by maximum likelihood and returns the log-time-ratio estimate \hat\beta_T. Handles right-, left-, and interval-censored observations via their appropriate survival/density likelihood contributions.

Recomputes the treatment estimate under subject/block-level Bayesian-bootstrap weights via weighted_weibull_bootstrap_surrogate_fit(), a fast weighted Weibull surrogate fit, as an approximation to the weighted AFT likelihood. Only supported for ordinary right-censored data: throws an error for left-/interval-censored designs, since the surrogate fit assumes ordinary right-censoring semantics.

Wald confidence interval for the log-time-ratio \beta_T using the fitted model's standard error; see InferenceAsymp for the shared Wald contract. Fits the model first if not already cached.

Two-sided Wald test of H_0: \beta_T = \code{delta} using the fitted model's standard error; see InferenceAsymp for the shared Wald contract. Fits the model first if not already cached.

Bayesian-bootstrap two-sided p-value; see Inference. Blocked outright (rather than letting the underlying weighted re-estimate fail once per replicate, inside a per-iteration tryCatch that silently converts the stop() in compute_estimate_with_bootstrap_weights() into an all-NA bootstrap distribution) under left-/interval-censored data.

Bartlett-corrected likelihood-ratio two-sided p-value; see Inference. Blocked outright under left-/ interval-censored data for the same reason as compute_bayesian_bootstrap_two_sided_pval(): the underlying simulate_under_lik_null() guard would otherwise only surface as a silently all-NA calibration distribution.

Randomization-test two-sided p-value; see Inference. A nonzero null shift (delta != 0) is blocked outright under left-/interval-censored data for the same reason as the other bootstrap guards in this class: the underlying shift-template guard in setup_randomization_template_and_shifts() would otherwise only surface as a silently all-NA randomization distribution.

Usage

SurvivalWeibullLikelihoodSource

Details

Source list for the SurvivalWeibullLikelihood component composed by InferenceSurvivalWeibullRegr.


Bisection loop for computing confidence interval bounds by inverting randomization tests

Description

This function implements the bisection algorithm to find CI bounds by inverting the randomization test. It repeatedly calls the p-value computation function until convergence.

Usage

bisection_ci_loop_cpp(
  pval_fn,
  r,
  l,
  u,
  pval_th,
  tol,
  transform_responses,
  lower
)

Arguments

pval_fn

R function that computes two-sided p-value given delta

r

Number of randomization iterations

l

Initial lower bound

u

Initial upper bound

pval_th

P-value threshold (typically alpha/2 for two-sided CI)

tol

Tolerance for convergence (in p-value space)

transform_responses

String: "none", "log", or "logit"

lower

Logical: TRUE for lower CI bound, FALSE for upper

Value

The CI bound value


Sequential computation of both CI bounds (called from R-level parallelism)

Description

This function is kept for backwards compatibility but the outer parallelism (running lower/upper bounds simultaneously) is now handled at the R level via parallel::mclapply, which is safe. Calling R functions from OpenMP threads is undefined behaviour in R and caused process crashes.

Usage

bisection_ci_parallel_cpp(
  pval_fn,
  r,
  l_lower,
  u_lower,
  l_upper,
  u_upper,
  pval_th,
  tol,
  transform_responses,
  num_cores = 1L
)

Arguments

pval_fn

R function: pval_fn(nsim, delta, transform_responses, num_cores)

r

Number of randomization iterations

l_lower

Initial lower bound for lower CI bound search

u_lower

Initial upper bound for lower CI bound search (typically the estimate)

l_upper

Initial lower bound for upper CI bound search (typically the estimate)

u_upper

Initial upper bound for upper CI bound search

pval_th

P-value threshold (typically alpha/2 for two-sided CI)

tol

Tolerance for convergence (in p-value space)

transform_responses

String: "none", "log", or "logit"

num_cores

Passed through to pval_fn for inner parallelism

Value

Numeric vector of length 2: [lower_bound, upper_bound]


Single-threaded helper for computing one CI bound

Description

Single-threaded helper for computing one CI bound

Usage

bisection_ci_single_bound_cpp(
  pval_fn,
  r,
  l,
  u,
  pval_th,
  tol,
  transform_responses,
  lower,
  num_cores = 1L
)

Build a Reusable Unstratified Cox Data Cache (C++ Backend)

Description

Precomputes and caches the sorted risk-set structure needed to evaluate the Cox partial-likelihood, score, and Hessian, so that repeated Newton-Raphson or L-BFGS fits on the same (X, y, dead) data (e.g. across bootstrap/randomization replicates of the treatment column, or successive estimate_only vs. full-variance calls) do not repeat the O(n \log n) sort and event-time tabulation on every call. This is the unstratified counterpart of build_stratified_cox_data_cache_cpp; the returned cache is consumed by fast_coxph_regression_prebuilt_cpp, which implements the same Cox partial-likelihood as fast_coxph_regression and fast_coxph_regression_cpp.

Usage

build_cox_data_cache_cpp(X, y, dead)

Arguments

X

A numeric matrix of predictor variables (no intercept column); n rows, one per subject.

y

A numeric vector of length n giving the observed (event or censoring) time for each subject.

dead

A numeric vector of length n with values in {0, 1} indicating event status (1 = event/death, 0 = right-censored).

Details

Model. No fitting happens here; this function only prepares the single risk set (stratum) used by the Breslow-tied Cox partial likelihood

\ell(\beta) = \sum_{k} \Big[ \big(\textstyle\sum_{i \in D_k} x_i\big)^\top \beta - d_k \log\!\big(\textstyle\sum_{j \in R_k} e^{x_j^\top \beta}\big) \Big],

where D_k is the set of subjects with an event at the k-th unique event time and R_k is the risk set (all subjects with y >= that event time). See fast_coxph_regression for the full model, scale, and optimizer contract; this page documents only the cache construction.

What is cached. Internally builds a single CoxData record that: (1) sorts subjects by ascending y (observed/censoring time), breaking ties by placing events (dead == 1) before censored observations at the same time; (2) stores the sort-permuted y, dead, and row-major copy of X; and (3) tabulates the vector of unique event times and, for each, the number of tied events (event_counts), which drives the Breslow tie-handling correction in the partial-likelihood, score, and Hessian.

Input conventions. X is an n \times p design matrix with one row per subject and no intercept column (Cox models are fit on the partial likelihood, which has no intercept). y is the observed time (event or censoring time), and dead is a 0/1 event indicator (1 = event, 0 = right-censored); both must have length n matching nrow(X). Tied event times are handled via the Breslow approximation (via event_counts), not the exact (Cox) or Efron method. There is no support for left- or interval-censored data at this layer; classes with more general censoring (see supports_interval_or_left_censored_data() in InferenceEngine) bypass this cache and dispatch to icenReg instead. NA/non-finite values in X, y, or dead and negative values of y are not checked or handled at this layer; callers are responsible for filtering or imputing upstream.

Return value and object lifetime. Returns an externalptr (Rcpp::XPtr) wrapping a heap-allocated std::vector<CoxData> of length 1 (a single unstratified stratum), so that fast_coxph_regression_prebuilt_cpp can share the same stratified interface as the stratified cache. The pointer owns its memory (finalizer registered via XPtr(..., true)) and is freed automatically by R's garbage collector; it must not be serialized (e.g. via saveRDS) or reused after the R session that created it exits. The cache is immutable once built and safe to reuse across many calls to fast_coxph_regression_prebuilt_cpp as long as X, y, and dead have not changed; callers (e.g. InferenceCoxPH$private$cox_data_cache) are responsible for invalidating and rebuilding the cache when the treatment assignment or covariates change.

Complexity. O(n \log n) for the sort plus O(n) for the tabulation pass, where n is the number of subjects; memory use is O(np) for the row-major copy of X plus O(n) for the sorted y/dead and event-time tables.

Value

An externalptr to a cached, sorted Cox risk-set representation of (X, y, dead), for use as the cox_data_xptr argument of fast_coxph_regression_prebuilt_cpp.

See Also

build_stratified_cox_data_cache_cpp for the stratified (multiple risk-set) analog, fast_coxph_regression_prebuilt_cpp for the fitting routine that consumes this cache, and fast_coxph_regression for the full Cox model documentation, including the partial-likelihood derivation, tie-handling, and references. Analogous Python API: lifelines CoxPHFitter.


Build a Reusable Stratified Cox Data Cache (C++ Backend)

Description

Precomputes and caches, per stratum, the sorted risk-set structure needed to evaluate the stratified Cox partial-likelihood, score, and Hessian, so that repeated Newton-Raphson or L-BFGS fits on the same (X, y, dead, strata) data (e.g. across bootstrap/randomization replicates of the treatment column) do not repeat the per-stratum sort and event-time tabulation on every call. This is the stratified counterpart of build_cox_data_cache_cpp; the returned cache is consumed by fast_coxph_regression_prebuilt_cpp, which implements the same partial-likelihood family as fast_coxph_regression and fast_coxph_regression_cpp, extended to sum the log-partial- likelihood, score, and Hessian across independent strata-specific risk sets while sharing a single regression coefficient vector \beta across all strata (a shared-\beta, per-stratum-baseline-hazard stratified Cox model).

Usage

build_stratified_cox_data_cache_cpp(X, y, dead, strata)

Arguments

X

A numeric matrix of predictor variables (no intercept column); n rows, one per subject.

y

A numeric vector of length n giving the observed (event or censoring) time for each subject.

dead

A numeric vector of length n with values in {0, 1} indicating event status (1 = event/death, 0 = right-censored).

strata

An integer vector of length n giving the stratum (block/cluster) label of each subject. Risk sets are formed separately within each distinct label.

Details

Model. No fitting happens here; this function only partitions subjects into per-stratum risk sets and prepares each one for the Breslow-tied stratified Cox partial likelihood

\ell(\beta) = \sum_{s} \sum_{k \in s} \Big[ \big(\textstyle\sum_{i \in D_{sk}} x_i\big)^\top \beta - d_{sk} \log\!\big(\textstyle\sum_{j \in R_{sk}} e^{x_j^\top \beta}\big) \Big],

where the outer sum runs over strata s (distinct baseline hazards), D_{sk} is the set of subjects with an event at the k-th unique event time within stratum s, and R_{sk} is the corresponding within-stratum risk set (subjects in stratum s with y >= that event time). Risk sets never cross strata, so subjects are only ever compared to other subjects in the same stratum; \beta is shared across strata while the baseline hazard is allowed to differ arbitrarily by stratum. See fast_coxph_regression for the unstratified model, scale, and optimizer contract; this page documents only the per-stratum cache construction.

What is cached. Subjects are grouped into strata by their integer strata label (via an ordered std::map<int, ...>, so strata are processed and stored in ascending label order). Within each stratum, a CoxData record is built exactly as in build_cox_data_cache_cpp: subjects are sorted by ascending y, ties broken by placing events before censored observations, and the per-stratum unique event times and tied-event counts (event_counts) are tabulated to drive the within-stratum Breslow tie-handling correction.

Input conventions. X is an n \times p design matrix with one row per subject and no intercept column. y is the observed time (event or censoring time) and dead is a 0/1 event indicator (1 = event, 0 = right-censored); both have length n matching nrow(X). strata is a length-n integer vector of stratum (block/cluster) labels; labels need not be contiguous or start at 1, and a stratum with a single subject or with no observed events contributes an empty or degenerate risk set that carries no information to the partial likelihood, score, or Hessian (its rows are still stored, but they will not affect the fit). Callers are expected to have already dropped uninformative strata upstream (see get_informative_rows() in the stratified-Cox inference class) when that matters for numerical stability; this function does not filter strata itself. As with build_cox_data_cache_cpp, tied event times use the Breslow approximation, there is no support for left- or interval-censored data at this layer, and NA/non-finite values in X, y, dead, or strata are not checked or handled here.

Return value and object lifetime. Returns an externalptr (Rcpp::XPtr) wrapping a heap-allocated std::vector<CoxData> with one element per distinct stratum label (ordered ascending by label), consumed by fast_coxph_regression_prebuilt_cpp exactly like the length-1 vector returned by build_cox_data_cache_cpp. Lifetime, garbage-collection, and non-serializability semantics are identical to build_cox_data_cache_cpp: the cache is immutable once built and safe to reuse across repeated fits as long as X, y, dead, and strata have not changed; callers (e.g. InferenceStratifiedCoxPH$private$strat_cox_data_cache) are responsible for invalidating and rebuilding it when the treatment assignment, covariates, or stratification changes.

Complexity. O(n \log(n/S)) for the per-stratum sorts plus O(n) for the tabulation passes, where n is the number of subjects and S is the number of distinct strata; memory use is O(np) for the per-stratum row-major copies of X plus O(n) for the sorted y/dead and event-time tables.

Value

An externalptr to a cached, per-stratum sorted Cox risk-set representation of (X, y, dead, strata), for use as the cox_data_xptr argument of fast_coxph_regression_prebuilt_cpp.

See Also

build_cox_data_cache_cpp for the unstratified (single risk-set) analog, fast_coxph_regression_prebuilt_cpp for the fitting routine that consumes this cache, and fast_coxph_regression for the full Cox model documentation, including the partial-likelihood derivation, tie-handling, and references. Analogous Python API: lifelines CoxPHFitter (stratified fits via the strata argument) and statsmodels duration models.


Check Whether a Suggested Package Is Installed (Memoized)

Description

Tests whether package_name is installed via requireNamespace and memoizes the result in a package-level environment (package_cache), so that repeated checks for the same package within an R session pay the namespace-lookup cost only once. EDI uses this to guard optional code paths that depend on Suggests-only packages (e.g. quantreg, betareg, nbpMatching, geepack, icenReg) that are not installed automatically with the package, issuing an informative stop()/warning() and falling back to an internal implementation when the dependency is absent, rather than failing with an opaque "could not find function" error.

Usage

check_package_installed(package_name)

Arguments

package_name

Character scalar. The name of the package to check (as passed to requireNamespace()).

Details

Caching / mutation semantics. This function has a side effect: on the first call for a given package_name in the current R session, it assigns the boolean result of requireNamespace() into the package-global environment package_cache, keyed by package_name. All subsequent calls for the same package_name (from anywhere in the package, or from user code) read the cached value directly and do not re-query the namespace registry. This means the result reflects whether the package was installed at the time of the first call; installing or removing package_name later in the same R session will not be picked up. The cache is a plain environment (not an R6 object) shared by all callers within the process and is not reset between calls; it is, however, re-initialized fresh in each new R session.

Determinism. For a fixed installed-package state, check_package_installed() is deterministic and side-effect-free beyond the memoization described above; it does not consume random number generator state and involves no numerical computation.

Lifecycle. Internal utility (exported for reuse within the package's own R6 classes across files, not intended as a general-purpose user-facing API); prefer requireNamespace() directly for one-off checks in user code.

Value

Logical scalar. TRUE if the package is installed and its namespace can be loaded, FALSE otherwise.

See Also

requireNamespace, on which this function is a thin memoizing wrapper.


Delete this machine's saved EDI tuning and return to shipped defaults

Description

Removes the per-user config file written by tune_EDI_for_this_machine (so the next library(EDI) starts from the package's built-in performance-policy defaults) and resets the in-session cold-start, warm-start, and optimizer-algorithm dispatch policies to those defaults right away.

Usage

clear_local_EDI_optimization()

Value

Invisibly, TRUE if a saved tuning existed and was removed, FALSE if there was none.

See Also

tune_EDI_for_this_machine, get_local_EDI_optimization.

Examples


clear_local_EDI_optimization()


Conditional logistic regression for matched pairs

Description

Internal method. Replaces bclogit::clogit. For matched pairs (exactly 2 subjects per stratum), the conditional log-likelihood depends only on discordant pairs. Within each discordant pair the contribution reduces to ordinary logistic regression on signed within-pair differences with no intercept.

Usage

clogit_helper(y_m, X_m, w_m, strata_m)

Arguments

y_m

Binary outcome vector (0/1) for matched subjects.

X_m

Covariate matrix or data.frame (may have 0 columns).

w_m

Treatment indicator (0/1) for matched subjects.

strata_m

Integer stratum IDs (pair labels) for matched subjects.

Value

Result list from fast_logistic_regression_with_var: b[1] = beta_T, ssq_b_j = Var(beta_T). NULL on failure.


Fast Bai Adjusted T Statistic for Multiple Permutations

Description

Fast Bai Adjusted T Statistic for Multiple Permutations

Usage

compute_bai_distr_parallel_cpp(
  w_mat,
  m_mat,
  y,
  delta,
  halves_idx,
  convex_flag,
  num_cores
)

Arguments

w_mat

Integer matrix of permuted treatment assignments (n x r).

m_mat

Integer matrix of match indicators (n x r).

y

Numeric response vector.

delta

Null treatment effect shift.

halves_idx

Integer matrix of half-sample indices.

convex_flag

Logical flag for convex combination.

num_cores

Number of OpenMP threads.

Value

Numeric vector of Bai adjusted T statistics.


Randomization/Bootstrap Reference Distribution of the Treatment Log-Hazard-Ratio for a Treatment-Only Cox PH Model (C++ Backend, Single-Covariate)

Description

Builds an empirical reference (null or shifted-null) distribution of the treatment log-hazard-ratio \hat\beta_T from a treatment-only (single-covariate, no other adjustment covariates) Cox proportional-hazards model, by refitting the model on B = ncol(i_mat) pre-generated resample-and-reassign draws. This backs bootstrap randomization test (BRT) inversion for confidence intervals and p-values on the Cox log-hazard-ratio: repeated calls at different delta values (or a single call at delta = 0 for a null/reference distribution) let the caller invert the empirical distribution of \hat\beta_T against a target quantile or tail probability. This is a single-covariate special case; compute_coxph_rand_bootstrap_parallel_cpp (used by fast_coxph_regression's survival inference class) generalizes this to models with additional adjustment covariates and to Gaussian smoothing noise on the resampled log-hazard-ratio, and is the version actually wired into InferenceCoxPH; this treatment-only function currently has no in-package caller and should be treated as a lighter-weight standalone utility or superseded building block rather than part of the primary inference path.

Usage

compute_coxph_rand_bootstrap_cpp(y0, dead, i_mat, w_mat, delta, num_cores)

Arguments

y0

Numeric vector of original survival times (event or censoring time), length n; not itself resampled — i_mat indexes into this vector per draw.

dead

Numeric vector of length n with values in {0, 1} giving the event indicator (1 = event, 0 = right-censored) for each original row, indexed by i_mat per draw.

i_mat

Integer matrix (n \times B) of 1-based row indices into y0/ dead, one resampled dataset per column.

w_mat

Integer matrix (n \times B) of treatment assignments in {0, 1}, one resampled assignment vector per column, aligned with the corresponding column of i_mat.

delta

Sharp-null log-time shift applied multiplicatively (e^\delta) to the working survival time of treated (w == 1) resampled subjects; delta = 0 leaves times unshifted. The shift is applied to event and censoring times alike, with dead carried over unchanged — the residual construction of rank-based AFT inference (Tsiatis 1990, doi:10.1214/aos/1176347504; Wei, Ying and Lin 1990, doi:10.1093/biomet/77.4.845; Jin, Lin, Wei and Ying 2003, doi:10.1093/biomet/90.2.341), valid under independent censoring and exact in finite samples only when censoring times share the accelerated clock. Note that \delta is a log time-ratio while the returned \hat\beta_T is a log hazard ratio; the two coincide only under a parametric link the Cox model does not supply (see InferenceSurvivalStratCoxPHRegr's refusal of the randomization CI and package_metadata/new_feature_plans/randomization_ci_construction_audit.md).

num_cores

Number of OpenMP threads to use for parallelizing across draws (ignored, and draws run sequentially, when the package is built without OpenMP support).

Details

Per-draw model. For each draw b = 1, \dots, B, this function forms a resampled dataset of size n = length(y0) by taking row indices i_mat[, b] (1-based, into the original y0/dead) and treatment labels w_mat[, b], applies the sharp-null time shift (see below), and fits an unstratified, single-covariate (treatment-only) Cox partial-likelihood model — the p = 1 case of the model documented in build_cox_data_cache_cpp and fast_coxph_regression — via Newton-Raphson (cox_fit() with estimate_only = true, maxit = 20, tol = 1e-9, no warm start, smart_cold_start = false). Only the fitted coefficient \hat\beta_T (the log hazard ratio for treatment) is returned per draw; no variance-covariance matrix is computed (this function is for building a resampling distribution, not for single-fit inference).

Sharp-null shift. delta encodes a sharp null hypothesis of a constant multiplicative shift on the time scale for treated subjects: for draw b, subject i with resampled treatment label w_i \in \{0, 1\}, the working survival time is y_i \cdot e^{\delta} if w_i = 1 and y_i (unchanged) if w_i = 0; dead status is carried over from the original row unchanged. This is an accelerated-failure-time-style sharp null (\delta = 0 recovers the unshifted resample), not a proportional-hazards sharp null; delta is on the same log-time scale used elsewhere in the package's AFT/Weibull machinery, not the log-hazard-ratio scale of the returned \hat\beta_T.

Resampling scheme. i_mat and w_mat are assumed pre-generated by the caller (e.g. via the package's bootstrap or randomization-draw machinery) and are not validated here: i_mat need not be a permutation (indices may repeat, as in a nonparametric bootstrap draw with replacement, or may be a permutation, as in a randomization test) and w_mat need not respect any particular design's assignment-probability structure — whatever exchangeability or randomization-validity properties the resulting reference distribution has are entirely a property of how the caller generated i_mat/w_mat, not of this function.

Non-convergence and missingness. A draw's entry in the output is NA if the Newton-Raphson fit fails to converge within maxit iterations or the fitted coefficient is non-finite; callers must handle NA entries (e.g. by omission) when computing empirical quantiles or tail probabilities from the returned vector.

Parallelism and reproducibility. When compiled with OpenMP support and num_cores > 1, draws are processed in parallel via #pragma omp parallel for schedule(dynamic); each draw is a self-contained fit with no shared mutable state across draws (aside from writing to disjoint output slots), so results are deterministic given i_mat/w_mat regardless of the number of threads or scheduling order — this function consumes no RNG state itself, since the randomness lives entirely in how the caller generated i_mat and w_mat.

Complexity. O(B \cdot n \log n) for the B independent single- covariate Cox fits (dominated by the per-draw CoxData sort), parallelized across num_cores threads when available; memory use is O(n) per in-flight draw.

Value

Numeric vector of length B = ncol(i_mat) with the fitted treatment log-hazard-ratio \hat\beta_T for each draw, or NA_real_ for draws whose Cox fit failed to converge.

See Also

fast_coxph_regression for the underlying Cox partial-likelihood model and its references; build_cox_data_cache_cpp for the per-draw sorted risk-set construction each draw performs internally. See also randomization test and bootstrap for background on resampling-based reference distributions.


Fast KK Wilcoxon Statistic for Multiple Permutations

Description

Fast KK Wilcoxon Statistic for Multiple Permutations

Usage

compute_matching_wilcox_distr_parallel_cpp(
  w_mat,
  m_mat,
  y,
  delta,
  transform_code,
  zero_one_logit_clamp,
  is_fixed_matching,
  num_cores
)

Arguments

w_mat

Integer matrix of permuted treatment assignments (n x r).

m_mat

Integer matrix of match indicators (n x r).

y

Numeric response vector.

delta

Null treatment effect shift.

transform_code

Integer code for response transformation.

zero_one_logit_clamp

Clamp value for logit transformation.

is_fixed_matching

Logical flag for fixed matching designs.

num_cores

Number of OpenMP threads.

Value

Numeric vector of KK Wilcoxon statistics.


Parallel Stereotype Logit Randomization Distribution

Description

Parallel Stereotype Logit Randomization Distribution

Usage

compute_stereotype_logit_distr_parallel_cpp(X, y, w_mat, delta, num_cores)

Arguments

X

Matrix of covariates (without intercept or treatment).

y

Numeric vector of response values (pre-null-shifted for treated).

w_mat

Integer matrix of permuted treatment assignments (n x nsim).

delta

Null treatment effect (additive shift).

num_cores

Number of OpenMP threads.

Value

Numeric vector of length nsim with treatment coefficients.


Parallel BRT kernel for KM-diff (median) and RMST-diff. Each replicate resamples rows i_mat(.,b) and pairs them with assignment w_mat(.,b). Sharp-null shift is multiplicative on treated times (exp(delta)). Uses an inline pure-C++ KM calculator — no R objects inside the loop, so OpenMP is safe.

Description

Parallel BRT kernel for KM-diff (median) and RMST-diff. Each replicate resamples rows i_mat(.,b) and pairs them with assignment w_mat(.,b). Sharp-null shift is multiplicative on treated times (exp(delta)). Uses an inline pure-C++ KM calculator — no R objects inside the loop, so OpenMP is safe.

Usage

compute_survival_stat_diff_rand_bootstrap_parallel_cpp(
  y0,
  dead,
  i_mat,
  w_mat,
  delta,
  do_rmst,
  noise_mat,
  num_cores
)

Arguments

do_rmst

TRUE for RMST-diff, FALSE for median (KM-diff).


Compute automatic survival strata IDs from low-cardinality covariates

Description

Selects numeric covariate columns with a small number of observed levels and combines them into a single all-subject stratum identifier. This is used to support automatic stratified Cox models when the design object stores observed covariates but no explicit stratum variable.

Usage

compute_survival_strata_ids_cpp(
  X,
  max_unique_per_col = 4L,
  max_strata_cols = 4L,
  min_count_per_level = 2L
)

Arguments

X

Numeric covariate matrix.

max_unique_per_col

Maximum number of unique values allowed for a column to be considered a stratification candidate.

max_strata_cols

Maximum number of candidate columns to combine.

min_count_per_level

Minimum frequency required for every level in a candidate column.

Value

A list with 'strata_id', 'selected_cols', and 'num_strata'.


Fast Wilcoxon HL Statistic for Multiple Permutations

Description

Fast Wilcoxon HL Statistic for Multiple Permutations

Usage

compute_wilcox_hl_distr_parallel_cpp(
  w_mat,
  y,
  delta,
  transform_code,
  zero_one_logit_clamp,
  num_cores
)

Arguments

w_mat

Integer matrix of permuted treatment assignments (n x r).

y

Numeric response vector.

delta

Null treatment effect shift.

transform_code

Integer code for response transformation.

zero_one_logit_clamp

Clamp value for logit transformation.

num_cores

Number of OpenMP threads.

Value

Numeric vector of HL statistics.


Build an Intercept-Free, Full-Rank Covariate Design Matrix from a Formula

Description

Expands formula against data (via model.matrix) into a purely numeric covariate design matrix suitable for the package's own fast_* GLM/survival/ordinal fitting routines, which manage their own intercept and treatment columns separately rather than relying on the formula/model-matrix machinery for them. This is the standard covariate-matrix builder used throughout EDI's inference classes (e.g. Inference$private$X) whenever adjustment covariates need to go from a user-facing formula/data-frame representation to a numeric matrix the C++ backends can consume.

Usage

create_model_matrix_from_features(formula, data)

Arguments

formula

A formula object giving the covariate specification to expand (e.g. ~ age + sex + age:sex); should not include the response.

data

A data frame or data table supplying the variables referenced in formula, with one row per subject.

Details

What it does. (1) If data has zero columns, returns a numeric nrow(data) x 0 matrix immediately (no covariates to expand). (2) Otherwise calls model.matrix(formula, data = data), which performs standard formula expansion: factor variables are dummy-coded against their reference level (the first level of levels, or the level ordering already present in data), interactions (a:b, a*b) are expanded to product columns, and any model.matrix contrasts option in effect at call time applies. (3) If the first resulting column is named "(Intercept)" (i.e. the formula was not given - 1 / + 0), that column is dropped — this function always returns a covariate-only matrix with no intercept column, since EDI's design and inference classes add their own intercept/treatment columns at a fixed position. (4) The result is passed through drop_linearly_dependent_cols, which detects the numeric rank of the matrix (via matrix_rank_cpp() at tolerance 1e-7) and, if the matrix is rank-deficient, greedily retains a full-rank subset of columns using the pivot order from qr(M, tol = 1e-7) (dropping the same tolerance's worth of redundant/aliased columns, e.g. from collinear dummy expansions or an over-specified interaction structure); this rank-reduction step is silent — no warning is issued when columns are dropped, and the dropped columns' identity/names are not returned to the caller, only the reduced matrix.

Input conventions. data is expected to already be free of missing values at call time (imputation, when configured, happens upstream in the design/ inference class before this function is called); this function does not impute or warn about NAs, and model.matrix will drop incomplete rows or error, per its own na.action default, if NAs remain. Column order and names in the returned matrix follow model.matrix's expansion order (all factor/interaction columns for a term before the next term), possibly reduced by the rank-deficiency step; callers relying on stable column identity (e.g. warm-starting coefficients across calls) should not assume the set or order of columns is invariant if data's factor levels or rank change between calls.

Failure semantics. If drop_linearly_dependent_cols detects the matrix is non-numeric or contains non-finite values, it returns the matrix unchanged (rank reduction is skipped rather than erroring); a downstream fitting routine operating on a rank-deficient or non-finite design matrix may then fail to converge or report non-finite coefficients/standard errors, which is where such problems will actually surface to the user.

Value

A numeric matrix with nrow(data) rows and one column per retained, full-rank expanded covariate term (no intercept column). Has zero columns if data has zero columns.

See Also

model.matrix, which performs the formula expansion this function wraps; drop_linearly_dependent_cols (internal, same file) for the rank-deficiency cleanup step. Analogous Python API: patsy/ statsmodels formula API for formula-based design matrix construction.


Return EDI Build Information (C++ Backend)

Description

Returns the compiler and package build metadata that was baked into the currently loaded EDI shared object at compile time (via preprocessor macros defined in edi_build_flags.h, generated by the package's build tooling), not anything queried live from the running system. This is intended for benchmark reports and reproducibility audits where the exact compiler, flags, and build environment used to produce the installed binary matter (e.g. explaining performance differences between two installations of the same package version, or confirming a binary was built with a particular optimization/vectorization configuration).

Usage

edi_build_info_cpp()

Details

Every field is a fixed string or boolean baked in when the C++ source was compiled (via preprocessor macro substitution); calling this function multiple times within the same R session always returns identical values, and the values reflect the build environment, not the environment the function happens to be called from. If EDI is reinstalled/recompiled, existing R sessions that already loaded the old shared object continue to report the old build's metadata until they restart and load the new one.

Value

A named list with the following fields, all length-1 character strings unless noted otherwise:

capture_method

How the build metadata below was captured by the package's build tooling (EDI_BUILD_CAPTURE_METHOD).

build_timestamp

Timestamp of compilation (EDI_BUILD_TIMESTAMP).

build_host

Hostname of the machine that compiled this binary (EDI_BUILD_HOST).

r_home

R_HOME of the R installation used to build the package.

r_version

R version string used to build the package.

r_cxx20

The C++20 compiler command R's build configuration reported.

r_cxx20std

The C++ standard flag (e.g. -std=gnu++20) used.

r_cxx20flags

Additional C++20 compiler flags from R's configuration.

r_shlib_openmp_cxxflags

OpenMP C++ flags R's configuration supplies for shared-library builds (empty if OpenMP support was not available/enabled).

env_edi_portable, env_edi_disable_vectorization, env_edi_native_speed, env_edi_native_lto

The values (or default/unset indicator) of the corresponding EDI_* environment variables that were set at build time to control portable-vs-native code generation, vectorization, and link-time optimization; see the package's build documentation for what each controls.

pkg_cppflags, pkg_cxxflags, pkg_libs

The package-level PKG_CPPFLAGS/ PKG_CXXFLAGS/PKG_LIBS used when compiling/linking EDI's C++ sources.

compiler

The compiler identification string (__VERSION__).

compiler_optimize_macro

Logical; TRUE iff the compiler defined __OPTIMIZE__ (i.e. optimizations were enabled) when this translation unit was compiled.

compiler_fast_math_macro

Logical; TRUE iff __FAST_MATH__ was defined (i.e. non-IEEE-compliant fast-math optimizations were enabled), which can affect floating-point reproducibility/edge-case behavior (NaN/Inf handling, exact rounding) relative to a standard-compliant build.

eigen_dont_vectorize_macro

Logical; TRUE iff EIGEN_DONT_VECTORIZE was defined, disabling Eigen's SIMD vectorization for this build (e.g. for portability to CPUs lacking the vector instructions a native build would target).

Examples

info = edi_build_info_cpp()
info$pkg_cxxflags

Re-bind already-installed lazy-component methods on a freshly cloned Inference object to that clone's own self/private.

Description

install_lazy_inference_component() permanently binds each real (non-stub) implementation it installs to whichever object triggered the install, via environment(value) = parent.frame(). R6's clone() correctly rebinds every method present in the class generator's original method list, but a lazily-installed method is injected into private/self at runtime and is invisible to that bookkeeping, so a clone keeps calling back into the ORIGINAL object's data (e.g. a Bayesian-bootstrap worker clone silently reading the pre-clone object's current_bayesian_bootstrap_context, always NULL, instead of its own). Call this right after self$clone() to repoint every already-installed lazy-component method (public and private) at the clone's own enclosing environment; state fields (owns_state) are left untouched since clone() already copies their current values correctly.

Usage

edi_rebind_lazy_components_after_clone(i, source_private = NULL)

Arguments

i

The freshly cloned Inference object.

source_private

The pre-clone source object's own private environment, if available (NULL if not). clone() does not preserve environment-level attributes, so the "already installed" marker for a lazy component installed while the private environment was locked cannot be read from the clone's own private environment; when supplied, it is also read from source_private so those attribute-only markers are not missed. See the implementation comment below for the full mechanism.


Exact Two-Group Jonckheere-Terpstra Test via Full Randomization Enumeration (C++ Backend)

Description

Computes the exact randomization-distribution p-value and a probabilistic-index effect size for the two-group Jonckheere-Terpstra statistic — which, with exactly two groups (w in {0, 1}), coincides with the Wilcoxon-Mann-Whitney U statistic generalized to handle ties (repeated ordinal levels in y):

U = \sum_{k} t_k \big(2 L_k + (n_k - t_k)\big) / 2,

summed over the K distinct observed levels of y (in increasing order), where n_k is the total count at level k, t_k is the observed count of w == 1 subjects at level k, and L_k is the number of subjects at strictly lower levels (this is the standard "number of favorable comparisons" Mann-Whitney statistic, adapted for tied/grouped ordinal data — a tie at the same level contributes 1/2 rather than 0 or 1). The exact (not asymptotic, not Monte Carlo) null/reference distribution of this statistic under the sharp null of no treatment effect is obtained by enumerating, via a dynamic-programming recursion over levels (recurse_jt_distribution()), every way to distribute n_treat "treated" labels among the n subjects consistent with the fixed per-level totals n_k — i.e. the exact multivariate hypergeometric randomization distribution of the statistic conditional on the observed level counts, with each configuration's probability computed in log-space from log-binomial-coefficient weights to avoid overflow for larger n.

Usage

exact_jonckheere_terpstra_pval_cpp(y, w)

Arguments

y

Integer (or integer-coercible) vector of length n giving each subject's ordinal response value; ties (repeated values) are handled via the grouped Mann-Whitney formula above. Must not contain NA.

w

Integer (or integer-coercible) vector of length n with values in {0, 1} giving each subject's group membership; both groups must be non-empty and no NA is permitted.

Details

Randomization test, not a model-based test. This is a randomization (permutation) test in the Fisherian sense: it conditions on the observed marginal level counts n_k and the group sizes n_treat/n_control, and asks how extreme the observed statistic is relative to every other way those same labels could have been randomly assigned, so its validity does not depend on any distributional assumption about y beyond exchangeability under the null. The two-sided p-value is p_{\mathrm{exact}} = \min(1, 2 \min(p_{\mathrm{lower}}, p_{\mathrm{upper}})), where p_{\mathrm{lower}}/p_{\mathrm{upper}} are the exact one-sided tail probabilities of the randomization distribution at or below / at or above the observed statistic.

Effect size (superiority). superiority is the probabilistic index \Pr(Y_T > Y_C) + \tfrac{1}{2}\Pr(Y_T = Y_C) (equivalently U rescaled to [0, 1] by dividing by n_{\mathrm{treat}} \cdot n_{\mathrm{control}}), the probability a randomly chosen treated subject's ordinal outcome exceeds a randomly chosen control subject's, counting ties as half a win; 0.5 indicates no stochastic ordering between groups, and 1/0 indicate the treated group's outcomes are uniformly higher/lower.

Input conventions. y is coerced to integer and treated as an ordinal (or any orderable-by-integer-value) response with an arbitrary number of tied levels; w must be an integer/coercible-to-integer vector of {0, 1} values with both groups non-empty. NA in either y or w is not permitted and raises an error, as does a non-{0,1} value in w or an empty input.

Complexity. The recursion's state space scales with the number of distinct possible statistic values (O(n_{\mathrm{treat}} \cdot n_{\mathrm{control}}) many), and thread-local buffers are reused (not reallocated) across repeated calls within the same thread for the same or smaller problem sizes; this exact enumeration is exponential in the number of distinct levels/group sizes in the worst case (unlike an asymptotic normal-approximation JT test), so it is intended for small-to- moderate n where exactness matters more than raw speed.

Value

A list with components stat2 (twice the observed U statistic, an integer, used internally to keep the enumeration in integer arithmetic), n_treat, n_control (the two group sizes), superiority (the probabilistic-index effect size described above), p_lower, p_upper (the exact one-sided randomization-distribution tail probabilities), and p_exact (the exact two-sided p-value).

See Also

Mann-Whitney U test and Jonckheere's trend test for background; analogous Python API: SciPy mannwhitneyu (method="exact" for the same exact-enumeration approach, though SciPy's exact path does not handle ties the same way).


Expand Ordinal Data into Stacked Binary Comparisons for Adjacent-Category Logit Regression (C++ Backend)

Description

Reshapes an ordinal response y (levels 1, \dots, K) into the stacked binary-outcome, per-cut-stratified form required to fit an adjacent-category logit model as a single conditional (stratified) logistic regression, so the package's existing binary/conditional-logit fitting backends can be reused unchanged for ordinal adjacent-category models rather than needing a bespoke ordinal solver.

Usage

expand_adjacent_category_data_cpp(y, w, strata, K)

Arguments

y

Integer vector of length n: each subject's ordinal category label, in 1:K.

w

Integer vector of length n: a covariate (typically treatment assignment) carried through unchanged into each stacked row for that subject.

strata

Integer vector of length n: positive-integer stratum/block labels; max(strata) is used as the per-cut stratum-ID offset (see Details).

K

Integer; the number of ordinal categories (so there are K - 1 adjacent-category cuts).

Details

Model. The adjacent-category logit model compares each pair of consecutive categories j and j+1 (j = 1, \dots, K-1) via

\log\frac{\Pr(Y = j+1 \mid Y \in \{j, j+1\})}{\Pr(Y = j \mid Y \in \{j, j+1\})} = \alpha_j + \beta^\top x,

i.e. a logistic model for "category j+1 vs. category j" fit using only the subjects actually observed in one of those two categories, with a cut-specific intercept \alpha_j and covariate effects \beta constrained equal across all K-1 cuts (the proportional/parallel adjacent-category assumption). This differs from the cumulative-logit (proportional-odds) model, which instead compares Y \le j vs. Y > j using every subject at every cut.

Expansion mechanics. For each subject i and each cut j = 1, \dots, K-1 (n_alpha = K - 1), a stacked row is emitted only if y[i] equals j or j+1; subjects at any other level contribute nothing to that cut's comparison (so each subject contributes to at most 2 of the K-1 cuts: the ones immediately adjacent to their observed level, and exactly 1 cut if at an extreme level). The stacked binary outcome is 1 if y[i] == j + 1 (upper category) and 0 if y[i] == j (lower category). The stacked stratum ID is strata[i] + (j - 1) * num_strata (where num_strata = max(strata)), i.e. the original stratum crossed with the cut index j: fitting a conditional logistic regression stratified on this combined ID and pooling all stacked rows together estimates a single shared treatment coefficient \beta across all cuts, while allowing each (original stratum, cut) combination to absorb its own nuisance intercept via strata conditioning (the same stratified-conditional-logit trick used elsewhere in the package, e.g. for continuation-ratio models via expand_continuation_ratio_data_cpp()).

Input conventions. y must take integer values in 1:K (1-based category labels); w is typically the treatment indicator/covariate to estimate a coefficient for, passed through unchanged per stacked row (not itself expanded/transformed); strata must be positive integers with max(strata) == num_strata (no gaps assumed beyond that maximum, since combined stratum IDs are computed by simple integer arithmetic on num_strata, not by re-indexing distinct values). No input validation is performed at this layer (no range/type checks on y/strata); passing out-of-range values silently produces incorrect stratum IDs or drops rows rather than erroring.

Value

A list with components y (stacked 0/1 binary outcome), w (stacked covariate, passed through unchanged), and strata (stacked combined stratum-by-cut ID); all three are integer vectors of the same, generally-longer-than-n length (each subject contributes 0, 1, or 2 stacked rows depending on their observed category).

See Also

expand_continuation_ratio_data_cpp() for the analogous expansion used by continuation-ratio ordinal models. Ordinal regression for orientation; analogous Python API: statsmodels discrete models (no direct adjacent-category equivalent; the closest analog is fitting the expanded data as a conditional/grouped logit).


Expand Ordinal Data into Stacked Binary Comparisons for Continuation-Ratio Regression (C++ Backend)

Description

Reshapes an ordinal response y (levels 1, \dots, K) into the stacked binary-outcome, per-cut-stratified form required to fit a (forward) continuation- ratio logit model as a single conditional (stratified) logistic regression — the discrete-time-hazard analog for ordinal data — so the package's existing binary/conditional-logit fitting backends can be reused unchanged rather than needing a bespoke ordinal solver. This is the continuation-ratio counterpart of expand_adjacent_category_data_cpp(); the two share the same stacking and combined-stratum trick but differ in which rows each subject contributes (see Details).

Usage

expand_continuation_ratio_data_cpp(y, w, strata, K)

Arguments

y

Integer vector of length n: each subject's ordinal category label, in 1:K.

w

Integer vector of length n: a covariate (typically treatment assignment) carried through unchanged into each stacked row for that subject.

strata

Integer vector of length n: positive-integer stratum/block labels; max(strata) is used as the per-cut stratum-ID offset (see Details).

K

Integer; the number of ordinal categories (so there are K - 1 continuation-ratio cuts).

Details

Model. The continuation-ratio model treats reaching each successive category as a sequence of conditional "continue past this cut" events, analogous to a discrete-time survival/hazard model: for cut j = 1, \dots, K-1, among subjects who have reached at least category j (Y \ge j),

\log\frac{\Pr(Y > j \mid Y \ge j)}{\Pr(Y = j \mid Y \ge j)} = \alpha_j + \beta^\top x,

i.e. the log-odds of "continuing" past category j versus "stopping" (being observed) exactly there, given the subject has reached at least j, with a cut-specific intercept \alpha_j and covariate effects \beta constrained equal across cuts (the proportional continuation-ratio assumption). This orientation — numerator is the "continue" event — keeps a positive \beta meaning "pushes toward higher categories of y", matching fast_continuation_ratio_regression_cpp and every other ordinal estimator in the package. Unlike the adjacent-category model (which only compares the two categories immediately flanking a cut), every subject contributes to every cut up to and including the one at which they are observed to stop.

Expansion mechanics. For each subject i with observed category y[i], a stacked row is emitted for every cut j = 1, \dots, \min(\code{y[i]}, K-1): the stacked binary outcome is 0 ("stopped here") if y[i] == j, and 1 ("continued past") for every earlier cut the subject passed through. A subject observed at the top category (y[i] == K) contributes a 1 at every one of the K - 1 cuts (having "survived" all of them without stopping); a subject observed at category j <= K - 1 contributes 1s for cuts 1:(j-1) and a single 0 at cut j, then no further rows (later cuts are irrelevant once a subject has already stopped). As in expand_adjacent_category_data_cpp(), the stacked stratum ID is strata[i] + (j - 1) * num_strata (with num_strata = max(strata)): fitting a conditional logistic regression stratified on this combined ID and pooling all stacked rows estimates a single shared treatment coefficient \beta across all cuts, while each (original stratum, cut) combination absorbs its own nuisance intercept via strata conditioning.

Input conventions. y must take integer values in 1:K; w is passed through unchanged into each stacked row for that subject (typically the treatment indicator/covariate to estimate a coefficient for); strata must be positive integers, with max(strata) used as the per-cut stratum-ID offset. No input validation is performed at this layer.

Value

A list with components y (stacked 0/1 "continued past this cut" outcome), w (stacked covariate, passed through unchanged), and strata (stacked combined stratum-by-cut ID); all three are integer vectors of the same, generally-longer-than-n length (each subject contributes between 1 and K - 1 stacked rows, depending on their observed category).

See Also

expand_adjacent_category_data_cpp() for the analogous expansion used by adjacent-category ordinal models. Ordinal regression for orientation; analogous Python API: statsmodels discrete models (no direct continuation-ratio equivalent; the closest analog is fitting the expanded data as a conditional/grouped logit, or discrete-time survival packages).


Fast Adjacent-Category Logit Regression, Direct MLE (C++ Backend)

Description

Fits the adjacent-category logit ordinal regression model

\log\frac{\Pr(Y = k+1)}{\Pr(Y = k)} = \alpha_k + \beta^\top x, \quad k = 1, \dots, K-1,

by direct maximum likelihood on the full multinomial likelihood of y, rather than via the stacked-binary / stratified-conditional-logit reduction implemented by expand_adjacent_category_data_cpp() elsewhere in the package. \beta (the covariate effects, shared across all K - 1 cuts) and the K - 1 cut-specific intercepts \alpha_k are estimated jointly by numerically optimizing the exact multinomial log-likelihood, which is generally more accurate and can be faster than fitting the row-stacked expansion as a stratified logistic regression, at the cost of a custom (rather than reused) optimizer implementation.

Usage

fast_adjacent_category_logit_cpp(
  X,
  y,
  maxit = 100L,
  tol = 1e-08,
  smart_cold_start = TRUE,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL,
  warm_start_params = NULL,
  warm_start_beta = NULL
)

Arguments

X

A numeric matrix of predictors, n \times p, with no intercept column (the model's cut-specific intercepts \alpha_k serve that role).

y

A numeric vector of length n giving each subject's ordinal category; need not be pre-coded 1:K (see Details for the rank-based remapping).

maxit

Maximum number of optimizer iterations.

tol

Convergence tolerance.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

fixed_idx

Optional integer indices (into the c(alpha, beta) parameter layout described in Details) of parameters to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at; must be the same length as fixed_idx.

optimization_alg

Optimization algorithm; see Details.

warm_start_fisher_info

Optional initial Fisher Information matrix (over the full c(alpha, beta) parameter vector) to warm-start curvature information.

warm_start_params

Optional starting values for the full parameter vector c(alpha, beta). If provided, smart_cold_start is ignored.

warm_start_beta

Optional starting values for just the covariate coefficients \beta (cut intercepts \alpha are still initialized separately). If provided, smart_cold_start is ignored.

Details

Category coding. y need not already be coded 1:K: the distinct values of y are extracted and sorted (get_levels()), and each observation is remapped to its 1-based rank among those sorted distinct values (map_y_to_1K()) — e.g. y = c(10, 30, 20, 10) is treated identically to y = c(1, 3, 2, 1), with K = 3. K is therefore the number of distinct observed values, not any externally supplied category count, and requires at least 2 (an error is raised otherwise).

Parameterization and likelihood. Internally, category probabilities are computed via a numerically stable log-space recurrence: unnormalized log-probabilities \log \tilde p_k = \sum_{j=k}^{K-2} (\alpha_j - \eta) (\eta = x^\top \beta) are accumulated additively — never by exponentiating \alpha_k or -\eta directly — and normalized with a standard log-sum-exp, so every exp() call sees an argument \le 0 and \Pr(Y = k) is bounded to [0, 1] regardless of how extreme \alpha/\eta get during optimization (fixed 2026-08-27: the prior right-to-left product recurrence in terms of raw e^{-\eta} and e^{\alpha_k} could each individually overflow to Inf before normalization, corrupting the objective/gradient to Inf/NaN and leaving the optimizer's line search unable to recover — confirmed via direct testing to reliably exhaust the full iteration budget without converging on ordinary synthetic data at every sample size and seed tried, and to diverge outright to NaN parameters from an all-zero start). The returned neg_loglik is the resulting exact multinomial negative log-likelihood (-\sum_i \log \Pr(Y_i = y_i)), with the analytic gradient computed in the same pass and used internally for optimization. See fast_adjacent_category_logit_with_var_cpp for the variant that additionally returns the variance-covariance matrix of the estimates.

Parameter vector layout. The optimizer's parameter vector (returned as params) is c(alpha_1, ..., alpha_{K-1}, beta_1, ..., beta_p) — the K - 1 cut intercepts first, then the p shared covariate coefficients (p = ncol(X)).

Optimization. Optimized via optimization_alg ("lbfgs" default; see .normalize_optimizer_algorithm for the supported set), for at most maxit iterations at tolerance tol. When no warm start is supplied, smart_cold_start = TRUE (default) seeds the optimizer from an OLS-based initial guess rather than a naive zero/arbitrary start; supplying warm_start_params (the full parameter vector) or warm_start_beta (just the covariate coefficients, with cut intercepts initialized separately) overrides smart_cold_start entirely. fixed_idx/fixed_values allow holding specific parameters (by index into the layout above) fixed at supplied values during optimization rather than estimating them, and warm_start_fisher_info allows reusing a previously computed Fisher information matrix to warm-start curvature information for faster convergence.

Value

A list with components b (the shared covariate coefficients \hat\beta, length p), alpha (the K - 1 estimated cut intercepts \hat\alpha_k), params (the full c(alpha, b) parameter vector, as optimized), neg_loglik (the multinomial negative log-likelihood at convergence), and converged (logical).

See Also

fast_adjacent_category_logit_with_var_cpp for the variance-augmented variant; expand_adjacent_category_data_cpp() for the alternative stacked-binary reduction of the same model. Ordinal regression for orientation.


Fast Adjacent-Category Logit with Variance (C++)

Description

Adjacent-category logit model fitting with full variance-covariance matrix.

Usage

fast_adjacent_category_logit_with_var_cpp(
  X,
  y,
  maxit = 100L,
  tol = 1e-08,
  smart_cold_start = TRUE,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL,
  warm_start_params = NULL,
  warm_start_beta = NULL
)

Arguments

X

A numeric matrix of predictors.

y

A numeric vector of responses (categorical).

maxit

Maximum number of iterations. Fast Adjacent-Category Logit Regression with Variance, Direct MLE (C++ Backend)

Fits the same adjacent-category logit model as fast_adjacent_category_logit_cpp (see that page for the model, category-coding/remapping, parameter layout, and optimizer contract, all shared unchanged here) and additionally computes the observed-information-based variance-covariance matrix of the fitted parameters.

tol

Convergence tolerance.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

fixed_idx

Optional integer indices (into the c(alpha, beta) parameter layout described in Details) of parameters to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at; must be the same length as fixed_idx.

optimization_alg

Optimization algorithm; see Details.

warm_start_fisher_info

Optional initial Fisher Information matrix (over the full c(alpha, beta) parameter vector) to warm-start curvature information.

warm_start_params

Optional starting values for the full parameter vector c(alpha, beta). If provided, smart_cold_start is ignored.

warm_start_beta

Optional starting values for just the covariate coefficients \beta (cut intercepts \alpha are still initialized separately). If provided, smart_cold_start is ignored.

Details

Variance computation. The observed Fisher information (the Hessian of the negative log-likelihood, via AdjacentCategoryLogitNegLogLik::hessian()) is evaluated at the fitted parameter vector over all n_alpha + p parameters (returned in full as fisher_information), then restricted to the free (non-fixed_idx) parameters and inverted via a rank-aware (symmetric_pseudo_inverse(), not a plain Cholesky/LDLT solve) inverse before being expanded back to full (n_alpha + p) x (n_alpha + p) size as vcov. The pseudo-inverse is used deliberately: adjacent-category fits can have an estimable treatment effect even when nuisance columns make the full information matrix rank-deficient, a case where a standard Cholesky/LDLT solve can report spurious success with an invalid (sometimes negative) variance rather than failing cleanly. vcov is only populated when converged is TRUE; otherwise it is NULL.

First-covariate variance shortcut. ssq_b_1 (aliased as ssq_b_j for interface consistency with the package's other fast_*_with_var_cpp functions) is the variance of \hat\beta_1, the coefficient on the first column of X — by the package's usual convention, the treatment-effect column — extracted directly from the free-parameter covariance block rather than requiring the caller to index into the full vcov matrix; it is NA if that coefficient was fixed (via fixed_idx) or if its estimated variance is non-finite or non-positive.

Value

A list with all the components of fast_adjacent_category_logit_cpp (b, alpha, params, neg_loglik, converged), plus ssq_b_1 (equivalently ssq_b_j, the variance of the first covariate's coefficient), vcov (the full parameter variance-covariance matrix, or NULL if not converged), and fisher_information (the full observed information matrix at the fitted parameters, over all parameters regardless of fixed_idx).

See Also

fast_adjacent_category_logit_cpp for the estimate-only variant and the full model/parameterization documentation.


Fast Beta Regression (R Wrapper)

Description

Fits the beta regression model of Ferrari and Cribari-Neto (2004) for a continuous response strictly in (0, 1), with mean linked to the covariates via the logit link and a single (constant) precision parameter \phi. See fast_beta_regression_cpp for the full model equation, parameter layout, and optimizer contract implemented by the C++ backend this function wraps; this page documents only the R-level fallback chain and response-scale conventions.

Usage

fast_beta_regression(
  X,
  y,
  start_phi = 10,
  optimization_alg = "lbfgs",
  warm_start_beta = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the response variable, with values strictly between 0 and 1. See sanitize_beta_response() (internal) for how boundary values (exact 0s/1s) are handled before fitting.

start_phi

A numeric value, the starting value for the precision parameter phi. Defaults to 10.

optimization_alg

Optimization algorithm: "lbfgs" (default) or "newton_raphson"; see .normalize_optimizer_algorithm.

warm_start_beta

Optional starting values for coefficients \beta.

warm_start_fisher_info

Optional initial Fisher Information matrix, used to warm-start curvature information for the optimizer.

Details

Fallback chain. The primary implementation uses the C++ backend (fast_beta_regression_cpp). If that fails to converge or errors, the function falls back to betareg, which is listed in Suggests and is not installed automatically with EDI. If betareg is also unavailable (or itself fails), a final fallback of OLS on logit(y) is used — this last resort is always available (no external dependency) but does not respect the beta distribution's mean-variance relationship or estimate \phi at all, so its coefficients should be treated as an approximate, non-model-based summary rather than a true beta-regression fit. Install betareg manually to enable the intermediate fallback. A warning() is issued whenever a fallback stage is used, naming which stage and the triggering error, so callers can detect when the primary fit failed even though a result was still returned.

Value

A list containing the following components:

b

A numeric vector of the estimated beta regression coefficients \hat\beta (on the logit-of-mean scale: plogis(X %*% b) gives the fitted mean \hat\mu), from whichever stage of the fallback chain (see Details) ultimately succeeded.

phi

The estimated precision parameter \hat\phi (only present when the C++ backend or the betareg fallback succeeds; absent from the final OLS-on-logit(y) fallback, which has no precision parameter).

fisher_information

The working-weights Fisher information matrix from the C++ backend (see fast_beta_regression_cpp); only present when that backend succeeds.

Examples

X = matrix(rnorm(500), 100, 5)
y = runif(100)
fast_beta_regression(X, y)

Fast Beta Regression (C++ Backend)

Description

Fits the beta regression model of Ferrari and Cribari-Neto (2004) for a continuous response strictly between 0 and 1 (proportions, rates, and similar bounded outcomes), via direct maximum likelihood on the reparameterized beta density

f(y_i; \mu_i, \phi) = \frac{\Gamma(\phi)}{\Gamma(\mu_i \phi)\Gamma((1-\mu_i)\phi)} y_i^{\mu_i \phi - 1} (1 - y_i)^{(1-\mu_i)\phi - 1}, \quad 0 < y_i < 1,

with mean E[Y_i] = \mu_i and variance \mathrm{Var}(Y_i) = \mu_i(1-\mu_i) / (1 + \phi), where \phi > 0 is a single (constant-across-observations) precision parameter and the mean is linked to the covariates via the logit link \mathrm{logit}(\mu_i) = x_i^\top \beta (fixed; no alternative link functions are supported by this backend).

Usage

fast_beta_regression_cpp(
  X,
  y,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  start_phi = 10,
  compute_std_errs = FALSE,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors, n \times p; include an explicit intercept column if desired (the model has no implicit intercept).

y

A numeric vector of responses, strictly in (0, 1) (values at or beyond the boundary are not valid beta-distributed outcomes; see fast_beta_regression for boundary-handling guidance at the R wrapper level).

warm_start_beta

Optional starting values for coefficients \beta. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

start_phi

Starting value for the precision parameter \phi (on its natural, not log, scale).

compute_std_errs

Deprecated; has no effect on this estimate-only entry point. Use fast_beta_regression_with_var_cpp for standard errors.

fixed_idx

Optional integer indices (into the c(beta, log(phi)) parameter layout) of parameters to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at; must be the same length as fixed_idx.

optimization_alg

Optimization algorithm; see Details.

warm_start_fisher_info

Optional initial Fisher Information matrix (over c(beta, log(phi))) to warm-start curvature information for the first optimizer iteration.

estimate_only

Logical; if TRUE, may skip work not needed to produce point estimates (kept in sync with the package's other fast_* estimate- vs-inference split).

Details

Parameterization and optimization. The optimizer's parameter vector is c(beta, log(phi)) — \phi is optimized on the log scale to keep it unconstrained (\phi > 0 enforced automatically by exponentiating back), initialized from start_phi (or a value derived from it via smart_cold_start). \mu_i is clipped to [10^{-8}, 1 - 10^{-8}] internally during likelihood/gradient/Hessian evaluation to avoid boundary blowup when x_i^\top \beta is extreme; this affects only numerical evaluation, not the returned \hat\beta itself. Optimized via optimization_alg ("lbfgs" default; see .normalize_optimizer_algorithm); when no warm_start_beta is supplied, smart_cold_start = TRUE (default) seeds \beta from an OLS-based initial guess. fixed_idx/fixed_values allow holding specific parameters (by index into the c(beta, log(phi)) layout) fixed rather than estimated. compute_std_errs is a legacy/deprecated argument with no effect in this estimate-only entry point; use fast_beta_regression_with_var_cpp to obtain standard errors.

Reported likelihood and information. neg_loglik is the exact beta negative log-likelihood re-evaluated at the fitted parameters (not merely the optimizer's internal objective trace); fisher_information is the X^\top W X-style working-weights curvature matrix from the fit (fit.XtWX) — the same expected-information approximation classical IRLS uses for GLM standard errors, and exactly what fast_beta_regression_with_var_cpp inverts to produce vcov — rather than a fresh evaluation of the exact observed-information Hessian from get_beta_regression_hessian_cpp. It is also suitable for warm-starting a subsequent fit via warm_start_fisher_info.

Value

A list containing the following components:

coefficients

A numeric vector of the estimated beta regression coefficients.

phi

The estimated precision parameter phi.

neg_ll

The negative log-likelihood at the final iteration.

converged

A logical value indicating whether the algorithm converged.

A list with components coefficients (\hat\beta, length p), phi (\hat\phi, on its natural scale), neg_loglik (the exact beta negative log-likelihood at the fitted parameters), converged (logical), and fisher_information (an approximate working curvature matrix; see Details).

References

Ferrari, S., and Cribari-Neto, F. (2004). "Beta regression for modelling rates and proportions." Journal of Applied Statistics, 31(7), 799-815, doi:10.1080/0266476042000214501. Analogous Python API: statsmodels GLM (via the Beta family, statsmodels.othermod.betareg).

See Also

fast_beta_regression_weighted_cpp for the row-weighted variant; fast_beta_regression_with_var_cpp for the variance-augmented variant; fast_beta_regression for the R-level wrapper with betareg fallback; get_beta_regression_score_cpp/ get_beta_regression_hessian_cpp for standalone score/Hessian evaluation at arbitrary parameter values.

Examples

X = matrix(rnorm(100), 10, 10)
y = runif(10)
fast_beta_regression_cpp(X, y)

Fast Weighted Beta Regression, Estimate Only (C++ Backend)

Description

Fits the same beta regression model as fast_beta_regression_cpp (see that page for the full model, parameterization, and optimizer contract), with each observation's contribution to the log-likelihood, score, and Hessian multiplied by a nonnegative row weight weights[i]. Setting all weights to 1 recovers fast_beta_regression_cpp exactly; this is the backend the package's Inference classes use whenever the beta regression must be fit on bootstrap-reweighted or otherwise weighted data (e.g. Bayesian bootstrap weights) without physically resampling rows.

Usage

fast_beta_regression_weighted_cpp(
  X,
  y,
  weights,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  start_phi = 10,
  compute_std_errs = FALSE,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors, n \times p.

y

A numeric vector of responses, strictly in (0, 1).

weights

A nonnegative, finite numeric vector of length nrow(X) giving each row's weight; must sum to a positive value (see Details).

warm_start_beta

Optional starting values for coefficients \beta. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when no warm start is provided.

start_phi

Starting value for the precision parameter \phi (natural scale).

compute_std_errs

Deprecated; has no effect on this estimate-only entry point.

fixed_idx

Optional integer indices (into the c(beta, log(phi)) parameter layout) of parameters to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm; see fast_beta_regression_cpp.

warm_start_fisher_info

Optional initial Fisher Information matrix to warm-start curvature information.

estimate_only

If TRUE, skip Fisher information calculation.

Details

Input validation. weights must have length nrow(X), be finite and non-negative, and sum to a strictly positive value; violating any of these raises an error immediately rather than producing a degenerate fit. A weight of 0 for a given row contributes nothing to the likelihood (effectively excludes that row) without changing n in downstream index bookkeeping.

Value

A list with the same components as fast_beta_regression_cpp: coefficients, phi, neg_loglik (the weighted negative log-likelihood), converged, and fisher_information.

See Also

fast_beta_regression_cpp for the unweighted model and full parameterization documentation; fast_beta_regression_with_var_cpp for the (unweighted) variance-augmented variant.


Fast Beta Regression with Variance Calculation (R Wrapper)

Description

Fits the same beta regression model as fast_beta_regression (see fast_beta_regression_cpp for the full model equation and parameterization) and additionally reports the estimated variance of a caller-selected coefficient, extracted from the fitted parameter variance-covariance matrix (see fast_beta_regression_with_var_cpp for how that matrix is computed and its plain-inverse numerical caveat on rank-deficient designs).

Usage

fast_beta_regression_with_var(
  X,
  y,
  start_phi = 10,
  j = 2,
  optimization_alg = "lbfgs",
  warm_start_beta = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the response variable, with values strictly between 0 and 1. See sanitize_beta_response() (internal) for boundary-value handling.

start_phi

A numeric value, the starting value for the precision parameter phi. Defaults to 10.

j

The 1-based index (into X's columns, i.e. into \beta) of the coefficient to report the variance of as ssq_b_j. Defaults to 2 (the package's usual convention for the treatment-effect column when an intercept occupies column 1).

optimization_alg

Optimization algorithm: "lbfgs" (default) or "newton_raphson"; see .normalize_optimizer_algorithm.

warm_start_beta

Optional starting values for coefficients \beta.

warm_start_fisher_info

Optional initial Fisher Information matrix, used to warm-start curvature information for the optimizer.

Details

The primary implementation uses a C++ backend. If that fails, the function falls back to betareg, which is listed in Suggests and is not installed automatically with EDI. If betareg is also unavailable, a final fallback of OLS on logit(y) is used. Install betareg manually to enable the intermediate fallback.

Value

A list containing the following components:

b

A numeric vector of the estimated beta regression coefficients \hat\beta (logit-of-mean scale).

ssq_b_j

The estimated variance (squared standard error) of the j-th coefficient, \widehat{\mathrm{Var}}(\hat\beta_j), i.e. the j-th diagonal entry of vcov. NA if the primary C++ fit failed and a fallback stage without a variance estimate was used (see Details).

ssq_b_2

The estimated variance of the second coefficient specifically (\widehat{\mathrm{Var}}(\hat\beta_2)), regardless of the j argument — provided as a convenience since column 2 is the package's usual treatment-effect position. Identical to ssq_b_j when j = 2.

Examples

X = matrix(rnorm(100), 10, 10)
y = runif(10)
fast_beta_regression_with_var(X, y)

Fast Beta Regression with Variance Calculation (C++ Backend)

Description

Fits the same beta regression model as fast_beta_regression_cpp (see that page for the full model, parameterization, and optimizer contract) and additionally computes the variance-covariance matrix and standard errors of the fitted parameters, via the same working-weights (X^\top W X) curvature matrix documented there.

Usage

fast_beta_regression_with_var_cpp(
  X,
  y,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  start_phi = 10,
  compute_std_errs = TRUE,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors, n \times p.

y

A numeric vector of responses, strictly in (0, 1).

warm_start_beta

Optional starting values for coefficients. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

start_phi

Starting value for the precision parameter \phi (natural scale).

compute_std_errs

Deprecated; standard errors are always computed by this entry point regardless of this argument's value.

fixed_idx

Optional integer indices (into the c(beta, log(phi)) parameter layout) of parameters to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm; see fast_beta_regression_cpp.

Details

Variance computation. The fit's working-information matrix (fit.XtWX, over all p + 1 parameters c(beta, log(phi))) is restricted to the free (non-fixed_idx) parameters and inverted via a plain matrix inverse (.inverse(), not a rank-aware pseudo-inverse as used by, e.g., fast_adjacent_category_logit_with_var_cpp) before being expanded back to the full (p + 1) x (p + 1) size as vcov; std_errs is sqrt(diag(vcov)). Because this uses a plain inverse, a rank-deficient or near-singular X (after restricting to free parameters) will produce numerically unstable or NaN standard errors rather than a graceful fallback — callers should ensure X is full rank on the free parameters (e.g. via the package's shared drop_linearly_dependent_cols() preprocessing) before calling this function if that is not already guaranteed.

Value

A list containing the following components:

coefficients

A numeric vector of the obtained Poisson regression coefficients.

phi

The estimated precision parameter phi.

vcov

The variance-covariance matrix.

neg_ll

The negative log-likelihood at the final iteration.

converged

A logical value indicating whether the algorithm converged.

A list with components coefficients (\hat\beta), phi (\hat\phi), neg_loglik, vcov (the full (p + 1) x (p + 1) parameter variance-covariance matrix), std_errs (sqrt(diag(vcov))), converged (logical), and fisher_information (the working-weights curvature matrix vcov was inverted from).

See Also

fast_beta_regression_cpp for the estimate-only variant and the full model/parameterization documentation; fast_beta_regression_weighted_cpp for the row-weighted estimate-only variant.

Examples

X = matrix(rnorm(100), 10, 10)
y = runif(10)
fast_beta_regression_with_var_cpp(X, y)

Fast Continuation-Ratio Regression, Direct MLE via Row Augmentation (C++ Backend)

Description

Fits the (forward) continuation-ratio logit ordinal regression model: for cut j = 1, \dots, K-1, \log \Pr(Y > j \mid Y \ge j) / \Pr(Y = j \mid Y \ge j) = \alpha_j + \beta^\top x, i.e. the log-odds of continuing past cut j (rather than stopping there), among subjects who have reached it. This orientation — numerator is the higher-category event — keeps a positive \beta meaning "pushes toward higher categories of y", consistent with every other ordinal estimator in the package (contrast the cumulative-logit -x^T \beta convention in fast_ordinal_regression.cpp and the adjacent-category model's \Pr(Y = j+1 \mid \cdot) numerator), and matches expand_continuation_ratio_data_cpp() (a separate, standalone row-expansion utility not used by this backend, but documenting the same "continue past this cut" = 1 orientation). This backend fits the model as a single unconditional logistic regression MLE on an internally-built augmented design: build_continuation_ratio_augmented_data() constructs an augmented matrix X_aug with one dummy column per cut (n_alpha = K - 1 columns) followed by the original p covariate columns, and an augmented binary response z (1 = "continued past this cut", 0 = "stopped here"). Because the cut effects \alpha_j are simply K - 1 ordinary coefficients on dummy columns (not nuisance parameters requiring conditioning), an unconditional logistic fit on the augmented data is exactly equivalent to the continuation-ratio likelihood — no stratification/conditioning machinery is needed for this standalone use case.

Usage

fast_continuation_ratio_regression_cpp(
  X,
  y,
  maxit = 100L,
  tol = 1e-08,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors, n \times p (no intercept column; threshold intercepts are estimated internally).

y

A numeric vector of length n giving each subject's ordinal category; need not be pre-coded 1:K (see Details).

maxit

Maximum number of optimizer iterations.

tol

Convergence tolerance.

warm_start_beta

Optional starting values for the full c(alpha, beta) parameter vector.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when no warm start is provided.

fixed_idx

Optional integer indices (into the c(alpha, beta) parameter layout) of parameters to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm; see Details.

warm_start_fisher_info

Optional initial Fisher Information matrix (over the full c(alpha, beta) parameter vector) to warm-start curvature information.

Details

Category coding. As in expand_continuation_ratio_data_cpp(), distinct values of y are extracted and sorted; K is the number of distinct observed values (not an externally supplied count), and each observation contributes min(observed_level + 1, K - 1) augmented rows.

Parameter vector layout. The optimizer's parameter vector (returned as params/beta_full) is c(alpha_1, ..., alpha_{K-1}, beta_1, ..., beta_p). Optimized via optimization_alg ("lbfgs" default), for at most maxit iterations at tolerance tol; when no warm start is supplied, smart_cold_start = TRUE seeds the optimizer via OLS on the augmented binary response z. fixed_idx/fixed_values hold specific parameters (by index into this layout) fixed rather than estimated, and warm_start_fisher_info warm-starts curvature information.

Degenerate case. If y has fewer than 2 distinct observed values (K < 2), no model can be fit: the function returns early with b zeroed (length p) and an empty alpha, without attempting optimization or setting converged/neg_loglik/etc.

Value

A list with components b (the shared covariate coefficients \hat\beta, length p), alpha (the K - 1 estimated cut intercepts), params/beta_full (the full c(alpha, b) parameter vector, identical to each other), neg_loglik (the augmented-data logistic negative log-likelihood, which equals the continuation-ratio model's negative log-likelihood), X_aug/ z (the augmented design matrix and binary response actually fit, exposed for reuse, e.g. by get_continuation_ratio_regression_hessian_cpp()), converged (logical), and fisher_information (the exact observed information Hessian at the fitted parameters). See Details for the degenerate fewer-than-2-categories case, which returns a reduced subset of these fields.

See Also

expand_continuation_ratio_data_cpp() for the full continuation- ratio model equation and the shared row-augmentation logic; fast_continuation_ratio_regression_with_var_cpp for the variance-augmented variant; fast_adjacent_category_logit_cpp for the analogous direct-MLE fit of the adjacent-category (rather than continuation-ratio) ordinal model. Ordinal regression for orientation.


Fast Weighted Continuation-Ratio Regression, Direct MLE (C++ Backend)

Description

Fits the same continuation-ratio likelihood as fast_continuation_ratio_regression_cpp(), weighting every augmented binary row for subject i by that subject's nonnegative weight w_i. This entry point is intended for bootstrap and other weighted refits whose estimates must retain the continuation-ratio coefficient convention.

Usage

fast_continuation_ratio_regression_weighted_cpp(
  X,
  y,
  weights,
  maxit = 100L,
  tol = 1e-08,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors, n \times p (no intercept column; threshold intercepts are estimated internally).

y

A numeric vector of length n giving each subject's ordinal category; need not be pre-coded 1:K (see Details).

weights

A finite, nonnegative subject-level weight vector of length nrow(X) containing at least one positive value.

maxit

Maximum number of optimizer iterations.

tol

Convergence tolerance.

warm_start_beta

Optional starting values for the full c(alpha, beta) parameter vector.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when no warm start is provided.

fixed_idx

Optional integer indices (into the c(alpha, beta) parameter layout) of parameters to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm; see Details.

warm_start_fisher_info

Optional initial Fisher Information matrix (over the full c(alpha, beta) parameter vector) to warm-start curvature information.

Value

The same result fields as fast_continuation_ratio_regression_cpp(), plus weights_aug, the weights copied onto the augmented binary rows.


Export of C++ function fast_continuation_ratio_regression_with_var_cpp

Description

Fits the same continuation-ratio model as fast_continuation_ratio_regression_cpp (see that page for the full model, row-augmentation mechanics, category coding, and parameter layout) and additionally computes the variance of the first covariate coefficient and (when converged) the full parameter variance-covariance matrix, from the same observed information Hessian.

Usage

fast_continuation_ratio_regression_with_var_cpp(
  X,
  y,
  maxit = 100L,
  tol = 1e-08,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors, n \times p (no intercept column; threshold intercepts are estimated internally).

y

A numeric vector of length n giving each subject's ordinal category; need not be pre-coded 1:K (see Details).

maxit

Maximum number of optimizer iterations.

tol

Convergence tolerance.

warm_start_beta

Optional starting values for the full c(alpha, beta) parameter vector.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when no warm start is provided.

fixed_idx

Optional integer indices (into the c(alpha, beta) parameter layout) of parameters to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm; see Details.

warm_start_fisher_info

Optional initial Fisher Information matrix (over the full c(alpha, beta) parameter vector) to warm-start curvature information.

Details

Variance computation. The observed information (Hessian of the augmented-data logistic negative log-likelihood, evaluated at the fitted parameters over all n_alpha + p parameters) is restricted to the free (non-fixed_idx) parameters. ssq_b_j — the variance of \hat\beta_1 (the coefficient on the first covariate column of X, the package's usual treatment-effect position) — is obtained via a single targeted diagonal-entry inversion (compute_diagonal_inverse_entry()), not a full matrix inverse, and is NA if that coefficient is fixed via fixed_idx. The full vcov (over all n_alpha + p parameters, expanded back from the free-parameter block) is computed only when converged is TRUE (via covariance_from_information()); otherwise vcov is NULL.

Degenerate case. As in fast_continuation_ratio_regression_cpp, if y has fewer than 2 distinct observed values, the function returns early with b = NA_real_, ssq_b_j = NA_real_, and converged = FALSE, without vcov/params/fisher_information.

Value

A list with components b (the shared covariate coefficients \hat\beta), ssq_b_j (the variance of the first covariate's coefficient), neg_loglik, vcov (the full parameter variance-covariance matrix, or NULL if not converged), converged (logical), params (the full c(alpha, b) parameter vector — cut intercepts are recoverable as params[1:n_alpha] but are not returned as a separate alpha field, unlike fast_continuation_ratio_regression_cpp), and fisher_information (the full observed information Hessian). See Details for the degenerate fewer-than-2-categories case, which returns a reduced subset of these fields.

See Also

fast_continuation_ratio_regression_cpp for the estimate-only variant and the full model/row-augmentation documentation.


Fast Cox Proportional Hazards Regression (R Wrapper)

Description

Fits the Cox proportional-hazards partial-likelihood model documented in full at build_cox_data_cache_cpp (model equation, Breslow tie-handling, and input conventions). This R-level wrapper dispatches to either the package's own native C++ implementation (fast_coxph_regression_cpp, the default and recommended path) or, for cross-checking or when the Rcpp path is unavailable, an elastic-net-with-zero-penalty Cox fit via glmnet (use_rcpp = FALSE) — not survival::coxph, despite that being the more commonly used reference implementation for Cox models in R.

Usage

fast_coxph_regression(
  X,
  y,
  dead,
  use_rcpp = TRUE,
  estimate_only = FALSE,
  optimization_alg = "lbfgs",
  warm_start_beta = NULL,
  warm_start_fisher_info = NULL,
  smart_cold_start = TRUE
)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept term is handled implicitly by the Cox model and should not be included in X.

y

A numeric vector representing the observed time (event time or censoring time).

dead

A numeric vector (0 or 1) indicating event status (1 for event, 0 for censored).

use_rcpp

Logical. If TRUE (default), use the optimized Rcpp implementation (fast_coxph_regression_cpp). If FALSE, use glmnet's Cox path at zero penalty (glmnet(..., family = "cox", lambda = 0)) instead.

estimate_only

Logical. If TRUE, skip variance-covariance matrix calculation for speed. Only affects the use_rcpp = TRUE path; the glmnet fallback path does not compute a variance-covariance matrix at all (vcov is never populated when use_rcpp = FALSE, regardless of estimate_only).

optimization_alg

Optimization algorithm: "newton_raphson" (default) or "lbfgs". Only affects the use_rcpp = TRUE path; unused when use_rcpp = FALSE.

warm_start_beta

Optional starting values for coefficients. If provided, smart_cold_start is ignored. Only affects the use_rcpp = TRUE path.

warm_start_fisher_info

Optional initial Fisher Information matrix. Only affects the use_rcpp = TRUE path.

smart_cold_start

Logical. If TRUE (default), use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if warm_start_beta is provided. Only affects the use_rcpp = TRUE path.

Details

Failure semantics. If the C++ fit (use_rcpp = TRUE) errors or fails to report converged, this function stops with an error rather than silently falling back to glmnet — the two code paths are alternative caller choices, not an automatic fallback chain (contrast with, e.g., fast_beta_regression's automatic betareg fallback).

glmnet dependency. When use_rcpp = FALSE, this function requires the glmnet package, which is listed in Suggests and is not installed automatically with EDI; it errors immediately if glmnet is not installed.

Value

A list. When use_rcpp = TRUE (default), a list with components

b, coefficients

A numeric vector of the estimated log-hazard-ratio coefficients \hat\beta (b and coefficients are identical; both are populated for interface consistency with the package's other fast_* wrappers).

vcov

The variance-covariance matrix of \hat\beta, or NULL when estimate_only = TRUE.

neg_log_lik

The negative Cox partial log-likelihood at the fitted coefficients.

fisher_information

The Hessian of the negative partial log-likelihood at the fitted coefficients.

When use_rcpp = FALSE, only a single component, b (the glmnet-fitted coefficient vector via coef(), in glmnet's own sparse-matrix representation rather than a plain numeric vector) — none of coefficients/vcov/neg_log_lik/ fisher_information are present on this path.

See Also

build_cox_data_cache_cpp for the full Cox partial-likelihood model, Breslow tie-handling, and input conventions; fast_coxph_regression_cpp for the native C++ backend this wrapper calls by default.

Examples

X = matrix(rnorm(500), 100, 5)
y = runif(100)
dead = rbinom(100, 1, 0.5)
fast_coxph_regression(X, y, dead)

Fast Cox Proportional Hazards Regression, One-Shot Fit (C++ Backend)

Description

Fits the unstratified Cox proportional-hazards partial-likelihood model documented in full at build_cox_data_cache_cpp — the same model, Breslow tie-handling, and input conventions — in a single call that internally builds the sorted risk-set cache, runs the optimizer, and discards the cache afterward. Use this entry point for a one-off fit; use build_cox_data_cache_cpp plus fast_coxph_regression_prebuilt_cpp instead when fitting the same (X, y, dead) repeatedly (e.g. across bootstrap/randomization replicates), to avoid rebuilding the risk-set cache on every call. fast_coxph_regression is the R-level wrapper around this backend (with an survival-free-of-Rcpp fallback path via glmnet).

Usage

fast_coxph_regression_cpp(
  X,
  y,
  dead,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  estimate_only = FALSE,
  maxit = 20L,
  tol = 1e-9,
  cluster = NULL,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "newton_raphson",
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictor variables (no intercept column; see build_cox_data_cache_cpp).

y

Numeric vector of observed (event or censoring) times.

dead

Numeric vector with values in {0, 1}: event indicator (1 = event, 0 = right-censored).

warm_start_beta

Optional starting values for the coefficients \beta.

smart_cold_start

Logical. If TRUE (default) and no warm_start_beta is supplied, use an OLS-based initial guess rather than a zero cold start.

estimate_only

Logical. If TRUE, skip variance-covariance matrix calculation for speed.

maxit

Maximum number of Newton-Raphson/L-BFGS iterations.

tol

Convergence tolerance.

cluster

Optional clustering variable; when supplied, the returned variance-covariance matrix uses a cluster-robust (grouped) sandwich correction instead of the naive model-based inverse-information variance, i.e. one that remains asymptotically valid under within-cluster correlation of the martingale residuals.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at; must be the same length as fixed_idx.

optimization_alg

Optimization algorithm: "newton_raphson" (default) or "lbfgs".

warm_start_fisher_info

Optional initial Fisher Information matrix to warm-start curvature information for the optimizer.

Value

A list containing the following components:

coefficients

A numeric vector of the estimated log-hazard-ratio coefficients \hat\beta.

vcov

The variance-covariance matrix of \hat\beta (naive inverse-information, or cluster-robust sandwich if cluster is supplied); omitted/not computed when estimate_only = TRUE.

neg_ll

The negative Cox partial log-likelihood at the final iteration.

converged

A logical value indicating whether the algorithm converged.

iterations

The number of optimizer iterations performed.

fisher_information

The Hessian of the negative partial log-likelihood at the fitted coefficients (the observed information matrix).

gradient_norm

The norm of the score (gradient) vector at convergence, a diagnostic of how tightly the convergence criterion was met.

See Also

build_cox_data_cache_cpp for the full Cox partial-likelihood model, Breslow tie-handling, and input conventions this function implements; fast_coxph_regression_prebuilt_cpp for the cache-reusing variant; fast_coxph_regression for the R-level wrapper.


Fast Cox Proportional Hazards Regression, Cache-Reusing Fit (C++ Backend)

Description

Fits the same unstratified-or-stratified Cox partial-likelihood model documented at build_cox_data_cache_cpp / build_stratified_cox_data_cache_cpp, but takes a pre-built risk-set cache (cox_data_xptr, an externalptr produced by one of those two functions) instead of raw (X, y, dead) data, skipping the sort/tabulation step on every call. This is the entry point the package's Cox inference classes (e.g. InferenceCoxPH, InferenceStratifiedCoxPH) use for repeated fits on the same data (successive estimate_only vs. full-variance calls, or bootstrap/ randomization replicates that only change the treatment column of X, rebuilding the cache only when the covariates or assignment actually change). fast_coxph_regression_cpp is the equivalent one-shot entry point that builds and discards the cache internally, for callers that only need a single fit.

Usage

fast_coxph_regression_prebuilt_cpp(
  cox_data_xptr,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  estimate_only = FALSE,
  maxit = 20L,
  tol = 1e-9,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "newton_raphson",
  warm_start_fisher_info = NULL
)

Arguments

cox_data_xptr

An externalptr to a cached Cox risk-set representation, as returned by build_cox_data_cache_cpp (unstratified) or build_stratified_cox_data_cache_cpp (stratified).

warm_start_beta

Optional starting values for the coefficients \beta.

smart_cold_start

Logical. If TRUE (default) and no warm_start_beta is supplied, use an OLS-based initial guess rather than a zero cold start.

estimate_only

Logical. If TRUE, skip variance-covariance matrix calculation for speed.

maxit

Maximum number of Newton-Raphson/L-BFGS iterations.

tol

Convergence tolerance.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at; must be the same length as fixed_idx.

optimization_alg

Optimization algorithm: "newton_raphson" (default) or "lbfgs".

warm_start_fisher_info

Optional initial Fisher Information matrix to warm-start curvature information for the optimizer.

Details

Because cox_data_xptr carries a fixed, already-sorted risk-set structure (whether unstratified — one risk set — or stratified — one risk set per stratum, depending on which cache-building function produced it), the number and identity of subjects/strata are entirely determined by the cache; only the optimization behavior (warm starts, convergence, algorithm) is configurable through this function's own arguments. Passing a stale cache (built from data that has since changed) silently fits the model to the cached data, not the caller's current X/y/dead — callers are responsible for invalidating and rebuilding the cache when the underlying data changes; see build_cox_data_cache_cpp for the exact caching/mutation contract.

Value

A list containing the following components:

coefficients

A numeric vector of the estimated log-hazard-ratio coefficients \hat\beta.

vcov

The variance-covariance matrix of \hat\beta; omitted when estimate_only = TRUE.

neg_ll

The negative Cox partial log-likelihood at the final iteration.

converged

A logical value indicating whether the algorithm converged.

iterations

The number of optimizer iterations performed.

fisher_information

The Hessian of the negative partial log-likelihood at the fitted coefficients.

gradient_norm

The norm of the score (gradient) vector at convergence.

See Also

build_cox_data_cache_cpp/ build_stratified_cox_data_cache_cpp for building the required cache and the full Cox partial-likelihood model documentation; fast_coxph_regression_cpp for the one-shot (build-and-discard) variant.


Fast Combined Conditional-Poisson + Poisson Regression for KK Matched-Pair/ Reservoir Designs, with Variance (C++ Backend)

Description

Jointly fits a single treatment-effect coefficient \beta_T (and shared covariate effects \beta_{xs}) across two structurally different count likelihoods at once — the matched-pair (conditional Poisson) component from subjects paired on-the-fly by a KK matching design (e.g. DesignSeqOneByOneKK14) and the marginal Poisson component from unmatched "reservoir" subjects — rather than fitting the two subsets separately and combining estimates afterward (as an inverse-variance-weighted combination does elsewhere in the package). This one-likelihood joint fit is what backs InferenceCountKKCondPoissonOneLik-style estimators.

Usage

fast_cpoisson_combined_with_var_cpp(
  yT_v_r,
  n_k_v_r,
  X_diff_v_r,
  y_r_r,
  w_r_r,
  X_r_r,
  maxit = 100L,
  tol = 1e-08,
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_fisher_info = NULL,
  warm_start_params = NULL,
  warm_start_beta = NULL,
  estimate_only = FALSE
)

Arguments

yT_v_r

Numeric vector of length n_{\mathrm{pairs}}: the treated member's count for each matched pair.

n_k_v_r

Numeric vector of length n_{\mathrm{pairs}}: the total (treated + control) count for each matched pair.

X_diff_v_r

Numeric matrix, n_{\mathrm{pairs}} \times p: each pair's covariate difference (treated minus control); p = 0 (zero columns) is valid (no covariate adjustment).

y_r_r

Numeric vector of length n_R: reservoir subjects' counts.

w_r_r

Numeric vector of length n_R with values in {0, 1}: reservoir subjects' treatment indicators.

X_r_r

Numeric matrix, n_R \times p: reservoir subjects' covariates (same p as X_diff_v_r).

maxit

Maximum number of Newton iterations.

tol

Convergence tolerance (on the norm of the parameter update step).

fixed_idx

Optional integer indices (into the c(beta_0, beta_T, beta_xs) parameter layout) of parameters to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

warm_start_fisher_info

Optional initial Fisher Information matrix (over the full p + 2 parameters) to warm-start the first Newton iteration.

warm_start_params

Optional starting values for the full parameter vector c(beta_0, beta_T, beta_xs).

warm_start_beta

Optional starting values for just c(beta_T, beta_xs) (length p + 1); beta_0 is still initialized separately. Ignored if warm_start_params is supplied.

estimate_only

Logical; if TRUE, skip score/information/variance computation after optimization (see Details).

Details

Matched-pair component (conditional Poisson). For pair k with total count n_k (sum of both members' counts) and treated-member count y_{T,k}, conditioning on n_k (the sufficient statistic that eliminates the pair's nuisance baseline rate) reduces the joint Poisson likelihood of the pair to a Binomial: y_{T,k} \mid n_k \sim \mathrm{Binomial}(n_k, p_k), p_k = \mathrm{logit}^{-1}(\beta_T + x_{\Delta,k}^\top \beta_{xs}), where x_{\Delta,k} is the pair's covariate difference (treated minus control). This is exactly the count-response analog of conditional logistic regression for matched pairs — no per-pair intercept is estimated (it is conditioned out entirely), so only \beta_T and \beta_{xs} appear in this component.

Reservoir component (marginal Poisson). Unmatched reservoir subjects contribute an ordinary Poisson log-linear likelihood, y_i \sim \mathrm{Poisson}(\mu_i), \log \mu_i = \beta_0 + w_i \beta_T + x_i^\top \beta_{xs}, sharing the same \beta_T and \beta_{xs} as the pair component but additionally estimating an intercept \beta_0 (which the conditional pair likelihood has no use for).

Combined likelihood and optimization. The total log-likelihood is the simple sum of the pair (conditional Poisson/Binomial) and reservoir (Poisson) log-likelihoods, jointly maximized over c(beta_0, beta_T, beta_xs) (length p + 2) via Newton's method using the analytic Fisher information as the Hessian (quadratic convergence near the optimum, typically very few iterations). fixed_idx/fixed_values hold specific parameters fixed rather than estimated; warm_start_params (full vector) or warm_start_beta (either the full vector, or just c(beta_T, beta_xs) when of length p + 1, in which case beta_0 is initialized separately) seed the optimizer, with a log-mean-based default cold start for beta_0 when neither is supplied.

Variance. ssq_b_j is the variance of \hat\beta_T specifically (index 1, 0-based, in the parameter layout — the package's usual single-treatment-coefficient convention), obtained via a targeted diagonal inverse of the observed/Fisher information restricted to free parameters; NA if \beta_T was itself fixed via fixed_idx.

Estimate-only mode. If estimate_only = TRUE, optimization still runs to convergence but the score/information/variance computation is skipped entirely, returning only b, params, and converged.

Value

A list with components b/params (the fitted c(beta_0, beta_T, beta_xs) vector), converged (logical), and, unless estimate_only = TRUE: ssq_b_j (the variance of \hat\beta_T), score (the score vector at the fitted parameters), observed_information/fisher_information/ information (three aliases for the same Fisher information matrix, also tagged by information_type = "fisher"), hessian (the Hessian of the negative log-likelihood, i.e. -information), neg_loglik/neg_ll (aliases for the combined negative log-likelihood at the fitted parameters), and loglik (its negation).

See Also

Conditional logistic regression for the matched-pair likelihood's structural analog; Poisson regression for the reservoir component; analogous Python API: statsmodels ConditionalPoisson for the conditional-Poisson matched-set likelihood alone (not the combined pair+reservoir model implemented here).


Fast Digamma Function, Vectorized (C++ Backend)

Description

Computes the digamma function \psi(x) = d/dx \log \Gamma(x) elementwise over x, via an asymptotic expansion with a recurrence (reflection) shift for small arguments to keep the expansion accurate — the standard technique for evaluating digamma/trigamma to double precision without a lookup table. Used internally inside the package's negative-binomial, beta, zero-inflated/hurdle, and KK21 count-response likelihood, score, and Hessian kernels (wherever a Poisson/NegBin/Beta log-likelihood derivative requires \psi), and exported standalone because it is consistently faster than base R's digamma — measured at 6.78x on a length-5000 vector (see the "Utility / Math Kernel Performance" benchmark report for the full methodology and per-kernel results).

Usage

fast_digamma_vec_cpp(x)

Arguments

x

Numeric vector of arguments (should be finite and, per the digamma function's domain, not a non-positive integer, where \psi has poles; no domain validation is performed by this function).

Value

A numeric vector of \psi(x) values, the same length as x.

References

Abramowitz, M., and Stegun, I. A. (1972). Handbook of Mathematical Functions, Section 6.3, for the asymptotic expansion and recurrence relation used. See also digamma function for orientation. Analogous Python API: SciPy digamma.


Fast Mean-Parameterized Negative-Binomial Density, Vectorized (C++ Backend)

Description

Computes the negative-binomial probability mass function, in its mean/dispersion parameterization,

f(x; \mathrm{size}, \mu) = \binom{x + \mathrm{size} - 1}{x} \left(\frac{\mathrm{size}}{\mathrm{size} + \mu}\right)^{\mathrm{size}} \left(\frac{\mu}{\mathrm{size} + \mu}\right)^{x},

elementwise over x, with E[X] = \mu and \mathrm{Var}(X) = \mu + \mu^2/\mathrm{size} (size is the dispersion/shape parameter; smaller size means more overdispersion relative to Poisson). This matches R::dnbinom_mu(x, size, mu, give_log) semantics exactly, but evaluates the three required lgamma calls per observation via fast_lgamma_vec_cpp's kernel instead of R's own lgamma dispatch, making it faster than base R's stats::dnbinom(x, size, mu = mu, log = ...) — measured at 1.35x on a length-5000 vector (see the "Utility / Math Kernel Performance" benchmark report) — while returning numerically identical values. Used internally inside the package's negative-binomial regression likelihood, score, and Hessian kernels.

Usage

fast_dnbinom_mu_vec_cpp(x, size, mu, return_log)

Arguments

x

Numeric vector of non-negative integer counts (non-integer or negative values are not validated by this function and will produce incorrect or non-finite results, matching R::dnbinom_mu's own lack of input validation at the C level).

size

Dispersion (shape) parameter > 0 (single value, recycled against every element of x).

mu

Mean parameter > 0 (single value, recycled against every element of x).

return_log

Logical. If TRUE, return the log-density instead of the density.

Value

A numeric vector of (log-)density values, the same length as x.

References

Negative binomial distribution for the mean/dispersion parameterization used here. Analogous Python API: SciPy stats distributions index (scipy.stats.nbinom, in its number-of-successes/probability parameterization — convert via p = \mathrm{size}/(\mathrm{size}+\mu)).

See Also

fast_lgamma_vec_cpp, whose kernel this function calls three times per observation.


Fast Hurdle Negative-Binomial Regression, with Variance (C++ Backend)

Description

Fits a two-part hurdle negative-binomial model for count data with excess zeros: (1) a hurdle part — logistic regression of the binary indicator I(Y_i > 0) on X_hurdle — models whether the hurdle is crossed at all, and (2) a count part — a zero-truncated negative-binomial regression fit only on the subset of subjects with Y_i > 0, using X — models the count given the hurdle is crossed. Unlike a zero-inflated model (which mixes a point mass at zero with an untruncated count distribution that can itself also produce zeros), the hurdle model's two parts are a clean partition: every zero comes from the hurdle part, and every positive count's distribution is exactly the negative-binomial conditional on being positive (left-truncated at 1).

Usage

fast_hurdle_negbin_with_var_cpp(
  X_r,
  y_r,
  X_hurdle_r,
  j = 2L,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  maxit = 1000L,
  tol = 1e-08,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL,
  warm_start_hurdle_fisher_info = NULL
)

Arguments

X_r

Numeric matrix of predictors for the count component (the zero-truncated negative-binomial part), n \times p.

y_r

Numeric vector of length n: observed non-negative integer counts (zeros are handled by the hurdle part; only the positive subset is passed to the truncated count part).

X_hurdle_r

Numeric matrix of predictors for the hurdle (zero-vs- positive) logistic component, n \times p_{\mathrm{hurdle}}; may differ from X_r (a different covariate set for "does an event occur at all" vs. "how many, given at least one").

j

1-based index (into the count model's p coefficients) of the coefficient to report ssq_b_j for.

warm_start_params

Optional starting values for the count model's full c(beta, log(theta)) parameter vector. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

maxit

Maximum number of count-model optimizer iterations.

tol

Convergence tolerance (count model).

fixed_idx

Optional integer indices (into the count model's c(beta, log(theta)) layout) of parameters to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm for the count model; the hurdle logistic part always uses the same algorithm internally.

warm_start_fisher_info

Optional initial Fisher Information matrix for the count model's first optimizer iteration.

warm_start_hurdle_fisher_info

Optional initial Fisher Information matrix for the hurdle logistic model's first optimizer iteration.

Details

Hurdle part. Fit via fast_logistic_regression_cpp's internal engine on y_pos_ind = as.numeric(y > 0) regressed on X_hurdle; if y_pos_ind has no variation (all-zero or all-positive y), the hurdle part is skipped (hurdle_b is all NA, hurdle_converged = FALSE) rather than erroring.

Count part. Fit on the positive-count subset (n_+ = \sum_i I(y_i > 0) rows) via maximum likelihood on the zero-truncated negative-binomial density with mean-parameterized dispersion \theta: parameter vector c(beta, log(theta)) (length p + 1), with \theta optimized on the log scale for positivity and reported back as theta_hat = exp(params[p+1]). If n_+ \le p (too few positive observations to identify the count-model coefficients), the count part returns b as all NA and converged = FALSE with an explanatory failure_message, while the hurdle part (which does not depend on n_+) is still fit and returned normally.

Variance. ssq_b_j/ssq_b_2 are the variances of the j-th and 2nd count-model coefficients (from the count part's observed information, restricted to free/non-fixed_idx parameters); hurdle_ssq_b_j/hurdle_ssq_b_2 are the analogous variances for the hurdle-model coefficients (index j into that model's own coefficient vector, from the hurdle logistic regression's own information matrix — no fixed_idx applies to the hurdle part). Both use a targeted diagonal-entry inversion rather than a full matrix inverse, and are NA if the relevant coefficient was fixed, out of range, or its model failed to converge/produce a finite information matrix.

Value

A list with components b (count-model coefficients \hat\beta), theta_hat (the zero-truncated NB dispersion), converged (count model), hurdle_b (hurdle logistic coefficients), hurdle_converged, ssq_b_j/ssq_b_2 (count-model coefficient variances), hurdle_ssq_b_j/ hurdle_ssq_b_2 (hurdle-model coefficient variances), observed_information/fisher_information/information (three aliases for the count model's observed information, over c(beta, log(theta))), information_type = "observed", hessian (the negative of that information), hurdle_fisher_information (the hurdle model's own information matrix), and failure_message (empty on success, otherwise an explanatory string for a degenerate count-part fit).

See Also

fast_logistic_regression_cpp for the hurdle component's fitting engine. Negative binomial distribution for orientation. Analogous Python API: statsmodels discrete models (HurdleCountModel with a negative-binomial count distribution).


Fast Identity-Link Binomial Regression, Estimate Only (C++ Backend)

Description

Fits a binary-response GLM with the identity link (a linear probability / risk-difference model), \mu_i = \Pr(Y_i = 1) = x_i^\top \beta (constrained to (10^{-8}, 1 - 10^{-8}); no other link transformation is applied), via Fisher scoring (IRLS) with a step-halving line search that rejects any Newton step whose resulting \eta_i = x_i^\top \beta would leave the valid probability range or decrease the log-likelihood — this boundary-constrained line search, not a link-function transform, is what keeps fitted probabilities in (0, 1) for this otherwise-unconstrained linear-in-\beta model. Regression coefficients on the identity-link scale are directly interpretable as risk differences: \beta_j is the change in \Pr(Y = 1) per unit change in covariate j, in contrast to fast_log_binomial_regression_cpp's log-link coefficients (interpretable as log relative risks) or a standard logit-link model's log-odds-ratio coefficients.

Usage

fast_identity_binomial_regression_cpp(
  X,
  y_r,
  maxit = 100L,
  tol = 1e-06,
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors, n \times p; include an explicit intercept column if desired (no implicit intercept).

y_r

A binary (0/1) numeric vector of responses, length n.

maxit

Maximum number of Fisher-scoring iterations.

tol

Convergence tolerance, on the relative norm of the coefficient update step.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

warm_start_beta

Optional starting values for coefficients. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE (default) and no warm_start_beta is supplied, use an OLS-based initial guess.

warm_start_weights

Optional initial working weights for the first IRLS iteration.

warm_start_fisher_info

Optional initial Fisher Information matrix for the first IRLS iteration.

Details

Optimization. Each Fisher-scoring iteration solves a weighted least-squares step using working weights w_i = 1 / \max(\mu_i(1-\mu_i), 10^{-8}) (the inverse Bernoulli variance, clamped away from 0 for stability near the boundary), then backtracks (halving the step size, down to a minimum step of 10^{-8}) until the resulting \eta stays within (10^{-8}, 1-10^{-8}) for every observation and the log-likelihood does not decrease; a step that cannot be accepted at any halving depth terminates iteration without converged = TRUE. fixed_idx/fixed_values hold specific coefficients fixed rather than estimated; warm_start_beta (or, when absent, an OLS-based guess if smart_cold_start = TRUE) seeds the first iteration, and warm_start_weights/warm_start_fisher_info warm-start the first IRLS working-weights/curvature computation.

No guarantee of a feasible solution. Because the identity link has no inherent boundary protection, some (X, y) configurations (e.g. extreme covariate values, near-perfect separation, or an ill-conditioned X) may have no interior maximum-likelihood solution reachable by this constrained line search; such cases surface as converged = FALSE rather than a silently invalid (out-of-range) fitted probability.

Value

A list with components b (estimated coefficients \hat\beta, on the risk-difference/identity scale), mu_hat (fitted probabilities \hat\mu_i, length n), working_weights (the final IRLS weights w_i), iterations (number of Fisher- scoring iterations performed), converged (logical; also requires all of b, mu_hat, working_weights to be finite), and fisher_information (the working-weights curvature matrix X^\top W X).

See Also

fast_identity_binomial_regression_with_var_cpp for the variance-augmented variant; fast_identity_binomial_regression_weighted_cpp for the row-weighted variant; fast_log_binomial_regression_cpp for the log-link (relative-risk) analog of this model. Generalized linear model for orientation. Analogous Python API: statsmodels GLM (families.Binomial(link=identity())).


Fast Weighted Identity-Link Binomial Regression, Estimate Only (C++ Backend)

Description

Fits the same identity-link (risk-difference) binomial regression as fast_identity_binomial_regression_cpp (see that page for the full model, boundary-constrained IRLS line search, and interpretation), with each observation's contribution to the log-likelihood and IRLS working weights multiplied by a nonnegative row weight weights_r[i]. Setting all weights to 1 recovers fast_identity_binomial_regression_cpp exactly; this is the backend used when the identity-link model must be fit on bootstrap-reweighted or otherwise weighted data.

Usage

fast_identity_binomial_regression_weighted_cpp(
  X,
  y_r,
  weights_r,
  maxit = 100L,
  tol = 1e-06,
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors, n \times p.

y_r

A binary (0/1) numeric vector of responses, length n.

weights_r

A nonnegative numeric vector of length n giving each row's weight.

maxit

Maximum number of Fisher-scoring iterations.

tol

Convergence tolerance.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

warm_start_beta

Optional starting values for coefficients. If provided, smart_cold_start is ignored.

warm_start_weights

Optional initial working weights for the first IRLS iteration.

warm_start_fisher_info

Optional initial Fisher Information matrix for the first IRLS iteration.

Value

A list with the same components as fast_identity_binomial_regression_cpp: b, mu_hat, working_weights, iterations, converged, and fisher_information (all reflecting the weighted log-likelihood).

See Also

fast_identity_binomial_regression_cpp for the unweighted model and full documentation; fast_identity_binomial_regression_with_var_cpp for the (unweighted) variance-augmented variant.


Fast Identity-Link Binomial Regression with Targeted Variance (C++ Backend)

Description

Fits the same identity-link (risk-difference) binomial regression as fast_identity_binomial_regression_cpp (see that page for the full model and boundary-constrained IRLS line search) and additionally computes the variance of a single caller-selected coefficient, via a targeted diagonal-entry inversion of the working-weights Fisher information — this entry point does not compute or return a full variance-covariance matrix or a vector of standard errors for every coefficient, despite its name; only the one coefficient named by j gets a variance (ssq_b_j).

Usage

fast_identity_binomial_regression_with_var_cpp(
  X,
  y_r,
  j = 2L,
  maxit = 100L,
  tol = 1e-06,
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors, n \times p.

y_r

A binary (0/1) numeric vector of responses, length n.

j

1-based index (into X's columns) of the coefficient to compute ssq_b_j for.

maxit

Maximum number of Fisher-scoring iterations.

tol

Convergence tolerance.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

warm_start_beta

Optional starting values for coefficients. If provided, smart_cold_start is ignored.

warm_start_weights

Optional initial working weights for the first IRLS iteration.

warm_start_fisher_info

Optional initial Fisher Information matrix for the first IRLS iteration.

Details

Variance computation. The IRLS working-weights Fisher information X^\top W X (reused from the underlying fit if finite and correctly sized, else recomputed from the final working weights) is restricted to the free (non-fixed_idx) parameters and factorized via LDLT; ssq_b_j is then obtained from a single targeted diagonal-entry inversion (compute_diagonal_inverse_entry()) at the free-parameter position corresponding to j, not a full matrix inverse. If the underlying fit did not converge, or the LDLT factorization fails (e.g. a rank-deficient free-parameter information matrix), the function returns early with converged = FALSE, ssq_b_j = NA, and empty (zero-length/zero-dimension) vcov/std_err/z_vals placeholders — these three fields are only ever populated as empty placeholders, on both the success and failure paths; no caller should rely on them containing actual values.

Value

A list with components b (estimated coefficients \hat\beta), ssq_b_j (the variance of \hat\beta_j, or NA on failure), converged (logical), fisher_information (the working-weights curvature matrix used for ssq_b_j, present only on the success path), neg_ll/logLik (the negative/positive log-likelihood at \hat\beta, present only on the success path), and the always-empty vcov/std_err/z_vals placeholders described in Details.

See Also

fast_identity_binomial_regression_cpp for the estimate-only variant and the full model documentation.


Fast Log-Beta Function, Vectorized (C++ Backend)

Description

Computes the log of the Beta function, \log B(a, b) = \log\Gamma(a) + \log\Gamma(b) - \log\Gamma(a+b), elementwise, via three calls into fast_lgamma_vec_cpp's kernel rather than R's own lgamma dispatch — faster than base R's lbeta — measured at 2.43x on a length-5000 vector (see the "Utility / Math Kernel Performance" benchmark report) — while returning numerically identical values (up to the Lanczos/Stirling approximation's own precision). Used internally inside the package's beta-regression and beta-distribution-based (zero-one-inflated beta) likelihood, score, and Hessian kernels, wherever a Beta-density normalizing constant is required.

Usage

fast_lbeta_vec_cpp(a, b)

Arguments

a

Numeric vector of first shape arguments (should be positive; not validated by this function).

b

Numeric vector of second shape arguments (should be positive; not validated), recycled against a elementwise — must be the same length as a; unlike R's own vectorized arithmetic, this function does not perform R-style shorter-vector recycling.

Value

A numeric vector of \log B(a, b) values, the same length as a/b.

References

Beta function for orientation. Analogous Python API: SciPy betaln.

See Also

fast_lgamma_vec_cpp, whose kernel this function calls.


Fast Log-Gamma Function, Vectorized (C++ Backend)

Description

Computes \log \Gamma(x) elementwise over x, via a Lanczos approximation (with a Stirling-series tail for large arguments) — faster than base R's lgamma while matching it to within the approximation's own precision. Used pervasively throughout the package's likelihood kernels (beta, negative-binomial, Poisson/count, and other Gamma-function-based densities) wherever a log-factorial-like normalizing term is required, and exported standalone for the same reason as fast_digamma_vec_cpp — measured at 2.18x over lgamma on a length-5000 vector (see the "Utility / Math Kernel Performance" benchmark report for the full methodology and per-kernel results).

Usage

fast_lgamma_vec_cpp(x)

Arguments

x

Numeric vector of arguments (should be positive, or a non-positive non-integer if the reflection formula is supported by the underlying kernel; not validated by this function — see the package's C++ source for the exact domain the Lanczos kernel handles).

Value

A numeric vector of \log \Gamma(x) values, the same length as x.

References

Lanczos approximation and Stirling's approximation for the numerical techniques used; see also Gamma function for orientation. Analogous Python API: SciPy gammaln.

See Also

fast_digamma_vec_cpp, fast_trigamma_vec_cpp, fast_lbeta_vec_cpp (built on this function's kernel).


Fast Log-Link Binomial Regression, Estimate Only (C++ Backend)

Description

Fits a binary-response GLM with the log link (a relative-risk model), \log \mu_i = \log \Pr(Y_i = 1) = x_i^\top \beta (equivalently \mu_i = e^{x_i^\top \beta}, constrained to stay below 1 - 10^{-8} so it remains a valid probability), via Fisher scoring (IRLS) with a step-halving line search that rejects any Newton step whose resulting \eta_i = x_i^\top \beta would push \mu_i out of range or decrease the log-likelihood — the same boundary-constrained-line-search mechanism documented in full at fast_identity_binomial_regression_cpp (see that page for the IRLS/line-search mechanics, which are shared verbatim between the log and identity links here; only the link function itself, and hence the coefficient scale, differs). Regression coefficients are directly interpretable as log relative risks: e^{\beta_j} is the multiplicative change in \Pr(Y = 1) per unit change in covariate j — in contrast to fast_identity_binomial_regression_cpp's risk-difference scale, or a logit-link model's odds-ratio scale.

Usage

fast_log_binomial_regression_cpp(
  X,
  y_r,
  maxit = 100L,
  tol = 1e-06,
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors, n \times p; include an explicit intercept column if desired (no implicit intercept).

y_r

A binary (0/1) numeric vector of responses, length n.

maxit

Maximum number of Fisher-scoring iterations.

tol

Convergence tolerance, on the relative norm of the coefficient update step.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

warm_start_beta

Optional starting values for coefficients. If provided, smart_cold_start is ignored.

warm_start_weights

Optional initial working weights for the first IRLS iteration.

warm_start_fisher_info

Optional initial Fisher Information matrix for the first IRLS iteration.

Value

A list with components b (estimated coefficients \hat\beta, on the log-relative-risk scale), mu_hat (fitted probabilities, length n), working_weights (final IRLS weights), iterations, converged (logical), and fisher_information (the working-weights curvature matrix X^\top W X).

See Also

fast_identity_binomial_regression_cpp for the identity-link (risk-difference) analog and the full IRLS/line-search mechanics; fast_log_binomial_regression_with_var_cpp for the variance-augmented variant; fast_log_binomial_regression_weighted_cpp for the row-weighted variant. Poisson regression's log link is the closest common orientation point for a log-link GLM. Analogous Python API: statsmodels GLM (families.Binomial(link=log())).


Fast Weighted Log-Link Binomial Regression, Estimate Only (C++ Backend)

Description

Fits the same log-link (relative-risk) binomial regression as fast_log_binomial_regression_cpp (see that page for the full model and boundary-constrained IRLS line search), with each observation's contribution to the log-likelihood and IRLS working weights multiplied by a nonnegative row weight weights_r[i]. Setting all weights to 1 recovers fast_log_binomial_regression_cpp exactly.

Usage

fast_log_binomial_regression_weighted_cpp(
  X,
  y_r,
  weights_r,
  maxit = 100L,
  tol = 1e-06,
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors, n \times p.

y_r

A binary (0/1) numeric vector of responses, length n.

weights_r

A nonnegative numeric vector of length n giving each row's weight.

maxit

Maximum number of Fisher-scoring iterations.

tol

Convergence tolerance.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

warm_start_beta

Optional starting values for coefficients. If provided, smart_cold_start is ignored.

warm_start_weights

Optional initial working weights for the first IRLS iteration.

warm_start_fisher_info

Optional initial Fisher Information matrix for the first IRLS iteration.

Value

A list with the same components as fast_log_binomial_regression_cpp: b, mu_hat, working_weights, iterations, converged, and fisher_information (all reflecting the weighted log-likelihood).

See Also

fast_log_binomial_regression_cpp for the unweighted model and full documentation.


Fast Log-Link Binomial Regression with Targeted Variance (C++ Backend)

Description

Fits the same log-link (relative-risk) binomial regression as fast_log_binomial_regression_cpp (see that page for the full model) and additionally computes the variance of a single caller-selected coefficient — the log-link analog of fast_identity_binomial_regression_with_var_cpp, sharing exactly the same targeted-diagonal-entry variance mechanism and the same caveat: this entry point does not compute or return a full variance-covariance matrix or per-coefficient standard errors, despite its name; only the coefficient named by j gets a variance (ssq_b_j), and the returned vcov/std_err/z_vals fields are always empty placeholders (see fast_identity_binomial_regression_with_var_cpp's Details for the exact mechanics, identical here up to the link function).

Usage

fast_log_binomial_regression_with_var_cpp(
  X,
  y_r,
  j = 2L,
  maxit = 100L,
  tol = 1e-06,
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors, n \times p.

y_r

A binary (0/1) numeric vector of responses, length n.

j

1-based index (into X's columns) of the coefficient to compute ssq_b_j for.

maxit

Maximum number of Fisher-scoring iterations.

tol

Convergence tolerance.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

warm_start_beta

Optional starting values for coefficients. If provided, smart_cold_start is ignored.

warm_start_weights

Optional initial working weights for the first IRLS iteration.

warm_start_fisher_info

Optional initial Fisher Information matrix for the first IRLS iteration.

Value

A list with components b, ssq_b_j, converged, fisher_information, neg_ll/logLik (present only on the success path), and the always-empty vcov/std_err/ z_vals placeholders; see fast_identity_binomial_regression_with_var_cpp for the exact field semantics (shared verbatim here).

See Also

fast_log_binomial_regression_cpp for the estimate-only variant; fast_identity_binomial_regression_with_var_cpp for the identity-link analog with the same targeted-variance mechanism.


Fast Log Standard Normal Density, Vectorized (C++ Backend)

Description

Computes \log \phi(x) = -\tfrac{1}{2}\log(2\pi) - x^2/2, the log-density of the standard normal distribution, elementwise over x, via a direct closed-form evaluation — no series expansion or special-function dispatch is needed since the standard normal log-density has an exact elementary closed form. Faster than base R's dnorm(x, log = TRUE) — measured at 5x on a length-5000 vector (see the "Utility / Math Kernel Performance" benchmark report) — while returning numerically identical values. Used internally inside the package's probit regression and other Gaussian-likelihood kernels wherever a standard normal log-density is required.

Usage

fast_log_dnorm_vec_cpp(x)

Arguments

x

Numeric vector of arguments.

Value

A numeric vector of \log \phi(x) values, the same length as x.

References

Normal distribution for orientation. Analogous Python API: SciPy stats distributions index (scipy.stats.norm.logpdf).

See Also

fast_log_pnorm_vec_cpp for the corresponding log-CDF kernel; fast_qnorm_vec_cpp for the standard normal quantile function.


Fast Log Standard Normal CDF, Vectorized (C++ Backend)

Description

Computes \log \Phi(x), the log of the standard normal cumulative distribution function, elementwise over x, via the complementary error function kernel fast_erfc (\Phi(x) = \tfrac{1}{2} \mathrm{erfc}(-x/\sqrt{2}), evaluated in a form stable for large negative x, where \Phi(x) underflows in ordinary (non-log) arithmetic long before the true log-probability does), avoiding R's own pnorm dispatch overhead. Faster than base R's pnorm(x, log.p = TRUE) — measured at 2.49x on a length-5000 vector (see the "Utility / Math Kernel Performance" benchmark report). Used internally inside the package's probit regression and other likelihood kernels that need a numerically stable normal log-CDF, e.g. for censored/truncated Gaussian contributions.

Usage

fast_log_pnorm_vec_cpp(x)

Arguments

x

Numeric vector of arguments.

Value

A numeric vector of \log \Phi(x) values, the same length as x.

References

Normal distribution for orientation. Analogous Python API: SciPy stats distributions index (scipy.stats.norm.logcdf).

See Also

fast_log_dnorm_vec_cpp for the corresponding log-density kernel; fast_qnorm_vec_cpp for the standard normal quantile function.


Fast Logistic Regression, Estimate Only (R Wrapper)

Description

Fits the logistic regression model documented in full at fast_logistic_regression_cpp (log-odds-ratio interpretation, IRLS/L-BFGS/Newton-Raphson optimization) via that C++ backend, returning only the point estimate \hat\beta — no variance-covariance matrix or per-coefficient standard errors are computed. Unlike fast_logistic_regression_with_var, this function does not attempt to detect or retry on (quasi-)complete separation; if the underlying C++ fit errors for any reason, this function silently returns b as a vector of NAs (of length ncol(X)) rather than raising an error or retrying with fewer covariates.

Usage

fast_logistic_regression(
  X,
  y,
  optimization_alg = "lbfgs",
  warm_start_beta = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the response variable, expected to be binary (0 or 1).

optimization_alg

Optimization algorithm: "lbfgs" (default), "newton_raphson", or "irls".

warm_start_beta

Optional starting values for the coefficients.

warm_start_fisher_info

Optional initial Fisher Information matrix.

Value

A list containing the following component:

b

A numeric vector of the estimated logistic regression coefficients \hat\beta, or a vector of NA_real_ (length ncol(X)) if the underlying fit errored.

See Also

fast_logistic_regression_cpp for the underlying backend and full model documentation; fast_logistic_regression_with_var for the variance- augmented, separation-retrying variant.

Examples

X = matrix(rnorm(500), 100, 5)
y = rbinom(100, 1, 0.5)
fast_logistic_regression(X, y)

Fast Logistic Regression, Estimate Only (C++ Backend)

Description

Fits the standard binary logistic regression model, \mathrm{logit}(\mu_i) = \Pr(Y_i = 1) \text{'s log-odds} = x_i^\top \beta, \mu_i = \mathrm{logit}^{-1}(x_i^\top \beta), via maximum likelihood. Coefficients are directly interpretable as log odds ratios: e^{\beta_j} is the multiplicative change in the odds \mu_i / (1 - \mu_i) per unit change in covariate j. This is the package's baseline binary-response fitting backend, used wherever an incidence/binary outcome needs a logit-link fit (as opposed to the log-link or identity-link constrained binomial models in fast_log_binomial_regression_cpp/ fast_identity_binomial_regression_cpp, which target relative risk / risk difference scales instead of odds ratios).

Usage

fast_logistic_regression_cpp(
  X,
  y,
  warm_start_beta = NULL,
  smart_cold_start = FALSE,
  maxit = 100L,
  tol = 1e-8,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "irls",
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the response variable, expected to be binary (0 or 1).

warm_start_beta

Optional starting values for coefficients \beta. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE and no warm_start_beta is supplied, use an OLS-based initial guess rather than a zero cold start. Default FALSE for this function (unlike most of the package's other fast_* fitters, which default this to TRUE), since IRLS (the default optimizer here) is typically robust enough from a zero start for well-behaved logistic regression problems.

maxit

Maximum number of iterations for the algorithm. Defaults to 100.

tol

Convergence tolerance. Defaults to 1e-8.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm: "irls" (default, classical iteratively-reweighted-least-squares Fisher scoring), "lbfgs", or "newton_raphson"; see .normalize_optimizer_algorithm.

warm_start_weights

Optional initial IRLS working weights for the first iteration.

warm_start_fisher_info

Optional initial Fisher Information matrix to warm-start curvature information.

estimate_only

Logical. If TRUE, skip the working-weights/ score/Fisher-information computation after convergence, returning only b, converged, num_iter, hit_iteration_cap, and gradient_norm.

Value

A list containing the following components:

b

A numeric vector of the estimated logistic regression coefficients \hat\beta.

w

The IRLS working weights \hat\mu_i(1-\hat\mu_i) at the final iteration (the Bernoulli variance function evaluated at the fitted probabilities); omitted when estimate_only = TRUE.

num_iter

The number of optimizer iterations performed.

fisher_information

The working-weights curvature matrix X^\top W X; omitted when estimate_only = TRUE.

score

The score (gradient of the log-likelihood) vector at the fitted coefficients; omitted when estimate_only = TRUE.

neg_ll

The negative log-likelihood at the fitted coefficients; omitted when estimate_only = TRUE.

converged

A logical value indicating whether the final gradient norm was below tol (gradient_norm < tol); uniform across the "irls"/"lbfgs" optimizers.

hit_iteration_cap

A logical value, mutually exclusive with converged: TRUE iff the optimizer exhausted maxit iterations without meeting the gradient-norm convergence criterion.

gradient_norm

The norm of the score vector at the returned coefficients, a diagnostic of how tightly the convergence criterion was met.

See Also

fast_logistic_regression_with_var_cpp for the variance-augmented variant; fast_logistic_regression for the R-level wrapper; fast_log_binomial_regression_cpp/ fast_identity_binomial_regression_cpp for the log-link/ identity-link analogs targeting relative-risk/risk-difference scales.


Fast Weighted Logistic Regression, Estimate Only (C++ Backend)

Description

Fits the same logistic regression model as fast_logistic_regression_cpp (see that page for the full model, log-odds-ratio interpretation, and optimizer contract), with each observation's contribution to the log-likelihood and IRLS working weights multiplied by a row weight weights[i]. Setting all weights to 1 recovers fast_logistic_regression_cpp exactly; this is the backend used when the logistic model must be fit on bootstrap-reweighted or otherwise weighted data.

Usage

fast_logistic_regression_weighted_cpp(
  X,
  y,
  weights,
  warm_start_beta = NULL,
  smart_cold_start = FALSE,
  maxit = 100L,
  tol = 1e-8,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "irls",
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the response variable, expected to be binary (0 or 1).

weights

A numeric vector of weights for each observation.

warm_start_beta

Optional starting values for coefficients \beta. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE and no warm_start_beta is supplied, use an OLS-based initial guess rather than a zero cold start.

maxit

Maximum number of iterations for the IRLS algorithm. Defaults to 100.

tol

Convergence tolerance. Defaults to 1e-8.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm: "irls" (default), "lbfgs", or "newton_raphson"; see .normalize_optimizer_algorithm.

warm_start_weights

Optional initial IRLS working weights for the first iteration.

warm_start_fisher_info

Optional initial Fisher Information matrix to warm-start curvature information.

Value

A list containing the following components:

b

A numeric vector of the estimated logistic regression coefficients \hat\beta.

mu

The fitted probabilities \hat\mu_i.

XtWX, fisher_information

Two aliases for the same working-weights curvature matrix X^\top W X at the final iteration.

score

The (weighted) score vector at the fitted coefficients.

neg_ll

The weighted negative log-likelihood at the fitted coefficients.

converged

A logical value indicating whether the final gradient norm was below tol (gradient_norm < tol); uniform across the "irls"/"lbfgs" optimizers.

num_iter

The number of optimizer iterations performed.

hit_iteration_cap

A logical value, mutually exclusive with converged: TRUE iff the optimizer exhausted maxit iterations without meeting the gradient-norm convergence criterion.

gradient_norm

The norm of the score vector at the returned coefficients.

See Also

fast_logistic_regression_cpp for the unweighted model and full documentation.


Fast Logistic Regression with Variance, Auto-Retrying on Separation (R Wrapper)

Description

Fits the logistic regression model documented in full at fast_logistic_regression_cpp (log-odds-ratio interpretation, IRLS/L-BFGS/Newton-Raphson optimization) via fast_logistic_regression_with_var_cpp, and additionally detects and automatically retries on (quasi-)complete separation — the well-known logistic-regression failure mode where the MLE does not exist because some linear combination of covariates perfectly (or near-perfectly) predicts the outcome, causing the optimizer's coefficient estimates to diverge to a large-but-finite value that would otherwise silently pass ordinary is.finite() convergence checks and corrupt downstream confidence intervals.

Usage

fast_logistic_regression_with_var(
  X,
  y,
  j = 2,
  optimization_alg = "lbfgs",
  warm_start_beta = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the response variable, expected to be binary (0 or 1).

j

The index of the coefficient to compute the variance for. Defaults to 2.

optimization_alg

Optimization algorithm: "lbfgs" (default), "irls", or "newton_raphson".

warm_start_beta

Optional starting values for the coefficients.

warm_start_fisher_info

Optional initial Fisher Information matrix.

Details

Separation detection and retry. After each fit attempt, is_separated_coefficient_magnitude() checks whether any fitted coefficient exceeds a fixed separation-detection threshold (EDI_SEPARATION_THRESHOLD); if so, the fit is treated as converged = FALSE regardless of what the underlying C++ optimizer itself reported. On non-convergence (including detected separation), the covariate (column index \ge 3, i.e. never the intercept in column 1 or the treatment column in column 2) with the largest absolute fitted coefficient is dropped, and the model is refit on the reduced design; this repeats until either the fit converges or only the intercept and treatment columns remain. If separation persists even with just those two columns, this function stop()s with an explicit "complete separation detected" error rather than returning a corrupted variance estimate.

Interpretation caveat. Because covariates can be silently dropped by this retry loop, the returned b may have fewer coefficients than ncol(X) implies, and the fitted model's covariate adjustment set can differ from what was requested; callers relying on a specific covariate being present in the final fit should check for this rather than assume it always is.

Value

A list containing the following components:

b

A numeric vector of the obtained logistic regression coefficients \hat\beta, from whichever (possibly covariate-reduced) fit in the retry sequence ultimately converged.

ssq_b_j

The squared standard error (variance) of the j-th estimated coefficient.

ssq_b_2

The squared standard error (variance) of the second estimated coefficient, which typically corresponds to the treatment effect.

See Also

fast_logistic_regression_with_var_cpp for the underlying single-fit (no retry) backend and its variance-computation details; fast_logistic_regression_cpp for the full model documentation.

Examples

X = matrix(rnorm(100), 10, 10)
y = rbinom(10, 1, 0.5)
fast_logistic_regression_with_var(X, y)

Fast Logistic Regression with Targeted Variance (C++ Backend)

Description

Fits the same logistic regression model as fast_logistic_regression_cpp (see that page for the full model and log-odds-ratio interpretation) and additionally computes the variance of two coefficients — the caller-selected j-th coefficient and, separately, the 2nd coefficient (the package's usual treatment-effect position) — via a targeted diagonal-entry inversion of the working-weights Fisher information, rather than a full matrix inverse. Unlike fast_logistic_regression_cpp, maxit and tol are not exposed here: they are fixed internally at 100 iterations and 1e-8 tolerance.

Usage

fast_logistic_regression_with_var_cpp(
  X,
  y,
  j = 2L,
  warm_start_beta = NULL,
  smart_cold_start = FALSE,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "irls",
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the response variable, expected to be binary (0 or 1).

j

1-based index (into X's columns) of the coefficient to compute ssq_b_j for. Defaults to 2.

warm_start_beta

Optional starting values for coefficients \beta. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE and no warm_start_beta is supplied, use an OLS-based initial guess rather than a zero cold start.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm: "irls" (default), "lbfgs", or "newton_raphson"; see .normalize_optimizer_algorithm.

warm_start_weights

Optional initial IRLS working weights for the first iteration.

warm_start_fisher_info

Optional initial Fisher Information matrix to warm-start curvature information.

Details

Variance computation. The working-weights Fisher information X^\top W X is restricted to the free (non-fixed_idx) parameters; ssq_b_j and ssq_b_2 are each obtained via a single targeted diagonal-entry inversion (compute_diagonal_inverse_entry()) at the free-parameter position corresponding to j and to column 2, respectively — NA if the relevant coefficient is out of range or was itself fixed via fixed_idx. ssq_b_2 is always computed (when ncol(X) >= 2) regardless of what j is, so a caller interested in the treatment effect's variance does not need to pass j = 2 explicitly.

Value

A list with components b/params (estimated coefficients \hat\beta, identical to each other), ssq_b_j/ssq_b_2 (the two targeted coefficient variances described in Details), score (the score vector at \hat\beta), observed_information/fisher_information/ information (three aliases for the same working-weights curvature matrix, tagged information_type = "fisher"), hessian (the negative of that matrix), neg_loglik/neg_ll (aliases for the negative log-likelihood), loglik (its negation), converged (logical, gradient_norm < tol, uniform across optimizers), num_iter, hit_iteration_cap (logical, mutually exclusive with converged), and gradient_norm.

See Also

fast_logistic_regression_cpp for the estimate-only variant and the full model documentation.


Fast Negative Binomial Regression, Estimate Only (C++ Backend)

Description

Fits the negative-binomial regression model in its mean/dispersion parameterization documented in full at fast_dnbinom_mu_vec_cpp: log link \log \mu_i = x_i^\top \beta (so e^{\beta_j} is a multiplicative change in the mean count, as in Poisson regression), with a single dispersion parameter \theta shared across all observations and \mathrm{Var}(Y_i) = \mu_i + \mu_i^2/\theta (smaller \theta means more overdispersion relative to Poisson; \theta \to \infty recovers Poisson). The optimizer's parameter vector is c(beta, log(theta)) (\theta optimized on the log scale for positivity).

High-performance negative binomial regression fitting.

Usage

fast_neg_bin_cpp(
  X,
  y,
  warm_start_params = NULL,
  smart_cold_start = FALSE,
  maxit = 1000L,
  eps_f = 1e-08,
  eps_g = 1e-06,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors.

y

A numeric vector of responses.

warm_start_params

Optional starting values for coefficients and dispersion. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

maxit

Maximum number of iterations.

eps_f

Convergence tolerance for function value.

eps_g

Convergence tolerance for gradient.

fixed_idx

Optional indices of fixed parameters.

fixed_values

Optional values for fixed parameters.

optimization_alg

Optimization algorithm.

warm_start_fisher_info

Optional initial Fisher Information matrix for the first IRLS iteration.

estimate_only

Logical; if TRUE, may skip work not needed to produce point estimates (kept in sync with the package's other fast_* estimate-vs-inference split).

Value

A list containing the following components:

b

A numeric vector of the estimated negative binomial regression coefficients \hat\beta.

theta_hat

The estimated dispersion parameter \hat\theta (natural scale).

logLik

The model's log-likelihood at the fitted parameters.

converged

A logical value indicating whether the final gradient norm was below the convergence tolerance (gradient_norm < tol); uniform across the "lbfgs"/"newton_raphson" optimizers.

num_iter

The number of optimizer iterations performed.

hit_iteration_cap

A logical value, mutually exclusive with converged: TRUE iff the optimizer exhausted maxit iterations without meeting the gradient-norm convergence criterion.

gradient_norm

The norm of the gradient at the returned parameters.

fisher_information

The working-weights curvature matrix used during fitting.

A list containing coefficients, theta, and convergence status.

See Also

fast_dnbinom_mu_vec_cpp for the mean/dispersion density parameterization used here; fast_neg_bin_with_var_cpp for the variance-augmented variant.

Examples

X = matrix(rnorm(100), 10, 10)
y = rpois(10, 2)
fast_neg_bin_cpp(X, y)

Fast Weighted Negative Binomial Regression, Estimate Only (C++ Backend)

Description

Fits the same mean/dispersion-parameterized negative-binomial regression as fast_neg_bin_cpp (see that page, and fast_dnbinom_mu_vec_cpp, for the full model), with each observation's contribution to the log-likelihood multiplied by a nonnegative row weight weights[i]. Setting all weights to 1 recovers fast_neg_bin_cpp exactly.

Usage

fast_neg_bin_weighted_cpp(
  X,
  y,
  weights,
  warm_start_params = NULL,
  smart_cold_start = FALSE,
  maxit = 1000L,
  eps_f = 1e-08,
  eps_g = 1e-06,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors, n \times p.

y

A numeric (integer-valued) vector of non-negative observed counts, length n.

weights

A nonnegative numeric vector of length n giving each row's weight.

warm_start_params

Optional starting values for coefficients and dispersion. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

maxit

Maximum number of optimizer iterations.

eps_f

Convergence tolerance on the objective (log-likelihood) value.

eps_g

Convergence tolerance on the gradient norm.

fixed_idx

Optional integer indices (into the c(beta, log(theta)) parameter layout) of parameters to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm: "lbfgs" (default) or "newton_raphson".

warm_start_fisher_info

Optional initial Fisher Information matrix to warm-start curvature information.

estimate_only

If TRUE, skip Fisher information calculation.

Value

A list with the same components as fast_neg_bin_cpp: b, theta_hat, logLik, converged, iterations, and fisher_information (all reflecting the weighted log-likelihood).

See Also

fast_neg_bin_cpp for the unweighted model and full documentation.

Examples

X = matrix(rnorm(100), 10, 10)
y = rpois(10, 2)
fast_neg_bin_weighted_cpp(X, y, weights = rep(1, 10))

Fast Negative Binomial Regression with Variance Calculation (C++ Backend)

Description

Fits the same mean/dispersion-parameterized negative-binomial regression as fast_neg_bin_cpp (see that page, and fast_dnbinom_mu_vec_cpp, for the full model) and additionally computes the full variance-covariance matrix of c(beta, log(theta)).

Usage

fast_neg_bin_with_var_cpp(
  X,
  y,
  warm_start_params = NULL,
  smart_cold_start = FALSE,
  maxit = 1000L,
  eps_f = 1e-08,
  eps_g = 1e-06,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors, n \times p.

y

A numeric (integer-valued) vector of non-negative observed counts, length n.

warm_start_params

Optional starting values for coefficients and dispersion. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

maxit

Maximum number of optimizer iterations.

eps_f

Convergence tolerance on the objective (log-likelihood) value.

eps_g

Convergence tolerance on the gradient norm.

fixed_idx

Optional integer indices (into the c(beta, log(theta)) parameter layout) of parameters to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm: "lbfgs" (default) or "newton_raphson".

warm_start_fisher_info

Optional initial Fisher Information matrix to warm-start curvature information.

Details

Variance computation. The working-weights curvature matrix res.XtWX (over all p + 1 parameters) is restricted to the free (non-fixed_idx) parameters and inverted via a plain matrix inverse (.inverse(), not a rank-aware pseudo-inverse as used by, e.g., fast_adjacent_category_logit_with_var_cpp) before being expanded back to the full (p + 1) x (p + 1) size as vcov. A rank-deficient or near-singular design (after restricting to free parameters) will therefore produce numerically unstable or NaN variances rather than a graceful fallback.

Value

A list containing the following components:

b

A numeric vector of the obtained Poisson regression coefficients.

hess_fisher_info_matrix

The Fisher information matrix.

neg_ll

The negative log-likelihood at the final iteration.

converged

A logical value indicating whether the final gradient norm was below the convergence tolerance (gradient_norm < tol); uniform across the "lbfgs"/"newton_raphson" optimizers.

num_iter

The number of optimizer iterations performed.

hit_iteration_cap

A logical value, mutually exclusive with converged: TRUE iff the optimizer exhausted maxit iterations without meeting the gradient-norm convergence criterion.

gradient_norm

The norm of the gradient at the returned parameters.

A list with components b (\hat\beta), theta_hat (\hat\theta), logLik, vcov (the full (p + 1) x (p + 1) parameter variance-covariance matrix), converged, iterations, and hess_fisher_info_matrix (the working-weights curvature matrix vcov was inverted from).

See Also

fast_neg_bin_cpp for the estimate-only variant and full model documentation; fast_neg_bin_weighted_cpp for the row-weighted estimate-only variant.

Examples

X = matrix(rnorm(100), 10, 10)
y = rpois(10, 2)
fast_neg_bin_with_var_cpp(X, y)

Fast Negative Binomial Regression, Estimate-Only (R Wrapper)

Description

This function provides a fast implementation of mean/dispersion-parameterized negative binomial regression, wrapping a C++ backend (fast_neg_bin_cpp; see that page, and fast_dnbinom_mu_vec_cpp, for the full model). It returns point estimates only (no standard errors) — see fast_negbin_regression_with_var for the variance-computing counterpart. Columns 1 and 2 of X (conventionally the intercept and treatment indicator) are always kept; this wrapper adds automatic, silent handling of rank-deficient or numerically unstable covariate sets beyond those first two columns: (1) if warm_start_params is supplied, a single fit is attempted on the full matrix using the warm start, and its result is returned if successful; (2) otherwise, upfront, any covariate columns (3 onward) found rank-deficient by qr are dropped before the first fit attempt; (3) if the C++ fit still fails (e.g. an L-BFGS line-search failure), covariates are dropped one at a time, in reverse QR-pivot order (most redundant first), retrying after each drop, until the fit succeeds or only the intercept and treatment columns remain — at which point, if it still fails, this function stop()s with an explicit error rather than returning a corrupted fit.

Usage

fast_negbin_regression(
  X,
  y,
  optimization_alg = "lbfgs",
  warm_start_params = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the response variable, representing count data.

optimization_alg

Optimization algorithm: "lbfgs" (default) or "newton_raphson".

warm_start_params

Optional starting values for coefficients and log_theta, passed straight through to the C++ backend for a single warm-started fit attempt (bypassing the QR-dropping retry sequence).

warm_start_fisher_info

Optional initial Fisher information matrix, used only together with warm_start_params.

Value

A list containing the following components:

b

A numeric vector of the estimated negative binomial regression coefficients \hat\beta (and log_theta), from whichever (possibly covariate-reduced) fit in the retry sequence ultimately succeeded.

fisher_information

The C++ backend's returned Fisher information matrix for the successful fit.

See Also

fast_negbin_regression_with_var for the variance-computing, similarly retry-hardened wrapper; fast_neg_bin_cpp for the underlying backend and full model documentation.

Examples

X = matrix(rnorm(100), 10, 10)
y = rpois(10, 2)
fast_negbin_regression(X, y)

Fast Negative Binomial Regression with Variance Calculation (R Wrapper)

Description

This function provides a fast implementation of negative binomial regression, wrapping a C++ backend (fast_neg_bin_with_var_cpp; see that page, and fast_dnbinom_mu_vec_cpp, for the full mean/dispersion negative-binomial model). Columns 1 and 2 of X (conventionally the intercept and treatment indicator) are always kept; this wrapper adds automatic, silent handling of rank-deficient or numerically unstable covariate sets beyond those first two columns: (1) upfront, any covariate columns (3 onward) found rank-deficient by qr are dropped before the first fit attempt; (2) if the C++ fit still fails (e.g. an L-BFGS line-search failure), covariates are dropped one at a time, in reverse QR-pivot order (most redundant first), retrying after each drop, until the fit succeeds or only the intercept and treatment columns remain — at which point, if it still fails, this function stop()s with an explicit error rather than returning a corrupted fit.

Usage

fast_negbin_regression_with_var(X, y, j = 2, optimization_alg = "lbfgs")

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the response variable, representing count data.

j

The index of the coefficient to compute the variance for. Defaults to 2.

optimization_alg

Optimization algorithm: "lbfgs" (default) or "newton_raphson".

Value

A list containing the following components:

b

A numeric vector of the estimated negative binomial regression coefficients \hat\beta, from whichever (possibly covariate-reduced) fit in the retry sequence ultimately succeeded.

ssq_b_j

The variance of the j-th estimated coefficient, computed by inverting the C++ backend's returned Hessian/Fisher-information matrix directly (via solve, not the backend's own vcov); NA if that inversion fails or j exceeds the number of columns remaining after covariate-dropping.

ssq_b_2

The variance of the second estimated coefficient specifically (typically the treatment effect), computed the same way, regardless of what j is.

See Also

fast_neg_bin_with_var_cpp for the underlying backend (no automatic rank-deficiency retry) and full model documentation; fast_negbin_regression for the estimate-only, similarly retry-hardened wrapper.

Examples

X = matrix(rnorm(100), 10, 10)
y = rpois(10, 2)
fast_negbin_regression_with_var(X, y)

Fast Ordinary Least Squares (OLS) Regression, Estimate-Only (C++ Backend)

Description

Solves the ordinary least squares normal equations \hat\beta = (X^\top X)^{-1} X^\top y via Eigen's LDLT Cholesky decomposition of X^\top X; if that decomposition fails (e.g. X is rank-deficient), it falls back to a column-pivoted QR decomposition of X directly (Eigen::ColPivHouseholderQR), which returns a minimum-norm least-squares solution even when X is not full rank. Estimate-only: no standard errors or covariance matrix are computed, only the coefficient vector \hat\beta.

Usage

fast_ols_cpp(X, y, fixed_idx = NULL, fixed_values = NULL)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the (continuous) response variable.

fixed_idx

Optional integer vector of 1-indexed columns of X whose coefficients should be held fixed at fixed_values rather than estimated.

fixed_values

Optional numeric vector, parallel to fixed_idx, of the fixed coefficient values.

Value

A list containing the following component:

b

A numeric vector of the estimated regression coefficients \hat\beta (with any fixed_idx entries set to fixed_values); if the solve produces non-finite values, this is instead a vector of NaN.

Fixed (offset) coefficients

fixed_idx (1-indexed columns of X) and fixed_values optionally hold a subset of coefficients at caller-supplied constant values rather than estimating them: those columns' contribution X_{\mathrm{fixed}} \beta_{\mathrm{fixed}} is subtracted out of y first, and only the remaining ("free") columns are fit by least squares; the fixed coefficients are then copied back into \hat\beta unchanged. This is used, e.g., to fit a model with a known/offset intercept without re-estimating it.

See Also

fast_ols_with_var_cpp for the variance-computing counterpart.


Fast Ordinary Least Squares (OLS) Regression with Variance (C++ Backend)

Description

As fast_ols_cpp, but additionally computes the classical OLS variance estimate \hat\sigma^2 = \mathrm{SSE} / (n - p) (with \mathrm{SSE} = y^\top y - \hat\beta^\top X^\top y on the fixed-parameter-adjusted response, and p the number of free — non-fixed — columns) and the sampling variance of two coefficients, \widehat{\mathrm{Var}}(\hat\beta_k) = \hat\sigma^2 \, [(X^\top X)^{-1}]_{kk}, obtained from the same Cholesky (LDLT) factorization used to solve for \hat\beta, without forming the full inverse matrix. If the LDLT decomposition fails (rank-deficient X), this function falls back to a QR solve exactly as fast_ols_cpp does, but in that case no variance quantities are computed: ssq_b_j, ssq_b_2, and XtX are omitted from the result and converged is FALSE.

Usage

fast_ols_with_var_cpp(X, y, j = 2L, fixed_idx = NULL, fixed_values = NULL)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the (continuous) response variable.

j

This function will compute the variance of the jth (1-indexed) coefficient estimator. Default is 2 (conventionally the treatment effect).

fixed_idx

Optional integer vector of 1-indexed columns of X whose coefficients should be held fixed at fixed_values rather than estimated.

fixed_values

Optional numeric vector, parallel to fixed_idx, of the fixed coefficient values.

Value

A list containing the following components (the last four only present when the LDLT solve succeeds):

b

A numeric vector of the estimated regression coefficients \hat\beta.

converged

TRUE if the LDLT solve succeeded, FALSE if the QR fallback was used.

sigma2_hat

The estimated residual variance \hat\sigma^2.

XtX

The (free-coefficient) X^\top X matrix, expanded back to full p \times p shape with zeros in the fixed-coefficient rows/columns.

ssq_b_j

The variance of the j-th coefficient estimator, NA if j indexes a fixed coefficient.

ssq_b_2

The variance of the second coefficient estimator specifically (typically the treatment effect), regardless of what j is; equal to ssq_b_j when j = 2. NA if the second column is a fixed coefficient.

Fixed (offset) coefficients

fixed_idx (1-indexed columns of X) and fixed_values optionally hold a subset of coefficients at caller-supplied constant values rather than estimating them: those columns' contribution X_{\mathrm{fixed}} \beta_{\mathrm{fixed}} is subtracted out of y first, and only the remaining ("free") columns are fit by least squares; the fixed coefficients are then copied back into \hat\beta unchanged. This is used, e.g., to fit a model with a known/offset intercept without re-estimating it.

See Also

fast_ols_cpp for the estimate-only counterpart.


Fast Cumulative Ordinal Regression with a Cauchit Link (C++)

Description

Fits a cumulative-link ordinal regression model with the cauchit link (the inverse CDF of the standard Cauchy distribution) via direct maximum likelihood, jointly optimizing the category thresholds and regression coefficients. y's distinct values (in sorted order, whatever their original coding) are treated as K ordered categories; for observation i in category k (k = 0, \ldots, K-1),

\Pr(Y_i \le k \mid x_i) = F(\alpha_k - x_i^\top \beta), \qquad F(z) = \frac{1}{2} + \frac{\arctan(z)}{\pi},

with \alpha_0 < \alpha_1 < \cdots < \alpha_{K-2} the (increasing) category thresholds and \beta the regression coefficients on X (no separate intercept column is needed — the thresholds serve that role); the category probability is the corresponding CDF difference, \Pr(Y_i = k \mid x_i) = F(\alpha_k - x_i^\top\beta) - F(\alpha_{k-1} - x_i^\top\beta) (with F(\alpha_{-1} - \cdot) := 0 and F(\alpha_{K-1} - \cdot) := 1 at the boundaries), each clamped below at 10^{-12} before taking logs for numerical safety. This is a direct-likelihood analogue of ordinal logistic regression with the logit link swapped for the heavier-tailed Cauchy CDF, which is more robust to outlying/misclassified extreme categories at the cost of less standard interpretability (no proportional-odds log-odds-ratio reading of \beta). If y has fewer than 2 distinct levels, an empty result list is returned (the model is degenerate).

Usage

fast_ordinal_cauchit_regression_cpp(
  X,
  y,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  maxit = 100L,
  tol = 1e-06,
  optimization_alg = "lbfgs",
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see Details).

y

A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding.

warm_start_params

Optional starting values for [\alpha, \beta]. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

maxit

Maximum number of optimizer iterations.

tol

Convergence tolerance.

optimization_alg

Optimization algorithm (default "lbfgs").

fixed_idx

Optional 1-indexed positions (into [\alpha,\beta]) of parameters to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

warm_start_fisher_info

Optional initial curvature (Fisher/observed information) matrix.

estimate_only

If TRUE, skip the post-fit Hessian/variance computation and return only point estimates (faster). If FALSE (the default), also compute and return the observed information matrix and, when the fit converged, the full variance-covariance matrix (identical to what fast_ordinal_cauchit_regression_with_var_cpp always computes).

Value

A list with components b (the \beta coefficients), alpha (the K-1 category thresholds), params (the concatenated [\alpha, \beta] vector), n_params, converged, and iterations; when estimate_only = FALSE (the default) and the fit converged, additionally neg_loglik, observed_information/fisher_information/information (all the same observed-information matrix), information_type (always "observed"), vcov (the parameter covariance matrix), and ssq_b_j (the variance of b[1], i.e. the coefficient on X's first column — conventionally the treatment effect, since X carries no separate intercept column here; the thresholds alpha play that role). Empty if y has fewer than 2 distinct levels.

Fixed parameters, warm starts, and optimization

fixed_idx (1-indexed into the combined [\alpha, \beta] parameter vector, thresholds first) and fixed_values optionally hold a subset of parameters at caller-supplied constant values rather than estimating them. warm_start_params supplies starting values for [\alpha, \beta] directly (skipping smart_cold_start); otherwise, when smart_cold_start = TRUE (the default), starting values come from an OLS-based heuristic, and when FALSE, thresholds start at \tan(\pi(k/K - 1/2)) (an inverse-cauchit spacing of the empirical marginal category proportions) with \beta at zero. Optimization runs via optimization_alg (default "lbfgs") for up to maxit iterations or until the parameter/gradient change falls below tol; warm_start_fisher_info, if supplied, seeds the first iteration's curvature estimate.

See Also

fast_ordinal_cauchit_regression_with_var_cpp, which always computes the variance quantities (equivalent to calling this function with estimate_only = FALSE) and additionally guards against the degenerate/non-converged case by returning NA placeholders instead of an empty list.


Fast Cumulative Ordinal Regression with a Cauchit Link, with Variance (C++)

Description

As fast_ordinal_cauchit_regression_cpp (see that page for the full cumulative cauchit-link model, \Pr(Y_i \le k \mid x_i) = F(\alpha_k - x_i^\top\beta)), but always fits with estimate_only = FALSE (equivalent to calling that function with its default), so the observed information matrix and variance-covariance matrix are always computed. It additionally guards the degenerate case: if y has fewer than 2 distinct levels (so the underlying fit is empty), this function returns list(b = NA, ssq_b_2 = NA) instead of an empty list.

Usage

fast_ordinal_cauchit_regression_with_var_cpp(
  X,
  y,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  optimization_alg = "lbfgs",
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see fast_ordinal_cauchit_regression_cpp).

y

A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding.

warm_start_params

Optional starting values for [\alpha, \beta]. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

optimization_alg

Optimization algorithm (default "lbfgs").

fixed_idx

Optional 1-indexed positions (into [\alpha,\beta]) of parameters to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

warm_start_fisher_info

Optional initial curvature (Fisher/observed information) matrix.

Value

A list with components b, alpha, params, n_params, neg_loglik, converged, iterations, observed_information/fisher_information/information (all the same observed-information matrix), information_type ("observed"), vcov, and ssq_b_2 (the variance of b[1], i.e. X's first-column coefficient — conventionally the treatment effect). vcov is omitted (and ssq_b_2 is NA) if the fit did not converge; list(b = NA, ssq_b_2 = NA) if y has fewer than 2 distinct levels.

Fixed parameters, warm starts, and optimization

fixed_idx (1-indexed into the combined [\alpha, \beta] parameter vector, thresholds first) and fixed_values optionally hold a subset of parameters at caller-supplied constant values rather than estimating them. warm_start_params supplies starting values for [\alpha, \beta] directly (skipping smart_cold_start); otherwise, when smart_cold_start = TRUE (the default), starting values come from an OLS-based heuristic, and when FALSE, thresholds start at \tan(\pi(k/K - 1/2)) (an inverse-cauchit spacing of the empirical marginal category proportions) with \beta at zero. Optimization runs via optimization_alg (default "lbfgs") for up to maxit iterations or until the parameter/gradient change falls below tol; warm_start_fisher_info, if supplied, seeds the first iteration's curvature estimate.

See Also

fast_ordinal_cauchit_regression_cpp for the estimate-only-capable variant and the full model documentation.


Fast Cumulative Ordinal Regression with a Complementary Log-Log Link (C++)

Description

Fits a cumulative-link ordinal regression model with the complementary log-log ("cloglog") link via direct maximum likelihood, jointly optimizing the category thresholds and regression coefficients. y's distinct values (in sorted order, whatever their original coding) are treated as K ordered categories; for observation i in category k (k = 0, \ldots, K-1),

\Pr(Y_i \le k \mid x_i) = F(\alpha_k + x_i^\top \beta), \qquad F(z) = 1 - \exp(-\exp(z)),

with \alpha_0 < \alpha_1 < \cdots < \alpha_{K-2} the (increasing) category thresholds and \beta the regression coefficients on X (no separate intercept column is needed — the thresholds serve that role); the category probability is the corresponding CDF difference, \Pr(Y_i = k \mid x_i) = F(\alpha_k + x_i^\top\beta) - F(\alpha_{k-1} + x_i^\top\beta) (with F(\alpha_{-1} + \cdot) := 0 and F(\alpha_{K-1} + \cdot) := 1 at the boundaries), each clamped below at 10^{-12} before taking logs for numerical safety. Unlike the symmetric logit/probit/cauchit links, cloglog is asymmetric: it is the natural link for a grouped/discretized proportional-hazards (continuation-ratio-free) survival model, appropriate when category probabilities are skewed toward the lower categories. If y has fewer than 2 distinct levels, an empty result list is returned (the model is degenerate).

Usage

fast_ordinal_cloglog_regression_cpp(
  X,
  y,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  maxit = 100L,
  tol = 1e-06,
  optimization_alg = "lbfgs",
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see Details).

y

A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding.

warm_start_params

Optional starting values for [\alpha, \beta]. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

maxit

Maximum number of optimizer iterations.

tol

Convergence tolerance.

optimization_alg

Optimization algorithm (default "lbfgs").

fixed_idx

Optional 1-indexed positions (into [\alpha,\beta]) of parameters to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

warm_start_fisher_info

Optional initial curvature (Fisher/observed information) matrix.

estimate_only

If TRUE, skip the post-fit Hessian/variance computation and return only point estimates (faster). If FALSE (the default), also compute and return the observed information matrix and, when the fit converged, the full variance-covariance matrix (identical to what fast_ordinal_cloglog_regression_with_var_cpp always computes).

Value

A list with components b (the \beta coefficients), alpha (the K-1 category thresholds), params (the concatenated [\alpha, \beta] vector), n_params, converged, and iterations; when estimate_only = FALSE (the default) and the fit converged, additionally neg_loglik, observed_information/fisher_information/information (all the same observed-information matrix), information_type (always "observed"), vcov (the parameter covariance matrix), and ssq_b_j (the variance of b[1], i.e. the coefficient on X's first column — conventionally the treatment effect, since X carries no separate intercept column here; the thresholds alpha play that role). Empty if y has fewer than 2 distinct levels.

Fixed parameters, warm starts, and optimization

fixed_idx (1-indexed into the combined [\alpha, \beta] parameter vector, thresholds first) and fixed_values optionally hold a subset of parameters at caller-supplied constant values rather than estimating them. warm_start_params supplies starting values for [\alpha, \beta] directly (skipping smart_cold_start); otherwise, when smart_cold_start = TRUE (the default), starting values come from an OLS-based heuristic, and when FALSE, thresholds start evenly spaced on (-1, 1) at -1 + 2(k+1)/K with \beta at zero. Optimization runs via optimization_alg (default "lbfgs") for up to maxit iterations or until the parameter/gradient change falls below tol; warm_start_fisher_info, if supplied, seeds the first iteration's curvature estimate.

See Also

fast_ordinal_cloglog_regression_with_var_cpp, which always computes the variance quantities (equivalent to calling this function with estimate_only = FALSE) and additionally guards against the degenerate/non-converged case by returning NA placeholders instead of an empty list.


Fast Cumulative Ordinal Regression with a Complementary Log-Log Link, with Variance (C++)

Description

As fast_ordinal_cloglog_regression_cpp (see that page for the full cumulative cloglog-link model, \Pr(Y_i \le k \mid x_i) = F(\alpha_k + x_i^\top\beta), F(z) = 1 - \exp(-\exp(z))), but always fits with estimate_only = FALSE (equivalent to calling that function with its default), so the observed information matrix and variance-covariance matrix are always computed. It additionally guards the degenerate case: if y has fewer than 2 distinct levels (so the underlying fit is empty), this function returns list(b = NA, ssq_b_2 = NA) instead of an empty list.

Usage

fast_ordinal_cloglog_regression_with_var_cpp(
  X,
  y,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  optimization_alg = "lbfgs",
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see fast_ordinal_cloglog_regression_cpp).

y

A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding.

warm_start_params

Optional starting values for [\alpha, \beta]. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

optimization_alg

Optimization algorithm (default "lbfgs").

fixed_idx

Optional 1-indexed positions (into [\alpha,\beta]) of parameters to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

warm_start_fisher_info

Optional initial curvature (Fisher/observed information) matrix.

Value

A list with components b, alpha, params, n_params, neg_loglik, converged, iterations, observed_information/fisher_information/information (all the same observed-information matrix), information_type ("observed"), vcov, and ssq_b_2 (the variance of b[1], i.e. X's first-column coefficient — conventionally the treatment effect). vcov is omitted (and ssq_b_2 is NA) if the fit did not converge; list(b = NA, ssq_b_2 = NA) if y has fewer than 2 distinct levels.

Fixed parameters, warm starts, and optimization

fixed_idx (1-indexed into the combined [\alpha, \beta] parameter vector, thresholds first) and fixed_values optionally hold a subset of parameters at caller-supplied constant values rather than estimating them. warm_start_params supplies starting values for [\alpha, \beta] directly (skipping smart_cold_start); otherwise, when smart_cold_start = TRUE (the default), starting values come from an OLS-based heuristic, and when FALSE, thresholds start evenly spaced on (-1, 1) at -1 + 2(k+1)/K with \beta at zero. Optimization runs via optimization_alg (default "lbfgs") for up to maxit iterations or until the parameter/gradient change falls below tol; warm_start_fisher_info, if supplied, seeds the first iteration's curvature estimate.

See Also

fast_ordinal_cloglog_regression_cpp for the estimate-only-capable variant and the full model documentation.


Fast Ordinal Cumulative-Logit Random-Intercept GLMM via Gauss-Hermite Quadrature (C++)

Description

Fits a cumulative-logit ordinal mixed model with a single Gaussian random intercept per group (e.g. a matched pair or singleton from a KK-style matched design):

\mathrm{logit}\,\Pr(Y_{ij} \le k \mid u_i) = \alpha_k - x_{ij}^\top\beta - u_i, \qquad u_i \sim N(0, \sigma^2),

for group i, member j, ordinal outcome y_{ij} \in \{1, \ldots, K\}, and increasing cutpoints \alpha_1 < \cdots < \alpha_{K-1} (X carries no separate intercept column — the cutpoints serve that role). The marginal likelihood for group i integrates the random intercept out,

L_i(\theta) = \int \prod_{j \in i} \Pr(Y_{ij} = y_{ij} \mid u_i) \, \phi(u_i / \sigma) \, du_i,

approximated by n_gh-point Gauss-Hermite quadrature (substituting u = \sqrt{2}\,\sigma\, z for quadrature node z), and optimized on the log scale by directly maximizing \sum_i \log L_i(\theta) (via optimize_fixed_likelihood, default optimization_alg = "lbfgs") over the reparameterized vector [\alpha_1, \log(\alpha_2-\alpha_1), \ldots, \log(\alpha_{K-1}-\alpha_{K-2}), \beta, \log\sigma] — cutpoints are recovered as successive partial sums of \alpha_1 and the exponentiated log-differences, which enforces \alpha_1 < \cdots < \alpha_{K-1} by construction rather than as a fitting constraint. Rows are stably sorted by group_id inside the kernel, so matched-group membership is invariant to input row order. Optimization uses supplied/cold, moderate-variance, and near-zero-variance starts, retains the smallest finite negative log-likelihood, and polishes that solution before applying a projected-gradient convergence check. If finite multistart L-BFGS stops on function decrease while its projected score remains above max(1e-5, eps_g), the kernel performs a local damped-Newton polish using its numerical Hessian. A Newton trial is retained only when its parameters, objective, and gradient are finite and its objective does not exceed the L-BFGS objective. At a valid lower log_sigma boundary, the KKT-satisfied variance coordinate is excluded from the Newton system, so fixed-effect convergence can be established without rejecting a near-zero random-effect variance. log_sigma is evaluated within [-\code{max\_abs\_log\_sigma}, \code{max\_abs\_log\_sigma}] with a quadratic penalty on excursions beyond that interval whose analytic gradient matches the bounded objective; variance_boundary_hit in the return value flags whether the fitted log_sigma landed at that boundary (a sign the random-intercept variance is being driven to (near-)zero or is unbounded, and that ssq_b_T/fisher_information should be treated with caution). The Hessian used for inference is a numerical (central finite-difference, step 10^{-4}, symmetrized) second derivative of the analytic gradient, not a closed-form expression. At the valid near-zero variance boundary, treatment variance is computed conditional on that boundary by excluding the nonregular variance-parameter row and column.

Usage

fast_ordinal_glmm_cpp(
  X,
  y,
  group_id,
  K,
  j_T,
  smart_cold_start = TRUE,
  estimate_only = FALSE,
  n_gh = 20L,
  max_abs_log_sigma = 8,
  maxit = 300L,
  eps_g = 1e-06,
  warm_start_params = NULL,
  warm_start_beta = NULL,
  optimization_alg = "lbfgs",
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors, one row per observation (member-level, not group-level); no intercept column (see Details).

y

Integer vector of 1-indexed ordinal outcomes (1, \ldots, K), one per row of X.

group_id

Integer vector of group (e.g. matched-pair) identifiers, one per row of X; assumed contiguous per group after internal sorting-by-value into blocks.

K

The number of ordinal levels.

j_T

0-based column index of X whose coefficient's variance (ssq_b_T) should be computed — typically the treatment indicator.

smart_cold_start

Logical. If TRUE and no warm start is given, initialize cutpoints evenly at 0 (all log-differences zero) and \beta via a naive OLS fit of y (treated as numeric) on X; log_sigma starts at -3. If FALSE, \beta instead starts at zero.

estimate_only

If TRUE, skip the (converged-fit-only) covariance calculation for ssq_b_T — point estimates and the Hessian are still returned regardless.

n_gh

Number of Gauss-Hermite quadrature nodes used to integrate out the random intercept.

max_abs_log_sigma

Symmetric clamp bound for log_sigma during optimization (default 8).

maxit

Maximum number of optimizer iterations.

eps_g

Gradient-norm convergence tolerance.

warm_start_params

Optional starting values for the full reparameterized vector [\alpha_1, \log\text{-diffs}, \beta, \log\sigma]; if its length doesn't match, falls back to zero cutpoints/\beta and log_sigma = -3. Takes precedence over warm_start_beta and smart_cold_start.

warm_start_beta

Optional starting values either for the full parameter vector (same length as warm_start_params above) or for \beta alone (length p, with cutpoints zeroed and log_sigma = -3); ignored if warm_start_params is supplied.

optimization_alg

Optimization algorithm (default "lbfgs").

fixed_idx

Optional 1-indexed positions (into the reparameterized parameter vector) to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

warm_start_fisher_info

Optional initial curvature matrix for the first optimizer iteration.

Value

A list with components b (\hat\beta), alpha (the K-1 cutpoints, recovered from the reparameterization), params (the full fitted reparameterized vector), log_sigma, ssq_b_T (variance of b[j_T], NA unless estimate_only = FALSE and the fit converged and the resulting information matrix inverts successfully), converged, neg_loglik, fisher_information (the numerical Hessian, always returned), score (the log-likelihood score at the returned parameters), gradient_norm, newton_polish_attempted, newton_polish_accepted, and newton_polish_iterations (diagnostics for the conditional damped-Newton fallback), and variance_boundary_hit (TRUE/FALSE, or NA if the optimizer itself threw an exception, in which case converged = FALSE and all other quantities besides b/alpha/log_sigma are NA).

References

Pinheiro, J. C., and Bates, D. M. (1995). "Approximations to the Log-Likelihood Function in the Nonlinear Mixed-Effects Model." Journal of Computational and Graphical Statistics, 4(1), 12-35, doi:10.1080/10618600.1995.10474663, for Gauss-Hermite quadrature as an approximation to the random-effect marginal likelihood integral used throughout this package's GLMM backends (fast_poisson_glmm_cpp, fast_logistic_glmm_cpp, fast_weibull_frailty_cpp, and this function).


Fast Cumulative Ordinal Regression with a Probit Link (C++)

Description

Fits a cumulative-link ordinal regression model with the probit link (the standard normal CDF) via direct maximum likelihood, jointly optimizing the category thresholds and regression coefficients. y's distinct values (in sorted order, whatever their original coding) are treated as K ordered categories; for observation i in category k (k = 0, \ldots, K-1),

\Pr(Y_i \le k \mid x_i) = \Phi(\alpha_k - x_i^\top \beta),

with \alpha_0 < \alpha_1 < \cdots < \alpha_{K-2} the (increasing) category thresholds and \beta the regression coefficients on X (no separate intercept column is needed — the thresholds serve that role); the category probability is the corresponding CDF difference, \Pr(Y_i = k \mid x_i) = \Phi(\alpha_k - x_i^\top\beta) - \Phi(\alpha_{k-1} - x_i^\top\beta) (with \Phi(\alpha_{-1} - \cdot) := 0 and \Phi(\alpha_{K-1} - \cdot) := 1 at the boundaries), each clamped below at 10^{-12} before taking logs for numerical safety. This is the ordinal generalization of probit regression, and the thin-tailed counterpart to fast_ordinal_cauchit_regression_cpp and the standard-logit ordinal model. If y has fewer than 2 distinct levels, an empty result list is returned (the model is degenerate).

Usage

fast_ordinal_probit_regression_cpp(
  X,
  y,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  maxit = 100L,
  tol = 1e-06,
  optimization_alg = "lbfgs",
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see Details).

y

A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding.

warm_start_params

Optional starting values for [\alpha, \beta]. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

maxit

Maximum number of optimizer iterations.

tol

Convergence tolerance.

optimization_alg

Optimization algorithm (default "lbfgs").

fixed_idx

Optional 1-indexed positions (into [\alpha,\beta]) of parameters to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

warm_start_fisher_info

Optional initial curvature (Fisher/observed information) matrix.

estimate_only

If TRUE, skip the post-fit Hessian/variance computation and return only point estimates (faster). If FALSE (the default), also compute and return the observed information matrix and, when the fit converged, the full variance-covariance matrix (identical to what fast_ordinal_probit_regression_with_var_cpp always computes).

Value

A list with components b (the \beta coefficients), alpha (the K-1 category thresholds), params (the concatenated [\alpha, \beta] vector), n_params, converged, and iterations; when estimate_only = FALSE (the default) and the fit converged, additionally neg_loglik, observed_information/fisher_information/information (all the same observed-information matrix), information_type (always "observed"), vcov (the parameter covariance matrix), and ssq_b_j (the variance of b[1], i.e. the coefficient on X's first column — conventionally the treatment effect, since X carries no separate intercept column here; the thresholds alpha play that role — set to NA if it comes out non-finite or non-positive). Empty if y has fewer than 2 distinct levels.

Fixed parameters, warm starts, and optimization

fixed_idx (1-indexed into the combined [\alpha, \beta] parameter vector, thresholds first) and fixed_values optionally hold a subset of parameters at caller-supplied constant values rather than estimating them. warm_start_params supplies starting values for [\alpha, \beta] directly (skipping smart_cold_start); otherwise, when smart_cold_start = TRUE (the default), starting values come from an OLS-based heuristic, and when FALSE, thresholds start at \Phi^{-1}(k/K) (inverse-normal spacing of the empirical marginal category proportions) with \beta at zero. Optimization runs via optimization_alg (default "lbfgs") for up to maxit iterations or until the parameter/gradient change falls below tol; warm_start_fisher_info, if supplied, seeds the first iteration's curvature estimate.

See Also

fast_ordinal_probit_regression_with_var_cpp, which always computes the variance quantities (equivalent to calling this function with estimate_only = FALSE) and additionally guards against the degenerate/non-converged case by returning NA placeholders instead of an empty list.


Fast Cumulative Ordinal Regression with a Probit Link, with Variance (C++)

Description

As fast_ordinal_probit_regression_cpp (see that page for the full cumulative probit-link model, \Pr(Y_i \le k \mid x_i) = \Phi(\alpha_k - x_i^\top\beta)), but always fits with estimate_only = FALSE (equivalent to calling that function with its default), so the observed information matrix and variance-covariance matrix are always computed. It additionally guards the degenerate case: if y has fewer than 2 distinct levels (so the underlying fit is empty), this function returns list(b = NA, ssq_b_2 = NA) instead of an empty list.

Usage

fast_ordinal_probit_regression_with_var_cpp(
  X,
  y,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  optimization_alg = "lbfgs",
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see fast_ordinal_probit_regression_cpp).

y

A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding.

warm_start_params

Optional starting values for [\alpha, \beta]. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

optimization_alg

Optimization algorithm (default "lbfgs").

fixed_idx

Optional 1-indexed positions (into [\alpha,\beta]) of parameters to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

warm_start_fisher_info

Optional initial curvature (Fisher/observed information) matrix.

Value

A list with components b, alpha, params, n_params, neg_loglik, converged, iterations, observed_information/fisher_information/information (all the same observed-information matrix), information_type ("observed"), vcov, and ssq_b_2 (the variance of b[1], i.e. X's first-column coefficient — conventionally the treatment effect; NA if it comes out non-finite or non-positive). vcov is omitted (and ssq_b_2 is NA) if the fit did not converge; list(b = NA, ssq_b_2 = NA) if y has fewer than 2 distinct levels.

Fixed parameters, warm starts, and optimization

fixed_idx (1-indexed into the combined [\alpha, \beta] parameter vector, thresholds first) and fixed_values optionally hold a subset of parameters at caller-supplied constant values rather than estimating them. warm_start_params supplies starting values for [\alpha, \beta] directly (skipping smart_cold_start); otherwise, when smart_cold_start = TRUE (the default), starting values come from an OLS-based heuristic, and when FALSE, thresholds start at \Phi^{-1}(k/K) (inverse-normal spacing of the empirical marginal category proportions) with \beta at zero. Optimization runs via optimization_alg (default "lbfgs") for up to maxit iterations or until the parameter/gradient change falls below tol; warm_start_fisher_info, if supplied, seeds the first iteration's curvature estimate.

See Also

fast_ordinal_probit_regression_cpp for the estimate-only-capable variant and the full model documentation.


Fast Cumulative Ordinal Regression with a Logit Link, i.e. Proportional-Odds Regression (C++)

Description

Fits the classical proportional-odds ordinal regression model (a cumulative-link model with the logit link) via direct maximum likelihood, jointly optimizing the category thresholds and regression coefficients. y's distinct values (in sorted order, whatever their original coding) are treated as K ordered categories; for observation i in category k (k = 0, \ldots, K-1),

\mathrm{logit}\,\Pr(Y_i \le k \mid x_i) = \alpha_k - x_i^\top \beta,

with \alpha_0 < \alpha_1 < \cdots < \alpha_{K-2} the (increasing) category thresholds and \beta the regression coefficients on X (no separate intercept column is needed — the thresholds serve that role); the category probability is the corresponding CDF difference, each clamped below at 10^{-12} before taking logs for numerical safety. The "proportional odds" name reflects that \beta does not depend on k: the odds ratio \exp(-\beta_j) for a unit increase in covariate j is the same across every cumulative cutpoint. If y has fewer than 2 distinct levels, fitting still proceeds with K = 1, n_alpha = 0 (degenerate: no thresholds to estimate); unlike the cauchit/probit/cloglog variants in this package, this function does not special-case that as an early return.

Usage

fast_ordinal_regression_cpp(
  X,
  y,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  maxit = 100L,
  tol = 1e-06,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see Details).

y

A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding.

warm_start_params

Optional starting values for [\alpha, \beta]. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

maxit

Maximum number of optimizer iterations.

tol

Convergence tolerance.

fixed_idx

Optional 1-indexed positions (into [\alpha,\beta]) of parameters to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

optimization_alg

Optimization algorithm (default "lbfgs").

warm_start_fisher_info

Optional initial curvature (Fisher/observed information) matrix.

estimate_only

If TRUE, skip the post-fit Hessian/variance computation and return only point estimates (faster). If FALSE (the default), also compute and return the observed information matrix and (if the resulting free-parameter information submatrix is invertible via Eigen::FullPivLU) the full variance-covariance matrix.

Value

A list with components b (the \beta coefficients), alpha (the K-1 category thresholds), params (the concatenated [\alpha, \beta] vector), n_params, converged, and iterations; when estimate_only = FALSE (the default), additionally neg_loglik, observed_information/fisher_information/information (all the same observed-information matrix), information_type (always "observed"), ssq_b_j (the variance of b[1], i.e. the coefficient on X's first column — conventionally the treatment effect, since X carries no separate intercept column here), and vcov (the full parameter covariance matrix) — the latter two are NA/omitted (vcov becomes NULL) if the free-parameter information matrix is not invertible.

Fixed parameters, warm starts, and optimization

fixed_idx (1-indexed into the combined [\alpha, \beta] parameter vector, thresholds first) and fixed_values optionally hold a subset of parameters at caller-supplied constant values rather than estimating them. warm_start_params supplies starting values for [\alpha, \beta] directly (skipping smart_cold_start); otherwise thresholds always start evenly spaced on (-1, 1) at -1 + 2(k+1)/K, and \beta starts at either zero, or (when smart_cold_start = TRUE, the default) an OLS fit of the rank-rescaled response (y - 1)/(K - 1) on X — falling back silently to zero if that OLS solve is not well-posed. When smart_cold_start = TRUE and no warm_start_fisher_info is supplied, the Hessian at the starting values is additionally used to seed the optimizer's first-iteration curvature estimate. Optimization runs via optimization_alg (default "lbfgs") for up to maxit iterations or until the parameter/gradient change falls below tol.

See Also

fast_ordinal_regression_weighted_cpp for the observation-weighted variant; fast_ordinal_regression_with_var_cpp, which additionally guards the non-invertible case with explicit NA placeholders.


Fast Cumulative Ordinal Regression with a Logit Link, Weighted (C++)

Description

As fast_ordinal_regression_cpp (see that page for the full proportional-odds model), but each observation's log-likelihood contribution is multiplied by a nonnegative weights[i] (negative weights are clamped to zero internally by the underlying weighted log-likelihood). Always fits with estimate_only = FALSE (equivalent to that function's default), so the observed information and, when invertible, the full variance-covariance matrix are always computed.

Usage

fast_ordinal_regression_weighted_cpp(
  X,
  y,
  weights,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  maxit = 100L,
  tol = 1e-06,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see fast_ordinal_regression_cpp).

y

A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding.

weights

A numeric vector of observation weights, length nrow(X) (negative entries are treated as zero).

warm_start_params

Optional starting values for [\alpha, \beta]. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

maxit

Maximum number of optimizer iterations.

tol

Convergence tolerance.

fixed_idx

Optional 1-indexed positions (into [\alpha,\beta]) of parameters to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

optimization_alg

Optimization algorithm (default "lbfgs").

warm_start_fisher_info

Optional initial curvature (Fisher/observed information) matrix.

Value

A list with components b, alpha, params, n_params, converged, iterations, neg_loglik, observed_information/fisher_information/information (all the same observed-information matrix), information_type ("observed"), ssq_b_j (the variance of b[1]), and vcov — the latter two are NA/omitted if the free-parameter information matrix is not invertible.

Fixed parameters, warm starts, and optimization

fixed_idx (1-indexed into the combined [\alpha, \beta] parameter vector, thresholds first) and fixed_values optionally hold a subset of parameters at caller-supplied constant values rather than estimating them. warm_start_params supplies starting values for [\alpha, \beta] directly (skipping smart_cold_start); otherwise thresholds always start evenly spaced on (-1, 1) at -1 + 2(k+1)/K, and \beta starts at either zero, or (when smart_cold_start = TRUE, the default) an OLS fit of the rank-rescaled response (y - 1)/(K - 1) on X — falling back silently to zero if that OLS solve is not well-posed. When smart_cold_start = TRUE and no warm_start_fisher_info is supplied, the Hessian at the starting values is additionally used to seed the optimizer's first-iteration curvature estimate. Optimization runs via optimization_alg (default "lbfgs") for up to maxit iterations or until the parameter/gradient change falls below tol.

See Also

fast_ordinal_regression_cpp for the unweighted variant and the full model documentation.


Fast Cumulative Ordinal Regression with a Logit Link, with Variance (C++)

Description

Identical to fast_ordinal_regression_cpp (see that page for the full proportional-odds model) called with estimate_only = FALSE — this is simply a convenience export that hardcodes that default rather than exposing the flag, so the observed information and, when invertible, the full variance-covariance matrix are always computed. Unlike the cauchit/probit/cloglog families' ⁠_with_var⁠ variants, this function does not add any extra degenerate-case guarding beyond what fast_ordinal_regression_cpp already does.

Usage

fast_ordinal_regression_with_var_cpp(
  X,
  y,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  maxit = 100L,
  tol = 1e-06,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see fast_ordinal_regression_cpp).

y

A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding.

warm_start_params

Optional starting values for [\alpha, \beta]. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided.

maxit

Maximum number of optimizer iterations.

tol

Convergence tolerance.

fixed_idx

Optional 1-indexed positions (into [\alpha,\beta]) of parameters to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

optimization_alg

Optimization algorithm (default "lbfgs").

warm_start_fisher_info

Optional initial curvature (Fisher/observed information) matrix.

Value

A list with components b, alpha, params, n_params, converged, iterations, neg_loglik, observed_information/fisher_information/information (all the same observed-information matrix), information_type ("observed"), ssq_b_j (the variance of b[1], i.e. the coefficient on X's first column — conventionally the treatment effect), and vcov — the latter two are NA/omitted if the free-parameter information matrix is not invertible.

Fixed parameters, warm starts, and optimization

fixed_idx (1-indexed into the combined [\alpha, \beta] parameter vector, thresholds first) and fixed_values optionally hold a subset of parameters at caller-supplied constant values rather than estimating them. warm_start_params supplies starting values for [\alpha, \beta] directly (skipping smart_cold_start); otherwise thresholds always start evenly spaced on (-1, 1) at -1 + 2(k+1)/K, and \beta starts at either zero, or (when smart_cold_start = TRUE, the default) an OLS fit of the rank-rescaled response (y - 1)/(K - 1) on X — falling back silently to zero if that OLS solve is not well-posed. When smart_cold_start = TRUE and no warm_start_fisher_info is supplied, the Hessian at the starting values is additionally used to seed the optimizer's first-iteration curvature estimate. Optimization runs via optimization_alg (default "lbfgs") for up to maxit iterations or until the parameter/gradient change falls below tol.

See Also

fast_ordinal_regression_cpp for the estimate-only-capable variant and the full model documentation; fast_ordinal_regression_weighted_cpp for the observation-weighted variant.


Fast Poisson Regression, Estimate-Only (C++ Backend)

Description

Fits a Poisson regression with the canonical log link, Y_i \sim \mathrm{Poisson}(\mu_i), \mu_i = \exp(x_i^\top\beta) (with \eta_i = x_i^\top\beta clamped above at 700 before exponentiating, to avoid overflow), by maximum likelihood. By default (optimization_alg = "irls"), fitting uses iteratively reweighted least squares: at each iteration the Fisher-scoring (Poisson canonical-link, so Fisher = observed) system X^\top W X \, \delta = X^\top(y - \mu) is solved via Eigen::LDLT, with a backtracking step-halving line search (up to 10 halvings) accepting the step only if it does not increase the negative log-likelihood; convergence is declared when the score norm falls below tol. Passing optimization_alg = "lbfgs" or "newton_raphson" instead routes through the generic likelihood optimizer (.normalize_optimizer_algorithm) on the raw (non-IRLS) negative log-likelihood/gradient/Hessian.

Usage

fast_poisson_regression_cpp(
  X,
  y,
  warm_start_beta = NULL,
  smart_cold_start = FALSE,
  maxit = 100L,
  tol = 1e-8,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "irls",
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the response variable, expected to be nonnegative-integer counts.

warm_start_beta

Optional starting values for coefficients \beta. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE and no warm_start_beta is supplied, use a Poisson-specific heuristic initial guess rather than a zero cold start.

maxit

Maximum number of iterations. Defaults to 100.

tol

Convergence tolerance. Defaults to 1e-8.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm: "irls" (default), "lbfgs", or "newton_raphson".

warm_start_weights

Accepted but unused; see Details.

warm_start_fisher_info

Optional initial curvature (information) matrix to warm-start the first iteration.

estimate_only

If TRUE, skip computing mu, XtWX/fisher_information, and score, returning only b, converged, num_iter, hit_iteration_cap, and gradient_norm.

Value

A list containing the following components:

b

A numeric vector of the estimated Poisson regression coefficients \hat\beta.

mu

(omitted if estimate_only = TRUE) The fitted means \hat\mu_i.

XtWX, fisher_information

(omitted if estimate_only = TRUE) Two aliases for the same curvature matrix X^\top W X (W = \mathrm{diag}(\hat\mu_i)) at the fitted coefficients — the Fisher information, which for the canonical log link coincides with the observed information.

score

(omitted if estimate_only = TRUE) The score vector X^\top(y - \hat\mu) at the fitted coefficients.

neg_ll

(omitted if estimate_only = TRUE) The negative log-likelihood at the fitted coefficients.

converged

A logical value indicating whether the final gradient norm was below tol (gradient_norm < tol); uniform across the "irls"/"lbfgs"/"newton_raphson" optimizers.

num_iter

The number of iterations performed.

hit_iteration_cap

A logical value, mutually exclusive with converged: TRUE iff the optimizer exhausted maxit iterations without meeting the gradient-norm convergence criterion.

gradient_norm

The norm of the score vector at the returned coefficients.

Fixed parameters, warm starts

fixed_idx and fixed_values optionally hold a subset of coefficients fixed at caller-supplied constant values (their contribution is folded into the linear predictor as an offset) rather than estimated. warm_start_beta supplies starting coefficients directly; otherwise, if smart_cold_start = TRUE, a Poisson-specific heuristic start is used, and if warm_start_fisher_info is also supplied (or, absent that, when smart_cold_start = TRUE), it seeds the curvature estimate used for the very first IRLS step (or first quasi-Newton step, for the non-IRLS algorithms). warm_start_weights is accepted for interface parity with sibling functions but is not consulted anywhere in this function's fitting logic.

See Also

fast_poisson_regression_weighted_cpp for the observation-weighted variant; fast_poisson_regression_with_var_cpp for the variance-computing variant; fast_quasipoisson_regression_with_var_cpp for the overdispersion-corrected variant.


Fast Weighted Poisson Regression (C++ Backend)

Description

Fits the same Poisson log-link model as fast_poisson_regression_cpp (see that page for the full model and optimizer contract), with each observation's contribution to the log-likelihood, score, and IRLS working weights multiplied by a row weight weights[i]. Setting all weights to 1 recovers fast_poisson_regression_cpp exactly. Always fits with estimate_only = FALSE (there is no flag to skip the post-fit mu/ XtWX/score computation for this variant).

Usage

fast_poisson_regression_weighted_cpp(
  X,
  y,
  weights,
  warm_start_beta = NULL,
  smart_cold_start = FALSE,
  maxit = 100L,
  tol = 1e-8,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "irls",
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the response variable, expected to be nonnegative-integer counts.

weights

A numeric vector of nonnegative weights, one per observation.

warm_start_beta

Optional starting values for coefficients \beta. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE and no warm_start_beta is supplied, use a Poisson-specific heuristic initial guess rather than a zero cold start.

maxit

Maximum number of iterations. Defaults to 100.

tol

Convergence tolerance. Defaults to 1e-8.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm: "irls" (default), "lbfgs", or "newton_raphson".

warm_start_weights

Accepted but unused; see fast_poisson_regression_cpp Details.

warm_start_fisher_info

Optional initial curvature (information) matrix to warm-start the first iteration.

Value

A list containing the following components:

b

A numeric vector of the estimated Poisson regression coefficients \hat\beta.

mu

The fitted means \hat\mu_i.

XtWX, fisher_information

Two aliases for the same (weighted) curvature matrix X^\top W X, W = \mathrm{diag}(\code{weights}_i \hat\mu_i), at the fitted coefficients.

score

The weighted score vector at the fitted coefficients.

neg_ll

The weighted negative log-likelihood at the fitted coefficients.

converged

A logical value indicating whether the algorithm converged.

iterations

The number of iterations performed.

gradient_norm

The norm of the score vector at convergence.

See Also

fast_poisson_regression_cpp for the unweighted variant and full model documentation.


Fast Poisson Regression with Variance Calculation (C++ Backend)

Description

Fits the same Poisson log-link model as fast_poisson_regression_cpp (see that page for the full model and optimizer contract; always with estimate_only = FALSE), and additionally inverts the fitted Fisher information matrix (Eigen::LDLT on the free-coefficient submatrix, via compute_diagonal_inverse_entry) to report the variance of two coefficients.

Usage

fast_poisson_regression_with_var_cpp(
  X,
  y,
  j = 2L,
  warm_start_beta = NULL,
  smart_cold_start = FALSE,
  maxit = 100L,
  tol = 1e-8,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "irls",
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the response variable, expected to be nonnegative-integer counts.

j

The 1-indexed coefficient whose variance to compute in ssq_b_j. Defaults to 2.

warm_start_beta

Optional starting values for coefficients \beta. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE and no warm_start_beta is supplied, use a Poisson-specific heuristic initial guess rather than a zero cold start.

maxit

Maximum number of iterations. Defaults to 100.

tol

Convergence tolerance. Defaults to 1e-8.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm: "irls" (default), "lbfgs", or "newton_raphson".

warm_start_weights

Accepted but unused; see fast_poisson_regression_cpp Details.

warm_start_fisher_info

Optional initial curvature (information) matrix to warm-start the first iteration.

Value

A list containing the following components:

b, params

A numeric vector of the estimated Poisson regression coefficients \hat\beta (both aliases of the same vector).

ssq_b_j

The variance of the j-th coefficient estimator; NA if j indexes a fixed coefficient.

ssq_b_2

The variance of the second coefficient estimator specifically (typically the treatment effect), regardless of what j is; equal to ssq_b_j when j = 2. NA if the second column is a fixed coefficient.

mu

The fitted means \hat\mu_i.

converged

A logical value indicating whether the final gradient norm was below tol (gradient_norm < tol); uniform across the "irls"/"lbfgs"/"newton_raphson" optimizers.

num_iter

The number of iterations performed.

score

The score vector at the fitted coefficients.

observed_information, fisher_information, information

Three aliases for the same X^\top W X curvature matrix (information_type is always "fisher").

hessian

The negative of that same matrix (the actual Hessian of the log-likelihood).

neg_loglik, neg_ll

The negative log-likelihood at the fitted coefficients (two aliases).

loglik

The log-likelihood (-neg_ll), or NA if neg_ll is non-finite.

hit_iteration_cap

A logical value, mutually exclusive with converged: TRUE iff the optimizer exhausted maxit iterations without meeting the gradient-norm convergence criterion.

gradient_norm

The norm of the score vector at the returned coefficients.

See Also

fast_poisson_regression_cpp for the estimate-only variant and full model documentation; fast_quasipoisson_regression_with_var_cpp for the overdispersion-corrected variant.


Fast Probit Regression, Estimate-Capable (C++)

Description

Fits binary probit regression, Y_i \sim \mathrm{Bernoulli}(\Phi(\eta_i)), \eta_i = x_i^\top\beta, by maximum likelihood, using a numerically stable log-scale evaluation of \Phi (log_pnorm_lower/log_pnorm_upper, matching pnorm's log.p = TRUE for |\eta| < 6 via a erfc-based identity, and falling back to a wider-range series approximation beyond that). By default (optimization_alg = "irls"), fitting uses iteratively reweighted least squares with working weights w_i = \phi(\eta_i)^2 / \max(\Phi(\eta_i)(1-\Phi(\eta_i)), 10^{-15}) (the standard probit Fisher-scoring weight) and generalized residual r_i = y_i\,\phi(\eta_i)/\Phi(\eta_i) - (1-y_i)\,\phi(\eta_i)/(1-\Phi(\eta_i)): each iteration solves X^\top W X\, \delta = X^\top r via Eigen::LDLT and takes the full Newton step (no step-halving line search), declaring convergence when either the score norm or the step norm falls below tol. Any optimization_alg value other than "lbfgs" runs this IRLS path; optimization_alg = "lbfgs" instead minimizes the exact negative log-likelihood directly via a bespoke L-BFGS driver with backtracking strong-Wolfe line search (mirroring RcppNumerical's optim_lbfgs defaults), bypassing IRLS entirely — in that path, warm_start_fisher_info is not consulted.

Usage

fast_probit_regression_cpp(
  X,
  y,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  maxit = 100L,
  tol = 1e-08,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "irls",
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors (including an intercept column, if desired).

y

A numeric vector of binary responses (0/1).

warm_start_beta

Optional starting values for coefficients \beta. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE (the default) and no warm_start_beta is supplied, use an OLS-based initial guess (see Details).

maxit

Maximum number of iterations (IRLS path only).

tol

Convergence tolerance.

fixed_idx

Optional indices of fixed parameters.

fixed_values

Optional values for fixed parameters.

optimization_alg

Optimization algorithm: any value other than "lbfgs" runs IRLS (default "irls"); "lbfgs" runs direct likelihood minimization.

warm_start_weights

Accepted but unused; see Details.

warm_start_fisher_info

Optional initial curvature matrix for the first IRLS iteration (IRLS path only).

estimate_only

If TRUE, skip the post-fit weight/curvature computation and return only b, converged, and iterations.

Value

If estimate_only = TRUE: a list with b, converged, iterations. Otherwise: a list additionally containing w (the final probit IRLS working weights, evaluated at the fitted \hat\beta regardless of which optimization_alg was used), fisher_information (X^\top W X at those weights), score, and neg_ll (the negative log-likelihood).

Fixed parameters, warm starts

fixed_idx and fixed_values optionally hold a subset of coefficients fixed at caller-supplied constant values (folded into the linear predictor as an offset) rather than estimated. warm_start_beta supplies starting coefficients directly; otherwise, if smart_cold_start = TRUE (the default), an OLS fit of the probit-transformed response \Phi^{-1}((y+0.5)/2) on X seeds the start. warm_start_fisher_info, if supplied, seeds the curvature matrix used for the IRLS path's very first iteration only (see above for the "lbfgs" exception). warm_start_weights is accepted for interface parity with sibling functions but is not consulted anywhere in this function's fitting logic.

See Also

fast_probit_regression_weighted_cpp() for the observation-weighted variant; fast_probit_regression_with_var_cpp for the variance-computing variant.


Export of C++ function fast_probit_regression_with_var_cpp

Description

Fits the same probit model as fast_probit_regression_cpp (see that page for the full model and optimizer contract; always with estimate_only = FALSE, and maxit/tol hardcoded to 100/10^{-8} — no override arguments here), and additionally inverts the fitted Fisher information matrix (compute_diagonal_inverse_entry on the free-coefficient submatrix) to report the variance of two coefficients.

Usage

fast_probit_regression_with_var_cpp(
  X,
  y,
  j = 2L,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "irls",
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors (including an intercept column, if desired).

y

A numeric vector of binary responses (0/1).

j

The 1-indexed coefficient whose variance to compute in ssq_b_j. Defaults to 2.

warm_start_beta

Optional starting values for coefficients \beta. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE (the default) and no warm_start_beta is supplied, use an OLS-based initial guess; see fast_probit_regression_cpp Details.

fixed_idx

Optional indices of fixed parameters.

fixed_values

Optional values for fixed parameters.

optimization_alg

Optimization algorithm: any value other than "lbfgs" runs IRLS (default "irls"); "lbfgs" runs direct likelihood minimization; see fast_probit_regression_cpp Details.

warm_start_weights

Accepted but unused; see fast_probit_regression_cpp Details.

warm_start_fisher_info

Optional initial curvature matrix for the first IRLS iteration (IRLS path only).

Value

A list with components b, params (the fitted coefficients \hat\beta, two aliases of the same vector), ssq_b_j (the variance of the j-th coefficient, NA if j indexes a fixed coefficient), ssq_b_2 (the variance of the second coefficient specifically, regardless of j; NA if the second column is fixed), score, observed_information / fisher_information / information (three aliases for the same X^\top W X curvature matrix; information_type is always "fisher"), hessian (the negative of that same matrix), neg_loglik/neg_ll (two aliases for the negative log-likelihood), loglik (-neg_ll, or NA if non-finite), converged, and iterations.

See Also

fast_probit_regression_cpp for the estimate-only-capable variant and full model documentation.


Fast Standard Normal Quantile Function, Vectorized (C++ Backend)

Description

Computes \Phi^{-1}(p), the standard normal quantile (inverse CDF), elementwise over p, via Peter Acklam's rational (minimax) approximation, accurate to within roughly 1.2 \times 10^{-9} relative error over the representable range of p — faster than base R's qnorm while matching it to that approximation precision. Used as the cold-start heuristic in several of this package's ordinal- and binary-response regression fitters (e.g. probit-family threshold initialization) wherever an approximate normal quantile is needed on a hot path, and exported standalone for the same reason as fast_digamma_vec_cpp and friends: to let performance-sensitive R or Python callers bypass qnorm's per-call dispatch overhead when evaluating many quantiles at once. Benchmarked at roughly 2.33x the speed of base R's vectorized qnorm() on this package's benchmark suite; see the "Utility / Math Kernel Performance" section of the benchmark report for the full methodology and current measured multiple.

Usage

fast_qnorm_vec_cpp(p)

Arguments

p

Numeric vector of probabilities in ⁠(0, 1)⁠; behavior at exactly 0, 1, or outside that range follows the underlying Acklam approximation's boundary handling, not necessarily -Inf/Inf/NaN exactly as base R's qnorm would return.

Value

A numeric vector of standard normal quantiles, the same length as p.

See Also

fast_log_pnorm_vec_cpp and fast_log_dnorm_vec_cpp for the corresponding forward (CDF/density) kernels.


Fast Quasi-Poisson Regression with Variance Calculation (C++ Backend)

Description

Fits the same Poisson log-link mean model as fast_poisson_regression_cpp (see that page for the full model and optimizer contract; point estimates \hat\beta are identical to what that function would return), but instead of the plain Poisson Fisher-information-based variance, scales it by an estimated overdispersion parameter to obtain quasi-likelihood standard errors that are robust to variance-mean deviations from the strict Poisson assumption \mathrm{Var}(Y_i) = \mu_i. The dispersion is the Pearson-statistic-based moment estimator,

\hat\phi = \frac{1}{n-p}\sum_{i=1}^n \frac{(y_i - \hat\mu_i)^2}{\hat\mu_i},

computed only when the residual degrees of freedom n - p > 0; the reported coefficient variances are then \widehat{\mathrm{Var}}(\hat\beta_k) = \hat\phi \, [(X^\top \hat W X)^{-1}]_{kk} (the ordinary Poisson Fisher information scaled by \hat\phi), matching the standard quasi-Poisson GLM correction (as in stats::glm(family = quasipoisson())). If n \le p, or \hat\phi comes out non-finite or non-positive, ssq_b_j/ssq_b_2/dispersion are left at their NA defaults (point estimates b and mu are still returned).

Usage

fast_quasipoisson_regression_with_var_cpp(
  X,
  y,
  j = 2L,
  warm_start_beta = NULL,
  smart_cold_start = FALSE,
  maxit = 100L,
  tol = 1e-8,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "irls",
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

y

A numeric vector of the response variable, expected to be nonnegative-integer counts.

j

The 1-indexed coefficient whose variance to compute in ssq_b_j. Defaults to 2.

warm_start_beta

Optional starting values for coefficients \beta. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE and no warm_start_beta is supplied, use a Poisson-specific heuristic initial guess rather than a zero cold start.

maxit

Maximum number of iterations. Defaults to 100.

tol

Convergence tolerance. Defaults to 1e-8.

fixed_idx

Optional integer indices of coefficients to hold fixed rather than estimate.

fixed_values

Optional values to fix the parameters named by fixed_idx at.

optimization_alg

Optimization algorithm: "irls" (default), "lbfgs", or "newton_raphson".

warm_start_weights

Accepted but unused; see fast_poisson_regression_cpp Details.

warm_start_fisher_info

Optional initial curvature (information) matrix to warm-start the first iteration.

Value

A list containing the following components:

b

A numeric vector of the estimated Poisson regression coefficients \hat\beta (point estimates, unaffected by the dispersion correction).

ssq_b_j

The dispersion-scaled variance of the j-th coefficient estimator; NA if j indexes a fixed coefficient or the dispersion estimate is unusable.

ssq_b_2

The dispersion-scaled variance of the second coefficient estimator specifically (typically the treatment effect), regardless of what j is; equal to ssq_b_j when j = 2.

dispersion

The estimated Pearson-based overdispersion parameter \hat\phi, or NA if the residual degrees of freedom are not positive.

mu

The fitted means \hat\mu_i.

converged

A logical value indicating whether the algorithm converged.

iterations

The number of iterations performed.

gradient_norm

The norm of the score vector at convergence.

See Also

fast_poisson_regression_with_var_cpp for the plain (non-overdispersion-corrected) variance variant; fast_poisson_regression_cpp for the estimate-only variant and full mean-model documentation.


Fast Robust (M/MM-Estimator) Linear Regression (C++)

Description

Fits a robust linear regression by iteratively reweighted least squares (IRLS), minimizing \sum_i \rho(r_i / \hat\sigma) for residuals r_i = y_i - x_i^\top\beta and a fixed robustness scale \hat\sigma, rather than ordinary least squares' \sum_i r_i^2. method = "M" uses Huber's weight function w(u) = 1 for |u| \le c and w(u) = c/|u| otherwise (c, default 1.345, tuned for 95% efficiency under normality); any other value of method (including the default, "MM") uses Tukey's bisquare weight w(u) = (1 - (u/c_b)^2)^2 for |u| \le c_b and 0 otherwise, with c_b = 4.685 hardcoded (not settable through this exported wrapper, though the internal fitter accepts it). The scale \hat\sigma is fixed once at the start as the normalized median absolute deviation of the OLS residuals, \hat\sigma = \mathrm{median}(|r_i|) / 0.6745 (the internal fitter also accepts a caller-supplied fixed scale, but this wrapper always estimates it). Each IRLS iteration re-weights and re-solves the weighted normal equations X^\top W X\, \beta = X^\top W y via Eigen::LDLT; convergence is declared when the relative change in \beta falls below tol. Column 1 of a real design typically holds the intercept, but no columns are treated specially except via fixed_idx.

Usage

fast_robust_regression_cpp(
  X,
  y,
  warm_start_beta = NULL,
  smart_cold_start = TRUE,
  method = "MM",
  j = 2L,
  c = 1.345,
  maxit = 50L,
  tol = 1e-07,
  fixed_idx = NULL,
  fixed_values = NULL,
  warm_start_weights = NULL,
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors.

y

A numeric vector of responses.

warm_start_beta

Optional starting values for coefficients. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE (the default) and no warm_start_beta is supplied, use an OLS (QR) initial guess; see Details.

method

Robust estimation method: "M" for Huber weighting, anything else (default "MM") for Tukey bisquare weighting; see Details.

j

1-based index of the coefficient whose asymptotic variance to return in ssq_b_j.

c

Huber tuning constant (default 1.345; only used when method = "M").

maxit

Maximum number of IRLS iterations.

tol

Relative parameter-change convergence tolerance.

fixed_idx

Optional indices of fixed parameters.

fixed_values

Optional values for fixed parameters.

warm_start_weights

Optional initial working weights for the first IRLS iteration.

warm_start_fisher_info

Optional initial curvature (X^\top W X) matrix for the first IRLS iteration.

estimate_only

If TRUE, skip the post-fit asymptotic-variance computation and return only coefficients, scale, converged, and iterations.

Value

If estimate_only = TRUE: a list with coefficients, scale (the fixed MAD-based robustness scale \hat\sigma), converged, iterations. Otherwise, additionally: ssq_b_j and fisher_information (the final IRLS X^\top W X curvature matrix). ssq_b_j is computed only if the fit converged or ran the full maxit iterations, as the standard M-estimator asymptotic variance \widehat{\mathrm{Var}}(\hat\beta_j) = \left(\frac{n}{n-p}\right) \frac{\sum_i \psi(r_i)^2}{n\,\bar\psi'^2} \, [(X^\top X)^{-1}]_{jj}, where \psi is the derivative of \rho (i.e. \psi(r) = w(r/\hat\sigma)\,r) and \bar\psi' is the mean of \psi' across observations, matching the classical Huber (1981) sandwich-free M-estimator variance formula; NA if j indexes a fixed coefficient or the fit neither converged nor exhausted maxit.

Fixed parameters, warm starts

fixed_idx and fixed_values optionally hold a subset of coefficients fixed at caller-supplied constant values (subtracted out of y as an offset) rather than estimated. warm_start_beta supplies a starting coefficient vector directly; otherwise, if smart_cold_start = TRUE (the default), an ordinary QR least-squares fit seeds the start (and, when variance will later be requested via j, also caches the QR-based [(X^\top X)^{-1}]_{jj} entry for reuse in the variance formula below). warm_start_weights seeds the IRLS weights for the first iteration only (skipping that iteration's Huber/bisquare weight computation); warm_start_fisher_info similarly seeds the first iteration's X^\top W X curvature matrix.


Fast Stereotype (Reduced-Rank Multinomial) Logistic Regression (C++)

Description

Fits Anderson's stereotype logit model — a reduced-rank multinomial logit for a categorical (nominal or ordinal) response with K distinct observed levels, using a single linear predictor \eta_i = x_i^\top\beta scaled by a category-specific "score" \phi_k \in [0, 1]:

\Pr(Y_i = k \mid x_i) = \frac{\exp(\alpha_k + \phi_k \eta_i)} {\sum_{l=1}^K \exp(\alpha_l + \phi_l \eta_i)}, \qquad \alpha_1 := 0,\ \phi_1 := 0,\ \phi_K := 1,

with free intercepts \alpha_2, \ldots, \alpha_K and free interior scores \phi_2, \ldots, \phi_{K-1} reparameterized via unconstrained \gamma_1, \ldots, \gamma_{K-2} as cumulative softmax-style partial sums, \phi_{j} = \left(\sum_{r \le j-2} e^{\gamma_r}\right) \big/ \left(1 + \sum_r e^{\gamma_r}\right) for j = 2, \ldots, K-1, which guarantees 0 = \phi_1 \le \phi_2 \le \cdots \le \phi_{K-1} \le \phi_K = 1 without an explicit constraint. A single \hat\beta therefore governs the covariate effect for every category, with the fitted \hat\phi_k determining how much of that effect applies to category k — collapsing categories with similar fitted scores are "stereotyped" together, which is the model's namesake use case (a parsimony-inducing alternative to full multinomial or ordinal cumulative-link models when categories are not clearly ordered but the covariate effect is plausibly one-dimensional). K = 2 reduces exactly to ordinary binary logistic regression (\phi_2 = 1 by construction, no \gamma parameters). At least 2 distinct observed outcome categories are required; fewer throws an error. Fitting optimizes the joint parameter vector [\alpha_2, \ldots, \alpha_K, \beta, \gamma_1, \ldots, \gamma_{K-2}] via optimization_alg (default "newton_raphson"), using this model's analytic gradient and Hessian.

Usage

fast_stereotype_logit_cpp(
  X,
  y,
  maxit = 100L,
  tol = 1e-08,
  smart_cold_start = TRUE,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "newton_raphson",
  warm_start_fisher_info = NULL,
  warm_start_params = NULL,
  warm_start_beta = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors (no intercept column needed; the category intercepts alpha serve that role).

y

A numeric vector of categorical (nominal or ordinal) responses; only the set of distinct values matters (mapped to 1, \ldots, K by sorted rank), not their numeric coding or order.

maxit

Maximum number of optimizer iterations.

tol

Convergence tolerance.

smart_cold_start

Present for interface parity with sibling functions but currently has no effect: starting values always come from initialize_params() (log empirical category-count ratios for alpha, zero for beta/gamma) unless overridden by warm_start_params/warm_start_beta.

fixed_idx

Optional 1-indexed positions (into the joint parameter vector, alpha first) of parameters to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

optimization_alg

Optimization algorithm (default "newton_raphson").

warm_start_fisher_info

Optional initial curvature matrix for the first optimizer iteration.

warm_start_params

Optional starting values for the full joint parameter vector. Takes precedence over warm_start_beta.

warm_start_beta

Optional starting values for \beta alone (ignored if warm_start_params is supplied).

estimate_only

If TRUE, skip computing the observed-information matrix and return only point estimates.

Value

A list with components b (\hat\beta), alpha (the K-1 free intercepts), scores_raw (the raw \hat\gamma reparameterization parameters, length \max(0, K-2); recover \hat\phi from these via the cumulative-softmax formula above), params (the full fitted joint parameter vector), neg_loglik, converged, and, unless estimate_only = TRUE, fisher_information (the negative Hessian of the log-likelihood at the fitted parameters — despite the name, this is the observed, not expected, information).

See Also

fast_stereotype_logit_with_var_cpp for the variance-computing variant.


Fast Stereotype (Reduced-Rank Multinomial) Logistic Regression with Variance (C++)

Description

As fast_stereotype_logit_cpp (see that page for the full stereotype logit model), but always computes the observed information and the variance of \hat\beta_1 (the first, and typically only meaningfully identified, regression coefficient — conventionally the treatment effect). The primary variance estimate is [(-H)^{-1}]_{\beta_1\beta_1} from the observed information (negative Hessian) at the fitted parameters. If \beta_1 is not held fixed (via fixed_idx) and that entry comes out non-finite (e.g. the information matrix is singular), this function falls back to a profile-likelihood variance: it re-optimizes all nuisance parameters (everything except \beta_1) at \hat\beta_1, \hat\beta_1 \pm h (h = \max(10^{-4}, 10^{-3}(|\hat\beta_1| + 1))), takes the central second-difference of the resulting profile log-likelihood to approximate the profile information I(\hat\beta_1), and reports 1/I(\hat\beta_1) if that comes out finite and positive (otherwise the variance remains NA). vcov is never populated (always NULL/missing) — only the single \beta_1 variance is available from this function.

Usage

fast_stereotype_logit_with_var_cpp(
  X,
  y,
  maxit = 100L,
  tol = 1e-08,
  smart_cold_start = TRUE,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "newton_raphson",
  warm_start_fisher_info = NULL,
  warm_start_params = NULL,
  warm_start_beta = NULL,
  estimate_only = FALSE
)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see fast_stereotype_logit_cpp).

y

A numeric vector of categorical (nominal or ordinal) responses; only the set of distinct values matters, not their numeric coding or order.

maxit

Maximum number of optimizer iterations.

tol

Convergence tolerance.

smart_cold_start

Present for interface parity but currently has no effect; see fast_stereotype_logit_cpp Details.

fixed_idx

Optional 1-indexed positions (into the joint parameter vector, alpha first) of parameters to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

optimization_alg

Optimization algorithm (default "newton_raphson").

warm_start_fisher_info

Optional initial curvature matrix for the first optimizer iteration.

warm_start_params

Optional starting values for the full joint parameter vector. Takes precedence over warm_start_beta.

warm_start_beta

Optional starting values for \beta alone (ignored if warm_start_params is supplied).

estimate_only

Accepted for interface parity but ignored: this function always computes the observed information and ssq_b_1/ssq_b_j regardless of its value.

Value

A list with components b (\hat\beta), alpha (the K-1 free intercepts), params (the full fitted joint parameter vector), ssq_b_1 and ssq_b_j (identical aliases for the variance of \hat\beta_1, computed as described above; NA if unavailable), vcov (always missing/NULL), converged, and fisher_information (the observed information, i.e. negative Hessian, at the fitted parameters).

See Also

fast_stereotype_logit_cpp for the estimate-only-capable variant and the full model documentation.


Stereotype Logit Profile Log-Likelihood for a Fixed Treatment Coefficient (C++)

Description

Computes the profile log-likelihood of the fast_stereotype_logit_cpp stereotype logit model (see that page for the full model) at a caller-fixed value of \beta_1 (the first regression coefficient, conventionally the treatment effect): all other parameters (intercepts \alpha, the remaining \beta columns if p > 1, and the score reparameterization \gamma) are re-optimized by damped Newton's method (up to 50 iterations, gradient-norm tolerance 10^{-8}, backtracking step-halving line search — both hardcoded and not controlled by the maxit/tol arguments below, which are accepted but currently unused) to maximize the log-likelihood conditional on \beta_1, and the resulting maximized log-likelihood is returned. This is the building block used by fast_stereotype_logit_with_var_cpp's profile-likelihood variance fallback (finite-differencing this function's output in \beta_1), and is exported standalone for constructing profile-likelihood confidence intervals or diagnostic profile plots directly.

Usage

fast_stereotype_profile_loglik_cpp(
  X,
  y,
  beta_fixed,
  maxit = 100L,
  tol = 1e-08,
  warm_start_params = NULL,
  warm_start_beta = NULL
)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see fast_stereotype_logit_cpp).

y

A numeric vector of categorical (nominal or ordinal) responses; only the set of distinct values matters, not their numeric coding or order.

beta_fixed

The fixed value at which to profile \beta_1.

maxit

Accepted but currently unused; see Details.

tol

Accepted but currently unused; see Details.

warm_start_params

Optional starting values for the full joint parameter vector (used to initialize the nuisance-parameter optimization). Takes precedence over warm_start_beta.

warm_start_beta

Optional starting values for \beta alone (ignored if warm_start_params is supplied).

Value

The maximized profile log-likelihood at beta_fixed, a single number.

See Also

fast_stereotype_logit_with_var_cpp, whose profile-likelihood variance fallback calls this function three times per fallback invocation.


Fast Trigamma Function, Vectorized (C++ Backend)

Description

Computes \psi'(x), the trigamma function (the second derivative of \log\Gamma(x), i.e. the derivative of the digamma function fast_digamma_vec_cpp), elementwise over x, via an asymptotic series expansion combined with the recurrence relation \psi'(x) = \psi'(x+1) + 1/x^2 (shifting small arguments up into the expansion's accurate range before applying it) — faster than base R's trigamma while matching it to within the approximation's own precision. Used wherever this package's likelihood kernels need the variance of a log-Gamma-based sufficient statistic or a Fisher-information second derivative involving \log\Gamma (e.g. negative-binomial dispersion-parameter curvature), and exported standalone for the same reason as fast_digamma_vec_cpp and friends. Benchmarked at roughly 19.3x the speed of base R's vectorized trigamma() on this package's benchmark suite; see the "Utility / Math Kernel Performance" section of the benchmark report for the full methodology and current measured multiple.

Usage

fast_trigamma_vec_cpp(x)

Arguments

x

Numeric vector of arguments (per the trigamma function's domain, should not be a non-positive integer, where \psi' has poles; no domain validation is performed by this function).

Value

A numeric vector of \psi'(x) values, the same length as x.

References

Trigamma function for orientation. Analogous Python API: SciPy polygamma(1, x).

See Also

fast_digamma_vec_cpp for the corresponding first-derivative kernel this function's recurrence builds on.


Fast Weibull AFT Regression (R Wrapper: Rcpp Backend or survival)

Description

Fits the Weibull accelerated failure time model documented in full at fast_weibull_regression_general_cpp (\log T_i = \eta_i + \sigma W_i, \eta_i = x_i^\top\beta, right-censoring only), via either that C++ backend (use_rcpp = TRUE, the default) or survreg with dist = "weibull" (use_rcpp = FALSE) as a fallback/cross-check implementation.

Usage

fast_weibull_regression(
  y,
  dead,
  X,
  use_rcpp = TRUE,
  estimate_only = FALSE,
  optimization_alg = "lbfgs",
  warm_start_params = NULL,
  warm_start_fisher_info = NULL
)

Arguments

y

Observed survival/censoring times (must be positive).

dead

Event indicator: 1 for an exactly observed event, 0 for right-censored (survival known only to exceed y[i]).

X

A numeric matrix of predictor variables. It is assumed that an intercept column (e.g., a column of ones) is already included in X if desired.

use_rcpp

Logical. If TRUE (default), use the optimized Rcpp implementation. If FALSE, use survreg.

estimate_only

Logical. If TRUE, skip variance-covariance matrix calculation for speed. Only has an effect when use_rcpp = TRUE.

optimization_alg

Optimization algorithm: "lbfgs" (default) or "newton_raphson". Only has an effect when use_rcpp = TRUE.

warm_start_params

Optional starting values for [\beta, \log\sigma]. Only has an effect when use_rcpp = TRUE.

warm_start_fisher_info

Optional initial curvature (Fisher/observed information) matrix. Only has an effect when use_rcpp = TRUE.

Details

When use_rcpp = TRUE, an intercept column is prepended to X automatically if not already present (detected as a first column of all 1s), fitting always starts from a zero cold start (smart_cold_start = FALSE is hardcoded, regardless of whether warm_start_params is supplied), and a non-converged C++ fit is escalated to an R-level stop() rather than returned silently.

When use_rcpp = FALSE, any existing intercept column is stripped and survreg is left to add its own; remaining covariate columns are first passed through drop_linearly_dependent_cols to remove collinear columns before fitting (silently — no error or warning is raised for dropped columns). estimate_only, optimization_alg, warm_start_params, and warm_start_fisher_info have no effect on this path — survreg always computes the full variance-covariance matrix, and std_errs (from sqrt(diag(vcov))) is included only in this path's return value, not the Rcpp path's. Both non-finite coefficients and (unless estimate_only = TRUE, which is ignored on this path regardless) non-finite variance-covariance entries from survreg are escalated to an R-level stop().

Value

A list containing the following components:

coefficients

A numeric vector of the estimated Weibull regression coefficients \hat\beta, including the intercept.

log_sigma

The logarithm of the fitted scale parameter \hat\sigma of the Weibull AFT distribution.

vcov

The variance-covariance matrix of the estimated coefficients, or NULL if estimate_only = TRUE (Rcpp path only).

neg_log_lik

(Rcpp path only) The negative log-likelihood at the fitted parameters.

fisher_information

(Rcpp path only) The observed information matrix, or NULL if estimate_only = TRUE.

std_errs

(survival path only) The coefficient standard errors, sqrt(diag(vcov)).

See Also

fast_weibull_regression_general_cpp for the full model documentation and Rcpp backend contract.

Examples

X = matrix(rnorm(500), 100, 5)
y = runif(100)
dead = rbinom(100, 1, 0.5)
fast_weibull_regression(y, dead, X)

Fast Weibull AFT Regression (C++)

Description

Weibull Accelerated Failure Time model fitting, exact/ right-censored responses only. See fast_weibull_regression_general_cpp for the left-/interval-censored extension.

Usage

fast_weibull_regression_cpp(
  X,
  y,
  dead,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  estimate_only = FALSE,
  maxit = 100L,
  tol = 1e-08,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors.

y

A numeric vector of survival times.

dead

A numeric vector of event indicators (1=event, 0=censored).

warm_start_params

Optional starting values for coefficients.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess.

estimate_only

Logical. If TRUE, do not compute variance-covariance.

maxit

Maximum number of iterations.

tol

Convergence tolerance.

fixed_idx

Optional indices of fixed parameters.

fixed_values

Optional values for fixed parameters.

optimization_alg

Optimization algorithm.

warm_start_fisher_info

Optional initial Fisher Information matrix.

Value

A list containing coefficients, log_sigma, and convergence status.


Fast Weibull Regression with General Censoring (C++ Backend)

Description

Weibull Accelerated Failure Time model fitting extended to left-, right-, and interval-censored responses (TODO-3 in interval_censored_survival_response.md). Zero-regression by construction: exact/right-censored-only input uses the same likelihood contributions as the corresponding survival::Surv() response.

Usage

fast_weibull_regression_general_cpp(
  X,
  y,
  y_L,
  y_R,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  estimate_only = FALSE,
  maxit = 100L,
  tol = 1e-08,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL
)

Arguments

X

A numeric matrix of predictors.

y

Exact survival times, NA for censored subjects.

y_L

Censored-interval lower bounds, NA for exact subjects; 0 for left-censored.

y_R

Censored-interval upper bounds, NA for exact subjects; Inf for right-censored.

warm_start_params

Optional starting values for coefficients.

smart_cold_start

Logical. If TRUE, use an initial OLS-based guess.

estimate_only

Logical. If TRUE, do not compute variance-covariance.

maxit

Maximum number of iterations.

tol

Convergence tolerance.

fixed_idx

Optional indices of fixed parameters.

fixed_values

Optional values for fixed parameters.

optimization_alg

Optimization algorithm.

warm_start_fisher_info

Optional initial Fisher Information matrix.

Value

A list containing the following components:

coefficients

A numeric vector of the estimated Weibull regression coefficients, including the intercept.

log_sigma

The logarithm of the scale parameter from the Weibull distribution.

vcov

The variance-covariance matrix of the estimated coefficients.

neg_ll

The negative log-likelihood at the final iteration.

converged

A logical value indicating whether the algorithm converged.

A list containing coefficients, log_sigma, and convergence status.


Fast Zero-Inflated or Hurdle Poisson Regression (C++)

Description

Fits, by direct maximum likelihood, a two-component count model with a Poisson log-link count component (\lambda_i = \exp(x_i^\top\beta_{\mathrm{cond}}), X) and a logit-link binary component (\pi_i = \mathrm{logit}^{-1}(x_{\mathrm{zi},i}^\top\beta_{\mathrm{zi}}), Xzi). The two supported models differ in how \pi_i enters the likelihood:

Both branches share one likelihood/gradient/(analytic and expected) Hessian implementation, switched at each observation by is_hurdle. Optimizes the joint parameter vector [\beta_{\mathrm{cond}}, \beta_{\mathrm{zi}}] via optimization_alg (default "lbfgs"). If the optimizer throws an exception internally, this function does not propagate an R error: it returns a diagnostic list with converged = FALSE, evaluated at the optimizer's starting values and including the caught exception_message (see Value), so callers must check converged before using the estimates.

Usage

fast_zero_augmented_poisson_cpp(
  X,
  y,
  Xzi,
  is_hurdle,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  estimate_only = FALSE,
  maxit = 1000L,
  tol = 1e-08,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL
)

Arguments

X

Matrix of predictors for the conditional (Poisson count) component.

y

Vector of nonnegative-integer count responses.

Xzi

Matrix of predictors for the zero-inflation/hurdle (logistic) component.

is_hurdle

If TRUE, fit a hurdle model; if FALSE, fit a zero-inflated model. See Details for the distinction.

warm_start_params

Optional starting values for the joint [\beta_{\mathrm{cond}}, \beta_{\mathrm{zi}}] vector. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE (the default) and no warm_start_params is supplied, use a model-specific heuristic initial guess; see Details.

estimate_only

If TRUE, skip the post-fit Hessian/variance computation and return only params, converged, neg_ll/neg_loglik, and gradient_norm.

maxit

Maximum number of optimizer iterations.

tol

Convergence tolerance.

fixed_idx

Optional 1-indexed positions of parameters to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

optimization_alg

Optimization algorithm (default "lbfgs").

warm_start_fisher_info

Optional initial curvature (Fisher/observed information) matrix.

Value

On optimizer failure (an internal exception, never an R error), and regardless of estimate_only: a list with converged = FALSE, num_iter = 0, hit_iteration_cap = FALSE, params (the starting values the optimizer began from, after applying fixed_idx/fixed_values), neg_ll/neg_loglik, observed_information/fisher_information/information (all evaluated at those starting values), information_type (always "observed"), hessian, gradient_norm (NA if not finite), min_eigenvalue_information (NA), params_origin (a message that terminal parameters are unavailable after the exception) and exception_message (the caught exception text). Otherwise, if estimate_only = TRUE: a list with params (the joint fitted [\hat\beta_{\mathrm{cond}}, \hat\beta_{\mathrm{zi}}] vector), converged, neg_ll/neg_loglik (two aliases), and gradient_norm. Otherwise, additionally: vcov (the joint variance-covariance matrix), observed_information/fisher_information/ information (three aliases for the same observed-information matrix; information_type is always "observed"), hessian (the negative of that same matrix), and coefficients (a list with cond and zi sub-vectors splitting params back into its two components).

Fixed parameters, warm starts

fixed_idx (1-indexed into the joint parameter vector, X's coefficients first) and fixed_values optionally hold a subset of parameters fixed at caller-supplied constant values rather than estimated. warm_start_params supplies the full starting vector directly; otherwise, if smart_cold_start = TRUE (the default), a model-specific heuristic start is used, and if FALSE, all parameters start at zero except the first conditional-model coefficient, initialized to \log(\bar y) (if \bar y > 0). warm_start_fisher_info, if supplied, seeds the curvature estimate used for the optimizer's first iteration.


Fast Zero/One-Inflated Beta Regression (C++)

Description

Fits, by direct maximum likelihood, a three-component mixture model for a response Y_i \in [0, 1] (e.g. a bounded proportion with excess exact 0s and 1s): with probability \pi_{0,i} the response is exactly 0, with probability \pi_{1,i} it is exactly 1, and with probability \pi_{b,i} = 1 - \pi_{0,i} - \pi_{1,i} it falls strictly inside (0, 1) and follows a mean-precision Beta distribution. The three category probabilities come from a multinomial-logit-style softmax over two linear predictors on X_zero_one, with the "interior" (Beta) category as the implicit zero baseline:

\pi_{0,i} = \frac{e^{\eta_{0,i}}}{e^{\eta_{0,i}} + e^{\eta_{1,i}} + 1}, \quad \pi_{1,i} = \frac{e^{\eta_{1,i}}}{e^{\eta_{0,i}} + e^{\eta_{1,i}} + 1}, \quad \eta_{0,i} = x_{\mathrm{zo},i}^\top\gamma_0,\ \ \eta_{1,i} = x_{\mathrm{zo},i}^\top\gamma_1,

and, conditional on 0 < Y_i < 1, Y_i \sim \mathrm{Beta}(\mu_i\phi, (1-\mu_i)\phi) with \mu_i = \mathrm{logit}^{-1}(x_i^\top\beta) (clamped to [10^{-8}, 1-10^{-8}]) and precision \phi = e^{\log\phi} — matching the mean-precision Beta regression parameterization documented at fast_beta_regression_cpp. Optimizes the joint parameter vector [\beta, \log\phi, \gamma_0, \gamma_1] via optimization_alg (default "lbfgs"), for up to 1500 iterations at gradient-norm tolerance 10^{-6} — both hardcoded, with no maxit/tol arguments exposed by this function.

Usage

fast_zero_one_inflated_beta_cpp(
  X,
  X_zero_one,
  y,
  warm_start_params = NULL,
  smart_cold_start = TRUE,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

Matrix of predictors for the interior Beta (mean) component.

X_zero_one

Matrix of predictors for the zero- and one-inflation (mixture-probability) components; shared between \gamma_0 and \gamma_1.

y

Vector of responses in ⁠[0, 1]⁠.

warm_start_params

Optional starting values for the joint [\beta, \log\phi, \gamma_0, \gamma_1] vector. If provided, smart_cold_start is ignored.

smart_cold_start

Logical. If TRUE (the default) and no warm_start_params is supplied, use an OLS-based initial guess; see Details.

fixed_idx

Optional 1-indexed positions of parameters to hold fixed.

fixed_values

Optional values, parallel to fixed_idx, of the fixed parameters.

optimization_alg

Optimization algorithm (default "lbfgs").

warm_start_fisher_info

Optional initial curvature (Fisher/observed information) matrix.

estimate_only

If TRUE, skip the post-fit Hessian/variance computation and return only b, log_phi, zero_one_b0, zero_one_b1, params, neg_loglik, and converged.

Value

A list with components b (\hat\beta), log_phi, zero_one_b0/zero_one_b1 (\hat\gamma_0/\hat\gamma_1), params (the full fitted joint vector), neg_loglik, converged; unless estimate_only = TRUE, additionally vcov (the joint variance-covariance matrix — an all-NA matrix if the observed information is non-finite or its free-parameter submatrix is not invertible, rather than an error), observed_information/fisher_information/information (three aliases for the same observed-information matrix; information_type is always "observed"), and hessian (the negative of that same matrix).

Fixed parameters, warm starts

fixed_idx (1-indexed into [\beta, \log\phi, \gamma_0, \gamma_1]) and fixed_values optionally hold a subset of parameters fixed at caller-supplied constant values rather than estimated. warm_start_params supplies the full starting vector directly; otherwise, if smart_cold_start = TRUE (the default): \beta starts from an OLS fit of \mathrm{logit}(y_i) on X restricted to interior observations (0 < y_i < 1), falling back to zero if fewer such observations than ncol(X) are available; \log\phi starts at 2; and \gamma_0/ \gamma_1 each start from a separate OLS fit of X_zero_one on the corresponding 0/1 indicator (\mathbb{1}[y_i = 0], \mathbb{1}[y_i = 1]). If smart_cold_start = FALSE, all parameters start at zero except \log\phi, which still starts at 2. warm_start_fisher_info, if supplied, seeds the curvature estimate used for the optimizer's first iteration.


Fast Zero-Inflated Negative Binomial Regression (C++)

Description

High-performance zero-inflated negative binomial model fitting via L-BFGS.

Usage

fast_zinb_cpp(
  X,
  Xzi,
  y,
  warm_start_params = NULL,
  maxit = 1000L,
  tol = 1e-08,
  fixed_idx = NULL,
  fixed_values = NULL,
  optimization_alg = "lbfgs",
  smart_cold_start = TRUE,
  warm_start_fisher_info = NULL,
  estimate_only = FALSE
)

Arguments

X

Numeric matrix of predictors for the count component (including intercept).

Xzi

Numeric matrix of predictors for the zero-inflation component (including intercept).

y

Numeric vector of non-negative integer count responses.

warm_start_params

Optional starting values for all parameters.

maxit

Maximum number of iterations.

tol

Convergence tolerance.

fixed_idx

Optional indices of fixed parameters.

fixed_values

Optional values for fixed parameters.

optimization_alg

Optimization algorithm (default "lbfgs").

smart_cold_start

Logical. If TRUE, use a heuristic initial guess.

warm_start_fisher_info

Optional initial Fisher Information matrix.

estimate_only

Logical. If TRUE, skip variance computation and return only coefficients.

Value

A list containing coefficients and convergence status.


Fast G-Computation (Standardization) Point Estimate for a Logit-Link Model (C++)

Description

Computes the G-computation (regression standardization) point estimate of the marginal treatment effect under a fitted logit-link model — logistic regression for a binary outcome, or the algebraically identical fractional logit / quasi-binomial model for a proportion outcome in [0, 1], since the standardization formula depends only on the fitted linear predictor and link function, not the response's distributional assumptions. For each subject i in the fitted sample, the fitted linear predictor is decomposed into its treatment-free baseline \eta_{\mathrm{base},i} = x_i^\top\hat\beta - \hat\beta_{j_{\mathrm{treat}}}\, x_{i,j_{\mathrm{treat}}} and the two counterfactual predictions \widehat{\Pr}(Y_i = 1 \mid \mathrm{do}(T=1)) = \mathrm{logit}^{-1}(\eta_{\mathrm{base},i} + \hat\beta_{j_{\mathrm{treat}}}) and \widehat{\Pr}(Y_i = 1 \mid \mathrm{do}(T=0)) = \mathrm{logit}^{-1}(\eta_{\mathrm{base},i}) are computed by setting every subject's treatment column to 1 (respectively 0) while leaving all other covariates at their observed values — this is standard G-computation / standardization: average the model-implied outcome over the empirical covariate distribution under each counterfactual treatment assignment. The two averages (mean1, mean0) and their difference (md, the standardized average treatment effect on the risk-difference scale) are returned.

Usage

gcomp_fractional_logit_point_estimate_cpp(X_fit, coef_hat, j_treat)

Arguments

X_fit

Numeric matrix of predictors used to fit the model, including an intercept column if the model has one.

coef_hat

Numeric vector of fitted model coefficients \hat\beta, same length and column order as X_fit.

j_treat

1-based column index of the treatment indicator in X_fit.

Value

A list with elements mean1 (standardized mean outcome under T=1 for everyone), mean0 (standardized mean outcome under T=0 for everyone), and md (mean1 - mean0, the standardized risk difference).

See Also

gcomp_logistic_point_estimate_cpp, which computes the identical quantity (it delegates directly to this function) under the "logistic regression" framing.


Fast G-Computation (Standardization) Point Estimate for Logistic Regression (C++)

Description

Computes the standardized (G-computation) marginal risk difference under a fitted logistic regression model. This is a thin alias: it delegates directly to gcomp_fractional_logit_point_estimate_cpp (see that page for the full standardization formula and counterfactual-averaging methodology, which is identical for logistic and fractional-logit/quasi-binomial models), passing its arguments through unchanged.

Usage

gcomp_logistic_point_estimate_cpp(X_fit, coef_hat, j_treat)

Arguments

X_fit

Numeric matrix of predictors used to fit the model, including an intercept column if the model has one.

coef_hat

Numeric vector of fitted logistic regression coefficients \hat\beta, same length and column order as X_fit.

j_treat

1-based column index of the treatment indicator in X_fit.

Value

A list with elements mean1 (standardized mean risk under T=1 for everyone), mean0 (standardized mean risk under T=0 for everyone), and md (mean1 - mean0, the standardized risk difference).

See Also

gcomp_fractional_logit_point_estimate_cpp for the full documentation of the underlying computation.


Export of C++ function gcomp_logistic_post_fit_cpp

Description

Given an already-fitted logistic regression, computes a Huber-White/Eicker sandwich (heteroskedasticity-robust, "HC0") coefficient covariance matrix and, by the delta method, standard errors for the G-computed standardized risk difference and log risk ratio (see gcomp_logistic_point_estimate_cpp for the standardization point estimates this builds inference around). The sandwich covariance is \widehat{\mathrm{Var}}(\hat\beta) = B\,M\,B, with "bread" B = (X^\top W X)^{-1} (W = \mathrm{diag}(\hat\mu_i(1-\hat\mu_i)), the model-based Fisher information weights) and "meat" M = X^\top \mathrm{diag}((y_i - \hat\mu_i)^2) X (the empirical score outer-product, making this robust to model misspecification, not just relying on the working Bernoulli variance). The risk-difference standard error is obtained by propagating this sandwich covariance through the standardized risks' gradients with respect to \beta (\nabla_\beta \bar{\mathrm{risk}}_1 - \nabla_\beta \bar{\mathrm{risk}}_0, each a population-averaged logistic-derivative weighted design-matrix sum, with the treatment column's gradient entry replaced by the sum of the standardized-risk derivative directly since every subject's treatment indicator is held fixed at 1 or 0 in the counterfactual averages); the log risk ratio's standard error is obtained the same way via the gradient of \log(\overline{\mathrm{risk}}_1) - \log(\overline{\mathrm{risk}}_0), and is only computed (non-NA) when both standardized risks are strictly positive. Aborts with an R error (rather than returning NAs) if mu_hat contains non-finite or boundary (0 or 1) values, if the weighted design crossproduct X^\top W X is not invertible, if the sandwich covariance comes out non-finite anywhere, or if the treatment coefficient's variance is non-positive.

Usage

gcomp_logistic_post_fit_cpp(X_fit, y, coef_hat, mu_hat, j_treat)

Arguments

X_fit

Numeric matrix of predictors used to fit the model, including an intercept column if the model has one.

y

The observed binary (0/1) response used to fit the model.

coef_hat

Numeric vector of fitted logistic regression coefficients \hat\beta, same length and column order as X_fit.

mu_hat

Numeric vector of fitted probabilities \hat\mu_i = \mathrm{logit}^{-1}(x_i^\top\hat\beta), one per row of X_fit.

j_treat

1-based column index of the treatment indicator in X_fit.

Value

A list with components vcov (the p \times p sandwich covariance matrix), std_err and z_vals (per-coefficient standard errors and Wald z-statistics, NA for any coefficient with non-finite or non-positive variance), risk1/risk0 (the standardized mean risks under T=1/T=0 for everyone), rd (risk1 - risk0) and se_rd (its delta-method standard error), and log_rr/rr/ se_log_rr (the log risk ratio, risk ratio, and the log risk ratio's delta-method standard error — all NA if either standardized risk is not strictly positive).

See Also

gcomp_logistic_point_estimate_cpp for the point-estimate computation this function's variances are built around; gcomp_fractional_logit_post_fit_cpp() for the analogous fractional-logit/quasi-binomial post-fit inference.


Export of C++ function gcomp_ordinal_proportional_odds_post_fit_cpp

Description

Computes the standardized (G-computation) marginal difference in expected ordinal category score under a fitted proportional-odds (cumulative logit) model for a K-category ordinal outcome coded 1, \ldots, K: \mathrm{logit}\,\Pr(Y_i \le k \mid x_i) = \alpha_k - x_i^\top\beta, k = 1, \ldots, K-1. For each subject i, the fitted linear predictor is decomposed into its treatment-free baseline \eta_{\mathrm{base},i} = x_i^\top\hat\beta - \hat\beta_{j_{\mathrm{treat}}}\, x_{i,j_{\mathrm{treat}}} and the counterfactual predictors \eta_{1,i} = \eta_{\mathrm{base},i} + \hat\beta_{j_{\mathrm{treat}}} (treatment column set to 1 for everyone) and \eta_{0,i} = \eta_{\mathrm{base},i} (set to 0 for everyone). The subject-level expected category score under each counterfactual is recovered from the fitted cumulative probabilities via the identity E[Y_i] = \sum_{k=1}^K k\,\Pr(Y_i=k) = 1 + \sum_{k=1}^{K-1} \Pr(Y_i > k) = 1 + \sum_{k=1}^{K-1} \left(1 - \mathrm{logit}^{-1}(\hat\alpha_k - \eta_i)\right), and the two population-averaged expected scores (mean1, mean0) and their difference (md) are returned — the ordinal analogue of gcomp_logistic_point_estimate_cpp's risk difference. Unlike gcomp_logistic_post_fit_cpp, this function computes point estimates only; no sandwich covariance, standard errors, or inferential quantities are returned despite the ⁠_post_fit⁠ name.

Usage

gcomp_ordinal_proportional_odds_post_fit_cpp(
  X_fit,
  coef_hat,
  alpha_hat,
  j_treat
)

Arguments

X_fit

Numeric matrix of predictors used to fit the model, including an intercept column if the model has one (conventionally absorbed into the thresholds alpha_hat rather than X_fit for a proportional-odds model, but this function does not enforce that).

coef_hat

Numeric vector of fitted proportional-odds regression coefficients \hat\beta, same length and column order as X_fit.

alpha_hat

Numeric vector of the K-1 fitted, increasing category thresholds \hat\alpha_1, \ldots, \hat\alpha_{K-1}.

j_treat

1-based column index of the treatment indicator in X_fit.

Value

A list with elements mean1 (standardized expected category score under T=1 for everyone), mean0 (standardized expected category score under T=0 for everyone), and md (mean1 - mean0).

See Also

gcomp_logistic_point_estimate_cpp for the binary-outcome analogue; gcomp_logistic_post_fit_cpp for an example of the sandwich-variance inference this function does not provide.


Generate Synthetic Simulation Covariates and Continuous Response

Description

A helper function to generate synthetic covariates and a latent continuous response identical to the logic used within SimulationFramework. Covariates may be supplied directly via X_mat or drawn randomly via cov_draw_method; exactly one of the two must be non-NULL.

Usage

generate_covariate_dataset(
  n,
  p,
  cond_exp_func_model = c("linear", "nonlinear"),
  norm_sq_beta_vec = 1,
  X_mat = NULL,
  cov_draw_method = stats::rnorm,
  cov_draw_method_args = list(mean = 0, sd = 1)
)

Arguments

n

Integer. Sample size (number of rows).

p

Integer. Number of covariates (number of columns).

cond_exp_func_model

Character scalar. Either "linear" (latent response is a weighted linear combination of covariates) or "nonlinear" (Friedman 1991 function applied to the first five covariates; requires p >= 5).

norm_sq_beta_vec

Positive numeric scalar. The desired squared Euclidean norm of the coefficient vector, i.e. sum(beta^2). The coefficient vector (or the overall Friedman scale) is rescaled so that this quantity equals norm_sq_beta_vec. Default 1.

X_mat

Numeric matrix of dimensions n x p, or NULL (default). When supplied, this matrix is used directly as the covariate matrix and cov_draw_method must be NULL.

cov_draw_method

A function used to draw n * p i.i.d. covariate values, or NULL. The function must accept the number of draws as its first positional argument followed by any named arguments in cov_draw_method_args. Default stats::rnorm. Must be NULL when X_mat is supplied.

cov_draw_method_args

Named list of additional arguments forwarded to cov_draw_method beyond the sample-size first argument. Default is list(mean = 0, sd = 1).

Details

Two conditional-expectation models are supported, both rescaled so that the latent response y_cont has a controlled signal magnitude:

"linear"

y_i = x_i^\top \beta, with \beta a fixed evenly spaced sequence from 1 to -1 across the p covariates (seq(1, -1, length.out = p)) — i.e. the first covariate gets the strongest positive effect, the last the strongest negative effect, and (for even p) covariates near the middle get effects near zero — rescaled so that \|\beta\|_2^2 = \code{norm\_sq\_beta\_vec}.

"nonlinear"

The Friedman (1991) MARS benchmark function applied to the first five covariates (requires p \ge 5; covariates beyond the fifth do not enter the response at all),

y_i = c \left(10 \sin(\pi x_{i1} x_{i2}) + 20 (x_{i3} - 0.5)^2 + 10 x_{i4} + 5 x_{i5}\right),

with the overall scale c chosen so that c^2 \sum_k \beta_{\mathrm{friedman},k}^2 = \code{norm\_sq\_beta\_vec} for the nominal term coefficients \beta_{\mathrm{friedman}} = (10, 20, 10, 5) (the \sin term's coefficient is not part of this nominal vector, so the realized \|c \cdot (10,20,10,5)\|_2^2 may not exactly equal norm_sq_beta_vec once the \sin term's own variance is accounted for — norm_sq_beta_vec calibrates the linear-in-covariate part of the scale, not the full response variance). This function assumes each covariate lies roughly in [-1, 1], matching Friedman's original specification; covariates drawn or supplied on a very different scale will not reproduce the intended signal shape.

Value

A list with two elements: X (a data frame of covariates) and y_cont (a numeric vector of the latent continuous response).

References

Friedman, J. H. (1991). "Multivariate Adaptive Regression Splines." The Annals of Statistics, 19(1), 1-67, doi:10.1214/aos/1176347963, for the nonlinear benchmark function used when cond_exp_func_model = "nonlinear".

Examples

generate_covariate_dataset(n = 10, p = 5)

Generate Atkinson optimal-design randomization permutations

Description

See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.

Usage

generate_permutations_atkinson_cpp(X_sexp, n, p_raw, prob_T, nsim)

Arguments

X_sexp

Numeric matrix: the design's covariate model matrix for the first n subjects, in arrival order.

n

Number of subjects.

p_raw

Number of raw covariate columns (before model-matrix expansion); the first p_raw + 3 subjects are assigned by coin flip before Atkinson's rule applies.

prob_T

Probability of assignment to treatment.

nsim

Number of randomization draws (columns of the returned matrix) to generate.

Value

A list with w_mat, an n x nsim integer matrix whose columns are independent 0/1 treatment-assignment draws, and m_mat (always NULL).


Generate Bernoulli randomization permutations

Description

See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.

Usage

generate_permutations_bernoulli_cpp(n, nsim, prob_T)

Arguments

n

Number of subjects.

nsim

Number of randomization draws (columns of the returned matrix) to generate.

prob_T

Probability of assignment to treatment.

Value

A list with w_mat, an n x nsim integer matrix whose columns are independent 0/1 treatment-assignment draws, and m_mat (always NULL).


Generate blocked randomization permutations

Description

See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.

Usage

generate_permutations_blocking_cpp(n, nsim, prob_T, strata_indices)

Arguments

n

Number of subjects.

nsim

Number of randomization draws (columns of the returned matrix) to generate.

prob_T

Probability of assignment to treatment.

strata_indices

List of integer vectors, one per stratum, holding the (1-based) indices of the subjects in that stratum.

Value

A list with w_mat, an n x nsim integer matrix whose columns are independent 0/1 treatment-assignment draws, and m_mat (always NULL).


Generate cluster randomization permutations

Description

See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.

Usage

generate_permutations_cluster_cpp(n, nsim, prob_T, cluster_indices)

Arguments

n

Number of subjects.

nsim

Number of randomization draws (columns of the returned matrix) to generate.

prob_T

Probability of assignment to treatment.

cluster_indices

List of integer vectors, one per cluster, holding the (1-based) indices of the subjects in that cluster; each cluster is assigned to one arm as a whole.

Value

A list with w_mat, an n x nsim integer matrix whose columns are independent 0/1 treatment-assignment draws, and m_mat (always NULL).


Generate Efron biased-coin randomization permutations

Description

See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.

Usage

generate_permutations_efron_cpp(n, nsim, prob_T, weighted_coin_prob)

Arguments

n

Number of subjects.

nsim

Number of randomization draws (columns of the returned matrix) to generate.

prob_T

Probability of assignment to treatment.

weighted_coin_prob

Efron's biased-coin probability: the chance of assigning the currently under-represented arm.

Value

A list with w_mat, an n x nsim integer matrix whose columns are independent 0/1 treatment-assignment draws, and m_mat (always NULL).


Generate IBCRD randomization permutations

Description

See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.

Usage

generate_permutations_ibcrd_cpp(n, nsim, prob_T)

Arguments

n

Number of subjects.

nsim

Number of randomization draws (columns of the returned matrix) to generate.

prob_T

Probability of assignment to treatment.

Value

A list with w_mat, an n x nsim integer matrix whose columns are independent 0/1 treatment-assignment draws, and m_mat (always NULL).


Generate matched-pair randomization permutations

Description

Every generate_permutations_*_cpp function in this file is seeded from one R::unif_rand() draw into edi_rng::RRng (RNG.h), a portable re-implementation of R's own Mersenne-Twister generator – a given seed therefore produces identical draws in R and in any future binding (e.g. Python) using the same core and the same seed.

Usage

generate_permutations_matching_cpp(m_vec, nsim, prob_T)

Arguments

m_vec

Integer vector of match ids, one per subject: subjects sharing a positive id form a matched pair (randomized within the pair); 0 marks an unmatched (reservoir) subject, assigned by an independent coin flip.

nsim

Number of randomization draws (columns of the returned matrix) to generate.

prob_T

Probability of assignment to treatment.

Value

A list with w_mat, an n x nsim integer matrix whose columns are independent 0/1 treatment-assignment draws, and m_mat (always NULL).


Generate Pocock-Simon minimization randomization permutations

Description

See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.

Usage

generate_permutations_pocock_simon_cpp(
  x_levels_matrix,
  num_levels_total,
  weights,
  p_best,
  prob_T,
  nsim
)

Arguments

x_levels_matrix

Integer matrix with one row per subject and one column per stratification covariate; each entry is the subject's level for that covariate as a (1-based) index into 1..num_levels_total.

num_levels_total

Total number of levels across all stratification covariates.

weights

Numeric vector of per-covariate imbalance weights (one per column of x_levels_matrix).

p_best

Probability of assigning the arm that minimizes the weighted imbalance.

prob_T

Probability of assignment to treatment.

nsim

Number of randomization draws (columns of the returned matrix) to generate.

Value

A list with w_mat, an n x nsim integer matrix whose columns are independent 0/1 treatment-assignment draws, and m_mat (always NULL).


Generate stratified permuted-block randomization (SPBR) permutations

Description

See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.

Usage

generate_permutations_spbr_cpp(strata_keys, block_size, prob_T, nsim)

Arguments

strata_keys

Character vector giving each subject's stratum label, in arrival order.

block_size

Size of the permuted blocks used within each stratum.

prob_T

Probability of assignment to treatment.

nsim

Number of randomization draws (columns of the returned matrix) to generate.

Value

A list with w_mat, an n x nsim integer matrix whose columns are independent 0/1 treatment-assignment draws, and m_mat (always NULL).


Beta Regression Hessian, Standalone (C++)

Description

Computes the Hessian matrix (second derivatives with respect to [\beta, \log\phi]) of the log-likelihood of the mean-precision Beta regression model documented in full at fast_beta_regression_cpp, at arbitrary caller-supplied parameters params (not necessarily the MLE). Exported standalone — independent of any optimizer run — for direct numerical diagnostics (e.g. checking curvature or building a custom variance estimate at a specific parameter value) and for use by get_beta_regression_score_cpp's sibling relationship in optimizer/inference code that needs both quantities at the same point.

Usage

get_beta_regression_hessian_cpp(X, y, params)

Arguments

X

A numeric matrix of predictors, as used to fit the model.

y

A numeric vector of responses in ⁠(0, 1)⁠.

params

A numeric vector [\beta, \log\phi]: the mean-model coefficients followed by the log-precision parameter.

Value

The (p+1) \times (p+1) Hessian matrix of the log-likelihood (i.e. the negative of the observed information) at params.

See Also

get_beta_regression_score_cpp for the corresponding gradient at the same point; fast_beta_regression_cpp for the full mean-precision Beta regression model documentation.


Compute Beta Regression Score

Description

Calculates the score vector (gradient of the log-likelihood) for a beta regression model.

Usage

get_beta_regression_score_cpp(X, y, params)

Arguments

X

A numeric matrix of predictors.

y

A numeric vector of responses (in (0, 1)).

params

A numeric vector of parameters [beta, log_phi].

Value

A numeric vector representing the score.


Get the default bootstrap dispatch policy

Description

Returns EDI's built-in policy table, consulted by the internal (non-exported) dispatcher edi_bootstrap_dispatch_policy(), for choosing which bootstrap confidence-interval type — "bca" (bias-corrected and accelerated) or "percentile" — an inference class uses by default.

Usage

get_bootstrap_dispatch_policy()

Details

The dispatcher resolves a type for a given inference-class name (and, if available, the fitted inference object) in this precedence order, returning the first match:

  1. If the object's experimental design class matches a key of design_class_overrides (via is), and the inference class name matches one of that design's named regular-expression patterns, use the associated type.

  2. Otherwise, if the inference class name matches one of inference_class_overrides's named regular-expression patterns (checked in list order, first match wins), use the associated type.

  3. Otherwise, fall back to default_type ("bca").

Whichever type is resolved by that process, a final safety check applies: if the resolved type is "bca" and the fitted object reports (via its private jackknife_block_size_gt_one_unsupported() method) that BCa's required jackknife computation is unsupported for its current data (e.g. a block size greater than 1), the type is silently downgraded to "percentile" instead. This override table exists because BCa is the generally preferred default (it corrects for both bias and skewness in the bootstrap distribution), but is empirically unreliable or computationally unsupported for specific inference/ design class combinations — the "percentile" overrides listed here were added as those cases were identified, not derived from a general rule.

Value

A named list describing the default bootstrap type configuration, with components default_type (the fallback type, "bca"), inference_class_overrides (a named character vector: regular-expression pattern names to bootstrap-type values, matched against the inference class name), and design_class_overrides (a named list keyed by experimental design class name, each value itself a named character vector of pattern-to-type overrides scoped to that design).

See Also

get_parallel_dispatch_policy for the analogous policy controlling forced-serial dispatch; get_optimization_dispatch_policy for the analogous policy controlling default optimizer algorithm choice.

Examples

get_bootstrap_dispatch_policy()

Get the default cold-start dispatch policy

Description

Returns EDI's built-in policy table, consulted by the internal (non-exported) dispatcher edi_cold_start_dispatch_policy(), for the smart_cold_start default used by each inference class's C++ model-fitting backend. A TRUE entry means the solver initializes via an OLS (or otherwise model-appropriate heuristic) warm-up before iterating; FALSE means a plain zero-vector cold start. Benchmarks show the OLS warm-up is net-negative for logistic and Poisson IRLS at typical trial sizes (the one extra OLS solve costs more than the IRLS iterations it saves), so those families — along with several G-computation-based incidence/proportion inference classes — default to FALSE here.

Usage

get_cold_start_dispatch_policy()

Details

The dispatcher checks the inference class name against inference_class_overrides's named regular-expression patterns in list order, returning the associated logical value at the first match; if none match, it falls back to default (TRUE). Unlike get_bootstrap_dispatch_policy, there is no separate design-class-scoped override table here — only a single flat pattern list.

Value

A named list with default (logical, the fallback when no override pattern matches; TRUE in the built-in policy) and inference_class_overrides (a named logical vector: regular-expression pattern names to TRUE/FALSE values, matched against the inference class name). These TRUE/FALSE defaults are empirical performance judgments computed on the maintainer's machine, not correctness facts — the same heuristic can be net-positive or net-negative depending on your hardware's core count, cache sizes, and BLAS backend. Run tune_EDI_for_this_machine to re-measure this axis on your own machine and persist any better setting it finds.

See Also

get_bootstrap_dispatch_policy and get_optimization_dispatch_policy for the analogous policies controlling bootstrap CI type and default optimizer algorithm; set_cold_start_dispatch_policy to override this policy at runtime; tune_EDI_for_this_machine to re-benchmark it on your own hardware.

Examples

get_cold_start_dispatch_policy()

Combined Conditional-Poisson/Poisson Hessian, Standalone (C++)

Description

Computes the Hessian matrix of the log-likelihood of the combined KK matched-pair conditional-Poisson (conditional-Binomial) plus reservoir marginal-Poisson model documented in full at fast_cpoisson_combined_with_var_cpp, at arbitrary caller-supplied params_r (not necessarily the MLE). Internally reuses the same score-and-information computation as get_cpoisson_combined_score_cpp (a single shared routine computes both at once) and returns the negative of the resulting information matrix, i.e. the actual Hessian of the log-likelihood. Exported standalone — independent of any optimizer run — for direct numerical diagnostics at a specific parameter value.

Usage

get_cpoisson_combined_hessian_cpp(
  yT_v_r,
  n_k_v_r,
  X_diff_v_r,
  y_r_r,
  w_r_r,
  X_r_r,
  params_r
)

Arguments

yT_v_r

Treated-subject outcome count per matched pair.

n_k_v_r

Total (treated + control) outcome count per matched pair.

X_diff_v_r

Covariate differences (treated minus control) between the members of each matched pair.

y_r_r

Reservoir (unmatched) subjects' outcomes.

w_r_r

Reservoir subjects' treatment indicators.

X_r_r

Reservoir subjects' covariates.

params_r

A numeric vector of model parameters at which to evaluate the Hessian.

Value

The Hessian matrix of the log-likelihood (the negative of the information matrix) at params_r.

See Also

get_cpoisson_combined_score_cpp for the corresponding gradient at the same point; fast_cpoisson_combined_with_var_cpp for the full model documentation.


Combined Conditional-Poisson/Poisson Score, Standalone (C++)

Description

Computes the score vector (gradient of the log-likelihood) of the combined KK matched-pair conditional-Poisson (conditional-Binomial) plus reservoir marginal-Poisson model documented in full at fast_cpoisson_combined_with_var_cpp, at arbitrary caller-supplied params_r (not necessarily the MLE). Exported standalone — independent of any optimizer run — for direct numerical diagnostics (e.g. verifying convergence, or building a custom estimating-equation solver) at a specific parameter value.

Usage

get_cpoisson_combined_score_cpp(
  yT_v_r,
  n_k_v_r,
  X_diff_v_r,
  y_r_r,
  w_r_r,
  X_r_r,
  params_r
)

Arguments

yT_v_r

Treated-subject outcome count per matched pair.

n_k_v_r

Total (treated + control) outcome count per matched pair.

X_diff_v_r

Covariate differences (treated minus control) between the members of each matched pair.

y_r_r

Reservoir (unmatched) subjects' outcomes.

w_r_r

Reservoir subjects' treatment indicators.

X_r_r

Reservoir subjects' covariates.

params_r

A numeric vector of model parameters at which to evaluate the score.

Value

The score vector (gradient of the log-likelihood) at params_r.

See Also

get_cpoisson_combined_hessian_cpp for the corresponding Hessian at the same point; fast_cpoisson_combined_with_var_cpp for the full model documentation.


Effective capabilities for an inference class or instance

Description

Effective capabilities for an inference class or instance

Usage

get_effective_capabilities(name, des_obj = NULL, live_obj = NULL)

Arguments

name

Either a class name (character), or an already-constructed inference object. Passing an object additionally refines the static answer with that object's own live-checkable capability gates (see EDI_INFERENCE_LIVE_CAPABILITY_GATES) – some supports_*() private methods are conditional on constructor arguments (e.g. InferenceSurvivalCoxPHRegr's use_rcpp), which a class-name-only, cached lookup can never reflect. The name-only path's cost and cached result are unaffected by this – a live object is never written into EDI_INFERENCE_EFFECTIVE_CAPABILITIES_CACHE, since the answer for one specific instance must not be silently handed to every future name-only caller for that class.

des_obj

Optional design object. Capabilities the design rules out (see get_design_excluded_inference_capabilities()) are removed from the answer. If NULL and a live inference object is available, its own design is used.

live_obj

Optional: an already-constructed inference object to use for the live-gate refinement, separate from name. Use this (with name still a character string) when the correct registry key isn't simply class(live_obj)[1] – e.g. an external/test subclass whose own leaf class isn't registered, where the caller has already resolved name to the nearest registered ancestor (see Inference$capabilities()) and passing the object as name instead would silently look up the wrong (unregistered) key. Ignored if name is itself a non-character object (that overload already derives its own live_obj).


Identity-Link (Risk-Difference) Binomial Regression Hessian, Standalone (C++)

Description

Computes a numerical (central finite-difference 4-point stencil, step h = 10^{-4}) approximation of the Hessian matrix of the log-likelihood of the constrained identity-link binomial regression model documented in full at fast_identity_binomial_regression_cpp, at arbitrary caller-supplied beta (not necessarily the MLE) — not an analytic second derivative. Exported standalone — independent of any optimizer run — for direct numerical diagnostics at a specific parameter value.

Usage

get_identity_binomial_regression_hessian_cpp(X, y_r, beta)

Arguments

X

A numeric matrix of predictors.

y_r

A binary (0/1) numeric vector of responses.

beta

A numeric vector of coefficients \beta at which to evaluate the Hessian.

Value

The finite-difference-approximated Hessian matrix of the log-likelihood at beta.

See Also

get_identity_binomial_regression_score_cpp for the corresponding (also finite-difference) gradient at the same point; fast_identity_binomial_regression_cpp for the full model documentation, including the probability-boundary constraint this Hessian is evaluated without enforcing.


Identity-Link (Risk-Difference) Binomial Regression Score, Standalone (C++)

Description

Computes a numerical (central finite-difference, step h = 10^{-6}) approximation of the score vector (gradient of the log-likelihood) of the constrained identity-link binomial regression model documented in full at fast_identity_binomial_regression_cpp, at arbitrary caller-supplied beta (not necessarily the MLE) — not an analytic derivative. Exported standalone — independent of any optimizer run — for direct numerical diagnostics (e.g. verifying convergence, or cross-checking an analytic gradient elsewhere) at a specific parameter value.

Usage

get_identity_binomial_regression_score_cpp(X, y_r, beta)

Arguments

X

A numeric matrix of predictors.

y_r

A binary (0/1) numeric vector of responses.

beta

A numeric vector of coefficients \beta at which to evaluate the score.

Value

The finite-difference-approximated score vector at beta.

See Also

get_identity_binomial_regression_hessian_cpp for the corresponding (also finite-difference) Hessian at the same point; fast_identity_binomial_regression_cpp for the full model documentation.


Weighted Identity-Link (Risk-Difference) Binomial Regression Hessian, Standalone (C++)

Description

Computes the observation-weighted Hessian matrix of the weighted log-likelihood of the constrained identity-link binomial regression model documented in full at fast_identity_binomial_regression_cpp, at arbitrary caller-supplied beta (not necessarily the MLE), with each observation's contribution multiplied by weights_r[i], via a numerical (central finite-difference 4-point stencil, step h = 10^{-4}) approximation — not an analytic second derivative. Exported standalone — independent of any optimizer run — for direct numerical diagnostics at a specific parameter value.

Usage

get_identity_binomial_regression_weighted_hessian_cpp(X, y_r, weights_r, beta)

Arguments

X

A numeric matrix of predictors.

y_r

A binary (0/1) numeric vector of responses.

weights_r

A nonnegative numeric vector of observation weights.

beta

A numeric vector of coefficients \beta at which to evaluate the Hessian.

Value

The finite-difference-approximated weighted Hessian matrix at beta.

See Also

get_identity_binomial_regression_weighted_score_cpp for the corresponding weighted gradient at the same point; get_identity_binomial_regression_hessian_cpp for the unweighted version; fast_identity_binomial_regression_cpp for the full model documentation.


Weighted Identity-Link (Risk-Difference) Binomial Regression Score, Standalone (C++)

Description

Computes the observation-weighted score vector (gradient of the weighted log-likelihood) of the constrained identity-link binomial regression model documented in full at fast_identity_binomial_regression_cpp, at arbitrary caller-supplied beta (not necessarily the MLE), with each observation's contribution multiplied by weights_r[i], via a numerical (central finite-difference, step h = 10^{-6}) approximation — not an analytic derivative. Exported standalone — independent of any optimizer run — for direct numerical diagnostics at a specific parameter value.

Usage

get_identity_binomial_regression_weighted_score_cpp(X, y_r, weights_r, beta)

Arguments

X

A numeric matrix of predictors.

y_r

A binary (0/1) numeric vector of responses.

weights_r

A nonnegative numeric vector of observation weights.

beta

A numeric vector of coefficients \beta at which to evaluate the score.

Value

The finite-difference-approximated weighted score vector at beta.

See Also

get_identity_binomial_regression_weighted_hessian_cpp for the corresponding weighted Hessian at the same point; get_identity_binomial_regression_score_cpp for the unweighted version; fast_identity_binomial_regression_cpp for the full model documentation.


Show this machine's saved EDI tuning, if any

Description

Reads the per-user config file written by tune_EDI_for_this_machine and returns it as an EDILocalMachineTuning object (whose print method shows when and how it was produced, the hardware fingerprint it was measured on, and every policy deviation it stores). Does not apply anything – application happens inside tune_EDI_for_this_machine() itself and at package load.

Usage

get_local_EDI_optimization()

Value

Invisibly, the saved EDILocalMachineTuning object, or NULL (with a message) if no valid saved tuning exists.

See Also

tune_EDI_for_this_machine, clear_local_EDI_optimization.

Examples


get_local_EDI_optimization()


Log-Link (Relative-Risk) Binomial Regression Hessian, Standalone (C++)

Description

Computes a numerical (central finite-difference 4-point stencil, step h = 10^{-4}) approximation of the Hessian matrix of the log-likelihood of the constrained log-link binomial regression model documented in full at fast_log_binomial_regression_cpp, at arbitrary caller-supplied beta (not necessarily the MLE) — not an analytic second derivative. Exported standalone — independent of any optimizer run — for direct numerical diagnostics at a specific parameter value.

Usage

get_log_binomial_regression_hessian_cpp(X, y_r, beta)

Arguments

X

A numeric matrix of predictors.

y_r

A binary (0/1) numeric vector of responses.

beta

A numeric vector of coefficients \beta at which to evaluate the Hessian.

Value

The finite-difference-approximated Hessian matrix of the log-likelihood at beta.

See Also

get_log_binomial_regression_score_cpp for the corresponding (also finite-difference) gradient at the same point; fast_log_binomial_regression_cpp for the full model documentation, including the probability-boundary constraint this Hessian is evaluated without enforcing.


Log-Link (Relative-Risk) Binomial Regression Score, Standalone (C++)

Description

Computes a numerical (central finite-difference, step h = 10^{-6}) approximation of the score vector (gradient of the log-likelihood) of the constrained log-link binomial regression model documented in full at fast_log_binomial_regression_cpp, at arbitrary caller-supplied beta (not necessarily the MLE) — not an analytic derivative. Exported standalone — independent of any optimizer run — for direct numerical diagnostics (e.g. verifying convergence) at a specific parameter value.

Usage

get_log_binomial_regression_score_cpp(X, y_r, beta)

Arguments

X

A numeric matrix of predictors.

y_r

A binary (0/1) numeric vector of responses.

beta

A numeric vector of coefficients \beta at which to evaluate the score.

Value

The finite-difference-approximated score vector at beta.

See Also

get_log_binomial_regression_hessian_cpp for the corresponding (also finite-difference) Hessian at the same point; fast_log_binomial_regression_cpp for the full model documentation.


Weighted Log-Link (Relative-Risk) Binomial Regression Hessian, Standalone (C++)

Description

Computes the observation-weighted Hessian matrix of the weighted log-likelihood of the constrained log-link binomial regression model documented in full at fast_log_binomial_regression_cpp, at arbitrary caller-supplied beta (not necessarily the MLE), with each observation's contribution multiplied by weights_r[i], via a numerical (central finite-difference 4-point stencil, step h = 10^{-4}) approximation — not an analytic second derivative. Exported standalone — independent of any optimizer run — for direct numerical diagnostics at a specific parameter value.

Usage

get_log_binomial_regression_weighted_hessian_cpp(X, y_r, weights_r, beta)

Arguments

X

A numeric matrix of predictors.

y_r

A binary (0/1) numeric vector of responses.

weights_r

A nonnegative numeric vector of observation weights.

beta

A numeric vector of coefficients \beta at which to evaluate the Hessian.

Value

The finite-difference-approximated weighted Hessian matrix at beta.

A numeric matrix representing the weighted Hessian.

See Also

get_log_binomial_regression_weighted_score_cpp for the corresponding weighted gradient at the same point; get_log_binomial_regression_hessian_cpp for the unweighted version; fast_log_binomial_regression_cpp for the full model documentation.


Weighted Log-Link (Relative-Risk) Binomial Regression Score, Standalone (C++)

Description

Computes the observation-weighted score vector (gradient of the weighted log-likelihood) of the constrained log-link binomial regression model documented in full at fast_log_binomial_regression_cpp, at arbitrary caller-supplied beta (not necessarily the MLE), with each observation's contribution multiplied by weights_r[i], via a numerical (central finite-difference, step h = 10^{-6}) approximation — not an analytic derivative. Exported standalone — independent of any optimizer run — for direct numerical diagnostics at a specific parameter value.

Usage

get_log_binomial_regression_weighted_score_cpp(X, y_r, weights_r, beta)

Arguments

X

A numeric matrix of predictors.

y_r

A binary (0/1) numeric vector of responses.

weights_r

A nonnegative numeric vector of observation weights.

beta

A numeric vector of coefficients \beta at which to evaluate the score.

Value

The finite-difference-approximated weighted score vector at beta.

See Also

get_log_binomial_regression_weighted_hessian_cpp for the corresponding weighted Hessian at the same point; get_log_binomial_regression_score_cpp for the unweighted version; fast_log_binomial_regression_cpp for the full model documentation.


Negative Binomial Regression Hessian, Standalone (C++)

Description

Computes the (analytic) Hessian matrix of the log-likelihood of the mean/dispersion-parameterized negative binomial regression model documented in full at fast_neg_bin_cpp (see also fast_dnbinom_mu_vec_cpp for the underlying density), at arbitrary caller-supplied params (not necessarily the MLE). Exported standalone — independent of any optimizer run — for direct numerical diagnostics at a specific parameter value.

Usage

get_negbin_regression_hessian_cpp(X, y, params)

Arguments

X

A numeric matrix of predictors, as used to fit the model.

y

A numeric vector of nonnegative-integer count responses.

params

A numeric vector [\beta, \log\theta]: the mean-model coefficients followed by the log-dispersion parameter, at which to evaluate the Hessian.

Value

The (p+1) \times (p+1) Hessian matrix of the log-likelihood at params.

See Also

get_negbin_regression_score_cpp for the corresponding gradient at the same point; fast_neg_bin_cpp for the full model documentation.


Compute Negative Binomial Regression Score

Description

Calculates the score vector (gradient of the log-likelihood) for a negative binomial regression model.

Usage

get_negbin_regression_score_cpp(X, y, params)

Arguments

X

A numeric matrix of predictors.

y

A numeric vector of responses (non-negative integers).

params

A numeric vector of parameters [beta, log_theta].

Value

A numeric vector representing the score.


Get the maximum number of threads for OpenMP

Description

Get the maximum number of threads for OpenMP

Usage

get_omp_max_threads_cpp()

Value

Integer.


Get the default optimization dispatch policy

Description

Returns EDI's built-in policy table, consulted by the internal (non-exported) dispatcher edi_optimization_dispatch_policy(), for choosing which optimization algorithm ("newton_raphson", "lbfgs", or "irls") an inference class's C++ model-fitting backend uses by default.

Usage

get_optimization_dispatch_policy()

Details

The dispatcher checks the inference class name against inference_class_overrides's named regular-expression patterns in list order, returning the associated algorithm string at the first match; if none match, it falls back to default_alg ("newton_raphson"). Unlike get_bootstrap_dispatch_policy, there is no separate design-class-scoped override table here — only a single flat pattern list. The built-in overrides are empirical, chosen per model family based on which algorithm converges fastest/most reliably for that likelihood surface in practice — e.g. plain-vanilla generalized linear models with a canonical or near-canonical link (Poisson, quasi-Poisson, robust Poisson, various incidence models) default to "irls", most non-canonical-link and ordinal/survival models default to "lbfgs", and stratified Cox PH and most matched (KK*GLMM-adjacent) models default to "newton_raphson".

Value

A named list with components default_alg (the fallback algorithm, "newton_raphson" in the built-in policy) and inference_class_overrides (a named character vector: regular-expression pattern names to algorithm-name values, matched against the inference class name). Which algorithm converges fastest/most reliably per family is partly a hardware fact (relative cost of Hessian solves vs. L-BFGS iterations depends on BLAS and cache), so these defaults, computed on the maintainer's machine, are not necessarily optimal on yours. Run tune_EDI_for_this_machine to re-measure this axis on your own machine — it will only switch a family's algorithm when the candidate converges on every benchmark replicate, never trading speed for a convergence failure.

See Also

get_bootstrap_dispatch_policy and get_cold_start_dispatch_policy for the analogous policies controlling bootstrap CI type and cold-start behavior; .normalize_optimizer_algorithm for how a resolved algorithm string is validated/normalized before being passed to a C++ backend; tune_EDI_for_this_machine to re-benchmark this policy on your own hardware.

Examples

get_optimization_dispatch_policy()

Proportional-Odds Ordinal Regression Hessian, Standalone (C++)

Description

Computes the (analytic) Hessian matrix of the log-likelihood of the logit-link cumulative (proportional-odds) ordinal regression model documented in full at fast_ordinal_regression_cpp, at arbitrary caller-supplied params (not necessarily the MLE). Exported standalone — independent of any optimizer run — for direct numerical diagnostics at a specific parameter value.

Usage

get_ordinal_regression_hessian_cpp(X, y, params)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see fast_ordinal_regression_cpp).

y

A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding.

params

A numeric vector [\alpha, \beta]: the category thresholds followed by the regression coefficients, at which to evaluate the Hessian.

Value

The Hessian matrix of the log-likelihood at params.

See Also

get_ordinal_regression_score_cpp for the corresponding gradient at the same point; fast_ordinal_regression_cpp for the full model documentation.


Proportional-Odds Ordinal Regression Score, Standalone (C++)

Description

Computes the (analytic) score vector (gradient of the log-likelihood) of the logit-link cumulative (proportional-odds) ordinal regression model documented in full at fast_ordinal_regression_cpp, at arbitrary caller-supplied params (not necessarily the MLE). Exported standalone — independent of any optimizer run — for direct numerical diagnostics (e.g. verifying convergence) at a specific parameter value.

Usage

get_ordinal_regression_score_cpp(X, y, params)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see fast_ordinal_regression_cpp).

y

A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding.

params

A numeric vector [\alpha, \beta]: the category thresholds followed by the regression coefficients, at which to evaluate the score.

Value

The score vector (gradient of the log-likelihood) at params.

See Also

get_ordinal_regression_hessian_cpp for the corresponding Hessian at the same point; fast_ordinal_regression_cpp for the full model documentation.


Get the default parallel dispatch policy

Description

Returns EDI's built-in blocklist-first policy table, consulted by an internal (non-exported) dispatcher, for deciding whether a given "bootstrap" or "rand_ci" (randomization confidence interval) operation is forced to run serially rather than in parallel, for a given inference class and response type. "Blocklist-first" means every operation is parallel-eligible by default; only combinations explicitly matched below are forced serial.

Usage

get_parallel_dispatch_policy()

Details

For a given operation ("bootstrap" or "rand_ci"), the dispatcher looks up that operation's sub-list (e.g. bootstrap) and forces serial execution if either the inference class name matches any of serial_inference_class_patterns (regular expressions, PCRE-flavored — serial_inference_class_patterns supports lookahead, e.g. the built-in ⁠"^InferenceSurvival(?!.*KK)"⁠ pattern matches non-KK survival inference classes but excludes matched-design (KK) survival variants) or the response type matches any of serial_response_types exactly. In the built-in policy, incidence-response inference generally runs serially for both operations (its bootstrap/randomization resampling is not currently parallel-safe or not worth parallelizing at typical trial sizes), and non-KK survival inference and InferenceAllKKWilcoxIVWC are additionally forced serial for bootstrapping specifically.

Value

A named list with two components, bootstrap and rand_ci, each itself a list with serial_inference_class_patterns (a character vector of PCRE-flavored regular expressions matched against the inference class name) and serial_response_types (a character vector of exact response-type strings) — any match in either forces that operation to run serially rather than in parallel. Unlike get_cold_start_dispatch_policy, get_warm_start_dispatch_policy, and get_optimization_dispatch_policy, this table is a correctness fact, not a performance one — every entry exists because that operation is not currently parallel-safe for that class or response type, not because it happens to be slower in parallel. tune_EDI_for_this_machine benchmarks a related but separate question — at what sample size parallel execution starts to beat serial, and which core count wins — and will never propose un-serializing anything named here.

See Also

get_bootstrap_dispatch_policy, get_cold_start_dispatch_policy, and get_optimization_dispatch_policy for the analogous policies controlling bootstrap CI type, cold-start behavior, and default optimizer algorithm (none of which control parallel-vs-serial dispatch); tune_EDI_for_this_machine for the machine-dependent parallel-crossover/core-count benchmark this blocklist constrains but is never itself a target of.

Examples

get_parallel_dispatch_policy()

Calculates the standard error of the difference in restricted mean survival times

Description

Calculates the standard error of the difference in restricted mean survival times

Usage

get_restricted_mean_se_diff(y, dead, w)

Arguments

y

Numeric vector of survival times.

dead

Integer vector of event indicators (1=event, 0=censored).

w

Integer vector of treatment assignments (1=treatment, 0=control).

Value

The standard error of the difference.


Calculates standard variance using the formula from Uno et al

Description

Var(RMST) = \sum_j A(t_j)^2 d_j / (n_j (n_j - d_j)) where A(t_j) = \int_{t_j}^{\tau} S(u) du is the remaining area under the KM curve from event time t_j to the last observation \tau. Here d_j is the number of events at t_j, and n_j is the number at risk just before t_j. Terms where n_j == d_j are omitted: S drops to 0 there, so A(t_j) = 0 and the contribution is 0 in the limit regardless of the undefined Greenwood denominator.

Usage

get_restricted_mean_se_for_group(y, dead)

Arguments

y

Numeric vector of survival times.

dead

Integer vector of event indicators (1=event, 0=censored).

Value

The standard error of the restricted mean.


Stereotype Logit Regression Hessian, Standalone (C++)

Description

Computes the (analytic) Hessian matrix of the log-likelihood of the stereotype (reduced-rank multinomial) logistic regression model documented in full at fast_stereotype_logit_cpp, at arbitrary caller-supplied params (not necessarily the MLE). Exported standalone — independent of any optimizer run — for direct numerical diagnostics at a specific parameter value.

Usage

get_stereotype_logit_hessian_cpp(X, y, params)

Arguments

X

A numeric matrix of predictors (no intercept column needed; see fast_stereotype_logit_cpp).

y

A numeric vector of categorical (nominal or ordinal) responses; only the set of distinct values matters, not their numeric coding or order.

params

A numeric vector of the full joint parameter vector [\alpha, \beta, \gamma], at which to evaluate the Hessian.

Value

The Hessian matrix of the log-likelihood at params.

See Also

get_stereotype_logit_score_cpp for the corresponding gradient at the same point; fast_stereotype_logit_cpp for the full model documentation.


Compute Stereotype Logit Score

Description

Calculates the score vector (gradient of the log-likelihood) for a stereotype logit model.

Usage

get_stereotype_logit_score_cpp(X, y, params)

Arguments

X

A numeric matrix of predictors.

y

A numeric vector of responses.

params

A numeric vector of parameters.

Value

A numeric vector representing the score.


Calculates the difference in a survival statistic (median or restricted mean) between two groups (treatment vs control)

Description

Calculates the difference in a survival statistic (median or restricted mean) between two groups (treatment vs control)

Usage

get_survival_stat_diff(y, dead, w, requested_stat)

Arguments

y

Numeric vector of survival times.

dead

Integer vector of event indicators (1=event, 0=censored).

w

Integer vector of treatment assignments (1=treatment, 0=control).

requested_stat

A string, either "median" or "restricted_mean".

Value

The difference in the statistic (treatment - control).


Calculates the median or restricted mean survival time for a single group

Description

Calculates the median or restricted mean survival time for a single group

Usage

get_survival_stat_for_group(y, dead, requested_stat)

Arguments

y

Numeric vector of survival times.

dead

Integer vector of event indicators (1=event, 0=censored).

requested_stat

A string, either "median" or "restricted_mean".

Value

The calculated statistic.


Get the default warm-start dispatch policy

Description

Returns EDI's built-in policy table for choosing whether warm starts (reusing a previous fit's parameters/curvature to seed the next fit, e.g. across bootstrap or randomization replicates) are enabled for a given inference class during a given resampling or simulation operation (one of "jackknife", "non_param_boot", "bayesian_boot", "param_boot", or "rand").

Usage

get_warm_start_dispatch_policy()

Details

The internal (non-exported) dispatcher, edi_warm_start_dispatch_policy(inference_class, operation, n), consults this table's default plus, per operation, two override layers: inference_class_overrides (a sample-size-independent pattern table, as before) and n_conditioned_overrides (a list of list(pattern, value, n_min, n_max) rules — each pattern only applies when the current sample size n falls in [n_min, n_max) — encoding empirical findings like "the extra bookkeeping only pays off once resampling is expensive enough per replicate, so disable below n=200/500/1000 for these families"). Both layers are returned by this function and are both reachable via set_warm_start_dispatch_policy.

Value

A named list with default (logical, TRUE in the built-in policy) and one component per operation (jackknife, non_param_boot, bayesian_boot, param_boot, rand), each itself a list containing inference_class_overrides (a named logical vector: regular-expression pattern names to TRUE/FALSE values, matched against the inference class name, applied regardless of sample size) and n_conditioned_overrides (a list of list(pattern, value, n_min, n_max) rules, applied only when n is supplied and falls in [n_min, n_max)); see Details. Like the cold-start table, every TRUE/FALSE entry here (including the n-thresholds) is an empirical performance judgment computed on the maintainer's machine, not a correctness fact. Run tune_EDI_for_this_machine to re-measure this axis, per resampling operation and sample size, on your own machine.

See Also

get_cold_start_dispatch_policy for the analogous (simpler, single-layer) policy governing the initial cold-start heuristic rather than cross-replicate warm-starting; set_warm_start_dispatch_policy to override this table at runtime; tune_EDI_for_this_machine to re-benchmark it on your own hardware.

Examples

get_warm_start_dispatch_policy()

Compute Weibull Regression Hessian (General Censoring)

Description

Hessian matrix for the Weibull AFT log-likelihood extended to left-, right-, and interval-censored responses. See get_weibull_regression_general_score_cpp() for the input convention.

Usage

get_weibull_regression_general_hessian_cpp(X, y, y_L, y_R, params)

Arguments

X

A numeric matrix of predictors.

y

Exact survival times, NA for censored subjects.

y_L

Censored-interval lower bounds, NA for exact subjects; 0 for left-censored.

y_R

Censored-interval upper bounds, NA for exact subjects; Inf for right-censored.

params

A numeric vector of parameters [beta, log_sigma].

Value

A numeric matrix representing the Hessian.


Compute Weibull Regression Score (General Censoring)

Description

Score vector for the Weibull AFT log-likelihood extended to left-, right-, and interval-censored responses (see TODO-3 in interval_censored_survival_response.md). Exactly one of y[i] or (y_L[i], y_R[i]) must be finite per subject: NA in the unused slot(s).

Usage

get_weibull_regression_general_score_cpp(X, y, y_L, y_R, params)

Arguments

X

A numeric matrix of predictors.

y

Exact survival times, NA for censored subjects.

y_L

Censored-interval lower bounds, NA for exact subjects; 0 for left-censored.

y_R

Censored-interval upper bounds, NA for exact subjects; Inf for right-censored.

params

A numeric vector of parameters [beta, log_sigma].

Value

A numeric vector representing the score.


Compute Weibull Regression Hessian

Description

Calculates the Hessian matrix (second derivatives of the log-likelihood) for a Weibull AFT regression model.

Usage

get_weibull_regression_hessian_cpp(X, y, dead, params)

Arguments

X

A numeric matrix of predictors.

y

A numeric vector of survival times.

dead

A numeric vector of event indicators.

params

A numeric vector of parameters [beta, log_sigma].

Value

A numeric matrix representing the Hessian.


Compute Weibull Regression Score

Description

Calculates the score vector (gradient of the log-likelihood) for a Weibull AFT regression model.

Usage

get_weibull_regression_score_cpp(X, y, dead, params)

Arguments

X

A numeric matrix of predictors.

y

A numeric vector of survival times.

dead

A numeric vector of event indicators.

params

A numeric vector of parameters [beta, log_sigma].

Value

A numeric vector representing the score.


Generic aliased overrides for the incidence g-computation component

Description

Uses the shared randomization two-sided p-value contract; see InferenceRand.

Uses the shared Wald testing-type contract; see InferenceAsymp. 'Wald' is not composed directly (its private 'get_standard_error' clashes with ‘IncidenceGComputation'’s public 'get_standard_error', an R6-forbidden same-name-both-slots collision), so only this piece is aliased.

Usage

incidence_gcomp_generic_alias_overrides

Details

Shared self$-aliased public overrides for methods of the IncidenceGComputation component whose bodies call super$...; composed by InferenceIncidGCompRiskDiff and InferenceIncidGCompRiskRatio.


Short display label for an inference class

Description

Short display label for an inference class

Usage

inference_class_short_label(name, tau = NA_real_)

Arguments

name

Inference class name (character), e.g. "InferenceContinOLS".

tau

Quantile level for a quantile-regression class's display label (e.g. '"InferenceContinQuantileRegr"', '"InferenceContinKKQuantileRegrOneLik"' – every concrete class whose wordified label contains the bare word "Quantile" is tagged the '"quantile_regression_effect"' estimand, so this is safe as an unconditional word-level substitution rather than a per-class allowlist). 'NA_real_' (default) or '0.5' renders "Quantile" as "Median" (e.g. "KK Quantile Regr" -> "KK Median Regr"); any other value renders it as "Quantile (<tau*100> so the actual tested quantile is visible when more than one might be compared – mirrors ‘estimand_short_label()'’s own tau handling (per user request, 2026-08-27: same rewording, now applied to the inference *class* name, not just the estimand label).


Inverse Logit (Logistic) Function

Description

Computes the inverse logit (standard logistic sigmoid) function, \mathrm{logit}^{-1}(x) = 1/(1 + e^{-x}), the canonical mean function for binomial/logistic-family models throughout this package (mapping a linear predictor on the log-odds scale back to a probability). The result is clamped to [\code{zero\_one\_logit\_clamp}, 1 - \code{zero\_one\_logit\_clamp}] before being returned, so an extreme x (e.g. from a poorly identified or diverging fit) cannot produce an exact 0 or 1 probability that would later cause a -Inf/ NaN when log-transformed downstream (e.g. in a log-likelihood).

Usage

inv_logit(x, zero_one_logit_clamp = .Machine$double.eps)

Arguments

x

Any real number (or vector), typically a fitted linear predictor \eta = x_i^\top\beta on the log-odds scale.

zero_one_logit_clamp

The clamping distance from the 0/1 boundaries applied to the result. Default .Machine$double.eps.

Value

The inverse-logit-transformed value(s), in ⁠(0, 1)⁠, the same length as x.

See Also

logit for the forward transform.

Examples

inv_logit(0)

Logit (Log-Odds) Transform

Description

Computes the logit (log-odds) function \mathrm{logit}(p) = \log\left(p / (1-p)\right), the canonical link function for binomial/logistic-family models throughout this package. p is first clamped to [\code{zero\_one\_logit\_clamp}, 1 - \code{zero\_one\_logit\_clamp}] before transforming, so exact 0 or 1 inputs (which would otherwise map to -\infty/\infty) instead return a large but finite value; this is what lets proportion/fractional responses with mass exactly at the boundary be used as pseudo-continuous inputs to logit-scale machinery elsewhere in the package (e.g. fast_ols_cpp-backed logit-transform-then-OLS shortcuts) without producing non-finite values.

Usage

logit(p, zero_one_logit_clamp = .Machine$double.eps)

Arguments

p

The value(s) to transform, nominally in ⁠(0, 1)⁠ (values outside that range, or exactly 0/1, are clamped rather than rejected).

zero_one_logit_clamp

The clamping distance from the 0/1 boundaries applied to p before transforming. Default .Machine$double.eps.

Value

The logit-transformed value(s) as a real number (or vector), the same length as p.

See Also

inv_logit for the inverse transform.

Examples

logit(0.25)

LRT confidence interval by Newton-Raphson + bisection (Rcpp implementation)

Description

Implements the bracket search + NR+bisection loop entirely in C++, calling back into R only for fit_null_fn, neg_loglik_fn, and score_fn. The derivative dp/d\delta = 2 f_{\chi^2}(T) \cdot \mathrm{score}[j] follows from the envelope theorem.

Usage

lrt_ci_nr_cpp(
  fit_null_fn,
  neg_loglik_fn,
  score_fn,
  est,
  full_negloglik,
  alpha,
  step,
  lower_seed,
  upper_seed,
  j,
  max_bracket = 60L,
  max_nr_iter = 25L,
  tol_p = 1e-07,
  tol_bracket = 1e-08
)

Arguments

fit_null_fn

R function delta -> list (constrained null fit)

neg_loglik_fn

R function fit -> double (negative log-likelihood)

score_fn

R function fit -> numeric (score vector)

est

Point estimate of the treatment effect

full_negloglik

Negative log-likelihood of the unrestricted model

alpha

Significance level (e.g. 0.05)

step

Initial step size for exponential bracket search

lower_seed

Initial lower-bound candidate, typically the Wald lower CI

upper_seed

Initial upper-bound candidate, typically the Wald upper CI

j

1-indexed position of the treatment coefficient in the score vector

max_bracket

Maximum exponential bracket search iterations (default 60)

max_nr_iter

Maximum NR+bisection iterations per bound (default 25)

tol_p

P-value convergence tolerance (default 1e-7)

tol_bracket

Bracket-width convergence tolerance (default 1e-8)

Value

Unnamed numeric vector of length 2: [lower_bound, upper_bound]


Fast Mean Calculation

Description

Calculates the mean of a numeric vector using Rcpp for speed.

Usage

mean_cpp(x)

Arguments

x

A numeric vector.

Value

The mean of the vector.


Miettinen-Nurminen Confidence Interval for Risk Difference

Description

Computes an approximate Miettinen-Nurminen confidence interval for the risk difference by inverting the score test with a bisection search.

Usage

mn_ci_cpp(x_t, n_t, x_c, n_c, p_t_obs, p_c_obs, alpha, pval_epsilon)

Arguments

x_t

Number of events in treatment.

n_t

Number of subjects in treatment.

x_c

Number of events in control.

n_c

Number of subjects in control.

p_t_obs

Observed treatment-arm risk.

p_c_obs

Observed control-arm risk.

alpha

The confidence level is 1 - alpha.

pval_epsilon

Bisection tolerance in p-value space.

Value

A length-2 numeric vector containing the lower and upper CI bounds.


Constrained MLE for Risk Difference (Miettinen-Nurminen)

Description

Solves the likelihood equations for p_C subject to p_T - p_C = delta. Uses bisection on the score function (derivative of log-likelihood) which is monotonic and well-behaved.

Usage

mn_constrained_mle_pc_cpp(x_t, n_t, x_c, n_c, delta)

Export of C++ function mn_pvalue_cpp

Description

Evaluates the two-sided asymptotic p-value for the Miettinen-Nurminen restricted-maximum-likelihood score test of H_0: p_T - p_C = \delta in two independent binomial samples (see InferenceIncidMiettinenNurminenRiskDiff for the class that consumes this function). Internally: mn_z_statistic_cpp computes the score z statistic using the constrained MLEs \tilde p_C, \tilde p_T = \tilde p_C + \delta (mn_constrained_mle_pc_cpp, found by bisecting the constrained score equation to zero) in place of the unconstrained sample proportions in the variance formula, with a small-sample correction factor (n_T+n_C)/(n_T+n_C-1) applied to the naive binomial variance; this function then returns 2\,\Phi(-|z|), the two-sided normal-tail p-value. Returns NA if either arm is empty, \delta is outside (-1, 1), or the resulting z is not finite (e.g. the constrained variance estimate is non-positive).

Usage

mn_pvalue_cpp(x_t, n_t, x_c, n_c, delta, p_t_obs, p_c_obs)

Arguments

x_t

Number of events in treatment.

n_t

Number of subjects in treatment.

x_c

Number of events in control.

n_c

Number of subjects in control.

delta

Null risk difference.

p_t_obs

Observed treatment-arm risk.

p_c_obs

Observed control-arm risk.

Value

The two-sided p-value.


Miettinen-Nurminen Score Z Statistic for Risk Difference

Description

Computes the score-style z statistic for testing the null risk difference p_T - p_C = delta in two independent binomial samples, using the Miettinen-Nurminen constrained nuisance estimates.

Usage

mn_z_statistic_cpp(x_t, n_t, x_c, n_c, delta, p_t_obs, p_c_obs)

Arguments

x_t

Number of events in treatment.

n_t

Number of subjects in treatment.

x_c

Number of events in control.

n_c

Number of subjects in control.

delta

Null risk difference.

p_t_obs

Observed treatment-arm risk.

p_c_obs

Observed control-arm risk.

Value

The asymptotic z statistic.


Export of C++ function newcombe_independent_ci_cpp

Description

Computes Newcombe's "Method 10" hybrid confidence interval for the difference between two independent proportions p_1 - p_2 (see InferenceIncidNewcombeRiskDiff for the class that consumes this function). Separate Wilson score intervals [\ell_1, u_1] and [\ell_2, u_2] are computed for each proportion individually (via wilson_score_interval_cpp), then combined as

\left[\,(p_1-p_2) - \sqrt{(p_1-\ell_1)^2 + (u_2-p_2)^2},\ \ (p_1-p_2) + \sqrt{(u_1-p_1)^2 + (p_2-\ell_2)^2}\,\right],

clamped to [-1, 1]. This avoids the boundary/coverage problems of the naive normal-approximation (Wald) interval on a risk difference while remaining closed-form (no iterative score-test inversion). Returns c(NA, NA) if either sample size is non-positive.

Usage

newcombe_independent_ci_cpp(x1, n1, x2, n2, alpha)

Arguments

x1

Number of events in group 1.

n1

Number of subjects in group 1.

x2

Number of events in group 2.

n2

Number of subjects in group 2.

alpha

The confidence level is 1-\alpha.

Value

A length-2 numeric vector containing the lower and upper CI bounds for p_1 - p_2.

References

Newcombe, R. G. (1998). "Interval Estimation for the Difference Between Independent Proportions: Comparison of Eleven Methods." Statistics in Medicine, 17(8), 873-890, doi:10.1002/(SICI)1097-0258(19980430)17:8<873::AID-SIM779>3.0.CO;2-I.

See Also

newcombe_paired_ci_cpp for the matched-pair generalization of this same hybrid-score method.


Newcombe Hybrid Score Interval for Paired Proportions

Description

Newcombe Hybrid Score Interval for Paired Proportions

Usage

newcombe_paired_ci_cpp(n11, n10, n01, n00, alpha)

Export of C++ function ols_hc2_post_fit_cpp

Description

Given an already-fitted ordinary least squares model, computes the HC2 heteroskedasticity-consistent sandwich covariance matrix (MacKinnon-White) for the coefficients: \widehat{\mathrm{Var}}(\hat\beta) = B\,M\,B, with "bread" B = (X^\top X)^{-1} and "meat" M = X^\top \mathrm{diag}(\omega_i) X, where \omega_i = r_i^2 / (1 - h_{ii}) — the squared OLS residual r_i leverage-corrected by dividing by 1 - h_{ii} (h_{ii} the i-th diagonal of the OLS hat matrix X(X^\top X)^{-1}X^\top), unlike the plain HC0 sandwich (gcomp_logistic_post_fit_cpp's logistic analogue, or this function's own uncorrected r_i^2 meat) which does not correct for leverage. HC2 is unbiased under homoskedasticity for balanced designs and generally has better small-sample properties than HC0/HC1 when leverage is uneven. Internally, this function first computes the (design-only) "setup" quantities bread/hat via ols_hc2_setup_cpp, then calls ols_hc2_post_fit_precomputed_cpp; callers who already have those precomputed (e.g. across repeated resampling on the same fixed design) can call the precomputed variant directly instead to skip recomputing the (X^\top X)^{-1} bread and leverage each time.

Usage

ols_hc2_post_fit_cpp(X_fit, y, coef_hat, j_treat)

Arguments

X_fit

A numeric matrix of predictors, as used to fit the model.

y

A numeric vector of responses.

coef_hat

A numeric vector of fitted OLS coefficients \hat\beta, same length and column order as X_fit.

j_treat

1-based column index of the treatment indicator in X_fit.

Value

A list with components beta_hat (\hat\beta_{j_{\mathrm{treat}}}), ssq_hat (its HC2 variance), se (its HC2 standard error), vcov (the full p \times p HC2 covariance matrix), std_err (per-coefficient HC2 standard errors), and z_vals (per-coefficient Wald z-statistics, \hat\beta_j / \widehat{\mathrm{SE}}(\hat\beta_j)).

References

MacKinnon, J. G., and White, H. (1985). "Some Heteroskedasticity-Consistent Covariance Matrix Estimators with Improved Finite Sample Properties." Journal of Econometrics, 29(3), 305-325, doi:10.1016/0304-4076(85)90158-7, for the HC2 estimator used here.


Fast G-Computation (Standardization) Point Estimate and Model-Based Inference for a Proportional-Odds Ordinal Model (C++)

Description

Computes the same standardized (G-computation) marginal difference in expected ordinal category score as gcomp_ordinal_proportional_odds_post_fit_cpp — under a fitted proportional-odds (cumulative logit) model \mathrm{logit}\,\Pr(Y_i \le k \mid x_i) = \alpha_k - x_i^\top\beta, k = 1, \ldots, K-1 — but additionally supplies model-based inferential quantities that the other function omits. The fitted linear predictor for each subject is recomputed with the treatment column (j_treat) forced to 1 (\eta_{1,i}) and to 0 (\eta_{0,i}), the standardized expected scores mean1, mean0, and their difference md are formed exactly as in gcomp_ordinal_proportional_odds_post_fit_cpp, and then:

Unlike the sandwich-based post-fit helpers elsewhere in this package (e.g. gcomp_logistic_post_fit_cpp), the covariance here comes purely from the model's observed information and does not attempt to be robust to misspecification.

Usage

ordinal_gcomp_post_fit_cpp(X_fit, y, coef_hat, alpha_hat, j_treat)

Arguments

X_fit

Numeric matrix of predictors used to fit the model.

y

Numeric vector of the ordinal responses used to fit the model (needed to reconstruct the OrdinalRegression likelihood for the Hessian; only categorization, not numeric coding, matters).

coef_hat

Numeric vector of fitted proportional-odds regression coefficients \hat\beta, same length and column order as X_fit.

alpha_hat

Numeric vector of the K-1 fitted, increasing category thresholds \hat\alpha_1, \ldots, \hat\alpha_{K-1}.

j_treat

1-based column index of the treatment indicator in X_fit.

Value

A list with elements vcov (model-based covariance of \hat\beta), std_err, z_vals (both length p, one per column of X_fit), mean1, mean0, md (mean1 - mean0), and se_md (delta-method SE of md). Errors (via stop()) if j_treat is out of bounds, the dimensions of y/coef_hat are inconsistent with X_fit, the Hessian is not invertible, or the resulting covariance has any non-finite entry.

See Also

gcomp_ordinal_proportional_odds_post_fit_cpp for the point-estimate-only variant (no y, no Hessian, no inference) that this function's point estimates match; fast_ordinal_regression_cpp for the fitting routine that produces coef_hat/alpha_hat.


Pocock-Simon Covariate-Adaptive Minimization: Assign and Update Counts (C++)

Description

The stateful wrapper around pocock_simon_assign_cpp actually used to drive a one-subject-at-a-time Pocock-Simon minimization design: it computes the next assignment using the identical imbalance-score/biased-coin logic documented on pocock_simon_assign_cpp (see that page for the full model), then mutates counts in place, incrementing, for every covariate in subject_levels_idx, the count of the assigned arm at that covariate's level — so the running covariate-by-arm counts stay correct for the next subject's assignment.

Usage

pocock_simon_assign_and_update_cpp(
  counts,
  subject_levels_idx,
  weights,
  p_best,
  prob_T
)

Arguments

counts

A numeric matrix, one row per stratification-covariate level and one column per treatment arm (2 columns); modified in place to record the new assignment.

subject_levels_idx

An integer vector (1-indexed) giving, for the subject being assigned, which row of counts each covariate's current level corresponds to.

weights

A numeric vector, parallel to subject_levels_idx, of the relative weight placed on each covariate's imbalance.

p_best

The probability of assigning the arm that minimizes the combined imbalance score (the biased coin).

prob_T

The Bernoulli probability used to break an exact imbalance tie.

Value

The assigned treatment arm, 0 or 1. As a side effect, counts is incremented at row subject_levels_idx[j], column w (the returned assignment) for every covariate j.

Note

Reproducibility notes: see pocock_simon_assign_cpp.

References

Pocock, S. J. and Simon, R. (1975). "Sequential Treatment Assignment with Balancing for Prognostic Factors in the Controlled Clinical Trial." Biometrics, 31(1), 103-115.

See Also

pocock_simon_assign_cpp for the underlying non-mutating decision logic and the full imbalance-score model.


Pocock-Simon Covariate-Adaptive Minimization: Assignment Decision (C++)

Description

Decides the next subject's treatment assignment under Pocock and Simon's (1975) covariate-adaptive minimization algorithm, without modifying any state (see pocock_simon_assign_and_update_cpp for the state-updating wrapper actually used by the stepwise design). counts stacks, for every level of every stratification covariate, the number of previously-assigned subjects at that level currently in each treatment arm (one row per covariate level, one column per arm — the design supports exactly 2 arms). The subject to be assigned belongs to one level per covariate, given by subject_levels_idx (1-indexed rows into counts).

Usage

pocock_simon_assign_cpp(counts, subject_levels_idx, weights, p_best, prob_T)

Arguments

counts

A numeric matrix, one row per stratification-covariate level (across all covariates, stacked) and one column per treatment arm (2 columns), of counts of subjects previously assigned to that level/arm combination.

subject_levels_idx

An integer vector (1-indexed, length = number of stratification covariates) giving, for the subject being assigned, which row of counts each covariate's current level corresponds to.

weights

A numeric vector, parallel to subject_levels_idx, of the relative weight w_j placed on each covariate's imbalance in the combined score G_k.

p_best

The probability of assigning the arm that minimizes G_k (the biased coin); 1 is deterministic minimization, values near 0.5 approach simple randomization.

prob_T

The Bernoulli probability used to break an exact tie in G_0 = G_1 (assigns arm 1 with this probability); typically 0.5.

Details

For each candidate arm k \in \{0,1\}, a marginal imbalance score is computed as the weighted sum, over the subject's covariates j, of the variance across arms of the level's counts if the subject were assigned to arm k:

G_k = \sum_j w_j \, \mathrm{Var}_t\big(n_{j,t} + \mathbb{1}\{t=k\}\big),

where n_{j,t} is the current count for the subject's level of covariate j in arm t, and the variance is taken over the (here, 2) arms with the usual n-1 divisor — so G_k is smallest for the arm that would leave covariate margins most balanced. The arm with the smaller G_k (best_trt) is then assigned with a biased-coin probability p_best (and the other arm with probability 1 - p_best); if G_0 = G_1 exactly (perfect tie), the assignment is instead a single Bernoulli(prob_T) draw for arm 1. Setting p_best = 1 recovers Taves' deterministic minimization; p_best strictly between 0.5 and 1 (Pocock and Simon's recommendation, e.g. 0.75-0.85) retains most of minimization's balancing power while preserving some unpredictability of the next assignment.

Value

The assigned treatment arm, 0 or 1.

Note

Seeded from one R::unif_rand() draw into edi_rng::RRng (RNG.h), a portable re-implementation of R's own Mersenne-Twister generator – a given seed therefore produces identical draws in R and in any future binding (e.g. Python) using the same core and the same seed, even though this call does not continue R's own live session stream bit- for-bit (see pocock_simon_redraw_w_cpp for the one function in this file where that distinction matters and is handled).

References

Pocock, S. J. and Simon, R. (1975). "Sequential Treatment Assignment with Balancing for Prognostic Factors in the Controlled Clinical Trial." Biometrics, 31(1), 103-115.

See Also

pocock_simon_assign_and_update_cpp for the assign-and-mutate-counts wrapper; pocock_simon_redraw_w_cpp for the batch re-derivation of a whole assignment sequence from scratch.


Pocock-Simon Covariate-Adaptive Minimization: Batch Redraw of a Whole Assignment Sequence (C++)

Description

Replays the same Pocock-Simon minimization decision logic as pocock_simon_assign_cpp (see that page for the imbalance-score and biased-coin model) over an entire sequence of n subjects in one call, starting from empty covariate-by-arm counts (all zero) and updating them internally after each subject, rather than being called once per subject with externally-maintained counts. This is used to re-derive (“redraw”) a full sequence of assignments in a single vectorized pass — e.g. for simulation, or for reconstructing what a one-at-a-time run would have produced — while consuming R's uniform random stream in exactly the same order a subject-by-subject loop calling pocock_simon_assign_cpp would have.

Usage

pocock_simon_redraw_w_cpp(
  x_levels_matrix,
  num_levels_total,
  weights,
  p_best,
  prob_T
)

Arguments

x_levels_matrix

An integer matrix with one row per subject and one column per stratification covariate; entry (i, j) is the 1-indexed level (row of the internal counts table) of covariate j for subject i.

num_levels_total

The total number of distinct covariate levels across all covariates (i.e. the number of rows the internal counts table has).

weights

A numeric vector, one per covariate column of x_levels_matrix, of the relative weight placed on each covariate's imbalance in the combined score.

p_best

The probability of assigning the arm that minimizes the combined imbalance score (the biased coin).

prob_T

The Bernoulli probability used to break an exact imbalance tie.

Value

An integer vector of length n, the treatment arm (0 or 1) assigned to each subject in row order of x_levels_matrix.

Note

Continues R's actual live .Random.seed stream via edi_rng::RRng (RNG.h) rather than seeding independently – output is bit-identical to what calling R's own unif_rand() directly, in this same loop, would have produced (verified against a pure-R reference implementation in test-pocock-simon-redraw-buffers.R). Requires RNGkind ("Mersenne-Twister", "Inversion"), R's default.


Prints the results table from an InferenceSuite run_all_inference() call – the same table screen = TRUE prints during the call itself, so a user who assigned the return value and later types its name (or calls print() on it) sees a readable table rather than a raw nested list dump. The table itself is rendered by run_all_inference_format_pretty_table(): rows sorted by estimand, with a double rule under the header and a single rule between estimand groups and at the bottom, class names and estimand values shortened for display (never the underlying results_table values), and a cov_model letter-key legend appended when applicable – see that function's own documentation for the exact column-by-column rendering rules.

Description

Prints the results table from an InferenceSuite run_all_inference() call – the same table screen = TRUE prints during the call itself, so a user who assigned the return value and later types its name (or calls print() on it) sees a readable table rather than a raw nested list dump. The table itself is rendered by run_all_inference_format_pretty_table(): rows sorted by estimand, with a double rule under the header and a single rule between estimand groups and at the bottom, class names and estimand values shortened for display (never the underlying results_table values), and a cov_model letter-key legend appended when applicable – see that function's own documentation for the exact column-by-column rendering rules.

Usage

## S3 method for class 'EDIInferenceSuiteResults'
print(x, ...)

Arguments

x

An EDIInferenceSuiteResults object, as returned by InferenceSuite$run_all_inference().

...

Ignored; present for S3 consistency with the generic.

Value

x, invisibly.


CI inversion by p-value bracket search + bisection (Rcpp implementation)

Description

Used by invert_test_pval_confidence_interval (score CI and any other CI that inverts a scalar p-value function with no derivative available). The Wald bounds are tried first as bracket candidates before falling back to exponential search. Bisection then polishes to tol in delta-space.

Usage

pval_invert_ci_cpp(
  pval_fn,
  est,
  alpha,
  step,
  lower_seed,
  upper_seed,
  max_bracket = 60L,
  max_bisect = 60L,
  tol = 1e-06
)

Arguments

pval_fn

R function delta -> double two-sided p-value

est

Point estimate of the treatment effect

alpha

Significance level (e.g. 0.05)

step

Initial step size for exponential bracket search

lower_seed

Wald lower CI bound (used as first bracket candidate; pass NA_real_ to skip)

upper_seed

Wald upper CI bound (same)

max_bracket

Maximum exponential bracket search iterations (default 60)

max_bisect

Maximum bisection iterations (default 60)

tol

Bracket-width convergence tolerance in delta-space (default 1e-6)

Value

Unnamed numeric vector of length 2: [lower_bound, upper_bound]


Robust Negative Binomial Regression with Backward Column-Dropping Fallback

Description

Fits a negative-binomial GLM via glm.nb (log link, joint ML estimation of the regression coefficients and the dispersion parameter \theta), falling back to a smaller model when the fit throws an error (typically non-convergence of \theta, or a singular design). On each failure, the last column of data_obj is dropped and the fit is retried against the same form_obj (which must resolve to y ~ . or similar so that its right-hand side tracks the shrinking column set); this repeats until a fit succeeds or every predictor column has been removed, at which point NA is returned. Because columns are dropped strictly from the right, callers should order data_obj's columns from most to least important a priori, or accept that this is a best-effort robustness measure rather than a principled model-selection procedure.

Usage

robust_negbinreg(form_obj, data_obj)

Arguments

form_obj

The model formula, typically y ~ . so its right-hand side automatically tracks data_obj's shrinking column set across retries.

data_obj

The data frame to run negative-binomial regression on; its last column is dropped on each retry, in order, until a fit converges or no columns remain.

Value

The fitted glm.nb model object, or NA if no column subset (down to and including the response alone) produced a successful fit.

Examples

dat = data.frame(y = rpois(10, 2), x1 = rnorm(10), x2 = rnorm(10))
robust_negbinreg(y ~ ., dat)

Robust Parametric Survival Regression from Response/Censoring Vectors

Description

Convenience wrapper around robust_survreg_with_surv_object that builds the Surv object from separate response and censoring vectors first. See that function for the full description of the warm-start-then-random-restart fitting strategy used to make survreg converge reliably even from poor or near-singular starting points.

Usage

robust_survreg(
  y,
  dead,
  cov_matrix_or_vector,
  dist = "weibull",
  num_max_iter = 50
)

Arguments

y

The (possibly right-censored) response vector (event/censoring time).

dead

The event indicator (1 if the event was observed/uncensored, 0 if right-censored at y).

cov_matrix_or_vector

The design matrix (or a single covariate vector) of predictors, excluding the intercept (one is added by the internal ~ . formula).

dist

The parametric AFT distribution family passed to survreg (default "weibull"); see that function's dist argument for the full list of supported families.

num_max_iter

Maximum number of random-restart attempts if the direct fit fails or does not converge (default 50); see robust_survreg_with_surv_object.

Value

The fitted survreg model object, or NULL if no attempt converged to a fit with no NA coefficients within num_max_iter tries.

Examples

X = matrix(rnorm(500), 100, 5)
y = runif(100)
dead = rbinom(100, 1, 0.5)
robust_survreg(y, dead, X)

Robust Parametric Survival Regression (AFT) with Warm-Start and Random-Restart Fallback

Description

Fits a parametric accelerated-failure-time (AFT) survival regression via survreg on surv_object ~ . over the columns of cov_matrix_or_vector, with two layers of robustness against survreg's well-known sensitivity to starting values and near-collinear design matrices:

  1. Preprocessing: near-collinear columns of the design matrix are dropped first via drop_highly_correlated_cols then drop_linearly_dependent_cols, before any fitting is attempted.

  2. Warm start (Weibull only): when dist = "weibull", a fast closed-form-gradient Weibull fit (fast_weibull_regression) is attempted first; if it succeeds and returns a finite log-likelihood, its coefficients and \log(\hat\sigma) are passed to survreg as the init vector, which typically converges the true MLE in a single survreg call. If this warm-started fit is unavailable, fails, or produces NA coefficients, fitting falls through to the general random-restart loop below (for all other dist values, this warm start is skipped entirely).

  3. Random-restart loop: starting from an all-zero init vector, survreg is called repeatedly (perturbing init by an independent standard-normal jitter, init + rnorm(length(init)), after every failed attempt) until a fit with no NA coefficients is obtained or num_max_iter attempts are exhausted, at which point NULL is returned.

survreg.control(maxiter = 100, rel.tolerance = 1e-9, outer.max = 10) is used throughout (tighter than survreg's own defaults) to reduce the chance of a spuriously "converged" fit at a poor optimum.

Usage

robust_survreg_with_surv_object(
  surv_object,
  cov_matrix_or_vector,
  dist = "weibull",
  num_max_iter = 50
)

Arguments

surv_object

The survival object (built from the response vector and censoring vector via Surv).

cov_matrix_or_vector

The design matrix (or a single covariate vector) of predictors, excluding the intercept (one is added by the internal ~ . formula).

dist

The parametric AFT distribution family passed to survreg (default "weibull"); only "weibull" triggers the closed-form warm start.

num_max_iter

Maximum number of random-restart attempts if the (possibly warm-started) direct fit fails or does not converge (default 50).

Value

The fitted survreg model object, or NULL if no attempt converged to a fit with no NA coefficients within num_max_iter tries.

Examples

X = matrix(rnorm(500), 100, 5)
y = runif(100)
dead = rbinom(100, 1, 0.5)
surv = survival::Surv(y, dead)
robust_survreg_with_surv_object(surv, X)

Sample Mode

Description

Thin R wrapper around sample_mode_cpp(), which returns the most frequently occurring value in data. Integer, logical, double, character, and factor vectors are all supported (dispatched internally on TYPEOF(data); factors preserve their class/levels attributes on the returned value). NA (and, for doubles, NaN as a category distinct from NA) is counted like any other value and can itself be returned as "the mode" if it is the most frequent entry. Ties are broken by first occurrence: among values tied for the highest count, the one that appears earliest in data is returned — this is a positional, not a numeric/lexicographic, tie-break rule.

Usage

sample_mode(data)

Arguments

data

A vector (integer, logical, double, character, or factor) to compute the mode of.

Value

A length-1 vector (same type as data) holding the most frequently occurring value, with ties broken in favor of whichever tied value occurs earliest in data.

Examples

sample_mode(c(1, 2, 2, 3))

Update the cold-start dispatch policy

Description

Overrides, queries, or resets the runtime policy consulted by edi_cold_start_dispatch_policy() for whether an inference class's C++ fitting backend defaults to an OLS-based smart_cold_start or a plain zero-vector start; see get_cold_start_dispatch_policy for the built-in default table and the rationale behind it.

Usage

set_cold_start_dispatch_policy(policy = NULL, reset = FALSE)

Arguments

policy

Either NULL (no change) or a named list of policy overrides merged into the current configuration (see Details).

reset

If TRUE, discard all overrides and restore the built-in default policy.

Details

Call with no arguments (policy = NULL, reset = FALSE) to retrieve the current configuration without changing it. Pass a named list to policy to merge new/overriding entries into the current configuration via modifyList (so default and/or inference_class_overrides can each be supplied independently, and entries not mentioned are left untouched — this cannot remove an existing override pattern, only add or replace one). Pass reset = TRUE to discard any accumulated overrides and restore the package's built-in default policy exactly as returned by get_cold_start_dispatch_policy.

Value

Invisible NULL when policy is supplied (a mutation), or invisibly the current policy configuration list when called for its side-effect-free query/reset value.

See Also

get_cold_start_dispatch_policy for the policy schema and built-in defaults; set_warm_start_dispatch_policy and set_optimization_dispatch_policy for the analogous setters governing warm-starting and optimizer-algorithm choice; tune_EDI_for_this_machine, which calls this setter with machine-measured overrides rather than hand-picked ones.

Examples

set_cold_start_dispatch_policy(reset = TRUE)

Set the number of cores for parallelization

Description

This function initializes a persistent parallel cluster (either a fork cluster on Unix-like systems or a mirai cluster on others) to be used by all Design and Inference objects. This avoids the overhead of creating clusters repeatedly.

Usage

set_num_cores(num_cores, force_mirai = FALSE)

Arguments

num_cores

Integer number of worker processes to make available.

force_mirai

If TRUE, forces the use of the mirai package even on systems where forking is available.

Details

set_num_cores() sets a global upper bound for parallel work. It does not guarantee that every inference routine will use all requested workers. EDI's inference dispatcher applies a blocklist-first heuristic informed by package benchmarks: workloads that have shown consistent multicore slowdowns are forced to run serially, while the remaining workloads are allowed to use their method-specific warmup heuristics and native thread caps.

The default forced-serial blocklist covers incidence randomization confidence intervals, bootstrap for non-regression KK Wilcoxon inference, bootstrap for non-KK survival procedures, and bootstrap for incidence procedures. Do not expect a universal "more cores is faster" rule.

If you want to change the default policy, use set_parallel_dispatch_policy().

Value

Invisible NULL.

Examples

set_num_cores(2)
unset_num_cores()

Set the number of threads for OpenMP, Eigen, and MKL

Description

Set the number of threads for OpenMP, Eigen, and MKL

Usage

set_omp_num_threads_cpp(n_threads)

Arguments

n_threads

Integer.


Update the optimization dispatch policy

Description

Overrides, queries, or resets the runtime policy consulted by edi_optimization_dispatch_policy() for which optimization algorithm ("newton_raphson", "lbfgs", or "irls") an inference class's C++ model-fitting backend uses by default; see get_optimization_dispatch_policy for the built-in default table and the empirical rationale behind its per-family choices.

Usage

set_optimization_dispatch_policy(policy = NULL, reset = FALSE)

Arguments

policy

Either NULL (no change) or a named list of policy overrides merged into the current configuration (see Details).

reset

If TRUE, discard all overrides and restore the built-in default policy.

Details

Call with no arguments (policy = NULL, reset = FALSE) to retrieve the current configuration without changing it. Pass a named list to policy to merge new/overriding entries into the current configuration via modifyList (default_alg and/or inference_class_overrides can each be supplied independently). Pass reset = TRUE to discard any accumulated overrides and restore the package's built-in default policy exactly as returned by get_optimization_dispatch_policy. Note that this only changes the default algorithm consulted when a fitting call does not itself specify optimization_alg explicitly; an explicit per-call argument still takes precedence.

Value

Invisible NULL when policy is supplied (a mutation), or invisibly the current policy configuration list when called for its side-effect-free query/reset value.

See Also

get_optimization_dispatch_policy for the policy schema and built-in defaults; set_cold_start_dispatch_policy and set_warm_start_dispatch_policy for the analogous setters governing cold-start and warm-start behavior; tune_EDI_for_this_machine, which calls this setter with machine-measured, convergence-checked overrides rather than hand-picked ones.

Examples

set_optimization_dispatch_policy(reset = TRUE)

Update the parallel dispatch policy

Description

EDI uses an empirical, blocklist-first dispatch policy to decide when an inference routine's "bootstrap" or "rand_ci" resampling should be forced serial even if multiple cores are available (see get_parallel_dispatch_policy for the built-in table and the PCRE pattern/response-type matching rules it encodes). This function lets the user update that policy at runtime without editing package internals.

Usage

set_parallel_dispatch_policy(policy = NULL, reset = FALSE)

Arguments

policy

Either NULL (no change), a named list of policy-section overrides, or a custom dispatch function (see Details).

reset

If TRUE, restore the built-in default policy and remove any custom function override.

Details

The policy can be updated in two mutually exclusive ways:

Use reset = TRUE to discard both the list-based overrides and any custom function, restoring the package's built-in default policy exactly as returned by get_parallel_dispatch_policy. Do not expect a universal "more cores is faster" rule — this policy exists because several resampling workloads have shown consistent multicore slowdowns in package benchmarks.

Value

Invisible NULL when policy is supplied (a mutation), or invisibly the current policy configuration list when called for its side-effect-free query/reset value.

See Also

get_parallel_dispatch_policy for the policy schema and built-in defaults (including why this table is a safety blocklist, not a performance one); set_num_cores for setting the actual worker-count upper bound this policy operates within; tune_EDI_for_this_machine for the separate, machine-dependent question of when parallel execution is worthwhile (never applied here — this policy is only ever overridden explicitly, by you).

Examples

set_parallel_dispatch_policy(reset = TRUE)


Update the warm-start dispatch policy

Description

Overrides, queries, or resets the runtime policy consulted by edi_warm_start_dispatch_policy() for whether an inference class reuses a previous fit's parameters/curvature to seed the next fit during resampling; see get_warm_start_dispatch_policy for the built-in default table and its jackknife/non_param_boot/ bayesian_boot/param_boot/rand operation schema (each with an inference_class_overrides layer and an n_conditioned_overrides layer).

Usage

set_warm_start_dispatch_policy(policy = NULL, reset = FALSE)

Arguments

policy

Either NULL (no change) or a named list of per-operation policy overrides merged into the current configuration (see Details).

reset

If TRUE, discard all overrides and restore the built-in default policy.

Details

Call with no arguments (policy = NULL, reset = FALSE) to retrieve the current configuration without changing it. Pass a named list to policy to merge new/overriding entries into the current configuration via modifyList (per-operation sub-lists, e.g. list(rand = list(inference_class_overrides = ...)) or list(rand = list(n_conditioned_overrides = ...)), are merged rather than replaced wholesale — note n_conditioned_overrides is a plain list of rules, so overriding it replaces the whole list for that operation, not a per-rule merge). Pass reset = TRUE to discard any accumulated overrides and restore the package's built-in default policy exactly as returned by get_warm_start_dispatch_policy. This function controls the full dispatch policy, including the sample-size-conditioned n_conditioned_overrides layer.

Value

Invisible NULL when policy is supplied (a mutation), or invisibly the current policy configuration list when called for its side-effect-free query/reset value.

See Also

get_warm_start_dispatch_policy for the policy schema and built-in defaults; set_cold_start_dispatch_policy for the analogous, simpler single-layer setter governing the initial cold-start heuristic; tune_EDI_for_this_machine, which calls this setter with machine-measured overrides rather than hand-picked ones.

Examples

set_warm_start_dispatch_policy(reset = TRUE)

Summarizes an InferenceSuite run_all_inference() result: counts by status, the estimate range across status == "ok" classes, and how many reject at alpha.

Description

Summarizes an InferenceSuite run_all_inference() result: counts by status, the estimate range across status == "ok" classes, and how many reject at alpha.

Usage

## S3 method for class 'EDIInferenceSuiteResults'
summary(object, ...)

## S3 method for class 'summary.EDIInferenceSuiteResults'
print(x, ...)

Arguments

object

An EDIInferenceSuiteResults object, as returned by InferenceSuite$run_all_inference().

...

Ignored; present for S3 consistency with the generic.

x

A summary.EDIInferenceSuiteResults object, as returned by summary.EDIInferenceSuiteResults.

Value

An object of class summary.EDIInferenceSuiteResults, printable via its own print method.

x, invisibly.


Lean GLM Summary (Skips Deviance Residual Quantiles)

Description

A drop-in replacement for summary.glm that produces the identical coefficient table, dispersion estimate, and (optionally) correlation matrix, but omits the five-number summary of the deviance residuals (summary(object$deviance.resid)) that summary.glm() always computes and stores in its deviance.resid component. That residual summary is cheap for a single fit but adds up when summarizing thousands of GLM fits in a resampling loop (e.g. bootstrap or randomization replicates elsewhere in this package), so this function skips it entirely; the returned object's deviance.resid component is simply absent rather than populated, which will matter to code that calls print.summary.glm() on the result or otherwise inspects that field. Every other computation — dispersion estimation (Pearson X^2/\mathrm{df} for Gaussian/Gamma/inverse-Gaussian families, fixed at 1 for Poisson/binomial, unless dispersion is supplied explicitly), the coefficient table (Wald z tests when dispersion is fixed/known, t tests with df.residual degrees of freedom when dispersion is estimated), and the optional correlation/symbolic.cor outputs, is identical to summary.glm.

Usage

summary_glm_lean(
  object,
  dispersion = NULL,
  correlation = FALSE,
  symbolic.cor = FALSE,
  ...
)

Arguments

object

A fitted glm object.

dispersion

The dispersion parameter for the fitting family; if NULL (default), estimated as in summary.glm (fixed at 1 for poisson/binomial, else the Pearson-residual-based moment estimate).

correlation

Logical; if TRUE, the estimated correlation matrix of the coefficients is returned and printed. Default FALSE.

symbolic.cor

Logical; if TRUE and correlation = TRUE, the correlation matrix is printed in symbolic form (see symnum) rather than as numbers. Default FALSE.

...

Currently unused; present only for signature compatibility with summary.glm.

Value

An object of class c("summary.glm") with the same components as summary.glm's return value except deviance.resid, which is not computed and is absent from the result.

See Also

summary.glm, of which this is a residual-summary-skipping variant.

Examples

fit = glm(rbinom(10, 1, 0.5) ~ rnorm(10), family = binomial)
summary_glm_lean(fit)

Toggle the execution of assertions throughout the package

Description

This function enables or disables the internal input validation checks (assertions) for the rest of the R session. It does not modify options(); setting options(edi.run_asserts = FALSE) yourself also disables them. Disabling assertions can provide a significant performance boost in heavy simulations (often 10x-20x speedup), but it removes the safety rails that prevent invalid data from reaching the internal algorithms.

Warning: If assertions are disabled, passing malformed or invalid data to package functions may result in cryptic R errors, incorrect statistical results, or even hard system crashes (SEGFAULTs) at the C++ layer. Only disable assertions if you are certain your data is pre-validated and follows the package requirements exactly.

Usage

toggle_asserts(on = TRUE)

Arguments

on

Logical scalar. If TRUE (default), assertions are executed. If FALSE, they are skipped.


Transform continuous latent signal to the response type scale

Description

A helper function to transform a latent continuous signal to the scale appropriate for a given response_type, identical to the logic used within SimulationFramework but not used within SimulationFramework.

Usage

transform_cont_y_based_on_response_type(
  y_cont,
  response_type,
  n_ordinal_levels = 4L,
  proportion_epsilon = 1e-06,
  survival_min_time = 0.1,
  count_min_rate = 0L,
  count_shift = 0
)

Arguments

y_cont

Numeric vector. The latent continuous response signal.

response_type

Character scalar. One of "continuous", "incidence", "proportion", "count", "survival", "ordinal".

n_ordinal_levels

Positive integer. Number of ordinal categories when response_type = "ordinal". Default 4L.

proportion_epsilon

Numeric scalar. Small value added to proportion to avoid 0 and 1. Default 1e-6.

survival_min_time

Numeric scalar. Minimum survival time and shift. Default 0.1.

count_min_rate

Integer scalar. Minimum baseline rate for count response. Default 0L.

count_shift

Numeric scalar. Constant added to counts after zero-centering. Default 0.

Value

A numeric vector of transformed responses on the appropriate scale.

Examples

transform_cont_y_based_on_response_type(rnorm(10), 'incidence')

Benchmark this machine and tune EDI's performance-policy defaults to it

Description

Every performance-policy default shipped in EDI (whether an inference class uses a smart cold start, whether resampling reuses warm starts and at what sample sizes, which optimizer algorithm a family uses, and at what sample size parallel bootstrapping starts to beat serial) was measured empirically on the maintainer's machine. Yours differs – core count, cache, BLAS, compiler flags – so a policy that is net-positive there can be net-negative here, and vice versa. This function re-runs those benchmarks on your hardware, decides the winning setting per axis, saves the result to a per-user config file, and applies it immediately; every later library(EDI) re-applies it. Hardware changed? Re-run; it overwrites.

Usage

tune_EDI_for_this_machine(
  effort = c("standard", "quick", "thorough"),
  axes = NULL,
  families = NULL,
  n_grid = NULL,
  reps = NULL,
  num_cores_grid = NULL,
  converged_fn = NULL,
  quiet = FALSE,
  dry_run = FALSE,
  force = FALSE
)

Arguments

effort

One of "standard" (default; moderate sample-size grid and replicate count), "quick" (coarser grid, fewer replicates, warm-start families narrowed to those the shipped tables already name, parallel axis on the bootstrap operation only), or "thorough" (full grid, more replicates).

axes

Which axes to tune: any subset of "cold_start", "warm_start", "optimizer", "parallel". NULL (default) means all that are available here – see Details for when the optimizer and parallel axes are included.

families

Optional character vector of inference class names to restrict every axis to; NULL (default) tunes every class each axis governs.

n_grid

Optional integer vector of sample sizes overriding the effort tier's grid.

reps

Optional replicate count per timed cell overriding the effort tier's.

num_cores_grid

Optional integer vector of core counts (each \ge 2) for the parallel axis; default c(2, detectCores()).

converged_fn

function(inf) -> logical(1), required for the optimizer axis (see Details).

quiet

If TRUE, print nothing (no preamble, no progress bar, no summary).

dry_run

If TRUE, run every benchmark and print the would-be policy changes, but write no file and apply nothing.

force

If FALSE (default), refuse to run when the machine looks busy (see Details: "Idle machine") – in an interactive session you are asked whether to proceed anyway; non-interactively it is an error. TRUE skips the check. Applies under dry_run too, since a dry run still benchmarks.

Details

What is tuned. Four axes, each against the corresponding get_*_dispatch_policy() table: cold start (get_cold_start_dispatch_policy), warm start per resampling operation (get_warm_start_dispatch_policy), optimizer algorithm (get_optimization_dispatch_policy), and the parallel-vs-serial crossover sample size per family (get_parallel_dispatch_policy). The bootstrap confidence-interval type policy is a statistical-validity table, not a performance one, and is never touched; nor are entries in the parallel policy's serial blocklist that exist for parallel safety.

How a deviation is accepted. Per axis, per (family, sample size) cell, both settings are timed on identical synthetic data – interleaved (A/B/A/B) for the cold-start/warm-start/optimizer axes, and blocked (all serial reps, then all parallel reps) for the parallel axis, whose fork-cluster setup cost rules out per-replicate interleaving. A candidate displaces the shipped default only if its median time is at least 5% better and that improvement exceeds twice the candidate's own interquartile spread; ties keep the shipped default. The optimizer axis additionally requires the candidate to have converged on every replicate of the cell – speed never trumps a convergence failure. Only deviations from the shipped defaults are stored, in the exact shape the matching set_*_dispatch_policy() setter accepts, and they are merged into (never replacing) the shipped tables.

Progress. A single progress bar with a running estimated-time-left is redrawn in place as each benchmark cell completes – the same bar InferenceSuite$run_all_inference() shows.

Correctness gate. A timing win alone does not displace a shipped default: every accepted deviation is re-fit once under both settings on identical synthetic data and the outputs compared (point estimates for cold start/optimizer/parallel; the resampling operation's own output, RNG-matched, for warm start). A disagreement – or an unverifiable comparison – discards the deviation, with a warning() naming it; discarded deviations are available on the returned object via attr(x, "discarded_by_correctness_gate") and are never written to the config file or applied.

The optimizer axis and converged_fn. There is not yet a generic, class-independent accessor on an inference object that reports whether its fit converged, so the optimizer axis needs you to supply one as converged_fn(inf). When axes is left NULL the optimizer axis is included only if converged_fn is given; asking for it explicitly without one is an error. An always-TRUE converged_fn disables the convergence guard and is not appropriate for a real tuning run.

The parallel axis. Runs only on Unix-alikes with at least two logical cores (it benchmarks a real fork cluster). Its preferred core count is recorded only – never applied at package load – you still opt into parallelism with set_num_cores.

Idle machine. Run this on an otherwise idle machine; a tuning run under contention measures the contention, not the hardware. Before benchmarking, the function checks the 1-minute load average against the core count and times a small fixed calibration operation for noise; if either says the machine is busy it refuses to run (interactively, it asks first) unless force = TRUE.

Value

Invisibly, an EDILocalMachineTuning object: the policy diffs (policy_diffs), the raw per-axis deviations (raw_deviations), the hardware fingerprint, the effort/grid/reps used, timing, and the config file path – printable via print().

See Also

get_local_EDI_optimization to see what is saved, clear_local_EDI_optimization to return to shipped defaults; the underlying tables: get_cold_start_dispatch_policy, get_warm_start_dispatch_policy, get_optimization_dispatch_policy, get_parallel_dispatch_policy.

Examples


# See what would change without writing anything. force = TRUE skips the
# idle-machine contention guard (see Details/`force` above) -- an example
# must not fail just because the machine running R CMD check happens to
# be busy (e.g. a shared CI runner); the guard itself has its own
# dedicated tests (test-local-machine-tuning-assembly.R). Scoped to one
# axis/family/n rather than the full live registry effort = "quick"
# otherwise walks (narrowing "quick" itself to a handful of
# high-effect-size families across every axis is still open, see
# edi_tuning_effort_presets()'s docs) -- this keeps the example a
# few-second sanity check instead of a multi-minute benchmark run.
res = tune_EDI_for_this_machine(effort = "quick", dry_run = TRUE, force = TRUE,
                                 axes = "cold_start", families = "InferenceIncidLogRegr",
                                 n_grid = 50L, reps = 1L)
print(res)


Unset the number of cores and stop parallel clusters

Description

This function stops any global fork or mirai clusters stored in the package environment and resets the core count to serial execution.

Usage

unset_num_cores()

Value

Invisible NULL.

Examples

set_num_cores(2)
unset_num_cores()

Fast Variance Calculation

Description

Calculates the variance of a numeric vector using Rcpp for speed.

Usage

var_cpp(x)

Arguments

x

A numeric vector.

Value

The variance of the vector.


Wilson Score Interval for a Single Proportion

Description

Wilson Score Interval for a Single Proportion

Usage

wilson_score_interval_cpp(x, n, alpha)

Zhang Exact Inference Helpers

Description

Internal method. Standalone functions to support Zhang (2026) exact test-inversion inference. These functions handle the bisection solver, p-value combination rules, and component-wise p-value logic.

Usage

zhang_combine_exact_pvals(p_M, p_R, m, nRT, nRC, method)

Arguments

p_M

Matched p-value.

p_R

Reservoir p-value.

m

Number of matches.

nRT

Number of treated in reservoir.

nRC

Number of control in reservoir.

method

Combination method (Fisher or Stouffer).

mirror server hosted at Truenetwork, Russian Federation.