---
title: "Getting started with cWise"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Getting started with cWise}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
  %\VignetteDepends{cWise}
---

```{r, include = FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
```

```{r setup}
library(cWise)
```

## The crosswise model

Crosswise questions let respondents answer a sensitive question without
revealing their answer directly. A respondent is shown a sensitive statement
and a second statement whose prevalence is known (`p`). They report only
whether their answers to the two statements are the same or different. The
crosswise response is therefore not itself the sensitive-trait indicator.

The design protects privacy, but it also makes inattentive responses important:
respondents who answer at random pull the observed crosswise proportion toward
one half. `cWise` uses an *anchor* crosswise question with known sensitive-item
prevalence to estimate attentiveness and correct that bias.

The package includes simulated data from this design. `Y` is the primary
crosswise response, `A` is the anchor response, and `p` and `p.prime` are the
known randomization probabilities.

```{r data}
head(cmdata)
```

## Estimate prevalence

Use `bc_est()` to obtain both the naive and bias-corrected prevalence estimates.
Here, the anchor makes it possible to estimate the share of attentive
respondents as well.

```{r prevalence}
estimate <- bc_est(Y = Y, A = A, p = 0.15, p.prime = 0.15, data = cmdata)
estimate
```

`Results` contains point estimates, standard errors, and confidence intervals.
`Stats` reports the estimated attentive-response rate and the analysis sample
size. The bias-corrected row is usually the estimand of interest when the
anchor-question assumptions are credible.

## Sensitivity analysis without an anchor

When no anchor question was fielded, `cmBound()` shows how the estimate changes
over a plausible range of inattentive-response rates. The following calculation
uses only the observed primary-question proportion, its randomization
probability, and the sample size.

```{r bounds, fig.width = 7, fig.height = 4}
cmBound(lambda.hat = mean(cmdata$Y), p = 0.15, N = nrow(cmdata))
```

An optional direct-question estimate can be supplied with `dq` and `N.dq`; it
is then drawn as a reference line with its uncertainty interval.

## Simulation and power planning

Power-analysis functions are deliberately not evaluated while this vignette is
built: they run nested bootstrap simulations. Start with a small exploratory
run and increase `N.sim` for a study-planning analysis.

```{r power, eval = FALSE}
simulation <- sim_cwdata(
  N.sim = 200, sample = 500, prevalence = 0.10, p = 0.15, p.prime = 0.15,
  gamma = 0.80, direct = 0.05
)
simulation$Results

sim_power(
  N.sim = 500, sample = 1000, pi.null = 0.05, pi.alt = 0.10,
  p = 0.15, p.prime = 0.15, gamma = 0.80, direct = 0.05
)
```

`sim_estimates()` plots the simulated estimates, while `sim_power_N()` compares
the package's standard sample-size grid. These functions are intended for
interactive study planning rather than a package build.


The sensitive trait in a crosswise survey is latent: we observe crosswise and
anchor responses, not each respondent's trait status. `cmreg()` and
`cmreg_p()` fit likelihood-based models that account for this measurement
structure.

## The latent trait as an outcome

`cmreg()` models the probability of possessing the sensitive trait as a
function of covariates. The formula contains the primary crosswise response on
the left; supply the anchor response separately.

```{r trait-outcome}
outcome_fit <- cmreg(
  Y ~ female + age,
  anchor = A, p = 0.10, p.prime = 0.15,
  data = cmdata2, n.start = 1L
)
summary(outcome_fit)
```

The first coefficient block describes the latent trait. The auxiliary block
describes the probability of attentive responding. Predict trait prevalence at
specific covariate combinations with uncertainty from a parametric bootstrap.

```{r trait-prediction}
trait_predictions <- cmpredict(
  outcome_fit,
  newdata = data.frame(female = c(0, 1), age = c(30, 30)),
  nsim = 300, seed = 20260825
)
trait_predictions
```

The plot shows the estimated prevalence and its 95% parametric-bootstrap
interval for each covariate scenario.

```{r trait-prediction-plot, fig.width=6, fig.height=4}
trait_labels <- c("Female = 0", "Female = 1")
plot(
  seq_len(nrow(trait_predictions)), trait_predictions$estimate,
  ylim = c(0, 1), xlim = c(0.75, 2.25), xaxt = "n",
  xlab = "Covariate scenario", ylab = "Predicted trait prevalence",
  pch = 19, col = "#1B6CA8"
)
axis(1, at = seq_along(trait_labels), labels = trait_labels)
arrows(
  x0 = seq_len(nrow(trait_predictions)), y0 = trait_predictions$conf.low,
  x1 = seq_len(nrow(trait_predictions)), y1 = trait_predictions$conf.high,
  angle = 90, code = 3, length = 0.05, col = "#1B6CA8"
)
```

## The latent trait as a predictor

`cmreg_p()` models a continuous observed outcome while treating the sensitive
trait as a latent predictor. The formula now contains the observed outcome;
the crosswise and anchor responses are supplied by name.

```{r trait-predictor}
predictor_fit <- cmreg_p(
  V ~ age + female,
  crosswise = Y, anchor = A, p = 0.10, p.prime = 0.15,
  data = cmdata3, n.start = 1L
)
summary(predictor_fit)
```

Predictions are returned for both latent-trait states. Each supplied covariate
scenario therefore produces an “absent” and a “present” row.

```{r outcome-prediction}
outcome_predictions <- cmpredict_p(
  predictor_fit,
  newdata = data.frame(age = 30, female = 1),
  nsim = 300, seed = 20260825
)
outcome_predictions
```

This plot compares the model-implied observed outcome when the latent trait is
absent versus present, with 95% parametric-bootstrap intervals.

```{r outcome-prediction-plot, fig.width=6, fig.height=4}
state_labels <- sub("^[0-9]+: ", "", rownames(outcome_predictions))
ylim <- range(outcome_predictions$conf.low, outcome_predictions$conf.high)
plot(
  seq_len(nrow(outcome_predictions)), outcome_predictions$estimate,
  ylim = ylim + c(-0.05, 0.05), xlim = c(0.75, 2.25), xaxt = "n",
  xlab = "Latent-trait scenario", ylab = "Predicted observed outcome",
  pch = 19, col = "#A23A2E"
)
axis(1, at = seq_along(state_labels), labels = state_labels)
arrows(
  x0 = seq_len(nrow(outcome_predictions)), y0 = outcome_predictions$conf.low,
  x1 = seq_len(nrow(outcome_predictions)), y1 = outcome_predictions$conf.high,
  angle = 90, code = 3, length = 0.05, col = "#A23A2E"
)
```

The uncertainty intervals reflect sampling uncertainty in the fitted
parameters. They do not turn an individual respondent's latent trait into an
observed value; instead, they compare model-implied outcomes under the two
trait scenarios.

